Display image generation device and image display method

The display image generating device and method address alignment issues by switching intermediate image rendering based on the display world's state, achieving precise alignment and enhancing user interaction in AR and MR systems.

JP2026022091APending Publication Date: 2026-02-12SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024123461
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately aligning 2D captured images with 3D virtual objects, especially when the display field of view changes, leading to misalignment and discomfort during user interaction in AR and MR systems.

Method used

A display image generating device and method that includes a captured image acquisition unit, an object placement unit, and a display image generation unit, which switches between drawing virtual objects via an intermediate image based on the display world's state, ensuring precise alignment and accurate composition of CG and photographed images.

Benefits of technology

Enables high-accuracy synthesis of CG and photographed images with reduced computational load, facilitating user interaction and manipulation in three-dimensional environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022091000001_ABST
    Figure 2026022091000001_ABST
Patent Text Reader

Abstract

To appropriately display a position of an image of an object or a real object in a virtual space.SOLUTION: The information processor 10 generates a see-through image using captured images and a CG image of virtual objects (S10). When it becomes necessary to display an instruction object for a user to perform an instruction operation (Y in S12), an information processor generates a CG image after generating an intermediate image representing the instruction object from a viewpoint of an imaging device if an instruction target is not a processing non-target object, and generates the CG image without generating the intermediate image if the instruction target is the processing non-target object (S20).SELECTED DRAWING: Figure 17
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a display image generating device and an image display method for combining and displaying a photographed image and CG. [Background technology]

[0002] Image display systems that allow users to view a target space from any viewpoint are becoming widespread. For example, systems have been developed that display panoramic images on a head-mounted display, displaying images according to the line of sight of the user wearing the head-mounted display. By displaying stereo images with parallax for the left and right eyes on the head-mounted display, the displayed images appear three-dimensional to the user, enhancing the sense of immersion in the image world.

[0003] Furthermore, technology has been put into practical use to realize Augmented Reality (AR) and Mixed Reality (MR) by attaching a camera that captures real space to a head-mounted display and synthesizing the captured image with computer graphics (CG).The captured image can also be displayed on a sealed head-mounted display, which is useful for users to check their surroundings or set up a game play area. Summary of the Invention [Problem to be solved by the invention]

[0004] In technologies such as AR and MR that combine CG images of virtual objects with captured images, the accuracy of alignment between the image of the real object and the CG significantly affects the quality of the content. However, it is not easy to precisely align the captured image, which is originally 2D information, with the virtual object, which is 3D information. In particular, in situations where the display field of view can change significantly depending on the user's movement, the composition must be performed to track this, making precise alignment even more difficult.

[0005] Furthermore, regardless of whether or not a captured image is synthesized, when a user manipulates a virtual object to point to an object in the displayed world or to generate an interaction, if the positional relationships set in the three-dimensional space are not accurately represented, the user may be unable to perform the intended operation or may feel uncomfortable. This problem becomes more pronounced as the types and specifications of virtual objects included in the display become more diverse.

[0006] The present invention has been made in consideration of these problems, and its purpose is to provide a technology for compositing CG and photographed images with high accuracy and with little load. Another purpose of the present invention is to provide a technology for allowing a user to appropriately operate a virtual object in a displayed world regardless of the situation. [Means for solving the problem]

[0007] To solve the above problems, one aspect of the present invention relates to a display image generating device, which includes a captured image acquisition unit that acquires data of an image captured by a camera, an object placement unit that places a virtual object operated by a user in a virtual three-dimensional space, a display image generation unit that draws an image of the virtual object and generates a display image by combining the image with the captured image, and an output unit that outputs the display image data, wherein the display image generation unit switches whether to draw the image of the virtual object via an intermediate image that represents the image from the viewpoint of the camera, depending on the state of the display world including the virtual object.

[0008] Another aspect of the present invention relates to an image display method, which includes the steps of acquiring data of an image captured by a camera, arranging a virtual object to be operated by a user in a virtual three-dimensional space, drawing an image of the virtual object and combining it with the captured image to generate a display image, and outputting data of the display image, wherein the step of generating the display image is characterized in that whether or not the image of the virtual object is drawn via an intermediate image that represents the image from the viewpoint of the camera is switched depending on the state of the display world including the virtual object.

[0009] Any combination of the above components, and any transformation of the present invention into a method, device, system, computer program, data structure, recording medium, etc., are also valid aspects of the present invention. [Effects of the Invention]

[0010] According to the present invention, it is possible to synthesize CG and a photographed image with high accuracy and with little load, and also to facilitate user manipulation of the three-dimensional displayed world. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a diagram illustrating an example of the configuration of an information processing system according to an embodiment of the present invention. [Figure 2] 1A and 1B are diagrams illustrating an example of the external shape of a head-mounted display according to an embodiment of the present invention. [Figure 3] FIG. 2 is a diagram showing functional blocks of the head-mounted display according to the present embodiment. [Figure 4] 1A and 1B are diagrams illustrating examples of the external shape of an input device according to an embodiment of the present invention. [Figure 5] FIG. 2 is a diagram showing functional blocks of an input device according to the present embodiment. [Figure 6] 1 is a diagram for explaining the relationship between a three-dimensional space that forms the display world of a head-mounted display and a display image generated from a captured image in this embodiment. [Figure 7] 10A and 10B are diagrams for explaining differences that may occur in a see-through image and the real world in the present embodiment. [Figure 8] 10A and 10B are diagrams for explaining the principle of positional deviation when CG is combined with a see-through image in the present embodiment. [Figure 9] 10A to 10C are diagrams for explaining a method for matching CG to an image of a real object in the present embodiment. [Figure 10] 10A and 10B are diagrams illustrating an example in which a user interferes with the displayed world via a virtual object in the present embodiment. [Figure 11] 10A and 10B are diagrams showing examples of images displayed on a head-mounted display when a user sets a play area using an instruction object in this embodiment. [Figure 12] 10A and 10B are diagrams schematically illustrating an example of a display image including an object for which an intermediate image cannot be generated in the present embodiment. [Figure 13] 10A and 10B are diagrams schematically illustrating an example of a display image including an object for which an intermediate image cannot be generated in the present embodiment. [Figure 14] 10A and 10B are diagrams illustrating an example of setting whether or not to use intermediate images when the information processing device draws a virtual object in the present embodiment. [Figure 15] 1 is a diagram showing an internal circuit configuration of an information processing device according to an embodiment of the present invention; [Figure 16] FIG. 2 is a diagram showing a functional block configuration of the information processing device according to the present embodiment. [Figure 17] 10 is a flowchart showing a processing procedure for generating a see-through image on which CG of a virtual object can be composited, performed by the information processing device according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] 1 shows an example of the configuration of an information processing system 1 according to the present embodiment. The information processing system 1 includes an information processing device 10, a recording device 11, a head-mounted display 100, and an input device 16. The recording device 11 records applications such as system software and content software used by the information processing device 10 for information processing.

[0013] The information processing device 10 loads software stored in the recording device 11, processes the content, and generates a display image. Typically, the information processing device 10 identifies the position of the viewpoint and the direction of the line of sight based on the position and posture of the head of a user wearing the head-mounted display 100, and generates a display image with a corresponding field of view. For example, while progressing through an electronic game, the information processing device 10 generates an image representing a virtual world in which the game is set, thereby realizing VR (Virtual Reality).

[0014] However, in this embodiment, the type and purpose of content processed by the information processing device 10 are not particularly limited. The information processing device 10 may be connected to a server via a network (not shown) and acquire content software, data of images to be displayed, and the like from the server. The information processing device 10, the head-mounted display 100, and the input device 16 may be connected using a known wireless communication protocol or may be connected by cable.

[0015] The head-mounted display 100 is a display device worn by a user on the head, which displays images on display panels located in front of the user's eyes. The head-mounted display 100 displays an image for the left eye on the left-eye display panel and an image for the right eye on the right-eye display panel. By displaying images with parallax as the images for the left eye and the right eye, stereoscopic vision can be achieved. The head-mounted display 100 also includes eyepieces for expanding the viewing angle. The information processing device 10 generates parallax image data that has been subjected to inverse correction to eliminate optical distortion caused by the eyepieces, and transmits the data to the head-mounted display 100.

[0016] The head mounted display 100 is equipped with a plurality of imaging devices 14. The imaging devices 14 are attached to the front of the head mounted display 100 at different positions and in different orientations so that the total imaging range of each imaging device 14 covers the user's field of view. The imaging devices 14 capture images of the real space at a predetermined cycle (for example, 120 frames per second) in a synchronized manner. The head mounted display 100 sequentially transmits data of the captured images to the information processing device 10.

[0017] The head mounted display 100 also includes an inertial measurement unit (IMU) including a three-axis acceleration sensor and a three-axis angular velocity sensor. The head mounted display 100 transmits sensor data to the information processing device 10 at a predetermined cycle (for example, 800 Hz).

[0018] The input device 16 has a plurality of operation members such as operation buttons, and a user operates the operation members with their fingers while holding the input device 16. When the information processing device 10 executes a game, the input device 16 is used as a game controller. The input device 16 has an IMU including a three-axis acceleration sensor and a three-axis angular velocity sensor, and transmits sensor data to the information processing device 10 at a predetermined cycle (for example, 800 Hz).

[0019] In this embodiment, not only operation information of the operating members of the input device 16 but also the position, speed, and attitude of the input device 16 are handled as operation information and reflected in the movement of virtual objects in the displayed world. For example, the information processing device 10 displays a CG image of a laser pointer that appears to be generated from the input device 16, and changes the position and attitude in conjunction with the position and attitude of the input device 16. This allows the user to point to objects or areas in the displayed world with the same operational feel as an actual laser pointer.

[0020] In order to track the position and orientation of the input device 16, the input device 16 may be provided with a plurality of markers that can be photographed by the imaging device 14 of the head-mounted display 100. The information processing device 10 may have a function of analyzing an image of the input device 16 and estimating the position and orientation of the input device 16 in real space.

[0021] The information processing device 10 may further include a function of analyzing sensor data transmitted from the input device 16 and estimating the position and orientation of the input device 16. In this case, the information processing device 10 may integrate the estimation result based on the marker image and the estimation result based on the sensor data to derive the position and orientation of the input device 16. This allows the state of the input device 16 at each time to be estimated with high accuracy.

[0022] 2 shows an example of the external shape of the head mounted display 100. The head mounted display 100 is composed of an output mechanism unit 102 and a wearing mechanism unit 104. The wearing mechanism unit 104 includes a wearing band 106 that goes around the head when worn by the user and secures the head mounted display 100 to the head. The wearing band 106 has a material or structure that allows its length to be adjusted to fit the user's head circumference.

[0023] The output mechanism unit 102 includes a housing 108 shaped to cover the left and right eyes when the user wears the head-mounted display 100, and includes a display panel inside that faces the eyes when worn. The display panel may be a liquid crystal panel, an organic EL panel, or the like. The housing 108 also includes a pair of left and right eyepieces that expand the user's field of view. The head-mounted display 100 may further include speakers or earphones at positions corresponding to the user's ears, and may be configured to allow external headphones to be connected.

[0024] Four image capturing devices 14a, 14b, 14c, and 14d are provided on the front outer surface of the housing 108. By installing multiple image capturing devices 14 in this manner and setting the directions of their optical axes to be different from one another, the combined image capturing range of each device can cover the user's field of view. However, the number and arrangement of image capturing devices 14 in this embodiment are not limited to those shown in the drawings.

[0025] 3 shows functional blocks of the head mounted display 100. The control unit 120 is a main processor that processes and outputs various data such as image data, audio data, and sensor data, as well as commands. The storage unit 122 temporarily stores the data and commands processed by the control unit 120. The IMU 124 acquires sensor data related to the movement of the head mounted display 100. The IMU 124 may include at least a three-axis acceleration sensor and a three-axis angular velocity sensor. The IMU 124 detects the values ​​of each axial component (sensor data) at a predetermined cycle (for example, 800 Hz).

[0026] The communication control unit 128 transmits data output from the control unit 120 to the external information processing device 10 by wired or wireless communication via a network adapter or an antenna. The communication control unit 128 also receives data from the information processing device 10 and outputs it to the control unit 120.

[0027] When the control unit 120 receives image data and audio data from the information processing device 10, it supplies the data to the display panel 130 for display and to the audio output unit 132 for audio output. The display panel 130 is composed of a left-eye display panel 130a and a right-eye display panel 130b, and a pair of parallax images is displayed on each display panel. The control unit 120 also causes the communication control unit 128 to transmit sensor data from the IMU 124, audio data from the microphone 126, and captured image data from the imaging device 14 to the information processing device 10.

[0028] 4A and 4B show examples of the external shape of the input device 16. The left-handed input device 16a shown in FIG. 4A includes a case body 20, multiple operation members 22a, 22b, 22c, and 22d operated by the user, and multiple markers 30 that emit light to the outside of the case body 20. The operation members 22 may include an analog stick that is operated by tilting, a push-button, or the like. The case body 20 has a grip portion 21 and a curved portion 23 that connects the top and bottom of the case body, and the user inserts their left hand into the curved portion 23 to grip the grip portion 21. While gripping the grip portion 21, the user operates the operation members 22a, 22b, 22c, and 22d with the thumb of their left hand.

[0029] An input device 16b for a right hand shown in (b) includes a case body 20, a plurality of operation members 22e, 22f, 22g, and 22h operated by the user, and a plurality of markers 30 that emit light to the outside of the case body 20. The operation members 22 may include an analog stick that is operated by tilting, a push-button, or the like. The case body 20 has a grip portion 21 and a curved portion 23 that connects the top and bottom of the case body, and the user inserts their right hand into the curved portion 23 to grip the grip portion 21. While gripping the grip portion 21, the user operates the operation members 22e, 22f, 22g, and 22h with the thumb of their right hand.

[0030] The marker 30 is a light-emitting portion that emits light to the outside of the case body 20, and includes a resin portion that diffuses and emits light from a light source such as an LED (Light Emitting Diode) element to the outside on the surface of the case body 20. The marker 30 is photographed by the imaging device 14 and used for tracking processing of the input device 16.

[0031] 5 shows functional blocks of the input device 16. The control unit 50 receives operation information input to the operation members 22. The control unit 50 also receives sensor data detected by the IMU 32 and sensor data detected by the touch sensor 24. The touch sensor 24 is attached to at least some of the multiple operation members 22, and detects a state in which the user's finger is in contact with the operation member 22.

[0032] The IMU 32 acquires sensor data related to the movement of the input device 16 and includes an acceleration sensor 34 that detects acceleration data along at least three axes and an angular velocity sensor 36 that detects angular velocity data along three axes. The acceleration sensor 34 and the angular velocity sensor 36 detect the values ​​(sensor data) of each axial component at a predetermined cycle (e.g., 800 Hz). The control unit 50 supplies the received operation information and sensor data to a communication control unit 54. The communication control unit 54 transmits the operation information and sensor data to the information processing device 10 via a network adapter or an antenna by wired or wireless communication.

[0033] The input device 16 includes a plurality of light sources 58 for lighting up a plurality of markers 30. The light sources 58 may be LED elements that emit light in a predetermined color. When the communication control unit 54 receives a light-emitting instruction from the information processing device 10, the control unit 50 causes the light sources 58 to emit light based on the light-emitting instruction, thereby lighting up the markers 30. Note that in the example shown in FIG. 5, one light source 58 is provided for one marker 30, but one light source 58 may light up a plurality of markers 30.

[0034] In this embodiment, a mode is provided in which moving images captured by the imaging device 14 of the head mounted display 100 are displayed with little delay, allowing the user to see the state of the real space in the direction the user is facing. Hereinafter, this mode will be referred to as a "see-through mode." For example, the head mounted display 100 automatically switches to the see-through mode during a period in which a content image is not being displayed.

[0035] This allows the user to check the surrounding situation before the content starts, ends, or is interrupted without removing the head-mounted display 100. The see-through mode may also be started when the user explicitly performs an operation, or may be started or ended depending on the situation, such as when a play area is set or when the user deviates from the play area. Here, the play area is the range of the real world in which a user viewing a virtual world through the head-mounted display 100 can move around, and is, for example, a range in which safe movement without colliding with surrounding objects is guaranteed.

[0036] Images captured by the imaging device 14 can also be used as content images. For example, AR and MR can be realized by compositing and displaying CG of a virtual object with the captured image, with the position, posture, and movement matching that of a real object in the field of view of the imaging device 14. Furthermore, regardless of whether the captured image is included in the display, the captured image can be analyzed and the results can be used to determine the position, posture, and movement of the object to be drawn.

[0037] For example, stereo matching may be performed on the captured images to extract corresponding points on the image of the subject and obtain the distance to the subject using the principle of triangulation. Alternatively, a well-known technique such as Visual SLAM (Simultaneous Localization and Mapping) may be used to obtain the position and orientation of the head-mounted display 100, and ultimately the user's head, relative to the surrounding space. Visual SLAM is a technique that obtains three-dimensional position coordinates of feature points on the surface of an object based on corresponding points extracted from stereo images, and tracks the feature points in chronologically ordered frames to obtain the position and orientation of the image capture device 14 and an environmental map in parallel.

[0038] FIG. 6 is a diagram for explaining the relationship between the three-dimensional space forming the display world of the head-mounted display 100 and the display image generated from the captured image. In the following explanation, the captured image converted into the display image will be referred to as the see-through image, regardless of whether it is in see-through mode or not. The upper part of the figure shows a bird's-eye view of the virtual three-dimensional space (hereinafter referred to as the display world) configured when the display image is generated. Virtual cameras 260a and 260b are virtual rendering cameras for generating the display image, and correspond to the left and right viewpoints of the user. The upward direction of the figure represents the depth direction (the distance from the virtual cameras 260a and 260b).

[0039] The see-through images 268a and 268b correspond to images captured by the imaging device 14 of the interior of the room in front of the head mounted display 100, and show one frame of display images for the left eye and the right eye. Naturally, if the user changes the direction of their face, the field of view of the see-through images 268a and 268b also changes. To generate the see-through images 268a and 268b, the head mounted display 100 or the information processing device 10, for example, positions the captured image 264 at a predetermined distance Di in the display world.

[0040] More specifically, the head mounted display 100 displays, for example, captured images 264 of the left and right viewpoints captured by the imaging device 14 on the inner surface of a sphere of radius Di centered on each of the virtual cameras 260a and 260b. The head mounted display 100 then generates see-through images 268a and 268b for the left and right eyes by rendering images obtained by viewing the captured images 264 from the virtual cameras 260a and 264b.

[0041] As a result, image 264 captured by imaging device 14 is converted into an image from the viewpoint of the user viewing the displayed world. Here, the image of the same subject appears shifted to the right in see-through image 268a for the left eye and shifted to the left in see-through image 268b for the right eye. Since the images captured from the left and right viewpoints are originally captured with parallax, the images of the subject appear with various amounts of deviation in see-through images 268a and 268b depending on their actual positions (distances). This allows the user to perceive a sense of distance from the image of the subject.

[0042] In this way, if the captured image 264 is displayed on a uniform virtual surface and the state of the captured image 264 as seen from a viewpoint corresponding to the user is used as the displayed image, the captured image can be displayed with a sense of depth without constructing a three-dimensional virtual world that accurately traces the placement and structure of the subject. Also, if the surface that displays the captured image 264 (hereinafter referred to as the projection surface) is a spherical surface that maintains a predetermined distance from the virtual camera 260, the image of an object existing within an expected range can be displayed with uniform quality regardless of the direction. As a result, it is possible to achieve both low latency and a sense of presence with a small processing load.

[0043] On the other hand, when compared to the state of viewing the real world directly, slight differences may occur in the image of a real object produced by the display method shown in the figure. While these differences are difficult to notice when only the see-through image is displayed, they become apparent as misalignment with the CG when compositing with CG. CG generally represents the appearance of a 3D model of a virtual object as seen from the user's viewpoint, whereas the see-through image is originally data obtained separately as a 2D photographed image, which causes this misalignment. Therefore, in this embodiment, a composite image with minimal misalignment is displayed by drawing the CG while assuming the position of the image of the real object in the see-through image.

[0044] FIG. 7 is a diagram illustrating differences that may occur in a see-through image from the real world in this embodiment. This figure shows a side view of the three-dimensional space of the display world shown in the upper part of FIG. 6, and shows one of the left and right virtual cameras, virtual camera 260a, as well as the corresponding camera of imaging device 14. As described above, the see-through image represents an image captured by imaging device 14, projected onto projection surface 272, as viewed from virtual camera 260a. Projection surface 272 is, for example, the inner surface of a sphere with a radius of 2 m centered on virtual camera 260a. However, the shape and size of the projection surface are not limited to this.

[0045] The virtual camera 260a and the imaging device 14 are linked to the movement of the head-mounted display 100, and ultimately to the movement of the user's head. For example, when a rectangular parallelepiped real object 276 enters the field of view of the imaging device 14, its image is projected on the projection surface 272 near a position 278 where a line of sight 280 from the imaging device 14 to the real object 276 intersects. In a see-through image viewed from the virtual camera 260a, the real object 276, which should actually be in the direction of line of sight 282, is displayed in the direction of line of sight 284. As a result, the real object 276 appears to the user as if it were located a distance D in front of the user (displayed real object 286).

[0046] Fig. 8 is a diagram for explaining the principle of positional deviation when compositing CG with a see-through image. This figure assumes that a virtual object 290 is expressed in CG as if it exists on a real object 276 in the environment shown in Fig. 7. In this case, generally, the three-dimensional position coordinates of the real object 276 are first obtained, and then the virtual object 290 in the display world is positioned to correspond to them.

[0047] Then, the state of virtual object 290 as seen from virtual camera 260a is rendered as a CG image and composited with the see-through image. According to this procedure, virtual object 290 on the display is naturally expressed as existing in the direction of line of sight 292 from virtual camera 260a to virtual object 290. On the other hand, as described with reference to Fig. 7, real object 276 is expressed as real object 286 on the display existing at distance D in front, and therefore the two appear to the user to be misaligned.

[0048] This phenomenon occurs due to the difference in the optical centers and optical axis directions of the imaging device 14 and the virtual camera 260a. In other words, the real object 276 is projected onto the screen coordinate system of the virtual camera 260a via the screen coordinate system corresponding to the imaging surface of the imaging device 14 and the projection surface 272, whereas the virtual object 290 is projected directly onto the screen coordinate system of the virtual camera 260a, which causes a positional discrepancy between the two. Therefore, in this embodiment, a process is incorporated in which the virtual object 290 is first projected onto the screen coordinate system of the imaging device 14 or onto the projection surface 272, and then its image (CG) is aligned with the image of the real object 276.

[0049] 9 is a diagram illustrating a technique for matching CG with an image of a real object. In this case, as in the case of FIG. 8, the three-dimensional position coordinates of real object 276 are obtained, and virtual object 290 is positioned corresponding to them. On the other hand, in this embodiment, an intermediate image of virtual object 290 is generated so as to follow the projection that real object 276 goes through before being displayed as a see-through image.

[0050] Specifically, the virtual object 290 is projected onto a screen coordinate system 298 of the imaging device 14, thereby representing the state of the virtual object 290 as seen from the imaging device 14 as an intermediate image. Alternatively, the image of the virtual object 290 as seen from the imaging device 14 may be projected onto a projection surface 272 and directly represented near a position 299 as an intermediate image. In either case, according to these intermediate images, the virtual object 290 is represented in the direction of a line of sight 294 as seen from the imaging device 14.

[0051] In other words, since the viewpoint is unified with that of the captured image, the remaining processing is then performed in the same way as generating a see-through image, and by combining the two at some stage, an image with no positional deviation between the CG and real object images can be displayed. In this case, virtual object 290 is displayed in the direction of line of sight 296 from virtual camera 260a. In other words, as in the case of Fig. 7, virtual object 290 appears to the user to exist at a distance D in front (displayed virtual object 297), but since the positional deviation with real object 286 is eliminated on the display, this is difficult for the user to notice, and it is possible to make it appear as if a highly accurate composite image is being displayed.

[0052] FIG. 10 illustrates an example of a mode in which a user interacts with the display world via a virtual object in this embodiment. The figure shows a virtual situation in which a user wearing a head-mounted display 100 is in a three-dimensional space 300. The three-dimensional space 300 is, for example, the user's room, and the display of a see-through image allows the user to look around the surroundings with the same sensation as if the user were not wearing the head-mounted display 100. In a situation in which the user makes some kind of instruction selection for the three-dimensional space 300, such as setting up a play area, the information processing device 10 causes an instruction object 302 that can be operated by the user to appear.

[0053] In the figure, the instruction object 302 is shown as a ray, but the shape is not particularly limited. The information processing device 10 displays the instruction object 302 in the three-dimensional space 300 as if it were extending in a predetermined direction from a predetermined position of one input device 16. This allows the user to easily specify a desired position in the three-dimensional space 300 by changing the position and orientation of the input device 16.

[0054] For example, if the user indicates a position using the instruction object 302 and then presses the operation member of the input device 16, the information processing device 10 accepts the position or the indicated object as a selection target. Alternatively, if the user draws a closed curve with the instruction object 302 while holding down the operation member of the input device 16, the information processing device 10 accepts the interior of that curve as a selection area. It will be understood by those skilled in the art that there are various other possible input operations that can be realized with the instruction object 302.

[0055] FIG. 11 shows a schematic example of an image displayed on the head-mounted display 100 when the user sets the play area using the instruction object. Note that although one display image is shown in the figure, in reality, images with parallax for the left and right eyes are displayed, as described above. The display image shown is based on a see-through image 304 captured in real time of the user's room. The see-through image 304 also includes an image 306 of the user's own hand and an image 308 of the input device being held.

[0056] When setting the play area, the information processing device 10 additionally displays an instruction object 310 in the see-through image 304. In detail, the information processing device 10 places a three-dimensional model of the instruction object 310 in three-dimensional space based on the position and orientation of the input device 16, and then draws the view of the instruction object 310 as seen from a virtual camera for display. The user uses the input device 16 to move the destination of the instruction object 310, thereby drawing the boundary line of the play area on the floor of the living room. The information processing device 10 further draws a line 312 representing the trajectory and a pattern (e.g., pattern 314) representing the interior of the area.

[0057] When the user completes the setting, the information processing device 10 stores the area on the floor corresponding to the area within the drawn boundary as the play area. The stored play area information is used to warn the user when they are about to leave the play area while playing a VR game, for example. This prevents users who cannot see the real space around them from colliding with furniture or the like.

[0058] 9, the information processing device 10 first generates an intermediate image in which the instruction object 310 and the trajectory line 312 are displayed in the screen coordinate system of the imaging device 14 or on the projection surface of the see-through image, and then displays the intermediate image in the screen coordinate system of the virtual camera. This results in no apparent misalignment between the instruction object 310 and the trajectory line 312 and the image 308 of the input device or the image on the floor. Note, however, that this process is merely a display positioning process, and the position of the instruction object 310 itself and its destination are calculated in three-dimensional space.

[0059] On the other hand, for objects that are specified to be drawn directly in the virtual camera's screen coordinate system, such as template objects provided by middleware, generating intermediate images becomes difficult. Figures 12 and 13 show schematic examples of display images that include objects for which intermediate images cannot be generated. In these examples, a dialog box 320 for instructing the user to set the play area is added to the display image shown in Figure 11.

[0060] For example, as shown in Fig. 12, the user checks the play area setting method by looking at the text and images shown in the dialog 320, and draws the boundaries of the play area on the floor using the instruction object 310. Next, as shown in Fig. 13, the user inputs completion of the play area setting by using the instruction object 310 to point to a GUI (Graphical User Interface) 322 displayed in the dialog 320 that reads "Done."

[0061] Like other virtual objects, the dialogue 320 is basically an object placed at a predetermined position in three-dimensional space, and is rendered as seen from a virtual camera used for display. On the other hand, if a template for which it is difficult to generate an intermediate image is used as the dialogue 320, the image of the dialogue 320 is rendered directly in the screen coordinate system of the virtual camera using a general method. Then, using the same principle as shown in Figure 8, the positional relationship between the dialogue 320 and real objects or other objects will appear different from the state set in three-dimensional space.

[0062] As a result, even if the instruction object 310 is used to point to the GUI 322, a collision is not detected in the calculations, and the play area setting operation cannot be completed, making it difficult to operate the GUI 322. This type of problem is not limited to the operation of the GUI 322, but can occur in any interaction between the instruction object 310 and the dialog 320.

[0063] Therefore, the information processing device 10 switches whether to draw the instruction object 310 via an intermediate image according to a predetermined condition. An example of the switching condition is the attribute of the pointing target. For example, as shown in FIG. 12, when the pointing target is a see-through image, the information processing device 10 draws the instruction object 310 via an intermediate image. On the other hand, as shown in FIG. 13, when the pointing target is a dialog 320, the information processing device 10 draws the instruction object 310 directly in the screen coordinate system of the virtual camera.

[0064] This ensures that the positional relationship between the instruction object 310 and the pointing target is represented in the same way as the positional relationship in three-dimensional space, enabling stable pointing operations regardless of the pointing target. Note that during the period when the instruction object 310 is not drawn via an intermediate image, a positional deviation may occur between the instruction object 310 and the image of the real object in the see-through image, due to the principle shown in Fig. 8. For example, it is conceivable that the base of the instruction object 310 and the image 308 of the input device may be misaligned, but due to the nature of pointing operations, the user is likely to be focusing on the pointing target, so the deviation is unlikely to be noticed and poses little problem.

[0065] An object that is difficult to render via an intermediate image, such as the dialog 320, will be referred to hereinafter as a "non-processing object." The type of non-processing object is not limited to the dialog shown in the figure, and the reason why it cannot be rendered via an intermediate image is not limited. For example, a non-processing object may be an object that instantly displays a 3D model sent from an external device using an existing program, such as an avatar of a communication partner.

[0066] Furthermore, the object for which switching is performed to determine whether or not to generate an intermediate image is not limited to the instruction object. For example, when expressing an interaction by detecting a collision between an object that reflects the movement of a part of the user's body, such as a hand, and another object, the information processing device 10 may switch whether or not to use an intermediate image when drawing the former object, depending on whether the latter object is an object that does not support processing. In this embodiment, a medium that can be operated by the user to interfere with the displayed world, even if it is not a strict instruction, is called an "instruction object," and an object that comes into contact with the instruction object is called a "pointed object."

[0067] The condition for switching whether to use intermediate images in drawing the instruction object is not limited to the attributes of the instruction target. For example, the information processing device 10 may stop using intermediate images when one of the following conditions is met: predetermined content, a predetermined scene within the content, a period during which a non-editable object is displayed, or a mode selected by the user. Furthermore, when stopping using intermediate images when the instruction target becomes a non-editable object, the trigger need not necessarily be the instruction object coming into contact with the non-editable object, but may also be the instruction object entering a predetermined range with a predetermined margin from the non-editable object. In summary, the information processing device 10 switches whether to use intermediate images in drawing the instruction object depending on the state of the display world including the instruction object.

[0068] 14 shows an example of a setting for whether or not to use intermediate images when the information processing device 10 draws a virtual object. First, when the drawing target is a "general object" that is not an instruction object or an object that does not support processing, the information processing device 10 draws the image of the object via intermediate images. That is, the information processing device 10 first generates an intermediate image representing the general object and then represents it in the screen coordinate system of the virtual camera. This allows the image of the general object to steadily fit the real object in the see-through image.

[0069] When the drawing target is a "pointing object," the information processing device 10 switches whether or not to use an intermediate image depending on the attributes of the pointing target. Specifically, when the pointing target is a "real object" or a "general object" in the see-through image, the information processing device 10 draws an image of the pointing object via an intermediate image. When the pointing target is a "non-processing compatible object," the information processing device 10 draws an image of the pointing object directly in the screen coordinate system of the virtual camera without using an intermediate image. When the drawing target is a "non-processing compatible object," an intermediate image cannot be generated, so the information processing device 10 draws an image of the object directly in the screen coordinate system of the virtual camera.

[0070] 15 shows the internal circuit configuration of the information processing device 10. The information processing device 10 includes a CPU (Central Processing Unit) 222, a GPU (Graphics Processing Unit) 224, and a main memory 226. These components are connected to one another via a bus 230. An input / output interface 228 is further connected to the bus 230. A communication unit 232, an output unit 236, an input unit 238, and a recording medium drive unit 240 are connected to the input / output interface 228.

[0071] The communication unit 232 includes a peripheral device interface such as USB or IEEE1394, and a network interface such as a wired LAN or wireless LAN. The output unit 236 outputs data to the head mounted display 100 and the recording device 11. The input unit 238 acquires data from the head mounted display 100, the input device 16, and the recording device 11. The recording medium drive unit 240 drives a removable recording medium such as a magnetic disk, an optical disk, or a semiconductor memory.

[0072] The CPU 222 controls the entire information processing device 10 by executing an operating system loaded from the recording device 11 to the main memory 226. The CPU 222 also executes various programs (e.g., VR game applications) that are read from the recording device 11 or a removable recording medium and loaded to the main memory 226, or that are downloaded via the communication unit 232. The GPU 224 has the functions of a geometry engine and a rendering processor, performs drawing processing in accordance with drawing commands from the CPU 222, and outputs the drawing results to the output unit 236. The main memory 226 is configured from RAM (Random Access Memory), and stores programs and data necessary for processing.

[0073] Fig. 16 shows the configuration of functional blocks of information processing device 10 in this embodiment. The functional blocks shown in the figure can be realized in hardware terms by the circuit configuration shown in Fig. 15, and in software terms by programs that perform various functions such as data input function, data storage function, image processing function, and communication function, which are loaded from recording device 11 to main memory 226. Therefore, it will be understood by those skilled in the art that these functional blocks can be realized in various forms using only hardware, only software, or a combination thereof, and are not limited to any one of them.

[0074] Furthermore, the information processing device 10 may have the functions of processing various electronic contents and communicating with a server as described above, but the diagram shows a functional configuration for combining CG with a see-through image and displaying it on the head-mounted display 100. From this perspective, the information processing device 10 may be a display image generating device. Note that some of the illustrated functional blocks may be provided in the head-mounted display 100.

[0075] The information processing device 10 includes a data acquisition unit 70 that acquires various data from the head-mounted display 100 and the input device 16, a display image generation unit 76 that generates display image data, and an output unit 78 that outputs the display image data. The information processing device 10 further includes an object surface detection unit 80 that detects the surface of a real object, an object surface data storage unit 82 that stores object surface data, an object placement unit 84 that places virtual objects in the display world, an object data storage unit 86 that stores virtual object data, and a pointing target detection unit 90 that detects a target pointed to by the pointing object.

[0076] The data acquisition unit 70 continuously acquires various data required for generating a display image from the head mounted display 100 and the input device 16. In detail, the data acquisition unit 70 includes a captured image acquisition unit 72, a sensor data acquisition unit 74, and an operation information acquisition unit 75. The captured image acquisition unit 72 acquires data of images captured by the imaging device 14 from the head mounted display 100 at a predetermined frame rate.

[0077] The sensor data acquisition unit 74 acquires sensor data detected by the IMU 124 included in the head mounted display 100 and the touch sensor 24 and IMU 32 included in the input device 16 at a predetermined rate. The IMU sensor data may be measurement values ​​such as acceleration and angular acceleration, or may be translational motion and rotational motion derived using the sensor data, and even position and orientation data at each time. In the former case, the sensor data acquisition unit 74 derives the position and orientation of the head mounted display 100 and the input device 16 at a predetermined rate using the acquired measurement values. When a user operates the operation member 22 included in the input device 16, the operation information acquisition unit 75 acquires operation information indicating the content of the operation.

[0078] The object surface detection unit 80 detects the surfaces of real objects around the user in the real world. For example, the object surface detection unit 80 generates data of an environment map that represents the distribution of feature points on the object surface in a three-dimensional space. In this case, the object surface detection unit 80 sequentially acquires captured image data from the captured image acquisition unit 72 and executes the above-mentioned Visual SLAM to generate data of the environment map. However, the detection method performed by the object surface detection unit 80 and the representation format of the detection results are not particularly limited. The object surface data storage unit 82 stores data indicating the detection results by the object surface detection unit 80, for example, data of the environment map.

[0079] The object data storage unit 86 stores the arrangement rules of the virtual objects to be displayed and the data of the 3D models to be represented by CG. The attributes of the virtual objects to be displayed can be a general object, an instruction object, or an object that does not support manipulation, as shown in FIG.

[0080] The line 312 and pattern 314 in Fig. 11 belong to the general object. The instruction object 310 in Fig. 11 and the dialog 320 in Fig. 12 belong to the instruction object and the non-processing object, respectively. The object data storage unit 86 also stores information that distinguishes the attributes of each object in association with its model.

[0081] The object placement unit 84 identifies a virtual object to be displayed based on the operation information acquired by the data acquisition unit 70, and then places the virtual object in the three-dimensional space of the display world based on the information stored in the object data storage unit 86. When displaying a virtual object in accordance with the position and movement of a real object as shown in Fig. 8, the object placement unit 84 acquires three-dimensional position information of the object surface, such as an environmental map, from the object surface data storage unit 82, and determines the three-dimensional position and orientation of the virtual object accordingly.

[0082] The pointing target detection unit 90 detects an object that the user is pointing to using the pointing object. More specifically, the pointing target detection unit 90 acquires the position and orientation of the pointing object in three-dimensional space from the object placement unit 84, and identifies the position coordinates of the pointing destination. Note that the unit of the pointing target detected by the pointing target detection unit 90 is not limited to position coordinates, and may be an object unit, a GUI unit, or other unit having an area. Alternatively, the unit may be a type of image, such as a see-through image / CG.

[0083] As described above, the pointing target detection unit 90 may determine that a detection unit is a pointing target when the target of the pointing object reaches a predetermined area that includes the image of the detection unit. Alternatively, the pointing target detection unit 90 may predict the arrival of the target based on the movement of the pointing object and determine the pointing target.

[0084] The display image generation unit 76 generates a see-through image using the captured images successively acquired by the captured image acquisition unit 72 of the data acquisition unit 70, and generates a display image by combining the see-through image with CG. In detail, the display image generation unit 76 includes a see-through image generation unit 94, an object drawing unit 96, and a composition unit 98. The see-through image generation unit 94 projects the captured image onto a projection surface of a predetermined shape, and then displays the image as seen from a virtual camera for display as a see-through image.

[0085] The object drawing unit 96 draws an image of the virtual object in the three-dimensional space arranged by the object arrangement unit 84 as seen from the virtual camera for display. The object drawing unit 96 includes an intermediate image generation unit 97. As shown in FIG. 14, the object drawing unit 96 operates the intermediate image generation unit 97 when drawing a general object and when drawing an instruction object that indicates an object other than an object that does not support processing. When the intermediate image generation unit 97 is not operated, the object drawing unit 96 directly draws the object to be drawn in the screen coordinate system of the virtual camera.

[0086] The intermediate image generation unit 97 generates an intermediate image in which the viewpoint of the virtual object placed in the three-dimensional space by the object placement unit 84 is aligned with the viewpoint of the real object shown in the captured image. When the intermediate image generation unit 97 is activated, the object drawing unit 96 draws the virtual object in one of the following two procedures, for example.

[0087] (Step a) The intermediate image generation unit 97 generates an intermediate image by drawing the virtual object in the screen coordinate system of the imaging device 14. As a result, the viewpoint for the captured image and the viewpoint for the virtual object to be drawn are already aligned. In this case, the object drawing unit 96 projects the intermediate image onto a projection surface (for example, projection surface 272 in FIG. 9 ) in the same way as when generating a see-through image, and displays the intermediate image as seen from the virtual camera 260, thereby obtaining a final image of the virtual object.

[0088] (Step b) The intermediate image generation unit 97 draws (projects) an image of the virtual object as seen from the imaging device 14 onto a projection surface (for example, projection surface 272 in FIG. 9 ) onto which the captured image is projected when generating a see-through image, to generate an intermediate image. That is, the intermediate image generation unit 97 draws the virtual object in the same state as the captured image projected onto the projection surface 272. In this case, the object drawing unit 96 obtains a final image of the virtual object by representing the intermediate image as seen from the virtual camera 260.

[0089] The composition unit 98 composes the see-through image generated by the see-through image generation unit 94 with the image of the virtual object drawn by the object drawing unit 96 to generate a display image. When the intermediate image generation unit 97 is operated in step a, the composition unit 98 may compose the intermediate image with the captured image at the stage when the intermediate image is generated in the screen coordinate system of the imaging device 14. In this case, instead of the object drawing unit 96, the composition unit 98 projects the composited image onto the projection surface and generates a final display image by representing the appearance as seen from the virtual camera 260. In this way, by once drawing and combining the virtual object according to the viewpoint of the imaging device 14, a natural display image can be generated regardless of the density of polygons.

[0090] Furthermore, the rendering of the virtual object on the projection surface in step b can be performed by well-known projective transformation without first rendering it in the screen coordinate system of the image capture device 14, thereby speeding up the processing. In either procedure, by operating the intermediate image generation unit 97, it is possible to ultimately generate a display image in which there is no positional deviation between the image of the captured real object and the image of the virtual object.

[0091] The position of the viewpoint and the direction of the line of sight of the imaging device 14, the position and orientation of the projection surface, and the position and orientation of the virtual camera 260, which are used when the display image generation unit 76 generates the intermediate image and the display image, depend on the movement of the head mounted display 100 and, ultimately, the user's head. Therefore, the display image generation unit 76 determines these parameters at a predetermined rate based on the data acquired by the data acquisition unit 70.

[0092] The intermediate image generation unit 97 is not limited to actually drawing CG as an intermediate image, but may simply generate information determining the position and orientation of the image. For example, the intermediate image generation unit 97 may represent only the vertex information of a virtual object on an image plane to create an intermediate image. Here, the vertex information may be data commonly used in CG drawing, such as position coordinates, normal vectors, colors, and texture coordinates. In this case, the display image generation unit 76 may draw the actual image while appropriately converting the viewpoint based on the intermediate image, for example, at the stage of combining with a see-through image. This reduces the load required for generating the intermediate image, allowing for faster generation of a composite image.

[0093] The output unit 78 acquires display image data from the display image generation unit 76, performs processing required for display, and sequentially outputs the data to the head-mounted display 100. The display image is composed of a pair of images for the left and right eyes. The output unit 78 may correct the display image in a direction that cancels out distortion aberration and chromatic aberration so that an undistorted image is viewed when viewed through the eyepiece. The output unit 78 may also perform various data conversions corresponding to the display panel of the head-mounted display 100.

[0094] Next, the operation of the information processing device 10, which can be realized by the above configuration, will be described. FIG. 17 is a flowchart showing the processing steps performed by the information processing device 10 to generate a see-through image onto which CG of a virtual object can be synthesized. This flowchart is executed, for example, during the play area setting period, but the content and purpose of the display are not limited to this. While the diagram shows only the steps directly related to the generation of the display image, the data acquisition unit 70 of the information processing device 10 simultaneously acquires necessary data from the head-mounted display 100 and the input device 16 as appropriate. Furthermore, the object surface detection unit 80 acquires the position and orientation of the object surface as appropriate and stores the data in the object surface data storage unit 82.

[0095] First, the display image generation unit 76 generates a see-through image based on the most recent captured image at that time (S10). That is, the display image generation unit 76 projects the captured image onto a predetermined projection surface in three-dimensional space, and then generates a see-through image that shows the image as seen from a virtual camera for display. Note that, depending on the purpose of display, the time and location of the image to be displayed as the see-through image are not limited, and an image captured in advance may be used as the display object. Furthermore, the display object is not limited to the captured image, and may be a separately generated CG image, or an image in which the captured image and CG are combined.

[0096] In the process of S10, the display image generation unit 76 also generates CG images of general objects and non-editable objects as needed. In this case, the display image generation unit 76 represents each object placed in three-dimensional space by the object placement unit 84 in the screen coordinate system of the virtual camera. Here, for general objects, the display image generation unit 76 first generates an intermediate image and then represents it in the screen coordinate system of the virtual camera. For non-editable objects, the object is directly drawn in the screen coordinate system of the virtual camera.

[0097] If it is not necessary to display the instruction object (N in S12), the display image generation unit 76 and the output unit 78 cooperate to appropriately combine the see-through image generated in S10 with the CG image and output the result to the head-mounted display (S22). If it is necessary to display the instruction object (Y in S12), the object placement unit 84 places the instruction object in three-dimensional space (S14). The position and orientation of the instruction object are determined, for example, according to the position and orientation of the input device 16 held by the user.

[0098] Next, the pointing target detection unit 90 checks whether the target pointed to by the pointing object is a non-manufacturing object (S16). If the target pointed to is not a non-manufacturing object (Y in S16), the display image generation unit 76 generates an intermediate image (S18) in the same way as for a general object, and then generates a CG image that represents the intermediate image in the screen coordinate system of the virtual camera (S20). If the target pointed to is a non-manufacturing object (N in S16), the display image generation unit 76 directly draws the pointing object in the screen coordinate system of the virtual camera, in the same way as for a non-manufacturing object (S20).

[0099] The display image generation unit 76 and the output unit 78 then cooperate to composite the image of the instruction object into the see-through image and output it to the head-mounted display 100 (S22). Note that the process of composite the CG image into the see-through image by the display image generation unit 76 may be performed at the stage of generating the intermediate image, as described above. Furthermore, in S16, the display image generation unit 76 may determine whether or not to pass through the intermediate image using criteria other than whether or not the target of the instruction object is an object that cannot be manipulated.

[0100] During the period when it is not necessary to stop the display of the see-through image (N in S24), the information processing device 10 repeats the processes from S10 to S22 at a predetermined rate, for example. When it becomes necessary to stop the display of the see-through image, the information processing device 10 ends all processes (Y in S24).

[0101] According to the present embodiment described above, when a captured image and an image of a 3D virtual object are synthesized and displayed, an intermediate image representing the image of the virtual object from the camera's viewpoint is first generated, and then the image from the viewpoint for display is generated. This makes it possible to generate a synthesized image with no positional misalignment between the images of the virtual object and the real object, without going through a high-load process of strictly associating the captured image with the 3D real space structure.

[0102] Furthermore, for specific objects, such as instruction objects, which are the medium for realizing user interaction with the displayed world, it is possible to switch whether or not to use intermediate images. This allows for the realization of intended instructions and interactions, just like with other objects, even when objects for which intermediate images cannot be generated are included in the display. As a result, even with a small processing load, it is possible to continuously display the captured image and the 3D object in a stable and appropriate positional relationship. It also enables user operations using virtual objects to be performed appropriately regardless of the situation.

[0103] The present invention has been described above based on the embodiments. The embodiments are merely examples, and it will be understood by those skilled in the art that various modifications are possible in the combination of the components and treatment processes, and that such modifications are also within the scope of the present invention.

[0104] For example, as long as the viewpoints of the captured image and the displayed image are different, the display device is not limited to a head-mounted display and can be applied to other devices.

[0105] The present disclosure may include the following aspects. [Item 1] A content server, comprising: a circuit configured to: The circuitry comprises: Acquires data from images captured by a camera, The virtual object that the user operates is placed in a virtual 3D space. An image of the virtual object is drawn, and a display image is generated by combining the image with the captured image; outputting the display image data; When generating the display image, whether or not to draw an image of the virtual object via an intermediate image that represents the image from the viewpoint of the camera is switched depending on a state of the display world including the virtual object. Display image generation device. [Item 2] The circuitry comprises: Item 1. The display image generating device according to item 1, wherein when generating the display image, if the virtual object enters a predetermined range from another virtual object that cannot be drawn via the intermediate image, the display image generating device switches to drawing that does not go via the intermediate image. [Item 3] The circuitry comprises: When placing the virtual object in the virtual three-dimensional space, an instruction object that allows a user to indicate a position in the display world is placed as the virtual object; Item 1. A display image generating device according to item 1, wherein when generating the display image, if the instruction object indicates another virtual object that cannot be drawn via the intermediate image, the device switches to drawing that does not go via the intermediate image. [Item 4] The circuitry comprises: Item 3. The display image generating device according to item 3, wherein when generating the display image, if a virtual object using a template provided by middleware is specified as the other virtual object, the display image generating device switches to drawing that does not go through the intermediate image. [Item 5] The circuitry comprises: Item 1. A display image generating device according to item 1, wherein, when generating the display image, the captured image projected onto a projection surface set in the three-dimensional space is displayed on the plane of the display image, and when the image of the virtual object is drawn via the intermediate image, the intermediate image displayed on the projection surface is displayed on the plane of the display image. [Item 6] The circuitry comprises: Acquire data of the image captured by a camera provided in the head-mounted display; Item 1. The display image generating device according to item 1, which outputs data of the display image to the head-mounted display. [Item 7] Acquires data from images captured by a camera, placing a virtual object to be operated by a user in a virtual three-dimensional space; An image of the virtual object is drawn, and a display image is generated by combining the image with the captured image; outputting the display image data; When generating the display image, whether or not to draw an image of the virtual object via an intermediate image that represents the image from the viewpoint of the camera is switched depending on a state of the display world including the virtual object. Image display method. [Item 8] A function to acquire data of images taken by the camera, placing a virtual object to be operated by a user in a virtual three-dimensional space; a function of drawing an image of the virtual object and generating a display image by combining the image with the captured image; a function of outputting data of the display image; This is realized by a computer, The function of generating the display image switches whether or not to draw the image of the virtual object via an intermediate image that represents the image from the viewpoint of the camera, depending on the state of the display world including the virtual object. A recording medium on which a program is recorded. [Explanation of symbols]

[0106] 1 Information processing system, 10 Information processing device, 11 Recording device, 14 Imaging device, 16 Input device, 70 Data acquisition unit, 72 Captured image acquisition unit, 74 Sensor data acquisition unit, 75 Operation information acquisition unit, 76 Display image generation unit, 78 Output unit, 80 Object surface detection unit, 82 Object surface data storage unit, 84 Object placement unit, 86 Object data storage unit, 90 Pointed object detection unit, 94 See-through image generation unit, 96 Object drawing unit, 97 Intermediate image generation unit, 98 Synthesis unit, 100 Head-mounted display, 222 CPU, 224 GPU.

Claims

1. a captured image acquisition unit that acquires data of an image captured by a camera; an object placement unit that places a virtual object to be operated by a user in a virtual three-dimensional space; a display image generation unit that draws an image of the virtual object and generates a display image by combining the image with the captured image; an output unit that outputs data of the display image; Equipped with The display image generation device is characterized in that the display image generation unit switches whether or not to draw an image of the virtual object via an intermediate image that represents the image from the camera's viewpoint, depending on the state of the display world including the virtual object.

2. 2. The display image generating device according to claim 1, wherein the display image generating unit switches to drawing without going through the intermediate image when the virtual object enters a predetermined range from another virtual object that cannot be drawn via the intermediate image.

3. the object placement unit places an instruction object, which is used by a user to indicate a position in the display world, as the virtual object; The display image generating device according to claim 1, characterized in that when the instruction object indicates another virtual object that cannot be drawn via the intermediate image, the display image generating unit switches to drawing that does not go via the intermediate image.

4. 4. The display image generating device according to claim 3, wherein when a virtual object using a template provided by middleware is specified as the other virtual object, the display image generating unit switches to drawing that does not go through the intermediate image.

5. The display image generating device according to any one of claims 1 to 4, characterized in that the display image generating unit displays the captured image projected onto a projection surface set in the three-dimensional space on the plane of the display image, and when drawing an image of the virtual object via the intermediate image, displays the intermediate image displayed on the projection surface on the plane of the display image.

6. the captured image acquisition unit acquires data of the captured image by a camera provided in the head-mounted display; 5. The display image generating device according to claim 1, wherein the output unit outputs the display image data to the head-mounted display.

7. acquiring data of an image captured by a camera; placing a virtual object to be operated by a user in a virtual three-dimensional space; a step of generating a display image by drawing an image of the virtual object and combining the image with the captured image; outputting data of the display image; Including, An image display method characterized in that the step of generating a display image involves switching whether or not to draw an image of the virtual object via an intermediate image that represents the image from the camera's viewpoint, depending on the state of the display world including the virtual object.

8. A function to acquire data of images taken by the camera, placing a virtual object to be operated by a user in a virtual three-dimensional space; a function of drawing an image of the virtual object and generating a display image by combining the image with the captured image; a function of outputting data of the display image; This is realized by a computer, A computer program characterized in that the function of generating the display image switches whether or not to draw the image of the virtual object via an intermediate image that represents the image from the viewpoint of the camera, depending on the state of the display world including the virtual object.