Information processing device
Patent Information
- Application Number
- JP2025028663
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2026-09-07
AI Technical Summary
【0012】 本発明によれば、小さい処理負荷で違和感のない複合現実空間(現実空間と仮想空間を融合した空間)の画像を得ることができる。
Smart Images

Figure 2026141904000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing apparatus, and particularly to a technology for generating images of a mixed reality space. [Background Art]
[0002] MR (Mixed Reality) technology is known as a technology for fusing the real world and the virtual world in real time and seamlessly. As a system using MR technology, an MR system including a video see-through head-mounted display (hereinafter simply referred to as HMD) has been proposed. In an MR system, a range that substantially matches the range of real space observed from the pupil position of a user wearing the HMD on the head is captured by an imaging unit provided on the HMD. A display image is then generated by compositing (overlaying) CG (Computer Graphics) on the obtained captured image. By displaying this display image on the HMD, the user can experience the MR space.
[0003] As a technology used in MR systems, when rendering CG, a technology has been proposed that renders a portion corresponding to the region gazed by the user (and the periphery thereof) at high resolution, and other portions at low resolution. This technology (method) is called foveated rendering. According to foveated rendering, the load of CG rendering processing can be reduced. Humans have the visual characteristic that the amount of recognizable information decreases as the distance from the focal position increases. Therefore, the effect of the resolution reduction caused by foveated rendering on the user's experience is extremely small (the resolution reduction does not affect the user's experience).
[0004] In MR systems, many processes are performed from image acquisition to display, including exposure and image processing to obtain captured images, detection of the HMD's position and orientation, image processing to obtain the display image (such as CG rendering and compositing of CG with captured images), and data transmission. The time required for these processes can cause the displayed image to lag behind the user's head movements, and this lag can cause discomfort to the user.
[0005] Patent Document 1 discloses a technique for reprojecting images after foveated rendering based on the position and orientation of the HMD. [Prior art documents] [Patent Documents]
[0006] [Patent Document 1] Japanese Patent Publication No. 2019-028368 [Overview of the project] [Problems that the invention aims to solve]
[0007] However, the technology disclosed in Patent Document 1 generates a single image that includes both high-resolution and low-resolution regions, requiring data transmission for the same number of pixels in the low-resolution region as in the high-resolution region. Furthermore, reprojection of the image after foveated rendering can cause the high-resolution region to move away from the user's viewing area, and this movement can cause discomfort to the user.
[0008] The present invention aims to provide a technology that can obtain images of a seamless augmented reality space (a space that fuses the real and virtual spaces) with minimal processing load. [Means for solving the problem]
[0009] A first aspect of the present invention is a first reality image that represents real space and has a first resolution, and the Real image acquisition means for acquiring a second real image that represents a part of a first real image and has a second resolution higher than the first resolution; information acquisition means for acquiring position and orientation information which is information regarding the position and orientation of a head-mounted display (HMD); virtual image acquisition means for acquiring a first virtual image that corresponds to the first real image, represents a virtual object and has a third resolution, and a second virtual image that corresponds to the second real image, represents a part of the first virtual image and has a fourth resolution higher than the third resolution; and based on the difference between the position and orientation information of a first timing and the position and orientation information of a second timing later than the first timing, the first timing The information processing apparatus is characterized by having correction means for correcting a first virtual image at a first timing to correspond to a first real image at a third timing later than the timing, and correcting a second virtual image at a first timing to correspond to a second real image at a third timing, and combining means for generating a first composite image by combining the first virtual image corrected by the correction means with the first real image at a third timing, generating a second composite image by combining the second virtual image corrected by the correction means with the second real image at a third timing, and generating a display image to be displayed on the HMD by combining the second composite image with the first composite image.
[0010] A second aspect of the present invention is a method for obtaining a first reality image that represents real space and has a first resolution; a second reality image that represents a part of the first reality image and has a second resolution higher than the first resolution; a method for obtaining position and orientation information that is information relating to the position and orientation of a head-mounted display (HMD); a first virtual image that corresponds to the first reality image, represents a virtual object, and has a third resolution; a second virtual image that corresponds to the second reality image, represents a part of the first virtual image, and has a fourth resolution higher than the third resolution; and the difference between position and orientation information at a first timing and position and orientation information at a second timing later than the first timing. The control method for an information processing device is characterized by comprising the steps of: correcting a first virtual image at a first timing to correspond to a first real image at a third timing that is later than the first timing, based on minutes; correcting a second virtual image at a first timing to correspond to a second real image at a third timing, based on the difference; generating a first composite image by combining the corrected first virtual image with the first real image at the third timing; generating a second composite image by combining the corrected second virtual image with the second real image at the third timing; and generating a display image to be shown on the HMD by combining the second composite image with the first composite image.
[0011] A third aspect of the present invention is a program for causing a computer to function as one of the means of the information processing apparatus described above. [Effects of the Invention]
[0012] According to the present invention, it is possible to obtain images of a seamless mixed reality space (a space that fuses the real space and the virtual space) with a small processing load. [Brief explanation of the drawing]
[0013] [Figure 1] This is an external view of the MR system. [Figure 2]This is a block diagram of the MR system. [Figure 3] This is a schematic diagram of the imaging operation. [Figure 4] This is a schematic diagram of the gaze area, the overall area, the gaze virtual image, and the overall virtual image. [Figure 5] This is a schematic diagram of the processing in the gaze CG correction unit. [Figure 6] This is a schematic diagram of the processing in the overall CG correction section. [Figure 7] This is a schematic diagram of the processing in the fixation image synthesis unit and the overall image synthesis unit. [Figure 8] This is a schematic diagram of the processing in the display image synthesis unit. [Figure 9] This is a schematic diagram of an image transmitted from an image processing device to an HMD (Head-Mounted Display). [Figure 10] This is a block diagram of the MR system. [Modes for carrying out the invention]
[0014] (Embodiment 1) Embodiment 1 of the present invention will now be described. Figure 1 is an external view showing an example of the appearance of an MR (Mixed Reality) system according to Embodiment 1. The MR system in Figure 1 includes a head-mounted display (HMD) 101 and an image processing device 104. The image processing device 104 includes a controller 102 and a personal computer (PC) 103.
[0015] The HMD 101 is a video see-through head-mounted display. The HMD 101 captures images of real space and receives virtual images (images of virtual objects, CG images) from the controller 102. Then, the HMD 101 generates a display image, which is an image of a mixed reality space (a space merging real space and virtual space) obtained by combining (superimposing) a virtual image onto a real image (a captured image of real space), and displays the generated display image. The HMD 101 transmits the real image to the controller 102. Furthermore, the HMD 101 detects the position and orientation of the HMD 101, and transmits position and orientation information, which is the detection result, to the controller 102. The position and orientation information of the HMD 101 only needs to be information relating to the position and orientation of the HMD 101, and for example, is information indicating the position and orientation of the HMD 101.
[0016] Note that the method for supplying power to the HMD 101 is not particularly limited; the HMD 101 may be operated with power supplied from the controller 102, or may be operated with power supplied from a battery provided in the HMD 101. The connection between the HMD 101 and the controller 102 may be a wired connection or a wireless connection. The HMD 101 and the controller 102 may be connected via both wired and wireless connections. The controller 102 may be a part of the HMD 101, or may be a part of the PC 103. The image processing apparatus 104 (the controller 102 and the PC 103) may be a part of the HMD 101. A part of the processing of the HMD 101 may be performed by the controller 102 or the PC 103. That is, the information processing apparatus to which the present invention is applied may be included in the HMD 101, may be included in the controller 102, or may be included in the PC 103.
[0017] The controller 102 relays communication between the HMD 101 and the PC 103. The controller 102 performs various image processing (resolution conversion, color space conversion, distortion correction, encoding, etc.) on the real image transmitted from the HMD 101, and transmits the real image after image processing to the PC 103. Similarly, the controller 102 transmits the position and orientation information transmitted from the HMD 101 to the PC 103. Furthermore, the controller 102 performs various image processing (resolution conversion, color space conversion, distortion correction, encoding, etc.) on the virtual image transmitted from the PC 103, and transmits the virtual image after image processing to the HMD 101.
[0018] The PC 103 estimates the position and orientation of the HMD 101 (the position and orientation of the imaging unit included in the HMD 101) based on the captured image and position and orientation information received from the controller 102. Then, the PC 103 generates a virtual image, which is an image of a virtual object (virtual space) viewed from the estimated position and orientation, and transmits the generated virtual image to the controller 102.
[0019] FIG. 2 is a block diagram showing a configuration example of the MR system according to Embodiment 1. The HMD 101 includes an imaging unit 201, a position and orientation sensor 202, a display unit 203, an HMD-I / F 204, a gaze CG correction unit 205, a full-screen CG correction unit 206, a gaze image combining unit 207, a full-screen image combining unit 208, and a display image combining unit 209. The image processing apparatus 104 includes an image processing apparatus-I / F 211, a gaze area setting unit 212, a position and orientation estimation unit 213, a content DB 214, and a rendering unit 215.
[0020] The imaging unit 201 is a camera that captures images of the real world. In Embodiment 1, the imaging unit 201 has an imaging unit for the user's (the user wearing the HMD 101 on their head) left eye and an imaging unit for the user's right eye. The imaging unit for the left eye generates and outputs a real image for the left eye (a real image representing what the left eye sees). The imaging unit for the right eye generates and outputs a real image for the right eye (a real image representing what the right eye sees). In other words, the imaging unit 201 generates and outputs a stereo image that includes two images with parallax (a real image corresponding to the left eye and a real image corresponding to the right eye). The generation and output of the stereo image (a real image corresponding to the left eye and a real image corresponding to the right eye) is performed repeatedly at a predetermined frame rate.
[0021] Each of the imaging units, one for the left eye and one for the right eye, has an optical system (lens) and an image sensor. Light from the outside world enters the image sensor through the optical system, and the image sensor outputs an image corresponding to the incident light. Preferably, the optical axis direction (imaging direction) of the left eye imaging unit is approximately the same as the line of sight of the left eye, and preferably, the optical axis direction (imaging direction) of the right eye imaging unit is approximately the same as the line of sight of the right eye.
[0022] The image sensor used in the imaging unit 201 is determined by considering various parameters such as the number of pixels, image quality, noise, sensor size, power consumption, and cost. A rolling shutter image sensor or a global shutter image sensor may be used. Depending on the application of the real image, one of the rolling shutter image sensor or the global shutter image sensor may be selected and used. Both a rolling shutter image sensor and a global shutter image sensor may be used (combined) to generate a single real image. For example, when acquiring a real image used to generate a display image (a real image used to synthesize virtual images), a rolling shutter image sensor capable of acquiring a higher-quality real image may be used. When acquiring a real image used for various alignment purposes, a global shutter image sensor capable of acquiring a real image without image blur may be used. Image blur is a phenomenon that occurs in the rolling shutter system, where exposure processing is performed line by line. Because the timing of exposure processing for each line is different, when the imaging unit or the subject moves, the subject is deformed and recorded as if it is blurred. In the case of the global shutter system, exposure processing for all lines is performed simultaneously, so image blur does not occur.
[0023] In Embodiment 1, the imaging unit 201 performs foveated capture, which images both the user's gaze area and the entire area including the gaze area.
[0024] Generally, the more pixels an image has, the greater the processing load it places on the image. However, it is known that human vision makes it more difficult to perceive shapes and colors in the peripheral field of view compared to shapes and colors in the central field of view. Because of this visual characteristic, even if an image is displayed where the resolution of the peripheral field of view is lower than the resolution of the central field of view (gaze area), the user is unlikely to perceive that the peripheral field of view has lower resolution. By limiting the high-resolution area to the gaze area, the processing load can be reduced without compromising the user's MR experience. Therefore, in order to obtain a display image where the resolution of the peripheral field of view is lower than the resolution of the gaze area, foveated capture is used to acquire a high-resolution real image representing only the gaze area (gaze real image) and a low-resolution real image representing the entire area (overall real image).
[0025] Figure 3(A) is a schematic diagram showing an example of the imaging operation (accumulation and readout operation, actual image acquisition) of the imaging unit 201. The horizontal axis in Figure 3(A) represents time, and the vertical axis in Figure 3(A) represents the position (row) of the image sensor. When the image sensor of the imaging unit 201 is a general CMOS sensor (rolling shutter type), as shown in Figure 3(A), the timing of the imaging operation of each line is different, and the imaging operation of one frame (the imaging operation of one image) is represented by a parallelogram. For example, suppose that a gazed-on reality image is acquired such that one pixel of the imaging unit 201 (image sensor) corresponds to one pixel of the gazed-on reality image, and that a whole-real-world image is acquired such that four pixels in a 2x2 arrangement of the imaging unit 201 correspond to one pixel of the whole-real-world image. In this case, the resolution of the whole-real-world image will be 1 / 4 of the resolution of the gazed-on reality image. The values of multiple pixels of the imaging unit 201 may be combined to obtain the value of one pixel of the whole-real-world image, or some of the multiple pixels of the imaging unit 201 may be downsampled and the values of the remaining pixels may be used as the values of multiple pixels in the whole-real-world image.
[0026] In Figure 3(A), the imaging operation to acquire the overall reality image and the imaging operation to acquire the gaze reality image are performed alternately. In Figure 3(A), the imaging operation to acquire the overall reality image takes 4.54 ms. The imaging operation to acquire the gaze reality image also takes 4.54 ms. As will be explained in detail later, one display image is generated using one overall reality image and one gaze reality image, so the display image is generated and displayed at 110 fps.
[0027] Figure 3(B) is a schematic diagram showing an example of the gaze region, the overall region, the gazed-on real image, and the overall real image. In Figure 3(B), the central part of the overall region 301 is set as the gaze region 302. The gazed-on real image 303 represents the gaze region 302, and the overall real image 304 represents the overall region 301. In Figure 3(B), the gazed-on real image 303 is a 1:1 (1x) image, and the overall real image 304 is a 1 / 4 (1 / 4) image. When the magnification of the gazed-on real image 303 and the magnification of the overall real image 304 are combined, the resolution (number of pixels per inch) of the overall real image 304 becomes 1 / 4 of the resolution of the gazed-on real image 303. Note that in Figure 3(B), the central part of the overall region 301 is set as the gaze region 302, but the position and size of the gaze region 302 are not limited to those shown in Figure 3(B). The position and size of the gaze region are set by the gaze region setting unit 212.
[0028] The position and attitude sensor 202 detects the position and attitude of the HMD 101 and outputs position and attitude information of the HMD 101. The position and attitude sensor 202 is composed of, for example, a magnetic sensor (including a geomagnetic sensor), an ultrasonic sensor, an acceleration sensor, an angular velocity sensor, and the like.
[0029] The HMD-I / F204 is a communication interface that communicates with the image processing device 104. The HMD-I / F204 transmits the real image output from the imaging unit 201 and the position and orientation information output from the position and orientation sensor 202 to the image processing device 104. The HMD-I / F204 also receives a virtual image output from the drawing unit 215 from the image processing device 104 (virtual image acquisition). The HMD-I / F204 also transmits and receives setting information and control signals for each device.
[0030] The image processing unit-I / F211 is a communication interface that communicates with the HMD101. The image processing unit-I / F211 receives real images output from the imaging unit 201 and position and orientation information output from the position and orientation sensor 202 from the HMD101. The image processing unit-I / F211 also transmits virtual images output from the drawing unit 215 to the HMD101. The image processing unit-I / F211 also transmits and receives setting information and control signals for each device.
[0031] The position and orientation estimation unit 213 estimates the position and orientation of the left eye imaging unit and the right eye imaging unit based on the real image and position and orientation information received from the HMD 101 via the image processing device I / F 211. Various known techniques can be used for this estimation process.
[0032] Content DB214 is a database that pre-stores various data (virtual space data) necessary for rendering virtual images. For example, virtual space data includes data that defines virtual objects (for example, data that defines the geometry, color, texture, position, and orientation of virtual objects). Virtual space data also defines virtual light sources. This also includes data (for example, data defining the type, position, and orientation of virtual light sources).
[0033] The drawing unit 215 generates a virtual image using the virtual space data stored in the content DB 214. The drawing unit 215 generates a virtual image for the left eye and a virtual image for the right eye. The drawing unit 215 generates the image of the virtual object as seen from the position and orientation of the left eye imaging unit estimated by the position and orientation estimation unit 213 as the virtual image for the left eye. Similarly, the drawing unit 215 generates the image of the virtual object as seen from the position and orientation of the right eye imaging unit estimated by the position and orientation estimation unit 213 as the virtual image for the right eye.
[0034] In Embodiment 1, the drawing unit 215 performs foveated rendering, generating two virtual images corresponding to the user's gaze area and the overall area including the gaze area. The foveated rendering in Embodiment 1 differs from conventional foveated rendering, which generates a single image in which the resolution of the gaze area and the resolution of the other areas are different. In order to obtain a display image in which the resolution of the peripheral field of view is lower than the resolution of the gaze area, foveated rendering is used to obtain a high-resolution virtual image representing only the gaze area (gaze virtual image) and a low-resolution virtual image representing the overall area (overall virtual image). The gaze virtual image is a virtual image corresponding to the gaze real image, and the overall virtual image is a virtual image corresponding to the overall real image.
[0035] Figure 4 is a schematic diagram showing an example of a gaze region, an overall region, a gaze virtual image, and an overall virtual image. In Figure 4, the central part of the overall region 401 is set as the gaze region 405. The gaze virtual image 403 represents the gaze region 405, and the overall virtual image 404 represents the overall region 401. Region 402 is a region of the same position and size as the gaze region (the gaze region considered in foveated capture) represented by the gaze reality image. The gaze region 405 (the gaze region considered in foveated rendering) is larger than region 402 and includes the entirety of region 402. Thus, in Embodiment 1, the range of the gaze virtual image is wider than the range of the gaze reality image. In Figure 4, the gaze virtual image 403 is a 1:1 (1x) image, and the overall virtual image 404 is a 1 / 4 (1 / 4) image. When the magnification of the gazed virtual image 403 and the magnification of the overall virtual image 404 are combined, the resolution (number of pixels per inch) of the overall virtual image 404 becomes 1 / 4 of the resolution of the gazed virtual image 403. In Figure 4, the central part of the overall area 401 is set as the gazed area 405, but the position and size of the gazed area 405 are not limited to those shown in Figure 4. The position and size of the gazed area are set by the gazed area setting unit 212. The range of the gazed virtual image may be equal to the range of the gazed real image.
[0036] In Embodiment 1, the resolution of the gazed-on real image and the resolution of the gazed-on virtual image are assumed to be equal, but they may be different. Similarly, the resolution of the overall real image and the resolution of the overall virtual image are assumed to be equal, but they may be different.
[0037] The gaze area setting unit 212 sets the position (coordinates) and size of the gaze area used by the imaging unit 201, the drawing unit 215, the gaze CG correction unit 205, and the display image synthesis unit 209. For example, the gaze area setting unit 212 sets an area specified by the user as the gaze area. The gaze area may be a fixed area, such as the center of the overall area.
[0038] The gaze CG correction unit 205 corrects the gaze virtual image received from the image processing device 104 via the HMD-I / F 204 based on the change in position and orientation information output from the position and orientation sensor 202. Here, the timing at which the position and orientation information (and the corresponding real image) is acquired and the virtual image is generated is defined as timing T1, and the timing at which the gaze CG correction unit 205 (and the overall CG correction unit 206) performs processing is defined as timing T2. Timing T2 is after timing T1. The gaze CG correction unit 205 corrects the position and orientation information at timing T1 Based on the difference between the position and orientation information at timing T2, the gaze virtual image at timing T1 is corrected to correspond to the gaze real image at timing T2. The gaze CG correction unit 205 modifies the shape and size of the gaze virtual image so that an image of the virtual object as seen from the position and orientation of the HMD 101 (imaging unit) at timing T2 is obtained. The process of changing the shape and size of the gaze virtual image includes horizontal shift, vertical shift, enlargement, reduction, geometric transformation (e.g., homography transformation). Then, the gaze CG correction unit 205 extracts an image of the region that matches the region of the gaze real image at timing T2 (gaze region) from the gaze virtual image after the shape and size have been modified, and outputs the extracted image as the corrected gaze virtual image.
[0039] The gaze CG correction unit 205 may also correct the gaze virtual image at timing T1 to correspond to the gaze real image at timing T3, which is later than timing T2, based on the change in position and orientation information from timing T1 to timing T2. Timing T3 is, for example, the timing at which the gaze image synthesis unit 207 (and the overall image synthesis unit 208) performs processing.
[0040] Figures 5(A) to 5(D) are schematic diagrams showing an example of the processing of the gaze-on CG correction unit 205.
[0041] Figure 5(A) shows an example where the position and orientation information does not change. Image 501 is the gaze-on virtual image before correction, and region 502 is the region that corresponds to the region of the gaze-on real image at the time of generation of the gaze-on virtual image 501 (timing T1). If the position and orientation information does not change, the imaging range (angle of view) of the imaging unit 201 does not change, so the corrected gaze-on virtual image 503 is obtained by extracting the image of region 502 from the gaze-on virtual image 501. The same procedure may be followed when the amount of change in position and orientation information is small (below the threshold).
[0042] Figure 5(B) shows an example where the position and orientation information changes so that the imaging range of the imaging unit 201 shifts to the upper right. Image 504 is the gaze virtual image before correction, and region 505 is the region that matches the region of the gaze real image at the time of generation of the gaze virtual image 504 (timing T1). Region 506 is the region that matches the region of the gaze real image after the shift in the imaging range (timing T2). Region 506 is determined by shifting region 505 to the upper right based on the change in position and orientation information. Then, the corrected gaze virtual image 507 is obtained by extracting the image of region 506 from the gaze virtual image 504. Alternatively, the gaze virtual image 504 may be shifted to the lower left based on the change in position and orientation information, and the corrected gaze virtual image 507 may be obtained by extracting the image of region 505 from the shifted gaze virtual image 504.
[0043] Figure 5(C) shows an example where the size of the uncorrected gaze virtual image is equal to the size of the gaze real image. Image 508 is the uncorrected gaze virtual image, and the region of gaze virtual image 508 matches the region of the gaze real image at the time gaze virtual image 508 is generated (timing T1). Region 509 is the region that matches the region of the gaze real image after the shift of the imaging range (timing T2). Similar to the case of Figure 5(B), region 509 is determined by shifting the region of gaze virtual image 508 in the upper right direction based on the change in position and orientation information. Then, the corrected gaze virtual image 510 is obtained by extracting the portion of gaze virtual image 508 in region 509. Gaze virtual image 510 includes region 511 (non-drawn region) that is not gaze virtual image 508. Thus, when the size of the uncorrected gaze virtual image is equal to the size of the gaze real image, the corrected gaze virtual image is likely to include a non-drawn region. If the corrected gaze-on virtual image contains non-drawn areas, a display image with some parts missing will be generated. Therefore, in Embodiment 1, the size of the gaze-on virtual image before correction is made larger than the size of the gaze-on real image.
[0044] Figure 5(D) shows the case where the positional information changes so that a certain position is imaged from a different direction. An example of this is shown. Image 512 is the gaze-on virtual image before correction, and region 513 is the region that matches the region of the gaze-on real image at the time of generation of the gaze-on virtual image 512 (timing T1). Region 515, which is part of the virtual object 514, is depicted in the gaze-on virtual image 512. In Figure 5(D), the outline of the virtual object 514 is shown with a dashed line outside the gaze-on virtual image 512 so that the entire virtual object 514 can be seen. Region 513 does not contain the virtual object 514. Image 516 is the gaze-on virtual image in which region 515 has been deformed to change the orientation of the virtual object 514 based on changes in position and orientation information. Region 517 is the region that matches the region of the gaze-on real image after the change in the position and orientation of the HMD 101 (timing T2). Here, for the sake of simplicity, we assume that region 517 is equal to region 513. Region 517 includes a portion of region 518 (part of virtual object 514), which is a modified version of region 515. By extracting the image of region 517 from the gazed virtual image 516, the corrected gazed virtual image 519 is obtained. In this way, even when the size of the gazed virtual image before correction is larger than the size of the gazed real image, and the virtual object falls within the region matching the region of the gazed real image, a corrected gazed virtual image in which the virtual object is suitably drawn can be obtained.
[0045] The overall CG correction unit 206 corrects the overall virtual image received from the image processing device 104 via the HMD-I / F 204 based on the change in position and orientation information output from the position and orientation sensor 202. Here, the timing at which the position and orientation information (and the corresponding real image) is acquired and the virtual image is generated is defined as timing T1, and the timing at which the overall CG correction unit 206 (and the gaze CG correction unit 205) performs processing is defined as timing T2. Timing T2 is after timing T1. Based on the difference between the position and orientation information at timing T1 and the position and orientation information at timing T2, the overall CG correction unit 206 corrects the overall virtual image at timing T1 to correspond to the overall real image at timing T2. The overall CG correction unit 206 changes the shape and size of the overall virtual image so that an image of the virtual object as seen from the position and orientation of the HMD 101 (imaging unit) at timing T2 is obtained. The process of changing the shape and size of the overall virtual image includes horizontal shift, vertical shift, enlargement, reduction, geometric transformation (e.g., homography transformation). The overall CG correction unit 206 then extracts images of the region that matches the region of the overall real image at timing T2 (overall region) from the overall virtual image after the shape and size have been changed, and outputs the extracted images as the corrected overall virtual image.
[0046] Figure 6 is a schematic diagram showing an example of the processing of the overall CG correction unit 206. The processing of the overall CG correction unit 206 is the same as the processing of the gaze CG correction unit 205. Figure 6 shows an example in which the position and orientation information changes so that the imaging range of the imaging unit 201 shifts to the upper right. Image 601 is the overall virtual image before correction, and the region of the overall virtual image 601 matches the region of the overall real image at the time of generation of the overall virtual image 601 (timing T1). Region 602 is the region that matches the region of the overall real image after the shift in the imaging range (timing T2). Region 602 is determined by shifting the region of the overall virtual image 601 to the upper right based on the change in position and orientation information. Then, the corrected overall virtual image 603 is obtained by extracting the portion of the overall virtual image 601 in region 602. The overall virtual image 603 includes region 604 (non-drawn region) which is not part of the overall virtual image 601, but since region 604 corresponds to the peripheral field of view, it is not easily perceived by the user. Region 604 may be displayed in black, or an image may be drawn in region 604 by external interpolation. Also, similar to the gaze region, the size of the overall virtual image before correction may be larger than the size of the overall real image.
[0047] The gaze image synthesis unit 207 generates a gaze composite image by combining the gaze reality image (the gaze reality image at timing T2 described above) output from the imaging unit 201 with the gaze virtual image corrected by the gaze CG correction unit 205. For example, the gaze virtual image is synthesized with the gaze reality image using methods such as chroma keying or alpha blending. Depth information of the real space Further information (depth information) of the virtual object and depth information (depth information) may be acquired, and a more advanced synthesis process that takes into account the spatial relationship between the real object and the virtual object may be performed using this depth information. Various known techniques can be used to acquire depth information.
[0048] Figure 7(A) is a schematic diagram showing an example of the processing of the gaze image synthesis unit 207. Figure 7(A) shows an example where the position and orientation information changes so that the imaging range of the imaging unit 201 shifts to the upper right. Image 705 is the gaze virtual image before correction by the gaze CG correction unit 205, image 701 is the gaze real image at the time of generation of the gaze virtual image 705 (timing T1), and region 703 is the region that matches the region of the gaze real image 701. Image 702 is the gaze real image after the shift in the imaging range (timing T2), and region 704 is the region that matches the region of the gaze real image 702. The gaze CG correction unit 205 obtains the corrected gaze virtual image 706 by extracting the image of region 704 from the gaze virtual image 705. In the gaze image synthesis unit 207, a gaze-synthesized image 707 is generated by synthesizing the gaze-virtual image 706 with the gaze-real image 702.
[0049] The overall image synthesis unit 208 generates an overall composite image by combining the overall real image output from the imaging unit 201 (the overall real image at timing T2 described above) with the overall virtual image corrected by the overall CG correction unit 206. For example, the overall virtual image is synthesized with the overall real image by methods such as chroma keying or alpha blending. Depth information of the real space and depth information of the virtual object may be acquired, and a more advanced synthesis process that takes into account the front-to-back relationship between the real object and the virtual object may be performed using this depth information. Various known techniques can be used to acquire depth information.
[0050] Figure 7(B) is a schematic diagram showing an example of the processing of the overall image synthesis unit 208. Similar to Figure 7(A), Figure 7(B) shows an example where the position and orientation information changes so that the imaging range of the imaging unit 201 shifts to the upper right. Image 710 is the overall virtual image before correction by the overall CG correction unit 206, and Image 708 is the overall real image at the time of generation of the overall virtual image 710 (timing T1), with the region of the overall virtual image 710 corresponding to the region of the overall real image 708. Image 709 is the overall real image after the shift in the imaging range (timing T2), with region 711 corresponding to the region of the overall real image 709. The overall CG correction unit 206 obtains the corrected overall virtual image 712 by extracting the portion of the overall virtual image 710 in region 711. The overall image synthesis unit 208 generates the overall synthesized image 713 by synthesizing the overall virtual image 712 with the overall real image 709.
[0051] The display image synthesis unit 209 generates a display image by combining the gaze-focused image synthesis unit 207 with the overall image synthesis unit 208.
[0052] Figure 8 is a schematic diagram showing an example of the processing of the display image synthesis unit 209. Image 801 is a gaze-focused composite image generated by the gaze-focused image synthesis unit 207, and image 802 is a whole-image composite image generated by the whole-image synthesis unit 208. Assume that the gaze-focused composite image 801 is a 1:1 (1x) image, and the whole-image composite image 802 is a 1 / 4 (1 / 4) image. The display image synthesis unit 209 scales the whole-image composite image 802 to the same magnification as the gaze-focused composite image 801, and then synthesizes the gaze-focused composite image 801 into the scaled whole-image composite image 803. The gaze-focused composite image 801 is synthesized at the position set by the gaze-focused area setting unit 212. This generates the display image 804.
[0053] Furthermore, the control signal determines whether the processing of the display image synthesis unit 209 is turned on (executed) or off (not executed). The display image synthesis unit 209 may be switched. Normally, when a user turns their head quickly, they cannot clearly perceive the real space. Therefore, in such cases, displaying a high-resolution gaze-composite image may cause discomfort. For this reason, the display image synthesis unit 209 may not perform processing if the amount of change in position and orientation information from timing T1 to timing T2 (the difference between the position and orientation information at timing T1 and the position and orientation information at timing T2) is greater than or equal to a threshold. If the display image synthesis unit 209 does not perform processing, the overall composite image is adopted as the display image. Also, the angular velocity of movement that generally causes motion sickness is 60 [deg / sec]. For this reason, the display image synthesis unit 209 may not perform processing if the amount of change in position and orientation information from timing T1 to timing T2 corresponds to an orientation change with an angular velocity of 60 [deg / sec] or more. The display image synthesis unit 209 may not perform processing if the corrected gaze-virtual image includes a non-drawn area.
[0054] The display unit 203 displays the display image generated by the display image synthesis unit 209. In Embodiment 1, the display unit 203 has a display unit for the left eye and a display unit for the right eye. The above-described process is performed for the left eye and the right eye respectively, and a display image for the left eye and a display image for the right eye are generated. The display image for the left eye is displayed on the display unit for the left eye, and the display image for the right eye is displayed on the display unit for the right eye.
[0055] By the way, in order to display the gaze-composite image at the set position, it is necessary to scale the entire gaze-composite image after it has been buffered. For this reason, it is preferable that the HMD 101 receives the gaze-virtual image from the image processing device 104 before the overall virtual image and generates the gaze-composite image before the overall composite image. It is also preferable that the image processing device 104 transmits the gaze-virtual image to the HMD 101 before the overall virtual image. Similarly, it is preferable that the imaging unit 201 outputs the gaze-real image before the overall real image. By doing so, the latency until display can be shortened.
[0056] Figure 9 is a schematic diagram showing an example of an image transmitted from the image processing device 104 to the HMD 101. Figure 9 shows an example where a gaze virtual image, depth information corresponding to the gaze virtual image, a whole virtual image, and depth information corresponding to the whole virtual image are transmitted from the image processing device 104 as a single image. In Figure 9, there is a gaze virtual image for the left eye, depth information corresponding to the gaze virtual image for the left eye, a whole virtual image for the left eye, and depth information corresponding to the whole virtual image for the left eye. Similarly, there is a gaze virtual image for the right eye, depth information corresponding to the gaze virtual image for the right eye, a whole virtual image for the right eye, and depth information corresponding to the whole virtual image for the right eye.
[0057] In Figure 9, the gaze-focused virtual image and the corresponding depth information (depth image) are arranged in the upper half of the image, while the overall virtual image and the corresponding depth information (depth image) are arranged in the lower half of the image. When data is transmitted line by line from the top to the bottom of the image, the arrangement in Figure 9 allows the gaze-focused virtual image to be transmitted before the overall virtual image.
[0058] Depth information is used, for example, in compositing processes that take into account the spatial relationships between real and virtual objects. Similar to depth information, transparency information for alpha blending may also be included.
[0059] Furthermore, for improved versatility, it is preferable that the size of the image to be transmitted (vertical V pixels × horizontal H pixels) be a size defined by the VESA standard.
[0060] Furthermore, while we have described an example where multiple data points are transmitted as a single image file, each data point may also be transmitted individually.
[0061] As described above, according to Embodiment 1, by using foveated capture and foveated rendering, images of a mixed reality space can be obtained with a small processing load. Furthermore, by processing the image of the gaze region separately from the image of the overall region, it is possible to obtain a mixed reality space image with high resolution in the gaze region and a natural appearance. In addition, since images of the gaze region and the overall region are generated and transmitted between devices, the amount of data transmitted between devices can be reduced compared to generating and transmitting a single image containing high-resolution and low-resolution regions between devices.
[0062] (Embodiment 2) Embodiment 2 of the present invention will now be described. In Embodiment 1, the gaze area was assumed to be a pre-set area. In Embodiment 2, user gaze information is acquired, and the gaze area is dynamically changed based on the gaze information. The gaze information is, for example, coordinate information indicating the position (gaze position) where the user's gaze is directed on the display surface of the display unit 203. The gaze information may also be angle information indicating the direction of the gaze (gaze direction).
[0063] Figure 10 is a block diagram showing an example configuration of the MR system according to Embodiment 2. The eyeball imaging unit 1020 of the HMD 101 is a camera that images the user's eyeballs and acquires eyeball images. The gaze information acquisition unit 1021 acquires the user's gaze information by analyzing the eyeball images acquired by the eyeball imaging unit 1020. In Embodiment 2, the area in real space that the user's gaze is directed towards is used as the gaze region. The imaging unit 1001 performs the same processing as the imaging unit 201 in Embodiment 1 (Figure 1). However, the imaging unit 1001 dynamically changes the gaze region (the area of the gazed real image) based on the gaze information acquisition unit 1021. The gaze region setting unit 1012 of the image processing device 104 receives the user's gaze information from the HMD 101 via the image processing device-I / F 211. The gaze area setting unit 1012 then sets or dynamically changes the position (coordinates) and size of the gaze area used by the drawing unit 215, gaze CG correction unit 205, and display image synthesis unit 209 based on the gaze information. Note that the processing other than the acquisition of gaze information and the dynamic modification of the gaze area is the same as in Embodiment 1.
[0064] As described above, according to Embodiment 2, since the gaze area is dynamically changed using the user's gaze information, the resolution of the gaze area is high, and images of a mixed reality space that feel natural can be obtained with higher accuracy.
[0065] The various controls described above may or may not be performed by a single piece of hardware (e.g., a processor or circuit). Multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) may share the processing to control the entire device.
[0066] Furthermore, the above-mentioned processors are processors in a broad sense, including general-purpose processors and specialized processors. General-purpose processors include, for example, CPUs (Central Processing Units), MPUs (Micro Processing Units), and DSPs (Digital Signal Processors). Specialized processors include, for example, GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices). Programmable logic devices include, for example, FPGAs (Field Programmable Gate Arrays) and CPLDs (Complex Programmable Logic Devices).
[0067] Furthermore, the embodiments described above (including variations) are merely examples and fall within the scope of the present invention. Configurations obtained by appropriately modifying or changing the above-described configurations are also included in the present invention. Configurations obtained by appropriately combining the above-described configurations are also included in the present invention.
[0068] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit that implements one or more functions.
[0069] This embodiment includes the following configurations, methods, and programs. (Composition 1) A first reality image that represents real space and has a first resolution, A second reality image that represents a portion of the first reality image and has a second resolution higher than the first resolution, A means for acquiring real images to obtain, Information acquisition means for acquiring positional and orientation information, which is information regarding the position and orientation of a head-mounted display (HMD), A first virtual image that corresponds to the first real image, represents a virtual object, and has a third resolution, A second virtual image that corresponds to the second real image, represents a part of the first virtual image, and has a fourth resolution higher than the third resolution. A virtual image acquisition means for acquiring an image, Based on the difference between the position and attitude information at the first timing and the position and attitude information at the second timing which is later than the first timing, The first virtual image at the first timing is corrected to correspond to the first real image at the third timing, which is later than the first timing. The second virtual image at the first timing is corrected to correspond to the second real image at the third timing. Correction means, A first composite image is generated by combining the first real image at the third timing with the first virtual image corrected by the correction means. A second composite image is generated by combining the second real image at the third timing with the second virtual image corrected by the correction means. The second composite image is composited onto the first composite image. By doing so, a synthesis means for generating a display image to be shown on the HMD and An information processing device characterized by having the following features. (Configuration 2) The size of the second virtual image before correction by the correction means is larger than the size of the second real image. The information processing device according to configuration 1, characterized by the above. (Composition 3) The correction of the second virtual image by the correction means includes a process of extracting from the second virtual image a portion of the second virtual image that corresponds to the region of the second real image at the third timing. An information processing device according to configuration 1 or 2, characterized by the above. (Composition 4) If the region corresponding to the region of the second real image at the third timing includes a region that is not the second virtual image before correction by the correction means, the synthesis means adopts the first synthesized image as the display image. The information processing apparatus according to configuration 3, characterized by the features described herein. (Composition 5) The aforementioned real-world image acquisition means further acquires depth information of the real-world space, The virtual image acquisition means further acquires depth information of the virtual object, The synthesis means generates the first composite image and the second composite image based on the depth information of the real space and the depth information of the virtual object. An information processing device according to any one of configurations 1 to 4, characterized by the above. (Composition 6) If the difference is greater than or equal to a threshold, the synthesis means adopts the first synthesized image as the display image. An information processing device according to any one of configurations 1 to 5, characterized by the above. (Composition 7) If the difference corresponds to a change in attitude with an angular velocity of 60 [deg / sec] or more, the synthesis means adopts the first synthesized image as the display image. An information processing device according to any one of configurations 1 to 6, characterized by the above. (Composition 8) The virtual image acquisition means acquires the second virtual image before the first virtual image. An information processing device according to any one of configurations 1 to 7, characterized by the above. (Composition 9) The synthesis means generates the second composite image before the first composite image. The information processing apparatus according to configuration 8, characterized by the above. (Composition 10) Eye-tracking means for acquiring eye-tracking information of a user wearing the aforementioned HMD It further possesses, The second reality image represents the portion of the real space that the user's gaze is directed towards, An information processing device according to any one of configurations 1 to 9, characterized by the above. (Composition 11) The second timing and the third timing are equal. An information processing device according to any one of configurations 1 to 10, characterized by the above. (Composition 12) The first resolution and the third resolution are equal. The earlier second resolution and the earlier fourth resolution are equal. An information processing device according to any one of configurations 1 to 11, characterized by the above. (method) A step of obtaining a first reality image that represents real space and has a first resolution, A step of obtaining a second reality image that represents a part of the first reality image and has a second resolution higher than the first resolution, The steps include acquiring positional and orientation information, which is information regarding the position and orientation of a head-mounted display (HMD), The steps include obtaining a first virtual image that corresponds to the first real image, represents a virtual object, and has a third resolution, A step of obtaining a second virtual image that corresponds to the second real image, represents a part of the first virtual image, and has a fourth resolution higher than the third resolution. A step of correcting the first virtual image at the first timing to correspond to the first real image at the third timing which is later than the first timing, based on the difference between the position and orientation information at the first timing and the position and orientation information at the second timing which is later than the first timing, The steps include correcting the second virtual image at the first timing to correspond to the second real image at the third timing based on the difference, The steps include generating a first composite image by combining the first real image at the third timing with the first corrected virtual image, The steps include generating a second composite image by combining the corrected second virtual image with the second real image at the third timing, The steps include generating a display image to be shown on the HMD by combining the second composite image with the first composite image, and A control method for an information processing device, characterized by having the following features. (program) A program for causing a computer to function as one of the means of the information processing device described in any of configurations 1 to 12. [Explanation of symbols]
[0070] 101: Head-mounted display (HMD) 201: Imaging unit 202: Position and orientation sensor 204: HMD-I / F 205: Focus CG Correction Section 206: Overall CG Correction Section 207: Focused Image Synthesis Unit 208: Overall Image Synthesis Unit 209: Display Image Synthesis Unit
Claims
1. A first reality image that represents real space and has a first resolution, A second reality image that represents a part of the first reality image and has a second resolution higher than the first resolution, A means for acquiring real images to obtain, Information acquisition means for acquiring positional and orientation information, which is information regarding the position and orientation of a head-mounted display (HMD), A first virtual image that corresponds to the first real image, represents a virtual object, and has a third resolution, A second virtual image that corresponds to the second real image, represents a part of the first virtual image, and has a fourth resolution higher than the third resolution. A virtual image acquisition means for acquiring an image, Based on the difference between the position and attitude information at the first timing and the position and attitude information at the second timing which is later than the first timing, The first virtual image at the first timing is corrected to correspond to the first real image at the third timing, which is later than the first timing. The second virtual image at the first timing is corrected to correspond to the second real image at the third timing. Correction means, A first composite image is generated by combining the first real image at the third timing with the first virtual image corrected by the correction means. A second composite image is generated by combining the second real image at the third timing with the second virtual image corrected by the correction means. The second composite image is composited onto the first composite image. By doing so, a synthesis means for generating a display image to be displayed on the HMD and An information processing device characterized by having the following features.
2. The size of the second virtual image before correction by the correction means is larger than the size of the second real image. The information processing apparatus according to feature 1.
3. The correction of the second virtual image by the correction means includes a process of extracting from the second virtual image a portion of the second virtual image that corresponds to the region of the second real image at the third timing. The information processing apparatus according to feature 1.
4. If the region corresponding to the region of the second real image at the third timing includes a region that is not the second virtual image before correction by the correction means, the synthesis means adopts the first synthesized image as the display image. The information processing apparatus according to claim 3.
5. The aforementioned real-world image acquisition means further acquires depth information of the real-world space, The virtual image acquisition means further acquires depth information of the virtual object, The synthesis means generates the first composite image and the second composite image based on the depth information of the real space and the depth information of the virtual object. The information processing apparatus according to feature 1.
6. If the difference is greater than or equal to a threshold, the synthesis means adopts the first synthesized image as the display image. The information processing apparatus according to feature 1.
7. If the difference corresponds to a change in attitude with an angular velocity of 60 [deg / sec] or more, the synthesis means adopts the first synthesized image as the display image. The information processing apparatus according to feature 1.
8. The virtual image acquisition means acquires the second virtual image before the first virtual image. The information processing apparatus according to feature 1.
9. The synthesis means generates the second composite image before the first composite image. The information processing apparatus according to feature 8.
10. Eye-gaze acquisition means for acquiring eye-gaze information of a user wearing the aforementioned HMD It further possesses, The second reality image represents the portion of the real space that the user's gaze is directed towards. The information processing apparatus according to feature 1.
11. The second timing and the third timing are equal. The information processing apparatus according to feature 1.
12. The first resolution and the third resolution are equal The aforementioned second resolution and the aforementioned fourth resolution are equal. The information processing apparatus according to feature 1.
13. A step of obtaining a first reality image that represents real space and has a first resolution, A step of obtaining a second reality image that represents a part of the first reality image and has a second resolution higher than the first resolution, The steps include acquiring positional and orientation information, which is information regarding the position and orientation of a head-mounted display (HMD), The steps include obtaining a first virtual image that corresponds to the first real image, represents a virtual object, and has a third resolution, A step of obtaining a second virtual image that corresponds to the second real image, represents a part of the first virtual image, and has a fourth resolution higher than the third resolution. A step of correcting the first virtual image at the first timing to correspond to the first real image at the third timing which is later than the first timing, based on the difference between the position and orientation information at the first timing and the position and orientation information at the second timing which is later than the first timing, The steps include correcting the second virtual image at the first timing to correspond to the second real image at the third timing based on the difference, The steps include generating a first composite image by combining the first real image at the third timing with the first corrected virtual image, The steps include generating a second composite image by combining the corrected second virtual image with the second real image at the third timing, The steps include generating a display image to be shown on the HMD by combining the first composite image with the second composite image, and A control method for an information processing device, characterized by having the following features.
14. A program for causing a computer to function as one of the means of an information processing apparatus described in any one of claims 1 to 12.
Citation Information
Patent Citations
Rendering device, head-mounted display, image transmission method, and image correction method
JP2019028368A