Information processing systems, methods for controlling information processing systems, programs

The system addresses the challenge of distinguishing between real-world and computer-generated graphics in MR by generating and synthesizing images with time-stamped data management, enhancing data processing accuracy in mixed reality systems.

JP2026055409APending Publication Date: 2026-03-31CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing MR technologies struggle to accurately determine whether a user is observing the real-world or computer-generated graphics, leading to improper data management and processing of gaze detection results.

Method used

An information processing system that generates a virtual image based on a reference position, synthesizes it with a real image, and manages input information with time stamps to accurately associate user inputs with the correct image type, enabling precise data management.

Benefits of technology

Enables appropriate processing of input information by accurately determining whether the user is observing real-world or computer-generated graphics, allowing for improved data management and processing in mixed reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026055409000001_ABST
    Figure 2026055409000001_ABST
Patent Text Reader

Abstract

The system manages data to enable more appropriate processing based on input information related to the displayed image. [Solution] The information processing system includes: generation means for generating a virtual image by drawing a virtual object based on a reference position determined at a first time; synthesis means for generating a display image by combining the virtual image and a first image; display control means for displaying the display image on a display means at a second time; input means for acquiring user input information to the display means at the second time; and control means for storing the input information, the display image, the reference position information, first information which is time information relating to the reference position, and second information which is time information relating to the first image in a storage means in association with each other.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system, a control method for an information processing system, and a program.

Background Art

[0002] In recent years, as a technology for seamlessly integrating the real space and the virtual space in real time, a technology with a composite sense of reality, so-called MR (Mixed Reality) technology, is known. As one of the MR technologies, there is a technology of displaying an image in which CG (Computer Graphics) is superimposed on a captured image of the real space using a video see-through type HMD (Head Mounted Display).

[0003] At this time, there is a technology of detecting the user's line-of-sight direction using a camera that captures the user's pupils and specifying the user's observation position in the display image. Also, there is a known function such as a line-of-sight log that saves the movement of the position where the user is looking in the display image by saving the display image of the HMD and the observation position. In order to realize the line-of-sight log function, it is necessary to accurately associate the subject in the display image with the position of line-of-sight detection. Patent Document 1 describes a method of specifying the display image used for line-of-sight detection and the line-of-sight detection result based on the set drive mode among the drive modes of imaging and display.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the gaze logging function, the objects the user was observing in the MR image, which is a fusion of real-world images and computer graphics (CG), are determined by scoring and weighting. In this case, it is necessary to accurately identify whether the user was observing the real-world domain or the CG domain, and at what point in time the real-world image and CG image were being observed. However, there is usually a difference between the time when the real-world image used in the MR image is generated and the time when the CG is generated.

[0006] While Patent Document 1 can identify the displayed image and the gaze detection result, it cannot identify the subject information displayed in the image. In other words, it was difficult to determine whether the user was observing the real world or computer graphics. As a result, the data was not managed in a way that allowed for proper processing based on input information such as gaze detection results.

[0007] The present invention aims to provide a data management technology that enables more appropriate processing in response to input information on displayed images. [Means for solving the problem]

[0008] One aspect of the present invention is, A generation means for generating a virtual image by drawing a virtual object based on a reference position determined at a first time point, A synthesis means for generating a display image by combining the aforementioned virtual image and a first image, A display control means that displays the display image on the display means at a second time, An input means for acquiring user input information at the second time point for the display means, The input information, the display image, the reference position information, and the time information relating to the reference position A control means that stores in a storage means a first piece of information and a second piece of information which is time information relating to the first image, in correspondence with each other. This is an information processing system characterized by having [a certain feature].

[0009] One aspect of the present invention is, A generation step of generating a virtual image by drawing a virtual object based on a reference position determined at a first time point, A synthesis step of generating a display image by combining the virtual image and the first image, A display control step of displaying the display image on the display means at a second time; An input step of acquiring user input information at the second time point for the display means, A control step of storing the input information, the display image, the reference position information, the first information which is time information relating to the reference position, and the second information which is time information relating to the first image in a storage means in association with each other, This is a control method for an information processing system characterized by having [a certain feature]. [Effects of the Invention]

[0010] According to the present invention, it becomes possible to manage data in a way that allows for more appropriate processing in response to input information on the displayed image. [Brief explanation of the drawing]

[0011] [Figure 1] This is a diagram showing the configuration of the information processing system according to Embodiment 1. [Figure 2] This is a diagram showing the configuration of the HMD and generation device according to Embodiment 1. [Figure 3] This is a flowchart showing the process according to Embodiment 1. [Figure 4] This is a time chart showing the processing at each time point according to Embodiment 1. [Figure 5] This is a flowchart showing the process according to Embodiment 2. [Figure 6] This is a diagram illustrating the reprojection process according to Embodiment 3. [Figure 7] This flowchart shows the processing of the log storage unit according to Embodiment 3. [Modes for carrying out the invention]

[0012] Hereinafter, embodiments of the present invention will be described in detail based on the accompanying drawings.

[0013] <Embodiment 1> FIG. 1 is a diagram showing the system configuration according to Embodiment 1. In FIG. 1, the information processing system 1 includes a head-mounted image display device (Head Mounted Display, hereinafter referred to as “HMD”) 100 and a generation device 200.

[0014] The HMD 100 is worn on the user's head. The HMD 100 allows the user to experience Mixed Reality by displaying an image on the display.

[0015] The HMD 100 captures the real space with an outward-facing camera to obtain a real image. The HMD 100 displays an image synthesized from the virtual image generated by the generation device 200 and the real image on the display as an MR image.

[0016] The generation device 200 generates a virtual image, which is an image of the virtual space that the user experiences using the HMD 100. Specifically, the generation device 200 performs rendering of the virtual image by calculating the drawing position of the virtual object based on the real image and the tracking data. In order to obtain an appropriate position of the virtual object in the real space, a method of detecting the position (marker position) of the marker 500 from the real image is used. As a result, a virtual image in which a virtual object corresponding to “coordinate Z1” uniquely determined based on the marker 500 placed on the floor is arranged is generated, and an MR image synthesized from the virtual image and the real image is generated.

[0017] The generation device 200 is an information processing device such as a PC (Personal Computer), for example. The generation device 200 may be connected to a server via a network. Further, the generation device 200 may be a portable device that can be carried together with the HMD 100.

[0018] Interface 300 connects the HMD 100 and the generation device 200 via a wired cable. This allows the HMD 100 and the generation device 200 to exchange data with each other. The data exchanged is not limited to image data, but also includes sensor data (data acquired by an acceleration sensor or angular velocity sensor), audio data, and control data for controlling the HMD 100. Furthermore, while Embodiment 1 describes Interface 300 as an interface that achieves connection via a wired connection, it may also be an interface that achieves connection via a wireless connection.

[0019] Figure 2 is a configuration diagram of the HMD 100 and generation device 200 shown in Figure 1. As shown in Figure 2, the HMD 100 includes a reality imaging unit 101, a display unit 102, a gaze detection unit 103, a posture detection unit 104, a time detection unit 105, a synthesis unit 106, a control unit 107, an interface unit (IF unit) 108, and a memory 109. These components are connected via a system bus 110.

[0020] The reality imaging unit 101 acquires a reality image by imaging the real space. The reality imaging unit 101 sends the reality image to the generation device 200 and the synthesis unit 106.

[0021] The display unit 102 displays the MR image (mixed reality image) generated by the synthesis unit 106. This allows the user to view the MR image.

[0022] The gaze detection unit 103 determines the movement of the user's pupils (eyes) based on images of the user's pupils captured by the camera. Based on the pupil movements, the gaze detection unit 103 identifies the direction of the user's gaze. The identified gaze direction information is taken in as the user's gaze point information (viewpoint information) when viewing the display unit 102 and sent to the generation device 200 via the interface unit 108. The method of gaze detection is generally known, so a detailed explanation is omitted. Alternatively, information on the position the user is looking at on the display surface of the display unit 102 may be detected as gaze point information by the gaze detection unit 103.

[0023] The attitude detection unit 104 acquires the attitude of the HMD 100. The acquired attitude information is taken in as position and attitude information indicating the movement of the HMD 100 and sent to the generation device 200 via the interface unit 108.

[0024] The time detection unit 105 manages the acquisition time of information for each configuration. Embodiment 1 describes a method for generating timestamp information within the HMD 100 for time management and identifying the acquisition time of each function. The acquired time information is sent to the generation device 200 via the interface unit 108 along with the data acquired by each function.

[0025] The synthesis unit 106 synthesizes the real image acquired by the real image acquisition unit 101 with the virtual image sent from the generation device 200 to generate an MR image. The generated MR image is sent to the display unit 102 and displayed on the display unit 102.

[0026] The control unit 107 is a processing unit such as a CPU (Central Processing Unit). The control unit 107 manages the operation and sequence of each function.

[0027] The HMD 100 may also consist of an information processing device (display control device) having a reality imaging unit 101, a display unit 102, and other components. In this case, the control unit 107 included in the information processing device operates as a display control unit that controls the display of the display unit 102, and also as an imaging control unit that controls the imaging of the reality imaging unit 101. Furthermore, the HMD 100 may have all or part of the components of the generation device 200.

[0028] As shown in Figure 2, the generation device 200 includes a position calculation unit 201, a rendering unit 202, a content DB 203, a reprojection unit 204, a log storage unit 205, a control unit 207, an interface unit (IF unit) 208, and a memory 209. These components are connected to each other via a system bus 210.

[0029] The position calculation unit 201 recognizes the marker 500 in real space from the real image acquired by the real image acquisition unit 101. The position calculation unit 201 then detects the "coordinate Z1" (position) of the marker 500 in a camera coordinate system based on the position and orientation of the HMD 100. The position calculation unit 201 also detects the position and orientation of the HMD 100 in real space in order to more accurately position the virtual image in the image viewed by the user of the HMD 100. The position calculation unit 201 may also detect the relative position and orientation of the HMD 100 by detecting an optical sensor attached to the HMD 100 with a tracking system placed in real space. The position calculation unit 201 is not particularly limited to inside-out or outside-in tracking methods, and any configuration that can track the HMD 100 is acceptable.

[0030] One method for determining "coordinate Z1" is to detect a marker 500 that has been pre-placed in real space from a real-world image, thereby determining "coordinate Z1" in the camera coordinate system. Specifically, the position calculation unit 201 determines "coordinate Z1" in the camera coordinate system based on the coordinates of the marker 500 in real space. Since the marker 500 in real space does not move, even if the user of the HMD 100 moves around, the coordinates of the marker 500 in real space remain fixed. Therefore, the position calculation unit 201 can convert the position of the marker 500 in "real space" to the camera coordinate system of the HMD 100 based on the position of the marker 500 in the "real-world image". In other words, the position calculation unit 201 can determine "coordinate Z1" in the camera coordinate system from the relative positional relationship between the HMD 100 and the marker 500 in real space.

[0031] Furthermore, the position calculation unit 201 is not limited to the marker 500; it is sufficient to identify a reference position (reference coordinate) for the placement of the virtual image. For this reason, the position calculation unit 201 may identify the coordinates in the camera coordinate system of the feature points of an object fixed in real space.

[0032] The rendering unit 202 is a generation unit that generates virtual images. First, the rendering unit 202 reads content from the content DB 203 according to the position of "coordinate Z1" detected by the position calculation unit 201. The rendering unit 202 renders a virtual object (CG) based on the read content. There are many types of algorithms for rendering virtual objects, but in Embodiment 1, a polygon-based calculation method (hereinafter referred to as "polygon rendering"), which is widely used in a field called real-time rendering, is used. Polygon rendering is a widely known method and is also a method executed by a rendering engine, so a detailed explanation is omitted. The processing time by the rendering engine depends on the performance of the GPU (not shown) installed in the generation device 200 and the content capacity stored in the content DB 203. Therefore, depending on the GPU performance and content capacity, a long processing time may be required. If the processing time is long, the view observed on the HMD 100 will be Because the images are displayed with a delay, it can cause discomfort for users who are moving their heads while experiencing mixed reality.

[0033] The reprojection unit 204 performs reprojection processing of virtual objects to counteract the processing delays caused by the rendering unit 202.

[0034] The log storage unit 205 incorporates a storage medium (storage unit). The log storage unit 205 is a storage control unit that associates the gaze detection result with the information displayed on the display unit 102 at the time of gaze detection and stores it on the storage medium.

[0035] Memory 109 and memory 209 temporarily store data acquired by each functional unit. Memory 109 and memory 209 may be non-volatile or volatile storage units.

[0036] Interface units 108 and 208 communicate data between the HMD 100 and the generation device 200 via interface 300. Communication may be implemented by either a wired or wireless method.

[0037] The process of Embodiment 1 will be explained using the flowcharts in Figures 3A to 3F. Although the flowcharts in Figures 3A to 3F operate independently, parts where the processes are related are indicated by dotted arrows. Furthermore, while each flowchart is assumed to show the processing of one frame of image, in reality, images are input continuously as moving images. Therefore, the processing in these flowcharts is performed continuously, and the process from start to finish is repeated each time an image is input.

[0038] Figure 3A is a flowchart showing the processing of the real-world imaging unit 101.

[0039] In step S301, the reality imaging unit 101 captures the real space at a predetermined frame rate and acquires a real image Ra.

[0040] In step S302, the real image imaging unit 101 obtains time information Ta, which indicates the acquisition time (imaging time) of the real image Ra, from the time detection unit 105. Time information Ta is timestamp information issued inside the HMD 100. In the HMD 100, time is managed by timestamp information.

[0041] In step S303, the real image imaging unit 101 sends the real image Ra and time information Ta to the synthesis unit 106.

[0042] In step S304, the real image unit 101 sends the real image Ra and time information Ta to the position calculation unit 201.

[0043] Figure 3B is a flowchart showing the processing of the position calculation unit 201.

[0044] In step S311, the position calculation unit 201 acquires the real image Ra acquired at time Ta' indicated by the time information Ta from the real image acquisition unit 101, along with the time information Ta.

[0045] In step S312, the position calculation unit 201 analyzes the real image Ra to detect (determine) the position of the marker placed in real space.

[0046] In step S313, the position calculation unit 201 determines the position of the marker in the camera coordinate system. This converts the position to a posable position. This allows the position calculation unit 201 to detect the relative positional relationship between the HMD 100 and the marker.

[0047] In step S314, the position calculation unit 201 sends marker position information Ca, which indicates the marker position converted to the camera coordinate system, and time information Ta to the rendering unit 202.

[0048] In step S315, the position calculation unit 201 sends (stores) the marker position information Ca and time information Ta to the memory 209.

[0049] Figure 3C is a flowchart showing the processing of the rendering unit 202.

[0050] In step S321, the rendering unit 202 obtains marker position information Ca and time information Ta from the position calculation unit 201.

[0051] In step S322, the rendering unit 202 reads the rendering settings for the virtual object. The rendering settings include, for example, information indicating the orientation and size of the virtual object to be rendered. The rendering settings depend on the application performing the rendering. Based on the rendering settings, the rendering unit 202 determines the orientation of the virtual object relative to the marker during rendering.

[0052] In step S323, the rendering unit 202 reads content information from the content DB 203.

[0053] In step S324, the rendering unit 202 renders a virtual object based on content information, drawing settings, and marker position information Ca.

[0054] In step S325, the rendering unit 202 generates an image representing the rendered virtual object as a virtual image Va.

[0055] In step S326, the rendering unit 202 sends the virtual image Va and time information Ta (information indicating the time of acquisition of the real image for obtaining marker position information Ca) to the synthesis unit 106.

[0056] Figure 3D is a flowchart showing the processing in the synthesis unit 106.

[0057] In step S331, the synthesis unit 106 obtains the latest real image Rb and time information Tb indicating the acquisition time of the real image Rb from the real imaging unit 101. At this time, time information Tb and time information Ta show different times. This is because processing takes time in the position calculation unit 201 and the rendering unit 202. In other words, there is a difference (time lag) in the reference time between the real image Rb and the virtual image Va. Also, the difference between the time Tb' indicated by time information Tb and the time Ta' indicated by time information Ta changes depending on the situation.

[0058] In step S332, the synthesis unit 106 obtains the virtual image Va and time information Ta from the rendering unit 202.

[0059] In step S333, the synthesis unit 106 synthesizes the real image Rb and the virtual image Va to generate the MR image Mb.

[0060] In step S334, the synthesis unit 106 sends the MR image Mb to the display unit 102. As a result, the display unit 102 displays the MR image Mb.

[0061] In step S335, the synthesis unit 106 sends the time information Tb and time information Ta to the gaze detection unit 103.

[0062] In step S336, the synthesis unit 106 transmits the MR image Mb and time information Tb displayed on the display unit 102 to the memory 209. The MR image Mb and time information Tb are temporarily stored in the memory 209.

[0063] Figure 3E is a flowchart showing the processing of the gaze detection unit 103.

[0064] In step S341, the gaze detection unit 103 obtains time information Tb and time information Ta from the synthesis unit 106, which indicate the acquisition time of the real image Rb related to the MR image Mb.

[0065] In step S342, the gaze detection unit 103 acquires a gaze image that captures the user's eye (pupil) at the time the MR image Mb is displayed on the display unit 102.

[0066] In step S343, the gaze detection unit 103 generates a gaze detection result Gc as user input information based on the acquired gaze image. The gaze detection result Gc is two-dimensional coordinate information and represents the position (viewpoint position) that the user is looking at on the display surface of the display unit 102.

[0067] In step S344, the gaze detection unit 103 sends the gaze detection result Gc and time information Tb and Ta to the log storage unit 205.

[0068] Figure 3F is a flowchart showing the processing of the log storage unit 205.

[0069] In step S351, the log storage unit 205 obtains the gaze detection result Gc and time information Tb and Ta from the gaze detection unit 103. By referring to the time information Ta, the log storage unit 205 can identify the acquisition time of the real image used to generate the virtual image Va included in the MR image Mb. By referring to the time information Tb, the log storage unit 205 can identify the acquisition time of the real image included in the MR image Mb.

[0070] In step S352, the log storage unit 205 accesses the memory 209 and obtains marker position information Ca corresponding to the time information Ta.

[0071] In step S353, the log storage unit 205 accesses the memory 209 to obtain the MR image Mb corresponding to the time information Tb.

[0072] In step S354, the log storage unit 205 associates the "gaze detection result Gc", "marker position information Ca", "time information Ta", "MR image Mb", and "time information Tb" with each other and stores them as log information on the storage medium.

[0073] The processing at each time point in Embodiment 1 will be explained using the time chart in Figure 4. Figure 4 describes the processing at each time point for the time detection unit 105, reality imaging unit 101, position calculation unit 201, rendering unit 202, synthesis unit 106, display unit 102, gaze detection unit 103, memory 209, and log storage unit 205.

[0074] At time point 4210, the reality imaging unit 101 captures the real image of frame 1 by imaging the real space.

[0075] During period 4310, the position calculation unit 201 calculates the marker based on the captured real image. This performs calculations such as determining the coordinates (position).

[0076] During period 4410, the rendering unit 202 generates a virtual image by rendering a virtual object based on the coordinates (positions) of the markers.

[0077] During period 4510, the synthesis unit 106 generates an MR image by combining a virtual image with the most recent real image captured at time 4230.

[0078] At time point 4610, the display unit 102 begins to display the MR image generated during period 4510.

[0079] At time 4710, the gaze detection unit 103 acquires a gaze image, which is an image of the user's eye (pupil) looking at the display unit 102.

[0080] At time 4810, the log storage unit 205 obtains (confirms) a gaze detection result indicating the user's viewing position as a result of performing gaze detection calculations on the detected gaze image. At this time, the display unit 102 displays the MR image that began to be displayed at time 4620. Each calculation process has a time lag, and the amount of this delay also fluctuates. Therefore, even by referring only to the gaze detection result confirmed at time 4810, it is not possible to determine at what point in time the user was viewing the displayed MR image, or at what point in time the real-world images captured in that MR image are based.

[0081] Therefore, the log storage unit 205 manages the time information of each processing point by linking it, as explained using the flowchart in Figure 3. This allows the log storage unit 205 to identify the time point for which each piece of information corresponds. As shown in step S351 in Figure 3, at time 4810, the log storage unit 205 has already obtained that "the MR image is information from time T3, and the position calculation result is information from time T1." Therefore, the log storage unit 205 accesses the memory 209 and, at time 4811, obtains the MR image linked to time T3, and at time 4812, obtains the position calculation result linked to time T1. Thus, the log storage unit 205 can appropriately link the temporal consistency between the gaze detection result and each acquired piece of information and save it as log information.

[0082] Through this process, the capture time of the real-world image used to generate the virtual image, and the capture time of the real-world image used to synthesize the composite image, can be determined from the log information. Based on these capture times, the MR image, marker positions, and gaze detection results, it is possible to accurately determine whether the user is looking at the virtual image or the real-world image in the MR image, and to identify the object the user is looking at in the MR image.

[0083] According to Embodiment 1, time information is issued as a timestamp value and managed in association with each acquired piece of information. This allows for highly accurate association of time information with each process, even when there is a time lag between the gaze detection result and each process. Therefore, since the object the user is observing can be identified based on the log information, it becomes possible to manage the data in a way that allows for more appropriate execution of processing according to the input information for the MR image (displayed image).

[0084] In Embodiment 1, there is one reality imaging unit that captures images of the real space, and the images used for generating MR images and the images used for position calculations are the same. However, the HMD100 may have multiple reality imaging units. Also, an optical sensor may be used for position calculations instead of a reality imaging unit.

[0085] Furthermore, instead of gaze detection results corresponding to the user's gaze, the results of user operations on the MR image may be used as input information. For example, instead of gaze detection results, the results of touch positions on a display unit having a touch panel may be used. Also, instead of gaze detection results, the results of gesture detection that point to a specific location on the MR image may be used.

[0086] <Embodiment 2> In Embodiment 1, the time detection unit 105 generates timestamp information for time management to identify the acquisition time of each data. In Embodiment 2, instead of generating timestamp information, the time detection unit 105 identifies the acquisition time of each data by predicting the time required for processing.

[0087] The process of Embodiment 2 will be explained using the flowchart in Figure 5. The basic flow is the same as the flowchart in Figure 3, but the differences will be explained by highlighting the relevant parts.

[0088] Figure 5A is a flowchart showing the processing of the time detection unit 105.

[0089] In step S501, the time detection unit 105 obtains processing load information L from the generation device 200. Processing time is unstable in the position calculation unit 201 and the rendering unit 202, etc. For example, the amount of content data (data used for rendering virtual objects) stored in the content DB 203 can be considered a factor that causes fluctuations in processing time. Therefore, in Embodiment 2, the processing load information L is the amount of content data. Alternatively, instead of the amount of content data, any parameter that causes fluctuations in the processing time of the rendering unit 202, etc., can be used for the processing load information L.

[0090] In step S502, the time detection unit 105 determines the predicted delay time D according to the processing load information L. The predicted time D is the time from the time the position calculation result of the real image for generating the virtual image is calculated to the time when the MR image Mb, into which the virtual image has been synthesized, is displayed. In Figure 4, the predicted time D is the time from the completion of the calculation (determination) of the marker position by the position calculation unit 201 (end of period 4310 = start of period 4410) to the display of the MR image Mb by the display unit 102 (time 4610 = time 4710). The predicted time D is a time measured (estimated) in advance according to the fluctuating value of the processing load information L, and is a value that is uniquely determined for the processing load information L.

[0091] In step S503, the time detection unit 105 obtains delay time E information from the gaze detection unit 103. The delay time E is the time from the time the gaze image is acquired (display time of MR image Mb) to the time when the gaze detection result is generated (acquired) from that gaze image. In Figure 4, the delay time E is the time from the acquisition of the gaze image (time 4710 = time 4610) to the end of the gaze detection calculation by the gaze detection unit 103 (starting point of the dotted arrow pointing to time 4810). The delay time E is a pre-measured processing time and is a constant value.

[0092] In step S504, the time detection unit 105 transmits information on the predicted delay time D and the delay time E to the log storage unit 205.

[0093] Figure 5B is a flowchart showing the processing of the log storage unit 205.

[0094] In step S511, the log storage unit 205 obtains information on the predicted time D and the delay time E from the time detection unit 105.

[0095] In step S512, the log storage unit 205 stores the gaze detection result Gc in the gaze detection unit 10 Obtain it from 3.

[0096] In step S513, the log storage unit 205 retrieves the MR image Mb corresponding to the delay time E from the memory 209. At this time, the log storage unit 205 refers to the data of the MR image Mb that was generated a delay time E in the past from the time the gaze detection result Gc was acquired (in Figure 4, the data generated during period 4510). Therefore, the log storage unit 205 can identify the acquisition time of the MR image Mb based on the delay time E. In addition, since the most recently captured image at the time of synthesis is used for the MR image Mb, the log storage unit 205 can identify the acquisition time of the real image included in the MR image Mb displayed on the display unit 102 based on the delay time E.

[0097] In step S514, the log storage unit 205 obtains marker position information Ca corresponding to the predicted time D from the memory 209. At this time, the log storage unit 205 refers to marker position information Ca (data obtained in period 4310 in Figure 4) that was obtained a time in the past of the predicted time D from the time when the MR image Mb identified in step S513 was stored in the memory 209. Therefore, the log storage unit 205 can determine the acquisition time of the marker position information Ca based on the sum of the delay time E and the predicted time D. Thus, the log storage unit 205 can determine the acquisition time of the real image used to generate the virtual image Va, which is included in the MR image Mb displayed on the display unit 102, because the most recent captured image at the time of acquisition of the marker position information Ca is used to generate the virtual image Va.

[0098] In step S515, the log storage unit 205 associates the "gaze detection result Gc", "marker position information Ca", and "predicted time D" with the "MR image Mb" and "delay time E" information and stores them as log information in the storage medium (storage unit).

[0099] In Embodiment 2, the time detection unit 105 predicts (estimates) the predicted time D and the delay time E. The log storage unit 205 then stores the information on the predicted time D, the information on the delay time E, the gaze detection information, and the MR image as log information. Therefore, since the object that the user is observing can be identified based on the log information, it becomes possible to manage the data in a way that allows for more appropriate processing according to the input information for the MR image (displayed image).

[0100] The predicted time D is used to identify the acquisition time of the real image used to generate the virtual image (to determine the marker position), etc. Therefore, by combining Embodiments 1 and 2, time information Ta may be stored by the log storage unit 205 instead of the predicted time D information. Also, the time information Ta and predicted time D are time information relating to the marker position (reference position), and instead of these, for example, information on the time of determination of the marker position (information at the end of period 4310 in Figure 4) may be stored. The time information Tb and delay time E are time information relating to the real image used for MR image synthesis, and instead of these, information on the acquisition time of the synthesized image using the real image (information at the end of period 4510 in Figure 4) may be stored.

[0101] <Embodiment 3> Embodiment 3 describes the operation when the reprojection unit 204 performs the reprojection process of the virtual image.

[0102] The reprojection unit 204 converts the image to cancel out the movement of the HMD 100 during processing that takes a long time, such as when processing by the position calculation unit 201 and the rendering unit 202. In this way, the reprojection unit 204 reduces the display delay of the image that becomes apparent due to processing time.

[0103] Figures 6A to 6E are diagrams illustrating Embodiment 3. Figure 6A shows the time information Ta The image shown is a real image acquired by the real imaging unit 101 at the indicated time Ta'. Figure 6A shows a real image in which real objects 1-3 and markers are captured. Figure 6C is a virtual image in which the rendering unit 202 places virtual objects based on the detection result of the markers shown in Figure 6A. At this time, since rendering takes time, it is assumed that the generation of the virtual image is completed at a time Tb' (the time Tb' indicated by the time information Tb) which is later than time Ta'.

[0104] Figure 6B is the real image acquired (imaged) by the real imaging unit 101 at time Tb'. As shown in Figure 6B, the HMD 100 moves from time Ta' to time Tb', so the marker position moves to the right in the real image. The reprojection unit 203 determines the amount of marker movement S (= amount of movement of the HMD 100) during the period between these two times, based on the position of the marker detected by the attitude detection unit 104. The reprojection unit 203 performs a reprojection process based on the amount of movement S to change the position of the virtual object. Figure 6D is a virtual image with the virtual object placed after the reprojection process. In Figure 6D, it can be seen that the virtual object has changed by the amount of movement S from the state shown in Figure 6C.

[0105] Figure 6E shows an MR image created by combining the real image shown in Figure 6B with the virtual image shown in Figure 6D. This synthesis corrects the virtual image so that the virtual object moves in the opposite direction to the HMD100's movement, based on the processing time between time Ta' and time Tb'. This makes it difficult for the user to perceive the impact of processing delays on the virtual image.

[0106] Here, we will explain the issues that arise when both the processing of the reprojection unit 204 and the processing of the log storage unit 205 are executed. Position Pc is the position the user is looking at at time Tc' when the MR image shown in Figure 6E is displayed. In embodiments 1 and 2, the virtual image associated with position Pc is the virtual image shown in Figure 6C, but the virtual image shown in Figure 6C is the virtual image before the reprojection process is executed. Therefore, the user does not actually see the object that is projected at position Pc in Figure 6C. Thus, since the virtual object actually moves by a displacement amount S, it is necessary to save the log information taking the displacement amount S into consideration.

[0107] Figure 7 is a flowchart showing the processing of the log storage unit 205 in Embodiment 3.

[0108] In step S701, the same process as in step S351 is performed.

[0109] In step S702, the log storage unit 205 obtains the reprojection parameter Wa of the virtual image at time Ta' indicated by the time information Ta from the reprojection unit 204. The reprojection parameter Wa is a parameter used in the reprojection process and corresponds to the movement vector of the marker (=movement of HMD100) in the time between time Ta' and time Tb'. The reprojection parameter Wa is, for example, a parameter such as a vector indicating the amount and direction of movement of the virtual object in the reprojection process. The reprojection parameter Wa may include not only movement on a 2D image, but also a homography transformation matrix or parameters of movement change on a 3D surface.

[0110] In step S703, the same process as in step S352 in Figure 3 is performed.

[0111] In step S704, the same process as in step S353 in Figure 3 is performed.

[0112] In step S705, the log storage unit 205 associates the "gaze detection result Gc", "marker position information Ca", "time information Ta", "reprojection parameter Wa", "MR image Mb", and "time information Tb" with each other and stores them as log information on the storage medium.

[0113] In Embodiment 3, the log storage unit 205 stores log information including the reprojection parameters used by the reprojection unit 204. This allows for high-precision matching of input information (gaze detection information) with other data, even when virtual objects move in response to the movement of the HMD 100.

[0114] Furthermore, in the above, "If A is greater than or equal to B, proceed to step S1; if A is less than (lower than) B, proceed to step S2" may be rephrased as "If A is greater than (higher than) B, proceed to step S1; if A is less than or equal to B, proceed to step S2." Conversely, "If A is greater than (higher than) B, proceed to step S1; if A is less than or equal to B, proceed to step S2" may be rephrased as "If A is greater than or equal to B, proceed to step S1; if A is less than (lower than) B, proceed to step S2." Therefore, as long as no contradiction arises, "greater than or equal to A" may be rephrased as "greater than (higher; longer; more) than A," and "less than or equal to A" may be rephrased as "less than (lower; shorter; fewer) than A." And "greater than (higher; longer; more) than A" may be rephrased as "greater than or equal to A," and "less than (lower; shorter; fewer) than A" may be rephrased as "less than or equal to A."

[0115] The various controls described above may or may not be performed by a single piece of hardware (e.g., a processor or circuit). Multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) may share the processing to control the entire device.

[0116] Furthermore, the above-mentioned processors are processors in a broad sense, including general-purpose processors and specialized processors. General-purpose processors include, for example, CPUs (Central Processing Units), MPUs (Micro Processing Units), and DSPs (Digital Signal Processors). Specialized processors include, for example, GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices). Programmable logic devices include, for example, FPGAs (Field Programmable Gate Arrays) and CPLDs (Complex Programmable Logic Devices).

[0117] Furthermore, although embodiments of the present invention have been described in detail, the present invention is not limited to these specific embodiments, and various forms that do not depart from the spirit of the invention are also included in the present invention. Moreover, each of the embodiments described above is merely one embodiment of the present invention, and it is possible to combine each embodiment as appropriate.

[0118] <Other Embodiments> The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit that implements one or more functions.

[0119] The above-disclosed embodiments include the following configurations, methods, and programs. (Composition 1) A generation means for generating a virtual image by drawing a virtual object based on a reference position determined at a first time point, A synthesis means for generating a display image by combining the aforementioned virtual image and a first image, A display control means that displays the display image on the display means at a second time, An input means for acquiring user input information at the second time point for the display means, , A control means that stores the input information, the display image, the reference position information, the first information which is time information relating to the reference position, and the second information which is time information relating to the first image in a corresponding manner in a storage means, An information processing system characterized by having the following features. (Configuration 2) The imaging means has an imaging mechanism that captures the real space at a third time after the first time, thereby acquiring the first image. The information processing system according to configuration 1, characterized by the features described above. (Composition 3) The imaging means acquires a second image of the real space captured at a fourth time before the first time, The information processing system has determination means for determining the reference position based on the second image at the first time, The information processing system according to configuration 2, characterized in that... (Composition 4) The synthesis means generates the display image by synthesizing the virtual image and the first image, which are reprojected onto the virtual object based on the movement of the imaging means between the fourth time and the third time. The control means stores the information relating to the reprojection process, the input information, the display image, the reference position information, the first information, and the second information in the storage means in association with each other. The information processing system according to configuration 3, characterized by the features described above. (Composition 5) The first piece of information is information indicating the fourth time. The information processing system according to configuration 3 or 4, characterized by the features described above. (Composition 6) The second piece of information is information indicating the third time. An information processing system according to any one of configurations 2 to 5, characterized by the above. (Composition 7) The input means acquires the input information at a fifth time point, The second piece of information indicates the time between the second time and the fifth time. An information processing system according to any one of configurations 1 to 4, characterized by the above. (Composition 8) The first piece of information indicates the time between the second time and the first time. The information processing system according to configuration 7, characterized by the features described above. (Composition 9) The first information indicates the time between the second time and the first time, estimated based on the amount of data in the content data used to generate the virtual image. The information processing system according to configuration 8, characterized by the above. (Composition 10) The input means acquires information relating to the user's gaze as input information based on an image of the user's eyes. An information processing system according to any one of configurations 1 to 9, characterized by the above. (method) A generation step of generating a virtual image by drawing a virtual object based on a reference position determined at a first time point, A synthesis step of generating a display image by combining the virtual image and the first image, A display control step of displaying the display image on the display means at a second time; An input step of acquiring user input information at the second time point for the display means, A control step of storing the input information, the display image, the reference position information, the first information which is time information relating to the reference position, and the second information which is time information relating to the first image in a storage means in association with each other, A control method for an information processing system, characterized by having the following features. (program) A program for causing a computer to function as one of the means of an information processing system described in any of configurations 1 to 10. [Explanation of Symbols]

[0120] 1: Information processing system, 100: HMD, 200: Generator, 103: Eye-line detection unit, 106: Synthesis unit 107: Control unit, 202: Rendering unit (virtual image generation unit)

Claims

1. A generation means for generating a virtual image by drawing a virtual object based on a reference position determined at a first time point, A synthesis means for generating a display image by combining the virtual image and the first image, A display control means that displays the display image on the display means at a second time, An input means for acquiring user input information at the second time point for the display means, A control means that stores the input information, the display image, the reference position information, the first information which is time information relating to the reference position, and the second information which is time information relating to the first image in a corresponding manner in a storage means, An information processing system characterized by having the following features.

2. The imaging means has an imaging mechanism that acquires the first image by imaging the real space at a third time after the first time, The information processing system according to feature 1.

3. The imaging means acquires a second image of the real space captured at a fourth time before the first time, The information processing system has determination means for determining the reference position based on the second image at the first time, The information processing system according to feature 2.

4. The synthesis means generates the display image by synthesizing the virtual image and the first image, which are reprojected onto the virtual object based on the movement of the imaging means between the fourth time and the third time. The control means stores the information relating to the reprojection process, the input information, the display image, the reference position information, the first information, and the second information in the storage means in association with each other. The information processing system according to feature 3.

5. The first piece of information is information indicating the fourth time. The information processing system according to feature 3.

6. The second piece of information is information indicating the third time. The information processing system according to feature 2.

7. The input means acquires the input information at a fifth time point, The second piece of information indicates the time between the second time and the fifth time. The information processing system according to feature 1.

8. The first piece of information indicates the time between the second time and the first time. The information processing system according to feature 7.

9. The first information indicates the time between the second time and the first time, estimated based on the amount of data in the content data used to generate the virtual image. The information processing system according to feature 8.

10. The input means acquires information relating to the user's gaze as input information based on an image of the user's eyes. The information processing system according to feature 1.

11. A generation step of generating a virtual image by drawing a virtual object based on a reference position determined at a first time point, A synthesis step of generating a display image by combining the virtual image and the first image, A display control step of displaying the display image on the display means at a second time; An input step of acquiring user input information at the second time point for the display means, A control step of storing the input information, the display image, the reference position information, the first information which is time information relating to the reference position, and the second information which is time information relating to the first image in a storage means in association with each other, A control method for an information processing system, characterized by having the following features.

12. A program for causing a computer to function as one of the means of an information processing system according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Electronic apparatus and control method thereof

    JP2021043368A