Data processing apparatus for virtual reality, data processing method, and computer software

JP2023099494A5Pending Publication Date: 2025-12-04SONY COMP ENTERTAINMENT EURO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022205471
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-31
Filing Date
2022-12-22
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing virtual reality head-mounted displays (HMDs) face delays in image processing that lead to inconsistencies between the tracked position of the user's head and the displayed image, causing disorientation and potential motion sickness due to mismatches in the virtual environment.

Method used

A data processing apparatus and method that evaluates tracking data and image frames to detect inconsistencies by comparing image features across frames, calculating viewpoint differences, and generating output data to correct for processing errors, ensuring synchronized image rendering with head movements.

Benefits of technology

The solution effectively reduces delays and inconsistencies in HMD image processing, enhancing user immersion and reducing the risk of motion sickness by accurately tracking head movements and adjusting the displayed image in real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a data processing apparatus and method for evaluating image frames for display generated using tracking data.SOLUTION: A data processing apparatus comprises: receiving circuitry to receive tracking data indicative of at least one of a tracked position and orientation of a head-mountable display; image processing circuitry to generate a sequence of image frames for display on the basis of the tracking data; detection circuitry to detect an image feature in a first image frame and to detect a corresponding image feature in a second image frame; and correlation circuitry to calculate a difference between the image feature in the first image frame and the corresponding image feature in the second image frame, generate difference data indicative of a difference between a viewpoint for the first image frame and a viewpoint for the second image frame on the basis of the difference between the image feature in the first image frame and the corresponding image feature in the second image frame, and generate output data on the basis of a difference data and the tracking data associated with the first and second image frames.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] ,

[0004] , , , , ,

[0006]

[0001] The present disclosure relates to systems and methods for virtual reality. In particular, the present disclosure relates to a data processing apparatus and method for evaluating generated tracking data and display image frames generated using such tracking data.

Background Art

[0002] The purpose of explaining "Background Art" in this specification is to generally show the concept of the present disclosure. The techniques and aspects of the disclosure by the inventors in this Background Art section do not mean, explicitly or implicitly, prior art to the present invention unless it is stated that it is prior art.

[0003] Head-mounted displays (HDMs) are an example of head-mounted devices for use in virtual reality systems. A person wearing an HMD views a virtual environment and is provided with an image or video display device, which can be worn on the head (or as part of a helmet), and small electrical display devices are provided for one or both eyes.

[0004] The initial development of HMDs and virtual reality was probably advanced for military or professional applications. However, HMDs have come to be more strongly supported by casual users such as in computer games and home computer applications.

[0005] The techniques discussed below are applicable to video signals including individual three-dimensional images or successive three-dimensional images. Therefore, when the term "image" is used in this specification, it includes the use of similar techniques in video signals.

[0006] The above description is only a general introduction and is not intended to limit the scope of the following claims. The best understanding of the embodiments and advantages will be obtained by reading the detailed description with reference to the accompanying drawings. [Overview of the project] [Means for solving the problem]

[0007] Various aspects and features of the present invention are defined by the language set forth in the claims and specification. [Brief explanation of the drawing]

[0008] This technology will be explained by illustration only, with reference to the attached drawings. [Figure 1] This is a schematic diagram of an HMD (Head-Mounted Display) worn by a user. [Figure 2] This is a schematic plan view of the HMD. [Figure 3] This is a schematic diagram illustrating the formation of a virtual image using an HMD (Head-Mounted Display). [Figure 4] This is a schematic diagram of another type of display used in HMDs. [Figure 5] This is a schematic diagram of a pair of three-dimensional images. [Figure 6] This is a schematic diagram illustrating the changes in the user's field of view when using an HMD (Head-Mounted Display). [Figure 7A] This is a schematic diagram of an HMD equipped with a motion detection mechanism. [Figure 7B] This is a schematic diagram of an HMD equipped with a motion detection mechanism. [Figure 8] This is a schematic diagram of a position sensor based on optical flow detection. [Figure 9] This is a schematic diagram illustrating the generation of images in response to position or motion detection in the HMD. [Figure 10] This is a schematic diagram illustrating image capture by a camera. [Figure 11] This is a schematic diagram illustrating the delay problem in HMD image display. [Figure 12] This is a schematic diagram of a data processing device. [Figure 13A] This is a schematic timing diagram. [Figure 13B] This is a schematic timing diagram. [Figure 14]This is a schematic diagram of an HMD display unit that displays stereo images. [Figure 15] This is a schematic diagram of another data processing device. [Figure 16] This is a schematic flowchart of the data processing method. [Modes for carrying out the invention]

[0009] In Figure 1, user 10 is wearing an HMD 20 on their head 30. The HMD comprises a frame 40 (formed in this example by a rear strap and an upper strap) and a display portion 50.

[0010] The HMD in Figure 1 completely blocks the user's view of the surrounding environment. The only thing the user can see is the pair of images displayed within the HMD.

[0011] The HMD has associated headphone transducers or earpieces 60 (which fit into the user's left and right ears 70). The earpieces 60 reproduce audio signals provided from an external sound source (which may be the same as the video signal source that gives the display a video signal).

[0012] During operation, the HMD provides a video signal for the display. This may be provided from an external video signal source 80 (e.g., a video game console or a data processing device such as a personal computer). In this case, the signal may be transmitted to the HMD by a wired or wireless connection 82. A suitable example of a wireless connection is a Bluetooth® connection. An audio signal for the earpiece 60 may be transmitted by the same connection. Similarly, any control signals sent from the HMD to the video (audio) signal source may be transmitted by the same connection.

[0013] Thus, the configuration of FIG. 1 gives an example of a head-mounted display system comprising a frame mounted on the viewer's head and a display element mounted with respect to the line-of-sight display position. The frame defines one or two line-of-sight display positions. The line-of-sight display positions are arranged in front of the viewer's eyes during use. The display element provides a virtual image of a video display signal from a video signal source towards the viewer's eyes.

[0014] FIG. 1 only shows an example of an HMD, and other forms are also possible. For example, the HMD may use a frame similar to conventional glasses. In that case, substantially horizontal legs extend rearward from the display to the top of the user's ears and bend downward behind the ears. In another (non-immersive) example, the user's view of the external environment may not actually be completely blocked. That is, the displayed image may be arranged to be superimposed on the external environment (as seen from the user's viewpoint). An example of such a configuration is shown in FIG. 4.

[0015] In the example of FIG. 1, separate displays are provided for each of the user's left and right eyes. FIG. 2 is a schematic plan view showing how this is realized. FIG. 2 shows the position 100 of the user's eyes and the relative position 110 of the user's nose. The display portion 50 generally comprises an external shield 120 for blocking ambient light from the user's eyes and an internal shield 130 for preventing the display seen by one eye from being seen by the other eye. With respect to the user's face, the external shield 120 and the internal shield 130 form two compartments 140 for each eye. Within each compartment, a display element 150 and one or more optical elements 160 are provided. FIG. 3 shows the optical path formed by the display element and the optical elements (by which the display is provided to the user).

[0016] Referring to FIG. 3, display element 150 generates a display image. (In this example) The display image is refracted by optical element 160 (shown schematically as a single convex lens, but may be a compound lens or the like). As a result, a virtual image 170 is generated. The virtual image 170 appears to the user to be larger and much farther away than the real image generated by display element 150. As an example, the virtual image may have an apparent image size (diagonal length) exceeding 1 m and may be disposed at a position more than 1 m from the user's eyes (or from the frame of the HMD). Generally, independent of the purpose of the HMD, it is desirable for the virtual image to be disposed far away from the user. For example, in the case of an HMD for movie viewing or the like, it is desirable that the user's eyes can relax during viewing. This requires a distance (to the virtual image) of at least several meters. In FIG. 3, solid lines (e.g., line 180) represent actual light rays, and dotted lines (e.g., line 190) represent virtual light rays.

[0017] FIG. 4 shows an alternative configuration. This configuration may be used when it is desirable that the user's view of the surrounding environment is not completely blocked. However, this configuration can also be applied to an HMD in which the user's view to the outside is completely blocked. In the configuration of FIG. 4, display element 150 and optical element 200 together provide an image projected onto mirror 210. Mirror 210 reflects the image towards the position 220 of the user's eyes. The user feels that the virtual image is at a forward position 230 of the user but is moderately distant from the user.

[0018] In the case of an HMD that completely blocks the user's view of the surrounding environment, the mirror 210 can be a virtually 100% reflective mirror. In this case, the configuration in Figure 4 has the advantage of being able to position the display and optical elements closer to the user's head's center of gravity and next to the user's eyes. This allows for a smaller HMD for the user wearing it. Alternatively, if the HMD is designed not to completely block the user's view of the surrounding environment, the mirror 210 may be a partially reflective mirror. This allows the user to see the surrounding environment through the mirror 210, and the virtual image is superimposed on the surrounding environment.

[0019] Stereoscopic images can be displayed by providing separate displays to the user's left and right eyes. Figure 5 shows an example of a pair of stereoscopic images to be displayed to the left and right eyes. These images are displaced laterally from each other. The displacement of the image features depends on the lateral (actual or simulated) spacing of the cameras that captured the images, the camera angle convergence, and the (actual or simulated) distance of each image feature from the camera position.

[0020] Note that the lateral displacement in Figure 5 may actually be reversed. That is, the image shown for the left eye may actually be the image for the right eye, and the image shown for the right eye may actually be the image for the left eye. This is because some stereoscopic image displays tend to shift the object to the right in the right-eye image and to the left in the left-eye image. This can simulate the feeling that the user is looking at the scenery behind them through a stereo window. However, some HMDs use the configuration shown in Figure 5 to give the user the impression that they are looking at the scenery through binoculars. The choice between these two configurations is left to the discretion of the designer.

[0021] In some situations, HMDs may be used solely for watching movies or other content. In this case, it is not necessary to change the apparent viewpoint of the displayed image when the user moves their head (for example, from one side to the other). However, in other uses such as virtual reality (VR) and augmented reality (AR) systems, the user's viewpoint needs to track the trajectory of their movement relative to the real or virtual space in which they exist.

[0022] Tracking is achieved by detecting the movement of the HMD and changing the apparent viewpoint of the displayed image. As a result, the apparent viewpoint tracks the movement. This will be discussed in more detail later.

[0023] Figure 6 schematically illustrates the effect of user head movements in VR or AR systems.

[0024] In Figure 6, the virtual environment is represented by a spherical shell 250 surrounding the user. To show this configuration on a two-dimensional plane, the shell is represented using a portion of a circle located at a certain distance from the user (i.e., a distance equal to the distance between the displayed virtual image and the user). Initially, the user is placed at a first position 260 and facing a portion 270 of the virtual environment. This portion 270 is what is displayed in the image shown on the display element 150 of the user's HMD.

[0025] Next, consider a scenario where the user moves their head to a new section and / or direction 280. To maintain a correct sense of virtual or augmented reality, at the end of the user's movement, the virtual reality display also moves to the new section 290 (i.e., the section displayed by the HMD).

[0026] Therefore, in this configuration, the apparent viewpoint within the virtual environment moves in conjunction with head movement. As shown in Figure 6, when the head turns to the right, the apparent viewpoint also moves to the right from the user's viewpoint. Considering this situation from the perspective of the displayed object (e.g., object 300), this effectively moves in the opposite direction to the head movement. Therefore, when the head moves to the right, the apparent viewpoint moves to the right. As a result, the object (e.g., object 300), which is stationary in the real environment, will move to the left of the displayed image and eventually disappear at the left edge of the displayed image. This is simply because the display portion of the virtual environment moved to the right, while the displayed object 300 did not move within the virtual environment. A similar consideration applies to the vertical component of movement.

[0027] Figures 7A and 7B schematically show an HMD equipped with a motion detection mechanism. These two figures are presented in the same format as those shown in Figure 2. That is, these two figures are schematic plan views of the HMD, with the display element 150 and optical element 160 indicated by simple rectangles. For simplification, many of the elements in Figure 2 are not shown in Figures 7A and 7B. Figures 7A and 7B show examples of an HMD equipped with a motion detection mechanism for detecting the movement of the HMD and generating corresponding tracking data for the HMD.

[0028] In Figure 7A, a forward-facing camera 320 is provided at the front of the HMD. This camera does not necessarily have to provide a display image to the user (although it may do so in configurations for augmented reality). Rather, the main objective of this embodiment is to realize motion detection. The technique used for image capture using the camera 320 for motion detection will be described below with reference to Figure 8. In this configuration, the motion detection mechanism comprises a camera fixed to move with the frame, and an image comparator that detects motion between images by comparing a sequence of images captured by the camera.

[0029] In Figure 7B, a hardware motion detector 330 is used to generate tracking data. This can be fixed to any location on the HMD. Suitable examples of hardware motion detectors are piezoelectric accelerometers and fiber optic gyroscopes. Of course, both hardware motion detectors and camera-based motion detectors can be used within the same device. In this case, one motion detector may be used as a backup in case of failure of the other. Alternatively, one motion detector (e.g., a camera) may provide data to change the apparent viewpoint of the displayed image, while the other (e.g., an accelerometer) provides data for image stabilization.

[0030] Figure 8 schematically shows an example of motion detection using the camera shown in Figure 7A.

[0031] Camera 320 is a video camera that captures images at a capture rate of, for example, 25 images per second. When each image is captured, it is sent to the image storage unit 400. Each image is then compared by the image comparator 410 to past images retrieved from the image storage unit 400. The comparison is made to determine whether substantially the entire image captured by camera 320 has moved since the previous image was captured. Known block matching techniques (so-called "optical flow" detection) may be used for this comparison. Local motion may represent the movement of an object within the field of view of camera 320. However, the macroscopic motion of substantially the entire image tends to represent the movement of the camera rather than individual features within the captured scene. In this example, the camera is fixed to the HMD. Therefore, the camera's movement corresponds to the movement of the HMD, and further, to the user's head movement.

[0032] The displacement between one image and the next, detected by the image comparator 410, is converted into a motion signal by the motion detector 420. If necessary, the motion signal is converted into a position signal by the integrator 430.

[0033] As described above, instead of (or in addition to) motion detection by detecting movement between images captured by a video camera associated with the HMD, the HMD may detect head movement using a mechanical or solid-state detector 330 (e.g., an accelerometer). In fact, given that the response time of video-based systems is longer than the round-trip time of the image capture rate, these detectors can provide a faster response in terms of motion representation. Therefore, in some examples, detector 330 is ideal for use in high-frequency motion detection. However, in other examples, such as when a high image rate camera (e.g., a camera with a capture rate of 200 Hz) is used, a camera-based system would be more appropriate. In Figure 8, detector 330 can be used instead of camera 320, image storage unit 400, and image comparator 410. In this case, a direct input to motion detector 420 is provided. Alternatively, detector 330 can be used instead of motion detector 420. In this case, a signal representing physical motion can be directly output.

[0034] Needless to say, other techniques can be used to detect position or motion. For example, a mechanical configuration may be used in which the HMD is linked to a predetermined fixed point (e.g., a data processing device or part of the equipment) using a pantograph arm. In this case, position and orientation sensors detect changes in the bending of the pantograph arm. In another embodiment, one or more transmit / receive systems attached to the HMD and the fixed point may be used to detect the position and orientation of the HMD using triangulation techniques. For example, the HMD may be equipped with one or more direction transmitters, and an array of receivers associated with a known (or fixed) point may detect the associated signals from one or more transmitters. Alternatively, the transmitters may be fixed and the receivers may be mounted on the HMD. Examples of transmitters and receivers include infrared transducers, ultrasonic transducers, and radio frequency transducers. Radio frequency transducers may be used to establish a radio frequency data link (e.g., a Bluetooth® link) to and / or from the HMD.

[0035] Figure 9 schematically shows the image processing performed in response to the detected position (or change in position) of the HMD.

[0036] As explained in relation to Figure 6, in applications such as virtual reality and augmented reality, the apparent viewpoint displayed to the user of the HMD changes in response to changes in the actual position or orientation of the HMD and the user's head.

[0037] Referring to Figure 9, this is done by a motion sensor 450 (e.g., the configuration in Figure 8 and / or the motion sensor 330 in Figure 7B) that supplies data representing motion and / or current position (tracking data) to the required position detector 460. The position detector 460 translates the actual position of the HMD into data to define the image to be displayed. If necessary, the image generation unit 480 accesses the image data stored in the image storage unit 470 and generates an image from the appropriate viewpoint required for display by the HMD. An external video signal source can also provide the functionality of the image generation unit 480. Furthermore, this video signal source can also function as a controller to compensate for the low-frequency components of the observer's head movement due to changes in the viewpoint of the displayed image. This allows the displayed image to move in the opposite direction to the detected motion, changing the observer's apparent viewpoint in the direction of the detected motion.

[0038] As shown below, the image generation unit 480 may operate based on metadata called view matrix data.

[0039] To schematically illustrate some of the general concepts related to this technology, Figure 10 shows a schematic representation of image capture by a camera.

[0040] In Figure 10, camera 500 captures an image of a portion 510 of a real-world scene. The field of view of camera 500 is schematically represented as a general triangle 520. Here, the camera is at one of the vertices of the general triangle, and the side adjacent to the camera schematically represents the left and right edges of the field of view. The side opposite the camera schematically represents the portion of the scene being captured. Camera 500 may be a still camera or a video camera that captures a series of images at regular time intervals. The images do not have to be images captured by a camera. All of these techniques are equally applicable to machine-generated images (for example, images generated by a computer game machine to be displayed to a user as part of the gameplay of a computer game).

[0041] Figure 11 schematically illustrates the latency issue related to HMD image display. As described above, the position and / or orientation of the HMD can be used. For example, as explained with reference to Figure 9, the displayed image is generated to be displayed according to the detected position and / or orientation of the HMD. When viewing a wider portion of a captured image, or when generating an image required as part of computer gameplay, the configuration explained with reference to Figure 9 includes a process for detecting the current position and / or orientation of the HMD, and a process for generating appropriate data for display.

[0042] However, the delays associated with this process may result in a corrupted image.

[0043] Figure 11 considers a situation where the user's viewpoint rotates from the first viewpoint 600 to the second viewpoint 610 (clockwise, as schematically shown in Figure 11) at time intervals on the order of the image repetition time period of the HMD image display (e.g., 1 / 25 second). The two representations shown in Figure 11 are shown next to each other, but this is for illustrative purposes only and does not necessarily indicate that the user's viewpoint is translating (although there may be some translational movement between the two viewpoints).

[0044] The HMD's position and / or orientation are detected when it is at viewpoint 600 to allow time to generate the next output image. The next image for display is then generated, but by the time the image is actually displayed, the viewpoint has rotated to viewpoint 610. As a result, the displayed image is no longer correct as it should be displayed at the user's viewpoint 610. This results in a subjectively poor quality image experience for the user, which may cause the user to lose their sense of direction or feel nauseous.

[0045] The techniques described below relate to evaluating tracking data and image frames generated using the tracking data for display, with the aim of detecting inconsistencies. As described above, in certain virtual reality and augmented reality configurations, the viewpoint on the generated image frames changes from one image frame to another in response to changes in the actual position and / or orientation of the HMD as well as the user's head movements. Differences in the tracked position and / or orientation of the HMD, and / or differences in the viewpoint on the generated image frames, can cause the user to lose immersion and, in some cases, cause motion sickness. Therefore, it is necessary to detect inconsistencies between the HMD tracking and the displayed images generated using the tracking data.

[0046] Refer to Figure 12. In embodiments of this disclosure, the data processing device 1200 comprises a receiving circuit 1210, an image processing circuit 1220, a detection circuit 1230, and a correlation circuit 1240. The receiving circuit 1210 receives tracking data representing the position and orientation of at least one tracked HMD. The image processing circuit 1220 uses the tracking data to generate a sequence of image frames for display. The detection circuit 1230 detects image features in a first image frame and the corresponding image features in a second image frame. The correlation circuit 1240 first calculates the difference between the image features in the first image frame and the corresponding image features in the second image frame. The correlation circuit 1240 then generates difference data representing the difference between the viewpoint for the first image frame and the viewpoint for the second image frame, based on the difference between the image features in the first image frame and the corresponding image features in the second image frame. The correlation circuit 1240 then generates output data based on the difference between the difference data and the tracking data for the first and second image frames.

[0047] The data processing unit 1200 may be provided as part of a general-purpose computer device, such as a personal computer. In this case, the receiving circuit 1210 receives HMD tracking data via wired or wireless communication (e.g., Bluetooth® or WiFi®) to and from the HMD. The data processing unit 1200 may also be provided as part of an HMD or game console (e.g., Sony® or PlayStation 5®) that generates display images for video games or other suitable content. Alternatively, the data processing unit 1200 may be implemented in a distributed manner using a combination of an HMD and a game console (which generates the display images for the HMD).

[0048] The receiving circuit 1210 is configured to receive tracking data representing the tracked position and / or orientation of the HMD. The receiving circuit 1210 can receive tracking data via wired or wireless communication with the HMD. Alternatively or additionally, an external camera mounted to capture an image including the HMD may be used to generate tracking data about the HMD using so-called outside-in tracking. Tracking data from an external tracker with one or more cameras may be received using the receiving circuit 1210. More generally, the tracking data represents the position and / or orientation of the tracked HMD relative to the real-world environment and may be acquired using a combination of image sensors given as part of the HMD and / or provided externally to the HMD and / or hardware motion sensors (as described with reference to Figures 8 and 9). The tracking data may be provided via the fusion of sensor data from one or more image sensors and one or more hardware motion sensors.

[0049] While the HMD is displaying a sequence of images, the receiving circuit 1210 may also receive tracking data. Therefore, while the image processing circuit 1220 is displaying a sequence of generated images, the receiving circuit 1210 may receive tracking data, and this received tracking data may be used to generate further image frames displayed by the HMD. The user's head movements while the HMD is displaying images (which cause changes in the HMD's position and / or orientation) can be detected using image-based tracking and / or inertial sensor data (as described with reference to Figures 8 and 9) as well as the corresponding tracking data received by the receiving circuit 1210. However, it may also be possible to save previously generated tracking data using HMD tracking during past game sessions (e.g., by the game console and / or HMD). In this case, the receiving circuit 1210 may also be configured to receive the saved tracking data from storage. In this case, the saved tracking data may be received by the data processing unit 1200. The received tracking data may be used to generate a sequence of image frames in order to perform an offline evaluation of inconsistency detection between the HMD tracking data and the display image generated by the image processing circuit 1220 using the stored tracking data.

[0050] Tracking data represents the position and / or orientation of the HMD for one cycle. Therefore, changes in the physical position and / or orientation of the HMD are reflected in the tracking data. For example, the receiving circuit 1210 may receive tracking data including the 3D position and orientation of the HMD, along with associated timestamps. Such data may be received periodically at a time rate corresponding to the sensors used for tracking. The position and / or orientation of the HMD with a given timestamp relative to a reference point in the real-world environment (or the position and orientation relative to the HMD with a given timestamp) may represent the change in the position and orientation of the HMD relative to the position and orientation of a past timestamp.

[0051] The image processing circuit 1220 is configured to generate a sequence of display image frames based on tracking data received by the receiving circuit 1210. The image processing circuit 1220 includes at least one CPU and GPU capable of performing image processing to generate display images. Image processing typically includes a process of processing model data or other predetermined image data to obtain the pixel values ​​of image pixels in the image frames. Unless otherwise specified, when an image frame is referred to herein, it means either a stereo image frame containing left and right images, or a single image frame viewed by the user with both eyes. The image processing circuit 1220 generates a sequence of image frames containing a plurality of consecutive image frames. In this case, the viewpoint of the image frames changes according to the change in the position and / or orientation of the HMD represented by the tracking data. The sequence of image frames may visually represent content such as virtual reality content or augmented reality content. Therefore, the user can change the position and / or orientation of the HMD by moving their head. Display image frames are generated accordingly. Therefore, the viewpoint of the image frames moves in conjunction with the user's head movements.

[0052] Therefore, the tracking data represents the first position and / or orientation of the HMD at a first time (T1) and the second position and / or orientation of the HMD at a second time (T2), where T2 is a time later than T1. The image processing circuit 1220 generates a predetermined display image frame based on the tracking data at the first time. This generates a predetermined image frame with a viewpoint that depends on the first position and / or orientation of the HMD. Next, the image processing circuit 1220 generates another display image frame based on the tracking data at the second time. This generates a predetermined image frame with a viewpoint that depends on the second position and / or orientation of the HMD. In this way, the viewpoint for each image frame changes according to the change in the position and / or orientation of the HMD. In particular, the image processing circuit 1220 calculates the viewpoint (or the change in viewpoint relative to a previously calculated viewpoint) based on the tracking data and generates a display image frame based on the calculated viewpoint. The calculated viewpoint is updated within this display image frame in response to the most recently received tracking data.

[0053] The image processing circuit 1220 is configured to generate a sequence of image frames at any preferred frame rate. For example, many HMDs use frame rates of 60, 90, or 120 Hz. In this case, the image processing circuit 1220 can generate display images at such frame rates. In this way, the image processing circuit 1220 generates a sequence of display image frames based on changes in tracking data. Each image frame has a viewpoint within this tracking data (which is derived from the tracking data based on the most recent position and / or direction).

[0054] The detection circuit 1230 is configured to detect image features in the sequence of image frames generated by the image processing circuit 1220 by detecting image features contained in the first image frame and corresponding image features contained in the second image frame. However, the first image frame is generated to be displayed at a first time and from a first viewpoint, and the second image frame is generated to be displayed at a second time and from a second viewpoint.

[0055] The image features detected by the detection circuit 1230 include points, edges, and virtual objects within the image frame. Image processing is performed by the image processing circuit 1220 as part of the execution of an application such as a computer game. Image processing may further include a process of processing model data or other predetermined image data according to an image processing pipeline in order to generate rendering image data to be displayed as an image frame. The image processing circuit 1220 may generate an image frame containing one or more virtual objects by performing image processing using various virtual object input data structures. Thus, the detection circuit 1230 may be configured to detect an image of a virtual object in a first image frame and an image of the corresponding virtual object in another image frame. For example, an image of a virtual tree in one image frame and a corresponding image of the same virtual tree in another image frame with a different viewpoint may be detected. This results in the same virtual tree being observed from two different viewpoints and appearing differently in the two image frames (due to this difference in viewpoint). Known computer vision techniques may be used to detect virtual objects in this way.

[0056] The detection circuit 1230 may also use feature point matching techniques. Known computer vision techniques can be used to detect sets of feature points relating to image features within an image frame. For example, corner detection algorithms such as FAST (Features from Accelerated Segment Tes) can be used to extract feature points corresponding to the corners of one or more elements in an image (e.g., the corners of a wall or chair). Feature point matching between image frames can be used to detect image features in one image frame and the same image features in another image frame.

[0057] The detection circuit 1230 may detect feature points within a predetermined image frame and generate a dataset containing multiple feature points detected within the predetermined image frame. In this dataset, each of the detected feature points is associated with image information representing an image feature for each feature point. The image features associated with the detected feature points can be compared with image features in other image frames (e.g., the next image frame or later image frames in the sequence). This makes it possible to detect when the detected feature points are contained within another image frame (which has a different viewpoint). In some examples, the image information may include an image patch extracted from the image frame. This image patch contains a small region of image data (i.e., a region smaller than the entire image frame). This small region can be used as a reference for detection when the detected feature points are contained within another image (e.g., a small region of pixel data). Thus, the image information represents an image feature for the detected feature point. Therefore, when detecting the same feature point in another image later, it can be used to refer to information about its visual appearance in the captured image.

[0058] Therefore, more generally, a given image feature can be detected in a first image frame, and the same given image feature (also called a corresponding image feature) can be detected in another image frame. The two image frames have different viewpoints, but each image frame contains at least one image feature common to both (when viewed from two different viewpoints). The difference in viewpoint between the two image frames is calculated based on the geometric difference between at least one image feature (which is common between the image frames). Furthermore, the same image feature may be detected across multiple image frames in a sequence of image frames. In this case, each image frame has a different viewpoint, and the image feature has a different position and / or orientation relative to the image frame. In the following discussion, we will refer to the first and second image frames and calculate the difference in viewpoint between the two image frames. However, the same technique can also be applied to any number of image frames to calculate the difference in viewpoint between each image frame.

[0059] In the case of image features corresponding to a part of a static virtual environment, the image features visible to the user have a fixed point within the virtual environment. Therefore, the movement of the user's head, which causes a change in the user's viewpoint relative to the virtual environment, results in image frames for display in which the position and / or orientation of the image features changes over time. This makes the image features appear static with respect to the virtual environment. As illustrated with reference to Figure 6, static image features appear to move within the virtual environment in a sequence of image frames. Therefore, due to different viewpoints, the geometric arrangement of the image features within the sequence of image frames changes.

[0060] The correlation circuit 1240 is configured to calculate the difference between an image feature detected in the first image frame and the corresponding image feature detected in the second image frame. The geometric difference between two corresponding image features detected in two or more corresponding image frames can also be calculated using the correlation circuit 1240, which is used to calculate the difference in viewpoints between two or more image frames. For example, sets of image feature points corresponding to the corners of buildings or the edges of mountains in the background of a virtual environment can be detected in each image frame and matched with each other to represent the same image feature. However, since these are viewed from different viewpoints, the difference in viewpoints can be calculated by calculating the difference in the geometric arrangement of these image feature points.

[0061] As explained above, one or more detected image features are, for example, virtual objects generated by the image processing circuit 1220 for display purposes. A virtual object, such as a virtual tree, may be contained in a first image frame having a first position and orientation, and may also be contained in a second image frame having a second position and orientation. The position and orientation of an image feature within an image frame depend on the viewpoint relative to the image frame. Therefore, with respect to image frames with different viewpoints, the same image feature will have different positions and / or orientations within those image frames. Feature matching can be performed to match virtual objects within each image frame, and the difference in position and / or orientation between two image frames can be calculated. Furthermore, the first image frame may contain multiple virtual objects. In this case, at least some of the multiple virtual objects may be contained in the second image frame. In this case, feature matching can be performed to match the virtual objects between image frames. Then, the difference in position and orientation between two image frames for multiple virtual objects can be calculated for use in calculating the difference between viewpoints relative to the image frames.

[0062] Thus, the correlation circuit 1240 calculates the difference between at least one image feature in the first image frame and the corresponding image feature in the second image frame, and generates difference data based on the difference between the image feature in the first image frame and the corresponding image feature in the second image frame. The difference data represents the difference between the viewpoint position and / or direction with respect to the first image frame and the viewpoint position and / or direction with respect to the second image frame. As described above, the difference data may be generated based on image features and corresponding image features detected in each of the two image frames with different viewpoints, and / or based on multiple image features and multiple corresponding image features detected in each of the two image frames.

[0063] The correlation circuit 1240 is configured to generate output data based on the difference between the difference data generated for the first and second image frames and the tracking data related to the first and second image frames. As described above, the first image frame is generated for the image based on the tracking data at the first time (T1). The second image frame is generated for the image based on the tracking data at the second time (T2). As explained with reference to Figure 11, the delay associated with the process of generating image frames for display is that when image frames are output for display consecutively, the viewpoint of the HMD at the time of image output is different from the viewpoint previously used to generate the image frame. In an ideal scenario, the tracking data would ideally reflect that the HMD's position has moved N meters (e.g., N=1) in a predetermined direction, so that the tracking data indicates that the HMD's position at the second time is N meters from its position at the first time. Similarly, based on the geometric difference between image features in the first image frame and the corresponding image features in the second image frame, it is desirable that the difference data generated by the correlation circuit 1240 ideally represents that the difference between the viewpoint for the first image frame and the viewpoint for the second image frame is N meters (e.g., N=1). For simplicity of explanation, the above discussion dealt with the example of movement in a single axis, but it should be noted that changes in position and / or direction in 3D space can be tracked in the same way. However, processing errors related to the generation of the display image frame may result in changes in the HMD's position and / or direction that do not match the changes in the viewpoint's position and / or direction relative to the generated image frame. In particular, errors in one or more stages of the image processing pipeline can lead to a mismatch between the geometric properties of the generated image frame and the changes in the HMD's position relative to the image frame. This causes the user to feel that the displayed image is not correctly tracking the user's head movements.

[0064] Therefore, the correlation circuit 1240 first calculates the geometric difference between two corresponding image features in each of the two image frames. The correlation circuit 1240 then generates difference data representing the difference between the viewpoint in the first image frame and the viewpoint in the second image frame, based on the calculated geometric difference. The correlation circuit 1240 then compares the difference data with tracking data for the first image frame and tracking data for the second image frame to calculate the difference between the change in the HMD's position and / or orientation and the change in the viewpoint's position and / or orientation relative to the first and second image frames. The correlation circuit 1240 then generates output data based on the difference between the change in the HMD's position and / or orientation and the change in the viewpoint's position and / or orientation relative to the first and second image frames. The output data thus generated gives a representation of the amount of mismatch between the change in the HMD viewpoint and the change in the viewpoint relative to the image frames (if any). More generally, if the HMD viewpoint changes relative to the real environment, the change in the associated virtual viewpoint relative to the virtual environment can be calculated using image feature detection in two or more consecutive image frames. Then, the change in the real-world viewpoint is compared with the change in the virtual viewpoint. Finally, output data is generated to represent the geometric difference (difference in position and / or direction) between the change in the real-world viewpoint and the change in the virtual viewpoint.

[0065] Therefore, the output data represents a positional or directional offset of the virtual viewpoint relative to the HMD viewpoint (which is indicated by the tracking data).

[0066] Figure 13a shows the change in the tracked HMD position and the change in the viewpoint position calculated with respect to image frames generated based on the tracking data (which represents the tracked HMD position). Figure 13a schematically shows a timing diagram. The increase in time is shown from left to right along the horizontal axis. The position with respect to a given 3D coordinate system is shown on the vertical axis (for simplicity of explanation, only one axis of the coordinate system is shown, but the position and orientation of the HMD can be tracked with respect to 3D space using the 3D coordinate system). In the illustrated example, the increasing position on the vertical axis represents the increase in distance with respect to the reference point. Data point 1310 represents the change in the HMD position represented by the tracking data. Thus, this represents the change in the position of the HMD viewpoint relative to the real-world environment caused by the user's head movement. Data point 1320 represents the change in the viewpoint position calculated with respect to the generated image frames (i.e., calculated based on the difference between image features). Thus, this represents the change in the position of the virtual viewpoint. In the example in Figure 13a, it can be seen that a given change in the real-world viewpoint is reflected in the amount of change in the corresponding virtual viewpoint. Therefore, the system operates ideally, and the position of the virtual viewpoint calculated by the correlation circuit 1240 lags behind the position of the real-world viewpoint represented by the tracking data. This lag is due to the delay associated with the generation of the display image frame. The amount of this lag is indicated by arrow 1305. Thus, Figure 13a schematically illustrates an ideally operating system with a delay associated with image frame generation (which causes an offset between pairs of data points).

[0067] Figure 13b schematically shows another timing diagram. For reference, data points 1310 and 1320 are also shown here. In Figure 13b, data point 1330 represents the change in viewpoint position (i.e., the change in virtual viewpoint position) with respect to the generated image frame. A discrepancy can be seen between the change in HMD position represented by the tracking data and the change in viewpoint position calculated by the correlation circuit 1240 based on the geometric properties of the image features within the generated image frame. In particular, in the illustrated example, the change in virtual viewpoint position calculated by the correlation circuit 1240 differs from the change in HMD position. Therefore, in this case, there is a discrepancy between the difference data generated by the correlation circuit 1240 and the HMD position represented by the tracking data. Consequently, the correlation circuit 1240 generates output data representing the difference between the difference data and the tracking data. In particular, in this example, the output data represents the difference between the HMD position and the field of view position calculated for the image frames at t1, t2, and t3. For simplicity, Figures 13a and 13b show only one axis of the coordinate system, but three axes may also be used. Similarly, the difference in direction may also be represented in the generated output data.

[0068] Therefore, more generally, the correlation circuit 1240 is configured to generate output data representing the difference between the pose of the virtual viewpoint with respect to a sequence of image frames (which is generated based on the geometric difference of corresponding image features in different image frames) and the pose of the HMD represented by tracking data. Thus, the output data may represent the difference between the position of the virtual viewpoint and the position of the HMD, or the difference between the direction of the virtual viewpoint and the direction of the HMD.

[0069] The output data may represent the difference between the virtual viewpoint pose for the image frame and the HMD pose represented by the tracking data used to generate the image frame. Therefore, in a complete system, the two poses should coincide, and any discrepancies between the poses can be assumed to be represented by the output data.

[0070] Selectively, the output data may represent the delay between the virtual viewpoint and the HMD viewpoint. As shown in Figure 13a, the position of the virtual viewpoint relative to a given image frame calculated by the correlation circuit 1240 lags behind the position of the real-world viewpoint represented by the tracking data. The amount of this delay is indicated by arrow 1305. This delay is due to the delay associated with the generation of the display image frame. The tracking data representing the HMD's orientation at the first time step is used to generate the display image frame. The image frame is then output and analyzed to calculate the viewpoint relative to the image frame (or the difference between the viewpoints of two image frames). Thus, by the time the image frame is output and analyzed, the tracking data now represents the updated orientation relative to the HMD (which is different from the viewpoint orientation calculated relative to the image frame). Therefore, the output data may include timing data, or offset data representing the time offset between the real-world viewpoint orientation and the virtual viewpoint orientation calculated based on image feature analysis.

[0071] Accordingly, the data processing unit 1200 receives tracking data related to the HMD, generates display image frames based on the tracking data analysis, performs custom image analysis on image features within the image frames (which calculates the pose of the virtual viewpoint associated with the image frames), and outputs output data representing the geometric difference between the ideally predicted pose with respect to the virtual viewpoint (which should be consistent with the tracking data) and the calculated pose with respect to the virtual viewpoint. The output data therefore indicates the presence of processing errors and inconsistencies in the image processing pipeline (which cause the actual viewpoint to deviate from the predicted field of view), and can be used as an aid in debugging.

[0072] In some embodiments of this disclosure, the detection circuit 1230 is configured to detect the position and orientation of an image feature in a first image frame and the position and orientation of a corresponding image feature in a second image frame. The orientation (position and orientation) of the image feature in the first image frame and the orientation of the image feature in the second image frame are detected by the detection circuit 1230. The orientation of an image feature contained in the first image frame depends on the viewpoint relative to the first image frame. Similarly, the orientation of an image feature contained in the second image frame depends on the viewpoint relative to the second image frame. In the above discussion, the image feature is a feature of a static image, which is static with respect to a virtual environment. The correlation circuit 1240 may be configured to apply viewpoint shift or viewpoint rotation to the image frame in order to adjust the viewpoint relative to the image frame. This adjusts the orientation of the image feature to correspond to the orientation of the image feature in other image frames. In particular, with respect to a specific viewpoint indicated by a predetermined rotation by position (x, y, z) and yaw, pitch, roll (α, β, γ), the viewpoint can be shifted to a new viewpoint by adjusting the above parameters in order to apply image distortion to the image frame. As a result, the orientation of an image feature matches the orientation of an image feature in the corresponding other image frame. For example, the geometric visual technique disclosed in "Finding the exact rotation between two images independently of the translation," L. Kneip et al., LNCS, Vol. 7577, can be used to establish viewpoint rotation and viewpoint movement between image frames. The contents of this document are incorporated herein by reference.

[0073] Therefore, more generally, the changes used to move and / or rotate the viewpoint for the purpose of matching the poses of image features represent the difference between the respective poses of the viewpoints in two image frames. Thus, the correlation circuit 1240 calculates the difference between the two viewpoints based on the changes and generates difference data. Although the above discussion has concerned the poses of two corresponding image features in each image frame, multiple image features may be detected in a given image frame, and distortion may be applied to multiple image features. This increases reliability. Therefore, more generally, in some embodiments of the present disclosure, the correlation circuit 1240 is configured to calculate the change between the viewpoint of the first image frame and the viewpoint of the second image frame based on the difference between the image feature in the first image frame and the corresponding image feature in the second image frame.

[0074] In some embodiments of this disclosure, the detection circuit 1230 is configured to detect image features in a first image frame and corresponding image features in a second image frame based on feature matching. The detection circuit may perform a variety of possible feature matching processes / algorithms, including Harris corner detectors, scale-invariant feature transformations, fast robust feature detectors, and Features from Accelerated Segment Test. More generally, image features in each image frame are detected and associated with descriptors containing image data relating to the image features. Matches of image features in different image frames are identified by comparing the descriptors to identify matching features. Then, differences in the geometric configuration of matching features in the image frames are calculated to compute viewpoint shifts and rotations between image frames.

[0075] In some embodiments of this disclosure, the sequence of image frames includes a sequence of stereo images. These stereo images include a left image and a right image. The image processing circuit 1220 can be configured to generate a stereo image that includes a left image and a right image for display to the left eye and right eye, respectively (see, for example, Figure 5).

[0076] In some embodiments of this disclosure, the correlation circuit 1240 is configured to calculate the difference between an image feature in the left or right image of a first stereo image frame and an image feature in the left or right image of a second stereo image frame (subsequently). Thus, in the case of a stereo image, an image feature can be detected in the left (or right) image of the first stereo image frame. The corresponding image feature can then be detected in the left (or right) image of another stereo image frame. Since the difference between the two image features can be calculated, the difference in viewpoint between the first stereo image frame and the second stereo image frame can be calculated. Thus, in some embodiments of this disclosure, the technique for generating output data representing the mismatch between the pose of a virtual viewpoint in a virtual environment and the pose of the HMD viewpoint in a real environment can be performed based on feature matching with respect to image features contained in the left or right image of the stereo image sequence.

[0077] By generating stereo images for display, an observer wearing an HMD experiences the illusion of depth thanks to the left and right images observed by the user's left and right eyes, respectively. As shown in Figure 2, stereo images can be displayed on the HMD's display unit. The depth of a given image feature (e.g., a virtual object) within the stereo image observed by the user can be adjusted to be in front of, behind, or approximately the same depth as the HMD's display screen, based on the parallax of the given image feature. For example, if the virtual object should be observed in front of the display screen, the left and right images are displayed with a parallax called reverse parallax. If the virtual object should be observed at the same depth as the display screen, there is typically no parallax between the left and right images of the virtual object. If the virtual object should be observed behind the display screen, the left and right images are typically displayed with the parallax of that virtual object.

[0078] Figure 14 schematically shows a display unit 540 that displays a stereo image. The display unit 540 includes a left image and a right image with binocular eccentricity distances such that the virtual object 500 is behind the display screen position (i.e., the virtual object 500 is displayed at a distance from the user than the display screen position). If the virtual object is a static image feature in a virtual environment, the depth of the position where the virtual object is observed by the user relative to the HMD's viewpoint changes in conjunction with the change in the HMD's viewpoint in the depth direction. For example, if the virtual viewpoint is observed to be Z meters away from the user in the depth direction (e.g., in the Z-axis direction), and the HMD moves substantially toward the virtual viewpoint in the depth direction, the virtual object will be observed by the user as substantially approaching the HMD in the depth direction.

[0079] In some embodiments of this disclosure, the correlation circuit 1240 is configured to calculate a first binocular divergence distance with respect to an image feature based on the image features of the left and right images of the first image frame. The detection circuit 1230 can detect image features within the first image frame by detecting the image features of the left and right images of the first image frame. Thus, the position of the image feature in the left image and the position of the image feature in the right image can be detected. Based on the lateral separation between the image features and the position of the user's eyes (which, in the case of an HMD, is a predetermined position with respect to the position of the display unit), the binocular divergence distance (where the lines of sight of both eyes intersect with respect to the image feature) can be calculated.

[0080] In some embodiments of this disclosure, the detection circuit 1230 is configured to calculate a second binocular divergence distance with respect to an image feature based on the image features of the left and right images of the second image frame. Thus, in addition to the first binocular divergence distance (the depth at which the image feature in the first image frame is observed), the second binocular divergence distance (the depth at which the image feature in the second image frame is observed) can also be calculated. Due to the movement of the HMD, the depth at which the image feature in the first image frame is observed differs from the depth at which the image feature in the second image frame is observed. For example, if the first image frame is generated using tracking data representing the position (0,0,0) and the second image frame is generated using tracking data representing the position (0,0,Z2), and the HMD moves Z2 in the Z-axis direction toward the observed image feature, then the first binocular divergence distance associated with the first image frame must be longer than the second binocular divergence distance associated with the second image frame. Conversely, if the first image frame is generated using tracking data representing the position (0,0,0), and the second image frame is generated using tracking data representing the position (0,0,-Z2), and the HMD moves Z2 in the Z-axis direction away from the observed image feature, then the first binocular divergence distance associated with the first image frame must be shorter than the second binocular divergence distance associated with the second image frame. Therefore, the correlation circuit 1240 can calculate the first and second binocular divergence distances, calculate the difference between these two binocular divergence distances, and compare the calculated difference with the tracking data associated with the first and second image frames. This allows it to determine whether the difference between the two binocular divergence distances corresponds to a change in the HMD's separation distance (particularly a change in the depth direction of the image feature). In an ideal system, the difference between the two binocular eccentricity distances should match the change in the HMD's separation distance (in the depth direction of the image features) (which is expressed as the difference between the tracking data for the first image frame and the tracking data for the second image frame). However, errors in one or more stages of the image processing pipeline can lead to a mismatch between the geometric properties of the generated image frames and the change in the HMD's position relative to the image frames.This causes the user to perceive that the displayed image is not correctly tracking the user's head movements. Therefore, the correlation circuit 1240 can further generate output data based on the difference between two binocular eccentricity distances and the change in the depth direction of the HMD position represented by the tracking data used for the first and second image frames. This allows for the identification of geometric errors. Thus, the output data can be used as a debugging aid because it indicates the presence of a processing error in the image processing pipeline (which causes the actual binocular eccentricity distance to deviate from the predicted binocular eccentricity distance).

[0081] Therefore, more generally, in some embodiments of the present disclosure, the detection circuit 1230 is configured to detect image features in the second image frame by detecting image features in the left and right images of the second image frame. The correlation circuit 1240 first calculates a second binocular divergence distance related to the image features based on the image features in the left and right images of the second image frame. The correlation circuit 1240 then generates binocular divergence distance difference data representing the difference between the first binocular divergence distance and the second binocular divergence distance. The correlation circuit 1240 then generates output data based on the difference between the binocular divergence distance difference data and tracking data related to the first and second image frames.

[0082] Moving on to Figure 15. In some embodiments of the present disclosure, the data processing device 1200 includes a test circuit 1250 configured to generate test tracking data. The image processing circuit 1220 is configured to generate at least one display image frame based on the test tracking data. The test circuit 1250 is configured to generate test tracking data representing the position and / or orientation of the HMD. The test tracking data may, for example, represent a three-dimensional position using a three-dimensional coordinate system, and / or represent an orientation relative to the three-dimensional coordinate system. For example, the three-dimensional position and / or orientation may be represented with respect to a reference point (origin). This reference point is the same as the reference point for the tracking data received by the receiving circuit 1210. Alternatively, the test tracking data may represent a relative change in the three-dimensional position and / or orientation, rather than representing a position and / or orientation relative to a reference point.

[0083] The image processing circuit 1220 is configured to generate a sequence of images based on tracking data received by the receiving circuit 1210, and to generate at least one image frame within the sequence of image frames based on test tracking data. Thus, the sequence of image frames includes multiple consecutive image frames. Each of these image frames has an associated viewpoint. At least one image frame within the sequence of image frames is generated along with a viewpoint based on the test tracking data.

[0084] At least one image frame may be generated using only test tracking data to determine the viewpoint of that at least one image frame. For example, the receiving circuit 1210 may receive a stream of tracking data that gives a representation of the position and / or orientation of the HMD at periodic time intervals. With respect to a given image frame, the image processing circuit 1220 may also be configured to use test tracking data instead of the received tracking data to generate the given image frame.

[0085] Alternatively, instead of replacing the tracking data with test tracking data to generate a predetermined image frame, the test tracking data may be used to modify the tracking data. In this case, the image processing circuit 1220 is configured to generate a predetermined image frame based on both the received tracking data and the test tracking data. For example, the received tracking data may include the 3D position (x, y, z), and the test tracking data may include the change in the 3D position, for example, (x1, y1, z1). Thus, the combination of the received tracking data and the test tracking data at a predetermined timestamp represents the 3D position (x+x1, y+y1, z+z1). Similarly, the test tracking data may include negative changes with respect to any axis of x, y, and z.

[0086] Therefore, more generally, test tracking data can be combined with the received tracking data to correct it. This allows for obtaining corrected tracking data. In this case, the image processing circuit 1220 can be configured to generate at least one image frame based on the corrected tracking data. Alternatively, the received tracking data as described above may be directly replaced with test tracking data.

[0087] In this way, the image processing circuit 1220 generates a sequence of display image frames. Within this sequence, at least one display image frame is generated based on the test tracking data. Therefore, the test circuit 1250 can be used to provide test HMD pose data (test tracking data). The test HMD pose data is provided to the image processing circuit 1220. As a result, a so-called "fake HMD pose" can be injected into the image processing pipeline. Thus, the test HMD pose data can be provided as input at a controlled time. The test HMD pose data is then reflected in the changes in viewpoint in the generated image frames. In particular, by inputting the test HMD pose data at a known time, detecting image features within the image frames using the above technique, and generating difference data representing the difference in viewpoint between two image frames, a sudden change in viewpoint caused by the test HMD pose data can be identified in a specific image frame generated by the image processing circuit 1220. For example, the test HMD pose data may be a step input that effectively adds (or subtracts) a predetermined amount from the received tracking data (e.g., by introducing an offset amount relating to the position on at least one axis of the 3D coordinate system). In this case, the difference data represents the change in viewpoint in a given image frame within a sequence of image frames. Therefore, the time difference between the input test HMD pose data and the output image frames (which reflect the change in viewpoint pose relative to the test HMD pose data) can be calculated.

[0088] Accordingly, in some embodiments of this disclosure, the correlation circuit 1240 is configured to first generate differential data and then calculate the delay for the image frame generation process by the image processing circuit 1220 based on the test tracking data and the differential data.

[0089] In the above discussion, test tracking data was used as a controlled input to the image processing circuit 1220 to quickly and reliably calculate the delay related to the generation of display image frames. However, alternatively, the correlation circuit 1240 can be configured to calculate this delay without using the test circuit 1250. As illustrated with reference to Figures 13a and 13b, the correlation circuit 1240 is configured to generate output data representing the difference between the pose of a virtual viewpoint with respect to a sequence of image frames (which is calculated based on the geometric difference of corresponding image features contained in different image frames) and the pose of the HMD represented by the tracking data. As shown in Figure 13a, the position of the virtual viewpoint calculated by the correlation circuit 1240 lags behind the position of the real-world viewpoint represented by the tracking data. This lag is due to the delay related to the generation of display image frames. The amount of this lag is indicated by arrow 1305. Thus, in some embodiments of this disclosure, the correlation circuit 1240 is configured to calculate the delay related to the image frame generation process by the image processing circuit based on the received tracking data and difference data, and to generate output data including timing data representing the delay related to the image frame generation process.

[0090] In some embodiments of this disclosure, the test circuit 1250 is configured to predict the orientation of the HMD based on tracking data received by the receiving circuit 1210 and to generate test tracking data including at least one predicted orientation for the HMD. The test circuit 1250 may also be configured to predict further position and / or orientation of the HMD based on the tracking data for the HMD. For this purpose, one or more suitable predictive tracking algorithms may be used. In some examples, the test circuit 1250 includes a machine learning model configured to generate test tracking data based on tracking data received by the receiving circuit 1210. The machine learning model is trained using past examples of HMD tracking data. In response to receiving an input including multiple HMD orientations, the machine learning model generates an output representing a predicted HMD orientation with respect to the input.

[0091] Therefore, the test circuit 1250 can also be configured to predict further HMD postures based on the received tracking data. For example, in response to receiving a stream of tracking data representing the trajectory of the HMD's motion, the test circuit 1250 can generate an output representing one or more predicted HMD postures for that trajectory. Thus, the test circuit 1250 can also be configured to generate test tracking data including at least one predicted HMD posture. In this case, the image processing circuit 1220 generates one or more display image frames based on the test tracking data. Thus, the test circuit 1250 can use insights into a calculated delay (e.g., time 1305 shown in Figure 13a) to generate test tracking data for input to the image processing circuit 1220 at a controlled time. As a result, the predicted HMD posture is input by a predetermined time before the received tracking data (obtained by tracking the HMD) representing such HMD posture is input. Then, one or more display image frames can be generated based on the test tracking data. This eliminates (or at least reduces) the delay that causes the virtual viewpoint to lag behind the HMD's real-world viewpoint. In other words, by generating test tracking data that includes at least one predicted HMD pose, the time difference (indicated by arrow 1305) between the HMD's tracked pose and the tracked pose reflected in the viewpoint of the display image frame output can be reduced, and in some cases reduced to the point where the delay is considered nonexistent.

[0092] In some embodiments of this disclosure, the image processing circuit 1220 is configured to generate each image frame for display by applying image distortion based on at least one optical element of the HMD. The image processing circuit 1220 can generate image frames for display by applying image distortion to the image frame to correct distortion related to the optical elements of the HMD. Light distortion can occur when light is guided (e.g., focuses or diverges) by an optical element comprising one or more lenses and / or one or more mirrors. For example, the geometry of a lens means that light incident on the lens may be reflected in different directions at different locations on the lens. Thus, the light may be guided in different directions depending on the location-specific characteristics of the lens. Therefore, the image processing circuit 1220 can apply an inverse distortion, which is the opposite of the distortion associated with the optical element, to neutralize the distortion of the optical element. Thus, inverse image distortion may be applied by rendering the image frame and distorting it as a post-processing effect. In particular, to neutralize pincushion distortion related to the optical elements of the HMD, the image processing circuit 1220 may apply barrel distortion such that the distortion is substantially eliminated. This allows the user to observe an image that is virtually distortion-free. Therefore, the detection circuit 1230 can be configured to detect image features within an image frame into which image distortion has been introduced, and to detect the differences between image features within the image frame. Alternatively, the detection circuit 1230 may be configured to detect image features within an image frame by detecting one or more image features within the image frame before image distortion has been introduced to the image frame.

[0093] Moving on to Figure 16. In some embodiments of this disclosure, the data processing method includes the following: (Step 1610) A step of receiving tracking data representing the tracked posture of a head-mounted display (HMD). (Step 1620) A step to generate a sequence of image frames for display based on tracking data. (Step 1630) A step of detecting image features in the first image frame and corresponding image features in the second image frame. (Step 1640) A step to calculate the difference between the image features in the first image frame and the corresponding image features in the second image frame. (Step 17650) A step to generate difference data representing the difference between the viewpoint for the first image frame and the viewpoint for the second image frame, based on the difference between the image features in the first image frame and the corresponding image features in the second image frame. (Step 1660) A step to generate output data based on the difference between the difference data and the tracking data associated with the first and second image frames.

[0094] The illustrated embodiments can be performed by computer software running on a general-purpose computer system such as a game machine. Examples of embodiments of the present disclosure also include computer software that, when run on a computer, causes the computer to perform any of the above methods. Similarly, some embodiments of the present disclosure are provided by a non-temporary, machine-readable storage medium on which the above software is recorded.

[0095] With respect to the above teachings, it goes without saying that various modifications and improvements to this disclosure are possible. Accordingly, unless otherwise specified, this disclosure is implementable without departing from the scope of the appended claims.

Claims

1. a receiving circuit for receiving tracking data representative of a position and orientation of at least one tracked head mounted display (HMD); an image processing circuit for generating a sequence of image frames for display based on the tracking data; a detection circuit for detecting an image feature in a first image frame and detecting a corresponding image feature in a second image frame; a correlation circuit; Equipped with The correlation circuit calculating a difference between an image feature in the first image frame and a corresponding image feature in the second image frame; generating difference data representing a difference between a viewpoint with respect to the first image frame and a viewpoint with respect to the second image frame based on a difference between an image feature in the first image frame and a corresponding image feature in the second image frame; A data processing apparatus for generating output data based on the difference data and tracking data relating to the first image frame and the second image frame.

2. 2. A data processing apparatus according to claim 1, wherein the detection circuitry is configured to detect the position and orientation of an image feature in the first image frame and to detect the position and orientation of a corresponding image feature in the second image frame.

3. 3. A data processing apparatus according to claim 1, wherein the correlation circuit is configured to calculate a change between a viewpoint relative to the first image frame and a viewpoint relative to the second image frame based on a difference between an image feature in the first image frame and a corresponding image feature in the second image frame.

4. the sequence of image frames comprises a sequence of stereo images; 2. The data processing apparatus according to claim 1, wherein each of the stereo images includes a left image and a right image.

5. 5. A data processing apparatus according to claim 4, wherein the correlation circuit is configured to calculate the difference between an image feature in either the left image or the right image of the first image frame and a corresponding image feature in either the left image or the right image of the second image frame.

6. 5. A data processing apparatus according to claim 4, wherein the detection circuitry is configured to detect image features in the first image frame by detecting image features in a left image and a right image of the first image frame.

7. 7. The data processing apparatus of claim 6, wherein the correlation circuit is configured to calculate a first vergence distance associated with an image feature based on the image feature in the left and right images of the first image frame.

8. the detection circuitry detects image features in the second image frame by detecting image features in a left image and a right image of the second image frame; The correlation circuit calculating a second vergence distance associated with an image feature based on the image feature in the left image and the right image of the second image frame; generating vergence distance difference data representing a difference between the first vergence distance and the second vergence distance; 8. A data processing apparatus according to claim 7, configured to generate the output data based on the vergence distance difference data and tracking data relating to the first image frame and the second image frame.

9. a test circuit configured to generate test tracking data; 2. The data processing apparatus of claim 1, wherein the image processing circuitry is configured to generate at least one image frame for display based on the test tracking data.

10. The test circuit Predicting the posture of the HMD based on the tracking data; 10. The data processing apparatus of claim 9, configured to generate test tracking data including at least one predicted pose for the HMD.

11. 2. A data processing apparatus according to claim 1, wherein said correlation circuit is configured to calculate a delay associated with the image frame generation process by said image processing circuit.

12. 2. The data processing apparatus of claim 1, wherein the image processing circuitry is configured to generate each image frame for display by applying image distortion based on at least one optical element of an HMD.

13. The data processing device of claim 1 , wherein the tracking data is generated based on captured images and / or one or more inertial sensors of an HMD.

14. receiving tracking data representing a tracked pose relative to a head mounted display (HMD); generating a sequence of image frames for display based on the tracking data; Detecting an image feature in a first image frame and a corresponding image feature in a second image frame; Calculating the difference between an image feature in the first image frame and a corresponding image feature in the second image frame. generating difference data representing a difference between a viewpoint relative to the first image frame and a viewpoint relative to the second image frame based on a difference between an image feature in the first image frame and a corresponding image feature in the second image frame; generating output data based on the difference between the difference data and tracking data associated with the first image frame and the second image frame; A data processing method comprising:

15. Computer software which, when executed on a computer, causes the computer to carry out the data processing method of claim 14.