A method and system for processing virtual reality content for preventing motion-sickness
The VR content processing method addresses motion sickness in VR users by filtering temporal contrasts in the peripheral vision area, improving user comfort and experience through effective luminance reduction.
Patent Information
- Application Number
- PCT/EP2024/083238
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
Virtual reality (VR) content often induces motion sickness in users due to incompatible sensory cues, particularly in the peripheral vision area, leading to discomfort and disorientation.
A computer-implemented method for processing VR content that involves determining the contrast in the peripheral vision area between video frames and applying a smoothing temporal filtering if the contrast exceeds a threshold, thereby reducing luminance and mitigating motion sickness.
The method effectively reduces motion sickness in VR users by smoothing temporal contrasts in the peripheral vision area, enhancing the overall VR experience by minimizing discomfort and disorientation.
Smart Images

Figure EP2024083238_30052025_PF_FP_ABST
Abstract
Description
A method and system for processing virtual reality content for preventing motionsickness
[0001] The present application claims priority from European Patent Application No. EP23383208.8 entitled “A method and system for processing virtual reality content for preventing motion-sickness” and filed on 24th November 2023.
[0002] The present disclosure is related to processing virtual reality content.BACKGROUND
[0003] Computer generated virtual reality (VR) allows a user to, for example, be immersed in a simulated real environment or an imaginary environment. A user may be able to interact with the simulated environment, comprising one or more of moving and / or looking around the VR environment, and interacting with objects within the VR environment.
[0004] VR systems may comprise a display to let the user view the VR environment, The display systems may include a computer monitor or display screen that is presented in front of the user.
[0005] VR systems may be included in various industries, such as the military, real estate, medicine, video gaming, etc. The VR environment may simulate a real environment for purposes of training or enjoyment. For example, a VR system may be used for surgical technique training within the medical industry.
[0006] The VR user’s experience of a user may vary with the type of VR content presented. In some cases, VR content may induce a sickness or discomfort in the user similar to motion sickness experienced by a user of a means of transportation. The discomfort by the user may be due to sensory cues (e.g., visual, auditory, etc.) that may not be compatible with the user.
[0007] Embodiments of this disclosure aim at processing VR content which reduces motion sickness.SUMMARY
[0008] The present disclosure describes a computer-implemented method for processing a virtual reality, VR, content, for reducing motion sickness, the VR content comprising a plurality of video frames, the plurality of video frames comprising at least two video frames. The processing comprises: determining a contrast of at least part of a peripheral vision area of a field of view, FOV, region between at least two video frames of the plurality of video frames; and, upon the contrast over a threshold, or in other words, if the contrast is over athreshold, filtering the at least part of the peripheral vision area of at least one video frame, wherein the filtering comprises a smoothing temporal filtering. The method may comprise providing the processed VR content. The method may comprise receiving the plurality of video frames. The video frames may be received, processed, and provided to be displayed by a VR display.
[0009] Virtual and augmented reality environments are generated by computers using, in part, data that describes the environment. The data may describe, for example, various objects with which a user may interact with. Examples of the objects include objects that are rendered and displayed for a user to see.
[0010] Generating virtual reality, VR, content may therefore comprise the generation, by a graphical processing unit, or by one or more computers, of images to be displayed on a screen or on a display. The graphical processing unit may be included or integrated in devices such as a head-mounted display, HMD. An HMD may comprise a display device, configured to be worn on the head of a user or as part of a helmet, and may comprise a relatively small virtual reality display, VR display, configured to be set in front of one eye, for monocular HMD, or in front of each eye, binocular HMD, when worn by a user.
[0011] Virtual reality, VR, systems may be useful for many applications, spanning the fields of scientific visualization, medicine and military training, engineering design and prototyping, tele-manipulation and tele-presence, and personal entertainment.
[0012] Processing VR content, or processing video frames, in the context of the present disclosure, may include modifying or manipulating video frames in some examples, or processing without manipulating or without modifying video frames in other examples. Processing VR content comprises the determination of a contrast between two or more areas or sections of two or more video frames and the filtering, or not, of at least one of the video frames which are subject to the comparison.
[0013] The computer-implemented methods of the disclosure comprise processing video frames which may, after processing, be displayed by a VR display. The plurality of video frames may be processed by a processing unit and may be displayed at a displaying frame rate. In some examples, the methods of the present disclosure may comprise the provision of VR content, the VR content comprising the processed plurality of videoframes. The provision may be to, for example, a display forming part of the VR content processing system or a more complex VR system. The provision may be to a different system or module. A VR system or a VR content processing system may provide or supply video frames at frame rates of approximately 90 frames per second, FPS.
[0014] The plurality of video frames may comprise a sequence of pictures forming a video or video sequence. Each -digital- video frame of the plurality of video frames may be regarded as a two-dimensional array or matrix of samples with intensity values. An elementor entry in the matrix, or sample, may also be referred to as pixel. The number of samples or pixels in horizontal and vertical direction, or axis, of the matrix or video frame define the size and / or resolution of the picture. For representation of color, typically three color components are employed, i.e. the video frame may be represented or include three sample arrays. In RGB format or color space a picture comprises a corresponding red, green and blue sample array. In video coding each pixel is typically represented in a luminance and chrominance format or color space, e.g. YCbCr, which comprises a luminance component indicated by Y or L and two chrominance components indicated by Cb and Cr. The luminance component Y or L represents the brightness or grey level intensity, while the two chrominance -or chroma- components Cb and Cr represent the chromaticity or color information components. Accordingly, a video frame in YCbCr format comprises a luminance sample matrix of luminance sample values (Y), and two chrominance sample arrays of chrominance values (Cb and Cr). Video frames in RGB format may be converted or transformed into YCbCr format and vice versa. If a video frame is monochrome, the video frame may comprise only a luminance sample matrix. Accordingly, a video frame may be, for example, an array of L samples in monochrome format or an array of L samples and two corresponding matrices of chroma samples. Luminance values in a display may be comprised in a range between 50 and 300 cd / m2- candela per square meter. Video frames in graphics memory are usually RGB values, and usually from 0 to 255; the exact conversion from RGB to cd / m2 may be known for displays that have been calibrated.
[0015] The computer-implemented methods of the present disclosure comprise filtering the part of the video frame which is to be displayed on a user’s retinal peripheral area or on part of the peripheral vision area of a user if the contrast, which may be defined as a variation of luminance between two or more video frames, within at least a region of the peripheral vision area, overpass a threshold. The filtering is referred to as temporal filtering since the average luminance in time is filtered or reduced for the processed video frames which are
[0016] In the context of the present invention, the peripheral vision area of a field of view, FOV, region of the video frame is the area of the video frame which is to be projected on the peripheral vision area of the retina of a user viewing the VR content. The peripheral vision area depends on a user’s gaze. Such peripheral regions may be pre-defined or predetermined in, at least, the luminance sample matrix of the video frame. The eye gaze or user’s gaze may be measured or determined by eye gaze tracking, where eye gaze tracking may comprise measuring and analyzing the movements of the user’s eyes to determine where they are looking at. Eye gaze tracking may comprise the use of sensors and algorithms to track the movement of the eyes and determine where the user is looking at in relation to their environment. Technologies that may be used for eye gaze tracking may include infrared eye tracking, which uses infrared light to detect eye movements, and video-based eye tracking, which uses a camera to track eye movements. These technologies may be used in combination with other sensors and algorithms to provide a more complete understanding of a person’s gaze and attention.
[0017] Starting from a center of gaze or fixation point, i.e., the point at which a user's gaze is directed, the human field of view, FOV, may consist of foveal vision or sharp vision, which may span less than or 3° from the center of gaze, and peripheral vision area which spans from 3° onwards from the center of gaze. Peripheral vision area may also sometimes be referred to as “side vision” or “indirect vision.” The peripheral vision area may be understood as the edge of a human’s visual field. The visual field for humans may comprise about 170 degrees: 60 degrees diameter for central vision, and the region outside this diameter for peripheral vision area. If a person has a visual field of 20 degrees, the person can see things that are right in front of the person without moving his / her eyes from side to side, but the person cannot see anything on either side (peripheral vision area).
[0018] As said, filtering is performed by the presented methods if the contrast within at least a region of the peripheral vision area, overpass a threshold. The threshold may be predefined or may be adaptative to the plurality of video frames.
[0019] The variation of luminance may comprise variation of luminance in a same predetermined number of pixels of at least two video frames. The variation of luminance may comprise variation of luminance in a same predetermined area or section of at least two video frames. “A same predetermined number of pixels” or “a same predetermined area or section” refers to a spatial section in a first video frame and the corresponding spatial section in a second or in a third video frame, as will be illustrated in the examples below, for example a corner of the shown video frame or a frame or a geometrical crown of the shown video frame, or may also include the compete peripheral vision area. In the present disclosure, the luminance of a video frame is a photometric measure of the luminous intensity per unit area of light emitted from a VR display when the VR display shows the video frames. The filtering or smoothing temporal filtering is to be applied before the video frame is provided or displayed; the variation of luminance between at least two video frames may therefore be compared through information or metadata which may be included in the video frames processed by the processing unit. For example, the luminance of a pixel in a video frame may be correlated or related to the voltage indicated in the luminance sample matrix for that pixel to be displayed. In cases where the “at least a part”, or the region or section whose contrast is to be determined, comprises a plurality of pixels, the luminance of the section may be defined as a function of the arithmetic mean of the luminance of the plurality of pixels.
[0020] Filtering the part of the video frame which is to be displayed on a user’s retinal peripheral area comprises applying a smoothing temporal filter. Smoothing temporal filteringmay comprise reducing the luminance of at least one video frame of the two or more video frames for which the contrast is determined, for example the two or more video frames for which a variation of luminance within at least a region of the peripheral vision area between the video frames overpass the predetermined threshold. The advantage compared to other solutions which may include spatial filtering is that the reduction of luminance is determined as a result of a temporal determination of contrast, i.e., an interframe contrast. Some other solutions may comprise an intraframe analysis. Interframe or temporal determination of contrast to reduce the luminance of an area of at least one of the video frames which luminance is compared, so that an average temporal luminance of the compared video frames is reduced, reduces motion sickness -also referred to as periphery induced VR- sickness- in a user.
[0021] In some aspects, a VR content processing system is provided. The VR content processing system comprises a processor or processing unit configured to carry out the computer-implemented methods of the disclosure to reduce motion sickness. The processing unit and VR content processing system according to the present disclosure may be implemented by computing means, electronic means or a combination thereof. The computing means may be a set of instructions (e.g., a computer program) and the processing unit and the VR content processing system may comprise a memory and a processor, embodying said set of instructions stored in the memory and executable by the processor. These instructions may comprise functionality or functionalities to execute corresponding control methods such as, e.g., the ones described in other parts of the disclosure.
[0022] In some aspects, the disclosure presents a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the steps of a method according to the disclosure.
[0023] The above description is only an overview of the technical solution of the present disclosure. In order to understand the technical features of the present disclosure more clearly, it can be implemented in accordance with the content of the specification. In the following, the preferred embodiments are cited in conjunction with the drawings, and the detailed description is as follows.DESCRIPTION OF THE FIGURES
[0024] Figure 1 shows a block diagram of a virtual reality, VR, content processing system according to the present disclosure.
[0025] Figure 2 shows a block diagram of the VR content processing system according to the disclosure;
[0026] Figure 3 shows an example of a VR content processing system according to the disclosure;
[0027] Figure 4 shows an example of a VR content processing system according to the disclosure.
[0028] Figure 5 shows a representation of two linear functions, f 1 , f2, which reduce the temporal power spectra;
[0029] Figure 6 represents a block diagram of an example of the computer-implemented method according to the disclosure.
[0030] Figure 7 represents a block diagram of an example of the computer-implemented method according to the disclosure.DETAILED DESCRIPTION
[0031] In order to further illustrate the technical features and effects of the present disclosure to achieve the intended purpose of the disclosure, the specific implementation, structure, features and effects of the present disclosure will be described in detail below with reference to the accompanying drawings and preferred embodiments.
[0032] Figure 1 shows a block diagram of a virtual reality, VR, content processing system 1 comprising a processor or processing unit 12 configured to carry out a method according to the embodiments of the disclosure. The VR content processing system 1 may be used to provide VR content. The system 1 may be configured to provide a video stream, the video stream comprising a plurality 2 of video frames 3, 3’, 4 to one or more clients via a network, wherein the network may comprise point to point link or routed links via gateways or routers. The VR content processing system 1 may include a Video Server System. The Video Server System may be configured to provide the video stream in a wide variety of alternative video formats, and the video stream may include video frames configured for presentation to a user at a wide variety of frame rates. Typical frame rates are 30 frames per second, or 60 frames per second, or 90 frames per second or 120 frames per second. Higher or lower frame rates are included in alternative embodiments of the disclosure.
[0033] Figure 1 further shows the plurality 2 of video frames 3, 4 being received by the VR content processing system 1. The processing unit 12 is configured to receive the plurality of video frames 2, the plurality of video frames comprising at least two video frames 3, 4, 3’; and, upon a contrast over a threshold in at least a part 7, 7’ of a peripheral vision area 5, 5’ of a FOV region 6 between at least two video frames 3, 3’ of the plurality of video frames, filtering the at least a part of the peripheral vision area of at least one video frame 3, 3’, wherein the filtering comprises a smoothing temporal filtering.
[0034] In examples, a contrast is measured as a difference of luminance in the same pixelsof the same region or differences of luminance in the same region of two or more video frames. In some examples, the contrast is measured as an arithmetic mean of two or more differences of luminance in the same pixels of the same region or differences of luminance in the same region of three or more video frames. Two or more differences may be stored in a temporal window of time and a mean or any other central tendency applicable to onedimensional data may be calculated. For example, the measure may include one or more of arithmetic mean or mean, a median, a mode, a generalized mean, a geometric mean, a weighted arithmetic mean or other.
[0035] In examples, the at least a part 7, 7’ of a peripheral vision area 5, 5’ of a FOV may be defined by one or more parameters, for example the at least a part 7, 7’ may be an angle or a range of angles of vision to be filtered, for example all the peripheral area from 60° to 120° of a second video frame of two consecutive video frames may be filtered. As seen in figure 1 , a “same number of pixels” or “a same area or section” refers to the part 7 in one video frame 3 and the same region 7’ in another video frame 3’. In some examples at least a part 7, 7’ comprises the complete peripheral vision area 5, 5’.
[0036] Some of the temporal filters which may be applied to one or more video frames may comprise a kernel or temporal convolution in at least part 7, 7’ of one or more of the video frames 3, 3’, 4. The convolution or kernel or smoothing filter results in a reduced displayed luminance over a time, for example a time comprising two or three video frames, relative to an original input luminance that the processed video frames are intended to have over the same time. The at least a part 7, 7’ is comprised within the peripheral vision area 5, 5’ or may represent the peripheral vision area 5, 5’. Advantageously, by filtering one or more of the video frames, high temporal contrast in the peripheral vision area is reduced, reducing thereby motion sickness or negative experience by a final viewer or user. For example, taking the illustration of figure 1 , where two consecutive video frames 3, 3’ are shown, the second video frame 3’ may be the video frame on which the smoothing filter is to be applied. The part 7’ which is to be filtered is represented as a triangular portion of the peripheral vision area 5’. The part 7’ which is to be filtered may also be an area covering, for example, a section, or crown section, resulting from the extraction of the image covering an angle (6’- A) of the video frame 3’, where 6’ represents the FOV in degrees and A is an angle to which the filter is not to be applied. The line L in figure 1 represents the line from a user’s eye to the center of gaze C. The angles are assumed to be given at a predetermined distance of eye relief.
[0037] The VR content processing system 1 may be a head mounted display, for example a head mounted display comprising Dual Mini LED LCD 2880 x 2720 px per eye displays with brightness calibrated to 150 NIT, with colors calibrated with coverage of 99% sRGB, with refresh rate: 90 Hz, or 100Hz, with Field of View, FOV, horizontal 115° and diagonal:134° at 12 mm eye relief.
[0038] Figure 2 shows a block diagram of the VR content processing system 1 further comprising an input interface 13 configured to receive tracking from a gaze tracking detector 14. The input interface 13 is shown in communication with the processing unit 12. Figure 2 shows the plurality 2 of video frames 3, 4, 3’ or stream, which are received by the system 1 . A video input interface may be provided in the system 1 to receive the plurality 2 of video frames forming a stream. The video input interface may comprise any of USB interface or Bluetooth or the input may be embedded or programmed by a program residing in a graphical processing unit GPU, also referred to as shader; the shader may be modified by a programming code, for example in C language or similar, so that received video frames and a gaze input are processed by the GPU. Figure 2 further shows a VR display 15 on which the VR content processed by the VR content processing system is displayed. The VR content comprises the plurality 2 of incoming video frames 3, 4, 3’ wherein some of the video frames have been filtered applying a smoothing temporal filter. The VR content processing system 1 , or in short, system 1 , may comprise connectivity interfaces such as USB or USB-C cables, one or more display ports, and may comprise, integrated, a positional tracking system or gaze tracking detector 14 which may provide tracking at a rate of 200 Hz with sub-degree accuracy. When the VR content processing system 1 receives tracking, from the gaze tracking detector 14, of a gaze of a user viewing the VR content, the gaze defines the FOV region. In such cases, the VR content processing system may determine the peripheral vision area based on the tracking, for example by considering that the peripheral vision area is separated 60 degrees from the center of gaze of the user.
[0039] The methods of the present disclosure may be executed by the processing unit or processing unit 12. The system 1 may comprise a non-transitory storage medium comprising instructions or a computer program product comprising instructions which, when being executed by the processing unit 12, cause said processing unit to perform the steps of the method of the disclosure. The computer program product may be also stored in a graphics memory of an integrated display of the system 1. A plugin may be used for integrating the computer program instructions into the VR content processing system 1.
[0040] The tracking or an updated tracking from a gaze tracking detector 14 may be obtained periodically at every gaze_time_cycle time units, for example every 8 milliseconds, ms. The tracking may correspond to a gaze of a user viewing the VR content. The system 1 may then, iteratively, determine a peripheral vision area for each video frame of the plurality of video frames, based on the updated tracking. The gaze_time_cycle may be greater or less or equal to an interframe_separation time units, wherein the interframe_separation is the separation in time or temporal separation between two received consecutive video frames in the plurality 2 of video frames.
[0041] Two video frames in the plurality 2 of video frames may be separated 8 ms or more, for example 10 ms or 11 ms for a frame rate of approximately 90 fps. The processing unit may execute the methods according to the present disclosure every time a video frame is received from the stream of video frames. The VR content processing system 1 may comprise transitory or non-transitory storage means in which the value or tracking received from the gaze tracking detector may be stored and updated at each input from the gaze tracking detector. The processing unit may execute the method using an available tracking in the storage means in the exact moment of execution of the method. The determination of a peripheral vision area for each video frame of the plurality of video frames, based on the updated tracking, is therefore performed using updated tracking.
[0042] In some examples, receiving the tracking comprises receiving a center of gaze, and the determining the peripheral vision area is performed by determining an area separated N degrees or more from the center of gaze, or fixation point, of a user’s gaze where N is a predefined parameter. A peripheral zone or peripheral vision area outside the center of gaze may be determined as 60 degrees separated from the center of gaze. In other words, the peripheral vision area may be determined by a circle 60° in radius or 120° in diameter, centered around the fixation point or the center of gaze. Receiving a center of gaze may comprise receiving coordinates of a point in a video frame to be displayed.
[0043] In some examples, the method implemented by the processing unit 12 may comprise determining the contrast between two consecutive video frames of the plurality of video frames in at least a part of the peripheral of the two consecutive video frames, wherein the two consecutive video frames comprise a first video frame and a second video frame and wherein the second video frame is set to be displayed after the first video frame. The method may further comprise, upon the contrast over a threshold, filtering at least part of the peripheral area of the FOV of the second video frame. Upon reception of a further video frame, or third video frame, the processing unit 12 may implement the methods of the disclosure by determining the contrast between at least a part of the filtered second video frame and the same part(s) of the third video frame. The method may therefore determine a contrast between two video frames, first and second video frame, decide to filter, or to not filter, the second video frame, and use the filtered, or the not filtered, second video frame to be the first video frame of the following pair of video frames to determine a new contrast. The third video frame will be the filtered, or not filtered, video frame of the pair comprising the second and the third videoframes.
[0044] Regarding the threshold, in some examples is received as an external input before or during the processing of VR content. For example, a user may experience motion sickness during the consumption of VR content processed by the VR content processing system and provide input to the VR content processing system by a button or any otherinput so that the threshold may be modified in a predetermined quantity or progressively until the user stops experiencing such motion sickness. In complementary or alternative examples, the method may -or may also- determine the threshold based on a standard deviation, STD, of the mean variation luminance between video frames displayed during the processing of the VR content. In complementary or alternative examples, the threshold may -or may also-be determined based on a different predefined criteria before or during the processing of VR content. In examples, the threshold may remain invariable during the processing of VR content or the threshold may vary during the processing of VR content.
[0045] Figure 3 shows an example of a VR content processing system 1 receiving a plurality 2 of video frames, where two consecutive video frames, first video frame 3 and second video frame 4 are being received. The first video frame and the second video frame are processed. The processing unit 12 is schematically shown comprising a comparison module C - schematic illustrative and non-limitative module- which implements the step of determining a contrast of at least part of a peripheral vision area of a FOV region between two video frames. The comparison module C in figure 3 determines the contrast by comparing the luminance L1 of at least a region of the peripheral vision area of the first video frame 3 with the luminance L2 of the same corresponding region of the second video frame 4. In the example shown the comparison measures the variation of luminance. The processed first video frame 3 is supplied to a buffer 16, which may be optional, and, if the variation of luminance is over the threshold Th, then the second video frame 4 is filtered by reducing the value of luminance to be displayed. After the filtering, the processed second video frame 4 is supplied to the buffer 16. If, alternatively, the variation of luminance is not over the threshold Th, then the second video frame 4 is supplied to the buffer 16 without filtering. After the processing, which may or may not comprise the filtering, the processed second video frame 4 is supplied to an entity, the entity being the buffer 16. The figure 3 shows supplying the VR content as the processed received video frames to an entity, the entity being the buffer 16. The buffer 16 may feed a display. The figure 3 shows the VR content processing system 1 comprising the display. The VR content processing system 1 of figure 3 may be a head mounted display, HMD. The VR content may also be supplied directly to a different entity, for example a display or HMD. As shown in figure 3, a switch 17 between a video frame connection 18, for example a cable configured to transport video frames, and the filter 19, represents a schematic illustrative and non-limitative module implementing the step of the method where, depending on the result of the comparison module C, the incoming video frames 3 and / or 4 and / or 4’ are directly provided to the buffer 16 without filtering or after being filtered by the filter 19. The comparison module C may comprise a processing unit or module or comparator or device that compares two voltages or currents and outputs a digital signal indicating which is larger, or may besoftware programmed, for example, a section of programming code, or may be a component of a Field Programmable Gate Array, i.e., FPGA. The switch may comprise an electronic switch, for example a diode or a transistor, controlled by an active electronic component or device, or may be software programmed or be part of a FPGA.
[0046] Figure 4 shows the VR content processing system 1 of figure 3, where a time window has advanced. The figure 4 shows the system 1 receiving the plurality 2 of video frames, where two consecutive video frames are received. Now, the current first video frame 4 is the second video frame of the preceding time window represented in figure 3, and the current second video frame 4’ is the video frame which follows the video frame 4 in the plurality 2 of incoming video frames. The current first video frame 4 and the current second video frame 4’ are processed. The processing unit 12 determines the contrast by comparing the luminance L3 of at least a region of the peripheral vision area of the current first video frame 4, with the luminance L4 of the same corresponding region of the current second video frame 4’. The processed current first video frame 4 is provided to the optional buffer 16 -which may be optional- and, if the variation of luminance is over the threshold Th, then the current second video frame 4’ is filtered by reducing the value of luminance to be displayed. If, alternatively, the variation of luminance is not over the threshold Th, then the current second video frame 4’ is supplied to the buffer 16 without filtering. After the processing, which may or may not comprise the filtering, the processed current second video frame 4’ is supplied to the buffer 16. The figure 4 shows providing or supplying the VR content as the processed received video frames to the buffer 16. The VR content may also be supplied directly to a display or to the HMD.
[0047] In some examples, the luminance values may be represented in a luminance matrix. One of the examples may comprise the following luminance matrices representing the luminance to be rendered or displayed on a screen:
[0048] First luminance matrix in RGB values at frame t=1231 252 233 246 243 249 249 250 245 237248 233 245 252 231 237 253 237 246 242247 235 253 238 248 241 243 232 243 239251 251 235 254 238 246 241 236 237 245251 239 237 253 248 249 252 243 231 231251 252 242 251 243 254 252 232 238 237244 252 237 249 248 233 243 244 239 253247 236 237 239 234 248 250 250 251 241232 248 244 237 246 238 234 233 252 237245 254 230 253 232 254 233 235 233 254
[0049] Second luminance matrix in RGB values at frame t=2225 221 214 210 222 219 242 211 217 240219 242 210 226 214 243 230 236 239 213222 210 232 221 245 237 216 213 237 239212 240 221 211 217 244 222 215 228 216245 213 242 220 221 231 223 227 237 242222 229 226 234 223 236 233 212 242 221211 224 231 230 232 229 211 234 242 234226 243 244 217 216 215 238 229 213 244233 229 227 234 218 221 214 232 227 215235 242 243 240 234 229 245 241 225 220
[0050] If the absolute value of the contrast in a peripheral vision area represented by a peripheral area of the luminance matrices is greater than a threshold, for example, greater than 10 RGB, then a filter is applied.
[0051] In some examples, the comparison is performed pixel by pixel in the peripheral vision area and a filter is applied to the luminance of each pixel of the peripheral vision area of the video frames represented by the luminance matrix. For example, the following matrix may represent TRUE values for those comparisons resulting in a contrast pixel by pixel greater than 10 between the first luminance matrix (Lfirst) and the second luminance matrix(Lsecond); as follows: contrast = abs(Lsecond_pixelab - Lfirst_pixelab) >10; with a referring to a specific position in rows and b referring to a position in columns of the matrices:FALSE TRUE TRUE TRUE TRUE TRUE FALSE TRUE TRUE FALSE TRUE FALSE TRUE TRUE TRUE FALSE TRUE FALSE FALSE TRUE TRUE TRUE TRUE TRUE FALSE FALSE TRUE TRUE FALSE FALSE TRUE TRUE TRUE TRUE TRUE FALSE TRUE TRUE FALSE TRUE FALSE TRUE FALSE TRUE TRUE TRUE TRUE TRUE FALSE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE FALSE TRUE TRUE TRUE FALSE TRUE TRUE FALSE TRUE FALSE FALSE TRUE TRUE FALSE FALSE TRUE TRUE TRUE TRUE TRUE TRUE FALSE FALSE TRUE TRUE FALSE TRUE TRUE TRUE FALSE TRUE TRUE FALSE TRUE TRUE TRUE FALSE TRUE TRUE FALSE FALSE TRUE
[0052] When the value is TRUE, and if the pixel corresponds to a pixel in the peripheral vision area of the video frame, then the luminance value of the corresponding pixel in Lsecondab is reduced so that abs(Lsecond_filteredab-Lfirstab) < 10 before the frame is rendered or displayed on a screen.
[0053] In other examples the comparison and calculation of contrast = abs(Lsecondab- Lfirstab) is only performed in the pixels of the peripheral vision area and not in the whole matrix.
[0054] In other examples the luminance values of the peripheral vision area in Lfirst and Lsecond are averaged, and if the difference between the averages of luminance between the peripheral vision areas of the first and second matrices is greater than 10, then all the luminance values of the pixels in the peripheral vision area of, for example the second matrix, are reduced or filtered, as follows: if (abs(avg_Lsecond-avg_Lfirst)>10) then filter luminance of all pixels of peripheral vision area of Lsecond.
[0055] Figure 5 shows a representation of two linear functions in log-log space of the power spectrum as a function of temporal frequency, f1 , f2, which includes a number of line slopes that are consistent with natural images: defined as p = 1 / (co ?), where co represents the temporal frequency or contrast between video frames, and ft represents the slope of either f1 or f2. The representation shows the filtering which may be applied to an incoming contrast. As the figure 5 shows, greater values of contrast between incoming video frames may receive more attenuation or reduction of contrast, and ft therefore grows with co. The power p is reduced in a greater scale for greater temporal frequencies co.
[0056] A human’s visual system is used to natural temporal statistics whose Temporal power spectra p = 1 / (co / 3), where usually is 1 < / ? < 3. Some stimuli on the periphery induced VR-sickness is identified for ft values under 1. As visual content or VR content in VR systems are not bounded by natural statistics -like in artistic creations- video frames may present ft values under 1 and over 3. These f1 and f2 show that modifying the contrast, by for example reducing luminance, ft is modified to lower the power (or contrast) in high temporal frequencies. High temporal frequencies may be temporal frequencies co over 5-6 Hz. In the figure 5 f1 corresponds to / ? = 3 and f2 corresponds to / ? =1. An objective of a VR content processing system according to the disclosure may be to provide VR content whose power remains between both f1 and f2, with 1 < / ? < 3.
[0057] Figure 6 represents a block diagram of an example of the computer-implemented method 600 according to the disclosure. In block 601 , the computer-implemented method 600 comprises determining a contrast of at least a part of a peripheral vision area of a FOV region between at least two video frames of a plurality of video frames; and, in block 602, upon the contrast over a threshold, filtering the at least a part of the peripheral vision area of at least one video frame, wherein the filtering comprises a smoothing temporal filtering.
[0058] Figure 7 represents a block diagram of an example of the computer-implemented method 600 according to the figure 6. The method 600 of figure 7 further comprises, in block 701 , receiving a plurality of video frames, the plurality of video frames comprising at least two video frames. In block 702, the computer-implemented method 600 further comprises providing VR content, VR content comprising the processed video frames.
[0059] The processing units of the present disclosure may comprise or may be implemented by electronic means, computing means or a combination of them, that is, electronic or computing means may be used interchangeably so that a part of the described means may be electronic means and the other part may be computing means, or all described means may be electronic means or all described means may be computing means. Examples of a processing unit comprising electronic means may comprise aprogrammable electronic device such as a Complex Programmable Logic Device, CPLD, a Field Programmable Gate Array, FPGA or an Application-Specific Integrated Circuit, ASIC.
[0060] Examples of a processing unit comprising computing means may comprise a computer system, which may comprise a memory and a processor, the memory being adapted to store a set of computer program instructions, and the processor being adapted to execute these instructions stored in the memory or non- transitory readable storage medium. The memory may be comprised in the processing unit, for example an EEPROM, or may be external, for example, data storage means such as magnetic disks, e.g., hard disks, optical disks, e.g., DVD or CD, memory cards, flash memory, e.g., pen drives, or solid-state drives, SSD based on RAM, based on flash, etc. A set of computer program instructions may be executable by the processor, such as a computer program, may be stored in a physical storage means, or may be carried by a carrier wave, or by a carrier medium. The processing unit may be any entity or device capable of carrying the program, such as electrical or optical, which can be transmitted via electrical or optical cable or by radio or other means.
[0061] The computer program may be in the form of source code, object code, a code intermediate source and object code such as in partially compiled form, or in any other form suitable for use in the implementation of the method. The carrier may be any entity or device capable of carrying the computer program.
[0062] The computer program may be embodied on a storage medium (for example, a CD- ROM, a DVD, a USB drive, a computer memory or a read-only memory) or carried on a carrier signal (for example, on an electrical or optical carrier signal).
[0063] All of the features disclosed in this disclosure, including any accompanying claims, abstract and drawings, and / or all of the steps of any computer-implemented method or process, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.
[0064] Each feature disclosed in this specification, including any accompanying claims, abstract and drawings, may be replaced by alternative features serving the same, equivalent or similar purpose, unless expressly stated otherwise. Thus, unless expressly stated otherwise, each feature disclosed is one example only of a generic series of equivalent or similar features.
Claims
CLAIMS1. A computer-implemented method for processing a virtual reality, VR, content for reducing motion sickness, the VR content comprising a plurality (2) of video frames, the plurality of video frames comprising at least two video frames (3, 4), the processing comprising: determining a contrast of at least part (7, 7’) of a peripheral vision area (5, 5’) of a field of view, FOV, region (6) between at least two video frames of the plurality of video frames; and upon the contrast over a threshold, filtering the at least part of the peripheral vision area of at least one video frame, wherein the filtering comprises a smoothing temporal filtering.
2. The computer-implemented method of claim 1 , further comprising- receiving a tracking, from a gaze tracking detector, of a gaze of a user viewing the VR content, wherein the gaze defines the FOV region (6); and- determining the peripheral vision area (5, 5’) based on the tracking.
3. The computer-implemented method of any one of claim 1 or claim 2 further comprising- receiving, from a gaze tracking detector, an updated tracking at a periodicity of gaze_time_cycle time units; and- iteratively determining a peripheral vision area for each video frame of the plurality (2) of video frames, based on the updated tracking.
4. The computer-implemented method of claim 2 or 3 wherein receiving the tracking comprises receiving a center of gaze, and the determining the peripheral vision area is performed by determining an area separated N degrees or more from the center of gaze, where N is a predefined parameter.
5. The computer-implemented method of any one of claims 3 to 4 wherein a temporal separation between at least two received consecutive video frames of the plurality (2) of video frames is interframe_separation time units and wherein the temporal separation interframe_separation is equal to or greater than a gaze_time_cycle time units.
6. The computer-implemented method of any one of claims 1 to 5: wherein determining a contrast comprises determining the contrast between two consecutive video frames of the plurality (2) of video frames in at least part of the peripheralvision area of the two consecutive video frames; wherein the two consecutive video frames comprise a first video frame and a second video frame and wherein the second video frame is set to be displayed after the first video frame; and wherein filtering comprises filtering at least part of the peripheral vision area of the FOV of the second video frame.
7. The computer-implemented method of any one of claims 1 to 6 wherein the computer- implemented method further comprises any one or more of the following:- receiving the threshold as an external input before or during processing VR content; and / or- determining the threshold based on a standard deviation, STD, of a mean variation of luminance between a set of video frames displayed during processing VR content; and / or- determining the threshold based on a different predefined criteria before processing VR content.
8. The computer-implemented method of any one of claims 1 to 7 wherein the threshold remains invariable during processing VR content or wherein the threshold varies during processing VR content.
9. The computer-implemented method of any one of claims 1 to 8 further comprising, before processing the VR content, receiving the plurality (2) of video frames; and / or after processing the VR content, providing processed VR content, wherein the processed VR content comprises the processed video frames.
10. A virtual reality, VR, content processing system (1) comprising a processor (12) configured to carry out the computer-implemented method of any one of claims 1 to 9.
11. The VR content processing system (1) of claim 10 further comprising an input interface (13) configured to receive tracking from a gaze tracking detector (14).
12. The VR content processing system (1) of claim 10 or 11 further comprising a video input interface configured to receive the plurality (2) of video frames.
13. The VR content processing system (1) of any of claims 10 to 12 further comprising a VR display.
14. The VR content processing system (1) of any one of claims 10 to 13, the VR content processing system being a head mounted display, HMD.
15. A computer program product comprising instructions which, when being executed by a processing unit, cause said processing unit to perform the steps of the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Display device and method for image processing
US20190244369A1
Selective peripheral vision filtering in a foveated rendering system
US20190384381A1
EP23383208A