A Real-Time Light Field Stereo Rendering Method for Immersive Head-Mounted Displays

By uniformly sampling and encoding the original light field data, combined with WebGL and Three.js technology, vertex and fragment shaders are designed, real-time light field three-dimensional rendering of immersive head-mounted displays is realized, solving the problem of insufficient rendering performance in traditional technologies, and improving real-time efficiency and cross-platformity.

CN119600239BActive Publication Date: 2025-07-01QUFU NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411704481.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-07-01
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Traditional light field rendering technology has poor real-time rendering performance in immersive environments, making it difficult to meet the needs of immersive head-mounted displays.

Method used

A real-time light field stereoscopic rendering method for immersive head-mounted displays is adopted to generate a multi-viewpoint image pseudo-video sequence by uniformly sampling the original light field data and YUV color space transformation, and compress and upload it using the H.264 video encoder. Use WebGL and Three.js to decompress to generate three-dimensional texture maps, design vertex and fragment shaders, and combine three-dimensional maps for rendering to realize the rendering of binocular viewpoint images.

Benefits of technology

It improves the real-time efficiency of light field stereo rendering, meets the needs of immersive head-mounted displays, and can run on mid- and low-end computers, has good cross-platformity, and reduces the requirements for high-performance platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600239B_ABST
    Figure CN119600239B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time light field stereoscopic rendering method for immersive head-mounted displays, belonging to the technical field of light field rendering, including: uniformly sampling the original light field data in the angular space, transforming the original data to the YUV color space, and generating a multi-viewpoint image pseudo-video sequence by zigzag scanning; compressing the image pseudo-video sequence with an H.264 video encoder and uploading it to the server; the client decompresses the light field data using WebGL and Three.js to generate a three-dimensional texture map; designing a vertex shader and a fragment shader; using a plane in the three-dimensional space as the proxy geometry of the scene, taking this geometric data as the input of the rasterization rendering pipeline, and using the vertex shader and the fragment shader, combined with the three-dimensional map, to render the binocular viewpoint image. By adopting the above method, the present invention improves the real-time efficiency of light field stereoscopic rendering and meets the requirements of immersive head-mounted displays.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of light field rendering, and in particular to a real-time light field stereoscopic rendering method for immersive head-mounted displays. Background Art

[0002] Light field rendering is an advanced computer graphics technology that simulates the behavior characteristics of light in the real world to generate realistic three-dimensional scenes. Its core lies in capturing and simulating the transmission of light in three-dimensional space, including direction, intensity, and color. Traditional research on light field rendering technology mainly focuses on view synthesis for a single viewpoint, using graphics rendering algorithms and neural network models to calculate or infer the images of the target viewpoints. To generate high-quality images, these technologies often rely on high-performance computing resources or even customized computing platforms to display new views, which not only increases costs but also has poor real-time rendering performance, restricting the wide application of light field technology, especially light field stereoscopic rendering in immersive environments.

[0003] Therefore, a light field stereoscopic rendering method for immersive environments is needed. Summary of the Invention

[0004] The object of the present invention is to provide a real-time light field stereoscopic rendering method for immersive head-mounted displays, which improves the real-time efficiency of light field rendering.

[0005] To achieve the above object, the present invention provides a real-time light field stereoscopic rendering method for immersive head-mounted displays, and the steps include:

[0006] S1. Uniformly sample the original light field data in the angular space, transform the original data to the YUV color space, and generate a pseudo-video sequence of multi-viewpoint images by zigzag scanning;

[0007] S2. Compress the pseudo-video sequence of images using an H.264 video encoder and upload it to the server;

[0008] S3. Load the light field data transmitted to the server by the client and decompress the light field data using WebGL and Three.js to generate a three-dimensional texture map;

[0009] S4. Design a vertex shader and a fragment shader, and the fragment shader is designed based on a refocusing algorithm in the spatial domain;

[0010] S5. Use a plane in three-dimensional space as a proxy geometry of the scene, use this geometry data as the input of the rasterization rendering pipeline, and use the vertex shader and the fragment shader to jointly render the binocular viewpoint images with the three-dimensional texture map.

[0011] Preferably, the functions of the vertex shader in step S4 include clipping vertex coordinates and calculating the starting point and direction of light rays. The formula for clipping vertex coordinates is as follows:

[0012] V clip = M projection * M view * M model * V;

[0013] In the formula, V = (x, y, z, 1) represents the local coordinates of the vertex, and V clip = (x, y, z, w) represents the transformed clipping coordinates. M projection represents the projection matrix of the virtual view point camera, which is defined by the field of view angle and the far and near planes of the projection. M view represents the view matrix, which is used to convert the world coordinates of the three-dimensional plane into camera coordinates. M model represents the model matrix, which is used to convert the local coordinates of the three-dimensional plane into world coordinates;

[0014] The starting point of the light ray is calculated according to the texture coordinates of the three-dimensional plane, and the direction of the light ray is calculated according to the position of the virtual camera and the local vertex coordinates of the three-dimensional plane. The formula is as follows:

[0015] r o = [u, v];

[0016] r d = normalize(p camera - p plane );

[0017] In the formula, r d represents the direction of the light ray, [u, v] represents the texture coordinates of the three-dimensional plane, r o represents the starting point of the light ray, p camera represents the position of the virtual camera, and p plane represents the local vertex coordinates of the three-dimensional plane.

[0018] Preferably, the fragment shader designed based on the refocusing algorithm in the spatial domain in step S4 includes:

[0019]

[0020] In the formula, i and j represent the coordinates of the imaging plane, M1M2 represents the sensor size, a represents the size measure of the aperture, f represents the focal length measure, F(x, y) represents the color value [R, G, B] of the corresponding sub-aperture image of the light field at the pixel coordinate x at the y coordinate, d represents the block distance vector between the current pixel point and the target pixel point, x screen , y screen represent the horizontal and vertical coordinates of the screen pixel, d x , d yRespectively represent the components of the distance vector d in the horizontal and vertical directions.

[0021] Preferably, step S5 includes:

[0022] S51. Use the plane in the three-dimensional space as the geometric proxy of the scene, and set the size of the plane according to the angular resolution of the light field;

[0023] S52. Input the plane in the three-dimensional space into the rasterization rendering pipeline, use the vertex shader to clip the vertex coordinates, and calculate the direction of the light ray;

[0024] S53. Convert the clipped coordinates into normalized device coordinates through perspective division;

[0025] S54. Use the fragment shader to convert the device coordinates into the final screen coordinates according to the current pose of the virtual camera, and generate the target view according to the refocused aperture and focal length parameters;

[0026] S55. Render the binocular viewpoint images according to the positions of the binocular viewpoints.

[0027] Preferably, the formula for converting the clipped coordinates into normalized device coordinates in step S53 is:

[0028]

[0029] In the formula, V NDC represents the device coordinates.

[0030] Preferably, the formula for converting the device coordinates into the final screen coordinates in step S54 is:

[0031]

[0032] In the formula, V screen represents the screen coordinates, width and height respectively represent the width and height of the viewport, x offset and y offset represent the offset of the lower left corner of the viewport on the screen.

[0033] Therefore, the real-time light field stereoscopic rendering method for an immersive head-mounted display of the present invention has the following beneficial effects:

[0034] (1) By controlling the binocular rendering parameters through the refocusing algorithm based on the spatial domain, the control of the scene depth, virtual camera focal length, and aperture is realized in the immersive head-mounted display, providing a high-quality visual experience and interactivity;

[0035] (2) The real-time efficiency of the light field stereoscopic rendering is improved, meeting the requirements of the immersive head-mounted display;

[0036] (3) The light field stereo rendering of the present invention can run on mid - to - low - end computers and has good cross - platform performance. It only requires a browser that supports WebGL, reducing the requirements of traditional light field calculations for high - performance platforms.

[0037] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings

[0038] Figure 1 is the flowchart of the method according to the embodiment of the present invention;

[0039] Figure 2 is the schematic diagram of the refocusing algorithm according to the embodiment of the present invention;

[0040] Figure 3 is the binocular structure diagram according to the embodiment of the present invention. Detailed Embodiments

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention claimed, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0042] Embodiment

[0043] Referring to Figure 1 , the present invention provides a real - time light field stereo rendering method for immersive head - mounted displays, and the steps include:

[0044] S1. Uniformly sample the original light field data in the angular space, transform the original data to the YUV color space, and generate a multi - view image pseudo - video sequence using zig - zag scanning.

[0045] The original light field data refers to the scene images obtained by using light field data acquisition devices such as array cameras. The main encoding methods are: multi-viewpoint image sequences and macro-pixel images. The multi-viewpoint images can be captured by a camera array located in a two-dimensional plane, and each image is a traditional perspective projection image. Different from the multi-viewpoint images, the macro-pixel images encode the angular information of the light field within the pixels, that is, the macro-pixel blocks are used to represent the projections of the same point in the three-dimensional space under different camera viewpoints. This indicates that there is a fixed conversion relationship between the multi-viewpoint images and the macro-pixel images. Assume that s and t represent the pixel coordinates of each sub-view in the multi-viewpoint image, and u and v represent the position coordinates of the cameras corresponding to the sub-views in the multi-viewpoint image on the two-dimensional plane. Then each pixel block of the macro-pixel image can be represented as extracting u*v pixels from the multi-viewpoint image with fixed s and t coordinates. Conversely, if a macro-pixel image is given, only by traversing each pixel block of the image and extracting pixels from the same coordinates s and t of each pixel block, the multi-viewpoint image can be generated.

[0046] S2. Compress the image pseudo-video sequence using the H.264 video encoder and upload it to the server. Based on the H.264 video encoder for compression, the compression rate can reach 98.76%. The H.264 video codec is an international standard widely used in video compression and network transmission. It is jointly developed by the International Telecommunication Union (ITU-T) and the International Organization for Standardization (ISO) to provide high compression efficiency while maintaining video quality.

[0047] S3. Load the light field data transmitted to the server through the client, and use WebGL and Three.js to decompress the light field data to generate a three-dimensional texture map. WebGL and Three.js libraries: WebGL is a technology for rendering interactive three-dimensional graphics in web browsers. Three.js is a JavaScript library based on WebGL for creating and displaying three-dimensional graphics.

[0048] S4. Design vertex shaders and fragment shaders. The vertex shader processes the data of each vertex, including transforming the vertex position to the clip space, applying lighting, etc. In the fragment shading stage, the fragment shader is mainly used to process the color and depth values of each pixel, and apply effects such as texture, lighting, and blending.

[0049] The functions of the vertex shader include clipping the vertex coordinates and calculating the starting point and direction of the light ray. The formula for clipping the vertex coordinates is as follows:

[0050] V clip =M projection *M view *M model *V (1)

[0051] Wherein, V = (x, y, z, 1) represents the local coordinates of the vertex, and V clip = (x, y, z, w) represents the transformed clip coordinates, and M projection represents the projection matrix of the virtual view point camera, which is defined by the field of view angle and the far and near planes of the projection, and M view represents the view matrix, which is used to convert the world coordinates of the three-dimensional plane into camera coordinates, and M model represents the model matrix, which is used to convert the local coordinates of the three-dimensional plane into world coordinates;

[0052] The starting point of the ray is calculated according to the texture coordinates of the three-dimensional plane, and the direction of the ray is calculated according to the position of the virtual camera and the local vertex coordinates of the three-dimensional plane. The formula is as follows:

[0053] r o = [u, v] (2)

[0054] r d = normalize(p camera - p plane ) (3)

[0055] Wherein, r d represents the direction of the ray, [u, v] represents the texture coordinates of the three-dimensional plane, and r o represents the starting point of the ray, p camera represents the position of the virtual camera, and p plane represents the local vertex coordinates of the three-dimensional plane.

[0056] The fragment shader is designed based on the refocusing algorithm in the spatial domain.

[0057] Referring to Figure 2 , the basic principle model of refocusing is described. Taking the two-dimensional projection of the light field and changing the focal length from F to F' as an example. Given a refocused image I(x, y), it can be calculated from the sub-aperture image L F (u, v) according to Equation (3).

[0058] I(x, y) = ∫∫L F (u(α - 1) + x, v(α - 1) + y)dudv (3)

[0059] Wherein, α is the refocusing factor, which represents the ratio of the original focal plane depth to the new focal plane depth. For the convenience of actual image calculation, the uv coordinates are converted into xy coordinates, and the new refocusing formula can be expressed as Equation (4).

[0060]

[0061] where, ij are the coordinates of the sensor plane, kl are the coordinates of the microlens array, M1M2 is the sensor size, k ∈ [1, N1], l ∈ [1, N2], N1N2 is the microlens array size, S ij (k, l) represents the sub-aperture image extracted from the ij coordinates. Therefore, the refocusing algorithm formula based on the spatial domain is as follows:

[0062]

[0063] In the formula, i, j represent the coordinates of the imaging plane, M1M2 represents the sensor size, a represents the size measure of the aperture, f represents the focal length measure, F(x, y) represents the color value [R, G, B] of the corresponding light field sub-aperture image at the pixel coordinate x at the y coordinate, d represents the block distance vector between the current pixel point and the target pixel point, x screen , y screen represent the horizontal and vertical coordinates of the screen pixels, d x , d y respectively represent the components of the distance vector d in the horizontal and vertical directions.

[0064] S5. Use the plane in the three-dimensional space as the proxy geometry of the scene, take this geometric data as the input of the rasterization rendering pipeline, and adopt vertex shaders and fragment shaders, combined with three-dimensional texture mapping, to render the binocular viewpoint images. Specifically, it includes:

[0065] S51. Use the plane in the three-dimensional space as the geometric proxy of the scene, and set the size of the plane according to the angular resolution of the light field.

[0066] S52. Input the plane in the three-dimensional space into the rasterization rendering pipeline, use formula (1) in the vertex shader to clip the vertex coordinates, and use formulas (2) and (3) to calculate the direction of the light ray. The range of this texture coordinate is [0, 1], the texture coordinate of the geometric center of the plane is (0.5, 0.5), and the texture coordinates of the four corner vertices are (0, 0), (1, 0), (1, 1), (0, 1) respectively, so as to construct the light ray parameters (r o , r d ).

[0067] S53. Convert the clipped coordinates into normalized device coordinates through formula (7) perspective division, and its formula is:

[0068]

[0069] In the formula, V NDC represents the device coordinates.

[0070] S54. Use equations (5) and (6) in the fragment shader to convert device coordinates to the final screen coordinates according to the current pose of the virtual camera, and generate the target view according to the aperture and focal length parameters of refocusing.

[0071] When converting device coordinates to the final screen coordinates, since the range of NDC is usually [-1, 1], the NDC coordinates are converted to the final screen coordinates through the viewport transformation in equation (8), and its formula is:

[0072]

[0073] In the formula, V screen represents the screen coordinates, width and height represent the width and height of the viewport respectively, and x offset and y offset represent the offset of the lower left corner of the viewport on the screen.

[0074] S55. Render the binocular viewpoint images according to the positions of the binocular viewpoints.

[0075] Through steps S51 - S54, a single viewpoint image can be rendered. However, an immersive head-mounted display has binocular characteristics and needs to provide views for the left and right eyes respectively to generate a stereoscopic immersive feeling. Therefore, the above method is extended to the binocular stereo rendering scenario, and in each rendering frame, a viewport and a rendering matrix are set for each view (corresponding to the left and right eyes). The binocular camera structure is as Figure 3 shown.

[0076] For each frame of image, it is rendered twice from the positions of the left and right cameras according to the refocusing rendering algorithm to generate left and right eye images with horizontal parallax, allowing the viewer's left and right eyes to view the two images simultaneously, thereby generating binocular stereoscopic feeling. Among them, the distance between the left and right cameras is set as the average interpupillary distance (IPD) of the human eye to generate a stereoscopic view that conforms to the human eye.

[0077] To verify the effectiveness of the method of the present invention, the present invention is compared with the traditional light field rendering method. Specifically: the average network transmission delay of the present invention is 43 ms, and the existing technology is 669 ms, which is about 20 times higher than the existing technology. At the same time, the average rendering time per frame is 0.7 ms, which is much higher than the requirement of 90HZ for immersive head-mounted displays and meets the real-time requirement of VR content.

[0078] Therefore, the present invention adopts the above-mentioned real-time light field stereo rendering method for immersive head-mounted displays, improves the real-time efficiency of light field stereo rendering, and meets the requirements of immersive head-mounted displays.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements do not enable the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A real-time light field stereo rendering method for an immersive head mounted display, characterized in that the steps include: S1, uniformly sampling the original light field data in the angle space, transforming the original data into the YUV color space, and using zigzag scanning to generate a multi-viewpoint image pseudo video sequence; S2, compressing the image pseudo video sequence using H.264 video encoder and uploading it to the server; S3, loading the light field data transmitted to the server through the client, and decompressing the light field data using WebGL and Three.js to generate a three-dimensional texture map; S4. Design a vertex shader and a fragment shader. The fragment shader is designed based on a refocusing algorithm in the spatial domain, specifically including: Where i, j represent the coordinates of the imaging plane, M1M2 represents the sensor size, a represents the aperture size measurement, f represents the focal length measurement, F(x,y) represents the color value [R, G, B] of the light field sub-aperture image at the pixel coordinate x at the y coordinate, d represents the street distance vector between the current pixel and the target pixel, and x screen ,y screen Represents the horizontal and vertical coordinates of screen pixels, d x , d y Respectively represent the horizontal and vertical components of the distance vector d; S5. Use the plane in the three-dimensional space as the proxy geometry of the scene, use the geometric data as the input of the rasterization rendering pipeline, use the vertex shader and the fragment shader, and combine the three-dimensional map to render the binocular viewpoint image.

2. A real-time light field stereo rendering method for an immersive head mounted display according to claim 1, characterized in that: The vertex shader in step S4 includes clipping vertex coordinates and calculating the starting point and direction of light. The clipping vertex coordinate formula is include: In clip =M projection *M view *M model *V; Where V = (x, y, z, 1) represents the local coordinates of the vertex, V clip =(x, y, z, w) represents the converted clipping coordinates, M projection Represents the projection matrix of the virtual viewpoint camera, which is defined by the field of view angle and the near and far planes of projection. view Represents the view matrix, which is used to convert the world coordinates of the three-dimensional plane into camera coordinates. model Represents the model matrix, which is used to transform the local coordinates of the three-dimensional plane into world coordinates; The starting point of the light is calculated according to the texture coordinates of the three-dimensional plane, and the direction of the light is calculated according to the position of the virtual camera and the local vertex coordinates of the three-dimensional plane. The formula is: r o =[u,v]; r d =normalize(p camera -p plane ); In the formula, r d represents the direction of the light, [u, v] represents the texture coordinates of the three-dimensional plane, r o represents the starting point of the light, p camera represents the position of the virtual camera, p plane Represents the local vertex coordinates of a 3D plane.

3. The real-time light field stereo rendering method for an immersive head mounted display according to claim 2, characterized in that: Step S5 includes: S51, using a plane in the three-dimensional space as a geometric proxy for the scene, and setting the size of the plane according to the angular resolution of the light field; S52, input the plane in the three-dimensional space into the rasterization rendering pipeline, use the vertex shader to clip the vertex coordinates, and calculate the direction of the light; S53, converting the clipped coordinates into normalized device coordinates through perspective division; S54, using a fragment shader to convert the device coordinates into final screen coordinates according to the current posture of the virtual camera, and generating a target view according to the aperture and focal length parameters of the refocusing; S55: Rendering a binocular viewpoint image according to the position of the binocular viewpoint.

4. The real-time light field stereo rendering method for an immersive head mounted display according to claim 3, characterized in that: Step S53 converts the clipped coordinates into normalized device coordinates by perspective division: Where V NDC Represents device coordinates.

5. A real-time light field stereo rendering method for an immersive head mounted display according to claim 4, characterized in that: The formula for converting the device coordinates to the final screen coordinates in step S54 is: Where V screen Represents screen coordinates, width and height represent the width and height of the viewport respectively, x offset and offset Indicates the offset of the lower left corner of the viewport on the screen.

Citation Information

Patent Citations

  • Interactive real-time autostereoscopic display method based on rendering pipeline

    CN108573524A

  • Multi-level parallel rendering method and system for light field image

    CN118537468A