Unity engine-based image fusion stereoscopic display method and system

By configuring the left-eye camera and the right-eye camera in the Unity engine and combining them with Compute Shader for image fusion processing, the problems of UI element collaboration, low resource adaptation efficiency, and poor hardware compatibility in 3D stereoscopic display are solved. Efficient image fusion and multi-device compatibility are achieved, improving user experience and rendering performance.

CN120725892AActive Publication Date: 2025-09-30JIANGSU LIREN TECH CO LTD

Patent Information

Application Number
CN202511159198.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-09-30
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing 3D stereoscopic display technology has a gap in the collaborative work between images and UI elements, low resource adaptation efficiency, poor hardware device compatibility, low rendering efficiency, inability to meet high resolution and high frame rate requirements, and insufficient visual personalization.

Method used

By configuring the left-eye camera and the right-eye camera in the Unity engine, generating left-eye rendering textures and right-eye rendering textures, combining Compute Shader for image fusion processing, dynamically adjusting the pupil distance, supporting the adaptation of UI elements in multiple Canvas modes, and using GPU accelerated rendering, we can achieve temporal continuous image fusion and multi-device compatibility.

Benefits of technology

It improves the smoothness and stability of 3D stereoscopic display, enhances user experience, supports compatibility with multiple stereoscopic display devices, shortens the development cycle, and improves the system's computing efficiency and visual personalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725892A_ABST
    Figure CN120725892A_ABST
Patent Text Reader

Abstract

The invention provides an image fusion stereoscopic display method and system based on a Unity engine, and relates to the technical field of computer graphics and virtual reality. According to the method, a left-eye camera and a right-eye camera are configured in a Unity scene, left and right view angle images are collected according to interpupillary distance parameters, and left-eye and right-eye rendering textures are generated; generating a UI texture according to the Canvas type of the UI element, executing parallax compensation, pixel alignment and image fusion processing through a Compute Shader, and superposing image layers according to the transparency of the UI texture to generate a synthetic display image; and introducing a previous frame result into the fused image, and realizing time continuous output based on an inter-frame smooth function. And outputting the synthesized texture to a device supporting a plurality of stereoscopic display modes through a graphic rendering pipeline. According to the method, efficient fusion of the 3D image and the UI is realized, the rendering efficiency and the visual immersion are improved, and the method is suitable for scenes such as virtual reality, augmented reality and digital twinning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer graphics and virtual reality technology, and in particular to a Unity engine-based image fusion stereoscopic display method and system. Background Art

[0002] With the rapid development of virtual reality (VR), augmented reality (AR) and 3D content creation (such as games, virtual simulations, digital twins, etc.), the Unity engine, as a mainstream real-time 3D development tool, is widely used in 3D stereoscopic display, game development and virtual reality. However, existing 3D stereoscopic display technology has multiple problems in terms of the coordination between images and UI elements, resource adaptation efficiency, hardware device compatibility, rendering efficiency and visual personalization. Specifically, the collaborative work between UI elements and 3D images in existing technologies is usually disconnected. Especially in Overlay mode, the UI often completely blocks the stereoscopic picture, resulting in poor display effects, and developers need to manually adjust the coordinates, hierarchy and other parameters of UI elements, which prolongs the development cycle. In addition, traditional hardware adaptation methods are also unable to effectively be compatible with different types of display devices, and are prone to screen tearing and synchronization delays, affecting the user experience.

[0003] At the same time, existing rendering methods mostly rely on the CPU for texture synthesis processing, which makes it impossible to meet the rendering requirements of high-resolution and high-frame-rate scenes (such as 4K / 120fps games), resulting in problems such as screen freezes and rendering delays. Existing technologies also fail to fully utilize the parallel computing capabilities of the GPU, limiting improvements in image processing speed and rendering efficiency, further affecting the smoothness and stability of the display effect. In addition, the existing binocular camera spacing is usually fixed (such as 0.065 meters), which is difficult to adapt to the needs of children, myopic users, or users with special viewing angles. This results in insufficient visual immersion and fails to meet the personalized needs of users.

[0004] Therefore, how to solve these problems and improve the flexibility of 3D stereoscopic display and UI collaboration, resource adaptation efficiency, hardware compatibility, rendering efficiency and visual personalization has become a core issue that needs to be urgently addressed in current technology. Summary of the Invention

[0005] The purpose of the present invention is to provide an image fusion stereoscopic display method and system based on the Unity engine to solve the limitations of the existing technology, such as the separation of 3D and UI collaboration, inefficient resource adaptation, weak stereoscopic display hardware adaptation, rendering efficiency bottlenecks and lack of visual personalization. Through innovative UI adaptation and image fusion technology, GPU accelerated rendering, dynamic pupil distance adjustment and multi-display device compatibility, the smoothness, compatibility and user experience of the display effect are significantly improved, while shortening the development cycle and improving the computing efficiency of the system.

[0006] The present invention achieves the above-mentioned purpose through the following technical solutions:

[0007] In one aspect, the present invention provides an image fusion stereoscopic display method based on the Unity engine, the method comprising:

[0008] Step S1: Configure the left-eye camera and the right-eye camera in the Unity scene, capture the left-view image and the right-view image respectively based on the set pupil distance parameters, and generate the left-eye rendering texture and the right-eye rendering texture;

[0009] Step S2: According to the Canvas type bound to each UI element in the user interface, a corresponding display mode flag is configured, wherein the display mode flag is used to indicate the display mode of the UI element in the left-eye image and the right-eye image. The display modes include binocular synchronous display, left-eye display only, and right-eye display only; and a user interface texture is generated based on the UI rendering process;

[0010] Step S3: Input the left-eye rendering texture, the right-eye rendering texture, and the user interface texture into the Compute Shader and perform image fusion processing, which includes: dynamically calculating the fusion weights of the left and right views based on image content features; performing a weighted operation based on the previous frame image and the current frame image, where the weight of the weighted operation is determined by the adaptive smoothing coefficient; performing parallax compensation and pixel alignment on the right-eye image based on the depth map; fusing the left and right eye images according to the fusion weights to generate a fused image; and overlaying the UI layer according to the transparency of the user interface texture to obtain a composite display image.

[0011] Step S4: writing the synthesized display image into a synthesized display texture, and outputting the synthesized display texture to a stereoscopic display device, wherein the stereoscopic display device supports a frame sequential, shutter or line alternating stereoscopic image display mode.

[0012] Preferably, the step S1 includes:

[0013] Duplicate the main camera to generate a left-eye camera and a right-eye camera. The left-eye camera and the right-eye camera are symmetrically offset relative to the main camera in the horizontal direction. The offset distance is an adjustable parameter with a default value of 0.065 meters and a value range of 0.03 meters to 0.08 meters.

[0014] Adjust the pupil distance parameter value in real time through the control interface, and dynamically update the horizontal position offset of the left-eye camera and the right-eye camera;

[0015] The field of view angles of the left-eye camera and the right-eye camera are both set to 60 degrees;

[0016] The left-eye camera and the right-eye camera collect left-view images and right-view images respectively, and render to generate left-eye rendering texture and right-eye rendering texture. The rendering texture is a two-dimensional texture that can be read and written by the GPU, and the data type is Texture2D <float4>.

[0017] Preferably, step S2 includes:

[0018] Identify the Canvas type of the UI element based on the display mode tag. When the Canvas type is Overlay mode, generate an independent user interface texture bound to the Render Texture object. When the Canvas type is World Space mode or Screen Space-Camera mode, the left-eye camera and the right-eye camera capture images separately and generate a unified user interface texture.

[0019] Automatically configure rendering parameters of the user interface texture according to the Canvas type, wherein the rendering parameters include: rendering priority, scaling factor, and occlusion level information;

[0020] The user interface texture is a two-dimensional texture that can be read and written by the GPU, and the data type is Texture2D <float4>.

[0021] Preferably, step S3 includes:

[0022] The left-eye rendering texture, the right-eye rendering texture, and the user interface texture are input to the Compute Shader, and image fusion processing is performed in the Compute Shader. The image fusion processing includes the following steps:

[0023] Dynamically calculate the fusion weights of the left and right views based on image content features;

[0024] Calculate the parallax compensation displacement of the right eye image based on the depth map and perform pixel alignment;

[0025] Fusing the aligned left and right images to obtain a fused image;

[0026] The user interface image is superimposed on the fused image according to its transparency information to generate the final display image. The fusion calculation formula is:

[0027] ;

[0028] in: is the left eye image at pixel location The color value of For the right eye image at pixel location The color value of For user interface images at pixel locations The color value of For user interface images at location The alpha channel value, ranging from [0,1], is used to control transparency; is the weight coefficient for left and right image fusion, which is dynamically calculated based on the image content and ranges from [0,1]; For the right eye image at position The parallax compensation displacement at comes from the depth map; is the pixel value after the left and right images are fused; For the final composite display image at pixel position The color value of .

[0029] The step S3 further includes:

[0030] Obtain the fusion weight change, parallax compensation displacement change, and user visual parameters between the current frame and the previous frame; calculate the adaptive inter-frame smoothing coefficient for each pixel position based on the basic smoothing coefficient, weight change adjustment coefficient, parallax change adjustment coefficient, and user parameter adjustment coefficient;

[0031] Based on the adaptive inter-frame smoothing coefficient, a weighted operation is performed on corresponding pixel values ​​of the current frame fusion image and the previous frame fusion image to generate an output pixel value of the current frame.

[0032] Preferably, step S4 includes:

[0033] The synthesized display texture is bound as an output texture to a target frame buffer of a graphics rendering pipeline and transmitted to a stereoscopic display device; the stereoscopic display device supports a frame sequential, shutter or line alternating stereoscopic image display mode.

[0034] Preferably, the left-eye rendering texture, right-eye rendering texture, user interface texture and synthetic display texture are all two-dimensional textures that can be read and written by the GPU, and the data type is Texture2D <float4>, and is configured as a continuous memory layout to support multi-threaded parallel access and fast processing of image data.

[0035] Preferably, the method further comprises a resource adaptation step, wherein the resource adaptation step comprises:

[0036] Configuring a metadata component including a display mode flag and a scaling factor for each user interface element, wherein the display mode flag is used to indicate a display mode of the interface element in the left-eye image and the right-eye image, wherein the display modes include binocular synchronization, left-eye display only, and right-eye display only;

[0037] In the resource preprocessing stage before image rendering, the editor extension script is used to identify the Canvas type bound to each user interface element, and the RectTransform property and Canvas Renderer property of the interface element are automatically adjusted according to the content in the metadata component; before image fusion processing, the horizontal pixel offset of each user interface element in the left and right images is calculated based on the user's pupil distance parameter and gaze point depth parameter. , used to adjust the position of interface elements in the user interface texture. The formula is:

[0038] ;

[0039] in: is the user's pupil distance parameter; is the gaze point depth parameter; The distance from the user's eyes to the screen; is the scaling factor of the interface elements; is the calibration parameter.

[0040] On the other hand, the present invention provides an image fusion stereoscopic display system based on the Unity engine, which is applied to the image fusion stereoscopic display method based on the Unity engine as described above, and is characterized by comprising:

[0041] The binocular rendering module is used to configure the left-eye camera and the right-eye camera in the Unity scene, and respectively capture the left-view image and the right-view image based on the set pupil distance parameters to generate the left-eye rendering texture and the right-eye rendering texture;

[0042] The interface texture generation module is used to identify the corresponding display mode mark according to the Canvas type bound to each UI element in the user interface, and generate the user interface texture;

[0043] an image fusion module that performs image fusion processing, including a weighted operation based on a previous frame image and a current frame image, wherein the weight of the weighted operation is determined by an adaptive smoothing coefficient, and performs superposition processing on the fused image based on the alpha channel information of the user interface texture to generate a synthetic display texture;

[0044] A time-continuous fusion module is used to perform weighted operations on corresponding pixel values ​​of the previous frame fusion image and the current frame fusion image in the image fusion module to achieve time-continuous image fusion;

[0045] The resource adapter module is used to configure metadata components including display mode flags and scaling factors for each UI element. During the resource preprocessing phase, it uses the editor extension script to identify the Canvas type bound to the UI element and adjust the RectTransform and Canvas Renderer properties.

[0046] A position control module is used to calculate the horizontal offset of each UI element in the composite image according to the user's pupil distance parameter and the gaze point depth parameter before the fusion process, and adjust the position of the corresponding UI element in the interface texture based on the offset;

[0047] The scheduling control module is used to calculate the number of two-dimensional thread groups according to the target display resolution and trigger ComputeShader to process image fusion tasks in parallel;

[0048] An output module is used to bind the synthesized display texture to a target frame buffer of a graphics rendering pipeline and transmit it to a stereoscopic display device that supports a frame sequential, shutter or line alternating display mode.

[0049] A computer-readable storage medium stores a computer program, which, when executed by a processor, is used to implement the image fusion stereoscopic display method based on the Unity engine as described above.

[0050] The beneficial effects of the present invention are as follows: Through innovative binocular image rendering and user interface (UI) adaptation technology, the present invention significantly improves the 3D stereoscopic display effect and user experience. By configuring left-eye and right-eye cameras in a Unity scene and capturing left and right perspective images according to the set interpupillary distance parameters, the present invention can accurately generate left and right eye rendering textures while supporting dynamic adaptation of UI elements in different Canvas modes. Using Compute Shader for image fusion processing, the present invention combines the depth map to perform parallax compensation on the right-eye image and calculates fusion weights based on image content features to ensure precise alignment and fusion of the left and right eye images. Furthermore, the fused image of the previous frame is used for temporally continuous fusion with the current frame, effectively avoiding image jitter and delay, and improving the smoothness and stability of the stereoscopic display. Ultimately, the generated composite display image is compatible with a variety of stereoscopic display devices, including frame sequential, shutter, and line alternating display modes, ensuring wide device compatibility and optimized display effects. Therefore, the present invention not only improves the accuracy of image fusion and the adaptation efficiency of UI elements, but also optimizes the compatibility and rendering performance of display devices while enhancing user immersion, thus resolving multiple limitations of the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort. Among them: Figure 1 The present invention is a flow chart of the method of the present invention; Figure 2 Schematic diagram of the overall structure of the system of the present invention. DETAILED DESCRIPTION

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of the present invention.

[0053] Example 1

[0054] like Figure 1 FIG. 1 is an embodiment of the present invention, which provides an image fusion stereoscopic display method based on the Unity engine, including:

[0055] Step S1: Binocular perspective and UI texture acquisition

[0056] Binocular camera configuration: Configure the left-eye camera and the right-eye camera in the Unity scene, and capture left-view images and right-view images based on the set interpupillary distance parameters.

[0057] Furthermore, the main camera is copied to generate a left-eye camera and a right-eye camera. The left-eye camera and the right-eye camera are symmetrically offset relative to the main camera in the horizontal direction. The offset distance is an adjustable parameter with a default value of 0.065 meters and a value range of 0.03 meters to 0.08 meters.

[0058] The interpupillary distance parameters are adjusted in real time through the control interface, and the horizontal position offset of the left-eye camera and the right-eye camera is dynamically updated.

[0059] Field of view setting: The field of view angles of the left-eye camera and the right-eye camera are both set to 60 degrees.

[0060] Texture generation: The left-eye camera and the right-eye camera collect the view images and rendering textures respectively. The rendering texture is a two-dimensional texture that can be read and written by the GPU. The data type is Texture2D <float4>

[0061] Step S2: According to the Canvas type bound to each UI element in the user interface, configure the corresponding display mode mark, which is used to indicate the display mode of the UI element in the left-eye image and the right-eye image. The display modes include binocular synchronous display, left-eye display only, and right-eye display only; generate the user interface texture based on the UI rendering process.

[0062] Display mode flag: The display mode flag is used to indicate the display mode of UI elements in the left-eye image and the right-eye image. The flags include: binocular synchronous display, left-eye display only, and right-eye display only.

[0063] Canvas type identification: Based on the display mode tag, the system identifies the Canvas type of the UI element and configures different rendering logic according to the different Canvas types:

[0064] When the Canvas type is Overlay mode, an independent user interface texture is generated and bound to the Render Texture object.

[0065] When the Canvas type is World Space mode or Screen Space-Camera mode, the left-eye camera and the right-eye camera respectively capture images and generate a unified user interface texture.

[0066] Rendering parameter configuration of user interface texture: According to the Canvas type, automatically configure the rendering parameters of the user interface texture, including: rendering priority, scaling factor, occlusion level information, user interface texture generation:

[0067] The user interface texture is a two-dimensional texture that can be read and written by the GPU, and its data type is Texture2D <float4>, used for subsequent image fusion processing.

[0068] Step S3: Image fusion processing

[0069] The left-eye rendering texture, the right-eye rendering texture, and the user interface texture are input to the compute shader for image fusion processing. The fusion processing includes: dynamically calculating the fusion weights of the left and right views based on image content features; performing a weighted operation based on the previous frame image and the current frame image, where the weight of the weighted operation is determined by the adaptive smoothing coefficient; performing parallax compensation and pixel alignment on the right-eye image based on the depth map; fusing the left and right eye images according to the fusion weights to generate a fused image; and overlaying the UI layer according to the transparency of the user interface texture to obtain a composite display image.

[0070] Furthermore, the pixel data of the left-eye image, the right-eye image, and the user interface image are input to Compute Shader, and image fusion processing is performed in Compute Shader. The image fusion processing includes the following steps:

[0071] Dynamically calculate the fusion weights of the left and right views based on image content features;

[0072] Calculate the parallax compensation displacement of the right eye image based on the depth map and perform pixel alignment;

[0073] Fusing the aligned left and right images to obtain a fused image;

[0074] The user interface image is superimposed on the fused image according to its transparency information to generate the final display image. The fusion calculation formula is:

[0075] ;

[0076] in: is the left eye image at pixel location The color value of For the right eye image at pixel location The color value of For user interface images at pixel locations The color value of For user interface images at location The alpha channel value, ranging from [0,1], is used to control transparency; is the weight coefficient for left and right image fusion, which is dynamically calculated based on the image content and ranges from [0,1]; For the right eye image at position The parallax compensation displacement at comes from the depth map; is the pixel value after the left and right images are fused; For the final composite display image at pixel position The color value of .

[0077] In one embodiment of the present invention, the image fusion process is implemented using the Compute Shader in the Unity engine. The image processing module fuses the left-eye image, right-eye image, and UI image using GPU parallel computing. The following is a typical Compute Shader implementation code:

[0078] Texture2D <float4>InputLeftTexture; / / Left eye image texture

[0079] Texture2D <float4>InputRightTexture; / / right eye image texture

[0080] Texture2D <float4>UITexture; / / model

[0081] SamplerState linearSampler; / / Linear sampler

[0082] RWTexture2D <float4>OutputTexture; / / output image

[0083] int Width;

[0084] int Height;

[0085] [numthreads(8, 8, 1)]

[0086] void CSMain(uint3 threadID : SV_DispatchThreadID)

[0087] {

[0088] / / Step 1: Synthesize 3D texture alternately row by row

[0089] uint index = threadID.y % 2; / / odd-even row judgment

[0090] float4 stereoColor = (index == 1)

[0091] ?InputLeftTexture[threadID.xy]

[0092] :InputRightTexture[threadID.xy];

[0093] / / Step 2: Mode Texture Overlay ( mix)

[0094] float2 uv = float2(threadID.x / (float)Width, threadID.y / (float)Height);

[0095] float4 uiColor = UITexture.Sample(linearSampler, uv);

[0096] / / Calculate the final synthesis result

[0097] OutputTexture[threadID.xy] = lerp(stereoColor, uiColor, uiColor.a);

[0098] }

[0099] The above code demonstrates a typical implementation of this invention, achieving seamless integration of the UI and stereoscopic imagery without sacrificing frame rate or image quality, and adapting to various rendering modes and stereoscopic display devices. The specific implementation can be adjusted based on the performance requirements of the development platform and display terminal.

[0100] In one embodiment, the time-continuous fusion processing in step S3 specifically includes the following steps:

[0101] Step 1: Calculation of adaptive inter-frame smoothing coefficient

[0102] Obtain the fusion weight, parallax compensation displacement and user visual parameters of the current frame and the previous frame, including the user's pupil distance parameter and the gaze point depth parameter. Calculate the adaptive inter-frame smoothing coefficient for each pixel position based on the preset basic smoothing coefficient, weight change adjustment coefficient, parallax change adjustment coefficient and user parameter adjustment coefficient. .

[0103] The calculation formula of the adaptive inter-frame smoothing coefficient is:

[0104] in: is the basic smoothing coefficient; and The current frame and the previous frame in pixels The fusion weight value at ; and The current frame and the previous frame in pixels The parallax compensation displacement value at ; is the user's pupil distance parameter; is the user's gaze point depth parameter; 、 、 They are the adjustment coefficients of image content change, parallax change and user visual parameters respectively.

[0105] Step 2: Weighted operation of temporal continuous pixel values

[0106] According to the adaptive inter-frame smoothing coefficient, a weighted operation is performed on the corresponding pixel values ​​of the current frame fusion image and the previous frame fusion image to obtain the current frame output pixel value , and its weighted calculation formula is:

[0107] ;

[0108] in: The fused image of the previous frame in pixels The pixel value at ; The current frame fusion image is in pixels The pixel value at .

[0109] Step 3: Output pixel value range limitation

[0110] To ensure that the output data is within the legal range, the calculated output pixel value of the current frame is numerically restricted so that it is within the preset pixel valid value range. The restriction formula is:

[0111] ;

[0112] Through the above steps, time-continuous inter-frame pixel fusion processing can be completed in the Compute Shader.

[0113] In order to more clearly illustrate the technical solution of the present invention, an embodiment of the image fusion stereoscopic display method of the present invention is described in detail below in conjunction with specific computer program codes.

[0114] Step S4: writing the synthesized display image into a synthesized display texture, and outputting the synthesized display texture to a stereoscopic display device, wherein the stereoscopic display device supports a frame sequential, shutter or line alternating stereoscopic image display mode.

[0115] Furthermore, the step S4 includes:

[0116] The synthesized display texture is bound as an output texture to a target frame buffer of a graphics rendering pipeline and transmitted to a stereoscopic display device; the stereoscopic display device supports a frame sequential, shutter or line alternating stereoscopic image display mode.

[0117] In a further embodiment of the present invention, the left eye rendering texture, the right eye rendering texture, the user interface texture and the synthetic display texture are all two-dimensional textures that can be read and written by the GPU, and the data type is Texture2D <float4>, and is configured as a continuous memory layout to support multi-threaded parallel access and fast processing of image data.

[0118] This example further includes a resource adaptation step, which includes:

[0119] Configuring a metadata component including a display mode flag and a scaling factor for each user interface element, wherein the display mode flag is used to indicate a display mode of the interface element in the left-eye image and the right-eye image, wherein the display modes include binocular synchronization, left-eye display only, and right-eye display only;

[0120] In the resource preprocessing stage before image rendering, the editor extension script is used to identify the Canvas type bound to each user interface element, and the RectTransform property and Canvas Renderer property of the interface element are automatically adjusted according to the content in the metadata component; before image fusion processing, the horizontal pixel offset of each user interface element in the left and right images is calculated based on the user's pupil distance parameter and gaze point depth parameter. , used to adjust the position of interface elements in the user interface texture. The formula is:

[0121] ;

[0122] in: is the user's pupil distance parameter; is the gaze point depth parameter; The distance from the user's eyes to the screen; is the scaling factor of the interface elements; is the calibration parameter; The pixel value of the horizontal offset of the interface element.

[0123] like Figure 2 FIG. 1 is another embodiment of the present invention, which provides an image fusion stereoscopic display system based on the Unity engine, applied to the image fusion stereoscopic display method based on the Unity engine described above, including:

[0124] The binocular rendering module is used to configure the left-eye camera and the right-eye camera in the Unity scene, and respectively capture the left-view image and the right-view image based on the set pupil distance parameters to generate the left-eye rendering texture and the right-eye rendering texture;

[0125] The interface texture generation module is used to identify the corresponding display mode mark according to the Canvas type bound to each UI element in the user interface, and generate the user interface texture;

[0126] an image fusion module that performs image fusion processing, including a weighted operation based on a previous frame image and a current frame image, wherein the weight of the weighted operation is determined by an adaptive smoothing coefficient, and performs superposition processing on the fused image based on the alpha channel information of the user interface texture to generate a synthetic display texture;

[0127] A time-continuous fusion module is used to perform weighted operations on corresponding pixel values ​​of the previous frame fusion image and the current frame fusion image in the image fusion module to achieve time-continuous image fusion;

[0128] The resource adapter module is used to configure metadata components including display mode flags and scaling factors for each UI element. During the resource preprocessing phase, it uses the editor extension script to identify the Canvas type bound to the UI element and adjust the RectTransform and Canvas Renderer properties.

[0129] A position control module is used to calculate the horizontal offset of each UI element in the composite image according to the user's pupil distance parameter and the gaze point depth parameter before the fusion process, and adjust the position of the corresponding UI element in the interface texture based on the offset;

[0130] The scheduling control module is used to calculate the number of two-dimensional thread groups according to the target display resolution and trigger ComputeShader to process image fusion tasks in parallel;

[0131] An output module is used to bind the synthesized display texture to a target frame buffer of a graphics rendering pipeline and transmit it to a stereoscopic display device that supports a frame sequential, shutter or line alternating display mode.

[0132] A computer-readable storage medium stores a computer program, which, when executed by a processor, is used to implement the image fusion stereoscopic display method based on the Unity engine as described above.

[0133] Example 2

[0134] The specific implementation method provided in this embodiment is intended to support UI display fusion in different Canvas modes, and introduce adjustable pupil distance parameters to enhance the personalized control capability of the stereoscopic visual experience.

[0135] For stereo rendering, the system first creates a left-eye camera (LeftCamera) and a right-eye camera (Right Camera) by duplicating the main camera. These cameras are symmetrically offset along the X-axis relative to the main camera, with the offset dynamically controlled by the inter Pupillary Distance parameter. This parameter defaults to 0.065 meters and can be adjusted in real time using a slider control within a range of 0.03 to 0.08 meters. Each left-eye camera is bound to its own RenderTexture object (Left RT and Right RT), which serves as the Input Left Texture and Input Right Texture, respectively, in the subsequent image fusion process.

[0136] In terms of UI image acquisition and fusion, the system supports UI configuration in both Overlay and World Space modes. For Overlay mode, create a Canvas node named UI Overlay Canvas and add user interface elements such as buttons and health bars to it. The entire Canvas is bound to a Render Texture (UI Overlay RT) to generate the UITexture. For World Space mode, create a UI World Canvas and set it to World Space display mode. Place it in the 3D scene, such as one meter in front of the character, and ensure that the Culling Masks of the left and right eye cameras include the UI layer to achieve complete capture of the UI texture in 3D space.

[0137] For image fusion processing, the system uses a compute shader to achieve row-by-row fusion of the left and right images, overlaying the UI texture information on the fused image. The core shader logic is defined in the StereoDisplay.compute file, using a two-dimensional thread structure of [numthreads(8,8,1)]. In the CSMain main function, the left and right image source pixels are switched by determining the parity of the pixel row number, achieving row-interleaved synthesis of the 3D stereo image. Simultaneously, the pixel color and alpha channel values ​​of the corresponding positions in the UI texture are sampled, blended according to transparency using the lerp function, and output to the target texture, OutputTexture, achieving seamless fusion of the UI layer and the 3D image.

[0138] To improve development efficiency and UI resource adaptability, the system introduces a metadata component called Stereo UI Meta, allowing users to annotate each UI element with parameters such as the stereoscopic display mode (binocular synchronization, left only, right only) and scaling factor. A supporting editor extension script, StereoUIEditor.cs, has also been developed to support "one-click Canvas mode adaptation." This automatically identifies the Canvas type of a UI element and adjusts its RectTransform and CanvasRenderer parameters to ensure correct rendering and integration of UI elements in different Canvas modes.

[0139] While maintaining high real-time performance and high frame rate rendering efficiency, this embodiment solves the resource adaptation and integration problems under different UI display modes, realizes personalized pupil distance control and multi-device compatible display, and provides a good application foundation and scalability for virtual reality, augmented reality and three-dimensional interactive systems.

[0140] Example 3: Metadata-based UI intelligent adaptation and runtime pupil distance dynamic control method

[0141] To achieve automatic adaptation of UI elements in different Canvas display modes and enhance the interactivity and personalized experience of the stereoscopic display system, this embodiment designs and deploys a metadata-based UI adaptation mechanism and runtime interpupillary distance control function on the Unity platform. This solution can adapt to various rendering modes and display terminals, and has strong versatility and engineering practicality.

[0142] 1. UI display mode metadata identification mechanism

[0143] First, add a custom StereoUIMeta component to each UI element. This component inherits from the Unity engine's MonoBehaviour class and defines an enumeration type used to identify how the UI element is displayed in the left and right eye images. The display mode supports three types: binocular simultaneous display, left eye only display, and right eye only display. The typical implementation code is as follows:

[0144] / / StereoUIMeta.cs — UI metadata component definition

[0145] public class StereoUIMeta : MonoBehaviour

[0146] {

[0147] / / Enumeration type: definition The visible range of the element

[0148] public enum DisplayMode

[0149] {

[0150] BothEyes, / / Display both left and right eye images

[0151] LeftOnly, / / Display only in the left eye image

[0152] RightOnly / / Display only in the right eye image

[0153] }

[0154] / / Current The display mode of the element, the default is binocular display

[0155] public DisplayMode displayMode = DisplayMode.BothEyes;

[0156] }

[0157] This component provides an independent metadata tag for each UI element, which can be used during rendering to determine whether to draw it into the left eye image, the right eye image, or both images.

[0158] 2. UI elements automatically adapt to Canvas type editor extensions

[0159] To simplify the configuration process of UI resources in different Canvas modes and improve development efficiency, the system introduces the editor extension script StereoUIEditor.cs. By identifying the RenderMode property of the Canvas to which the UI element is attached, it implements the one-click adaptation of UI parameters. The core code is as follows:

[0160] / / StereoUIEditor.cs — Adaptive Editor Extension

[0161] [CustomEditor(typeof(StereoUIMeta))]

[0162] public class StereoUIEditor : Editor

[0163] {

[0164] public override void OnInspectorGUI()

[0165] {

[0166] base.OnInspectorGUI();

[0167] StereoUIMeta meta = (StereoUIMeta)target;

[0168] if (GUILayout.Button("Auto Adapt to Canvas Mode"))

[0169] {

[0170] Canvas canvas = meta.GetComponentInParent <canvas>();

[0171] switch (canvas.renderMode)

[0172] {

[0173] case RenderMode.Overlay:

[0174] meta.GetComponent <recttransform>().localScale = Vector3.one * 1.2f;

[0175] meta.GetComponent <canvasrenderer>().cull = false;

[0176] break;

[0177] case RenderMode.WorldSpace:

[0178] meta.GetComponent <recttransform>().position = new Vector3(0, 1.5f,5f);

[0179] break;

[0180] }

[0181] }

[0182] }

[0183] }

[0184] This mechanism achieves compatibility with multiple Canvas display modes by automatically adjusting the coordinates, scaling, and occlusion properties of UI elements, effectively reducing manual debugging costs.

[0185] 3. Dynamic adjustment control logic of pupil distance parameters during operation

[0186] To improve the system's visual adaptability to different users (such as children and people with myopia), the system provides dynamic control of the pupillary distance parameter. The core controller script, StereoDisplayController.cs, defines the pupillary distance field and adjusts the positions of the left and right eye cameras in real time at runtime, as follows:

[0187] / / StereoDisplayController.cs — IPD controller

[0188] public class StereoDisplayController : MonoBehaviour

[0189] {

[0190] public float interPupillaryDistance = 0.065f; / / Initial pupil distance, unit: meter

[0191] public Transform leftCameraTrans;

[0192] public Transform rightCameraTrans;

[0193] void Update()

[0194] {

[0195] float halfIPD = interPupillaryDistance / 2;

[0196] leftCameraTrans.localPosition = new Vector3(-halfIPD, 0, 0);

[0197] rightCameraTrans.localPosition = new Vector3(halfIPD, 0, 0);

[0198] }

[0199] }

[0200] 4. User interactive pupil distance adjustment interface binding

[0201] At the UI level, we use the Slider component to create an interactive IPD adjustment interface. We use the script IPDAdjuster.cs to bind the Slider's sliding value to the IPD parameter in the controller to implement user-defined adjustment functionality. The implementation code is as follows:

[0202] / / IPDAdjuster.cs — Slide adjuster

[0203] public class IPDAdjuster : MonoBehaviour

[0204] {

[0205] public StereoDisplayController stereoController; / / controller reference

[0206] public Slider ipdSlider; / / Component reference

[0207] void Start()

[0208] {

[0209] / / Sliding event binding: real-time update of pupil distance value

[0210] ipdSlider.onValueChanged.AddListener((value) => {

[0211] stereoController.interPupillaryDistance = value;

[0212] });

[0213] }

[0214] }

[0215] Users can flexibly set the pupil distance value between 0.03 and 0.08 meters by sliding the Slider control. The system automatically and synchronously updates the camera position, thereby changing the binocular parallax in real time and enhancing the personalized experience of stereo perception.

[0216] This embodiment introduces a metadata tagging mechanism, editor extensions, and runtime interaction logic to build a complete "UI intelligent adaptation + dynamic control of pupil distance" solution, achieving the following technical effects:

[0217] Realize automatic adaptation of UI elements in different Canvas modes to improve development efficiency;

[0218] Supports users to adjust the pupil distance value in a personalized way to improve the comfort of stereoscopic vision;

[0219] The system has good scalability and operational stability, and is compatible with multiple platforms and multi-resolution devices.

[0220] This solution significantly improves the interactivity, configurability and human-machine adaptability of the stereoscopic display system, and has high practical value and engineering feasibility.

[0221] In summary, the present invention solves a number of technical problems existing in the prior art, especially in the aspects of the collaborative work of 3D stereoscopic display and UI elements, resource adaptation efficiency, rendering performance and personalized adjustment. Through innovative image fusion processing, UI automatic adaptation and GPU acceleration technology, the present invention can provide efficient 3D display effects and good user experience, and is particularly suitable for fields such as virtual reality (VR), augmented reality (AR), game development and digital twins. In addition, the present invention has strong hardware adaptability and is compatible with different types of stereoscopic display devices, further expanding the application scenarios and market prospects. In general, the present invention has broad application value and good commercial prospects, can promote technological development in related fields, and provide innovative solutions for related industries.

[0222] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate different embodiments or examples, and features of different embodiments or examples, described in this specification, unless otherwise mutually incompatible.

[0223] Any process or method description in the flowchart or otherwise described herein can be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations in which the functions may be performed in a different order than shown or discussed, including in a substantially simultaneous manner or in a reverse order depending on the functions involved.

[0224] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.< / recttransform> < / canvasrenderer> < / recttransform> < / canvas>

Claims

1. An image fusion stereoscopic display method based on Unity engine, characterized in that: The method comprises: Step S1: Configure the left-eye camera and the right-eye camera in the Unity scene, capture the left-view image and the right-view image respectively based on the set pupil distance parameters, and generate the left-eye rendering texture and the right-eye rendering texture; Step S2: According to the Canvas type bound to each UI element in the user interface, a corresponding display mode flag is configured, wherein the display mode flag is used to indicate the display mode of the UI element in the left-eye image and the right-eye image. The display modes include binocular synchronous display, left-eye display only, and right-eye display only; and a user interface texture is generated based on the UI rendering process; Step S3: Input the left-eye rendering texture, the right-eye rendering texture, and the user interface texture into the Compute Shader and perform image fusion processing, which includes: dynamically calculating the fusion weights of the left and right views based on image content features; performing a weighted operation based on the previous frame image and the current frame image, where the weight of the weighted operation is determined by the adaptive smoothing coefficient; performing parallax compensation and pixel alignment on the right-eye image based on the depth map; fusing the left and right eye images according to the fusion weights to generate a fused image; and overlaying the UI layer according to the transparency of the user interface texture to obtain a composite display image. Step S4: writing the synthesized display image into a synthesized display texture, and outputting the synthesized display texture to a stereoscopic display device, wherein the stereoscopic display device supports a frame sequential, shutter or line alternating stereoscopic image display mode.

2. The image fusion stereoscopic display method based on the Unity engine according to claim 1, characterized in that: The step S1 comprises: The main camera is copied to generate a left-eye camera and a right-eye camera respectively. The left-eye camera and the right-eye camera are symmetrically offset relative to the main camera in the horizontal direction, and the offset distance is an adjustable parameter; Adjust the pupil distance parameter value in real time through the control interface, and dynamically update the horizontal position offset of the left-eye camera and the right-eye camera; The field of view angles of the left-eye camera and the right-eye camera are both set to 60 degrees; The left-eye camera and the right-eye camera respectively capture left-view images and right-view images, and render to generate left-eye rendering textures and right-eye rendering textures, where the rendering textures are two-dimensional textures that can be read and written by a GPU.

3. The image fusion stereoscopic display method based on the Unity engine according to claim 1, characterized in that: The step S2 comprises: Identify the Canvas type of the UI element based on the display mode tag. When the Canvas type is Overlay mode, generate an independent user interface texture bound to the Render Texture object. When the Canvas type is World Space mode or Screen Space-Camera mode, the left-eye camera and the right-eye camera capture images separately and generate a unified user interface texture. Automatically configure the rendering parameters of the user interface texture according to the Canvas type, wherein the rendering parameters include rendering priority, scaling factor, and occlusion level information.

4. The image fusion stereoscopic display method based on the Unity engine according to claim 1, characterized in that: The step S3 comprises: The left-eye rendering texture, the right-eye rendering texture, and the user interface texture are input to the Compute Shader, and image fusion processing is performed in the Compute Shader. The image fusion processing includes the following steps: Dynamically calculate the fusion weights of the left and right views based on image content features; Calculate the parallax compensation displacement of the right eye image based on the depth map and perform pixel alignment; Fusing the aligned left and right images to obtain a fused image; The user interface image is superimposed on the fused image according to its transparency information to generate the final display image. The fusion calculation formula is: ; in: is the left eye image at pixel location The color value of For the right eye image at pixel location The color value of For user interface images at pixel locations The color value of For user interface images at pixel locations The alpha channel value of is the weight coefficient for left and right image fusion; For the right eye image at pixel location Parallax compensation displacement at ; is the pixel value after the left and right images are fused; For the composite display image at pixel location The color value of .

5. The image fusion stereoscopic display method based on the Unity engine according to claim 4, characterized in that: The step S3 further comprises: Obtain the fusion weight change, parallax compensation displacement change, and user visual parameters between the current frame and the previous frame; calculate the adaptive inter-frame smoothing coefficient for each pixel position based on the basic smoothing coefficient, weight change adjustment coefficient, parallax change adjustment coefficient, and user parameter adjustment coefficient; Based on the adaptive inter-frame smoothing coefficient, a weighted operation is performed on corresponding pixel values ​​of the current frame fusion image and the previous frame fusion image to generate an output pixel value of the current frame.

6. The image fusion stereoscopic display method based on the Unity engine according to claim 1, characterized in that: The step S4 comprises: The synthesized display texture is bound as an output texture to a target frame buffer of a graphics rendering pipeline and transmitted to a stereoscopic display device; the stereoscopic display device supports a frame sequential, shutter or line alternating stereoscopic image display mode.

7. The image fusion stereoscopic display method based on the Unity engine according to claim 1, characterized in that: The left eye rendering texture, right eye rendering texture, user interface texture and synthetic display texture are all two-dimensional textures that can be read and written by the GPU, and the data type is Texture2D <float4> , and is configured as a continuous memory layout to support multi-threaded parallel access and fast processing of image data.

8. The image fusion stereoscopic display method based on the Unity engine according to claim 1, characterized in that: The method includes a resource adaptation step, which includes: Configuring a metadata component including a display mode flag and a scaling factor for each user interface element, wherein the display mode flag is used to indicate a display mode of the interface element in the left-eye image and the right-eye image, wherein the display modes include binocular synchronization, left-eye display only, and right-eye display only; In the resource preprocessing stage before image rendering, the editor extension script is used to identify the Canvas type bound to each user interface element, and the RectTransform property and Canvas Renderer property of the interface element are automatically adjusted according to the content in the metadata component; before image fusion processing, the horizontal pixel offset of each user interface element in the left and right images is calculated based on the user's pupil distance parameter and gaze point depth parameter. , used to adjust the position of interface elements in the user interface texture. The formula is: ; in: is the user's pupil distance parameter; is the gaze point depth parameter; The distance from the user's eyes to the screen; is the scaling factor of the interface elements; is the calibration parameter.

9. An image fusion stereoscopic display system based on the Unity engine, applied to the image fusion stereoscopic display method based on the Unity engine according to any one of claims 1 to 8, characterized in that: The system comprises: The binocular rendering module is used to configure the left-eye camera and the right-eye camera in the Unity scene, and respectively capture the left-view image and the right-view image based on the set pupil distance parameters to generate the left-eye rendering texture and the right-eye rendering texture; The interface texture generation module is used to identify the corresponding display mode mark according to the Canvas type bound to each UI element in the user interface, and generate the user interface texture; an image fusion module, configured to perform image fusion processing, the image fusion processing including a weighted operation based on a previous frame image and a current frame image, wherein the weight of the weighted operation is determined by an adaptive smoothing coefficient, and superimposing the fused images based on the alpha channel information of the user interface texture to generate a synthetic display texture; A time-continuous fusion module is used to perform weighted operations on corresponding pixel values ​​of the previous frame fusion image and the current frame fusion image in the image fusion module; The resource adapter module is used to configure metadata components including display mode flags and scaling factors for each UI element. During the resource preprocessing phase, it uses the editor extension script to identify the Canvas type bound to the UI element and adjust the RectTransform and Canvas Renderer properties. A position control module is used to calculate the horizontal offset of each UI element in the composite image according to the user's pupil distance parameter and the gaze point depth parameter before the fusion process, and adjust the position of the corresponding UI element in the interface texture based on the offset; The scheduling control module is used to calculate the number of two-dimensional thread groups according to the target display resolution and trigger ComputeShader to process image fusion tasks in parallel; An output module is used to bind the synthesized display texture to a target frame buffer of a graphics rendering pipeline and transmit it to a stereoscopic display device that supports a frame sequential, shutter or line alternating display mode.

10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, is used to implement the image fusion stereoscopic display method based on the Unity engine according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image processing method and image processing device

    CN103294453A

  • Naked eye 3D display system based on Unity3D game engine

    CN103957400A

  • Information display method and device, equipment and storage medium

    CN112965773A

  • Dynamic evolution analysis method suitable for graphical user side demand response strategy

    CN120387206A

  • Three-dimensional scene synthesis method fusing depth estimation and double-view video

    CN120431274A

Cited By

  • Remote 3D image data transmission and display method and device based on Internet

    CN121771380A

  • Display method and system, electronic equipment and storage medium

    CN122027777A