Stereoscopic image display system and method for visualizing a stereoscopic scene
Patent Information
- Application Number
- CN202510351262.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2026-09-25
AI Technical Summary
然而,目前市面上的3D图像内容并不充足,因而即便使用者具有一台3D显示器,但使用者还是无法充分且任意享受3D显示器带来的显示效果
[0006]基于上述,于本揭示实施例中,可根据单目深度估测的深度估测结果与立体图像对的立体深度信息来产生立体场景网格。因此,可通过单目深度估测来弥补双视角深度估测的遮挡与纹理重复问题,也可通过立体图像对的立体深度信息来优化单目深度估测的估测结果。基此,基于单目深度估测与立体图像对的立体深度信息产生的立体场景网格,不仅可让使用者感知到符合场景类型的景深,并可实现高精度、灵活且高效的3D场景生成与可视化。
Smart Images

Figure CN122824883A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an image processing technology, and more particularly to a stereoscopic image display system and a method for visualizing stereoscopic scenes. Background Technology
[0002] With advancements in display technology, stereoscopic displays supporting stereoscopic vision technology have become increasingly common. Stereoscopic vision technology allows viewers to experience the three-dimensionality of images, such as the three-dimensional features of a person and depth of field, effects that traditional 2D images cannot achieve. The principle of stereoscopic vision technology is to allow the viewer's left eye to see the left-eye image and their right eye to see the right-eye image, thus creating a 3D visual effect. 3D displays can provide separate left-eye and right-eye images to the viewer's left and right eyes respectively, providing a visually immersive experience. However, the current market supply of 3D image content is insufficient, meaning that even with a 3D display, users cannot fully and freely enjoy its display effects. While technologies exist for generating 3D content from monocular images, they often suffer from blurred or incomplete image edges. Furthermore, the viewing angle and depth of existing 3D image content are fixed, resulting in a relatively inflexible 3D visual effect. Summary of the Invention
[0003] This disclosure provides a stereoscopic image display system and a stereoscopic scene visualization method that can effectively solve the above problems.
[0004] This disclosed exemplary embodiment provides a method for visualizing a stereo scene, comprising the following steps: Acquiring a left-eye image and a right-eye image from a stereo image pair. Performing monocular depth estimation on the left-eye and right-eye images respectively to obtain a left-eye depth map and a right-eye depth map. Generating stereo depth information based on the disparity information between the left-eye and right-eye images. Generating a left-side stereo mesh for the left-eye image based on the left-eye depth map, and generating a right-side stereo mesh for the right-eye image based on the right-eye depth map. Generating a stereo scene mesh based on the stereo depth information, the left-side stereo mesh, and the right-side stereo mesh. Outputting stereo scene content based on the stereo scene mesh.
[0005] Another exemplary embodiment disclosed herein provides a stereoscopic image display system including a stereoscopic display and at least one processor. The processor is coupled to the stereoscopic display and configured to perform the following operations: acquiring a left-eye image and a right-eye image from a stereoscopic image pair; performing monocular depth estimation on the left-eye and right-eye images respectively to obtain a left-eye depth map and a right-eye depth map; generating stereoscopic depth information based on disparity information between the left-eye and right-eye images; generating a left-side stereo grid for the left-eye image based on the left-eye depth map, and generating a right-side stereo grid for the right-eye image based on the right-eye depth map; generating a stereoscopic scene grid based on the stereoscopic depth information, the left-side stereo grid, and the right-side stereo grid; and outputting stereoscopic scene content through the stereoscopic display based on the stereoscopic scene grid.
[0006] Based on the above, in this disclosed embodiment, a stereo scene mesh can be generated based on the depth estimation result of monocular depth estimation and the stereo depth information of the stereo image pair. Therefore, monocular depth estimation can compensate for the occlusion and texture duplication problems of dual-view depth estimation, and the stereo depth information of the stereo image pair can optimize the estimation result of monocular depth estimation. Accordingly, the stereo scene mesh generated based on monocular depth estimation and the stereo depth information of the stereo image pair not only allows users to perceive a depth of field consistent with the scene type, but also enables high-precision, flexible, and efficient 3D scene generation and visualization. Attached Figure Description
[0007] Figure 1 This is a schematic diagram of a stereoscopic image display system according to an embodiment of the present disclosure;
[0008] Figure 2 This is a schematic diagram of a stereoscopic display according to an embodiment of the present disclosure;
[0009] Figure 3 This is a flowchart of a three-dimensional scene visualization method according to an embodiment of this disclosure;
[0010] Figure 4 This is a schematic diagram of a stereoscopic scene visualization method according to an embodiment of this disclosure;
[0011] Figure 5 This is a schematic diagram illustrating the generation of stereoscopic scene content based on the left-eye image and the right-eye image according to an embodiment of this disclosure;
[0012] Figure 6 This is a flowchart illustrating the generation of a 3D scene mesh according to an embodiment of this disclosure;
[0013] Figure 7 This is a flowchart illustrating the determination of texture information according to an embodiment of this disclosure;
[0014] Figure 8 This is a schematic diagram illustrating the determination of texture information according to an embodiment of this disclosure;
[0015] Figure 9 This is a flowchart illustrating the output of a stereoscopic scene content according to an embodiment of this disclosure;
[0016] Figure 10 This is a flowchart of the output stereoscopic scene content according to an embodiment of the present disclosure. Detailed Implementation
[0017] Reference will now be made in detail to exemplary embodiments of the invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same component reference numerals are used in the drawings and description to denote the same or similar parts.
[0018] Figure 1 This is a schematic diagram of a stereoscopic image display system according to an embodiment of this disclosure. Please refer to... Figure 1 The stereoscopic image display system 100 may include a stereoscopic display 110, a storage device 120, and at least one processor 130. In different embodiments, the stereoscopic image display system 100 may be implemented as an integrated system or a discrete system. In some embodiments, the stereoscopic display 110, storage device 120, and processor 130 may be implemented as an all-in-one electronic device, such as a laptop computer, tablet computer, desktop computer, game console, portable electronic device, or other personal electronic device. Alternatively, in some embodiments, the stereoscopic display 110 may be connected to a calculator device including the storage device 120 and the processor 130 via a wired or wireless transmission interface.
[0019] The stereoscopic display 110 allows users to experience stereoscopic visual effects. In order to allow users to experience 3D visual effects through the stereoscopic display 110, the stereoscopic display 110 can, according to its hardware specifications and the 3D display technology it applies, allow the user's left and right eyes to view image content corresponding to different perspectives (i.e., left-eye image and right-eye image).
[0020] In some embodiments, the stereoscopic display 110 may be a naked-eye 3D display, such as a laptop screen, television, desktop screen, or electronic billboard, etc. In some embodiments, the left-eye image and the right-eye image may be displayed simultaneously based on stereoscopic image display technology, such as parallax barrier technology, lens technology, or directional backlight technology. Alternatively, in some embodiments, the stereoscopic display 110 may be a head-mounted display device, such as a virtual reality display device or a mixed reality display device, etc.
[0021] On the other hand, the stereoscopic display 110 may include a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, or other types of displays, and this disclosure is not limited thereto.
[0022] Storage device 120 is used to temporarily or permanently store data, such as images, instructions, program code, software modules, etc. Specifically, storage device 120 may include volatile storage circuitry. Volatile storage circuitry is used to store data in a volatile manner. For example, volatile storage circuitry may include random access memory (RAM) or similar volatile storage media. Alternatively, storage device 120 may include non-volatile storage circuitry. Non-volatile storage circuitry is used to store data in a non-volatile manner. For example, non-volatile storage circuitry may include read-only memory (ROM), solid-state drive (SSD), and / or traditional hard disk drive (HDD) or similar non-volatile storage media. The number of storage devices 120 may be one or more, and this disclosure is not limited thereto.
[0023] Processor 130 connects stereoscopic display 110 and storage device 120. For example, processor 130 may include a central processing unit (CPU), graphics processing unit (GPU) or other programmable general-purpose or special-purpose microprocessor, digital signal processor (DSP), programmable controller, application-specific integrated circuit (ASIC), programmable logic device (PLD), or other similar device or combination of these devices. The number of processors 130 may be one or more, and this disclosure does not limit this.
[0024] Figure 2 This is a schematic diagram of a stereoscopic display according to an embodiment of this disclosure. Please refer to... Figure 2In some embodiments, the stereoscopic display 110 may be a naked-eye stereoscopic display, which provides different images to the left and right eyes through the principle of lens refraction, allowing the viewer to experience a stereoscopic display effect. The stereoscopic display 110 may include a display panel 111 and a lens layer 112. The lens layer 112 is disposed above the display panel 111, and the viewer can see the image content provided by the display panel 111 through the lens layer 112. The stereoscopic display 110 may place the pixels of the first left-eye image and the pixels of the first right-eye image at the corresponding pixel positions on the display panel 111. The lens layer 112 refracts different display content (i.e., the left-eye image and the right-eye image) to different positions in space through light refraction, allowing the left and right eyes to receive two different images with parallax. As is known, in order to place the pixels of the left-eye image and the right-eye image at the corresponding pixel positions on the display panel 111, the left-eye image and the right-eye image need to undergo image weaving processing to produce a woven frame in which the pixels of the left-eye image and the pixels of the right-eye image are arranged alternately.
[0025] Figure 3 This is a flowchart of a three-dimensional scene visualization method according to an embodiment of this disclosure. Please refer to... Figure 3 The operation process of this embodiment is applicable to the stereoscopic image display system 100 in the above embodiment. The following describes the detailed steps of this embodiment in conjunction with the various components in the stereoscopic image display system 100.
[0026] In step S310, the processor 130 acquires the left-eye and right-eye images from a stereo image pair. In some embodiments, the stereo image pair can be generated by various devices, such as a stereo camera module, a VR / AR device, or a dual-lens module in a mobile device. These dual-lens devices capture the left-eye and right-eye images of a scene using preset baseline distances and camera parameters, forming a stereo image pair with parallax features. Generally, to ensure the accuracy of subsequent processing, the acquisition of the left-eye and right-eye images needs to meet the requirements of synchronization, resolution, and format consistency.
[0027] In some embodiments, the stereo image pair may be side-by-side (SBS) images conforming to a stereo image format. The SBS image includes left-eye and right-eye images from different perspectives.
[0028] In step S320, the processor 130 performs monocular depth estimation on the left-eye image and the right-eye image respectively to obtain a left-eye depth map and a right-eye depth map. Specifically, in some embodiments, the processor 130 may perform a monocular depth estimation on the left-eye image and the right-eye image respectively to obtain the depth information of the left-eye image and the right-eye image respectively. By performing monocular depth estimation, the processor 130 can estimate the depth information of the shooting scene based on a single-view image.
[0029] In some embodiments, processor 130 may input a left-eye image into a monocular depth estimation model to obtain a left-eye depth map. Furthermore, processor 130 may input a right-eye image into a monocular depth estimation model to obtain a right-eye depth map. That is, processor 130 may utilize a deep learning model to perform monocular depth estimation on the input image frame. Processor 130 may input an input image frame (i.e., a left-eye image or a right-eye image) into a trained monocular depth estimation model to obtain a depth map of the input image frame.
[0030] Alternatively, in some embodiments, the processor 130 may utilize other conventional vision algorithms to perform monocular depth estimation on the left-eye and right-eye images separately. For example, the processor 130 may analyze image features or object motion trajectories at different scales in the left-eye or right-eye image to estimate the depth information of the left-eye or right-eye image.
[0031] It should be noted that the depth values in the depth information obtained from monocular depth estimation have been normalized and fall within a preset range. For example, the depth values in the left-eye depth map and the right-eye depth map can be between 0 and 255.
[0032] In step S330, the processor 130 generates stereo depth information based on the disparity information between the left-eye and right-eye images. In some embodiments, the stereo depth information may be a depth map. In some embodiments, the disparity information may be obtained by disparity calculation between the left-eye and right-eye images to generate stereo depth information corresponding to the stereo scene. The stereo depth information can be used to describe the three-dimensional structure and spatial position of objects in the scene, providing a basis for subsequent mesh generation and visualization.
[0033] Regarding disparity calculation, the processor 130 can calculate the disparity value for each pair of corresponding pixels by matching the pixel positions in the left-eye and right-eye images. This disparity value represents the horizontal displacement of the same object in the left-eye and right-eye images and is inversely proportional to the depth of the object. The formula for calculating disparity is as follows (1).
[0034]
[0035] Where disp represents the parallax value, IPD represents the baseline distance of the camera (the distance between the left and right cameras), z represents the depth, and f represents the focal length of the camera.
[0036] In some embodiments, the processor 130 may perform disparity calculation using a stereo matching algorithm. This matching method is used to perform pixel-by-pixel matching between pixels in the left-eye image and pixels in the right-eye image, and may include block matching, global optimization, or deep learning-based methods. Block matching may search for the optimal matching position using a sliding window based on the texture features of the image region. Global optimization may be designed based on an energy function, considering the balance between matching cost and smoothness constraints.
[0037] In step S340, the processor 130 generates a left-side stereo grid of the left-eye image based on the left-eye depth map, and generates a right-side stereo grid of the right-eye image based on the right-eye depth map.
[0038] In some embodiments, the processor 130 can map multiple pixels of the left-eye image to multiple first stereo coordinates in a stereo coordinate system based on camera intrinsic parameters and the left-eye depth map, to construct a left-side stereo mesh of the left-eye image. Furthermore, the processor 130 can map multiple pixels of the right-eye image to multiple second stereo coordinates in a stereo coordinate system based on camera intrinsic parameters and the right-eye depth map, to construct a right-side stereo mesh of the right-eye image. Camera intrinsic parameters describe the internal geometry and optical characteristics of the camera and are primarily used to define the relationship between the image plane and the camera coordinate system. In some embodiments, the stereo coordinate system can be the camera coordinate system or other stereo coordinate systems established by transforming the camera coordinate system.
[0039] In some embodiments, the mapping between the left and right stereo meshes is performed within a preset depth range. That is, based on the preset depth range, the processor 130 can map the pixels of the left-eye image to multiple first stereo coordinates (e.g., camera coordinates) in a stereo coordinate system according to the left-eye depth map and camera intra-parameters. These first stereo coordinates can constitute the vertices of the left stereo mesh. Then, during the construction of the left stereo mesh, the processor 130 can further combine the relationships between neighboring pixels to generate a polygonal mesh surface to fully describe the structure of the left stereo mesh. In some embodiments, the processor 130 can use a similar method to generate the right stereo mesh.
[0040] In some embodiments, the processor 130 may obtain normalized depth values for each pixel in the left-eye image based on the left-eye depth map. The processor 130 may adjust the normalized depth values of each pixel according to a preset depth range in the projection parameters to generate adjusted depth values. The range of normalized depth values in the left-eye depth map will be scaled to the preset depth range. Based on the camera intra-parameters in the projection parameters and the adjusted depth values of each pixel, the processor 130 may project each left-eye pixel onto a stereo coordinate system to generate multiple first stereo coordinates corresponding to the multiple left-eye pixels. In some embodiments, the processor 130 may use a similar method to map multiple pixels of the right-eye image to multiple second stereo coordinates in a stereo coordinate system.
[0041] In detail, based on the back projection principle of the pinhole camera model, the processor 130 needs the camera's intrinsic parameters and pixel depth information to project the pixels in the left-eye and right-eye images onto a three-dimensional coordinate system. The camera's intrinsic parameters can be a camera intrinsic parameter matrix, including focal length information and principal point position in the x-axis and y-axis directions on the image plane.
[0042] Specifically, the processor 130 can project the pixels on the left-eye image and the right-eye image onto the stereo coordinate system according to the following formulas (2) to (4).
[0043] z′=z near +(z far -z near Formula (2)
[0044]
[0045] Where Znear represents the near-plane scene depth; Zfar represents the far-plane scene depth; z represents the normalized depth value generated by monocular depth estimation; z' represents the adjusted depth value; fx and fy represent the focal lengths along the x and y axes of the image plane in the intrinsic parameter matrix; cx and cy represent the principal point coordinates on the image plane. The processor 130 can project the pixels (x, y) in the left-eye image and the right-eye image to the first / second stereo coordinates (x', y', z') in the stereo coordinate system according to formulas (2) to (4).
[0046] In step S350, processor 130 generates a stereo scene mesh based on the stereo depth information, the left stereo mesh, and the right stereo mesh. In some embodiments, processor 130 can map the left stereo mesh and the right stereo mesh to the world coordinate system respectively to ensure that their coordinate references are consistent. This mapping process, combined with camera extrinsic parameters (rotation matrix R and translation vector T), transforms the vertices in the stereo coordinate system to a unified world coordinate system, as shown in formula (5) below.
[0047] P world =R·P camera +T formula (5)
[0048] Where P camera Let P be the first solid coordinate of the left solid grid or the second solid coordinate of the right solid grid, and P world These are the mapped world coordinates.
[0049] Next, the processor 130 can perform vertex matching and depth integration based on the world coordinates of the mapped left and right 3D meshes to generate a 3D scene mesh. The 3D scene mesh generated by the processor 130 can contain complete vertex, mesh surface, depth data and texture mapping information, which can support subsequent 3D scene visualization operations.
[0050] In some embodiments, the stereo depth information generated based on binocular depth estimation can be used to correct the depth components of the first stereo coordinates of the left stereo mesh and the second stereo coordinates of the right stereo mesh. Alternatively, in some embodiments, the stereo depth information generated based on binocular depth estimation can be used to perform a weighted operation with the depth components of the first stereo coordinates and the second stereo coordinates of the right stereo mesh to generate the depth of each vertex in the stereo scene mesh.
[0051] In step S360, processor 130 outputs stereoscopic scene content based on the stereoscopic scene mesh. Processor 130 can convert the generated stereoscopic scene mesh into a format suitable for display or further processing, and adjust the output content according to application requirements to meet different scene needs. For example, processor 130 can convert the stereoscopic scene mesh into a single-view 2D rendered image based on a specific viewing angle. Alternatively, processor 130 can generate two images based on the stereoscopic scene mesh, one from a left-eye perspective and the other from a right-eye perspective, so that the user can experience a stereoscopic visual effect.
[0052] Figure 4 This is a schematic diagram of a stereoscopic scene visualization method according to an embodiment of this disclosure. Figure 5 This is a schematic diagram illustrating the generation of stereoscopic scene content based on the left-eye and right-eye images according to an embodiment of this disclosure. Please refer to it as well. Figure 4 and Figure 5 .
[0053] In operation 41, processor 130 performs monocular depth estimation on the left-eye image Img_L to generate a left-eye depth map Dm_L. In operation 42, processor 130 performs monocular depth estimation on the right-eye image Img_R to generate a right-eye depth map Dm_R. In operation 42, processor 130 performs binocular depth estimation based on the left-eye image Img_L and the right-eye image Img_R to generate stereo depth information Dm_S.
[0054] In operation 44, processor 130 can map the left-eye image Img_L to a stereo coordinate system according to a preset depth range and the left-eye depth map Dm_L to generate a left-side stereo mesh msh_L1. In operation 45, processor 130 can map the right-eye image Img_R to a stereo coordinate system according to a preset depth range and the right-eye depth map Dm_R to generate a right-side stereo mesh msh_R1.
[0055] In operation 46, the processor 130 can combine the three-dimensional depth information Dm_S, the left three-dimensional mesh msh_L1 and the right three-dimensional mesh msh_R1 to generate the three-dimensional scene mesh msh_S1.
[0056] For details, please refer to Figure 6 This is a flowchart of generating a stereo scene mesh according to an embodiment of the present disclosure. In step S610, the processor 130 can map the left stereo mesh msh_L1 to a world coordinate system to obtain multiple left world coordinate points. In some embodiments, by utilizing the camera extrinsic parameters of the left lens, the processor 130 can transform the left stereo vertices in the camera coordinate system (i.e., the stereo coordinate system) to a unified world coordinate system.
[0057] In step S620, the processor 130 can map the right-side stereo mesh msh_R1 to the world coordinate system to obtain multiple right-side world coordinate points. In some embodiments, by utilizing the camera extrinsic parameters of the right-side lens, the processor 130 can transform the right-side stereo vertices in the camera coordinate system (i.e., the stereo coordinate system) to a unified world coordinate system.
[0058] In step S630, the processor 130 can merge multiple left-side world coordinate points and multiple right-side world coordinate points based on the matching relationship between them to generate a stereo scene mesh msh_S1. Specifically, during the matching process, the processor 130 needs to determine which left-side world coordinate points and right-side world coordinate points correspond to the same scene point, so as to generate a vertex in the stereo scene mesh msh_S1 based on a left-side world coordinate point and a right-side world coordinate point corresponding to the same scene point.
[0059] In some embodiments, when the Euclidean distance between the first left world coordinate point and the first right world coordinate point is less than a matching distance threshold, the processor 130 may match the first left world coordinate point and the first right world coordinate point with each other. The matched first left world coordinate point and the first right world coordinate point may be merged to generate a vertex in the stereo scene mesh msh_S1.
[0060] In some embodiments, the processor 130 may perform an average or weighted operation on the X-coordinate component of the first left world coordinate point and the X-coordinate component of the first right world coordinate point to generate the X-coordinate component of a vertex in the stereo scene mesh msh_S1. The processor 130 may also perform an average or weighted operation on the Y-coordinate component of the first left world coordinate point and the Y-coordinate component of the first right world coordinate point to generate the Y-coordinate component of a vertex in the stereo scene mesh msh_S1.
[0061] In some embodiments, since the left-eye image Img_L and the right-eye image Img_R correspond to different viewpoints, some scene points in the scene may only exist in either the left-eye image Img_L or the right-eye image Img_R due to occlusion effect. That is, a right-side world coordinate point in the right-side stereo mesh msh_R1 may not match any left-side world coordinate point in the left-side stereo mesh msh_L1. Alternatively, a left-side world coordinate point in the left-side stereo mesh msh_L1 may not match any right-side world coordinate point in the right-side stereo mesh msh_R1. In this case, the processor 130 can directly retain the unmatched right-side world coordinate point or the unmatched left-side world coordinate point in the stereo scene mesh msh_S1. For example, multiple left-side world coordinate points of the left edge block of the left-side stereo mesh msh_L1 can be directly retained in the stereo scene mesh msh_S1. Similarly, multiple right-side world coordinate points of the right edge block of the right-side stereo mesh msh_R1 can be directly retained in the stereo scene mesh msh_S1.
[0062] In some embodiments, the processor 130 can utilize stereo depth information to correct depth discrepancies between multiple left-side world coordinate points and multiple right-side world coordinate points. Specifically, because the left-eye image Img_L and the right-eye image Img_R may have inconsistencies, a depth discrepancy exists between the right-eye depth map Dm_R and the left-eye depth map Dm_L generated by the monocular depth estimation model. Furthermore, the depth value "100" in the right-eye depth map Dm_R and the depth value "100" in the left-eye depth map Dm_L actually correspond to different real-world scene depths. Therefore, the processor 130 can utilize stereo depth information Dm_S to correct the depth discrepancy between multiple left-side world coordinate points and multiple right-side world coordinate points, thereby making the depth values in the stereo scene mesh msh_S1 more accurate.
[0063] In some embodiments, the processor 130 may perform a weighted operation on the depth components of a matching right-side world coordinate point and a matching left-side world coordinate point, as well as the depth value in the stereo depth information Dm_S, to generate the depth component of a vertex in the stereo scene mesh msh_S1.
[0064] In some embodiments, the processor 130 may calculate the depth difference information between the stereo depth information Dm_S and the left-eye depth map Dm_L to generate a left-side depth scaling transformation function. The processor 130 may adjust all depth components of the left-side stereo mesh msh_L1 according to the left-side depth scaling transformation function. In some embodiments, the processor 130 may calculate the depth difference information between the stereo depth information Dm_S and the right-eye depth map Dm_R to generate a right-side depth scaling transformation function. The processor 130 may adjust all depth components of the right-side stereo mesh msh_R1 according to the right-side depth scaling transformation function. Then, the processor 130 may perform an averaging or weighted calculation on the scaled-adjusted depth components in the left-side stereo mesh msh_L1 and the scaled-adjusted depth components in the right-side stereo mesh msh_R1 to generate the depth component of a vertex in the stereo scene mesh msh_S1.
[0065] In some embodiments, the processor 130 can measure the near-plane scene depth z of the left stereo mesh msh_L1. min,L Depth z of the far-plane scene max,L The near-plane scene depth z of the right-side 3D mesh msh_R1 min,R Depth z of the far-plane scene max,R The target parameters are set as optimization parameters. Using an optimization algorithm (such as gradient descent), the processor 130 can obtain the near-plane scene depth z based on the left-eye image Img_L, the right-eye image Img_R, and the stereo depth information Dm_S. min,L The depth of the far-plane scene z max,L Near-planar scene depth zmin,R Depth z of the far-plane scene max,R The best solution.
[0066] Gradient descent, as a numerical optimization method, is suitable for solving nonlinear problems with multiple unknown parameters. In these embodiments, the processor 130 first sets an initial depth range parameter (i.e., the planar scene depth z). min,L The depth of the far-plane scene z max,L Near-planar scene depth z min,R Depth z of the far-plane scene max,R (Initial value). Next, the processor 130 can calculate the left pixel plane coordinates and right pixel plane coordinates corresponding to the same stereo scene coordinate point based on the depth range parameter, as shown in formula (6) and formula (7), and the stereo scene coordinate point can be calculated based on the stereo depth information Dm_S.
[0067]
[0068] Where f represents the camera focal length, IPD represents the baseline distance between the cameras (the distance between the left and right cameras), and x represents the distance between the left and right cameras. i The x-coordinate represents the 3D scene coordinates, while f represents the camera's focal length, and d... i,L The depth value represented by d in the left eye depth map. i,R This represents the depth value of the right eye depth map, where i represents the vertex index. i,L The X-coordinate representing the left pixel plane coordinates, u i,R The X-coordinate representing the plane coordinates of the right pixel.
[0069] Then, the processor 130 can calculate the error value based on the color of the left pixel plane coordinates in the left-eye image Img_L and the color of the right pixel plane coordinates in the right-eye image Img_R. Through multiple iterations, gradient descent gradually updates the parameter values to minimize the above error value until it converges to a preset threshold or reaches the maximum number of iterations.
[0070] In this way, processor 130 can determine the depth z of the near-plane scene. min,L The depth of the far-plane scene z max,L The left-side depth scaling transformation function is calculated using the minimum and maximum scene depths in the stereo depth information Dm_S. The processor 130 can then calculate the left-side depth scaling transformation function based on the near-plane scene depth z. min,R The depth of the far-plane scene z max,R The right-side depth scaling transformation function is calculated using the minimum and maximum scene depths in the stereo depth information Dm_S.
[0071] Alternatively, in some embodiments, the processor 130 may calculate a first average of the depth values in the stereo depth information Dm_S located within the near-field distance range, and a second average of the depth values in the stereo depth information Dm_S located within the far-field distance range. Furthermore, the processor 130 may calculate a third average of the depth values in the right stereo mesh msh_R1 located within the near-field distance range, and a fourth average of the depth values in the right stereo mesh msh_R1 located within the far-field distance range. Thus, the processor 130 may establish a right-side depth scaling transformation function based on the scaling ratio between the first and third averages and the scaling ratio between the second and fourth averages. Similarly, the processor 130 may calculate a fifth average of the depth values in the left stereo mesh msh_L1 located within the near-field distance range, and a sixth average of the depth values in the left stereo mesh msh_L1 located within the far-field distance range. Thus, the processor 130 may establish a left-side depth scaling transformation function based on the scaling ratio between the first and fifth averages and the scaling ratio between the second and sixth averages. The right-side depth scaling transformation function and the left-side depth scaling transformation function can be used to correct the depth of the right-side 3D mesh msh_R1 and the left-side 3D mesh msh_L1, respectively.
[0072] Back Figure 4 In operation 47, processor 130 can perform texture rendering on the stereo scene mesh msh_S1. Specifically, processor 130 can render textures on each vertex of the stereo scene mesh msh_S1 based on the texture information of the left eye image Img_L and the texture information of the right eye image Img_R.
[0073] For details, please refer to Figure 7This is a flowchart illustrating the determination of texture information according to an embodiment of the present disclosure. In step S710, the processor 130 obtains the normal vectors of the mesh surface corresponding to each vertex in the stereo scene mesh msh_S1. In step S720, the processor 130 determines the pixel texture information of each vertex based on the normal vectors corresponding to each vertex, a preset left-eye view, and a preset right-eye view. By comparing the directional similarity between the normal vectors corresponding to each vertex and the preset left-eye view, and the directional similarity between the normal vectors corresponding to each vertex and the preset right-eye view, the processor 130 determines whether to use the texture information of the left-eye image Img_L or the right-eye image Img_R. For example, if the directional similarity between the normal vector corresponding to a vertex and the preset left-eye view is high, the processor 130 can directly use the texture information of the corresponding pixel in the left-eye image Img_L for the texture information of that vertex. Alternatively, if the normal vector corresponding to a vertex has a high directional similarity to the preset right eye view, the processor 130 can multiply the texture information of the corresponding pixel in the left eye image Img_L by a larger weighting factor and multiply the texture information of the corresponding pixel in the right eye image Img_L by a smaller weighting factor to generate the texture information of that vertex.
[0074] In some embodiments, each vertex in the stereo scene mesh msh_S1 may be associated with multiple mesh surfaces, and the processor 130 may select one of the normal vectors of these mesh surfaces to determine the texture information. Alternatively, in some embodiments, each vertex in the stereo scene mesh msh_S1 may be associated with multiple mesh surfaces, and the processor 130 may choose to sum these normal vectors of these mesh surfaces to determine the texture information based on the summed normal vector.
[0075] In some embodiments, step S720 may be implemented as steps S721 to S723. In step S721, the processor 130 determines the left-side texture weight based on the angle between the normal vector corresponding to the first vertex and a preset left-eye viewpoint. In step S722, the processor 130 determines the right-side texture weight based on the angle between the normal vector corresponding to the first vertex and a preset right-eye viewpoint. In step S722, the processor 130 performs a weighted calculation on the texture information of the left-eye image Img_L and the texture information of the right-eye image Img_R based on the left-side and right-side texture weights to obtain the pixel texture information of the first vertex.
[0076] For example, the processor 130 may determine the pixel texture information of the first vertex according to the following formula (8).
[0077] I x,y =f(n) x,y,z ·d l )L x,y +g(n x,y,z ·dr )R x,y Formula (8)
[0078] Among them, I x,y Pixel texture information representing the first vertex; L x,y Represents the texture information of the left eye image Img_L; R x,y Represents the texture information of the right eye image Img_R; n x,y,z d represents the normal vector corresponding to the first vertex; l This represents the default left-eye perspective; while d r This represents the preset right-eye perspective. Processor 130 can calculate n. x,y,z With d l The inner product of n represents the angle information between the normal vector and the preset left-eye view. Processor 130 can calculate n x,y,z With d r The inner product of represents the angle information between the normal vector and the preset right-eye viewpoint. The right-side texture weight is equal to g(n). x,y,z ·d r The texture weight on the left side is equal to f(n). x,y,z ·d l g(·) is a predefined function, and f(·) is another predefined function.
[0079] For example, Figure 8 This is a schematic diagram illustrating the determination of texture information according to an embodiment of this disclosure. Please refer to... Figure 8 The processor 130 can calculate n x,y,z With d l The inner product of n represents the angle information between the normal vector and the preset left-eye view. Processor 130 can calculate n x,y,z With d r The inner product of n and n represents the angle between the normal vector and the preset right-eye view. In this example, n x,y,z With d l Because the angle is smaller, the texture weight on the left side is greater than the texture weight on the right side.
[0080] Back Figure 4 In operation 48, processor 130 can generate visualization content based on the 3D scene mesh msh_S1 with texture information to output 3D scene content VC1. Figure 5 In the example, the stereoscopic scene content (VC) can include side-by-side images of content from different perspectives (Img_SBS) or a two-dimensional rendered image (Img_2D) corresponding to a specific perspective.
[0081] Please refer to Figure 9This is a flowchart of the output of stereoscopic scene content according to an embodiment of the present disclosure. In step S910, the processor 130 generates a first-view image and a second-view image of side-by-side images Img_SBS based on the stereoscopic scene mesh msh_S1. In step S920, the processor 130 performs 3D display operations based on the side-by-side images Img_SBS using a stereoscopic display 110.
[0082] In detail, the processor 130 can generate a first-view image and a second-view image based on multiple scene stereo coordinates in the stereo scene mesh msh_S1, and generate a side-by-side image Img_SBS based on the first-view image and the second-view image. Then, the processor 130 can use a stereo display 110 to perform 3D display operations based on the side-by-side image.
[0083] In some implementations, the processor 130 can perform pinhole projection on the stereoscopic scene mesh msh_S1 based on the user-input binocular distance (i.e., the distance between the two virtual cameras) to generate a first-view image and a second-view image. In some implementations, the processor 130 can dynamically determine the binocular distance (i.e., the distance between the two virtual cameras) based on the scene content, and perform pinhole projection on the stereoscopic scene mesh msh_S1 according to the aforementioned binocular distance, ultimately generating a first-view image and a second-view image. For example, when the scene content includes close-up objects, the processor 130 can reduce the distance between the two virtual cameras to reduce viewing discomfort. When the scene content is distant scenery, the processor 130 can increase the distance between the two virtual cameras to enhance the stereoscopic effect.
[0084] In some embodiments, the processor 130 can control the stereoscopic display 110 to operate in a stereoscopic display mode to display a side-by-side image Img_SBS including a first-view image and a second-view image. Specifically, when the stereoscopic display 110 is a naked-eye stereoscopic display, the processor 130 can perform image weaving processing on the side-by-side image Img_SBS to obtain a woven image. This image weaving processing arranges the left-eye and right-eye image pixels of the side-by-side image Img_SBS alternately within the woven frame. Then, when the stereoscopic display 110 operates in stereoscopic display mode, the display panel 111 of the stereoscopic display 110 will display the woven image, and the refractive function of the lens layer 112 of the stereoscopic display 110 will be enabled, allowing the viewer to experience a stereoscopic visual effect.
[0085] Please refer to Figure 10This is a flowchart illustrating the output of a stereoscopic scene content according to an embodiment of the present disclosure. In step S1010, the processor 130 generates a two-dimensional rendered image from a single perspective based on the stereoscopic scene mesh msh_S1. In step S1020, the processor 130 displays the two-dimensional rendered image using a stereoscopic display 110. In some embodiments, the processor 130 can control the stereoscopic display 110 to operate in a two-dimensional display mode to display the two-dimensional rendered image. Furthermore, the processor 130 can determine a specific viewing angle based on user input and perform pinhole projection on the stereoscopic scene mesh msh_S1 according to this specific viewing angle to generate a two-dimensional rendered image. That is, the user can see images from different perspectives by controlling a specific viewing angle.
[0086] In summary, in this disclosed embodiment, a stereo scene mesh can be generated based on the depth estimation result of monocular depth estimation and the stereo depth information of stereo image pairs. Therefore, monocular depth estimation can compensate for the occlusion and texture duplication problems of dual-view depth estimation, and the stereo depth information of stereo image pairs can optimize the estimation result of monocular depth estimation. Based on this, the stereo scene mesh generated based on monocular depth estimation and the stereo depth information of stereo image pairs not only allows users to perceive a depth of field consistent with the scene type, but also enables high-precision, flexible, and efficient 3D scene generation and visualization. This disclosed embodiment proposes a stereo scene mesh generation method that combines monocular depth estimation and stereo depth information, overcoming the inherent limitations of dual-view depth estimation through complementary optimization techniques, while simultaneously improving the accuracy and stability of monocular depth estimation.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for visualizing a three-dimensional scene, characterized in that, include: Obtain the left-eye and right-eye images from a stereo image pair; Monocular depth estimation is performed on the left eye image and the right eye image respectively to obtain the left eye depth map and the right eye depth map; Stereo depth information is generated based on the parallax information between the left-eye image and the right-eye image; A left-side stereo grid of the left-eye image is generated based on the left-eye depth map, and a right-side stereo grid of the right-eye image is generated based on the right-eye depth map. A 3D scene mesh is generated based on the 3D depth information, the left 3D mesh, and the right 3D mesh. as well as Output the 3D scene content based on the 3D scene grid.
2. The method for visualizing a three-dimensional scene according to claim 1, characterized in that, The steps of performing monocular depth estimation on the left-eye image and the right-eye image respectively to obtain the left-eye depth map and the right-eye depth map include: The left-eye image is input into a monocular depth estimation model to obtain the left-eye depth map; and The right eye image is input into the monocular depth estimation model to obtain the right eye depth map.
3. The method for visualizing a three-dimensional scene according to claim 1, characterized in that, The steps of generating the left-side stereo mesh of the left-eye image based on the left-eye depth map and generating the right-side stereo mesh of the right-eye image based on the right-eye depth map include: Based on the camera's intra-camera parameters and the left-eye depth map, multiple pixels of the left-eye image are mapped to multiple first stereo coordinates in a stereo coordinate system to construct the left-side stereo mesh of the left-eye image; and Based on the camera's intrinsic parameters and the right-eye depth map, multiple pixels of the right-eye image are mapped to multiple second stereo coordinates in a stereo coordinate system to construct the right-side stereo grid of the right-eye image.
4. The method for visualizing a three-dimensional scene according to claim 3, characterized in that, The mapping between the left-side 3D mesh and the right-side 3D mesh is performed based on a preset depth range.
5. The method for visualizing a three-dimensional scene according to claim 1, characterized in that, The steps for generating the stereo scene mesh based on the stereo depth information, the left stereo mesh, and the right stereo mesh include: Multiple left-side world coordinate points are obtained by mapping the left-side 3D mesh to the world coordinate system; Multiple right-side world coordinate points are obtained by mapping the right-side 3D mesh to the world coordinate system; and Based on the matching relationship between the multiple left-side world coordinate points and the multiple right-side world coordinate points, the multiple left-side world coordinate points and the multiple right-side world coordinate points are merged to generate the three-dimensional scene mesh.
6. The method for visualizing a three-dimensional scene according to claim 5, characterized in that, The step of merging the multiple left-side world coordinate points and the multiple right-side world coordinate points to generate the 3D scene mesh, based on the matching relationship between the multiple left-side world coordinate points and the multiple right-side world coordinate points, includes: The depth deviation between the plurality of left-side world coordinate points and the plurality of right-side world coordinate points is corrected using the stereo depth information.
7. The method for visualizing a three-dimensional scene according to claim 1, characterized in that, The steps for generating the stereo scene mesh based on the stereo depth information, the left stereo mesh, and the right stereo mesh include: Obtain the normal vector of the mesh surface corresponding to each vertex in the 3D scene mesh; and The pixel texture information of each vertex is determined based on the normal vector corresponding to each vertex, the preset left eye view, and the preset right eye view.
8. The method for visualizing a three-dimensional scene according to claim 7, characterized in that, The steps for determining the pixel texture information of each vertex based on the normal vector corresponding to each vertex, the preset left-eye view, and the preset right-eye view include: The left texture weight is determined based on the angle between the normal vector corresponding to the first vertex and the preset left eye view. The right-side texture weight is determined based on the angle between the normal vector corresponding to the first vertex and the preset right-eye viewpoint; and Based on the left texture weight and the right texture weight, the texture information of the left eye image and the texture information of the right eye image are weighted and calculated to obtain the pixel texture information of the first vertex.
9. The method for visualizing a three-dimensional scene according to claim 1, characterized in that, The steps for outputting 3D scene content based on the 3D scene mesh include: A single-view 2D rendered image is generated based on the aforementioned 3D scene mesh; and The two-dimensional rendered image is displayed on a monitor.
10. The method for visualizing a three-dimensional scene according to claim 1, characterized in that, The steps for outputting 3D scene content based on the 3D scene mesh include: A first-view image and a second-view image are generated side-by-side based on the stereoscopic scene mesh; and A stereoscopic display is used to perform 3D display operations based on the side-by-side images.
11. A stereoscopic image display system, characterized in that, include: 3D display; as well as At least one processor is coupled to the stereoscopic display and configured to: Obtain the left-eye and right-eye images from a stereo image pair; Monocular depth estimation is performed on the left eye image and the right eye image respectively to obtain the left eye depth map and the right eye depth map; Stereo depth information is generated based on the parallax information between the left-eye image and the right-eye image; A left-side stereo grid of the left-eye image is generated based on the left-eye depth map, and a right-side stereo grid of the right-eye image is generated based on the right-eye depth map. A 3D scene mesh is generated based on the 3D depth information, the left 3D mesh, and the right 3D mesh. as well as Output the 3D scene content based on the 3D scene grid.
12. The stereoscopic image display system according to claim 11, characterized in that, The processor is configured to: The left-eye image is input into a monocular depth estimation model to obtain the left-eye depth map; and The right eye image is input into the monocular depth estimation model to obtain the right eye depth map.
13. The stereoscopic image display system according to claim 11, characterized in that, The processor is configured to: Based on the camera's internal parameters and the left-eye depth map, multiple pixels of the left-eye image are mapped to multiple first stereo coordinates in a stereo coordinate system to construct the left-side stereo grid of the left-eye image. as well as Based on the camera's intrinsic parameters and the right-eye depth map, multiple pixels of the right-eye image are mapped to multiple second stereo coordinates in a stereo coordinate system to construct the right-side stereo grid of the right-eye image.
14. The stereoscopic image display system according to claim 13, characterized in that, The mapping between the left-side 3D mesh and the right-side 3D mesh is performed based on a preset depth range.
15. The stereoscopic image display system according to claim 11, characterized in that, The processor is configured to: Multiple left-side world coordinate points are obtained by mapping the left-side 3D mesh to the world coordinate system; Multiple right-side world coordinate points are obtained by mapping the right-side 3D mesh to the world coordinate system; as well as Based on the matching relationship between the multiple left-side world coordinate points and the multiple right-side world coordinate points, the multiple left-side world coordinate points and the multiple right-side world coordinate points are merged to generate the three-dimensional scene mesh.
16. The stereoscopic image display system according to claim 15, characterized in that, The processor is configured to: The depth deviation between the plurality of left-side world coordinate points and the plurality of right-side world coordinate points is corrected using the stereo depth information.
17. The stereoscopic image display system according to claim 11, characterized in that, The processor is configured to: Obtain the normal vector of the mesh surface corresponding to each vertex in the 3D scene mesh; as well as The pixel texture information of each vertex is determined based on the normal vector corresponding to each vertex, the preset left eye view, and the preset right eye view.
18. The stereoscopic image display system according to claim 17, characterized in that, The processor is configured to: The left texture weight is determined based on the angle between the normal vector corresponding to the first vertex and the preset left eye view. The right-side texture weight is determined based on the angle between the normal vector corresponding to the first vertex and the preset right-eye view. as well as Based on the left texture weight and the right texture weight, the texture information of the left eye image and the texture information of the right eye image are weighted and calculated to obtain the pixel texture information of the first vertex.
19. The stereoscopic image display system according to claim 11, characterized in that, The processor is configured to: A single-view 2D rendered image is generated based on the aforementioned 3D scene mesh; and The two-dimensional rendered image is displayed on a monitor.
20. The stereoscopic image display system according to claim 11, characterized in that, The processor is configured to: A first-view image and a second-view image are generated side-by-side based on the stereoscopic scene mesh; and A stereoscopic display is used to perform 3D display operations based on the side-by-side images.