Transmission type display method and transmission type display system
By employing a permeable display method and utilizing image sensor and processor technology, the problem of adjusting the relationship between AR images and the user's position on large displays has been solved, achieving alignment between the displayed content and the actual scene and enabling an interactive experience.
Patent Information
- Application Number
- CN202410493196.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-10-24
AI Technical Summary
When AR technology is applied to a large display at a fixed position, the AR screen cannot be adjusted according to the relative position between the user and the display, resulting in the user being unable to effectively view the scene objects of interest.
The transmissive display method is adopted, which acquires user images through a first image sensor and scene images through a second image sensor. The processor determines the viewing frustum based on the user's position information and the display size, and outputs display frames on the display through a projection matrix to display the scene behind the display.
The content of the displayed frames can change accordingly with the user's movement and is well aligned with the actual scene around the monitor, providing a good interactive experience.
Smart Images

Figure CN120833720A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to an image display technology, and more particularly to a see-through display method and a see-through display system. BACKGROUND
[0002] With the advancement of technology, augmented reality (AR) applications have become increasingly popular. This technology not only breaks through in the field of entertainment, but also is widely used in the fields of business, education, medicine, etc. With the continuous maturity and popularity of AR technology, people can integrate virtual elements into the real world through AR glasses, smart phones, various handheld electronic devices, or various wearable electronic devices, to provide users with rich interactive experiences. In general, AR technology will continue to change the way of modern life and bring more convenience and rich experiences to modern people.
[0003] Generally speaking, the camera disposed on the back of the handheld electronic device can capture a real scene, and the AR picture displayed by the handheld electronic device includes a real scene image and a virtual element superimposed on the real scene image. Traditionally, it is unnecessary to determine the AR picture based on the tracking results of the user, because the mobility of the handheld electronic device and the relative position change between the user and the handheld electronic device are not obvious. However, when trying to apply AR technology to a large display located at a fixed position, if the relative position relationship between the user and the display is not considered, the scene content in the AR picture will not meet the user's needs. For example, the user may not be able to view the scene object of interest through the AR picture of the display. SUMMARY
[0004] Therefore, the present disclosure provides a see-through display method and a see-through display system to solve the above technical problems.
[0005] An exemplary embodiment of the present disclosure provides a see-through display method suitable for a see-through display system including a first image sensor, a second image sensor, and a display. The see-through display method includes: acquiring a user image through the first image sensor toward a front side of the display and acquiring a scene image through the second image sensor toward a back side of the display; acquiring user position information associated with a stereoscopic reference coordinate system according to the user image; acquiring scene position information associated with the stereoscopic reference coordinate system according to the scene image; determining a viewing frustum according to the user position information and an actual size of the display; generating a display frame projected on a display plane of the display according to the scene position information by using a projection matrix of the viewing frustum; and outputting the display frame through the display to display a scene at the back side of the display.
[0006] Another exemplary embodiment of the present application provides a see-through display system, which includes a first image sensor, a second image sensor, a display, and at least one processor. The processor is coupled to the first image sensor, the second image sensor, and the display, and is configured to: acquire a user image by the first image sensor toward a front side of the display, and acquire a scene image by the second image sensor toward a back side of the display; acquire user position information associated with a stereoscopic reference coordinate system according to the user image; acquire scene position information associated with the stereoscopic reference coordinate system according to the scene image; determine a viewing frustum according to the user position information and an actual size of the display; generate a display frame projected on a display plane of the display according to the scene position information by using a projection matrix of the viewing frustum; and output the display frame by the display to display a scene at the back side of the display.
[0007] Based on the above, in the embodiment of the present application, the user position information and the scene position information in the same stereoscopic reference coordinate system can be acquired according to the user image and the scene image. The viewing frustum for determining the display content of the display frame can be determined based on the user position information and the actual size of the display. Thus, the display scene content of the display frame output by the display can be changed according to the movement of the user and well aligned with the actual scene around the display. BRIEF DESCRIPTION OF DRAWINGS
[0008] Figures 1A-1C is a schematic diagram of a see-through display system according to an embodiment of the present application;
[0009] Figure 2 is a flowchart of a see-through display method according to an embodiment of the present application;
[0010] Figure 3 is a flowchart of acquiring scene position information according to an embodiment of the present application;
[0011] Figure 4 is a schematic diagram of a stereoscopic reference coordinate system according to an embodiment of the present application;
[0012] Figure 5 is a schematic diagram of determining a viewing frustum according to user position information according to an embodiment of the present application;
[0013] Figure 6A and Figure 6B is a schematic diagram of displaying a scene according to an embodiment of the present application;
[0014] Figure 7 is a schematic diagram of a see-through display system according to an embodiment of the present application;
[0015] Figure 8 is a flowchart of a see-through display method according to an embodiment of the present application.
[0016] Reference Signs List
[0017] 10, 70: see-through display system
[0018] 110: first image sensor
[0019] 120: second image sensor
[0020] 130: display
[0021] 140: storage device
[0022] 150: processor
[0023] 160: depth sensor
[0024] F1: display frame
[0025] Obj1: scene object
[0026] U1: user
[0027] 111, 112, 113: FOV
[0028] S1: display plane
[0029] DP1-DP4: vertex
[0030] VP1: user coordinate
[0031] 51: view cone
[0032] S210-S260, S310-S330, S810-S870: step DETAILED DESCRIPTION
[0033] Reference will now be made in detail to the exemplary embodiments of the present application, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numbers will be used in the drawings and the description to refer to the same or like parts.
[0034] Figures 1A-1C is a schematic diagram of a see-through display system according to an embodiment of the present application.
[0035] Reference will now be made in detail to the exemplary embodiments of the present application, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numbers will be used in the drawings and the description to refer to the same or like parts. Figure 1A In some embodiments, the see-through display system 10 can be implemented in, for example, a notebook computer, a tablet computer, a personal computer, a server, a game console, a portable electronic device, a desktop computer, or other electronic devices with image processing capability and data computing capability. The see-through display system 10 includes a first image sensor 110, a second image sensor 120, a display 130, a storage device 140, and at least one processor 150.
[0036] The processor 150 is responsible for all or part of the operation of the see-through display system 10. For example, the processor 150 can include a central processing unit (CPU), a graphic processing unit (GPU), or other programmable general purpose or special purpose microprocessors, digital signal processors (DSPs), programmable controllers, application specific integrated circuits (ASICs), programmable logic devices (PLDs), or other similar devices or a combination thereof. The number of processors 150 can be one or more, which is not limited in the present application.
[0037] The storage device 140 is connected to the processor 150 and is used to temporarily or permanently store data, such as images, instructions, program codes, software modules, and the like. Specifically, the storage device 140 can include volatile storage circuitry. The volatile storage circuitry is used to store data in a volatile manner. For example, the volatile storage circuitry can include a random access memory (RAM) or similar volatile storage medium. Alternatively, the storage device 140 can include non-volatile storage circuitry. The non-volatile storage circuitry is used to store data in a non-volatile manner. For example, the non-volatile storage circuitry can include a read only memory (ROM), a solid state drive (SSD), and / or a traditional hard disk drive (HDD), or similar non-volatile storage medium. The number of storage devices 140 can be one or more, which is not limited in the present application.
[0038] The display 130 is, for example, a Liquid Crystal Display (LCD), a Light-Emitting Diode (LED) display, an Organic Light-Emitting Diode (OLED), or other types of displays, which are not limited in the present application. In some embodiments, the display 130 can be a stereoscopic display, which provides different images to the left eye or the right eye of a user respectively to present a stereoscopic visual effect, but the present application is not limited thereto. For example, the display 130 can be a naked-eye 3D display or a glasses-type 3D display.
[0039] The first image sensor 110 is configured to acquire an image and includes a camera lens having a lens and a light-sensing component. The first image sensor 110 can be implemented as a camera module including the lens, the light-sensing component and other components, for example. The light-sensing component can be a charge coupled device (CCD), a complementary metal-oxide semiconductor (CMOS) component or other components, which are not limited in the present application. From another point of view, the first image sensor 110 can be an RGB image sensor.
[0040] The second image sensor 120 is configured to acquire an image and includes a camera lens having a lens and a light-sensing component. The second image sensor 120 can be implemented as a camera module including the lens, the light-sensing component and other components, for example. The light-sensing component can be a charge coupled device (CCD), a complementary metal-oxide semiconductor (CMOS) component or other components, which are not limited in the present application. From another point of view, the second image sensor 120 can be an RGB image sensor.
[0041] Please refer to Figure 1B In the embodiment of the present application, the user U1 is located in front of the display 130, and the first image sensor 110 is configured to acquire a user image toward the front of the display 130. According to eye tracking technology or face tracking technology well known to those skilled in the art, the user image can be used to track eye position information or face position information of the user U1, etc. The second image sensor 120 is configured to acquire a scene image toward the back of the display 130. Accordingly, the processor 150 can acquire user position information of the user U1 based on the user image acquired by the first image sensor 110, and the processor 150 can acquire scene position information, such as spatial position information of the scene object Obj1, based on the scene image acquired by the second image sensor 120.
[0042] In the embodiment of the present application, the processor 150 can determine a display frame F1 according to the user position information of the user U1. The display frame F1 is configured to display a scene behind the display 130, and the display scene of the display frame F1 can be substantially aligned with the actual surrounding scene. As shown in Figure 1B Although the user U1 cannot directly see the scene object Obj1 hidden by the display 130, the display frame F1 can include an image of the scene object Obj1 to realize the function of see-through display. It is worth noting that when the user U1 moves, the hidden scene content by the display 130 also changes accordingly. The display scene of the display frame F1 can also change in response to the movement of the user U1, so that the display scene of the display frame F1 can remain substantially aligned with the actual surrounding scene.
[0043] Please refer to Figure 1CIn an embodiment of the present disclosure, the see-through display system 10 can be a notebook computer. The first image sensor 110 can be a front camera module disposed above the display plane of the display 130, and the second image sensor 120 can be a rear camera module disposed on the upper cover of the notebook computer. When the user U1 operates the notebook computer, the display 130 of the notebook computer can display the context content that is blocked by the notebook computer according to the current position of the user U1.
[0044] In some embodiments, the field of view (FOV) 111 of the first image sensor 110 is smaller than the FOV 112 of the second image sensor 120 to ensure that the second image sensor 120 acquires sufficient scene content. In some embodiments, the lens of the second image sensor 120 can be implemented by a fisheye lens or a wide-angle lens with a large FOV.
[0045] In addition, in some embodiments, the range of the scene displayed by the display 130 can be determined according to the user position information of the user U1 and the actual size of the display 130. Further, under the condition that the user U1 is regarded as a virtual camera, the processor 150 can determine the FOV 113 and the view cone of the virtual camera according to the user position information of the user U1 and the actual size of the display 130.
[0046] Figure 2 is a flowchart of a see-through display method according to an embodiment of the present disclosure, and Figure 2 The method flow of Figure 1A may be implemented by the see-through display system 10 of In this way, the user can view the scene content behind the display 130 through the display 130 of the see-through display system 10.
[0047] In step S210, the processor 150 acquires a user image toward the front side of the display 130 through the first image sensor 110 and acquires a scene image toward the rear side of the display 130 through the second image sensor 120. Specifically, the first image sensor 110 is used to take a picture toward the user viewing the display 130, and the second image sensor 120 is used to take a picture toward the actual scene behind the display 130.
[0048] At step S220, the processor 150 obtains user position information associated with the stereoscopic reference coordinate system based on the user image. It is noted that the stereoscopic reference coordinate system is defined based on the display plane of the display 130. In some embodiments, the user position information can include distance information between the user and the display 130. The processor 150 can estimate the distance information between the user and the display 130 based on the face size, the interpupillary distance or other facial features in the user image. Alternatively, in other embodiments, the user position information can include three-dimensional user coordinates of the user in a three-dimensional coordinate system. The processor 150 can convert the image coordinates of the user in the user image to world coordinates in a world coordinate system (e.g. the stereoscopic reference coordinate system established based on the display 130) based on the intrinsic parameters and extrinsic parameters of the first image sensor 110.
[0049] In some embodiments, the first and second coordinate axes of the stereoscopic reference coordinate system are parallel to the display plane of the display 130, and the origin of the stereoscopic reference coordinate system is a reference point on the display plane. For example, the origin of the stereoscopic reference coordinate system can be the center point on the display plane. The X and Y axes of the stereoscopic reference coordinate system are on the display plane. That is, the display plane is the plane of Z=0 in the stereoscopic reference coordinate system.
[0050] In some embodiments, the first image sensor 110 can also be coupled with at least one depth sensor (not shown) or distance sensor (not shown) to perform image recognition positioning on the user to obtain three-dimensional user coordinates of the user in the stereoscopic reference coordinate system.
[0051] At step S230, the processor 150 obtains scene position information associated with the stereoscopic reference coordinate system based on the scene image. In some embodiments, the scene position information can include three-dimensional scene coordinates in a three-dimensional coordinate system. The processor 150 can convert the image coordinates in the scene image to world coordinates in a world coordinate system (e.g. the stereoscopic reference coordinate system established based on the display 130) based on the intrinsic parameters and extrinsic parameters of the second image sensor 120. In some embodiments, the image coordinates in the scene image for coordinate conversion can be sampled from the grid nodes of the three-dimensional grid.
[0052] In some embodiments, Figure 3 is a flowchart of obtaining scene position information according to an embodiment of the present application. Please refer to Figure 3 At step S310, the processor 150 establishes a stereoscopic reference coordinate system based on the display plane of the display 130.
[0053] For example, Figure 4 is a schematic diagram of a stereoscopic reference coordinate system according to an embodiment of the present application. Please refer toFigure 4 The processor 150 can set the origin (0, 0, 0) of the stereoscopic reference coordinate system as a center point on the display plane S1. The X-axis of the stereoscopic reference coordinate system is the display horizontal axis of the display 130, and the Y-axis of the stereoscopic reference coordinate system is the display vertical axis of the display 130. The Z-axis of the stereoscopic reference coordinate system passes through the origin (0, 0, 0) and is perpendicular to the display plane S1.
[0054] At step S320, the processor 150 determines the extrinsic parameters of the second image sensor 120 according to the spatial positional relationship between the second image sensor 120 and the display 130. The extrinsic parameters of the second image sensor 120 describe the position and orientation of the second image sensor 120, and the conversion relationship between the second image sensor 120 and the world coordinate system. These extrinsic parameters are usually used to define the spatial pose of the second image sensor 120, so as to map the points in the camera coordinate system to the world coordinate system, or map the points in the world coordinate system to the camera coordinate system.
[0055] In some embodiments, the processor 150 can define a stereoscopic reference coordinate system based on the display plane of the display 130, and take this stereoscopic reference coordinate system as the world coordinate system. In this case, according to the spatial positional relationship between the second image sensor 120 and the display 130, the processor 150 can obtain the coordinate position of the second image sensor 120 in the stereoscopic reference coordinate system. In addition, other extrinsic parameters of the second image sensor 120, such as the shooting direction, etc., can be obtained through a camera calibration program.
[0056] At step S330, the processor 150 performs coordinate conversion on the scene pixel coordinates in the scene image according to the intrinsic parameters and the extrinsic parameters of the second image sensor 120, so as to obtain the scene position information associated with the stereoscopic reference coordinate system. Specifically, the processor 150 can perform coordinate conversion based on the following formula (1) to convert the scene pixel coordinates in the scene image to the stereoscopic scene coordinates in the stereoscopic reference coordinate system. In some embodiments, the processor 150 can convert the image coordinates of the grid nodes of the grid of the scene image to the stereoscopic scene coordinates.
[0057]
[0058] Where (u, v) represents the image coordinates, and (X, Y, Z) represents the world coordinates.
[0059] The intrinsic parameter matrix of the second image sensor 120 is
[0060]
[0061] is an external parameter matrix of the second image sensor 120. The external parameter matrix includes a rotation matrix R and a translation vector T. The external parameter matrix of the second image sensor 120 can be used to represent the position and the shooting direction of the second image sensor 120 in the world coordinate system (i.e., the stereoscopic reference coordinate system). In this way, the processor 150 can convert the image coordinates in the scene image into the stereoscopic scene coordinates in the stereoscopic reference coordinate system through coordinate conversion.
[0062] In some embodiments, the second image sensor 120 can include a fisheye lens or a wide-angle lens, and the fisheye lens or the wide-angle lens is used to acquire the scene image. Before acquiring the scene position information associated with the stereoscopic reference coordinate system through coordinate conversion, the processor 150 can perform a distortion correction process on the scene image. In other words, the processor 150 can correct the image distortion for the fisheye image or the wide-angle image. Specifically, when the second image sensor 120 acquires the scene image through the wide-angle lens, the processor 150 can perform the distortion correction process through equation (2). When the second image sensor 120 acquires the scene image through the fisheye lens, the processor 150 can perform the distortion correction process through equation (3).
[0063]
[0064]
[0065] wherein, θ = tan -1 r, k n are radial distortion coefficients, and p n are tangential distortion coefficients.
[0066] Returning to Figure 2 , at step S240, the processor 150 determines a viewing frustum according to the user position information and the actual size of the display. The viewing frustum is a geometric body used to represent the visible region of the camera, and can also be referred to as a projection frustum. The viewing frustum is composed of six planes, which are the near plane, the far plane, the left plane, the right plane, the top plane, and the bottom plane. These planes define the region that the camera can see. That is, the viewing frustum can be used to determine the local scene image acquired from the scene image.
[0067] In an embodiment of the present invention, processor 150 may set the user's coordinates in the stereoscopic reference coordinate system as the coordinate position of the virtual camera and determine the viewing frustum based on these user coordinates. In some embodiments, the viewing frustum changes in response to changes in the user's position information. That is, when the user moves, the viewing frustum also changes accordingly.
[0068] In some embodiments, the user position information may include a user coordinate associated with a stereo reference coordinate system, and the viewing frustum is obtained by connecting the user coordinate with multiple vertices of the display plane of the display 130. In other words, the left plane, right plane, top plane, and bottom plane of the viewing frustum are determined according to the display range of the display 130. For example, Figure 5 This is a schematic diagram of determining the viewing cone based on user location information according to an embodiment of the present invention. Figure 5 , display plane S1 of display 130 includes vertices DP1, DP2, DP3, and DP4. After obtaining user coordinates VP1 in a stereoscopic reference coordinate system based on the user's shadow, processor 150 may connect user coordinates VP1 with the plurality of vertices DP1, DP2, DP3, and DP4 of display plane S1 to obtain a viewing frustum 51. Processor 150 may set the far plane and near plane of viewing frustum 51 based on a predetermined distance. It is understood that the aspect ratio of the viewport is the aspect ratio of display 130.
[0069] After determining the viewing frustum based on the actual size of the display 130 and the user coordinates, the processor 150 can derive the parameters of the projection matrix. Specifically, each plane of the viewing frustum defines parameters in the projection matrix, such as the viewing angle, the viewport aspect ratio, and the distance between the near and far planes. These parameters determine the values of the projection matrix. In some embodiments, this projection matrix can be an off-center perspective matrix. For example, the projection matrix P obtained by the processor 150 can be represented by Formula (5).
[0070]
[0071] Wherein, near represents the distance between the near plane and the user coordinates, far represents the distance between the far plane and the user coordinates, right represents the X coordinate of the right display border of the display 130, and left represents the X coordinate of the left display border of the display 130. top represents the Y coordinate of the top display border of the display 130, and bottom represents the Y coordinate of the bottom display border of the display 130.
[0072] At step S250, the processor 150 generates a display frame projected on the display plane of the display 130 according to the scene position information using the projection matrix of the frustum. Specifically, the processor 150 can multiply the four-dimensional homogeneous coordinates (x, y, z, 1) of the plurality of three-dimensional scene coordinates in the stereoscopic reference coordinate system by the projection matrix P to map the scene coordinates to the corresponding screen coordinates on the viewport (i.e., the display plane). The scene coordinates can be the three-dimensional coordinates of the plurality of mesh nodes in the stereoscopic reference coordinate system. In other words, the partial scene images in the scene image can be projected to the display plane via the projection matrix to generate the display frame.
[0073] At step S260, the processor 150 outputs the display frame via the display 130 to display the scene behind the display 130. Specifically, since the projection range of the scene image projected on the display plane of the display 130 is determined according to the user position information and the actual size of the display 130, the display frame output by the display 130 can not only present the scene behind the display 130, but also align the scene image in the display frame with the actual scene around the display 130. In addition, in response to the user movement, the scene content of the display frame of the display 130 will also change accordingly.
[0074] For example, Figure 6A With Figure 6B is a schematic diagram of displaying a scene according to an embodiment of the present application. Please refer to Figure 6A When the user U1 is at the first position, the display frame of the display 130 can include the scene content behind the display 130. For example, the container that is blocked by the display 130 can be displayed in the display 130. Please refer to Figure 6B When the user moves from the first position to the second position, the frustum determined according to the user position will change accordingly. Thus, the scene content obtained by the frustum will also change, causing the scene content displayed by the display 130 to be adjusted accordingly. In some embodiments, the processor 150 can also use the display frame as the background of an AR picture to provide an AR function or an AR application.
[0075] Figure 7 is a schematic diagram of a see-through display system according to an embodiment of the present application. Please refer to Figure 7The transparent display system 70 may include a first image sensor 110, a second image sensor 120, a display 130, a storage device 140, at least one processor 150, and a depth sensor 160. Different from the embodiment of Figure 1, the transparent display system 70 may also include a depth sensor 160 for sensing scene depth information. The depth sensor 160 can be implemented using active depth sensing technology and passive depth sensing technology. Active depth sensing technology can calculate depth information by actively emitting light sources, infrared rays, ultrasound waves, lasers, etc. as signals in combination with time-of-flight ranging technology. Passive depth sensing technology can use two image sensors to obtain two images in front of them at different perspectives to calculate depth information using the parallax of the two images.
[0076] Figure 8 is a flow chart of a transparent display method according to an embodiment of the present invention, and Figure 8 The method flow can be obtained by Figure 7 Here, the user can watch the scene content on the rear side of the display 130 through the display 130 of the transparent display system 70.
[0077] In this embodiment, the scene position information of the scene image includes 3D mesh scene information associated with a 3D reference coordinate system. The Z coordinate values of the mesh nodes in the 3D mesh scene information in the 3D reference coordinate system may be generated according to the scene depth.
[0078] In step S810, the processor 150 captures an image of the user from the front of the display 130 via the first image sensor 110 and captures an image of the scene from the rear of the display 130 via the second image sensor 120. In step S820, the processor 150 obtains user position information associated with a stereoscopic reference coordinate system based on the image of the user. These steps can be referred to in the description of the previous embodiment and are not repeated here.
[0079] In step S830, processor 150 obtains depth information corresponding to the scene image. For example, the value of each pixel (or position) in the depth map may indicate the depth value of the corresponding pixel (or position) in the scene image. Processor 150 may use the depth map as the depth information corresponding to the scene image.
[0080] In some embodiments, processor 150 may utilize depth sensor 160 to obtain depth information corresponding to the scene image. Alternatively, in other embodiments, processor 150 may perform image preprocessing on the scene image to generate an adjusted scene image that meets the input requirements of the deep learning model. Processor 150 may analyze the adjusted scene image using the deep learning model to obtain depth information.
[0081] In some embodiments, the storage 140 can store a deep learning model. The deep learning model is implemented based on a neural network structure, such as a convolutional neural network (CNN) or a neural network-like structure. The deep learning model is used to estimate (i.e., predict) the depth of each pixel (or location) in the scene image. In addition, the processor 150 can perform image pre-processing operations on the scene image to generate an adjusted scene image that satisfies the input requirements of the deep learning model. For example, in the image pre-processing operations, the processor 150 can adjust the size of the scene image and / or convert the format of the scene image to generate the adjusted scene image. The processor 150 analyzes the adjusted scene image by the deep learning model to obtain the depth information corresponding to the first image. For example, the processor 150 can input the adjusted scene image to the deep learning model and then receive the output depth map of the deep learning model with respect to the adjusted scene image.
[0082] In some embodiments, the processor 150 can determine whether the depth sensor 160 for sensing the depth information of the scene is available or determine whether the deep learning model for estimating the depth information of the scene is available. When the depth sensor 160 is available or the deep learning model is available, the processor 150 can obtain the depth information of the scene image.
[0083] When at least one of the depth sensor 160 and the deep learning model is available, in step S840, the processor 150 generates three-dimensional mesh scene information associated with the stereoscopic reference coordinate system according to the depth information and the scene image. In an embodiment, the height of the mesh node of the three-dimensional mesh in the Z-axis direction of the stereoscopic reference coordinate system can be considered as the depth of the three-dimensional mesh node. Otherwise, when neither the depth sensor 160 nor the deep learning model is available, the processor 150 can generate mesh information of a plane with the same height in the Z-axis direction.
[0084] At step S850a, the processor 150 determines the first frustum according to the right eye position information in the user position information and the actual size of the display 130. At step S850b, the processor 150 determines the second frustum according to the left eye position information in the user position information and the actual size of the display. The detailed operation of the processor 150 in determining the frustum can refer to the description of the foregoing embodiments. It should be particularly pointed out that in some embodiments, the processor 150 can locate the right eye position information of the right eye and the left eye position information of the left eye of the user according to the user image respectively. The right eye position information and the left eye position information can be left eye coordinates and right eye coordinates in the stereoscopic reference coordinate system respectively. Thus, the processor 150 can determine the first frustum and the second frustum respectively with the left eye coordinates and the right eye coordinates and the actual size of the display 130. It can be known that since the right eye coordinates are different from the left eye coordinates, the scene content obtained by the first frustum and the second frustum will also be different.
[0085] At step S860a, the processor 150 generates a right eye display frame projected on the display plane of the display 130 according to the scene position information by using the projection matrix of the first frustum. In detail, the processor 150 can project the three-dimensional grid nodes in the scene image obtained by the first frustum onto the display plane of the display 130 to render the right eye display frame.
[0086] At step S860b, the processor 150 generates a left eye display frame projected on the display plane of the display 130 according to the scene position information by using the projection matrix of the second frustum. In detail, the processor 150 can project the three-dimensional grid nodes in the scene image obtained by the second frustum onto the display plane of the display 130 to render the left eye display frame.
[0087] It should be pointed out that the processor 150 obtains the scene depth of the three-dimensional grid nodes by using the depth estimation technology or the depth sensing technology, and then projects the three-dimensional grid nodes with depth information onto the display plane of the display 130. Thus, the scene content presented by the display 130 can be more accurate, and the perceived position of the scene object will not be improperly offset due to the lack of depth information.
[0088] At step S870, the processor 150 outputs the left eye display frame and the right eye display frame through the display 130 to display the scene behind the display 130. In some embodiments, when the display 130 is a naked-eye 3D display, the processor 150 can perform image weaving processing on the left eye display frame and the right eye display frame to synchronously interleave the left eye display frame and the right eye display frame. When the display 130 is a glasses-type 3D display, the processor 150 can control the display 130 to alternately display the left eye display frame and the right eye display frame. In this way, the user can experience the stereoscopic visual effect.
[0089] Based on the above, in the embodiments of the present application, the user position information and the scene position information in the same stereoscopic reference coordinate system can be obtained according to the user image and the scene image. The view cone for determining the display content of the display frame can be determined based on the user position information and the actual size of the display. Thus, the display scene content of the display frame output by the display can change in response to the user movement and achieve good alignment with the actual scene around the display. In addition, when the scene depth information is used to project the content of the scene image to the display plane, the scene content presented by the display can be closer to the actual scene.
[0090] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A see-through display method suitable for a see-through display system comprising a first image sensor, a second image sensor, and a display, the method comprising: The method comprises: acquiring a user image through the first image sensor towards a front side of the display and acquiring a scene image through the second image sensor towards a back side of the display; acquiring user position information associated with a stereoscopic reference coordinate system according to the user image; acquiring scene position information associated with the stereoscopic reference coordinate system according to the scene image; determining a frustum according to the user position information and an actual size of the display; generating a display frame projected on a display plane of the display according to the scene position information by using a projection matrix of the frustum; and outputting the display frame through the display to display a scene at the back side of the display. The second image sensor comprises a fisheye lens or a wide-angle lens, and the fisheye lens or the wide-angle lens is used to acquire the scene image.
2. The see-through display method according to claim 1, wherein, Before the step of acquiring the scene position information associated with the stereoscopic reference coordinate system according to the scene image, the method further comprises: performing a distortion correction process on the scene image. A first coordinate axis and a second coordinate axis of the stereoscopic reference coordinate system are parallel to a display plane of the display, and an origin of the stereoscopic reference coordinate system is a reference point on the display plane.
3. The see-through display method according to claim 1, wherein, The frustum changes in response to a change in the user position information.
4. The see-through display method according to claim 1, wherein, The step of acquiring the scene position information associated with the stereoscopic reference coordinate system according to the scene image comprises:
5. The see-through display method according to claim 1, wherein, establishing the stereoscopic reference coordinate system based on the display plane of the display; determining an external parameter of the second image sensor according to a spatial position relationship between the second image sensor and the display; and performing coordinate conversion on scene pixel coordinates in the scene image according to an internal parameter of the second image sensor and the external parameter to acquire the scene position information associated with the stereoscopic reference coordinate system. The user position information comprises a user coordinate associated with the stereoscopic reference coordinate system, and the frustum is acquired by connecting the user coordinate and a plurality of vertices of the display plane.
6. The see-through display method according to claim 1, wherein, The scene position information comprises three-dimensional grid scene information associated with the stereoscopic reference coordinate system, and the step of acquiring the scene position information associated with the stereoscopic reference coordinate system according to the scene image comprises:
7. The see-through display method according to claim 1, wherein, acquiring depth information corresponding to the scene image; and generating the three-dimensional grid scene information associated with the stereoscopic reference coordinate system according to the depth information and the scene image. The step of acquiring the depth information corresponding to the scene image comprises:
8. The see-through display method according to claim 7, wherein, performing an image preprocessing operation on the scene image to generate an adjusted scene image satisfying an input requirement of a deep learning model; and obtaining the depth information by analyzing the adjusted scene image through the deep learning model. The step of acquiring the depth information corresponding to the scene image comprises:
9. The see-through display method according to claim 7, wherein, acquiring the depth information corresponding to the scene image by using a depth sensor. 10. The see-through display method of claim 1, wherein, The display includes a stereoscopic display providing a right-eye display frame and a left-eye display frame, and the step of determining the frustum according to the user position information and the actual size of the display includes: determining a first frustum according to right-eye position information in the user position information and the actual size of the display; and determining a second frustum according to left-eye position information in the user position information and the actual size of the display, wherein the step of generating the display frame projected on the display plane of the display according to the scene position information using the projection matrix of the frustum includes: generating the right-eye display frame projected on the display plane of the display according to the scene position information using the projection matrix of the first frustum; and generating the left-eye display frame projected on the display plane of the display according to the scene position information using the projection matrix of the second frustum.
11. A transmissive display system, characterized by comprise: a first image sensor; a second image sensor; a display; and at least one processor coupled to the first image sensor, the second image sensor and the display, wherein the at least one processor is configured to: acquire a user image through the first image sensor towards a front side of the display and acquire a scene image through the second image sensor towards a back side of the display; acquire user position information associated with a stereoscopic reference coordinate system according to the user image; acquire scene position information associated with the stereoscopic reference coordinate system according to the scene image; determine a frustum according to the user position information and the actual size of the display; generate a display frame projected on a display plane of the display according to the scene position information using a projection matrix of the frustum; and output the display frame through the display to display a scene at the back side of the display. The second image sensor includes a fisheye lens or a wide-angle lens configured to acquire the scene image, and the at least one processor is configured to:
12. The see-through display system of claim 11, wherein, perform a distortion correction process on the scene image. A first coordinate axis and a second coordinate axis of the stereoscopic reference coordinate system are parallel to a display plane of the display, and an origin of the stereoscopic reference coordinate system is a reference point on the display plane.
13. The see-through display system of claim 11, wherein, The frustum changes in response to a change in the user position information.
14. The see-through display system of claim 11, wherein, The at least one processor is configured to:
15. The see-through display system of claim 11, wherein, establish the stereoscopic reference coordinate system based on the display; determine an external parameter of the second image sensor according to a spatial position relationship between the second image sensor and the display; and perform coordinate conversion on scene pixel coordinates in the scene image according to an internal parameter of the second image sensor and the external parameter to acquire the scene position information associated with the stereoscopic reference coordinate system. The user position information includes user coordinates associated with the stereoscopic reference coordinate system, and the frustum is acquired by connecting the user coordinates and a plurality of vertices of the display plane. 16. The see-through display system of claim 11, wherein, 17. The see-through display system of claim 11, wherein, The scene position information comprises three-dimensional mesh scene information associated with the stereoscopic reference coordinate system, and the at least one processor is configured to: obtain depth information corresponding to the scene image; and generate the three-dimensional mesh scene information associated with the stereoscopic reference coordinate system according to the depth information and the scene image.
18. The see-through display system of claim 17, wherein, The at least one processor is configured to: perform image preprocessing operation on the scene image to generate an adjusted scene image satisfying input requirements of a deep learning model; and analyze the adjusted scene image by the deep learning model to obtain the depth information. Further comprising a depth sensor, and the at least one processor is connected to the depth sensor and configured to:
19. The see-through display system of claim 17, wherein, obtain the depth information corresponding to the scene image by the depth sensor.
20. The see-through display system of claim 11, wherein: the display comprises a stereoscopic display providing a right-eye display frame and a left-eye display frame, and the at least one processor is configured to: determine a first frustum according to right-eye position information in the user position information and an actual size of the display; and determine a second frustum according to left-eye position information in the user position information and the actual size of the display, wherein the at least one processor is configured to: generate the right-eye display frame projected on the display plane of the display according to the scene position information by using a projection matrix of the first frustum; and generate the left-eye display frame projected on the display plane of the display according to the scene position information by using a projection matrix of the second frustum.