Augmented reality near-to-eye display device and method
By combining monocular depth estimation and hologram generation algorithms with optical modulation and waveguide display, virtual-real fusion of augmented reality near-eye display devices has been achieved, solving the problems of difficult virtual-real fusion and large device size and weight in existing technologies, and improving wearing comfort.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-12
AI Technical Summary
Existing augmented reality near-eye display devices cannot achieve the fusion of auxiliary information with objects at different depths in real scenes, and the devices are large, heavy, and have poor wearing comfort.
The algorithm employs a monocular depth estimation algorithm and a hologram generation algorithm. It acquires images of the target scene through a monocular camera, determines the target scene information, and generates auxiliary information holograms. It uses an optical modulation module and a waveguide display module to achieve virtual-real fusion, avoids mechanical zoom and high-precision optical components, and uses computational holography technology for continuous depth zoom.
It achieves precise virtual-real fusion of auxiliary information and real objects at different depths, reduces the size and weight of the device, improves wearing comfort, and is suitable for industrial applications such as intelligent manufacturing and warehouse management.
Smart Images

Figure CN122018164A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of display technology, and in particular to an augmented reality near-eye display device and method. Background Technology
[0002] Augmented Reality (AR) near-eye display technology, by overlaying virtual information onto real-world scenes, has shown great potential in industrial applications such as smart manufacturing and warehouse management. Current technologies use microdisplay chips to load two-dimensional digital images, projecting the image onto the user's eyes via an optical relay system. However, such solutions often only project information to a fixed depth, lacking spatial depth and failing to achieve seamless integration of auxiliary information with objects at different depths in the real-world work environment. Furthermore, the optical calibration unit significantly increases the device's size and weight, reducing user comfort. Summary of the Invention
[0003] This invention provides an augmented reality near-eye display device and method to address the shortcomings of existing near-eye display devices, which can only project information to a fixed depth, cannot achieve virtual-real fusion with objects at different depths in real-world scenes, and suffer from large system size and poor wearing comfort. This invention enables precise virtual-real fusion of auxiliary information with real-world objects at different depths, reduces device size and weight, improves wearing comfort, and is suitable for industrial applications such as intelligent manufacturing and warehouse management.
[0004] The present invention provides an augmented reality near-eye display device, comprising: an acquisition module for acquiring a target scene image; a processing module for determining target object information based on the target scene image and generating an auxiliary information hologram corresponding to the target object based on the target object information; and a fusion module for fusing the auxiliary information hologram into the target scene for augmented reality near-eye display.
[0005] According to an augmented reality near-eye display device provided by the present invention, the processing module includes: a scene information determination module, used to segment target scenes from a target scene image based on a monocular depth estimation algorithm to determine the target scene information; the target scene information includes the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eye; and a hologram generation module, used to generate auxiliary information describing the target scene based on a hologram generation algorithm, according to the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eye, and to generate the auxiliary information hologram according to the auxiliary information and the three-dimensional coordinates.
[0006] According to an augmented reality near-eye display device provided by the present invention, the scene information determination module is used to classify the target scene image based on mixed features; the mixed features include color, edge, contour, vanishing line, and vanishing point; the image categories include distant image, mid-range image, and near-range image; the depth of the distant image is estimated using the horizontal edge gradient method, the depth of the mid-range image is estimated using the vanishing line gradient method, and the depth of the near-range image is estimated using the weighted depth superposition method to obtain the target scene information.
[0007] According to an augmented reality near-eye display device provided by the present invention, the scene information determination module is used to input the target scene image into a pre-trained depth gradient-assisted monocular depth estimation network model to obtain the target scene information output by the monocular depth estimation network model; wherein, the monocular depth estimation network model is trained by depth estimation training samples and depth gradient-assisted samples.
[0008] According to an augmented reality near-eye display device provided by the present invention, the hologram generation module is used to: perform image feature recognition based on the intensity distribution to determine the type features of the target scene; generate text or graphic auxiliary information containing type identifiers according to the type features, and generate numerical auxiliary information containing distance parameters according to the depth component in the three-dimensional coordinates; input the auxiliary information and the three-dimensional coordinates into a pre-trained model without convolution error to drive an autoencoder deep learning network to obtain the auxiliary information hologram output by the autoencoder deep learning network; wherein, the autoencoder deep learning network compensates for the encoding phase convolution error by extending the phase during the decoding stage.
[0009] According to an augmented reality near-eye display device provided by the present invention, the fusion module includes: a light source module for generating collimated polarized illumination light; an optical modulation module disposed at the output end of the light source module, the input end of the optical modulation module being electrically connected to the output end of the hologram generation module, for loading the auxiliary information hologram, and for diffraction occurring after the polarized illumination light illuminates the auxiliary information hologram, projecting the auxiliary information hologram onto a preset position of a nearby target scene; and a waveguide display module disposed at the diffracted light wave output end of the optical modulation module, for transmitting the diffracted light wave to the user's eye region to perform virtual-real fusion of the auxiliary information hologram and the target scene.
[0010] According to the present invention, an augmented reality near-eye display device is provided, wherein the light source module includes: a laser light source for emitting diverging spherical waves; a collimating lens disposed at the output end of the laser light source for converging the diverging spherical waves and outputting collimated plane waves; and a polarizer disposed at the collimated plane wave output end of the collimating lens, wherein the output end of the polarizer serves as the output end of the light source module for changing the polarization state of the collimated plane waves and outputting collimated polarized illumination light.
[0011] According to an augmented reality near-eye display device provided by the present invention, the optical modulation module is a spatial light modulator, which is used to load a phase-type hologram corresponding to auxiliary information, and the phase-type hologram is used to change the surface phase distribution of the spatial light modulator.
[0012] The present invention also provides an augmented reality near-eye display method, comprising: acquiring a target scene image; determining target object information based on the target scene image, and generating an auxiliary information hologram corresponding to the target object based on the target object information; and fusing the auxiliary information hologram into the target scene for augmented reality near-eye display.
[0013] According to an augmented reality near-eye display method provided by the present invention, the step of determining target scene information based on the target scene image and generating an auxiliary information hologram corresponding to the target scene based on the target scene information includes: segmenting the target scene from the target scene image based on a monocular depth estimation algorithm to determine the target scene information; the target scene information includes the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eye; generating auxiliary information describing the target scene based on the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eye based on a hologram generation algorithm, and generating the auxiliary information hologram based on the auxiliary information and the three-dimensional coordinates.
[0014] The present invention provides an augmented reality near-eye display device and method, which can achieve accurate virtual-real fusion of auxiliary information and real objects at different depths, reduce the size and weight of the device, improve wearing comfort, and is suitable for industrial applications such as intelligent manufacturing and warehouse management. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0016] Figure 1This is a schematic diagram of the structure of an augmented reality near-eye display device provided by the present invention.
[0017] Figure 2 This is a schematic diagram of the specific structure of an augmented reality near-eye display device provided by the present invention.
[0018] Figure 3 This is a flowchart illustrating an augmented reality near-eye display method provided by the present invention.
[0019] Figure 4 This is a schematic diagram illustrating the specific process of an augmented reality near-eye display method provided by the present invention.
[0020] Figure label: 1: Acquisition module; 2: Processing module; 3: Fusion module; 11: Camera; 21: Digital signal processing chip; 31: Laser diode; 32: Convex lens; 33: Polarizer; 34: Spatial light modulator; 35: Coupler grating; 36: Waveguide; 37: First coupler grating; 38: Second coupler grating; 4: Signal cable; 5: Battery; 6: Support frame; 71: First target scene; 72: Second target scene; 81: First holographic auxiliary information; 82: Second holographic auxiliary information. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0022] Augmented reality near-eye display devices in industrial settings are mainly single-focal-plane, making it difficult to achieve virtual-real fusion of the work scene and auxiliary information; a few devices can achieve three-dimensional display effects, but they are usually based on binocular parallax and light field display principles, which have convergence-focusing conflicts. Wearers may experience dizziness, fatigue and other discomfort after long-term use, affecting the virtual-real fusion experience.
[0023] Please refer to Figure 1 , Figure 1 This is a schematic diagram of the structure of an augmented reality near-eye display device provided by the present invention.
[0024] The present invention provides an augmented reality near-eye display device, comprising: an acquisition module 1 for acquiring a target scene image; a processing module 2 for determining target scene information based on the target scene image and generating an auxiliary information hologram corresponding to the target scene based on the target scene information; and a fusion module 3 for fusing the auxiliary information hologram into the target scene for augmented reality near-eye display.
[0025] The acquisition module 1 (e.g., a monocular camera 11) of this invention acquires images of the real work scene during the wearer's work. The processing module 2, based on monocular depth estimation technology, calculates target scene information (e.g., the distances of various real objects in the work scene image relative to the wearer) and generates an auxiliary information hologram corresponding to the target scene. The fusion module 3 projects the auxiliary information hologram to an arbitrary distance and achieves virtual-real fusion with the corresponding real objects. This invention avoids convergence-focusing conflicts, improving the virtual-real fusion experience. Simultaneously, this device can achieve continuous depth zoom without using a mechanical zoom module, which is expected to reduce the device's weight; it can also correct projection distortion without using high-precision optical components, which is expected to reduce the device's cost. In summary, this device has the advantages of system simplicity, compact design, continuous depth, and clear images. It is easy to implement and is expected to provide ideal augmented reality near-eye display effects, suitable for industrial applications such as warehouse management and intelligent manufacturing.
[0026] Please refer to Figure 2 , Figure 2 This is a schematic diagram of the specific structure of an augmented reality near-eye display device provided by the present invention.
[0027] In a preferred embodiment, processing module 2 includes: a scene information determination module, used to segment target objects from a target scene image based on a monocular depth estimation algorithm to determine target scene information; the target scene information includes the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eyes; and a hologram generation module, used to generate auxiliary information describing the target scene based on a hologram generation algorithm, according to the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eyes, and generate an auxiliary information hologram based on the auxiliary information and the three-dimensional coordinates.
[0028] Computational holography offers a novel approach to achieving continuously zoomed augmented reality near-eye displays. This technology calculates the hologram corresponding to the light intensity distribution on any focal plane using a diffraction algorithm, and then uses a spatial light modulator 34 to reproduce the corresponding light intensity distribution. Computational holography does not rely on a physical screen; the focal plane position, image size, and projection height of the projected pattern can all be flexibly adjusted through algorithms. The number of focal planes can theoretically be extended to infinity, and even ghosting, distortion, and astigmatism can be compensated for through algorithms. This avoids the use of mechanical zoom components and optical compensation components, and features simplicity, compactness, and strong virtual-real fusion capabilities, making it highly suitable for the field of augmented reality near-eye displays.
[0029] In this embodiment, a monocular camera 11 is used to acquire two-dimensional images of a real working scene; a monocular depth estimation algorithm is used to determine the distance of each real object in the real working scene relative to the wearer's eyes; a hologram generation algorithm is used to encode the auxiliary information corresponding to each real object into a phase-type hologram, and a spatial light modulator 34 is used to project the encoded auxiliary information in the hologram onto the depth of the corresponding real object, thus completing the virtual-real fusion of the auxiliary information and the real object. The device of this invention can achieve a continuously zooming augmented reality near-eye display effect.
[0030] It should be noted that, in order to illustrate the invention principle as intuitively as possible, only two target objects are shown in the target scene image, with the first target object 71 being a more distant object and the second target object 72 being a closer object. In reality, the number of target objects that this invention can collect is not limited to two, and the positional relationship between different target objects is not limited to the situation shown in the schematic diagram.
[0031] The first holographic auxiliary information 81 is the auxiliary information corresponding to the first target object 71 in the real scene. The second holographic auxiliary information 82 is the auxiliary information corresponding to the second target object 72 in the real scene. The first holographic auxiliary information 81 and the second holographic auxiliary information 82 are key to realizing the near-eye display of augmented reality through virtual-real fusion. After obtaining the intensity distribution and three-dimensional coordinates of the target object through the monocular depth estimation algorithm built into the processing module 2 (digital signal processing chip 21), the hologram generation algorithm will generate the corresponding auxiliary information. After obtaining the content of the auxiliary information, the hologram generation algorithm will use the three-dimensional coordinates of the target object relative to the wearer's eyes as calculation parameters to generate a phase-type hologram corresponding to the auxiliary information. After the phase-type hologram is loaded onto the surface of the spatial light modulator 34 through the signal cable 4, the loaded phase-type hologram will change the surface phase distribution of the spatial light modulator 34 and diffract under the illumination of polarized parallel light, projecting the image of the auxiliary information onto the vicinity of the corresponding target object. During use, the wearer can see both the first target object 71 and the second target object 72 in the real scene, as well as the first holographic auxiliary information 81 and the second holographic auxiliary information 82 simultaneously, achieving a fusion of virtual and real elements and providing the wearer with richer auxiliary information. The device does not require a zoom component or a reference object; it achieves multi-focal plane display through holograms, and the position of the focal planes can be freely configured according to the position of the real scene.
[0032] It should be noted that the auxiliary information shown in the holographic illustration is the name and distance of the target object in the real scene. In fact, the holographic auxiliary information that this invention can present is not limited to name and distance, and the positional relationship between the holographic auxiliary information and the target object is not limited to the situation shown in the illustration.
[0033] In a preferred embodiment, the scene information determination module is used to classify the target scene image based on mixed features; the mixed features include color, edge, contour, vanishing line and vanishing point; the image categories include distant image, mid-range image and close-up image; the depth of the distant image is estimated using the horizontal edge gradient method, the depth of the mid-range image is estimated using the vanishing line gradient method, and the depth of the close-up image is estimated using the weighted depth superposition method to obtain the target scene information.
[0034] In this embodiment, monocular depth estimation essentially extracts depth data through two-dimensional image features. During the depth estimation process, features such as edges, textures, colors, and motion parallax in the image can all play a role. Estimation techniques based on hybrid features have a wide range of applications and high estimation accuracy.
[0035] In a single-lens imaging system, the relationship between the size of an object in a photograph and its actual distance is as follows: , Among them, L dis S represents the actual distance of an object. siz The actual size of the object is represented by f, and the focal length of the lens is represented by S. siz ' represents the image size of an object in a two-dimensional photograph. When L dis When the image is infinitely large, the image of an object in a photograph approximates a single point, known as the "vanishing point." Any straight line drawn from the vanishing point is called a "vanishing line." Line segments connecting corresponding points on the vanishing line have approximately the same length in real space.
[0036] Vanishing points and vanishing lines are commonly found in long-range and medium-range images. Long-range images are those shot from a distance, usually outdoors, and include elements such as sky, land, and water. Medium-range images are those shot at a moderate distance, exhibiting perspective effects and displaying the "near objects appear larger, far objects smaller" characteristic. In long-range images, the vanishing point lies on the boundary between the sky and other physical elements. In medium-range images, the vanishing point is located near the center of the image and is the intersection of multiple vanishing lines.
[0037] It's important to note that, in addition to long-range and medium-range images, single-lens cameras can also capture close-up images. Close-up images, also known as extreme close-up shots, feature a subject whose size is close to or larger than the shooting distance, and other physical elements are obscured by the subject, making it difficult to extract the vanishing point and vanishing line from the positional relationships of the elements in the image.
[0038] The monocular depth estimation algorithm based on hybrid features first classifies the photo. The first step in the classification process is to convert the photo from RGB color to HSI color (hue, saturation, intensity). In the HSI color space, the pixel values of elements such as sky, land, and water are... V pix The following conditions must be met: , Here, H, S, and I represent hue, saturation, and intensity, respectively. When the number of pixels meeting the criteria exceeds 50% of the total number of pixels in the photo, the photo is classified as a distant view image.
[0039] When a photo is not classified as a distant view, the second step is to determine whether it is a mid-range view. If a vanishing point exists in the image, it is classified as a mid-range view; otherwise, it is classified as a close-up view.
[0040] To detect the presence of vanishing points in an image, the Canny operator is first used to extract edge features, and then the Hough transform is used to detect whether there are straight lines intersecting at a single point among these edge features. If many detected lines share exactly one common intersection point, then a vanishing point is considered to exist in the image.
[0041] After photo classification, different depth extraction models are needed to perform depth estimation for different photo categories. For distant images, the horizontal edge gradient method is used for depth estimation. This method assumes the sky is infinitely far away, and the actual distance of corresponding pixels in the image decreases linearly from the sky boundary to the bottom of the image. This process can be represented as: , in, B dep Represents the bit depth of the depth map. N This represents the number of pixels in the depth map in the vertical direction. y bo This represents the vertical coordinates of the boundary line. The larger the value of a pixel, the closer it is to the boundary. For example, when the depth map has a bit depth of 8 bits, the pixel representing the sky has a value of 0, the pixel representing the bottom of the image has a value of 255, and the pixel value increases linearly from 0 to 255 from the boundary line to the bottom of the image.
[0042] Mid-range images are obtained using the vanishing line gradient method, where the vanishing point is considered the farthest point. The vanishing line emanating from the vanishing point divides the image into horizontal and vertical planes, and the depth maps for these two planes need to be calculated separately. Therefore, the vanishing line gradient method first distinguishes between the vertical and horizontal directions, and then obtains the depth image using the following formula: in, D map_h and D map_v These represent the depth gradients on the horizontal and vertical planes, respectively. x vp , y vp() represents the coordinates of the vanishing point in a 2D photograph. M and N These represent the number of pixels in the horizontal and vertical directions of the depth map, respectively. In the horizontal direction, the pixel value of the depth map increases linearly from 0 to 255 along the column direction from the vanishing point to the image edge; in the vertical direction, the pixel value of the depth map increases linearly from 0 to 255 along the row direction from the vanishing point to the image edge.
[0043] The depth map of the close-up image is obtained using a weighted depth overlay method. This method assumes that closer image regions contain more detail, while farther image regions contain less detail. By counting the number of edge contours, the distance of different image regions relative to the photographer can be approximately determined. The method first uses the Canny algorithm to extract image contours and divides them into 5×5 rectangular regions. Then, it counts the number of edge contours in each rectangular region, defining the region with a higher number of edge contours than the average as the principal region. The sub-depth map of each principal region is elliptical in shape, with the pixel value at the center of the ellipse set to 255 and the pixel value on the circumference set to 0. The pixel value decreases linearly from the center to the circumference. All sub-depth maps are then weighted and overlaid using the following formula to obtain the global depth map: , in, D i Indicates the first i Sub-depth maps of the main regions D f This represents the depth map of the entire close-up image. A Indicates the number of main areas. N i 'Indicates the number of edge contours in each main region.
[0044] In a preferred embodiment, the scene information determination module is used to input the target scene image into a pre-trained monocular depth estimation network model with depth gradient assistance to obtain the target scene information output by the monocular depth estimation network model; wherein, the monocular depth estimation network model is trained by depth estimation training samples and depth gradient assistance samples.
[0045] In this embodiment, the Depth Gradient-Assisted Monocular Depth Estimation Network (DGE-CNN) achieves more generalizable monocular depth estimation. The DGE-CNN model consists of two network modules: an encoder (first encoder) and a decoder (first decoder). The encoder detects and acquires depth features from the 2D color image, while the decoder generates a depth map corresponding to the 2D image based on the acquired depth features. The encoder is a modified ResNet-50 network module. The 2D color image is input into the DGE-CNN model, first passing through a downsampling function block in the encoder, and then entering a network structure composed of several residual downsampling function blocks and residual projection function blocks connected in series. Both the residual downsampling function blocks and the residual projection function blocks consist of operations such as convolution, batch normalization (BN) operations, and linear rectified function (ReLU).
[0046] The feature information processed by the encoder then enters the decoder of the network model. The decoder is an upsampling network module, mainly composed of an upconvolutional residual function block, convolution, batch normalization (BN) operations, ReLU, and the Dropout 2D algorithm. The upconvolutional residual function block improves the network's predictive ability and generalization performance; the Dropout 2D algorithm alleviates overfitting during training, achieving a certain degree of regularization. After encoder processing, the output of the DGE-CNN model is a depth map of a two-dimensional color image.
[0047] The completed DGE-CNN model needs to be trained to obtain high-precision depth maps. This invention uses an input-output image pair training set to train the network. This invention also designs a DGE module to extract depth gradients, and incorporates these depth gradients as an auxiliary training set into the CNN model training. The DGE module utilizes two independent Sobel convolution operators. G x and G y This allows for the separate extraction of depth gradients in the horizontal and vertical directions. The depth gradient information used in the auxiliary training set is obtained by fusing the depth gradients in these two directions.
[0048] In a preferred embodiment, the hologram generation module is used to: perform image feature recognition based on intensity distribution to determine the type characteristics of the target scene; generate textual or graphic auxiliary information containing type identifiers based on the type characteristics, and generate numerical auxiliary information containing distance parameters based on the depth component in the three-dimensional coordinates; input the auxiliary information and three-dimensional coordinates into a pre-trained model without convolution error to drive an autoencoder deep learning network, and obtain an auxiliary information hologram output by the autoencoder deep learning network; wherein, the autoencoder deep learning network compensates for the encoding phase convolution error by extending the phase during the decoding stage.
[0049] To simultaneously ensure both accuracy and speed in the holographic algorithm, this embodiment uses a model without convolutional error to drive an autoencoder deep learning network, achieving a significant improvement in the high-precision hologram generation time. The target image to be displayed is the network input. After passing through the encoder (second encoder), it obtains the output phase, and then through the decoder (second decoder), it obtains the output amplitude. A loss function is used to establish the relationship between the input image and the output amplitude, and the network parameters are trained based on this. After the autoencoder deep learning network is trained, the encoder can be extracted separately as a tool for calculating pure phase holograms. The encoder can be a U-Net, where the input is the target image to be displayed, and the output is the pure phase hologram corresponding to the target image. The complete U-Net consists of a downsampling path and a corresponding upsampling path. The downsampling path increases the level of feature abstraction and consists of six downsampling modules and six corresponding residual modules. Each downsampling module consists of two sets of BN operations, a ReLU function, and a 3×3 convolutional layer. The upsampling path extracts high-level image features and restores the output to the same size as the input image. It consists of six upsampling modules. Each upsampling module is designed based on a subpixel convolution method, and its structure can be further decomposed into convolutional layers and transformation layers. The convolutional layers increase the number of channels, while the transformation layers restructure the data. The residual layers in the downsampling module are skipped to the corresponding upsampling modules to mitigate degradation issues during network training. After the upsampling path, the output data first passes through a tanh function layer to constrain the phase value of the output. Within the range.
[0050] The decoder's input information is the encoder's output phase, and its initial resolution is... N x × N y To avoid convolution errors, this phase distribution needs to be padded with zeros to achieve a resolution of 2. N x × 2 N yThe extended phase is used as the phase distribution on the holographic plane, while the intensity matrix, with all elements equal to 1, is used as the amplitude distribution on the holographic plane. The amplitude and phase together constitute the complex amplitude on the holographic plane, which, after being decoded, results in the complex amplitude on the object plane. The decoder is the forward propagation process of the non-convolutional angular spectrum model: , in, E O ( x , y , z 0) represents the amplitude distribution of the target image to be displayed. z 0 represents the distance between the object plane and the holographic plane. This represents the initial phase superimposed on the target image. The wavelength is represented by FT, which stands for Fourier Transform. u and v These represent the spatial frequencies in the horizontal and vertical directions, respectively. Typically, the reference plane in the angular spectrum model is the holographic plane, and its depth coordinates... z =0. Extract the amplitude distribution on the extract plane and clip its resolution to... N x × N y The final output of the network can then be obtained. The amplitude of this output can be correlated with the input image using a loss function. Based on the calculation results of the loss function, parameter training of an autoencoder deep learning network without convolutional error can be achieved.
[0051] The training process of a deep learning network first requires calculating the loss function between the input and output images, and then updating the relevant parameters of the convolutional kernel based on the calculation results. The specific loss function used in this invention is the negative Pearson correlation coefficient, which can increase the convergence probability of the network parameters. Its mathematical form can be expressed as follows: , in, I and I r These represent the intensity of the original image and the intensity of the reconstructed image, respectively, and their specific values can be obtained from the amplitude distribution on the object plane.
[0052] In a preferred embodiment, the fusion module 3 includes: a light source module for generating collimated polarized illumination light; an optical modulation module disposed at the output end of the light source module, the input end of the optical modulation module being electrically connected to the output end of the hologram generation module, for loading an auxiliary information hologram, which diffracts after being illuminated by the polarized illumination light, projecting the auxiliary information hologram onto a preset position of a nearby target scene; and a waveguide display module disposed at the diffracted light wave output end of the optical modulation module, for transmitting the diffracted light wave to the user's eye area to perform virtual-real fusion of the auxiliary information hologram and the target scene.
[0053] In a preferred embodiment, the light source module includes: a laser light source for emitting diverging spherical waves; a collimating lens disposed at the output end of the laser light source for converging the diverging spherical waves and outputting collimated plane waves; and a polarizer 33 disposed at the collimated plane wave output end of the collimating lens, the output end of which serves as the output end of the light source module for changing the polarization state of the collimated plane wave and outputting collimated polarized illumination light.
[0054] In a preferred embodiment, the optical modulation module is a spatial light modulator 34, which is used to load a phase-type hologram corresponding to auxiliary information. The phase-type hologram is used to change the surface phase distribution of the spatial light modulator 34.
[0055] In this embodiment, the acquisition module 1 is a camera 11, which can be a commonly used monocular camera 11, used to acquire two-dimensional images of the real scene. The camera 11 is connected to the digital signal processing chip 21 via the signal cable 4. The acquired two-dimensional images will be used as input to the monocular depth estimation algorithm to subsequently determine the distance of each real object in the real scene relative to the wearer's eyes.
[0056] Signal cable 4 is mainly used to connect components such as camera 11, laser diode 31 and spatial light modulator 34 to digital signal processing chip 21 to realize functions such as power supply and signal transmission.
[0057] Processing module 2 is a digital signal processing chip 21, mainly responsible for tasks such as power supply to various components, signal acquisition, monocular depth estimation, and hologram generation. Processing module 2 includes a scene information determination module and a hologram generation module.
[0058] The digital signal processing chip 21 is connected to the battery 5 via a power supply cable, and further supplies power to components such as the camera 11, laser diode 31, and spatial light modulator 34 via the signal cable 4.
[0059] The digital signal processing chip 21 is connected to the camera 11 via the signal cable 4. After the camera 11 acquires a two-dimensional image of the real scene, it transmits the image back to the digital signal processing chip 21 via the signal cable 4.
[0060] The scene information determination module incorporates a monocular depth estimation algorithm. After the two-dimensional image captured by camera 11 is processed by the monocular depth estimation algorithm, each target scene contained in the image is segmented, and the three-dimensional coordinates of each target scene relative to the wearer's eyes are obtained. The segmented target scenes and their corresponding three-dimensional coordinates will be used for subsequent hologram generation.
[0061] The hologram generation module has a built-in hologram generation algorithm. This algorithm first needs to acquire two types of data: the intensity distribution of each target object after image segmentation and the three-dimensional coordinates of each target object. After acquiring the intensity distribution and three-dimensional coordinates of the target objects, the hologram generation algorithm will generate corresponding auxiliary information. For example, if the hologram generation algorithm identifies the target object type as "star" and "water droplet," it will generate the text auxiliary information "STAR" and "DROP"; if it identifies the target object distance as "0.5m" and "0.6m," it will generate the text auxiliary information "Distance:0.5m" and "Distance:0.6m."
[0062] Considering that the projection position of the auxiliary information is near the corresponding target object and the projection angle should be within the wearer's eye box (user's eye area), after obtaining the content of the auxiliary information, the hologram generation algorithm will use the three-dimensional coordinates of the target object relative to the wearer's eyes as calculation parameters to generate a phase-type hologram corresponding to the auxiliary information.
[0063] The digital signal processing chip 21 is connected to the laser diode 31 via signal cable 4. Under the power supply and control of the digital signal processing chip 21, the laser diode 31 emits red, green, and blue illumination lasers as needed.
[0064] The digital signal processing chip 21 is connected to the spatial light modulator 34 (optical modulation module) via signal cable 4. Under the power supply and control of the digital signal processing chip 21, the spatial light modulator 34 loads the phase-type holograms corresponding to each color component as required.
[0065] Battery 5 is connected to digital signal processing chip 21 via power supply cable, and further connected to components such as camera 11, laser diode 31 and spatial light modulator 34 via signal cable 4, and supplies power to the above components.
[0066] The fusion module 3 includes a light source module, an optical modulation module, and a waveguide 36 display module. The light source module includes a laser light source, a collimating lens, and a polarizer 33.
[0067] The laser diode 31 (laser source) is connected to the digital signal processing chip 21 via the signal cable 4. It can emit illumination lasers with red, green and blue color components as needed, which are used to illuminate the phase-type hologram loaded on the spatial light modulator 34 and project corresponding auxiliary information near the target object through diffraction.
[0068] The collimating lens is a convex lens 32, whose function is to converge diverging spherical waves to obtain a collimated plane wave.
[0069] It should be noted that in this invention, the laser diode 31 is considered an ideal point light source, and the corresponding wavefront shape is a diverging spherical wave. The laser diode 31 is placed on the object-side focal plane of the convex lens 32, and the distance from the point light source to the convex lens 32 is exactly the focal length of the collimating lens. At this point, the diverging spherical wave is collimated into a plane wave. Depending on actual needs, a filter element can be added to the optical path to ensure that the emitted light wave is a diverging spherical wave.
[0070] The function of polarizer 33 is to change the polarization state of the collimated plane wave. The spatial light modulator 34 used in this invention typically has polarization selectivity. By rotating polarizer 33 and changing its fast axis angle, the polarization state of the collimated plane wave can be altered. After passing through polarizer 33, the plane wave emitted by convex lens 32 will have a specific polarization state.
[0071] The optical modulation module is a spatial light modulator 34, which is used to load holograms. The spatial light modulator 34 is connected to the digital signal processing chip 21 via signal cable 4. After the digital signal processing chip 21 generates a phase-type hologram corresponding to the auxiliary information, it loads the corresponding phase-type hologram onto the surface of the spatial light modulator 34 via signal cable 4. The loaded phase-type hologram changes the surface phase distribution of the spatial light modulator 34. When polarized parallel light emitted from polarizer 33 illuminates the surface of the spatial light modulator 34 with a specific phase distribution, diffraction occurs, thereby projecting the image of the auxiliary information onto the vicinity of the corresponding target object, and ensuring that the projected light reaches the wearer's eye box range after passing through the coupling grating 35, waveguide 36, first coupling grating 37, and second coupling grating 38.
[0072] The waveguide display module includes an input grating 35, a waveguide 36, a first output grating 37, and a second output grating 38.
[0073] The coupling grating 35 is used to couple the diffracted light wave that has passed through the spatial light modulator 34 into the waveguide 36 and propagate it in a specified direction.
[0074] Waveguide 36 utilizes the principle of total internal reflection to enable the light waves introduced by the coupled grating 35 to be transmitted over long distances inside the substrate, thus avoiding light leakage and loss.
[0075] The first coupling grating 37 is used to couple the light waves propagating inside the waveguide 36 into free space and into the wearer's left eye region.
[0076] The second coupling grating 38 is used to couple the light waves propagating inside the waveguide 36 into free space and into the wearer's right eye region.
[0077] The support frame 6 is used to fix and protect the relevant components in the near-eye display device.
[0078] The augmented reality near-eye display method provided by the present invention is described below. The augmented reality near-eye display method described below can be referred to in correspondence with the augmented reality near-eye display device described above.
[0079] Please refer to Figure 3 , Figure 3 This is a flowchart illustrating an augmented reality near-eye display method provided by the present invention.
[0080] The present invention also provides an augmented reality near-eye display method, comprising: 301: Acquire the target scene image; 302: Based on the target scene image, determine the target scene information, and based on the target scene information, generate an auxiliary information hologram corresponding to the target scene; 303: Integrate auxiliary information holograms into the target scene for augmented reality near-eye display.
[0081] As a preferred embodiment, the method involves determining target scene information based on a target scene image and generating an auxiliary information hologram corresponding to the target scene based on the target scene information. This includes: segmenting the target scene from the target scene image using a monocular depth estimation algorithm to determine the target scene information; the target scene information includes the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eyes; generating auxiliary information describing the target scene based on a hologram generation algorithm, according to the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eyes; and generating an auxiliary information hologram based on the auxiliary information and the three-dimensional coordinates.
[0082] Please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating the specific process of an augmented reality near-eye display method provided by the present invention.
[0083] The complete workflow of this invention is as follows: (1): Battery 5 is connected to digital signal processing chip 21 via power supply cable, and further connected to components such as camera 11, laser diode 31 and spatial light modulator 34 via signal cable 4, and supplies power to the above components. Camera 11 captures two-dimensional images of the real scene and transmits them to digital signal processing chip 21 via signal cable 4.
[0084] (2): The two-dimensional image of the real scene is processed by the monocular depth estimation algorithm built into the digital signal processing chip 21. After processing, the two-dimensional image of the real scene is divided into the first target scene 71 and the second target scene 72 in the real scene, and their three-dimensional coordinates relative to the wearer's eyes are obtained.
[0085] (3): Using the intensity distribution and three-dimensional coordinates of the first target object 71 and the second target object 72 in the real scene as input, the hologram generation algorithm built into the digital signal processing chip 21 generates their respective auxiliary information. For example, if the hologram generation algorithm identifies the type of the first target object 71 in the real scene as "star", then the text auxiliary information "STAR" is generated; if the monocular depth estimation algorithm obtains that the distance of the first target object 71 in the real scene is 0.5m, then the text auxiliary information "Distance:0.5m" is generated. Similarly, the auxiliary information corresponding to the second target object 72 in the real scene is "DROP, Distance:0.6m".
[0086] (4): The phase-type hologram corresponding to the auxiliary information is generated by the hologram generation algorithm built into the digital signal processing chip 21. Considering that the projection position of the auxiliary information should be near the corresponding target object and the angle of the projection light should be within the wearer's eye box, the hologram generation algorithm will use the three-dimensional coordinates of the target object as the calculation parameter and take into account the wearer's viewing angle to generate the phase-type hologram corresponding to the auxiliary information.
[0087] (5): Achieving optical projection of auxiliary information. The phase hologram generated by the digital signal processing chip 21 is loaded onto the surface of the spatial light modulator 34 via the signal cable 4. At the same time, the digital signal processing chip 21 is connected to the laser diode 31 via the signal cable 4 to synchronize the timing of the illumination light and the timing of the phase hologram. The laser emitted by the laser diode 31 passes through the convex lens 32 and the polarizer 33 to obtain polarized plane illumination light, which diffracts after illuminating the phase hologram loaded on the spatial light modulator 34. The diffracted light passes through the coupling grating 35, the waveguide 36, the first coupling grating 37 and the second coupling grating 38 and then reaches the eye box range of the wearer's left and right eyes, respectively.
[0088] (6) When the wearer looks straight ahead, they can simultaneously see the first target scene 71 and the second target scene 72 in the real scene, as well as their corresponding first holographic auxiliary information 81 and second holographic auxiliary information 82, thus achieving virtual-real fusion. At the same time, the first holographic auxiliary information 81 and the second holographic auxiliary information 82 supplement the first target scene 71 and the second target scene 72 in the real scene, respectively, and play a role in augmenting reality.
[0089] The present invention has the following beneficial effects: The device has a simple structure. In this invention, wavefront error can be compensated by optimizing the phase hologram without using composite holographic elements and without involving joint hardware and software optimization; the determination of projection distance and angle can be achieved by using a monocular camera and a monocular depth estimation algorithm without using a three-dimensional measurement device; the projection distance can be adjusted by changing the calculation parameters of the phase hologram without using zoom elements or superlenses; except for the collimating lens, it can be designed entirely based on a lensless scheme, and the distance of each projected image can be configured according to the position of the real scene.
[0090] A unique virtual-real fusion scheme. Based on computational holography, this invention enables the levitation projection of auxiliary information of any pattern at any spatial location, and has the advantages of continuous depth and clear image, providing an ideal virtual-real fusion augmented reality near-eye display experience.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An augmented reality near-eye display device, characterized in that, include: The acquisition module is used to acquire images of the target scene; The processing module is used to determine target scene information based on the target scene image, and generate an auxiliary information hologram corresponding to the target scene based on the target scene information; The fusion module is used to fuse the auxiliary information hologram into the target scene for augmented reality near-eye display.
2. The augmented reality near-eye display device according to claim 1, characterized in that, The processing module includes: The scene information determination module is used to segment target objects from the target scene image based on a monocular depth estimation algorithm and determine the target scene information; the target scene information includes the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eye; The hologram generation module is used to generate auxiliary information describing the target scene based on the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eye, and to generate the auxiliary information hologram based on the auxiliary information and the three-dimensional coordinates.
3. The augmented reality near-eye display device according to claim 2, characterized in that, The scene information determination module is used to classify the target scene image based on blended features; the blended features include color, edge, contour, vanishing line, and vanishing point; the image categories include distant image, mid-range image, and close-up image; The depth of the distant image is estimated using the horizontal edge gradient method, the depth of the mid-range image is estimated using the vanishing line gradient method, and the depth of the near-range image is estimated using the weighted depth stacking method, thereby obtaining the target scene information.
4. The augmented reality near-eye display device according to claim 2, characterized in that, The scene information determination module is used to input the target scene image into a pre-trained depth gradient-assisted monocular depth estimation network model to obtain the target scene information output by the monocular depth estimation network model. The monocular depth estimation network model is obtained by training depth estimation training samples and depth gradient auxiliary samples.
5. The augmented reality near-eye display device according to any one of claims 2 to 4, characterized in that, The hologram generation module is used for: Image feature recognition is performed based on the intensity distribution to determine the type characteristics of the target scene; Based on the type characteristics, generate textual or graphic auxiliary information containing type identifiers, and generate numerical auxiliary information containing distance parameters based on the depth component in the three-dimensional coordinates; The auxiliary information and the three-dimensional coordinates are input into a pre-trained model without convolution error to drive an autoencoder deep learning network, thereby obtaining the auxiliary information hologram output by the autoencoder deep learning network. The autoencoder deep learning network compensates for the coding phase convolution error during the decoding stage by extending the phase.
6. The augmented reality near-eye display device according to claim 5, characterized in that, The fusion module includes: A light source module is used to generate collimated polarized illumination light; An optical modulation module is disposed at the output end of the light source module. The input end of the optical modulation module is electrically connected to the output end of the hologram generation module. It is used to load the auxiliary information hologram. After the polarized illumination light illuminates the auxiliary information hologram, diffraction occurs, and the auxiliary information hologram is projected onto a preset position of a nearby target scene. A waveguide display module is disposed at the diffraction light wave output end of the optical modulation module, and is used to transmit the diffraction light wave to the user's eye area to fuse the auxiliary information hologram and the target scene in a virtual-real manner.
7. The augmented reality near-eye display device according to claim 6, characterized in that, The light source module includes: Laser light source, used to emit diverging spherical waves; A collimating lens is disposed at the output end of the laser source to converge the diverging spherical wave and output a collimated plane wave. A polarizer is disposed at the collimated plane wave output end of the collimating lens. The output end of the polarizer serves as the output end of the light source module, used to change the polarization state of the collimated plane wave and output the collimated polarized illumination light.
8. The augmented reality near-eye display device according to claim 6, characterized in that, The optical modulation module is a spatial light modulator used to load a phase-type hologram corresponding to auxiliary information. The phase-type hologram is used to change the surface phase distribution of the spatial light modulator.
9. An augmented reality near-eye display method, characterized in that, include: Acquire the target scene image; Based on the target scene image, target object information is determined, and based on the target object information, an auxiliary information hologram corresponding to the target object is generated; The auxiliary information hologram is fused into the target scene for augmented reality near-eye display.
10. The augmented reality near-eye display method according to claim 9, characterized in that, The step of determining target scene information based on the target scene image and generating an auxiliary information hologram corresponding to the target scene based on the target scene information includes: The target scene is segmented from the target scene image based on a monocular depth estimation algorithm to determine the target scene information; the target scene information includes the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eye. Based on the hologram generation algorithm, auxiliary information describing the target scene is generated according to the intensity distribution of the target scene and the three-dimensional coordinates of the target scene relative to the user's eyes, and the auxiliary information hologram is generated according to the auxiliary information and the three-dimensional coordinates.