Augmented reality head-up display device and method

By generating auxiliary information holograms through monocular depth estimation algorithms and deep learning networks, the problem of limited focal planes and poor virtual-real fusion capabilities in existing augmented reality head-up display devices is solved, and a simple and compact augmented reality head-up display effect is achieved.

CN122018165APending Publication Date: 2026-05-12TSINGHUA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2026-02-05
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing augmented reality head-up display devices have a limited number of focal planes, making it impossible to achieve accurate virtual-real fusion with real objects at different depths. The system structure is complex and has low reliability.

Method used

An augmented reality head-up display device based on monocular depth estimation algorithm and deep learning network is used to generate auxiliary information holograms by acquiring road scene and driver facial images, and to achieve virtual-real fusion by using optical components, avoiding mechanical zoom components and high-precision optical compensation elements.

Benefits of technology

It achieves precise virtual-real fusion of auxiliary information and real road scenery at any depth, without the need for mechanical zoom components. The system has a simple and compact structure and is suitable for intelligent cockpit environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018165A_ABST
    Figure CN122018165A_ABST
Patent Text Reader

Abstract

The invention provides an augmented reality head-up display device and method, and the device comprises an obtaining module which is used for obtaining a road scene image, a driver face image, and parameters of optical components; the processing module is used for determining target scenery information and driver eye position information according to the road scene image and the driver face image, and generating an auxiliary information hologram corresponding to the target scenery according to the parameters of the optical components, the target scenery information and the driver eye position information; and the fusion module is used for fusing the auxiliary information hologram to a road scene through an optical component and carrying out augmented reality head-up display. Accurate virtual-real fusion of auxiliary information and real road scenery at any depth can be realized, a mechanical zoom assembly is not needed, and the system is simple and compact in structure and suitable for an intelligent cabin environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of display technology, and more particularly to an augmented reality head-up display device and method. Background Technology

[0002] Head-up display (HUD) technology projects driving information into the driver's field of vision, allowing the driver to obtain key driving information without looking down, significantly improving driving safety. Augmented reality (AR) HUD technology goes a step further, merging virtual assistance information with real-world road scenes to provide drivers with more intuitive and comprehensive driving assistance functions.

[0003] Current augmented reality head-up display devices primarily use single-focal-plane or dual-focal-plane projection, which limits the depth range they can cover and results in poor virtual-real fusion capabilities. To achieve multi-depth coverage, existing solutions typically require the addition of mechanical zoom components or high-precision optical compensation elements, which not only increases the engineering integration difficulty of the device but also reduces its long-term reliability. Furthermore, while a few solutions using oblique projection technology can construct continuously variable focal planes, they usually only allow navigation arrows to move close to the ground, making it difficult to achieve accurate virtual-real fusion with other real-world objects (such as roadside trees, vehicles ahead, etc.). Summary of the Invention

[0004] This invention provides an augmented reality head-up display device and method to address the shortcomings of existing head-up display devices, such as limited focal planes, inability to achieve virtual-real fusion with real-world objects at different depths, complex system structure, and low reliability. This invention enables precise virtual-real fusion of auxiliary information with real-world road scenery at any depth, eliminates the need for mechanical zoom components, and features a simple and compact system structure, making it suitable for smart cockpit environments.

[0005] This invention provides an augmented reality head-up display device, comprising: an acquisition module for acquiring a road scene image, a driver's facial image, and parameters of optical components; a processing module for determining target object information and driver's eye position information based on the road scene image and the driver's facial image, and generating an auxiliary information hologram corresponding to the target object based on the parameters of the optical components, the target object information, and the driver's eye position information; and a fusion module for fusing the auxiliary information hologram into the road scene through the optical components to perform augmented reality head-up display.

[0006] According to an augmented reality head-up display device provided by the present invention, the processing module includes: a first information determination module, used to segment target objects from a road scene image based on a monocular depth estimation algorithm, and determine the target object information; the target object information includes the intensity distribution of the target object and the three-dimensional coordinates of the target object; a second information determination module, used to obtain the three-dimensional coordinates of the driver's eyes from a driver's face image based on a monocular depth estimation algorithm; and a hologram generation module, used to generate auxiliary information describing the target object based on the intensity distribution of the target object, the three-dimensional coordinates of the target object, and the three-dimensional coordinates of the driver's eyes, and to generate the auxiliary information hologram based on the auxiliary information, the three-dimensional coordinates of the target object, the three-dimensional coordinates of the driver's eyes, and the parameters of the optical components.

[0007] According to an augmented reality head-up display device provided by the present invention, the first information determination module is used to classify the road scene image according to the mixed features; the mixed features include color, edge, contour, vanishing line and vanishing point; the image categories include distant image, mid-range image and close-up image; the depth of the distant image is estimated by using the horizontal edge gradient method, the depth of the mid-range image is estimated by using the vanishing line gradient method, and the depth of the close-up image is estimated by using the weighted depth superposition method to obtain the target scene information.

[0008] According to an augmented reality head-up display device provided by the present invention, the second information determination module is used to input the driver's facial image into a pre-trained depth gradient-assisted monocular depth estimation network model to obtain the three-dimensional coordinates of the driver's eyes output by the monocular depth estimation network model; wherein, the monocular depth estimation network model is trained by depth estimation training samples and depth gradient-assisted samples.

[0009] According to an augmented reality head-up display device provided by the present invention, the hologram generation module is used to: perform image feature recognition based on the intensity distribution to determine the type characteristics of the target scene; generate textual or graphic auxiliary information containing type identifiers according to the type characteristics, and generate numerical auxiliary information containing distance parameters according to the depth component in the three-dimensional coordinates of the target scene and the three-dimensional coordinates of the driver's eyes; input the textual or graphic auxiliary information, numerical auxiliary information, the three-dimensional coordinates of the target scene, the three-dimensional coordinates of the driver's eyes, and the parameters of the optical components into a pre-trained model without convolution error to drive an autoencoder deep learning network, and obtain the auxiliary information hologram output by the autoencoder deep learning network; wherein, the autoencoder deep learning network compensates for the encoding phase convolution error by extending the phase during the decoding stage.

[0010] According to an augmented reality head-up display device provided by the present invention, the fusion module includes: a light source module for generating collimated polarized illumination light; an optical modulation module disposed at the output end of the light source module, the input end of the optical modulation module being electrically connected to the output end of the hologram generation module, for loading the auxiliary information hologram, and for diffraction occurring after the polarized illumination light illuminates the auxiliary information hologram, projecting the auxiliary information hologram onto a preset position of a nearby target scene; a display module disposed at the diffracted light wave output end of the optical modulation module, for reflecting the diffracted light wave to the driver's eye area to perform virtual-real fusion of the auxiliary information hologram and the road scene; and parameters of the optical components of the light source module, the optical modulation module, and the display module for fine-tuning the calculation parameters of the auxiliary information hologram.

[0011] According to an augmented reality head-up display device provided by the present invention, the light source module includes: a laser light source for emitting diverging spherical waves; a collimating lens disposed at the output end of the laser light source for converging the diverging spherical waves and outputting collimated plane waves; and a polarizer disposed at the collimated plane wave output end of the collimating lens, the output end of the polarizer serving as the output end of the light source module for changing the polarization state of the collimated plane waves and outputting collimated polarized illumination light.

[0012] According to an augmented reality head-up display device provided by the present invention, the optical modulation module is a spatial light modulator, which is used to load a phase-type hologram corresponding to auxiliary information, and the phase-type hologram is used to change the surface phase distribution of the spatial light modulator.

[0013] The present invention also provides an augmented reality head-up display method, comprising: acquiring a road scene image, a driver's facial image, and parameters of optical components; determining target object information and driver's eye position information based on the road scene image and the driver's facial image, and generating an auxiliary information hologram corresponding to the target object based on the parameters of the optical components, the target object information, and the driver's eye position information; and fusing the auxiliary information hologram into the road scene through the optical components to perform augmented reality head-up display.

[0014] According to an augmented reality head-up display method provided by the present invention, the step of determining target scene information and driver's eye position information based on the road scene image and the driver's facial image, and generating an auxiliary information hologram corresponding to the target scene based on the parameters of the optical components, the target scene information, and the driver's eye position information, includes: segmenting the target scene from the road scene image based on a monocular depth estimation algorithm to determine the target scene information; the target scene information includes the intensity distribution of the target scene and the three-dimensional coordinates of the target scene; obtaining the three-dimensional coordinates of the driver's eyes from the driver's facial image based on the monocular depth estimation algorithm; generating auxiliary information describing the target scene based on the intensity distribution of the target scene, the three-dimensional coordinates of the target scene, and the three-dimensional coordinates of the driver's eyes based on a hologram generation algorithm, and generating the auxiliary information hologram based on the auxiliary information, the three-dimensional coordinates of the target scene, the three-dimensional coordinates of the driver's eyes, and the parameters of the optical components.

[0015] The present invention provides an augmented reality head-up display device and method that can achieve accurate virtual-real fusion of auxiliary information and real road scenery at any depth, without the need for mechanical zoom components, and the system structure is simple and compact, making it suitable for smart cockpit environments. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the structure of an augmented reality head-up display device provided by the present invention.

[0018] Figure 2 This is a schematic diagram of the specific structure of an augmented reality head-up display device provided by the present invention.

[0019] Figure 3 This is a flowchart illustrating an augmented reality head-up display method provided by the present invention.

[0020] Figure 4 This is a schematic diagram illustrating the specific process of an augmented reality head-up display method provided by the present invention.

[0021] Figure label: 1: Acquisition module; 2: Processing module; 3: Fusion module; 11: First camera; 12: Second camera; 21: Signal processing and control board; 31: Laser emitter; 32: Optical fiber; 33: Laser output aperture; 34: Collimating lens; 35: Polarizer; 36: Spatial light modulator; 37: Windshield; 41: Target scene; 42: Auxiliary information; 43: Instrument panel. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0023] Currently, mass-produced augmented reality head-up display devices are mainly single-focal-plane, with a few iterating to dual-focal-plane. The depth range they can cover is relatively limited, and their ability to blend virtual and real objects is poor. A very small number of continuous zoom devices mainly achieve multi-depth coverage through oblique projection, but they can usually only make the navigation arrows close to the ground, making it difficult to achieve virtual-real fusion with other real objects.

[0024] Please refer to Figure 1 , Figure 1 This is a schematic diagram of an augmented reality head-up display device provided by the present invention.

[0025] This invention provides an augmented reality head-up display device, comprising: an acquisition module 1 for acquiring a road scene image and a driver's facial image; a processing module 2 for determining target scene information and driver's eye position information based on the road scene image, driver's facial image, and parameters of optical components, and generating an auxiliary information hologram corresponding to the target scene 41 based on the parameters of the optical components, the target scene information, and the driver's eye position information; and a fusion module 3 for fusing the auxiliary information hologram into the road scene through optical components to perform augmented reality head-up display.

[0026] The acquisition module 1 of this invention (which may include a first camera 11 and a second camera 12) acquires images of real road scenes and driver facial images during driving. Of course, the acquisition module 1 can also be used to acquire parameters of the optical components in the fusion module 3. The processing module 2, based on monocular depth estimation technology, calculates target scene information (e.g., the distances of various real objects in the road scene relative to the car) and driver eye position information (the three-dimensional coordinates of the driver's eyes relative to the car), and generates an auxiliary information hologram corresponding to the target scene 41 based on the parameters of the optical components, the target scene information, and the driver's eye position information. The fusion module 3 projects the auxiliary information hologram to any distance using optical components and achieves virtual-real fusion with the corresponding real scene. The device of this invention can achieve continuous depth zoom without the use of a mechanical zoom device and can correct projection distortion without the use of high-precision optical components. It has the advantages of simple system, compact device, continuous depth, and clear image, and is easy to implement in engineering, and is expected to provide an ideal user experience.

[0027] Please refer to Figure 2 , Figure 2 This is a schematic diagram of the specific structure of an augmented reality head-up display device provided by the present invention.

[0028] In a preferred embodiment, processing module 2 includes: a first information determination module, used to segment target scene 41 from a road scene image based on a monocular depth estimation algorithm, and determine target scene information; the target scene information includes the intensity distribution of target scene 41 and the three-dimensional coordinates of target scene 41; a second information determination module, used to obtain the three-dimensional coordinates of the driver's eyes from a driver's face image based on a monocular depth estimation algorithm; and a hologram generation module, used to generate auxiliary information 42 describing target scene 41 based on the intensity distribution of target scene 41, the three-dimensional coordinates of target scene 41 and the three-dimensional coordinates of the driver's eyes, and to generate an auxiliary information hologram based on the auxiliary information 42, the three-dimensional coordinates of target scene 41, the three-dimensional coordinates of the driver's eyes and the parameters of optical components.

[0029] Computational holography offers a new approach to realizing continuously zoomed augmented reality head-up displays. This technology calculates the hologram corresponding to the light intensity distribution on any focal plane using a diffraction algorithm, and then uses a spatial light modulator 36 to reproduce the corresponding light intensity distribution. Computational holography does not rely on a physical screen; the focal plane position, image size, and projection height of the projected pattern can all be flexibly adjusted through algorithms. The number of focal planes can theoretically be extended to infinity, and even ghosting, distortion, and astigmatism can be compensated for through algorithms. This avoids the use of mechanical zoom components and optical compensation components, and features a simple and compact system with strong virtual-real fusion capabilities, making it highly suitable for the field of augmented reality head-up displays.

[0030] In this embodiment, a first camera 11 (monocular camera) is used to acquire two-dimensional images of a real road scene; a second camera 12 (monocular camera) is used to acquire two-dimensional images of the driver's face. The monocular depth estimation algorithm in the first information determination module determines the three-dimensional coordinates (including intensity distribution and three-dimensional coordinates) of each real object relative to the car in the real road scene; the monocular depth estimation algorithm in the second information determination module determines the three-dimensional coordinates of the driver's eyes. The hologram generation module, combining the intensity distribution of the target object 41, the three-dimensional coordinates of the target object 41, the three-dimensional coordinates of the driver's eyes, and the parameters of the optical components, encodes the auxiliary information 42 corresponding to each real object into a phase-type hologram. A spatial light modulator 36 then projects the encoded auxiliary information 42 from the hologram onto the depth of the corresponding real object, completing the virtual-real fusion of the auxiliary information 42 and the real object. The device of this invention can achieve a continuously zooming augmented reality head-up display effect. The phase-type hologram generated by this invention can be fine-tuned according to the parameters of the optical components, thereby ensuring the clarity of the final augmented reality head-up display.

[0031] The device does not require a zoom component or a reference object; it can achieve multi-focal plane display through holograms, and the position of the focal plane can be freely configured according to the position of the real scene.

[0032] It should be noted that the auxiliary information 42 shown in the holographic illustration is the name and distance of the target object 41 in the real scene. In fact, the holographic auxiliary information 42 that can be presented by this invention is not limited to name and distance, and the positional relationship between the holographic auxiliary information 42 and the target object 41 is not limited to the situation shown in the illustration.

[0033] In a preferred embodiment, the first information determination module is used to classify road scene images based on mixed features; the mixed features include color, edge, contour, vanishing line and vanishing point; the image categories include distant image, mid-range image and close-up image; the depth of the distant image is estimated using the horizontal edge gradient method, the depth of the mid-range image is estimated using the vanishing line gradient method, and the depth of the close-up image is estimated using the weighted depth overlay method to obtain the target scene information.

[0034] In this embodiment, monocular depth estimation essentially extracts depth data through two-dimensional image features. During the depth estimation process, features such as edges, textures, colors, and motion parallax in the image can all play a role. Estimation techniques based on hybrid features have a wide range of applications and high estimation accuracy.

[0035] In a single-lens imaging system, the relationship between the size of an object in a photograph and its actual distance is as follows: , in,L dis Indicates the actual distance of the object. S siz Indicates the actual size of the object. f Indicates the focal length of the lens. S siz 'Indicates the image size of an object in a 2D photograph. When L dis When the image is infinitely large, the image of an object in the photograph approximates a single point, called the vanishing point. Any straight line drawn from the vanishing point is called a vanishing line. Line segments connecting corresponding points on the vanishing line have approximately the same length in real space.

[0036] Vanishing points and vanishing lines are commonly found in long-range and medium-range images. Long-range images are those taken at a considerable distance, typically outdoors, and include elements such as sky, land, and water. Medium-range images are those taken at a moderate distance, exhibiting perspective effects, where objects appear larger in the foreground and smaller in the background. In long-range images, the vanishing point lies on the boundary between the sky and other physical elements. In medium-range images, the vanishing point is located near the center of the image and is the intersection of multiple vanishing lines.

[0037] It's important to note that, in addition to long-range and medium-range images, single-lens cameras can also capture close-up images. Close-up images, also known as extreme close-up shots, feature a subject whose size is close to or larger than the shooting distance, and other physical elements are obscured by the subject, making it difficult to extract the vanishing point and vanishing line from the positional relationships of the elements in the image.

[0038] The monocular depth estimation algorithm based on hybrid features first classifies the photo. The first step in the classification process is to convert the photo from RGB color to HSI color (hue, saturation, intensity). In the HSI color space, the pixel values ​​of elements such as sky, land, and water are... V pix The following conditions must be met: , Here, H, S, and I represent hue, saturation, and intensity, respectively. When the number of pixels meeting the criteria exceeds 50% of the total number of pixels in the photo, the photo is classified as a distant view image.

[0039] When a photo is not classified as a distant view, the second step is to determine whether it is a mid-range view. If a vanishing point exists in the image, it is classified as a mid-range view; otherwise, it is classified as a close-up view.

[0040] To detect the presence of vanishing points in an image, the Canny operator is first used to extract edge features, and then the Hough transform is used to detect whether there are straight lines intersecting at a single point among these edge features. If many detected lines share exactly one common intersection point, then a vanishing point is considered to exist in the image.

[0041] After photo classification, different depth extraction models are needed to perform depth estimation for different photo categories. For distant images, the horizontal edge gradient method is used for depth estimation. This method assumes the sky is infinitely far away, and the actual distance of corresponding pixels in the image decreases linearly from the sky boundary to the bottom of the image. This process can be represented as: , in, B dep Represents the bit depth of the depth map. N This represents the number of pixels in the depth map in the vertical direction. y bo This represents the vertical coordinates of the boundary line. The larger the value of a pixel, the closer it is to the boundary. For example, when the depth map has a bit depth of 8 bits, the pixel representing the sky has a value of 0, the pixel representing the bottom of the image has a value of 255, and the pixel value increases linearly from 0 to 255 from the boundary line to the bottom of the image.

[0042] Mid-range images are obtained using the vanishing line gradient method, where the vanishing point is considered the farthest point. The vanishing line emanating from the vanishing point divides the image into horizontal and vertical planes, and the depth maps for these two planes need to be calculated separately. Therefore, the vanishing line gradient method first distinguishes between the vertical and horizontal directions, and then obtains the depth image using the following formula: in, D map_h and D map_v These represent the depth gradients on the horizontal and vertical planes, respectively. x vp , y vp () represents the coordinates of the vanishing point in a 2D photograph. M and N These represent the number of pixels in the horizontal and vertical directions of the depth map, respectively. In the horizontal direction, the pixel value of the depth map increases linearly from 0 to 255 along the column direction from the vanishing point to the image edge; in the vertical direction, the pixel value of the depth map increases linearly from 0 to 255 along the row direction from the vanishing point to the image edge.

[0043] The depth map of the close-up image is obtained using a weighted depth overlay method. This method assumes that closer image regions contain more detail, while farther image regions contain less detail. By counting the number of edge contours, the distance of different image regions relative to the photographer can be approximately determined. The method first uses the Canny algorithm to extract image contours and divides them into 5×5 rectangular regions. Then, it counts the number of edge contours in each rectangular region, defining the region with a higher number of edge contours than the average as the principal region. The sub-depth map of each principal region is elliptical in shape, with the pixel value at the center of the ellipse set to 255 and the pixel value on the circumference set to 0. The pixel value decreases linearly from the center to the circumference. All sub-depth maps are then weighted and overlaid using the following formula to obtain the global depth map: , in, D i Indicates the first i Sub-depth maps of the main regions D f This represents the depth map of the entire close-up image. A Indicates the number of main areas. N i 'Indicates the number of edge contours in each main region.

[0044] Of course, the three-dimensional coordinates of the driver's eyes can also be determined using the monocular depth estimation algorithm based on mixed features described in the above embodiments. Specifically, the second information determination module can classify the driver's facial image according to the mixed features. For example, for the driver's eye image classified as a close-up image, a weighted depth overlay method is used to estimate the depth of the driver's eye image to obtain the three-dimensional coordinates of the driver's eyes.

[0045] In a preferred embodiment, the second information determination module is used to input the driver's facial image into a pre-trained depth gradient-assisted monocular depth estimation network model to obtain the three-dimensional coordinates of the driver's eyes output by the monocular depth estimation network model; wherein, the monocular depth estimation network model is trained by depth estimation training samples and depth gradient-assisted samples.

[0046] In this embodiment, the Depth Gradient-Assisted Monocular Depth Estimation Network (DGE-CNN) achieves more generalizable monocular depth estimation. The DGE-CNN model consists of two network modules: an encoder (first encoder) and a decoder (first decoder). The encoder detects and acquires depth features from the 2D color image, while the decoder generates a depth map corresponding to the 2D image based on the acquired depth features. The encoder is a modified ResNet-50 network module. The 2D color image is input into the DGE-CNN model, first passing through a downsampling function block in the encoder, and then entering a network structure composed of several residual downsampling function blocks and residual projection function blocks connected in series. Both the residual downsampling function blocks and the residual projection function blocks consist of operations such as convolution, batch normalization (BN) operations, and linear rectified function (ReLU).

[0047] The feature information processed by the encoder then enters the decoder of the network model. The decoder is an upsampling network module, mainly composed of an upconvolutional residual function block, convolution, batch normalization (BN) operations, ReLU, and the Dropout 2D algorithm. The upconvolutional residual function block improves the network's predictive ability and generalization performance; the Dropout 2D algorithm alleviates overfitting during training, achieving a certain degree of regularization. After encoder processing, the output of the DGE-CNN model is a depth map of a two-dimensional color image.

[0048] The completed DGE-CNN model needs to be trained to obtain high-precision depth maps. This invention uses input-output images as the training set to train the network. This invention also designs a DGE module to extract depth gradients and incorporates these depth gradients as an auxiliary training set into the CNN model training. The DGE module utilizes two independent Sobel convolution operators. G x and G y This allows for the separate extraction of depth gradients in the horizontal and vertical directions. The depth gradient information used in the auxiliary training set is obtained by fusing the depth gradients in these two directions.

[0049] Of course, the method for determining the target scene information can also adopt the monocular depth estimation network model based on depth gradient assistance in the above embodiments. Specifically, the first information determination module can input the road scene image into the pre-trained monocular depth estimation network model based on depth gradient assistance to obtain the target scene information output by the monocular depth estimation network model.

[0050] In a preferred embodiment, the hologram generation module is used to: perform image feature recognition based on intensity distribution to determine the type characteristics of the target scene 41; generate textual or graphic auxiliary information 42 containing type identifiers based on the type characteristics, and generate numerical auxiliary information 42 containing distance parameters based on the depth component in the three-dimensional coordinates of the target scene 41 and the three-dimensional coordinates of the driver's eyes; input the textual or graphic auxiliary information 42, numerical auxiliary information 42, the three-dimensional coordinates of the target scene 41, the three-dimensional coordinates of the driver's eyes, and the parameters of the optical components into a pre-trained model without convolution error to drive an autoencoder deep learning network, thereby obtaining an auxiliary information hologram output by the autoencoder deep learning network; wherein, the autoencoder deep learning network compensates for the encoding phase convolution error by extending the phase during the decoding stage.

[0051] To simultaneously ensure the accuracy and speed of the holographic algorithm, this embodiment uses a model without convolutional error to drive an autoencoder deep learning network, achieving a significant improvement in the high-precision hologram generation time. The target image to be displayed is the network input. After passing through the encoder (second encoder), it obtains the output phase, and then through the decoder (second decoder), it obtains the output amplitude. A loss function is used to establish the relationship between the input image and the output amplitude, and the network parameters are trained based on this. After the autoencoder deep learning network is trained, the encoder can be extracted separately as a tool for calculating pure phase holograms. The encoder can be a U-Net, where the input is the target image to be displayed, and the output is the pure phase hologram corresponding to the target image. The complete U-Net consists of a downsampling path and a corresponding upsampling path. The downsampling path increases the level of feature abstraction and consists of six downsampling modules and six corresponding residual modules. Each downsampling module consists of two sets of BN operations, a ReLU function, and a 3×3 convolutional layer. The upsampling path extracts high-level image features and restores the output to the same size as the input image. It consists of six upsampling modules. The upsampling modules are designed based on the subpixel convolution method, and their structure can be further decomposed into convolutional layers and transformation layers. The convolutional layers increase the number of channels, while the transformation layers restructure the data. The residual layers in the downsampling modules are skipped to the corresponding upsampling modules to alleviate degradation problems during network training. After the upsampling path, the output data first passes through a tanh function layer to constrain the phase value of the output. Within the range.

[0052] The decoder's input information is the encoder's output phase, and its initial resolution is... N x × N y To avoid convolution errors, this phase distribution needs to be padded with zeros to achieve a resolution of 2.N x ×2 N y The extended phase is used as the phase distribution on the holographic plane, while the intensity matrix, with all elements equal to 1, is used as the amplitude distribution on the holographic plane. The amplitude and phase together constitute the complex amplitude on the holographic plane, which, after being decoded, results in the complex amplitude on the object plane. The decoder is the forward propagation process of the non-convolutional angular spectrum model: , in, E O ( x , y , z 0) represents the amplitude distribution of the target image to be displayed. z 0 represents the distance between the object plane and the holographic plane. This represents the initial phase superimposed on the target image. Indicates the recording wavelength. This indicates phase fine-tuning based on the parameters of the optical components within the fusion module; FT stands for Fourier Transform. u and v These represent the spatial frequencies in the horizontal and vertical directions, respectively. Typically, the reference plane in the angular spectrum model is the holographic plane, and its depth coordinates... z =0. Extract the amplitude distribution on the extract plane and clip its resolution to... N x × N y The final output of the network can then be obtained. The amplitude of this output can be correlated with the input image using a loss function. Based on the calculation results of the loss function, parameter training of an autoencoder deep learning network without convolutional error can be achieved.

[0053] The training process of a deep learning network first requires calculating the loss function between the input and output images, and then updating the relevant parameters of the convolutional kernel based on the calculation results. The specific loss function used in this invention is the negative Pearson correlation coefficient, which can increase the convergence probability of the network parameters. Its mathematical form can be expressed as follows: , in, I and I r These represent the intensity of the original image and the intensity of the reconstructed image, respectively, and their specific values ​​can be obtained from the amplitude distribution on the object plane.

[0054] In a preferred embodiment, the fusion module 3 includes: a light source module for generating collimated polarized illumination light; an optical modulation module disposed at the output end of the light source module, the input end of the optical modulation module being electrically connected to the output end of the hologram generation module, for loading an auxiliary information hologram, which diffracts after being illuminated by polarized illumination light, projecting the auxiliary information hologram onto a preset position of a nearby target object 41; a display module disposed at the diffracted light wave output end of the optical modulation module, for reflecting the diffracted light wave to the driver's eye area to perform virtual-real fusion of the auxiliary information hologram and the road scene; and parameters of the optical components of the light source module, optical modulation module, and display module for fine-tuning the calculation parameters of the auxiliary information hologram.

[0055] In a preferred embodiment, the light source module includes: a laser light source for emitting diverging spherical waves; a collimating lens 34 disposed at the output end of the laser light source for converging the diverging spherical waves and outputting collimated plane waves; and a polarizer 35 disposed at the collimated plane wave output end of the collimating lens 34, the output end of which serves as the output end of the light source module for changing the polarization state of the collimated plane wave and outputting collimated polarized illumination light.

[0056] In a preferred embodiment, the optical modulation module is a spatial light modulator 36, which is used to load a phase-type hologram corresponding to the auxiliary information 42. The phase-type hologram is used to change the surface phase distribution of the spatial light modulator 36.

[0057] In this embodiment, the acquisition module 1 is a camera, which can be a commonly used monocular camera, used to acquire two-dimensional images of the real road scene during driving. The first camera 11 is connected to the signal processing and control board 21 via a data cable (not shown in the figure). The acquired two-dimensional images will be used as input to the monocular depth estimation algorithm to subsequently determine the distance of each real object in the real road scene relative to the car.

[0058] Target object 41 is a specific real-world object in a real road scene during driving. In reality, there may be several target objects of various types in a real road scene; the types and quantities of target objects in the illustration are for illustrative purposes only.

[0059] The signal processing and control board 21 is mainly responsible for tasks such as system power supply, signal acquisition, monocular depth estimation, and hologram generation.

[0060] The signal processing and control board 21 is connected to the vehicle's power supply unit (not shown in the figure) via electrical cables and is used to power the first camera 11, the second camera 12, the laser emitter 31 and the spatial light modulator 36.

[0061] The signal processing and control board 21 is connected to the first camera 11 via a data cable (not shown in the figure). After the first camera 11 acquires a two-dimensional image of the real road scene, it transmits the image back to the signal processing and control board 21 via the data cable. The signal processing and control board 21 is connected to the second camera 12 via a data cable (not shown in the figure). After the second camera 12 acquires an image of the driver's face, it transmits the image back to the signal processing and control board 21 via the data cable. The signal processing and control board 21 includes a first information determination module, a second information determination module, and a hologram generation module.

[0062] The first information determination module incorporates a monocular depth estimation algorithm. After the road scene image acquired by the first camera 11 is processed by the monocular depth estimation algorithm, the various target objects contained in the image are segmented, and the three-dimensional coordinates of each target object relative to the vehicle are obtained. The segmented target objects and their corresponding three-dimensional coordinates will be used for subsequent hologram generation.

[0063] The second information determination module incorporates a monocular depth estimation algorithm. After the driver's facial image captured by the second camera 12 is processed by the monocular depth estimation algorithm, the three-dimensional coordinates of the driver's eyes are obtained. This three-dimensional coordinate data will also be used for subsequent hologram generation.

[0064] The hologram generation module incorporates a hologram generation algorithm. This algorithm first requires three pieces of data: the intensity distribution of each target object 41 after image segmentation, the three-dimensional coordinates of each target object 41, and the three-dimensional coordinates of the driver's eyes. After acquiring the intensity distribution and three-dimensional coordinates of the target objects 41, the algorithm generates corresponding auxiliary information 42. For example, if the algorithm identifies the target object 41 as a tree, it generates the text auxiliary information "Tree"; if it identifies the height and depth values ​​of the target object, it generates the text auxiliary information "Height:1m,Distance:1m". Considering that the projection position of the auxiliary information 42 is near the corresponding target object 41 and the projection angle should be within the driver's eye box (driver's eye area), after acquiring the content of the auxiliary information 42, the hologram generation algorithm uses the three-dimensional coordinates of the target object 41 and the three-dimensional coordinates of the driver's eyes as calculation parameters to generate a phase-type hologram corresponding to the auxiliary information 42.

[0065] The signal processing and control board 21 is connected to the laser emitter 31 via a data cable with power supply capability. Under the control of the signal processing and control board 21, the laser emitter 31 emits red, green, and blue illumination lasers as needed.

[0066] The signal processing and control board 21 is connected to the spatial light modulator 36 via a data cable with power supply capability. Under the control of the signal processing and control board 21, the spatial light modulator 36 loads the phase-type holograms corresponding to each color component as required.

[0067] The fusion module 3 includes a light source module, an optical modulation module, and a display module. The light source module includes a laser light source, a collimating lens 34, and a polarizer 35.

[0068] The laser emitter 31 (laser source) is connected to the signal processing and control board 21 via a data cable. It can emit illumination lasers with red, green and blue color components as needed, which are used to illuminate the phase-type hologram loaded on the spatial light modulator 36 and project corresponding auxiliary information 42 near the target scene 41 through diffraction.

[0069] The function of optical fiber 32 is to connect laser emitter 31 and laser exit aperture 33. After the coherent beam emitted by laser emitter 31 is coupled into optical fiber 32, it enters laser exit aperture 33 under the guidance of optical fiber 32.

[0070] The laser exit aperture 33 is a tiny aperture located at the center of an opaque metal sheet. Its function is to filter out stray light and obtain a more ideal point light source. The coherent beam emitted by the laser emitter 31 is guided into the laser exit aperture 33 through the optical fiber 32, and the emitted light can be regarded as a point light source with a very small geometric size.

[0071] The function of the collimating lens 34 is to converge diverging spherical waves to obtain collimated plane waves. After passing through the laser exit aperture 33, the coherent beam becomes an ideal point source, with a corresponding wavefront shape of a diverging spherical wave. When the laser exit aperture 33 is placed on the object-side focal plane of the collimating lens 34, the distance from the point source to the collimating lens 34 is exactly the focal length of the collimating lens 34. At this point, the diverging spherical wave is collimated into a plane wave.

[0072] The function of polarizer 35 is to change the polarization state of the collimated plane wave. The spatial light modulator 36 used in this invention typically has polarization selectivity. Rotating polarizer 35 and changing its fast axis angle can alter the polarization state of the collimated plane wave. After passing through polarizer 35, the plane wave emitted by collimating lens 34 will have a specific polarization state.

[0073] The optical modulation module is a spatial light modulator 36, which is used to load holograms. The spatial light modulator 36 is connected to the signal processing and control board 21 via a data cable with power supply. After generating the phase-type hologram corresponding to the auxiliary information 42, the signal processing and control board 21 loads the corresponding phase-type hologram onto the surface of the spatial light modulator 36 via the data cable. The loaded phase-type hologram changes the surface phase distribution of the spatial light modulator 36. When polarized parallel light emitted from the polarizer 35 illuminates the surface of the spatial light modulator 36 with its specific phase distribution, diffraction occurs, projecting the image of the auxiliary information 42 onto the vicinity of the corresponding target object 41, and ensuring that the projected light, after being reflected by the windshield 37, reaches the driver's field of vision.

[0074] The function of the windshield 37 is to reflect the diffracted light emitted by the spatial light modulator 36, project the image of the auxiliary information 42 onto the vicinity of the corresponding target scene 41, and ensure that the projected light can reach the driver's field of vision after reflection.

[0075] Auxiliary information 42 is crucial for realizing the virtual-real fusion augmented reality head-up display. After obtaining the intensity distribution and three-dimensional coordinates of the target scene 41 through the monocular depth estimation algorithm built into the signal processing and control board 21, the hologram generation algorithm will generate the corresponding auxiliary information 42. For example, if the algorithm identifies the type of the target scene 41 as "tree", it will generate the text auxiliary information "Tree"; if it identifies the height and depth values ​​of the target scene, it will generate the text auxiliary information "Height:1m,Distance:1m". After obtaining the content of auxiliary information 42, the hologram generation algorithm will use the three-dimensional coordinates of the target scene 41 and the three-dimensional coordinates of the driver's eyes as calculation parameters to generate a phase-type hologram corresponding to the auxiliary information 42. After the phase-type hologram is loaded onto the surface of the spatial light modulator 36 through the data cable, the loaded phase-type hologram will change the surface phase distribution of the spatial light modulator 36 and diffract under the illumination of polarized parallel light, projecting the image of the auxiliary information 42 onto the vicinity of the corresponding target scene 41. During driving, the driver can simultaneously see the real target scenery 41 and the projected auxiliary information 42, achieving a fusion of virtual and real. Meanwhile, the auxiliary information 42 supplements the target scenery 41, playing a role in augmenting reality.

[0076] The instrument panel 43 serves to fix and protect the relevant components. The projection optical path, consisting of the laser emitter 31, optical fiber 32, laser output aperture 33, collimating lens 34, and polarizer 35, is placed inside the instrument panel 43. The signal processing and control board 21 is also placed inside the instrument panel 43.

[0077] The augmented reality head-up display method provided by the present invention is described below. The augmented reality head-up display method described below can be referred to in correspondence with the augmented reality head-up display device described above.

[0078] Please refer to Figure 3 , Figure 3 This is a flowchart illustrating an augmented reality head-up display method provided by the present invention.

[0079] The present invention also provides an augmented reality head-up display method, comprising: 301: Acquire road scene images, driver facial images, and parameters of optical components; 302: Based on the road scene image and the driver's facial image, determine the target scene information and the driver's eye position information, and generate an auxiliary information hologram corresponding to the target scene 41 based on the parameters of the optical components, the target scene information and the driver's eye position information; 303: By using optical components, auxiliary information holograms are integrated into the road scene to create an augmented reality head-up display.

[0080] In a preferred embodiment, target scene information and driver's eye position information are determined based on a road scene image and a driver's facial image. An auxiliary information hologram corresponding to the target scene 41 is generated based on the parameters of the optical components, the target scene information, and the driver's eye position information. This includes: segmenting the target scene 41 from the road scene image using a monocular depth estimation algorithm to determine the target scene information; the target scene information includes the intensity distribution and three-dimensional coordinates of the target scene 41; obtaining the three-dimensional coordinates of the driver's eyes from the driver's facial image using a monocular depth estimation algorithm; generating auxiliary information 42 describing the target scene 41 based on the intensity distribution, three-dimensional coordinates, and three-dimensional coordinates of the driver's eyes using a hologram generation algorithm; and generating an auxiliary information hologram based on the auxiliary information 42, the three-dimensional coordinates of the target scene 41, the three-dimensional coordinates of the driver's eyes, and the parameters of the optical components.

[0081] Please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating the specific process of an augmented reality head-up display method provided by the present invention.

[0082] The complete workflow of this invention is as follows: (1): The first camera 11 captures two-dimensional images of the real road scene during driving and transmits them to the signal processing and control board 21 via a data cable; at the same time, the second camera 12 captures two-dimensional images of the driver's face and transmits them to the signal processing and control board 21 via a data cable.

[0083] (2): The two-dimensional images of the real road scene and the driver's face are processed by the monocular depth estimation algorithm built into the signal processing and control board 21. After processing, the two-dimensional images of the real road scene are segmented, and the three-dimensional coordinates of each target object relative to the car are obtained (only one example of the target object is shown in the schematic diagram); after processing, the three-dimensional coordinates of the driver's eyes are obtained.

[0084] (3): Using the intensity distribution and three-dimensional coordinates of the target scene 41 as input, the hologram generation algorithm built into the signal processing and control board 21 generates auxiliary information 42 corresponding to the target scene 41. For example, if the algorithm identifies the type of the target scene 41 as "tree", it generates the text auxiliary information "Tree"; if it identifies the height and depth values ​​of the target scene, it generates the text auxiliary information "Height:1m,Distance:1m".

[0085] (4): The phase-type hologram corresponding to the auxiliary information 42 is generated by the hologram generation algorithm built into the signal processing and control board 21. Considering that the projection position of the auxiliary information 42 should be near the corresponding target object 41, and the angle of the projection light should be within the driver's eye box, the hologram generation algorithm will use the three-dimensional coordinates of the target object 41, the three-dimensional coordinates of the driver's eyes, and the parameters of the optical components as calculation parameters to generate the phase-type hologram corresponding to the auxiliary information 42.

[0086] (5): Achieving optical projection of auxiliary information 42. The phase-type hologram generated by the signal processing and control board 21 is loaded onto the surface of the spatial light modulator 36 via a data cable. At the same time, the signal processing and control board 21 is connected to the laser emitter 31 via a cable to synchronize the timing of the illumination light and the timing of the phase-type hologram. The laser emitted by the laser emitter 31 passes through the optical fiber 32, the laser exit aperture 33, the collimating lens 34, and the polarizer 35 to obtain polarized plane illumination light, which diffracts after illuminating the phase-type hologram loaded on the spatial light modulator 36. The diffracted light is reflected by the windshield 37 and projected onto the vicinity of the corresponding target scene 41, and the projected light can reach the driver's eye box range after reflection.

[0087] (6): When driving, the driver looks ahead and can see both the real target scenery 41 and the projected auxiliary information 42, thus achieving virtual-real fusion. At the same time, the auxiliary information 42 supplements the target scenery 41 and plays a role in augmenting reality.

[0088] The present invention has the following beneficial effects: The device has a simple structure. In this invention, wavefront error can be compensated by optimizing the phase hologram without using composite holographic elements and without involving joint hardware and software optimization; the determination of projection distance and angle can be achieved by using a monocular camera and a monocular depth estimation algorithm without using a three-dimensional measurement device; the projection distance can be adjusted by changing the calculation parameters of the phase hologram without using zoom elements or superlenses; except for the collimating lens, it can be designed entirely based on a lensless scheme, and the distance of each projected image can be configured according to the position of the real scene.

[0089] A unique virtual-real fusion solution. Based on computational holography, this invention enables the levitation projection of auxiliary information of any pattern at any spatial location, and has the advantages of continuous depth and clear image, providing an ideal virtual-real fusion augmented reality head-up display experience.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An augmented reality head-up display device, characterized in that, include: The acquisition module is used to acquire road scene images, driver facial images, and parameters of optical components; The processing module is used to determine target scene information and driver eye position information based on the road scene image and the driver's facial image, and generate an auxiliary information hologram corresponding to the target scene based on the parameters of the optical components, the target scene information and the driver's eye position information. The fusion module is used to fuse the auxiliary information hologram into the road scene through the optical components for augmented reality head-up display.

2. The augmented reality head-up display device according to claim 1, characterized in that, The processing module includes: The first information determination module is used to segment target objects from the road scene image based on a monocular depth estimation algorithm and determine the target object information; the target object information includes the intensity distribution of the target object and the three-dimensional coordinates of the target object; The second information determination module is used to obtain the three-dimensional coordinates of the driver's eyes from the driver's facial image based on a monocular depth estimation algorithm; The hologram generation module is used to generate auxiliary information describing the target scene based on the intensity distribution of the target scene, the three-dimensional coordinates of the target scene and the three-dimensional coordinates of the driver's eyes, according to the hologram generation algorithm, and to generate the auxiliary information hologram based on the auxiliary information, the three-dimensional coordinates of the target scene, the three-dimensional coordinates of the driver's eyes and the parameters of the optical components.

3. The augmented reality head-up display device according to claim 2, characterized in that, The first information determination module is used to classify the road scene image based on mixed features; the mixed features include color, edge, contour, vanishing line and vanishing point; the image categories include distant image, mid-range image and close-up image; The depth of the distant image is estimated using the horizontal edge gradient method, the depth of the mid-range image is estimated using the vanishing line gradient method, and the depth of the near-range image is estimated using the weighted depth stacking method, thereby obtaining the target scene information.

4. The augmented reality head-up display device according to claim 2, characterized in that, The second information determination module is used to input the driver's facial image into a pre-trained depth gradient-assisted monocular depth estimation network model to obtain the three-dimensional coordinates of the driver's eyes output by the monocular depth estimation network model; The monocular depth estimation network model is obtained by training depth estimation training samples and depth gradient auxiliary samples.

5. The augmented reality head-up display device according to any one of claims 2 to 4, characterized in that, The hologram generation module is used for: Image feature recognition is performed based on the intensity distribution to determine the type characteristics of the target scene; Based on the type characteristics, textual or graphic auxiliary information containing type identifiers is generated, and numerical auxiliary information containing distance parameters is generated based on the three-dimensional coordinates of the target scene and the depth component in the three-dimensional coordinates of the driver's eyes. Textual or graphic auxiliary information, numerical auxiliary information, the three-dimensional coordinates of the target scene, the three-dimensional coordinates of the driver's eyes, and the parameters of the optical components are input into a pre-trained model without convolutional error to drive an autoencoder deep learning network, thereby obtaining the auxiliary information hologram output by the autoencoder deep learning network. The autoencoder deep learning network compensates for the coding phase convolution error during the decoding stage by extending the phase.

6. The augmented reality head-up display device according to claim 5, characterized in that, The fusion module includes: A light source module is used to generate collimated polarized illumination light; An optical modulation module is disposed at the output end of the light source module. The input end of the optical modulation module is electrically connected to the output end of the hologram generation module. It is used to load the auxiliary information hologram. After the polarized illumination light illuminates the auxiliary information hologram, diffraction occurs, and the auxiliary information hologram is projected onto a preset position of a nearby target scene. The display module is located at the diffraction light wave output end of the optical modulation module and is used to reflect the diffraction light wave to the driver's eye area to fuse the auxiliary information hologram and the road scene in a virtual and real manner. The parameters of the optical components in the light source module, the optical modulation module, and the display module are used to fine-tune the calculation parameters of the auxiliary information hologram.

7. The augmented reality head-up display device according to claim 6, characterized in that, The light source module includes: Laser light source, used to emit diverging spherical waves; A collimating lens is disposed at the output end of the laser source to converge the diverging spherical wave and output a collimated plane wave. A polarizer is disposed at the collimated plane wave output end of the collimating lens. The output end of the polarizer serves as the output end of the light source module, used to change the polarization state of the collimated plane wave and output the collimated polarized illumination light.

8. The augmented reality head-up display device according to claim 6, characterized in that, The optical modulation module is a spatial light modulator used to load a phase-type hologram corresponding to auxiliary information. The phase-type hologram is used to change the surface phase distribution of the spatial light modulator.

9. An augmented reality head-up display method, characterized in that, include: Acquire road scene images, driver facial images, and parameters of optical components; Based on the road scene image and the driver's facial image, target scene information and driver's eye position information are determined, and an auxiliary information hologram corresponding to the target scene is generated based on the parameters of the optical components, the target scene information, and the driver's eye position information. The auxiliary information hologram is fused into the road scene using the optical components to achieve augmented reality head-up display.

10. The augmented reality head-up display method according to claim 9, characterized in that, The step of determining target scene information and driver eye position information based on the road scene image and the driver's facial image, and generating an auxiliary information hologram corresponding to the target scene based on the parameters of the optical components, the target scene information, and the driver's eye position information, includes: The target scene is segmented from the road scene image based on a monocular depth estimation algorithm to determine the target scene information; the target scene information includes the intensity distribution of the target scene and the three-dimensional coordinates of the target scene. The three-dimensional coordinates of the driver's eyes are obtained from the driver's facial image based on a monocular depth estimation algorithm; Based on the hologram generation algorithm, auxiliary information describing the target scene is generated according to the intensity distribution of the target scene, the three-dimensional coordinates of the target scene and the three-dimensional coordinates of the driver's eyes. Then, the auxiliary information hologram is generated according to the auxiliary information, the three-dimensional coordinates of the target scene, the three-dimensional coordinates of the driver's eyes and the parameters of the optical components.