An optimization method and system for augmented reality display

CN121707839BActive Publication Date: 2026-08-21SCHOOL OF FOREIGN AFFAIRS CENT SOUTH FORESTRY UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511895637.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-08-21
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

一种是非光学方案:利用体积或全息显示器在物理空间直接呈现三维信息,但因硬件复杂度难以实用化;另一种是光学方案:通过微透镜阵列、空间光调制器、变焦显示器或多焦点显示器提供变焦能力,但仍存在系统昂贵、光路复杂的问题

Benefits of technology

本发明针对视觉辐辏调节冲突问题,设计了一种用于增强现实显示的优化方法及系统,通过将真实场景和虚拟场景融合后的光场信息进行采集,然后设计卷积神经网络对融合光场信息进行处理生成多焦平面虚拟场景。将生成的多焦平面虚拟场景与理想合成数据进行对比,通过对卷积神经网络进行训练,得到优化后的微透镜几何参数与卷积神经网络权重。借助优化后的微透镜几何参数与卷积神经网络权重,预生成不同焦平面的增强现实场景图像库。最后,基于眼动追踪实时捕捉人眼焦距,动态投射匹配焦平面的场景,实现虚拟信息与真实场景的生理自适应融合。在解决视觉辐辏调节冲突问题的同时,系统响应延迟低,设备成本降低,并可支持连续变焦显示。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121707839B_ABST
    Figure CN121707839B_ABST
Patent Text Reader

Abstract

The application discloses an optimization method and system for augmented reality display, and the method comprises the following steps: collecting light field information after fusing a real scene and a virtual scene, and then designing a convolutional neural network to process the fused light field information to generate a multi-focal plane virtual scene; comparing the generated multi-focal plane virtual scene with ideal synthetic data through a convolutional neural network training model, optimizing the micro-lens geometric parameters and the convolutional neural training model, and finally generating an augmented reality scene image library of different focal planes; and finally, based on the focal length of the human eye, selecting a scene with a matching focal plane from the image library for dynamic projection, realizing physiological adaptive fusion of virtual information and a real scene, and solving the problem of visual vergence accommodation conflict.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of display technology, and more specifically to an optimization method and system for augmented reality displays. Background Technology

[0002] Artificial intelligence, as a core driving force for the development of the global information industry, is propelling augmented reality (AR) technology into deep penetration into high-precision fields. AR technology overlays computer-generated virtual information onto the real environment in real time, enabling intraoperative anatomical navigation in the medical field, micrometer-level assembly guidance in industrial settings, and enhancing three-dimensional battlefield situational awareness in military applications. The core value of AR systems lies in the synergy between optical display and rendering algorithms to achieve visual perception fusion between physical scenes and virtual content. With the explosive expansion of application scenarios and the continuous upgrading of user demands for immersive experiences, unprecedentedly stringent requirements are being placed on the visual realism of systems.

[0003] However, current mainstream augmented reality solutions face a fundamental physiological bottleneck—the convergence-accommodation conflict. This conflict stems from the inherent coupling of the human visual neural mechanism: convergence drives the dynamic matching of the visual axis angle between the eyes to the target distance, while accommodation achieves precise focusing through changes in lens refractive power. When a virtual image is fixed to a single optical focal plane, these physiological processes are forced to decouple, causing the lens accommodation to lock at a fixed focal length. This neural signal conflict has been proven to cause visual fatigue and depth perception distortion.

[0004] To overcome the visual convergence conflict in augmented reality, researchers have proposed two technical approaches. One is a non-optical approach: using volumetric or holographic displays to directly present three-dimensional information in physical space, but this is difficult to implement due to hardware complexity. The other is an optical approach: providing zoom capabilities through microlens arrays, spatial light modulators, zoom displays, or multifocal displays, but these still suffer from problems such as high system cost and complex optical paths. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides an optimization method and system for augmented reality display, which resolves the visual convergence-accommodation conflict and achieves physiological adaptive fusion of virtual information and real-world scenes. This system features low response latency, reduced equipment costs, and supports continuous zoom display.

[0006] In a first aspect, the present invention provides an optimization method for augmented reality display, the technical solution of which includes the following steps: 1. Establish an optical transmission model for an augmented reality display system based on a microlens array to obtain fused light field data of the real and virtual scenes; 2. Design a convolutional neural network to optimize the fused light field data after fusing real and virtual scenes, and generate augmented reality scene images with different focal planes; 3. Optimize the geometric parameters of the microlens array and the training model of the convolutional neural network by comparing and training augmented reality scene images with ideal synthetic data at different focal planes; 4. Based on the optimized microlens array parameters, update the parameters of the microlens array built in step 1, and deploy the trained convolutional neural network to the hardware control system so that the training model of the convolutional neural network in the hardware control system matches the updated microlens array parameters. 5. Based on the eye-tracking device, the focal length of the human eye is captured in real time, and the captured focal length is input into the hardware control system. The hardware control system controls the training model of the convolutional neural network according to the captured focal length, and dynamically outputs the augmented reality scene corresponding to the focal plane to the human eye.

[0007] Furthermore, in step 1, the augmented reality display system based on the microlens array includes a microlens array, a CCD camera, a semi-transparent mirror, and a display. The CCD camera includes a lens and a CCD sensor. The computer-generated object is projected onto the real scene through the display. The real light field generated in the real environment and the light field data of the virtual object projected by the display are superimposed by the semi-transparent mirror to form a composite light field. After the composite light field passes through the lens and the microlens array, the fused light field formed by phase modulation is captured by the CCD sensor.

[0008] Furthermore, in step 1, the microlens array employs periodically arranged aspherical units.

[0009] Furthermore, in step 2, the convolutional neural network includes a super-resolution module and a zoom rendering module. The super-resolution module is based on the LF-InterNet network structure and reconstructs the input fused light field data to improve its spatial resolution and obtain a high-resolution multi-view image. The zoom rendering module can re-render the light field array sub-image after super-resolution enhancement into augmented reality scene images focused on different specific planes according to the specified refocusing parameters.

[0010] Furthermore, in step 2, the zoom rendering module calculates the refocusing integral formula for the image focused on a specific depth plane according to the following formula: in, For the rendered augmented reality scene image, To render the spatial coordinates of the image plane, These are the spatial coordinates of the microlens plane, used to represent the direction information of light rays; Recorded the fused light field passing through the aperture point Arrival at CCD sensor point The intensity of light; spatial coordinates pass The relationship is determined by the rendering coordinates With microlens coordinates Sure; It is the refocusing factor, determined by the focal plane. ,in The principal focal length of the CCD camera lens. The equivalent focal length corresponding to the target focusing plane. This indicates the angle between the ray and the normal to the CCD sensor. It is an attenuation factor used to compensate for the intensity attenuation caused by the angle of light.

[0011] Furthermore, in step 3, by comparing and training augmented reality scene images with ideal synthetic data at different focal planes, gradients are calculated based on the differentiable optical transmission model to achieve joint optimization of the geometric parameters of the microlens array and the weights of the neural network. The geometric parameters of the microlens array include: the arrangement period, the aperture size of each microlens, and the focal length of the microlens.

[0012] Furthermore, in step 3, the implementation method for comparative training and optimization includes the following steps: 3.1 Construct the optical transmission model as a differentiable computational graph, so that the fused optical field data input to the convolutional neural network is discretized and differentiable, and its gradient can be calculated.

[0013] 3.2 The convolutional neural network, based on the current weights and microlens parameters, fuses the four-dimensional light field data input to the convolutional neural network. A series of augmented reality scene images with different focal planes are predicted.

[0014] 3.3 During the training process, the predicted augmented reality scene images at different focal planes are compared with the expected images, the differences between the two are calculated, and the gradient of the loss function is simultaneously backpropagated to the weights of the convolutional neural network and the microlens geometric parameters in the optical transmission model.

[0015] 3.4. Through repeated iterations, the weights of the convolutional neural network and the geometric parameters of the microlenses in the optical transmission model are optimized to reduce the difference between the predicted augmented reality scene images at different focal planes and the desired images.

[0016] 3.5 Once the difference between the predicted augmented reality scene images at different focal planes and the desired image reaches a preset value, the optimized microlens array parameters and the trained convolutional neural network training model are obtained.

[0017] Furthermore, in step 5, after the eye-tracking device captures the user's eye's adjustment of focus, it calculates the convergence point of the user's gaze through the pupil-corneal reflection vector and maps it to the depth of the target focal plane in the scene; based on the depth of the target focal plane, it retrieves the corresponding scene data from a pre-generated augmented reality scene image library with different focal planes and inputs it into the human eye.

[0018] Secondly, the present invention also provides an augmented reality perspective optical system for implementing the above-described method, comprising: The virtual image module is used to generate virtual image light field data; The virtual-real scene capture module is used to fuse virtual object light field data and real light field scene data, and capture them using a light field camera to finally obtain the fused light field data to be processed. The rendering module uses the geometric parameters of the microlens array and the convolutional neural network to perform zoom rendering on the fused light field data to be processed, obtain augmented reality scene images of different focal planes, and establish an augmented reality scene image library of different focal planes. The training module compares and trains augmented reality scene images with ideal synthetic data at different focal planes to obtain the optimal geometric parameters of the microlens array and the trained convolutional neural network. The optimized microlens array parameters obtained through the training module, along with the trained convolutional neural network model, will be updated in the virtual and real scene capture module and the training module.

[0019] Furthermore, the augmented reality perspective optical system also includes a display module and an eye-tracking device that can capture the focal length of the human eye. The display module can select suitable augmented reality scene images from a pre-generated augmented reality scene image library based on the focal length information of the human eye, and dynamically input the suitable augmented reality scene images into the human eye.

[0020] Furthermore, the virtual scene capture module includes a lens, a microlens array, a CCD camera, and a semi-transparent mirror. The virtual image module includes a display. The computer-generated virtual information is projected through the display to generate virtual object light field data. The real light field generated in the real environment and the virtual object light field data projected by the display are superimposed by the semi-transparent mirror to form a composite light field. The composite light field is captured by the CCD sensor after passing through the CCD camera lens and the microlens array. The semi-transparent mirror is at a 45-degree angle to the horizontal plane. The real object and the virtual object displayed on the display are located on the two light-splitting surfaces of the semi-transparent mirror. The light field camera, which consists of the lens, the microlens array, and the CCD sensor, should be located on the light-combining surface of the semi-transparent mirror.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention addresses the problem of visual convergence-accommodation conflict by designing an optimized method and system for augmented reality (AR) displays. The method involves collecting light field information from a fused real-world and virtual scene, then using a convolutional neural network (CNN) to process this fused information and generate a multi-focal-plane virtual scene. The generated virtual scene is compared with ideal synthetic data, and the CNN is trained to obtain optimized microlens geometry parameters and CNN weights. Using these optimized parameters, an AR scene image library with different focal planes is pre-generated. Finally, eye-tracking is used to capture the human eye's focal length in real time, dynamically projecting a scene matching the focal plane to achieve physiologically adaptive fusion of virtual information and the real-world scene. This solution addresses the visual convergence-accommodation conflict while exhibiting low system response latency, reduced equipment costs, and support for continuous zoom display. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. The elements or parts in the drawings are not necessarily drawn to scale. Obviously, the drawings described below are some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0023] Figure 1 This is a schematic diagram illustrating the application of an optimization method and the construction of a system based on the optimization method of the present invention.

[0024] Figure 2 This is a schematic diagram of the structure of a perspective optical system for optimizing augmented reality displays according to the present invention.

[0025] Figure 3 The zoom rendering module of this invention re-renders the fused light field into augmented reality scene images focused on different specific planes. Detailed Implementation

[0026] The embodiments of the technical solution of this application will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of this application, and are therefore merely examples and should not be used to limit the scope of protection of this application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and the foregoing description of the accompanying drawings are intended to cover non-exclusive inclusion.

[0027] In this document, suffixes such as "module," "part," or "unit" used to denote elements are used only for the purpose of illustrative purposes and have no specific meaning in themselves. Therefore, "module," "part," or "unit" may be used interchangeably.

[0028] Figure 1 This is a schematic diagram illustrating the application and system construction of an optimization method based on the optimization method of the present invention, which can be roughly divided into an initial design stage, a manufacturing stage, and an application stage. Through the explanation of this embodiment, the specific usage process of the optimization method of the present invention, the configuration modules required to build the optimization system of the present invention, and the role played by each configuration module can be understood.

[0029] I. Design Phase 1. Establish an optical transmission model for an augmented reality display system based on a microlens array to obtain fused light field data of the real scene and the virtual scene.

[0030] The augmented reality display system based on a microlens array includes a microlens array, a CCD camera, a semi-transparent mirror, and a display. The CCD camera includes a lens and a CCD sensor. The microlens array uses periodically arranged aspherical units. A schematic diagram of this perspective optical system is shown below. Figure 2 As shown in the diagram, the semi-transparent mirror forms a 45-degree angle with the horizontal plane. The real object and the virtual object displayed on the monitor are located on the two beam-splitting surfaces of the semi-transparent mirror. The light field camera, consisting of a lens, a microlens array, and a CCD sensor, should be located on the light-combining surface of the semi-transparent mirror.

[0031] Computer-generated virtual objects are projected onto a real scene through a monitor. The real light field generated in the real environment and the light field data of the virtual object projected on the monitor are superimposed by a semi-transparent and semi-reflective mirror to form a composite light field. The composite light field is captured by a CCD sensor after passing through a lens and a microlens array.

[0032] Based on scalar diffraction theory and the principle of light field propagation, an optical field transmission equation is constructed to quantify the illumination transmission process in real and virtual scenes. The model focuses on describing the phase modulation effect of the microlens array and derives the mathematical expression for the fused light field distribution.

[0033] Specifically, the optical transmission model can be expressed as follows: Realistic light field acquisition model: Light field of a real-world scene After passing through a microlens array, the image is captured by a CCD sensor. This indicates the spatial coordinates of the light ray on the CCD sensor. This represents the coordinates of the direction of light rays reaching the CCD sensor. Each microlens unit samples the incident light field, forming a sub-aperture image array.

[0034] Virtual Object Light Field Data Generation Model: Computer-generated virtual information is projected onto a display screen to produce virtual object light field data. The light field data of the virtual object is reflected by a semi-transparent mirror, forming a fused light path with the real light field.

[0035] Light field fusion and transmission model: real light field With virtual object light field data By superimposing semi-transparent and semi-reflective mirrors, a composite light field is formed. It can be represented as: in and Let be the transmittance and reflectance of the semi-transparent and semi-reflective mirror, respectively, and satisfy . .

[0036] Microlens array phase modulation model: As a key phase modulation element, the phase function of the microlens array can be modeled as follows: in The wavelength of light. The focal length of the microlens. Represents the spatial coordinates of light The phase modulation is used to synthesize the optical field. Modulated output light field Represented as: Considering the propagation of the system's optical path, the fused light field that ultimately reaches the human eye or imaging plane. It can be obtained through diffraction integrals, by merging the light field. The expression form is: in Let be the system's optical field transmission function, which characterizes the optical field propagation characteristics from the modulation plane to the observation plane.

[0037] 2. Design a convolutional neural network to optimize the fused light field data after fusing real and virtual scenes, and generate augmented reality scene images with different focal planes.

[0038] A convolutional neural network primarily designed for zoom functionality was constructed. The input to this neural network was the fused light field data output from the aforementioned optical transmission model, modulated by a microlens array. The output of the convolutional neural network is a series of high-resolution augmented reality scene images corresponding to different focal planes. The designed convolutional neural network includes a super-resolution module and a zoom rendering module.

[0039] The super-resolution module, based on the LF-InterNet network structure, reconstructs the input light field data, improves its spatial resolution, and obtains high-resolution multi-view images. The zoom rendering module can re-render the light field array sub-image enhanced by the super-resolution module into augmented reality scene images focused on different specific planes according to the specified refocusing parameters. Figure 3 The zoom rendering module of this invention re-renders the fused light field into augmented reality scene images focused on different specific planes.

[0040] The principle of zoom rendering is mainly based on the light field refocusing theory, which synthesizes images of different depths by integrating pixel offsets from different viewpoints, and renders an image focused on a specific depth plane from the light field data. In this example, the zoom rendering module calculates the image focused on a specific depth plane according to the following refocusing integral formula: in, For the rendered augmented reality scene image, To render the spatial coordinates of the image plane, These are the spatial coordinates of the microlens plane, used to represent the direction information of light rays; It is the four-dimensional fused light field function input to the convolutional neural network, recording the fused light field passing through the aperture point. Arrival at CCD sensor point The intensity of light. In the refocusing integral formula, The first two parameters (i.e., spatial coordinates) The rendering coordinates are transformed by the following methods. With microlens coordinates Sure: Based on the above transformation equations, the original CCD sensor coordinates... A set mapping relationship is formed between the composite coordinates in the refocusing integral formula and the composite coordinates. It is the refocusing factor, determined by the focal plane. ,in The principal focal length of the CCD camera lens. The equivalent focal length corresponding to the target focusing plane. This indicates the angle between the ray and the normal to the CCD sensor. It is an attenuation factor used to compensate for the intensity attenuation caused by the angle of light.

[0041] 3. Optimize the geometric parameters of the microlens array and the neural network training model by comparing and training augmented reality scene images with ideal synthetic data at different focal planes.

[0042] The aforementioned refocusing integral formula is the core algorithm for achieving zoom functionality, and its input is four-dimensional fused light field data. The quality and characteristics of the light are directly determined by the geometric parameters of the microlens array. These geometric parameters, by influencing the sampling method of the light field, establish a fundamental link with the refocusing process, specifically manifested as follows: Alignment period: This determines the sampling reference of the light field in both the spatial and angular domains. The matching relationship between the alignment period of the microlens array and the pixel size of the main sensor directly determines the number of pixels covered by each microlens, thus defining the upper limit of the angular resolution of the light field. An inappropriate alignment period can lead to insufficient angular information (aliasing) or wasted spatial information.

[0043] Aperture size: The aperture of each microlens is equivalent to a sub-aperture. Aperture size affects the angular ambiguity and luminous flux of the light field. A larger aperture can collect more light but reduces angular resolution; a smaller aperture can provide more accurate angular information but may lead to a decrease in image signal-to-noise ratio.

[0044] Microlens focal length (radius of curvature): This, along with the distance from the image plane of the primary lens to the microlens array, determines the sharpness of the microlens image. Ideal matching ensures that each microlens forms a sharp sub-image on the sensor, thereby ensuring the accuracy of the light direction information extracted from that sub-image.

[0045] Therefore, optimizing these microlens geometric parameters during neural network training is essentially optimizing the light field data acquisition process. By adjusting these parameters through backpropagation, the system learns to construct an optimal "front-end optical encoder," enabling the acquired light field data, after being processed by a fixed refocusing algorithm, to generate the highest quality output image that is most beneficial for the neural network to complete the zoom task.

[0046] By comparing and training augmented reality scene images with ideal synthetic data at different focal planes, gradients are calculated based on a differentiable optical transmission model to achieve joint optimization of the geometric parameters of the microlens array and the weights of the neural network. The geometric parameters of the microlens array include: arrangement period, aperture size of each microlens, and focal length of the microlens.

[0047] Specifically, in this embodiment, the implementation method for comparative training and optimization includes the following steps: 3.1 Construct the optical transmission model as a differentiable computational graph, so that the fused optical field data input to the convolutional neural network is discretized and differentiable, and its gradient can be calculated.

[0048] 3.2 The convolutional neural network, based on the current weights and microlens parameters, fuses the four-dimensional light field data input to the convolutional neural network. A series of augmented reality scene images with different focal planes are predicted.

[0049] 3.3 During training, the predicted augmented reality scene images at different focal planes are compared with the expected images, the differences between the two are calculated, and the gradient of the loss function is backpropagated to the weights of the convolutional network and the microlens geometric parameters in the optical transmission model.

[0050] 3.4. Through repeated iterations, the weights of the convolutional neural network and the geometric parameters of the microlenses in the optical transmission model are optimized to reduce the difference between the predicted augmented reality scene images at different focal planes and the desired images.

[0051] 3.5 Once the difference between the predicted augmented reality scene images at different focal planes and the desired image reaches a preset value, the optimized microlens array parameters and the trained convolutional neural network training model are obtained.

[0052] II. Manufacturing Stage 4. Based on the optimized microlens array parameters, update the parameters of the microlens array built in step 1, and deploy the trained convolutional neural network control system to the hardware control system so that the convolutional neural training model in the hardware control system matches the updated microlens array parameters.

[0053] Based on the optimized microlens array parameters (including radius of curvature, aperture size, and arrangement period), a photolithography process can be used to fabricate the microlens array on an optical substrate. The process includes photoresist coating, mask exposure, development, and etching steps to ensure that the optical surface accuracy meets the requirements of the diffraction model.

[0054] III. Application Stage 5. Based on the eye-tracking device, the focal length of the human eye is captured in real time, and the captured focal length is input into the hardware control system. The hardware control system controls the convolutional neural training model according to the captured focal length and dynamically outputs the augmented reality scene corresponding to the focal plane to the human eye.

[0055] Referring to the optimized method usage process of this embodiment, this embodiment also provides an augmented reality perspective optical system that implements the above method, including: The virtual image module is used to generate virtual image light field data; The virtual-real scene capture module is used to fuse virtual object light field data and real light field scene data, and capture them using a light field camera to finally obtain the fused light field data to be processed. The rendering module uses the geometric parameters of the microlens array and the convolutional neural network to perform zoom rendering on the light field data to be processed, obtain augmented reality scene images of different focal planes, and establish an augmented reality scene image library of different focal planes. The training module compares and trains augmented reality scene images with ideal synthetic data at different focal planes to obtain the optimal geometric parameters of the microlens array and the trained convolutional neural network. The optimized microlens array parameters obtained through the training module, along with the trained convolutional neural network model, will be updated in the virtual and real scene capture module and the training module.

[0056] The augmented reality perspective optical system also includes a display module. Based on the focal length of the human eye captured by the eye-tracking device, the display module can select suitable augmented reality scene images from a pre-generated augmented reality scene image library and dynamically input the suitable augmented reality scene images into the human eye.

[0057] The virtual scene capture module includes a lens, a microlens array, a CCD camera, and a semi-transparent mirror. The virtual image module includes a display. The computer-generated virtual information is projected through the display to generate virtual object light field data. The real light field generated in the real environment and the virtual object light field data projected by the display are superimposed by the semi-transparent mirror to form a composite light field. The composite light field is captured by the CCD sensor after passing through the CCD camera lens and the microlens array.

[0058] The method steps in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory, flash memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a head-mounted display device or terminal device. Of course, the processor and storage medium can also exist as discrete components in a head-mounted display device or terminal device.

[0059] The above embodiments are merely illustrative of the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and they should all be covered within the scope of the claims and specification of this application. In particular, as long as there is no structural conflict, the various technical features mentioned in the embodiments can be combined in any way. This application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. An optimization method for augmented reality displays, characterized in that, The optimization method includes the following steps: Step 1: Establish an optical transmission model for an augmented reality display system based on a microlens array to obtain fused light field data of the real and virtual scenes; Step 2: Design a convolutional neural network to optimize the fused light field data after fusing real and virtual scenes, and generate augmented reality scene images with different focal planes; The convolutional neural network includes a super-resolution module and a zoom rendering module. The super-resolution module is based on the LF-InterNet network structure and reconstructs the input fused light field data to improve its spatial resolution and obtain high-resolution multi-view images. The zoom rendering module can re-render the light field array sub-images enhanced by super-resolution into augmented reality scene images focused on different specific planes according to specified refocusing parameters. The zoom rendering module calculates the refocusing integral formula for images focused on a specific depth plane according to the following formula: , in, For the rendered augmented reality scene image, To render the spatial coordinates of the image plane, These are the spatial coordinates of the microlens plane, used to represent the direction information of light rays; Recorded the fused light field passing through the aperture point Arrival at CCD sensor point The intensity of light; spatial coordinates pass The relationship is determined by the rendering coordinates With microlens coordinates Sure; It is the refocusing factor, determined by the focal plane. ,in The main focal length of the CCD camera lens. The equivalent focal length corresponding to the target focusing plane. This indicates the angle between the light ray and the normal to the CCD sensor. It is the attenuation factor; Step 3: Optimize the geometric parameters of the microlens array and the convolutional neural network training model by comparing and training augmented reality scene images with ideal synthetic data at different focal planes; Step 4: Based on the optimized microlens array parameters, update the parameters of the microlens array built in Step 1, and deploy the trained convolutional neural network to the hardware control system so that the training model of the convolutional neural network in the hardware control system matches the updated microlens array parameters. Step 5: Based on the eye-tracking device, capture the human eye's focal length in real time, and input the captured human eye's focal length into the hardware control system. The hardware control system controls the convolutional neural network training model according to the captured human eye's focal length, and dynamically outputs the augmented reality scene corresponding to the focal plane to the human eye.

2. The optimization method for augmented reality display as described in claim 1, characterized in that, In step 1, the augmented reality display system based on microlens array includes a microlens array, a CCD camera, a semi-transparent mirror, and a display. The CCD camera includes a lens and a CCD sensor. The computer-generated virtual object is projected onto the real scene through the display. The real light field generated in the real environment and the light field data of the virtual object projected by the display are superimposed by the semi-transparent mirror to form a composite light field. After the composite light field passes through the lens and microlens array, the fused light field formed by phase modulation is captured by the CCD sensor.

3. The optimization method for augmented reality display as described in claim 1, characterized in that, In step 3, by comparing and training augmented reality scene images with ideal synthetic data at different focal planes, gradients are calculated based on the differentiable optical transmission model to achieve joint optimization of the geometric parameters of the microlens array and the weights of the neural network. The geometric parameters of the microlens array include: arrangement period, aperture size of each microlens, and focal length of the microlens.

4. The optimization method for augmented reality display as described in claim 3, characterized in that, Step 3, the implementation method for comparative training and optimization includes the following steps: 3.1 Construct the optical transmission model as a differentiable computational graph, so that the fused optical field data input to the convolutional neural network is discretized and differentiable, and its gradient can be calculated; 3.2 The convolutional neural network, based on the current weights and microlens parameters, fuses the four-dimensional light field data input to the convolutional neural network. A series of augmented reality scene images with different focal planes are predicted; 3.3 During the training process, the predicted augmented reality scene images at different focal planes are compared with the expected images, the differences between the two are calculated, and the gradient of the loss function is simultaneously backpropagated to the weights of the convolutional neural network and the microlens geometric parameters in the optical transmission model. 3.

4. Through repeated iterations, the weights of the convolutional neural network and the geometric parameters of the microlenses in the optical transmission model are optimized to reduce the difference between the predicted augmented reality scene images at different focal planes and the desired images. 3.5 Once the difference between the predicted augmented reality scene images at different focal planes and the desired image reaches a preset value, the optimized microlens array parameters and the trained convolutional neural network training model are obtained.

5. The optimization method for augmented reality display as described in claim 1, characterized in that, In step 5, after the eye-tracking device captures the user's eye focusing, it calculates the convergence point of the user's gaze using the pupil-corneal reflection vector and maps it to the depth of the target focusing plane in the scene; Based on the depth of the target focal plane, the corresponding scene data is retrieved from a pre-generated library of augmented reality scene images of different focal planes and input into the human eye.

6. An augmented reality perspective optical system based on the optimization method of any one of claims 1-5, characterized in that, include: The virtual image module is used to generate virtual image light field data; The virtual-real scene capture module is used to fuse virtual object light field data and real light field scene data, and capture them using a light field camera to finally obtain the fused light field data to be processed. The rendering module uses the geometric parameters of the microlens array and the convolutional neural network to perform zoom rendering on the fused light field data to be processed, obtain augmented reality scene images of different focal planes, and establish an augmented reality scene image library of different focal planes. The training module compares and trains augmented reality scene images with ideal synthetic data at different focal planes to obtain the optimal geometric parameters of the microlens array and the trained convolutional neural network. The optimized microlens array parameters obtained through the training module, along with the trained convolutional neural network model, will be updated in the virtual and real scene capture module and the training module.

7. An augmented reality perspective optical system as described in claim 6, characterized in that, The augmented reality perspective optical system also includes a display module and an eye-tracking device that can capture the focal length of the human eye. The display module can select suitable augmented reality scene images from a pre-generated augmented reality scene image library based on the focal length information of the human eye, and dynamically input the suitable augmented reality scene images into the human eye.

8. An augmented reality perspective optical system as described in claim 7, characterized in that, The virtual-real scene capture module includes a lens, a microlens array, a CCD camera, and a semi-transparent mirror. The virtual image module includes a display. Computer-generated virtual information is projected through the display to generate virtual object light field data. The real light field generated in the real environment and the virtual object light field data projected by the display are superimposed by the semi-transparent mirror to form a composite light field. The composite light field is captured by the CCD sensor after passing through the CCD camera lens and the microlens array. The semi-transparent mirror is at a 45-degree angle to the horizontal plane. The real object and the virtual object displayed on the display are located on the two light-splitting surfaces of the semi-transparent mirror. The light field camera, which consists of a lens, a microlens array, and a CCD sensor, should be located on the light-combining surface of the semi-transparent mirror.

Citation Information

Patent Citations

  • Light field display system for augmented reality and augmented reality device

    CN110618529A

  • Light field three-dimensional imaging method and system based on neural network

    CN117934708A