A holographic display system based on AI visual recognition technology

By combining deep learning and computer graphics technology, a holographic display system is proposed, which solves the efficiency and accuracy problems of dynamic scenes and complex lighting processing in the existing technology, and realizes efficient and real three-dimensional holographic image generation.

CN119668413BActive Publication Date: 2025-06-20HUMKA (FUJIAN) DISPLAYS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510190784.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-20
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

Existing holographic display technologies face efficiency and accuracy problems when generating dynamic scenes and handling complex lighting and material changes, and rely on specific hardware devices, making it difficult to achieve real-time rendering and efficient computing.

Method used

Combining deep learning technology and computer graphics, a new image processing and rendering method is proposed. Image spatial information is extracted through deep learning models, combined with lighting and material information, and high-precision three-dimensional holographic images are generated, and image display is adjusted through line-of-sight tracing and user interaction.

Benefits of technology

It realizes efficient and real three-dimensional holographic image generation, improves spatial accuracy and image authenticity, optimizes computing efficiency, and supports real-time rendering and dynamic scene display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119668413B_ABST
    Figure CN119668413B_ABST
Patent Text Reader

Abstract

The present invention discloses a holographic display system based on AI vision recognition technology, which relates to the field of AI vision recognition technology. It includes obtaining real-time image data through at least one image sensor, and preprocessing the image data; using a deep learning model carried by an algorithm module to analyze the processed image data to generate three-dimensional spatial information, constructing a holographic image based on the three-dimensional spatial information, and rendering the holographic image in combination with lighting and material information; adjusting the display angle of the holographic image according to the user's position and perspective information, while monitoring changes in the scene, automatically adjusting the display parameters of the holographic image in the holographic image display system, and dynamically modifying the holographic image through user interaction input. The present invention realizes efficient and accurate three-dimensional holographic image generation, has broad application potential, and is particularly suitable for fields such as virtual reality, augmented reality, and holographic display that require dynamic display and strong interactivity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AI visual recognition technology, and particularly to a holographic display system based on AI visual recognition technology. Background Art

[0002] With the development of technologies such as virtual reality (VR), augmented reality (AR), and holographic imaging, three-dimensional image and holographic display technologies have become important research directions in modern technology. Although traditional three-dimensional display technologies can present the stereoscopic effect of objects, they often rely on specific hardware devices such as stereoscopic glasses and virtual reality helmets, which limits their application scope. Although holographic display can achieve medium-free and natural three-dimensional image display, the existing technologies still face great challenges in generating dynamic scenes and dealing with complex lighting and material changes. Especially in generating high-quality holographic images efficiently and in real time, the rendering efficiency, calculation accuracy, and image quality of the existing technologies are still difficult to meet the actual application requirements. Therefore, how to combine deep learning and computer graphics technologies to achieve efficient and realistic holographic image generation has become an important direction in the current technological development. Summary of the Invention

[0003] In view of the above existing problems, the present invention is proposed.

[0004] Therefore, the present invention provides a holographic display system based on AI visual recognition technology, aiming to solve several key problems in the existing holographic display technology. First, when the existing methods extract image spatial information, the accuracy is insufficient. Especially when dealing with the depth, position, size, and dynamic changes in complex scenes, the three-dimensional features of objects cannot be accurately obtained. Second, the existing holographic image rendering methods fail to effectively integrate lighting, material properties, and object geometric forms, resulting in the generated holographic images not being realistic enough, especially in terms of lighting changes and material detail performance. Finally, the existing technologies mostly rely on specific hardware devices, and the calculation process is relatively complex, making it difficult to achieve real-time rendering and efficient calculation. The present invention combines deep learning technology with computer graphics to propose a new image processing and rendering method, which can accurately extract spatial information, combine the lighting and material characteristics of objects, and quickly generate high-precision three-dimensional holographic images, thus effectively solving the above problems in the existing technologies.

[0005] To solve the above technical problems, the present invention provides the following technical solution. A holographic display system based on AI visual recognition technology includes:

[0006] Acquire real-time image data through at least one image sensor, and preprocess the image data to extract spatial feature information in the image, including the depth, position, size, motion trajectory, and lighting conditions of the object;

[0007] Use the deep learning model carried by the algorithm module to analyze the processed image data, generate three-dimensional space information from it, construct a holographic image based on the three-dimensional space information, and render the holographic image in combination with lighting and material information;

[0008] Through the gaze tracking system, gesture recognition system, eye movement tracking system and speech recognition system, adjust the display angle of the holographic image according to the position and perspective information of the user, and at the same time monitor the changes in the scene, automatically adjust the display parameters of the holographic image in the holographic image display system, and dynamically modify the holographic image through user interaction input.

[0009] As a preferred solution of the holographic display system based on AI vision recognition technology of the present invention, wherein: the image sensor is at least one of a camera, a depth sensor, and a lidar sensor, and is used to obtain video stream or static image data.

[0010] As a preferred solution of the holographic display system based on AI vision recognition technology of the present invention, wherein: the preprocessing includes using an adaptive denoising method to perform denoising operations on the input image, dynamically adjusting the denoising weight according to the local image gradient, assuming that the original image pixel value is , and the pixel value of the denoised image is , and the denoising operation is expressed as:

[0011]

[0012] Among them, is the adaptive filtering weight, calculated based on the image gradient, and the weight formula is:

[0013]

[0014] Among them, represents the position coordinates in the image; n represents the radius of the convolution window; represents the gradient of the image in the local area, is a parameter that controls the influence of the image gradient on the weight, is the balance term, is the normalization factor to ensure that the sum of the weights is 1;

[0015] After denoising, the local contrast of the image is not sufficient to clearly present the details, and a contrast enhancement method based on local statistics needs to be introduced. Calculate the local mean and standard deviation of the image, enhance the contrast of the image according to the local mean and standard deviation, and dynamically adjust the enhancement intensity through the gradient information of the image. The enhanced image can be expressed as:

[0016] Among them, and are the mean and standard deviation of the local region, is the contrast enhancement coefficient, is the adaptive contrast adjustment coefficient, expressed as:

[0017]

[0018] where, is the local gradient of the image, is the parameter that controls the influence of the gradient on the enhancement intensity, is a constant;

[0019] Edge detection is implemented using a convolutional neural network CNN, and the obtained edge image is calculated through a convolution operation:

[0020]

[0021] where, represents the edge detection output value of the image at position ; represents the pixel value of the image at position ; is the weight of the convolution kernel, is the bias term, is the activation function; The loss function consists of the difference between the edge image and the true edge image and the smoothness constraint of the gradient information:

[0022]

[0023] where, is the true edge image, is the regularization coefficient; is the gradient of the image ; is the gradient of the image ;

[0024] The depth map of the image is obtained from different perspectives , and the depth information of multiple perspectives is fused by combining the correlation of the images through weighted averaging. The fused depth map is calculated by the following formula:

[0025]

[0026] where, is the weighting coefficient of the k-th perspective, and the calculation formula is:

[0027]

[0028] where, is the Images from a perspective and the reference perspective The correlation metric between them is a parameter that controls the influence of image correlation on the weighting coefficient is the normalization factor, and K is the total number of perspectives

[0029] As a preferred embodiment of the holographic display system based on AI vision recognition technology described in the present invention, wherein: the deep learning model includes that the image is subjected to feature extraction through multiple convolutional layers of a convolutional neural network, and the output of each convolutional layer is expressed as

[0030] wherein is the output feature map after the th layer of convolutional operation is the weight of the convolutional kernel of the th layer is the bias term

[0031] By performing weighted summation on the outputs of multiple layers of the convolutional neural network, a depth map is generated, wherein the output of each layer is given different weights according to its importance and added with the bias term to obtain the depth map

[0032]

[0033] wherein is the generated depth map, representing the depth value of each pixel point is the th layer output feature map weight is the bias term is the total number of convolutional layers

[0034] As a preferred embodiment of the holographic display system based on AI vision recognition technology described in the present invention, wherein: the deep learning model further includes that the generative adversarial network GAN takes the image and the noise vector as inputs and generates three-dimensional spatial information through the generator :

[0035]

[0036] wherein is the three-dimensional spatial information output by the generator is the mapping function of the generator, responsible for converting the input image and noise vector into three-dimensional spatial information

[0037] Discriminator of the generative adversarial network Responsible for judging the three-dimensional information output by the generator Whether it is real, where the output of the discriminator is a probability value representing the likelihood that the generated data is real:

[0038]

[0039] Among them, is the judgment value of the discriminator for the generated three-dimensional information ; is the weight of the discriminator, is the bias term of the discriminator;

[0040] Loss function of the generator It is expressed as:

[0041]

[0042] Among them, is the loss function of the generator, representing the quality of the generated data; is the expectation;

[0043] Loss function of the discriminator It is expressed as:

[0044]

[0045] Among them, is the loss function of the discriminator, representing the classification ability of the discriminator; is the judgment value of the discriminator for the real data ;

[0046] As a preferred solution of the holographic display system based on the AI vision recognition technology described in the present invention, wherein: the construction of the holographic image includes, according to the three-dimensional space information generated by the deep learning model, combining the depth map and the spatial layout data of the object, reconstructing the geometric shape of the object in the image, mapping the depth and position data of the object into a three-dimensional coordinate system by using three-dimensional modeling technology, generating a mesh model of the object, defining the points, edges and faces on the surface of the object, and constituting the spatial structure of the object. This process depends on the spatial features extracted from the depth map, and generates a geometric model of the object through point cloud data or surface reconstruction technology;

[0047] According to the material properties, calculate the reflection and refraction behaviors of the object surface for different light sources; the material properties include the smoothness, roughness, transparency and refractive index of the object surface, which determine the interaction mode of light with the object surface; the light source, i.e., the lighting condition, includes ambient light, directional light, point light source and spotlight, which affect the reflection, refraction and shadow characteristics of the object surface;

[0048] Using the Phong lighting model in computer graphics, the interaction between light and the object surface is simulated by calculating the paths of incident light, reflected light, and refracted light from the light source to the object surface. During the rendering process, the intensity and direction of the incident light are calculated first, then the direction of the reflected light and the surface reflection coefficient are calculated, and the intensity of the reflected light is adjusted according to the material properties of the object. For transparent objects, the refracted light path is calculated using the refractive index and the refraction angle, and the brightness and color are adjusted according to the ambient light source and the glossiness of the object surface.

[0049] The generation of the image is completed by projecting the three-dimensional object onto a two-dimensional plane. The projection process transforms the position, size, and orientation of the object in three-dimensional space into two-dimensional image coordinates through perspective projection or orthographic projection.

[0050] As a preferred embodiment of the holographic display system based on AI vision recognition technology according to the present invention, wherein: the viewing angle information is obtained through a position sensor and a gaze tracking system, and the gaze tracking system is composed of a depth camera or an infrared sensor for monitoring the position and movement of the user's head or eyes.

[0051] The change in the monitored scene is obtained by the sensor acquiring information on the change in ambient light and the position of the object.

[0052] As a preferred embodiment of the holographic display system based on AI vision recognition technology according to the present invention, wherein: the user interaction input is obtained through a gesture recognition system, an eye movement tracking system, or a voice recognition system; the interaction input is processed in real time and fed back to the holographic image display system to adjust the objects or display content in the image.

[0053] A computer device includes a memory and a processor, the memory stores a computer program, and it is characterized in that when the processor executes the computer program, the steps of the holographic display system based on AI vision recognition technology are implemented.

[0054] A computer-readable storage medium stores a computer program, and it is characterized in that when the computer program is executed by a processor, the steps of the holographic display system based on AI vision recognition technology are implemented.

[0055] Advantages of the present invention: By introducing deep learning technology to accurately analyze image data, the present invention can efficiently extract spatial feature information in images, including the depth, position, size, and dynamic changes of objects. This innovation has significantly improved the spatial accuracy and authenticity of holographic image generation. Secondly, by combining lighting and material rendering technologies in computer graphics, the present invention can accurately render the surface features of objects based on the generated three-dimensional spatial data, considering factors such as lighting, shadows, and reflections, to generate more realistic and detailed holographic images. In addition, the present invention has optimized image processing and rendering algorithms, improving computational efficiency, enabling the image generation and rendering process to be completed in a short time, and meeting the holographic display requirements of real-time or dynamic scenarios. Generally speaking, the present invention realizes efficient and accurate three-dimensional holographic image generation, has broad application potential, and is particularly suitable for fields such as virtual reality, augmented reality, and holographic display that require dynamic display and strong interactivity. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 FIG. is a schematic flowchart of a holographic display system based on AI vision recognition technology provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] Please refer to Figure 1 , which is the first embodiment of the present invention. This embodiment provides a holographic display system based on AI vision recognition technology, including:

[0059] Obtain real-time image data through at least one image sensor, and preprocess the image data to extract spatial feature information in the image, including the depth, position, size, motion trajectory, and lighting conditions of the object;

[0060] Use the deep learning model carried by the algorithm module to analyze the processed image data, generate three-dimensional spatial information therefrom, construct a holographic image based on the three-dimensional spatial information, and render the holographic image in combination with lighting and material information;

[0061] Through a gaze tracking system, a gesture recognition system, an eye movement tracking system, and a speech recognition system, adjust the display angle of the holographic image according to the position and perspective information of the user, monitor changes in the scene, automatically adjust the display parameters of the holographic image in the holographic image display system, and dynamically modify the holographic image through user interaction input.

[0062] The image sensor is at least one of a camera, a depth sensor, and a lidar sensor, and is used to obtain video stream or static image data.

[0063] The preprocessing includes performing a denoising operation on the input image using an adaptive denoising method, dynamically adjusting the denoising weight according to the local image gradient. Assume the pixel value of the original image is , and the pixel value of the denoised image is . The denoising operation is expressed as:

[0064] where is the adaptive filtering weight, calculated based on the image gradient. The weight formula is:

[0065]

[0066] where represents the position coordinates in the image; n represents the radius of the convolution window; represents the gradient of the image in the local area, is a parameter that controls the influence of the image gradient on the weight, is the balance term, is the normalization factor to ensure that the sum of the weights is 1;

[0067] After denoising, the local contrast of the image is not sufficient to clearly present details. It is necessary to introduce a contrast enhancement method based on local statistics, calculate the local mean and standard deviation of the image, enhance the contrast of the image according to the local mean and standard deviation, and dynamically adjust the enhancement intensity through the gradient information of the image. The enhanced image can be expressed as:

[0068] where and are the mean and standard deviation of the local area respectively, is the contrast enhancement coefficient, is the adaptive contrast adjustment coefficient, expressed as:

[0069]

[0070] where is the local gradient of the image, is a parameter that controls the influence of the gradient on the enhancement intensity, is a constant;

[0071] Edge detection is implemented using a convolutional neural network (CNN), and the obtained edge image is calculated through a convolution operation:

[0072]

[0073] where, represents the edge detection output value of the image at position ; represents the pixel value of the image at position ; is the weight of the convolution kernel, is the bias term, is the activation function; the loss function consists of the difference between the edge image and the true edge image and the smoothness constraint of the gradient information:

[0074]

[0075] where, is the true edge image, is the regularization coefficient; is the gradient of the image at ; is the gradient of the image at ;

[0076] The depth map of the image is obtained from different perspectives , and the depth information of multiple perspectives is fused by combining the correlation of the images in a weighted average manner. The fused depth map is calculated by the following formula:

[0077]

[0078] where, is the weighting coefficient of the k-th perspective, and the calculation formula is:

[0079]

[0080] where, is the correlation measure between the image of the -th perspective and the reference perspective , is the parameter that controls the influence of the image correlation on the weighting coefficient, is the normalization factor, and K is the total number of perspectives.

[0081] The deep learning model includes that the image is subjected to feature extraction through multiple convolutional layers of the convolutional neural network, and the output of each convolutional layer

[0081] is expressed as:

[0082]

[0083] Among them, is the output feature map after the -th layer of convolutional operation, is the weight of the -th layer of convolutional kernel, and

[0084] is the bias term; By performing weighted summation on the outputs of multiple layers of the convolutional neural network, a depth map is generated, where the output of each layer is assigned different weights according to its importance, and the bias term

[0085]

[0086] is added to obtain the depth map: Among them, is the generated depth map, representing the depth value of each pixel point; is the weight of the output feature map of the -th layer, is the bias term;

[0087] The deep learning model further includes that the generative adversarial network GAN takes an image and a noise vector as inputs, and generates three-dimensional spatial information through a generator :

[0088]

[0089] Among them, is the three-dimensional spatial information output by the generator, is the mapping function of the generator, responsible for converting the input image and noise vector into three-dimensional spatial information;

[0090] The discriminator of the generative adversarial network is responsible for judging whether the three-dimensional information output by the generator is real, where the output of the discriminator is a probability value, representing the possibility that the generated data is real:

[0091]

[0092] Among them, is the judgment value of the discriminator on the generated three-dimensional information ; is the weight of the discriminator, is the bias term of the discriminator;

[0093] The loss function of the generator is expressed as:

[0094]

[0095] where, is the loss function of the generator, representing the quality of the generated data; is the expectation;

[0096] The loss function of the discriminator is expressed as:

[0097]

[0098] where, is the loss function of the discriminator, representing the classification ability of the discriminator; is the discriminator's judgment value for real data of.

[0099] The construction of the holographic image includes reconstructing the geometric shape of the object in the image according to the three-dimensional space information generated by the deep learning model, combining the depth map and the spatial layout data of the object, mapping the depth and position data of the object to a three-dimensional coordinate system using three-dimensional modeling technology to generate a mesh model of the object, defining the points, edges, and faces on the surface of the object to form the spatial structure of the object. This process depends on the spatial features extracted from the depth map, and a geometric model of the object is generated through point cloud data or surface reconstruction technology;

[0100] Calculating the reflection and refraction behaviors of the object surface for different light sources according to the material properties; the material properties include the smoothness, roughness, transparency, and refractive index of the object surface, which determine the interaction mode of light with the object surface; the light source, i.e., the lighting condition, includes ambient light, directional light, point light source, and spotlight, which affect the reflection, refraction, and shadow characteristics of the object surface;

[0101] Using the Phong model, a lighting model in computer graphics, to simulate the interaction of light with the object surface by calculating the paths of incident light, reflected light, and refracted light from the light source to the object surface. During the rendering process, first calculate the intensity and direction of the incident light, then calculate the direction of the reflected light and the surface reflection coefficient, and adjust the intensity of the reflected light according to the material properties of the object; for transparent objects, calculate the refracted light path using the refractive index and refraction angle, and adjust the brightness and color according to the ambient light source and the glossiness of the object surface;

[0102] The generation of the image is completed by projecting a three-dimensional object onto a two-dimensional plane. In the projection process, the position, size, and pose of the object in the three-dimensional space are transformed into two-dimensional image coordinates through perspective projection or orthographic projection.

[0103] The perspective information is obtained through a position sensor and a gaze tracking system. The gaze tracking system consists of a depth camera or an infrared sensor and is used to monitor the position and movement of the user's head or eyes.

[0104] Among them, the changes in the monitored scene are obtained by the sensor acquiring information on changes in ambient light and object position.

[0105] The user interaction input is obtained through a gesture recognition system, an eye movement tracking system, or a speech recognition system. The interaction input is processed in real time and fed back to the holographic image display system to adjust the objects or display content in the image.

[0106] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

[0107] Embodiment 2, the second embodiment of the present invention, which is different from the previous embodiment:

[0108] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution of the present invention, or the part that contributes to the prior art, or this part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks and other various media that can store program codes.

[0109] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or one or more of the blocks.

[0110] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or one or more of the blocks.

[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or one or more of the blocks.

[0112] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0113] Obviously, those skilled in the art can make various changes and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A holographic display method based on AI visual recognition technology, characterized in that: include, Acquire real-time image data through at least one image sensor, and pre-process the image data to extract spatial feature information in the image, including depth, position, size, motion trajectory and lighting conditions of the object; Use the deep learning model carried by the algorithm module to analyze the processed image data, generate three-dimensional spatial information from it, build a holographic image based on the three-dimensional spatial information, and render the holographic image in combination with lighting and material information; Through the gaze tracking system, gesture recognition system, eye tracking system and voice recognition system, the display angle of the holographic image is adjusted according to the user's position and viewing angle information, and at the same time, the changes in the scene are monitored, the display parameters of the holographic image in the holographic image display system are automatically adjusted, and the holographic image is dynamically modified through user interactive input; The preprocessing includes performing a denoising operation on the input image using an adaptive denoising method, dynamically adjusting the denoising weight according to the local image gradient, assuming that the original image pixel value is I'(x, y), and the denoised image pixel value is I(x, y), the denoising operation is expressed as: Among them, w(x,y,k,l) ​​is the adaptive filter weight, which is calculated based on the image gradient. The weight formula is: Among them, (x, y) represents the position coordinates in the image; n represents the radius of the convolution window; Represents the gradient of the image in the local area, α is the parameter that controls the influence of the image gradient on the weight, β is the balance term, and Z(x,y) is the normalization factor to ensure that the sum of the weights is 1; After denoising, the local contrast of the image is not enough to clearly present the details. It is necessary to introduce a contrast enhancement method based on local statistics, calculate the local mean and standard deviation of the image, enhance the image contrast according to the local mean and standard deviation, and dynamically adjust the enhancement intensity through the gradient information of the image. The enhanced image It is expressed as: Among them, μ(x, y) and ∝(x, y) are the mean and standard deviation of the local area, α1 is the contrast enhancement coefficient, and γ(x, y) is the adaptive contrast adjustment coefficient, which is expressed as: in, is the local gradient of the image, δ is the parameter that controls the influence of the gradient on the enhancement strength, and ∈ is a constant; Edge detection is achieved using a convolutional neural network (CNN), and the resulting edge image is calculated using a convolution operation: Where E(x,y) represents the edge detection output value of the image at position (x,y); represents the pixel value of the image at position (x+i, y+j); w ij is the weight of the convolution kernel, b is the bias term, and σ is the activation function; the loss function is composed of the difference between the edge image and the true edge image and the smoothness constraint of the gradient information: Among them, E gt (x, y) is the true edge image, λ is the regularization coefficient; is the gradient of the image E(x,y); is an image The gradient at Get the depth map D of the image from different perspectives k (x, y)(k=1,2,…,K), the depth information of multiple perspectives is fused by combining the correlation of images through weighted averaging. The fused depth map D(x, y) is calculated by the following formula: Among them, w k is the weighted coefficient of the kth perspective, and the calculation formula is: Among them, corr(I k (x,y),I1(x,y)) is the image I of the kth view k is the correlation measure between (x, y) and the reference view I1(x, y), α2 is a parameter that controls the effect of image correlation on the weighting coefficient, Z is a normalization factor, and K is the total number of views; The deep learning model includes: the image is subjected to feature extraction through multiple convolutional layers of a convolutional neural network, and the output C of each convolutional layer l It is expressed as: Among them, C l (x,y) is the output feature map after the l-th layer convolution operation, is the weight of the convolution kernel of the lth layer, b l is the bias term; The depth map D(x,y) is generated by weighted summing of the multi-layer outputs of the convolutional neural network, where the output C of each layer l (x,y) are assigned different weights w according to their importance l , and add the bias term b d , and get the depth map: Where D(x,y) is the generated depth map, which represents the depth value of each pixel; w l is the output feature map C of the lth layer l The weight of (x,y), b d is the bias term; L is the total number of convolutional layers; The deep learning model also includes a generative adversarial network GAN that takes the image and the noise vector z as input and generates three-dimensional spatial information through a generator G: in, is the three-dimensional space information output by the generator, It is the mapping function of the generator, which is responsible for converting the input image and noise vector z into three-dimensional spatial information; The discriminator D of the generative adversarial network is responsible for judging the three-dimensional spatial information output by the generator. Is it real? The output of the discriminator is a probability value indicating the possibility that the generated data is real: in, is the three-dimensional spatial information generated by the discriminator The judgment value of D is the weight of the discriminator, b D is the bias term of the discriminator; The loss function of the generator It is expressed as: in, is the loss function of the generator, which indicates the quality of the generated data; It is expectation; Loss function of the discriminator It is expressed as: in, is the loss function of the discriminator, which indicates the classification ability of the discriminator; D(I real ) is the discriminator for the real data I real The judgment value of .

2. A holographic display method based on AI visual recognition technology as claimed in claim 1, characterized in that: The image sensor is at least one of a camera, a depth sensor, and a lidar sensor, and is used to obtain video stream or static image data.

3. A holographic display method based on AI visual recognition technology as claimed in claim 2, characterized in that: The construction of the holographic image includes, based on the three-dimensional spatial information generated by the deep learning model, combining the depth map and the spatial layout data of the object, reconstructing the geometric shape of the object in the image, mapping the depth and position data of the object into a three-dimensional coordinate system by using three-dimensional modeling technology, generating a mesh model of the object, defining points, edges and faces on the surface of the object, and forming the spatial structure of the object; Calculate the reflection and refraction behavior of the object surface to different light sources based on material properties; the material properties include the smoothness, roughness, transparency, and refractive index of the object surface, which determine the interaction between light and the object surface; the light source, i.e., lighting conditions, includes ambient light, directional light, point light source, and spotlight, which affect the reflection, refraction, and shadow characteristics of the object surface; The Phong model, a lighting model in computer graphics, is used to simulate the interaction between light and the surface of an object by calculating the paths of incident light, reflected light, and refracted light from the light source to the surface of the object. In the rendering process, the intensity and direction of the incident light are calculated first, and then the direction of the reflected light and the surface reflection coefficient are calculated, and the intensity of the reflected light is adjusted according to the material properties of the object. For transparent objects, the path of the refracted light is calculated using the refractive index and refraction angle, and the brightness and color are adjusted according to the ambient light source and the glossiness of the object surface. The image is generated by projecting a three-dimensional object onto a two-dimensional plane. The projection process converts the position, size and posture of the object in the three-dimensional space into two-dimensional image coordinates through perspective projection or orthographic projection.

4. The holographic display method based on AI visual recognition technology according to claim 3, characterized in that: The viewing angle information is obtained through a position sensor and a gaze tracking system, wherein the gaze tracking system is composed of a depth camera or an infrared sensor for monitoring the position and movement of the user's head or eyes; The changes in the monitoring scene are carried out by obtaining information about the ambient light and the position of objects through sensors.

5. A holographic display method based on AI visual recognition technology as claimed in claim 4, characterized in that: The user interaction input is obtained through a gesture recognition system, an eye tracking system or a voice recognition system; the interaction input is processed in real time and fed back to the holographic image display system to adjust the objects in the image or the displayed content.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Holographic display method based on holographic sand table and related device

    CN117075739A

  • Method and apparatus for processing holographic image

    US20210279951A1