Neural network-based 3D naked eye image generation method and system

Through the 3D naked-eye image generation method based on neural network, combined with the Monodepth2 model and the U-Net model, the problem of difficult to balance image quality and real-time in the prior art is solved, and high-quality and low-energy 3D naked-eye image generation is achieved, enhancing the realism and immersion experience of the image.

CN120091122AActive Publication Date: 2025-06-03SHANGHAI YINYU DIGITAL TECH GRP CO LTD

Patent Information

Application Number
CN202510563421.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing 3D naked-eye image generation technology is difficult to take into account both image quality and real-time, especially when processing high-speed moving objects, high computational complexity leads to picture tearing or delay.

Method used

Using a neural network-based method, the stereoscopic view data set is obtained for preprocessing, the pre-trained Monodepth2 model is loaded to analyze the parallax, the depth map is calculated based on the principle of triangulation, and the depth map is input into the U-Net model for repeated iterative training to generate a 3D naked-eye image. At the same time, dynamic area enhancement technology and pruning optimization are applied to enable image generation to run smoothly on mobile devices with low energy consumption.

Benefits of technology

The comprehensive performance and practicality of the system are improved, and the generated 3D naked-eye images have high quality and realism, and can run efficiently on mobile devices, solving the problems of picture tearing and delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091122A_ABST
    Figure CN120091122A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D naked eye image generation method and system based on a neural network, and relates to the technical field of 3D naked eye image generation, and the method comprises the steps: inputting a disparity map and a preprocessed stereoscopic view into an end-to-end neural network frame, and a generator in the GANs is combined to generate a 3D naked eye image according to the input disparity map and the stereo view, a neural network model is used for the 3D naked eye image, and the neural network model is optimized, so that the 3D naked eye image can efficiently run on the mobile equipment, and the 3D naked eye image can be efficiently displayed. And performing color and brightness adjustment according to a 3D naked eye image generated by the optimized neural network model. Through accurate size adjustment and color correction, errors caused by original data differences are reduced, the accuracy of subsequent depth estimation is improved, important feature points can be more accurately positioned in a complex scene by adopting an attention mechanism algorithm, the accuracy of a disparity map is improved, and the sense of reality and the three-dimensional effect of a final 3D naked eye image are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of 3D naked-eye image generation, and particularly to a 3D naked-eye image generation method and system based on a neural network. Background Art

[0002] With the rapid development of information technology, 3D image display technology, as a key part of the human-computer interaction interface, has gradually become the focus of attention in the scientific research and industrial fields. Since the early 20th century, 3D imaging technology has gone through multiple stages of development, from stereoscopic vision, holographic imaging to modern computational imaging. In recent years, with the rise of deep learning, especially neural network technology, 3D image generation methods based on deep learning have received extensive attention due to their high efficiency and high quality.

[0003] However, the existing 3D naked-eye image generation technology still has some deficiencies. Especially in the aspect of naked-eye 3D display, the 3D naked-eye image generation technology solutions often have difficulty in simultaneously considering image quality and real-time performance. When dealing with fast-moving objects, due to the high computational complexity, the existing algorithms cannot update the view in time, resulting in phenomena such as screen tearing or delay. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a 3D naked-eye image generation method based on a neural network to solve the problem that the 3D naked-eye image cannot be updated in time due to high computational complexity.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: In the first aspect, the present invention provides a 3D naked-eye image generation method based on a neural network, which includes obtaining a stereoscopic view dataset and preprocessing the stereoscopic view; Loading a pre-trained Monodepth2 model to analyze the disparity between the left and right eye views in the stereoscopic view, and applying the principle of triangulation to calculate the depth value corresponding to each pixel to obtain a depth map; Inputting the depth map into a U-Net model and combining it with a generator for iterative training; After the iterative training of the U-Net model ends, obtaining new left and right eye views and a depth map, applying dynamic region enhancement, and then inputting them into the trained U-Net model to generate a 3D naked-eye image; Pruning and optimizing the U-Net model for generating the 3D naked-eye image so that the 3D naked-eye image can run smoothly with low energy consumption on a mobile device.

[0007] As a preferred solution of the 3D naked-eye image generation method based on a neural network according to the present invention, wherein: obtaining a stereoscopic view dataset and preprocessing the stereoscopic view, specifically: The stereoscopic view includes a left-eye view and a right-eye view; The preprocessing includes size adjustment and color correction.

[0008] As a preferred solution of the 3D naked-eye image generation method based on a neural network according to the present invention, wherein: loading the pre-trained Monodepth2 model to analyze the disparity between the left-eye view and the right-eye view in the stereoscopic view, specifically, Converting the preprocessed left-eye view and right-eye view into a tensor format and adjusting them to the size required by the Monodepth2 model; Inputting the tensors of the left-eye view and the right-eye view after format conversion and size adjustment into the Monodepth2 model for forward propagation to obtain the disparity; Calculating the disparity value by applying the sum of absolute values to the corresponding pixel windows of the left-eye view and the right-eye view, and comparing the similarity of the disparity values; Establish a disparity Figure 2 dimensional array, and fill the two-dimensional array with the disparity value with the smallest similarity to obtain a disparity map.

[0009] As a preferred solution of the 3D naked-eye image generation method based on a neural network according to the present invention, wherein: applying the principle of triangulation to calculate the depth value corresponding to each pixel to obtain a depth map, specifically, According to the size of the disparity map, use NumPy to create a depth map matrix as a container for the final depth map to store the depth values; Using the principle of triangulation of a binocular camera to calculate the disparity to calculate the depth value, and filling it into the depth map matrix to obtain a depth map; Applying the Matplotlib plotting tool to check the edge sharpness of the depth map and the accuracy of the object contour.

[0010] As a preferred solution of the 3D naked-eye image generation method based on a neural network according to the present invention, wherein: inputting the depth map into the U-Net model and combining it with the generator for iterative training, specifically: Construct an encoder using 3x3 convolutional kernels and 2x2 max-pooling layers, and establish a decoder through transposed convolution to obtain a U-Net model; Initialize all the weights of the generator and the discriminator using the standard normal distribution, and add the initialized transposed convolution of the generator to the output layer of the U-Net model; Input the depth map into the U-Net model, and design adversarial loss, content loss, and discriminator loss for iterative training.

[0011] As a preferred solution of the 3D naked-eye image generation method based on neural network according to the present invention, wherein: generating the 3D naked-eye image specifically includes: After the repeated iterative training is completed, new left and right eye views and depth maps are obtained from the stereoscopic view dataset; An additional CNN model is constructed using convolutional layers, pooling layers and fully connected layers, and the CNN model is initialized by Xavier and trained using the cross-entropy loss function; The depth map is input into the trained CNN model using dynamic enhancement technology to generate an attention weight map; The newly obtained left and right eye views, depth map and attention weight map are input into the U-Net model to generate a 3D naked-eye image.

[0012] As a preferred solution of the 3D naked-eye image generation method based on neural network according to the present invention, wherein: pruning and optimizing the U-Net model for generating the 3D naked-eye image to enable the 3D naked-eye image to run efficiently on mobile devices, specifically: The proportion of the weights of the U-Net model being set to zero; According to the set pruning steps, set the start and end times of pruning, and record all relevant pruning configuration parameters to guide the pruning process; The pruning configuration parameters include expected sparsity, initial sparsity at the start of pruning, final sparsity after pruning is completed, start step of pruning and end step of pruning; According to the expected sparsity level, input the start and end times of pruning into the U-Net model for step-by-step pruning; Use the TensorFlow Lite converter tool to convert the pruned U-Net model into the tflite format Integrate the converted tflite file into Android and iOS applications and input the 3D naked-eye image.

[0013] In a second aspect, the present invention provides a 3D naked-eye image generation system based on neural network, including: A data preprocessing module extracts stereoscopic views from a dataset containing stereoscopic views and performs preprocessing operations of size adjustment and color correction on the stereoscopic views; A depth map module analyzes the left eye view and the right eye view in the stereoscopic view using the Monodepth2 model to obtain a disparity, creates a depth map matrix, calculates depth values using the triangulation principle of a binocular camera, and fills the depth values into the depth map matrix to obtain a depth map; U-Net model training module, which inputs the depth map and the preprocessed left-eye view and right-eye view Figure 1 into the U-Net model together, and combines with the generator for iterative training until the discriminator can no longer distinguish the images generated by the U-Net model from the real images; Image generation module, after the U-Net model is trained, inputs the newly obtained left-eye and right-eye views and depth map after using the dynamic region enhancement technology into the trained U-Net model to generate high-quality 3D naked-eye images; U-Net model optimization module, which prunes and optimizes the U-Net model according to pruning parameters, target sparsity, initial sparsity, final sparsity, start step and end step.

[0014] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the 3D naked-eye image generation method based on neural network as described in the first aspect of the present invention is implemented.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the 3D naked-eye image generation method based on neural network as described in the first aspect of the present invention is implemented.

[0016] The beneficial effects of the present invention are as follows: Through the overall process optimization from data preprocessing to high-quality 3D naked-eye image generation, the comprehensive performance and practicability of the system are improved. Overall, an efficient, accurate and user-friendly 3D naked-eye image generation mechanism is jointly constructed. Among them, the dynamic region enhancement technology not only provides higher depth estimation accuracy and consistency in complex scene reconstruction, but also greatly enhances the realism and immersive experience of 3D naked-eye images. By using advanced disparity optimization, attention mechanism and generative adversarial network technology, it can particularly focus on and optimize the expressiveness of 3D naked-eye images, while ensuring the effective utilization of computing resources. In addition, the standardized data preprocessing steps lay a solid foundation for subsequent processing, making the entire process more robust, adaptable, and able to quickly identify and correct potential problems, and also bringing an unprecedented visual experience to users, promoting the technological progress and application expansion in related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a flowchart of the 3D naked-eye image generation method based on neural network in Embodiment 1.

[0019] Figure 2 It is a schematic diagram of the optimization of the mobile device in Embodiment 1.

[0020] Figure 3 It is a schematic diagram of depth map estimation and disparity calculation in Embodiment 1.

[0021] Figure 4 It is a schematic diagram of training the U-Net model architecture in Embodiment 1. Specific Embodiments

[0022] To make the above objects, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings of the specification.

[0023] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0024] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that excludes other embodiments.

[0025] Embodiment 1, referring to Figures 1 to 4 , this embodiment provides a 3D naked-eye image generation method based on neural network, including the following steps: S1. Obtain a stereoscopic view dataset and preprocess the stereoscopic views. Specifically: Extract stereoscopic views from a publicly available dataset containing stereoscopic views. This dataset usually contains multiple compressed packages, such as.zip files, which need to be decompressed to a specified directory on the local hard disk. Each sequence contains multiple subfolders, including left-eye perspective images and right-eye perspective images. Confirm that each image sequence is also accompanied by a corresponding calibration file, usually in.txt format. The calibration file contains the internal parameter matrix of the camera, distortion coefficients, the external parameter matrix of the calibration information, and the reprojection error. These calibration files are crucial for correctly parsing the depth information in the images.

[0026] Read stereoscopic view files from the specified directory manually and via scripts to ensure that stereoscopic views and image sequences can be correctly identified and loaded, and that the sequences of each pair of stereoscopic views are correctly matched. View the basic properties of the stereoscopic views, such as resolution and color space, and ensure that all stereoscopic views have the same dimensions and color representation.

[0027] Confirm whether the stereoscopic view is in RGB format. If not, it needs to be converted to RGB format for subsequent processing. In addition, record the width, height, and number of channels of each stereoscopic view (usually 3 for RGB).

[0028] Set a unified target size, width and height To ensure processing efficiency and consistency in model training, calculate the scaling ratio based on the original width and height of the stereoscopic view, as well as the target width and height, expressed as: ; where represents the scaling ratio, represents the target width, represents the original width, represents the target height, represents the original height, represents the minimum value operation; Adjust the actual size of the stereoscopic view according to the calculated scaling ratio, expressed as: ; ; where represents the new width after adjustment, represents the new height after adjustment, represents the original height, represents the original width; For each stereoscopic view to be processed, traverse all the pixels of the stereoscopic view, such as red, green, and blue, and calculate the average brightness value of the three color channels of red, green, and blue. Specifically, add the color values of red, green, and blue and finally divide by the total number of pixels in the stereoscopic view.

[0029] Based on the original color value of each pixel and the average brightness value of each color channel, the color value of each pixel is adjusted. Specifically, the new color value is obtained by multiplying the ratio of the original color value to the average brightness value of the color channel by the ideal grayscale mean, ensuring that the processed stereoscopic view has better color balance and visual effects. After completing the color correction of all pixels, save the processed stereoscopic view. In addition, download the training weight file of the super-resolution ESRGAN model from the GitHub repository, use the deep learning framework PyTorch to load the super-resolution ESRGAN model and training weights, and input the training weights into the super-resolution ESRGAN model for training. The super-resolution model ESRGAN has excellent image reconstruction and detail recovery capabilities. The resized and color-corrected stereoscopic view is normalized, that is, all pixel values ​​of the stereoscopic view are converted to the range of [0, 1] by removing 255 (the pixel values ​​of the stereoscopic view in RGB format are between [0, 255]). The normalized stereoscopic view is input into the super-resolution model ESRGAN. The early layers of the super-resolution ESRGAN model perform preliminary feature extraction (edges, textures), and the upsampling layers gradually increase the spatial resolution of the stereoscopic view. As the stereoscopic view flows in the network, the super-resolution ESRGAN model gradually constructs a high-resolution version of the stereoscopic view, thereby enhancing the expression of stereoscopic view details.

[0030] S2. Load the pre-trained Monodepth2 model to analyze the disparity of the left and right eye views in the stereoscopic view. Specifically: Install the Pillow dependency library and visit the official GitHub page of Monodepth2 to find and download the pre-trained model weight file containing the stereoscopic view dataset. Make sure that the selected Monodepth2 model has the ability to handle data types and sizes. Create a new Python script and import the torch module. Set the Monodepth2 model structure in the script and load the pre-trained weights downloaded from the official source. Then call the state dictionary of the Monodepth2 model object to load the weight file.

[0031] Using the image functions in the Pillow dependency library, open the pre - processed stereo views, namely the left - eye view and the right - eye view. Adjust each left - eye view and right - eye view to the size expected by the Monodepth2 model, such as 640x192. If the original left - eye view and right - eye view have different sizes, perform cropping and scaling operations. For each pixel value, subtract the mean [0.485, 0.456, 0.406] and divide by the standard deviation [0.229, 0.224, 0.225] for normalization, so that the input left - eye view and right - eye view conform to the distribution used during the training of the Monodepth2 model. Convert the resized and normalized left - eye view and right - eye view into the PyTorch tensor format so that they can be directly input into the Monodepth2 model for forward propagation. Input the tensors of the converted left - eye view and right - eye view into the Monodepth2 model with pre - loaded weights. At this time, it should be noted that if the Monodepth2 model is designed to support binocular input, both the left - eye view and the right - eye view need to be provided simultaneously. If it only supports monocular input, an image from either side perspective can be selected. Then call the forward propagation of the Monodepth2 model to analyze the input tensors of the left - eye view and the right - eye view. The output result is the disparity between the left - eye view and the right - eye view.

[0032] For each pixel point in the left - eye view, for example, the pixel at the position (100, 100), consider a 5x5 pixel - sized window around the pixel and search for the window in the right - eye view image that is most similar along the horizontal direction within the maximum disparity range, that is, moving up to 64 pixels to the left from the current position. Specifically, it is necessary to check each position from 0 to 64 pixels moved to the left from the current position to find the part that best matches the window in the left - eye view image. Use the sum of absolute differences to compare the current window in the left - eye view image with the corresponding window in the right - eye view image, and apply the sum of absolute values to measure the similarity between the pixels of the corresponding windows in the left - and right - eye view images, so that the disparity value with the minimum similarity is used as the best - matching disparity value for this pixel. For example, if the similarity score is the lowest when the disparity value is 15, then 15 is used as the best - matching disparity value for this pixel. Fill the best - matching disparity value of each pixel into a new two - dimensional array to form the final disparity map. The disparity map directly reflects the relative distance of each object in the scene with respect to the observer, that is, the depth information. For 3D autostereoscopic images, the depth information can ensure that the generated 3D autostereoscopic images have a realistic three - dimensional sense.

[0033] S3. Apply the principle of triangulation to calculate the depth value corresponding to each pixel to obtain the depth map. Specifically, Based on the previously obtained disparity map, create a new depth map matrix as the container for the final depth map. The size of the depth map matrix is the same as that of the disparity map and is used to store the depth values of each pixel point.

[0034] Calculate the depth value using the triangulation principle of the binocular camera, expressed as ; where represents the depth value, represents the camera focal length, represents the distance between the centers of the two cameras, represents the disparity; According to the calculated depth value, fill this depth value into the depth map matrix to obtain the depth map. Use the Matplotlib plotting tool to convert the optimized depth map into a visual heat map and pseudo-color map, which is convenient for intuitively checking the prediction quality and helps to quickly discover potential problems. When there is real depth data as a reference, the accuracy of the depth map can be quantified according to the error metrics between the predicted depth map and the real depth map, such as the root mean square error (RMSE) and the absolute relative error (Abs Rel). In addition, qualitative analysis of manually checking the edge sharpness and whether the object contours are accurate further ensures the precision of the depth map in detail processing.

[0035] S4. Input the depth map into the U-Net model and combine it with the generator for iterative training. Specifically: Use a 3x3 convolutional kernel to establish a convolutional layer and a 2x2 max pooling layer to establish an encoder. Add multiple convolutional layers as a bridging layer at the last layer of the encoder to connect the decoder constructed by transposed convolution. Finally, use a 1x1 convolutional layer as the output layer at the end of the decoder to create a U-Net model with an encoder-decoder structure. The encoder is used to extract image edge contours, textures, and color intensity features, while the decoder is responsible for generating the output image based on the image edge contours, textures, and color intensity features. An important feature of the U-Net model is the skip connection, which allows direct transfer of feature information from the encoder to the decoder at different levels. Feed the preprocessed left and right eye views and the depth map as inputs into the U-Net model. The left and right eye views and the depth map are merged into a tensor through the channel dimension. For example, for an RGB image and a single-channel disparity map, they can be merged into a 4-channel input tensor (RGB + disparity).

[0036] Add the transposed convolution of the generator to the output layer of the U-Net model. The task is to generate a realistic 3D stereoscopic image based on the input left and right eye views and the depth map. The task of the discriminator is to distinguish between real images and the images generated by the U-Net model.

[0037] Simultaneously design adversarial loss and content loss for the U-Net model. The adversarial loss encourages the generator to generate images that are difficult for the discriminator to distinguish, and the content loss ensures that the content of the generated image is consistent with the input disparity. Figure 1 The loss function of the discriminator aims to accurately distinguish between real images and images generated by the U-Net model.

[0038] Use random numbers, such as the standard normal distribution (mean = 0, standard deviation = 0.02), to initialize all the weights of the generator and the discriminator. Input a set of left and right eye views and depth maps into the U-Net model to generate a fake 3D stereoscopic image. Feed this fake 3D stereoscopic image together with a batch of real 3D stereoscopic images into the discriminator, and calculate the adversarial loss of the U-Net model, denoted as ; where represents the adversarial loss of the generator, represents the average value, represents the probability that this 3D stereoscopic image is judged to be real; According to the adversarial loss, update the weights of the U-Net model to make the generated images more realistic, denoted as ; where represents the updated weights, represents the learning rate, represents the current weights, represents the adversarial loss of the generator; Input a batch of real 3D stereoscopic images and the fake 3D stereoscopic images generated by the U-Net model into the discriminator, and calculate the loss of the discriminator, aiming to maximize the probability of correctly classifying real 3D stereoscopic images and fake 3D stereoscopic images, denoted as ; where represents the total loss of the discriminator, represents the batch size, represents the number of real 3D stereoscopic images, represents the probability that a 3D stereoscopic image is judged to be real, represents the fake 3D stereoscopic image; When the total loss of the discriminator is close to a score of 1, it is a real 3D stereoscopic image. When the total loss of the discriminator is close to 0, it is a fake 3D stereoscopic image.

[0039] Repeat the above process of updating the weights of the U-Net model and the discriminator's ability to distinguish 3D stereoscopic images, and alternately train the discriminator and the U-Net model according to a certain ratio, such as 1:1.

[0040] During this training process, in each iteration, the U-Net model creates 3D stereoscopic images that are increasingly difficult to be recognized as "forged", while also enhancing the discriminator's ability to distinguish between real and fake images. As the number of iterations increases, the U-Net model gradually learns how to generate highly realistic 3D stereoscopic images, and it becomes increasingly difficult for the discriminator to tell whether these 3D stereoscopic images are generated by the U-Net model.

[0041] S5. After the U-Net model finishes iterative training repeatedly, new left and right eye views and depth maps are obtained. After applying dynamic region enhancement, they are input into the trained U-Net model to generate 3D stereoscopic images. Specifically, When the U-Net model reaches the preset total number of iterations, at this time, the U-Net model can generate highly realistic images, and the discriminator can no longer effectively distinguish which are real images and which are images generated by the U-Net model, indicating that the training of the U-Net model has achieved the expected goal, and the discriminator has lost its function.

[0042] After the discriminator fails, the trained U-Net model is loaded according to the previous training framework PyTorch. New left and right eye views and depth maps are extracted from the stereoscopic view dataset and preprocessed to ensure that the sizes of the left and right eye views and depth maps are the same and meet the input requirements of the generator. In addition, a CNN model is constructed using convolutional layers, pooling layers, and fully connected layers. The CNN model is initialized by Xavier to keep the input-output variance consistent. At the same time, for the task of generating an attention weight map, a cross-entropy loss function is added. The depth map is input into the CNN model, and the CNN model outputs an attention weight map that is the same size as the depth map and has a single channel. The three channels of the left eye view and the three channels of the right eye view in the stereoscopic view are concatenated with the single-channel depth map and attention weight map to form a multi-channel input, so that the U-Net model can pay special attention to the edge and boundary regions and texture-rich regions when generating 3D stereoscopic images, and adjust the application intensity of the attention weight map in real time according to the content shown in each 3D stereoscopic image. For example, the attention weight is increased in the edge and texture complex regions, and appropriately reduced in the background and color-uniform regions. This dynamic adjustment can not only enhance the expressiveness of the edge and boundary regions and texture-rich regions, but also optimize the use efficiency of computing resources.

[0043] When multi-channel information (RGB image + depth map + attention weight map) is input into the trained U-Net model, the encoder part extracts various features including structural features, contour features, and dynamic object type features in the edge and boundary regions and texture-rich regions. The decoder then generates the final high-quality 3D stereoscopic images based on the structural features, contour features, and dynamic object type features.

[0044] Through dynamic region enhancement technology, especially by emphasizing the edge and boundary regions as well as the texture-rich regions, the realism of 3D naked-eye images is enhanced. This method can capture and reproduce the details of the original scene more accurately. Especially when displaying objects with complex geometries and textures, the advantages of 3D naked-eye images are particularly obvious, enabling observers to more easily perceive the depth of field and spatial relationships, increasing the three-dimensional effect and sense of reality of 3D naked-eye images.

[0045] S6. Prune and optimize the U-Net model for generating 3D naked-eye images so that the 3D naked-eye images can run smoothly with low energy consumption on mobile devices. Specifically: According to the requirements of the U-Net model, determine the desired sparsity level. Sparsity refers to the proportion of weights in the U-Net model that are set to zero. For example, if it is desired that 70% of the weights in the final U-Net model become zero, the final sparsity is set to 70%.

[0046] Determine the start and end of the pruning process, usually based on the number of training epochs. For example, if you plan to start pruning at the 2000th step and continue until the 4000th step during the training process, you need to clarify these two time points of starting and ending pruning. The initial sparsity refers to the sparsity degree of the U-Net model at the start of pruning, while the final sparsity is the target sparsity that the U-Net model should reach after pruning. For example, the initial sparsity can be set to 30%, which means that at the start of pruning, 30% of the weights in the U-Net model will be removed, and the final sparsity is set to 70%, indicating that after pruning, 70% of the weights in the U-Net model will be set to zero.

[0047] Record the above pruning-related configuration parameters, including the target sparsity, initial sparsity, final sparsity, start step, and end step. At the same time, input the target sparsity, initial sparsity, final sparsity, start step, and end step into the U-Net model. This process will gradually remove unimportant weight connections according to the target sparsity. Note that pruning is carried out step by step, not all at once. This means that between the start step and the end step, the sparsity of the U-Net model will gradually increase until it reaches the final sparsity. If it is found that the pruning speed is too fast or too slow, the pruning rate can be adjusted according to the actual situation. For example, if it is found that the performance of the U-Net model deteriorates rapidly during pruning, the pruning speed can be appropriately slowed down and the final sparsity can be reduced. After the pruning process is completed, save the pruned U-Net model. The pruned model is smaller in size and occupies less memory accordingly, which means that more 3D naked-eye images can run simultaneously without worrying about memory shortage problems, and at the same time, it reduces the risk of image crashes caused by memory overflow.

[0048] Use the TensorFlow Lite converter tool to convert the pruned U-Net model into the TensorFlow Lite format (.tflite file), integrate the converted.tflite file into Android and iOS applications, input the 3D naked-eye image into the U-Net model in the.tflite file, and increase and decrease the pruning intensity according to the speed and memory occupancy when the 3D naked-eye image runs through the U-Net model in Android and iOS applications to find the best performance and energy consumption ratio. According to the performance in actual applications, developers can flexibly adjust the pruning intensity to find the best balance between performance and energy consumption. For example, in tasks that require higher precision, a lower sparsity can be selected; while in the pursuit of higher efficiency, a certain degree of accuracy loss can be accepted in exchange for faster speed and lower energy consumption. The 3D naked-eye image generation function that runs efficiently on mobile devices can provide users with a more smooth and natural interaction experience. Whether for entertainment or professional applications, the 3D naked-eye image generation technology can let users feel the convenience and fun brought by technological progress.

[0049] This embodiment also provides a 3D naked-eye image generation system based on a neural network, including: A data preprocessing module that extracts stereoscopic views from a dataset containing stereoscopic views and performs preprocessing operations such as resizing and color correction on the stereoscopic views; A depth map module that uses the Monodepth2 model to analyze the left-eye view and right-eye view in the stereoscopic view to obtain the disparity, creates a depth map matrix, calculates the depth value using the triangulation principle of a binocular camera, and fills the depth value into the depth map matrix to obtain the depth map; A U-Net model training module that inputs the depth map and the preprocessed left-eye view and right-eye view Figure 1 into the U-Net model together and performs iterative training in combination with a generator until the discriminator can no longer distinguish the image generated by the U-Net model from the real image; An image generation module that, after the U-Net model training is completed, inputs the newly obtained left-eye and right-eye views and depth map and uses the dynamic region enhancement technology into the trained U-Net model to generate high-quality 3D naked-eye images; A U-Net model optimization module that prunes and optimizes the U-Net model according to pruning parameters, target sparsity, initial sparsity, final sparsity, start step, and end step.

[0050] This embodiment also provides a computer device, which is applicable to the case of the 3D naked-eye image generation method based on a neural network, and includes: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the 3D naked-eye image generation method based on a neural network as proposed in the above embodiment.

[0051] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0052] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the 3D naked-eye image generation method based on a neural network as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read Only Memory (EPROM for short), Programmable Red-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0053] In summary, the present invention optimizes the overall process from data preprocessing to the generation of high-quality 3D naked-eye images, enhancing the comprehensive performance and practicality of the system. Overall, it jointly constructs an efficient, accurate, and user-friendly 3D naked-eye image generation mechanism. Among them, the dynamic region enhancement technology not only provides higher depth estimation accuracy and consistency in complex scene reconstruction but also greatly enhances the realism and immersive experience of 3D naked-eye images. By using advanced disparity optimization, attention mechanism, and generative adversarial network technology, it can particularly focus on and optimize the expressiveness of 3D naked-eye images while ensuring the effective utilization of computing resources. In addition, the standardized data preprocessing steps lay a solid foundation for subsequent processing, making the entire process more robust, adaptable, and capable of quickly identifying and correcting potential problems, and also bringing an unprecedented visual experience to users, promoting technological progress and application expansion in related fields.

[0054] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for generating 3D naked eye images based on a neural network, characterized in that: include, Acquire a stereoscopic view data set and preprocess the stereoscopic view; Load the pre-trained Monodepth2 model to analyze the disparity between the left and right eye views in the stereoscopic view, and apply the triangulation principle to calculate the depth value corresponding to each pixel to obtain a depth map; Input the depth map into the U-Net model and combine it with the generator for repeated iterative training; After the U-Net model is repeatedly trained, the new left and right eye views and depth maps are obtained and dynamic area enhancement is applied. After that, they are input into the trained U-Net model to generate a 3D naked eye image. The U-Net model for generating 3D naked-eye images is pruned and optimized, so that 3D naked-eye images can run smoothly with low energy consumption on mobile devices.

2. The method for generating 3D naked-eye images based on a neural network according to claim 1, wherein: The step of obtaining a stereoscopic image data set and preprocessing the stereoscopic image is as follows: The stereoscopic view includes a left-eye view and a right-eye view; The preprocessing includes size adjustment and color correction.

3. The method for generating 3D naked-eye images based on a neural network as claimed in claim 2, characterized in that: The pre-trained Monodepth2 model is loaded to analyze the disparity between the left eye view and the right eye view in the stereoscopic view, specifically, Convert the pre-processed left-eye and right-eye views into tensor format and adjust them to the size required by the Monodepth2 model; Input the format-converted and size-adjusted left-eye and right-eye view tensors into the Monodepth2 model for forward propagation to obtain the disparity; The disparity value is calculated by applying the sum of absolute values ​​of the pixel windows corresponding to the left eye view and the right eye view, and the disparity value similarity is compared; A two-dimensional array of disparity maps is established, and the disparity values ​​with the smallest similarity are used to fill the two-dimensional array to obtain the disparity map.

4. The method for generating 3D naked eye images based on a neural network as claimed in claim 3, characterized in that: The depth map is obtained by calculating the depth value corresponding to each pixel using the triangulation principle, specifically, According to the size of the disparity map, use NumPy to create a depth map matrix as a container for the final depth map to store depth values; The parallax is calculated using the triangulation principle of the binocular camera to calculate the depth value, and then the depth map matrix is ​​filled in to obtain the depth map; Use Matplotlib drawing tools to check the edge sharpness of the depth map and the accuracy of object outlines.

5. The method for generating 3D naked-eye images based on a neural network as claimed in claim 4, characterized in that: The depth map is input into the U-Net model and combined with the generator for repeated iterative training, specifically: The encoder is constructed using 3x3 convolution kernels and 2x2 maximum pooling layers, and the decoder is built through transposed convolution to obtain the U-Net model; Initialize all weights of the generator and discriminator using standard normal distribution, and add the initialized generator transposed convolution to the output layer of the U-Net model; The depth map is input into the U-Net model, and adversarial loss, content loss, and discriminator loss are designed for repeated iterative training.

6. The method for generating 3D naked eye images based on a neural network as claimed in claim 5, characterized in that: The generating of the 3D naked eye image specifically comprises: After repeated iterative training is completed, new left and right eye views and depth maps are obtained from the stereoscopic view dataset; Build an additional CNN model using convolutional layers, pooling layers, and fully connected layers, initialize the CNN model using Xavier, and train the CNN model using the cross entropy loss function; Apply dynamic enhancement technology to input the depth map into the trained CNN model to generate an attention weight map; Generate 3D naked-eye images using the newly acquired left and right eye views, depth map, and attention weight map input into the U-Net model.

7. The method for generating 3D naked-eye images based on a neural network as claimed in claim 6, characterized in that: The U-Net model for generating 3D naked-eye images is pruned and optimized so that the 3D naked-eye images can be efficiently run on mobile devices, specifically, The proportion of U-Net model weights that are set to zero; Set the start and end time of pruning according to the set number of pruning steps, and record all relevant pruning configuration parameters to guide the pruning process; The pruning configuration parameters include expected sparsity, initial sparsity at the beginning of pruning, final sparsity after pruning is completed, number of pruning start steps and number of pruning end steps; According to the desired sparsity level, the start and end times of pruning are input into the U-Net model for step-by-step pruning; Use the TensorFlow Lite converter tool to convert the pruned U-Net model to tflite format Integrate the converted tflite file into Android and iOS applications and input 3D naked eye images.

8. A 3D naked eye image generation system based on a neural network, based on the 3D naked eye image generation method based on a neural network according to any one of claims 1 to 7, characterized in that: include, A data preprocessing module extracts a stereoscopic image from a data set containing the stereoscopic image and performs preprocessing operations of resizing and color correction on the stereoscopic image; The depth map module uses the Monodepth2 model to analyze the left-eye view and the right-eye view in the stereo view to obtain the disparity, creates a depth map matrix, calculates the depth value using the triangulation principle of the binocular camera, and fills the depth value into the depth map matrix to obtain the depth map; The U-Net model training module inputs the depth map and the pre-processed left-eye and right-eye views into the U-Net model, and performs iterative training with the generator until the discriminator cannot distinguish between the U-Net model generated images and the real images; Image generation module: After the U-Net model training is completed, the newly acquired left and right eye views and depth maps are input into the trained U-Net model using dynamic area enhancement technology to generate high-quality 3D naked eye images; The U-Net model optimization module performs pruning optimization on the U-Net model according to the pruning parameters, target sparsity, initial sparsity, final sparsity, starting steps and ending steps.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for generating a 3D naked-eye image based on a neural network according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for generating a 3D naked-eye image based on a neural network according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Image conversion method and device, depth map prediction method and device, model training method and device and electronic equipment

    CN110111244A

  • Electronic component three-dimensional reconstruction method based on semi-supervised learning

    CN113902807A

  • Building three-dimensional model reconstruction method based on improved MVSNet and reconstruction system thereof

    CN119107426A

  • 3D image generation method and device, equipment and storage medium

    CN119299645A

  • Computer-implemented method and associated device for modelling a joint of a patient

    WO2024213641A1

Cited By

  • XR head-mounted display video perspective method and system based on neural network

    CN120580390A

  • A neural network-based XR head-mounted display video perspective method and system

    CN120580390B