A 3D naked-eye image generation method and system based on a neural network
The method addresses the challenge of balancing image quality and real-time performance in 3D naked-eye image generation by preprocessing data, using Monodepth2 models, and optimizing U-Net models for mobile devices, resulting in efficient and immersive 3D image generation.
Patent Information
- Application Number
- CN202510563421.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Existing 3D naked-eye image generation technologies struggle to balance image quality and real-time performance, particularly when processing fast-moving objects, due to high computational complexity that leads to lag and tearing.
A method involving data preprocessing, using Monodepth2 models to analyze stereo views, constructing U-Net models with specific convolutional layers, and applying pruning optimization to enhance performance on mobile devices.
The method improves the efficiency and accuracy of 3D naked-eye image generation, ensuring high-quality, real-time performance and immersive experiences by optimizing computational resources and enhancing depth estimation precision.
Smart Images

Figure CN120091122B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of 3D naked-eye image generation, and particularly to a 3D naked-eye image generation method and system based on a neural network. Background Art
[0002] With the rapid development of information technology, as a key component of the human-computer interaction interface, 3D image display technology has gradually become the focus of attention in the scientific research and industrial fields. Since the early 20th century, 3D imaging technology has gone through multiple stages of development, from stereoscopic vision, holographic imaging to modern computational imaging. In recent years, with the rise of deep learning, especially neural network technology, 3D image generation methods based on deep learning have received extensive attention due to their high efficiency and high quality.
[0003] However, the existing 3D naked-eye image generation technology still has some deficiencies. Especially in the aspect of naked-eye 3D display, the 3D naked-eye image generation technology solutions often have difficulty in simultaneously considering image quality and real-time performance. When processing fast-moving objects, due to the high computational complexity, the existing algorithms cannot update the view in time, resulting in phenomena such as screen tearing or delay. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a 3D naked-eye image generation method based on a neural network to solve the problem that the 3D naked-eye image cannot be updated in time due to high computational complexity.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a 3D naked-eye image generation method based on a neural network, which includes obtaining a stereoscopic view dataset and preprocessing the stereoscopic view;
[0008] Loading a pre-trained Monodepth2 model to analyze the disparity between the left and right eye views in the stereoscopic view, and applying the principle of triangulation to calculate the depth value corresponding to each pixel to obtain a depth map;
[0009] Constructing an encoder using a 3x3 convolutional kernel and a 2x2 max pooling layer, and establishing a decoder through transposed convolution, and performing iterative training in combination with the depth map until the discriminator cannot distinguish the image generated by the U-Net model from the real image;
[0010] After the iterative training is completed, obtaining new left and right eye views, depth map and attention weight map and inputting them into the U-Net model to generate a 3D naked-eye image;
[0011] Prune and optimize the U-Net model for generating 3D naked-eye images, enabling the 3D naked-eye images to run smoothly on mobile devices with low energy consumption.
[0012] As a preferred embodiment of the 3D naked-eye image generation method based on neural network according to the present invention, wherein: the obtaining of the stereoscopic view dataset and the preprocessing of the stereoscopic view are specifically as follows:
[0013] The stereoscopic view includes a left-eye view and a right-eye view;
[0014] The preprocessing includes size adjustment and color correction.
[0015] As a preferred embodiment of the 3D naked-eye image generation method based on neural network according to the present invention, wherein: the loading of the pre-trained Monodepth2 model to analyze the disparity between the left-eye view and the right-eye view in the stereoscopic view is specifically as follows,
[0016] Convert the preprocessed left-eye view and right-eye view into tensor format and adjust them to the size required by the Monodepth2 model;
[0017] Input the tensors of the preprocessed left-eye view and right-eye view after format conversion and size adjustment into the Monodepth2 model for forward propagation to obtain the disparity;
[0018] Calculate the disparity value by applying the sum of absolute values to the corresponding pixel windows of the left-eye view and the right-eye view, and compare the similarity of the disparity values;
[0019] Establish a disparity Figure 2 dimensional array, and fill the two-dimensional array with the disparity value with the smallest similarity to obtain the disparity map.
[0020] As a preferred embodiment of the 3D naked-eye image generation method based on neural network according to the present invention, wherein: the application of the triangulation principle to calculate the depth value corresponding to each pixel to obtain the depth map is specifically as follows,
[0021] According to the size of the disparity map, use NumPy to create a depth map matrix as a container for the final depth map to store the depth values;
[0022] Utilize the triangulation principle of the binocular camera to calculate the depth value from the disparity and fill it into the depth map matrix to obtain the depth map;
[0023] Apply the Matplotlib plotting tool to check the edge sharpness of the depth map and the accuracy of the object contour.
[0024] As a preferred solution of the 3D naked-eye image generation method based on neural network according to the present invention, wherein: an encoder is constructed by using a 3x3 convolution kernel and a 2x2 max pooling layer, and a decoder is established by transposed convolution, and iterative training is carried out in combination with a depth map until the discriminator cannot distinguish the image generated by the U-Net model from the real image, specifically:
[0025] An encoder is constructed by using a 3x3 convolution kernel and a 2x2 max pooling layer, and a decoder is established by transposed convolution to obtain a U-Net model;
[0026] Initialize all the weights of the generator and the discriminator using the standard normal distribution, and add the initialized transposed convolution of the generator to the output layer of the U-Net model;
[0027] Input the depth map into the U-Net model, and design adversarial loss, content loss, and discriminator loss for iterative training.
[0028] As a preferred solution of the 3D naked-eye image generation method based on neural network according to the present invention, wherein: the generation of the 3D naked-eye image is specifically as follows
[0029] After the iterative training is completed, new left and right eye views and depth maps are obtained from the stereoscopic view dataset;
[0030] Construct an additional CNN model using convolutional layers, pooling layers, and fully connected layers, and initialize the CNN model by Xavier and train the CNN model using the cross-entropy loss function;
[0031] Apply the dynamic enhancement technology to input the depth map into the trained CNN model to generate an attention weight map;
[0032] Input the newly obtained left and right eye views, depth map, and attention weight map into the U-Net model to generate a 3D naked-eye image.
[0033] As a preferred solution of the 3D naked-eye image generation method based on neural network according to the present invention, wherein: pruning and optimizing the U-Net model for generating the 3D naked-eye image to enable the 3D naked-eye image to run efficiently on a mobile device, specifically:
[0034] Set the weights of the U-Net model to zero;
[0035] Set the start and end times of pruning according to the set pruning steps, and record all relevant pruning configuration parameters to guide the pruning process;
[0036] The pruning configuration parameters include expected sparsity, initial sparsity at the start of pruning, final sparsity after pruning, start step of pruning, and end step of pruning;
[0037] According to the expected sparsity level, input the start and end times of pruning into the U-Net model for progressive pruning;
[0038] Use the TensorFlow Lite converter tool to convert the pruned U-Net model into the tflite format
[0039] Integrate the converted tflite file into Android and iOS applications and input 3D stereoscopic images.
[0040] In a second aspect, the present invention provides a 3D stereoscopic image generation system based on a neural network, including,
[0041] A data preprocessing module extracts stereoscopic views from a dataset containing stereoscopic views and performs preprocessing operations of resizing and color correction on the stereoscopic views;
[0042] A depth map module analyzes the left-eye view and right-eye view in the stereoscopic view using the Monodepth2 model to obtain disparities, creates a depth map matrix, calculates depth values using the triangulation principle of a binocular camera, and fills the depth values into the depth map matrix to obtain a depth map;
[0043] A U-Net model training module inputs the depth map and the preprocessed left-eye view and right-eye view Figure 1 into the U-Net model together and performs iterative training in combination with a generator until the discriminator can no longer distinguish between the images generated by the U-Net model and real images;
[0044] An image generation module, after the U-Net model training is completed, inputs the newly obtained left-eye and right-eye views and depth map and uses dynamic region enhancement technology and then inputs them into the trained U-Net model to generate high-quality 3D stereoscopic images;
[0045] A U-Net model optimization module prunes and optimizes the U-Net model according to pruning parameters, the target sparsity, the initial sparsity, the final sparsity, the start step number, and the end step number.
[0046] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the 3D stereoscopic image generation method based on a neural network as described in the first aspect of the present invention is implemented.
[0047] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by a processor, any step of the method for generating a 3D naked-eye image based on a neural network as described in the first aspect of the present invention is implemented.
[0048] The beneficial effects of the present invention are as follows: Through the overall process optimization from data preprocessing to the generation of high-quality 3D naked-eye images, the comprehensive performance and practicality of the system are improved. Overall, an efficient, accurate and user-friendly 3D naked-eye image generation mechanism is jointly constructed. Among them, the dynamic area enhancement technology not only provides higher depth estimation accuracy and consistency in complex scene reconstruction, but also greatly enhances the realism and immersive experience of 3D naked-eye images. By using advanced parallax optimization, attention mechanism and generative adversarial network technology, the expressiveness of 3D naked-eye images can be particularly focused on and optimized, while ensuring the effective utilization of computing resources. In addition, the standardized data preprocessing steps lay a solid foundation for subsequent processing, making the entire process more robust, adaptable, and capable of quickly identifying and correcting potential problems, and also bringing an unprecedented visual experience to users, promoting the technological progress and application expansion in related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a flowchart of the method for generating a 3D naked-eye image based on a neural network in Embodiment 1.
[0051] Figure 2 It is a schematic diagram of the optimization of the mobile device in Embodiment 1.
[0052] Figure 3 It is a schematic diagram of depth map estimation and parallax calculation in Embodiment 1.
[0053] Figure 4 It is a schematic diagram of training the U-Net model architecture in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.
[0055] In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0056] Secondly, as used herein, an "embodiment" or "embodiments" refers to specific features, structures, or characteristics that may be included in at least one implementation of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other.
[0057] Embodiment 1, referring to Figures 1 to 4 , this embodiment provides a method for generating 3D naked-eye images based on a neural network, including the following steps:
[0058] S1. Obtain a stereoscopic view dataset and preprocess the stereoscopic views. Specifically:
[0059] Extract stereoscopic views from a publicly available dataset containing stereoscopic views. This dataset usually contains multiple compressed packages, such as.zip files, which need to be decompressed to a specified directory on the local hard drive. Each sequence contains multiple subfolders, including left-eye perspective images and right-eye perspective images. Confirm that each image sequence is also accompanied by a corresponding calibration file, usually in.txt format. The calibration file contains the internal parameter matrix of the camera, distortion coefficients, the external parameter matrix of the calibration information, and the reprojection error. These calibration files are crucial for correctly parsing the depth information in the images.
[0060] Read the stereoscopic view files from the specified directory manually and by script to ensure that the stereoscopic views and image sequences can be correctly identified and loaded, and the sequences of each pair of stereoscopic views are correctly matched. View the basic properties of the stereoscopic views, such as resolution and color space, and all stereoscopic views have the same size and color representation.
[0061] Confirm whether the stereoscopic views are in RGB format. If not, they need to be converted to RGB format for subsequent processing. In addition, record the width, height, and number of channels (usually 3 for RGB) of each stereoscopic view.
[0062] Set a unified target size. The width and height ensure the consistency of processing efficiency and model training. Calculate the scaling ratio based on the original width and height of the stereoscopic views and the target width and height, expressed as:
[0063] ;
[0064] Wherein, represents the scaling ratio. Represents the target width, Represents the original width, Represents the target height, Represents the original height, Represents the operation of taking the minimum value;
[0065] Adjust the actual size of the stereoscopic view according to the calculated scaling ratio, expressed as:
[0066] ;
[0067] ;
[0068] Among them, Represents the new width after adjustment, Represents the new height after adjustment, Represents the original height, Represents the original width;
[0069] For each stereoscopic view to be processed, traverse all the pixels of the stereoscopic view, such as red, green, and blue, and at the same time calculate the average brightness value of the three color channels of red, green, and blue. Specifically, add the color values of red, green, and blue and finally divide by the total number of pixels of the stereoscopic view.
[0070] Based on the original color value of each pixel and the average brightness value of each color channel, adjust the color value of each pixel. Specifically, obtain the new color value by multiplying the ratio of the original color value to the average brightness value of this color channel by the ideal gray mean value to ensure that the processed stereoscopic view has better color balance and visual effect. After completing the color correction of all pixels, save the processed stereoscopic view. In addition, download the training weight file of the super-resolution ESRGAN model from the GitHub repository, use the deep learning framework PyTorch to load the super-resolution ESRGAN model and the training weights, and input the training weights into the super-resolution ESRGAN model for training. The super-resolution model ESRGAN has excellent image reconstruction ability and detail restoration ability. Normalize the stereoscopic view after size adjustment and color correction, that is, divide all pixel values of the stereoscopic view by 255 and convert them to the range of [0, 1] (the pixel values of the stereoscopic view in RGB format are between [0, 255]), and input the normalized stereoscopic view into the super-resolution model ESRGAN. The early layers of the super-resolution ESRGAN model perform preliminary feature extraction (edges, textures), and the upsampling layer gradually increases the spatial resolution of the stereoscopic view. As the stereoscopic view flows through the network, the super-resolution ESRGAN model gradually constructs a high-resolution version of the stereoscopic view, thereby enhancing the detail expressiveness of the stereoscopic view.
[0071] S2. Load the pre-trained Monodepth2 model to analyze the disparity between the left-eye view and the right-eye view in the stereo view. Specifically:
[0072] Install the Pillow dependency library, visit the official GitHub page of Monodepth2, find and download the pre-trained model weight file containing the stereo view dataset, ensure that the selected Monodepth2 model has the ability to handle data types and dimensions, create a new Python script, import the torch module, set the Monodepth2 model structure in the script, load the pre-trained weights downloaded from the official source, and then call the state dictionary of the Monodepth2 model object to load the weight file.
[0073] Use the image function in the Pillow dependency library to open the pre-processed stereo views, namely the left-eye view and the right-eye view, and resize each left-eye view and right-eye view to the size expected by the Monodepth2 model, such as 640x192. If the original left-eye view and right-eye view have different sizes, perform cropping and scaling operations. For each pixel value, subtract the mean [0.485, 0.456, 0.406] and divide by the standard deviation [0.229, 0.224, 0.225] for normalization to make the input left-eye view and right-eye view conform to the distribution used during the training of the Monodepth2 model. Convert the resized and normalized left-eye view and right-eye view to the PyTorch tensor format so that they can be directly input into the Monodepth2 model for forward propagation. Input the tensors of the converted left-eye view and right-eye view into the Monodepth2 model that has already loaded the weights. At this time, it should be noted that if the Monodepth2 model is designed to support binocular input, both the left-eye view and the right-eye view need to be provided simultaneously. If it only supports monocular input, an image from either perspective can be selected. Then call the forward propagation of the Monodepth2 model to analyze the input tensors of the left-eye view and the right-eye view, and the output result is the disparity between the left-eye view and the right-eye view.
[0074] For each pixel point in the left-eye view, for example, the pixel at the position (100, 100), consider a 5x5 pixel-sized window around the pixel and search for the window in the right-eye view image that is most similar to the window in the left-eye view image along the horizontal direction within the maximum parallax range, that is, moving up to 64 pixels to the left from the current position. Specifically, it is necessary to check each position from moving 0 to 64 pixels to the left from the current position to find the part that best matches the window in the left-eye view image. Use the sum of absolute differences to compare the current window in the left-eye view image with the corresponding window in the right-eye view image, and apply the sum of absolute values to measure the similarity between the pixels of the corresponding windows in the left and right-eye view images, so that the parallax value with the minimum similarity is used as the best-matching parallax value for this pixel. For example, if the similarity score is the lowest when the parallax value is 15, then 15 is used as the best-matching parallax value for this pixel. Fill the best-matching parallax value of each pixel into a new two-dimensional array to form the final parallax map. The parallax map directly reflects the relative distance of each object in the scene relative to the observer, that is, the depth information. For 3D autostereoscopic images, the depth information can ensure that the generated 3D autostereoscopic images have a real three-dimensional sense.
[0075] S3. Apply the principle of triangulation to calculate the depth value corresponding to each pixel to obtain a depth map. Specifically,
[0076] According to the previously obtained parallax map, create a new depth map matrix as the container for the final depth map. The size of the depth map matrix is the same as that of the parallax map, and it is used to store the depth value of each pixel point.
[0077] Use the principle of triangulation of binocular cameras to calculate the depth value, expressed as,
[0078] ;
[0079] Among them, represents the depth value, represents the camera focal length, represents the distance between the centers of the two cameras, represents the parallax;
[0080] According to the calculated depth value, fill this depth value into the depth map matrix to obtain the depth map. Use the Matplotlib plotting tool to convert the optimized depth map into a visualized heat map and a pseudo-color map, which is convenient for visually checking the prediction quality and helps to quickly discover potential problems. When there is real depth data as a reference, the accuracy of the depth map can be quantified according to the error metrics between the predicted depth map and the real depth map, such as the root mean square error (RMSE) and the absolute relative error (Abs Rel). In addition, a qualitative analysis of manually checking the edge sharpness and whether the object contours are accurate further ensures the precision of the depth map in detail processing.
[0081] S4. Construct an encoder using a 3x3 convolutional kernel and a 2x2 max pooling layer, and build a decoder through transposed convolution. Combine with the depth map for iterative training until the discriminator can no longer distinguish the images generated by the U-Net model from the real images. Specifically:
[0082] Use a 3x3 convolutional kernel to build a convolutional layer and a 2x2 max pooling layer to build an encoder. Add multiple convolutional layers as a bridging layer to the last layer of the encoder to connect to the decoder constructed through transposed convolution. Finally, use a 1x1 convolutional layer as the output layer at the end of the decoder to create a U-Net model with an encoder-decoder structure. The encoder is used to extract image edge contours, textures, and color intensity features, while the decoder is responsible for generating the output image based on the image edge contours, textures, and color intensity features. An important feature of the U-Net model is the skip connection, which allows feature information to be directly passed from the encoder to the decoder at different levels. And feed the preprocessed left and right eye views and the depth map as inputs into the U-Net model. The left and right eye views and the depth map are merged into a tensor through the channel dimension. For example, for an RGB image and a single-channel disparity map, they can be merged into a 4-channel input tensor (RGB + disparity).
[0083] Add the transposed convolution of the generator to the output layer of the U-Net model. The task is to generate a realistic 3D stereoscopic image based on the input left and right eye views and the depth map. The task of the discriminator is to distinguish between real images and the images generated by the U-Net model.
[0084] At the same time, design adversarial loss and content loss for the U-Net model. The adversarial loss encourages the generator to generate images that are difficult to be distinguished by the discriminator. The content loss ensures that the content of the generated image is consistent with the input disparity Figure 1 consistent. And the loss function of the discriminator aims to accurately distinguish between real images and the images generated by the U-Net model.
[0085] Use random numbers, such as a standard normal distribution (mean = 0, standard deviation = 0.02), to initialize all the weights of the generator and the discriminator. Input a set of left and right eye views and depth maps into the U-Net model to generate a fake 3D stereoscopic image. Feed this fake 3D stereoscopic image together with a batch of real 3D stereoscopic images into the discriminator, and calculate the adversarial loss of the U-Net model, denoted as,
[0086] ;
[0087] where, denotes the adversarial loss of the generator, denotes the average value, denotes the probability that this 3D stereoscopic image is judged to be real;
[0088] According to the adversarial loss, update the weights of the U-Net model to make the generated images more realistic, denoted as,
[0089] ;
[0090] where, denotes the updated weights, denotes the learning rate, denotes the current weights, denotes the adversarial loss of the generator;
[0091] Input a batch of real 3D stereoscopic images and the fake 3D stereoscopic images generated by the U-Net model into the discriminator, and calculate the loss of the discriminator, aiming to maximize the probability of correctly classifying real 3D stereoscopic images and fake 3D stereoscopic images, denoted as,
[0092] ;
[0093] where, denotes the total loss of the discriminator, denotes the batch size, denotes the number of real 3D stereoscopic images, denotes the probability that a 3D stereoscopic image is judged to be real, denotes a fake 3D stereoscopic image;
[0094] When the total loss of the discriminator is close to a score of 1, it is a real 3D stereoscopic image, and when the total loss of the discriminator is close to 0, it is a fake 3D stereoscopic image.
[0095] Repeat the above process of updating the weights of the U-Net model and the discriminator's ability to distinguish 3D stereoscopic images, and alternately train the discriminator and the U-Net model according to a certain ratio, such as 1:1.
[0096] During this training process, in each iteration, the U-Net model is made to produce 3D naked-eye images that are more difficult to be recognized as "forged", while at the same time, the discriminator's ability to distinguish between real and fake images is enhanced. As the number of iterations increases, the U-Net model gradually learns how to generate highly realistic 3D naked-eye images, and it becomes increasingly difficult for the discriminator to tell whether these 3D naked-eye images are produced by the U-Net model.
[0097] S5. After the end of repeated iterative training, new left and right eye views, depth maps, and attention weight maps are obtained and input into the U-Net model to generate 3D naked-eye images. Specifically,
[0098] When the U-Net model reaches the preset total number of iterations, at this time, the U-Net model can generate highly realistic images, and the discriminator can no longer effectively distinguish which are real images and which are images generated by the U-Net model, indicating that the training of the U-Net model has achieved the expected goal, and the discriminator has lost its function.
[0099] After the discriminator fails, the trained U-Net model is loaded according to the previous training framework PyTorch. New left and right eye views and depth maps are extracted from the stereo view dataset and preprocessed to ensure that the sizes of the left and right eye views and the depth map are consistent and meet the input requirements of the generator. In addition, a CNN model is constructed using convolutional layers, pooling layers, and fully connected layers. The CNN model is initialized by Xavier to keep the input-output variance of the CNN model consistent. At the same time, for the task of generating attention weight maps, a cross-entropy loss function is added. The depth map is input into the CNN model, and the CNN model outputs an attention weight map with the same size as the depth map and a single channel. The three channels of the left eye view and the three channels of the right eye view in the stereo view are concatenated with the single-channel depth map and attention weight map to form a multi-channel input, so that the U-Net model can pay special attention to the edge and boundary regions and the texture-rich regions when generating 3D naked-eye images, and adjust the application intensity of the attention weight map in real time according to the content shown in each 3D naked-eye image. For example, the attention weight is increased in the edge and texture complex regions, and the weight is appropriately reduced in the background and color-uniform regions. This dynamic adjustment can not only enhance the expressiveness of the edge and boundary regions and the texture-rich regions, but also optimize the use efficiency of computing resources.
[0100] When multi-channel information (RGB image + depth map + attention weight map) is input into the trained U-Net model, the encoder part will extract various features including structural features, contour features, and dynamic object type features in the edge and boundary regions and the texture-rich regions, and the decoder will generate the final high-quality 3D naked-eye images based on the structural features, contour features, and dynamic object type features.
[0101] Through dynamic region enhancement technology, especially the emphasis on edge and boundary regions as well as texture-rich regions, the realism of 3D naked-eye images is enhanced. This method can capture and reproduce the details of the original scene more accurately. Especially when displaying objects with complex geometries and textures, the advantages of 3D naked-eye images are particularly obvious, enabling observers to more easily perceive the depth of field and spatial relationships, increasing the three-dimensional effect and sense of reality of 3D naked-eye images.
[0102] S6. Prune and optimize the U-Net model for generating 3D naked-eye images so that the 3D naked-eye images can run smoothly on mobile devices with low energy consumption. Specifically:
[0103] According to the requirements of the U-Net model, determine the desired sparsity level. Sparsity refers to the proportion of weights in the U-Net model that are set to zero. For example, if it is desired that 70% of the weights in the final U-Net model become zero, the final sparsity is set to 70%.
[0104] Determine the start and end of the pruning process, usually based on the number of training epochs. For example, if you plan to start pruning at the 2000th step and continue until the 4000th step during the training process, you need to clarify these two time points for starting and ending pruning. The initial sparsity refers to the sparsity level of the U-Net model at the start of pruning, while the final sparsity is the target sparsity that the U-Net model should reach after pruning. For example, the initial sparsity can be set to 30%, which means that at the start of pruning, 30% of the weights in the U-Net model will be removed. If the final sparsity is set to 70%, it means that after pruning, 70% of the weights in the U-Net model will be set to zero.
[0105] Record the above pruning-related configuration parameters, including the target sparsity, initial sparsity, final sparsity, start step, and end step. At the same time, input the target sparsity, initial sparsity, final sparsity, start step, and end step into the U-Net model. This process will gradually remove unimportant weight connections according to the target sparsity. Note that pruning is carried out step by step, not all at once. This means that between the start step and the end step, the sparsity of the U-Net model will gradually increase until it reaches the final sparsity. If it is found that the pruning speed is too fast or too slow, the pruning rate can be adjusted according to the actual situation. For example, if it is found that the performance of the U-Net model deteriorates rapidly during pruning, the pruning speed can be appropriately slowed down and the final sparsity can be reduced. After the pruning process is completed, save the pruned U-Net model. The pruned model is smaller in size and occupies less memory accordingly, which means that more 3D naked-eye images can run simultaneously without worrying about memory shortages, and at the same time, it reduces the risk of image crashes caused by memory overflow.
[0106] Use the TensorFlow Lite converter tool to convert the pruned U-Net model into the TensorFlow Lite format (.tflite file), integrate the converted.tflite file into Android and iOS applications, input the 3D naked-eye image into the U-Net model in the.tflite file, and increase and decrease the pruning intensity according to the speed and memory occupancy when the 3D naked-eye image runs through the U-Net model in Android and iOS applications to find the best performance and energy consumption ratio. According to the performance in actual applications, developers can flexibly adjust the pruning intensity to find the best balance between performance and energy consumption. For example, in tasks that require higher precision, a lower sparsity can be selected; while in the pursuit of higher efficiency, a certain degree of accuracy loss can be accepted in exchange for faster speed and lower energy consumption. The 3D naked-eye image generation function that runs efficiently on mobile devices can provide users with a more smooth and natural interaction experience. Whether for entertainment or professional applications, the 3D naked-eye image generation technology can enable users to feel the convenience and fun brought by technological progress.
[0107] This embodiment also provides a 3D naked-eye image generation system based on a neural network, including:
[0108] A data preprocessing module that extracts stereo views from a dataset containing stereo views and performs preprocessing operations such as size adjustment and color correction on the stereo views;
[0109] A depth map module that uses the Monodepth2 model to analyze the left-eye view and right-eye view in the stereo view to obtain the disparity, creates a depth map matrix, calculates the depth value using the triangulation principle of a binocular camera, and fills the depth value into the depth map matrix to obtain the depth map;
[0110] A U-Net model training module that inputs the depth map and the preprocessed left-eye view and right-eye view Figure 1 into the U-Net model together and performs iterative training in combination with a generator until the discriminator can no longer distinguish the images generated by the U-Net model from real images;
[0111] An image generation module that, after the U-Net model training is completed, inputs the newly obtained left and right eye views and depth map and uses the dynamic region enhancement technology to input them into the trained U-Net model to generate high-quality 3D naked-eye images;
[0112] A U-Net model optimization module that prunes and optimizes the U-Net model according to pruning parameters, target sparsity, initial sparsity, final sparsity, start step, and end step.
[0113] This embodiment also provides a computer device applicable to the case of a method for generating 3D naked-eye images based on a neural network, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for generating 3D naked-eye images based on a neural network as proposed in the above embodiment.
[0114] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0115] This embodiment also provides a storage medium on which a computer program is stored, and when the program is executed by a processor, it implements the method for generating 3D naked-eye images based on a neural network as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read Only Memory (EPROM for short), Programmable Red-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0116] In summary, the present invention optimizes the overall process from data preprocessing to the generation of high-quality 3D naked-eye images, improving the overall performance and practicality of the system. Overall, an efficient, accurate, and user-friendly 3D naked-eye image generation mechanism is jointly constructed. Among them, the dynamic region enhancement technology not only provides higher depth estimation accuracy and consistency in complex scene reconstruction, but also greatly enhances the realism and immersive experience of 3D naked-eye images. By using advanced parallax optimization, attention mechanism, and generative adversarial network technology, it can specifically focus on and optimize the expressiveness of 3D naked-eye images while ensuring the effective utilization of computing resources. In addition, the standardized data preprocessing steps lay a solid foundation for subsequent processing, making the entire process more robust, adaptable, and able to quickly identify and correct potential problems, and also bringing an unprecedented visual experience to users, promoting the technological progress and application expansion in related fields.
[0117] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A 3D naked-eye image generation method based on a neural network, characterized in that: including, obtaining a stereoscopic view dataset and preprocessing the stereoscopic views; loading a pre-trained Monodepth2 model to analyze the disparity between the left-eye view and the right-eye view in the stereoscopic views, and applying the triangulation principle to calculate the depth value corresponding to each pixel to obtain a depth map; constructing an encoder using a 3x3 convolutional kernel and a 2x2 max pooling layer, and establishing a decoder through transposed convolution to obtain a U-Net model; initializing all the weights of the generator and the discriminator using the standard normal distribution, and adding the transposed convolution of the initialized generator to the output layer of the U-Net model; inputting the depth map into the U-Net model, and designing adversarial loss, content loss, and discriminator loss for iterative training until the discriminator can no longer distinguish the images generated by the U-Net model from the real images; after the iterative training is completed, obtaining new left-eye and right-eye views and depth maps from the stereoscopic view dataset; constructing an additional CNN model using convolutional layers, pooling layers, and fully connected layers, and initializing the CNN model through Xavier and training the CNN model using the cross-entropy loss function; applying dynamic enhancement technology to input the depth map into the trained CNN model to generate an attention weight map; inputting the newly obtained left-eye and right-eye views, depth map, and attention weight map into the U-Net model to generate a 3D naked-eye image; pruning and optimizing the U-Net model that generates the 3D naked-eye image so that the 3D naked-eye image can run smoothly with low energy consumption on a mobile device.
2. The 3D naked-eye image generation method based on a neural network according to claim 1, characterized in that: The obtaining of the stereoscopic view dataset and the preprocessing of the stereoscopic views are specifically as follows: The stereoscopic views include a left-eye view and a right-eye view; The preprocessing includes size adjustment and color correction.
3. The 3D naked-eye image generation method based on a neural network according to claim 2, wherein: The loading of the pre-trained Monodepth2 model to analyze the disparity between the left-eye view and the right-eye view in the stereoscopic views is specifically as follows, converting the preprocessed left-eye view and right-eye view into tensor format and adjusting them to the size required by the Monodepth2 model; inputting the tensors of the left-eye view and the right-eye view after format conversion and size adjustment into the Monodepth2 model for forward propagation to obtain the disparity; calculating the disparity value by applying the sum of absolute values to the corresponding pixel windows of the left-eye view and the right-eye view, and comparing the similarity of the disparity values; establishing a two-dimensional array of the disparity map, and filling the two-dimensional array with the disparity value with the smallest similarity to obtain the disparity map.
4. The 3D naked-eye image generation method based on a neural network according to claim 3, characterized in that: The application of the triangulation principle to calculate the depth value corresponding to each pixel to obtain a depth map is specifically as follows, creating a depth map matrix using NumPy as a container for the final depth map to store the depth values according to the size of the disparity map; calculating the depth value by using the triangulation principle of the binocular camera to calculate the disparity, and filling the depth map matrix to obtain the depth map; applying the Matplotlib plotting tool to check the edge sharpness of the depth map and the accuracy of the object contour.
5. The 3D naked-eye image generation method based on a neural network according to claim 4, characterized in that: The pruning and optimizing of the U-Net model that generates the 3D naked-eye image so that the 3D naked-eye image can run efficiently on a mobile device is specifically as follows, setting the weights of the U-Net model to zero; Set the start and end times of pruning according to the set number of pruning steps, and record all relevant pruning configuration parameters to guide the pruning process; The pruning configuration parameters include the desired sparsity, the initial sparsity at the start of pruning, the final sparsity after pruning is completed, the start step of pruning, and the end step of pruning; According to the desired sparsity level, input the start and end times of pruning into the U-Net model for step-by-step pruning; Use the TensorFlow Lite converter tool to convert the pruned U-Net model into the tflite format, integrate the converted tflite file into Android and iOS applications, and input 3D stereoscopic images.
6. A 3D naked-eye image generation system based on a neural network, for performing the 3D naked-eye image generation method based on a neural network according to any one of claims 1 to 5, characterized in that: Including, A data preprocessing module that extracts stereoscopic views from a dataset containing stereoscopic views and performs preprocessing operations such as resizing and color correction on the stereoscopic views; A depth map module that uses the Monodepth2 model to analyze the left-eye view and right-eye view in the stereoscopic view to obtain the disparity, creates a depth map matrix, calculates the depth value using the triangulation principle of a binocular camera, and fills the depth value into the depth map matrix to obtain a depth map; A U-Net model training module that inputs the depth map together with the preprocessed left-eye view and right-eye view into the U-Net model and performs iterative training in combination with a generator until the discriminator can no longer distinguish between the images generated by the U-Net model and the real images; An image generation module that, after the U-Net model training is completed, inputs the newly obtained left and right eye views and depth map after using the dynamic region enhancement technique into the trained U-Net model to generate high-quality 3D stereoscopic images; A U-Net model optimization module that prunes and optimizes the U-Net model according to the pruning parameters, the target sparsity, the initial sparsity, the final sparsity, the start step, and the end step; 7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the neural network-based 3D stereoscopic image generation method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the neural network-based 3D stereoscopic image generation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Image conversion method and device, depth map prediction method and device, model training method and device and electronic equipment
CN110111244A
Electronic component three-dimensional reconstruction method based on semi-supervised learning
CN113902807A