Lightweight high-quality integrated imaging 3D video generation method
By applying adaptive block selection algorithms and convolutional neural networks in the microimage array, dense microimage array sequences are generated, and the problems of large amount of data and interpolation errors in integrated imaging 3D video generation are solved, and high-quality and high-frame rate 3D video generation is achieved.
Patent Information
- Application Number
- CN202510124671.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-03-04
AI Technical Summary
In the integrated imaging 3D video generation, the existing technology is difficult to effectively reduce the amount of data, resulting in difficulty in acquisition and processing. Traditional methods can only generate 3D videos with low frame rate or low resolution, and traditional interpolation methods will experience distortion and artifact errors when processing 3D images.
A lightweight, high-quality integrated imaging 3D video generation method is proposed. By taking into account the optical logic structure in the micro-image array, the adaptive block selection algorithm and convolutional neural network are used to pre-process and feature extraction of sparse micro-image array sequences to generate dense micro-image array sequences, and three-dimensional reconstruction with high frame rate is achieved.
It realizes the accurate generation of intermediate frames of the micro-image array without reducing the resolution of the micro-image array, which reduces the data volume and improves the quality and frame rate of 3D video, and solves the problems of huge data volume and interpolation errors in traditional methods.
Smart Images

Figure CN119967144A_ABST
Abstract
Description
1. Technical Field
[0001] The present invention relates to the field of video generation of image processing tasks, and more specifically, to a lightweight high-quality integrated imaging 3D video generation method. 2. Background Technology
[0002] Integrated imaging is a naked-eye three-dimensional stereo imaging technology with full parallax and full-color display. It uses lens arrays or pinhole arrays to achieve efficient storage and reconstruction of 3D image data, thereby increasing the information carrying capacity and display dimension. With the development of science and technology, integrated imaging 3D video can give people a more realistic visual experience. However, the amount of data required for integrated imaging video is huge and difficult to collect. How to reduce the amount of data collected and processed is a key issue in the generation of integrated imaging video. The interpolation method can densify the sparse 3D image sequence in time, greatly simplifying the acquisition and information processing process of 3D video.
[0003] In the research on 3D video, traditional methods use virtual viewpoint synthesis, depth matching, and data compression to reduce the amount of data involved in acquisition and transmission. However, these methods still face the problem that the amount of data for a complete 3D video is too large, and can only compromise to complete the generation of 3D videos with low frame rates or low resolutions. Although the method using light field cameras can capture 3D videos with high frame rates, the 3D videos obtained are limited by the low resolution of the sensor and the small parallax, which cannot be used for three-dimensional reconstruction. In the research on interpolation, traditional interpolation methods are only applicable to 2D images. Since 3D images contain a large number of discrete points of the same name, their pixel distribution is not continuous. This causes traditional interpolation methods to have a large number of distortions, artifacts, and other errors when processing 3D images. The generated intermediate frames cannot reconstruct three-dimensional objects, which hinders the understanding of spatial context. III. Summary of the invention
[0004] The present invention proposes a lightweight high-quality integrated imaging 3D video generation method, the core of which is to take the optical logic structure in the micro-image array into consideration to improve the accuracy of 3D video generation, convert the input sparse micro-image array sequence into a dense micro-image array sequence output, and realize lightweight high-quality integrated imaging 3D video generation. The method first performs sparse sampling on the four-dimensional light field in the time dimension. Subsequently, the range of the receptive field and the image block are respectively delineated with the target pixel as the center in the two adjacent frames before and after the collected sparse micro-image array sequence, and the pre-processed receptive field and image block are obtained by cutting and filling them using an adaptive block selection algorithm. Then, the pre-processed receptive field and image block are feature extracted using a motion estimation and pixel synthesis network to obtain an interpolation kernel K, which implicitly contains the position information and color value information at the target pixel. The interpolation kernel K is convolved with the pre-processed image block to obtain the pixel information of the target pixel position. This process traverses the entire micro-image array, and combines all the obtained pixel information to synthesize the micro-image array intermediate frame. In order to further improve the frame rate, the micro-image array sequence including the newly generated micro-image array will repeat the above steps again, and continuously iterate to generate a dense micro-image array sequence. Finally, a high-quality 3D video effect is reconstructed with the help of an integrated imaging display device.
[0005] The method includes six processes: acquisition of four-dimensional light field data, adaptive block selection, motion and color value feature extraction, target pixel filling, micro-image array sequence densification, and high frame rate three-dimensional reconstruction. The specific process is shown in the attached Figure 1 shown.
[0006] The process of acquiring the four-dimensional light field data can establish a real camera array to sparsely collect the real four-dimensional light field information on a time scale, or can specify the object motion trajectory in computer graphics software and establish a virtual camera array to acquire virtual four-dimensional light field information, and then generate a micro-image array through a pixel mapping algorithm.
[0007] The adaptive block selection process is shown in the attached Figure 2 As shown. With the target pixel coordinates (x, y) as the center, the range of the initial image block and the receptive field is framed, and the size of the initial image block and the receptive field is determined according to the size of the micro-image. Adaptive block selection determines whether the target pixel coordinates (x, y) are located in the center of the micro-image where it is located. According to the determination result, the adaptive block selection algorithm adopts different strategies to crop and fill the image block and the receptive field. Take the processing of the image block as an example. When the target pixel coordinates (x, y) are located in the center of the micro-image where it is located, the image block overlaps with the micro-image. At this time, the image block will not be processed and the image block will be directly output. When the target pixel coordinates (x, y) gradually deviate from the center of the micro-image, the image block overlaps with the surrounding micro-images. At this time, the image block is cropped in the third step to remove the extra pixels and reorganize to fill the empty part of the image block.
[0008] The process of cropping and filling image blocks by the adaptive block selection process is shown in the attached figure. Figure 3 When the target pixel coordinate (x1, y1) deviates from the center of the micro-image, except for the pixel block 2 provided by the micro-image where the target pixel coordinate (x1, y1) is located, the remaining pixels are all extra pixels. These extra pixels are cropped and filled with the pixel blocks 1, 3, and 4 where the target pixel (x1, y1) is located to complete the reorganization of the pixel blocks, as shown in the attached figure. Figure 3 In particular, when the target pixel (x2, y2) is located at the center of the micro-image, there are no extra pixels in the image block at this time, so the image block will be directly input without any additional processing, as shown in Figure 3 (b) as shown.
[0009] To avoid the receptive field collecting too much information from different micro-images, the adaptive block selection process will process the receptive field in the same way.
[0010] The motion and color value feature extraction process is as shown in the attached Figure 4 As shown. The two pre-processed receptive fields are input into the deep convolutional neural network in this process. The pre-processed receptive fields are divided into three-channel images of red, green and blue for common input. Convolutional layers and down-convolutional layers are arranged alternately. A series of convolutional layers simultaneously extract motion features and pixel features of two receptive fields at different scales. Batch Normlization is used for regularization in each convolutional layer, and Leaky ReLu is used as the activation function. Down-convolutional layers are used in the intervals of feature extraction to reduce feature loss while downsampling. The final output one-dimensional vector will pass through the spatial softmax layer to ensure that all values in the vector are non-negative. After processing, the one-dimensional vector will be reorganized into an interpolation kernel K of twice the size of the image block.
[0011] The micro-image array sequence densification process. The interpolation kernel K is convolved with two pre-processed image blocks to obtain the color value at the target pixel (x, y). This process traverses the entire micro-image array, and arranges the color value at the target pixel (x, y) obtained by convolving the interpolation kernel K with the corresponding image block each time according to the coordinates to obtain the micro-image array intermediate frame. All micro-image array intermediate frames are added to the original sparse micro-image array sequence in chronological order, and iterated to finally generate a dense micro-image array sequence.
[0012] In the high frame rate micro-image array three-dimensional reconstruction process, a dense micro-image array sequence reconstructs four-dimensional light field angle and spatial information with the help of an integrated imaging device, thereby displaying high-quality 3D video effects.
[0013] The present invention proposes a lightweight high-quality integrated imaging 3D video generation method. According to the optical logic structure inside the micro-image array, an adaptive block selection algorithm is designed and combined with a convolutional neural network. On the basis of strictly complying with the optical logic structure inside the micro-image array, the motion features and pixel features contained in the micro-image array sequence are fully explored, and the intermediate frames of the micro-image array are accurately generated without reducing the resolution of the micro-image array. At the same time, an iterative method is used to insert the intermediate frames of the micro-image array generated in each round into the micro-image array sequence in chronological order, and a new micro-image array intermediate frame is generated again by this method. Finally, the transformation from a sparse micro-image array sequence to a dense micro-image array sequence is realized, that is, lightweight high-quality integrated imaging 3D video generation is realized. IV. Description of the drawings
[0014] Attached Figure 1 This is a flow chart of a lightweight, high-quality integrated imaging 3D video generation method proposed by the present invention.
[0015] Attached Figure 2 Select a flowchart for the Adaptive block.
[0016] Attached Figure 3 Schematic diagram of the image block processing process for adaptive block selection.
[0017] Attached Figure 4 Flowchart for adjacent micro-image array motion and pixel feature extraction.
[0018] It should be understood that the above drawings are only schematic and are not drawn to scale. V. Specific implementation methods
[0019] A typical embodiment of a lightweight high-quality integrated imaging 3D video generation method proposed by the present invention is described in detail below, and the present invention is further described in detail. It is necessary to point out here that the following embodiments are only used to further illustrate the present invention and cannot be understood as limiting the protection scope of the present invention. Persons skilled in the art in this field may make some non-essential improvements and adjustments to the present invention based on the above content of the present invention, which still fall within the protection scope of the present invention.
[0020] The present invention proposes a lightweight high-quality integrated imaging 3D video generation method, which includes six processes: acquisition of four-dimensional light field data, adaptive block selection, motion and color value feature extraction, target pixel filling, dense micro-image array sequence generation, and high frame rate three-dimensional reconstruction. The specific process is shown in the attached figure. Figure 1 shown.
[0021] The process of acquiring the four-dimensional light field data can establish a real camera array to sparsely collect the real four-dimensional light field information on a time scale, or can specify the object motion trajectory in computer graphics software and establish a virtual camera array to acquire virtual four-dimensional light field information, and then generate a micro-image array through a pixel mapping algorithm.
[0022] The adaptive block selection process is shown in the attached Figure 2 As shown. With the target pixel coordinates (x, y) as the center, the range of the initial image block and the receptive field is framed, and the size of the initial image block and the receptive field is determined according to the size of the micro-image. Adaptive block selection determines whether the target pixel coordinates (x, y) are located in the center of the micro-image where it is located. According to the determination result, the adaptive block selection algorithm adopts different strategies to crop and fill the image block and the receptive field. Take the processing of the image block as an example. When the target pixel coordinates (x, y) are located in the center of the micro-image where it is located, the image block overlaps with the micro-image. At this time, the image block will not be processed and the image block will be directly output. When the target pixel coordinates (x, y) gradually deviate from the center of the micro-image, the image block overlaps with the surrounding micro-images. At this time, the image block is cropped in the third step to remove the extra pixels and reorganize to fill the empty part of the image block.
[0023] The process of cropping and filling image blocks by the adaptive block selection process is shown in the attached figure. Figure 3 When the target pixel coordinate (x1, y1) deviates from the center of the micro-image, except for the pixel block 2 provided by the micro-image where the target pixel coordinate (x1, y1) is located, the remaining pixels are all extra pixels. These extra pixels are cropped and filled with the pixel blocks 1, 3, and 4 where the target pixel (x1, y1) is located to complete the reorganization of the pixel blocks, as shown in the attached figure. Figure 3 In particular, when the target pixel (x2, y2) is located at the center of the micro-image, there are no extra pixels in the image block at this time, so the image block will be directly input without any additional processing, as shown in Figure 3 (b) as shown.
[0024] To avoid the receptive field collecting too much information from different micro-images, the adaptive block selection process will process the receptive field in the same way.
[0025] The motion and color value feature extraction process is as shown in the attached Figure 4As shown. The two pre-processed receptive fields are input into the deep convolutional neural network in this process. The pre-processed receptive fields are divided into three-channel images of red, green and blue for common input. Convolutional layers and down-convolutional layers are arranged alternately. A series of convolutional layers simultaneously extract motion features and pixel features of two receptive fields at different scales. Batch Normlization is used for regularization in each convolutional layer, and Leaky ReLu is used as the activation function. Down-convolutional layers are used in the intervals of feature extraction to reduce feature loss while downsampling. The final output one-dimensional vector will pass through the spatial softmax layer to ensure that all values in the vector are non-negative. After processing, the one-dimensional vector will be reorganized into an interpolation kernel K of twice the size of the image block.
[0026] The micro-image array sequence densification process. The interpolation kernel K is convolved with two pre-processed image blocks to obtain the color value at the target pixel (x, y). This process traverses the entire micro-image array, and arranges the color value at the target pixel (x, y) obtained by convolving the interpolation kernel K with the corresponding image block each time according to the coordinates to obtain the micro-image array intermediate frame. All micro-image array intermediate frames are added to the original sparse micro-image array sequence in chronological order, and iterated to finally generate a dense micro-image array sequence.
[0027] In the high frame rate micro-image array three-dimensional reconstruction process, a dense micro-image array sequence reconstructs four-dimensional light field angle and spatial information with the help of an integrated imaging device, thereby displaying high-quality 3D video effects.
[0028] The present invention proposes a lightweight high-quality integrated imaging 3D video generation method. According to the optical logic structure inside the micro-image array, an adaptive block selection algorithm is designed and combined with a convolutional neural network. On the basis of strictly complying with the optical logic structure inside the micro-image array, the motion features and pixel features contained in the micro-image array sequence are fully explored, and the intermediate frames of the micro-image array are accurately generated without reducing the resolution of the micro-image array. At the same time, an iterative method is used to insert the intermediate frames of the micro-image array generated in each round into the micro-image array sequence in chronological order, and a new micro-image array intermediate frame is generated again by this method. Finally, the transformation from a sparse micro-image array sequence to a dense micro-image array sequence is realized, that is, lightweight high-quality integrated imaging 3D video generation is realized.
Claims
1. A lightweight high-quality integrated imaging 3D video generation method, characterized in that The optical logic structure in the micro-image array is taken into account to improve the accuracy of 3D video generation, and the input sparse micro-image array sequence is converted into a dense micro-image array sequence output, so as to realize lightweight high-quality integrated imaging 3D video generation; the method includes six processes: acquisition of four-dimensional light field data, adaptive block selection, motion and color value feature extraction, target pixel filling, micro-image array sequence densification, and high frame rate three-dimensional reconstruction; firstly, the four-dimensional light field is sparsely sampled in the time dimension; Subsequently, the range of the receptive field and the image block are respectively delineated with the target pixel as the center in the two adjacent frames before and after the acquired sparse micro-image array sequence, and the preprocessed receptive field and the image block are obtained by cropping and filling them using the adaptive block selection algorithm; then, the preprocessed receptive field and the image block are feature extracted using the motion estimation and pixel synthesis network to obtain an interpolation kernel K, which implicitly contains the position information and color value information of the target pixel, and the interpolation kernel K is convolved with the preprocessed image block to obtain the pixel information of the target pixel position; This process traverses the entire micro-image array and combines all the obtained pixel information to synthesize the micro-image array intermediate frame; in order to further improve the frame rate, the micro-image array sequence including the newly generated micro-image array will repeat the above steps again, and continuously iterate to generate a dense micro-image array sequence; finally, with the help of an integrated imaging display device, a high-quality 3D video effect is reconstructed.
2. A lightweight high-quality integrated imaging 3D video generation method according to claim 1, characterized in that: A real camera array can be established to sparsely collect real four-dimensional light field information on a time scale, or the object motion trajectory can be specified in computer graphics software, and a virtual camera array can be established to obtain virtual four-dimensional light field information, and then a micro-image array can be generated through a pixel mapping algorithm.
3. The lightweight high-quality integrated imaging 3D video generation method according to claim 1, characterized in that: With the target pixel coordinates (x, y) as the center, the range of the initial image block and the receptive field is framed, and the size of the initial image block and the receptive field is determined according to the size of the micro-image; the adaptive block selection determines whether the target pixel coordinates (x, y) are located in the center of the micro-image where it is located. According to the determination result, the adaptive block selection algorithm adopts different strategies to crop and fill the image block and the receptive field; taking the processing of the image block as an example, when the target pixel coordinates (x, y) are located in the center of the micro-image where it is located, the image block overlaps with the micro-image. At this time, the image block will not be processed and the image block will be directly output; when the target pixel coordinates (x, y) gradually deviate from the center of the micro-image, the image block overlaps with the surrounding micro-images. At this time, the image block is cropped in the third step to remove the extra pixels and reorganize to fill the empty part of the image block.
4. The lightweight high-quality integrated imaging 3D video generation method according to claim 1, characterized in that: When the target pixel coordinate (x1, y1) deviates from the center of the micro-image, except for the pixel block 2 provided by the micro-image where the target pixel coordinate (x1, y1) is located, the remaining pixels are all extra pixels. These extra pixels are cropped and filled with pixel blocks 1, 3, and 4 where the target pixel (x1, y1) is located to complete the reorganization of the pixel block. In particular, when the target pixel (x2, y2) is located at the center of the micro-image, there are no extra pixels in the image block at this time, so the image block will be directly input without any additional processing.
5. The lightweight high-quality integrated imaging 3D video generation method according to claim 1, characterized in that: The interpolation kernel K is convolved with the two preprocessed image blocks to obtain the color value at the target pixel (x, y); this process traverses the entire micro-image array, and arranges the color value at the target pixel (x, y) obtained by each convolution of the interpolation kernel K with the corresponding image block according to the coordinates to obtain the intermediate frame of the micro-image array; all the generated intermediate frames of the micro-image array are added to the original sparse micro-image array sequence in chronological order, and iterated to finally generate a dense micro-image array sequence.
6. The lightweight high-quality integrated imaging 3D video generation method according to claim 1, characterized in that: The dense micro-image array sequence reconstructs the four-dimensional light field angle and spatial information with the help of an integrated imaging device, thereby displaying high-quality 3D video effects.
Citation Information
Patent Citations
Light field super-resolution three-dimensional reconstruction method and system
CN113870433A
High-precision integrated imaging 3D salient target detection method with texture features
CN119515824A
Image resolution conversion device, image resolution conversion method, and computer program
JP2017010115A
Method for compressing light field data using variable block-size four-dimensional transforms and bit-plane decomposition
US10687068B1
Sparse light field representation
US20140328535A1