A lightweight high-quality integrated imaging 3D video generation method

CN119967144BActive Publication Date: 2026-09-22XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510124671.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2026-09-22
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

[0003]在3D视频的相关研究中,传统方法利用虚拟视点合成、深度匹配、数据压缩的方法减少了采集以及传输当中涉及到的数据量,然而这些方法仍然面临着完整3D视频数据量过于庞大,只能妥协性地完成低帧率或低分辨率的3D视频生成

Benefits of technology

[0013]本发明提出一种轻量化的高质量集成成像3D视频生成方法。根据微图像阵列内部的光学逻辑结构,设计自适应块选择算法,并与卷积神经网络结合,在严格遵守微图像阵列内部光学逻辑结构的基础上,充分挖掘微图像阵列序列中蕴含的运动特征与像素特征,在不降低微图像阵列分辨率的情况下准确生成微图像阵列的中间帧。同时利用迭代的方法,将每一轮生成的微图像阵列中间帧按照时间顺序插入微图像阵列序列中,并再次通过本方法生成新的微图像阵列中间帧。最终实现稀疏微图像阵列序列到密集微图像阵列序列的转化,即实现轻量化的高质量集成成像3D视频生成。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967144B_ABST
    Figure CN119967144B_ABST
Patent Text Reader

Abstract

The application provides a light-weight high-quality integrated imaging 3D video generation method, which comprises four processes of four-dimensional light field data acquisition, adaptive block selection, motion and color value feature extraction, target pixel filling, micro image array sequence densification and high frame rate three-dimensional reconstruction, generates a sparse micro image array sequence from the collected four-dimensional light field data, obtains a pre-processing receptive field and an image block through an adaptive block selection algorithm and inputs the pre-processing receptive field and the image block into a motion and color value feature extraction network, fully excavates motion features and pixel features contained in the micro image array sequence on the basis of strictly complying with internal optical logic structure of the micro image array, accurately generates an intermediate frame of the micro image array without reducing resolution of the micro image array, realizes densification of the sparse micro image array sequence through an iterative method, and finally reconstructs a high-quality 3D video effect by means of an integrated imaging display device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video generation for image processing tasks, and more specifically, to a lightweight, high-quality integrated imaging 3D video generation method. Background Technology

[0002] Integrated imaging is a naked-eye 3D stereoscopic imaging technology characterized by full parallax and full-color display. It utilizes lens arrays or pinhole arrays to efficiently store and reconstruct 3D image data, increasing information capacity and display dimensionality. With the development of science and technology, integrated imaging 3D videos can provide a more realistic visual experience. However, integrated imaging videos require a massive amount of data, making acquisition difficult. Reducing the amount of data acquired and processed is a key issue in the generation of integrated imaging videos. Frame interpolation can densify temporally sparse 3D image sequences, greatly simplifying the acquisition and information processing of 3D videos.

[0003] In 3D video research, traditional methods utilize virtual viewpoint synthesis, depth matching, and data compression to reduce the amount of data involved in acquisition and transmission. However, these methods still face the challenge of the excessively large data volume of complete 3D videos, limiting them to low frame rate or low resolution 3D video generation as a compromise. While using light field cameras can acquire high frame rate 3D videos, the resulting 3D videos are limited by low sensor resolution and small parallax, making 3D reconstruction impossible. In frame interpolation research, traditional methods are only applicable to 2D images. 3D images, due to their numerous discrete corresponding points, exhibit discontinuous pixel distribution. This leads to significant distortions and artifacts in traditional frame interpolation methods when processing 3D images, resulting in intermediate frames that cannot reconstruct 3D objects, thus hindering the understanding of spatial context. Summary of the Invention

[0004] This invention proposes a lightweight, high-quality integrated imaging 3D video generation method. Its core lies in incorporating the optical logic structure within the micro-image array to improve the accuracy of 3D video generation. It transforms the input sparse micro-image array sequence into a dense micro-image array sequence as the output, achieving lightweight, high-quality integrated imaging 3D video generation. The method first performs sparse sampling of the four-dimensional light field in the time dimension. Then, in two adjacent frames of the acquired sparse micro-image array sequence, the receptive field and image block range are defined centered on the target pixel. An adaptive block selection algorithm is used to crop and fill the receptive field and image block, resulting in a preprocessed receptive field and preprocessed image block. Next, a motion and color value feature extraction network is used to extract features from the preprocessed receptive field, thereby obtaining an interpolation kernel K. This interpolation kernel K implicitly contains the position and color value information of the target pixel. The interpolation kernel K is convolved with the preprocessed image block to obtain the pixel information of the target pixel position. This process traverses the entire micro-image array and combines all the obtained pixel information to synthesize intermediate frames of the micro-image array. To further improve the frame rate, the micro-image array sequence, including the newly generated micro-image array, will repeat the above steps again, iterating continuously to generate a dense micro-image array sequence. Finally, the four-dimensional light field angle information and spatial information are reconstructed using an integrated imaging display device.

[0005] The method comprises six processes: acquisition of four-dimensional light field data, adaptive block selection, extraction of motion and color value features, target pixel filling, micro-image array sequence densification, and high frame rate three-dimensional reconstruction. The specific process is attached. Figure 1 As shown.

[0006] The process of acquiring the four-dimensional light field data involves establishing a real camera array to sparsely collect real four-dimensional light field information over time, or specifying the trajectory of an object in computer graphics software and establishing a virtual camera array to acquire virtual four-dimensional light field information, and then generating a micro-image array through a pixel mapping algorithm.

[0007] The adaptive block selection process is as follows: Figure 2As shown. The initial image block and initial receptive field are defined with the target pixel coordinates (x, y) as the center. The size of the initial image block and initial receptive field is determined based on the size of the micro-image. The initial image block and initial receptive field are cropped and filled based on whether the target pixel coordinates (x, y) are located at the center of their respective micro-images. When the target pixel coordinates (x, y) are located at the center of their respective micro-images, there are no extra pixels in the initial image block and initial receptive field that overlap with surrounding micro-images, and the initial image block and initial receptive field are directly output, resulting in the preprocessed image block and preprocessed receptive field. When the target pixel coordinates (x, y) gradually deviate from the center of the micro-image, the initial image block and initial receptive field overlap with surrounding micro-images. In this case, the pixels overlapping with the surrounding micro-images are cropped as extra pixels, and the corresponding pixels in the micro-image containing the target pixel are used to fill the cropped gaps, resulting in the preprocessed image block and preprocessed receptive field.

[0008] The adaptive block selection process for cropping and filling image blocks is shown in the attached figure. Figure 3 As shown in the figure. When the target pixel coordinates are (x1, y1), it deviates from the center of the micro-image. Pixels overlapping with the surrounding micro-images are cropped as extra pixels, and the missing parts after cropping are filled with pixels at the corresponding positions in the micro-image where the target pixel is located, thus completing the reorganization of the pixel block, as shown in the attached figure. Figure 3 As shown in (a), when the target pixel coordinates are (x2, y2), it is located at the center of the micro-image. At this time, there are no extra pixels in the image block, and the image block is directly input without any additional processing, as shown in (a). Figure 3 As shown in (b).

[0009] To avoid the receptive field acquiring too much information from different micro-images, the adaptive block selection process will process the receptive field in the same way.

[0010] The motion and color value feature extraction process is shown in the appendix. Figure 4 As shown, two preprocessed receptive fields are input into the deep convolutional neural network used in this process. The preprocessed receptive fields are input as a single input from three channels: red, green, and blue. Convolutional layers and downconvolutional layers are alternated. A series of convolutional layers simultaneously extract motion and pixel features at different scales from the two receptive fields. Batch Normalization is used for regularization in each convolutional layer, with Leaky ReLU as the activation function. Downconvolutional layers are used during feature extraction intervals to reduce feature loss while downsampling. The final output one-dimensional vector is passed through a spatial softmax layer to ensure that all values ​​in the vector are non-negative. After processing, the one-dimensional vector is reconstructed into an interpolation kernel K that is twice the size of the image patch.

[0011] The process of densifying the micro-image array sequence involves convolving an interpolation kernel K with two preprocessed image blocks to obtain the color value at the target pixel coordinates (x, y). This process iterates through the entire micro-image array, arranging the color values ​​at the target pixel coordinates (x, y) obtained by convolving the interpolation kernel K with the corresponding preprocessed image block according to their coordinates, thus obtaining intermediate frames of the micro-image array. All intermediate frames of the micro-image array are added to the original sparse micro-image array sequence in chronological order, and the process is iterated to finally generate a dense micro-image array sequence.

[0012] The high frame rate micro-image array three-dimensional reconstruction process uses a dense micro-image array sequence to reconstruct four-dimensional light field angle information and spatial information with the help of an integrated imaging display device.

[0013] This invention proposes a lightweight, high-quality integrated imaging 3D video generation method. Based on the internal optical logic structure of a micro-image array, an adaptive block selection algorithm is designed and combined with a convolutional neural network. While strictly adhering to the internal optical logic structure of the micro-image array, this method fully exploits the motion and pixel features inherent in the micro-image array sequence, accurately generating intermediate frames of the micro-image array without reducing its resolution. Simultaneously, an iterative method is used to insert the intermediate frames generated in each round into the micro-image array sequence in chronological order, and new intermediate frames are generated again using this method. Ultimately, this achieves the transformation from a sparse micro-image array sequence to a dense micro-image array sequence, thus realizing lightweight, high-quality integrated imaging 3D video generation. Attached Figure Description

[0014] Appendix Figure 1 This is a flowchart of a lightweight, high-quality integrated imaging 3D video generation method proposed in this invention.

[0015] Appendix Figure 2 Flowchart for selecting adaptive blocks.

[0016] Appendix Figure 3 A schematic diagram of the image block selection process for adaptive block selection.

[0017] Appendix Figure 4 This is a flowchart of the motion and pixel feature extraction process for adjacent micro-image arrays.

[0018] It should be understood that the above figures are only schematic and are not drawn to scale. Detailed Implementation

[0019] The following detailed description of a typical embodiment of the lightweight, high-quality integrated imaging 3D video generation method proposed in this invention further illustrates the invention. It is important to note that the following embodiments are for illustrative purposes only and should not be construed as limiting the scope of protection of this invention. Any non-essential improvements and adjustments made to this invention by those skilled in the art based on the above description are still within the scope of protection of this invention.

[0020] This invention proposes a lightweight, high-quality integrated imaging 3D video generation method, comprising six processes: acquisition of four-dimensional light field data, adaptive block selection, motion and color value feature extraction, target pixel filling, micro-image array sequence densification, and high frame rate 3D reconstruction. The detailed process is attached. Figure 1 As shown.

[0021] The process of acquiring the four-dimensional light field data involves establishing a real camera array to sparsely collect real four-dimensional light field information over time, or specifying the trajectory of an object in computer graphics software and establishing a virtual camera array to acquire virtual four-dimensional light field information, and then generating a micro-image array through a pixel mapping algorithm.

[0022] The adaptive block selection process is as follows: Figure 2 As shown. The initial image block and initial receptive field are defined with the target pixel coordinates (x, y) as the center. The size of the initial image block and initial receptive field is determined based on the size of the micro-image. The initial image block and initial receptive field are cropped and filled based on whether the target pixel coordinates (x, y) are located at the center of their respective micro-images. When the target pixel coordinates (x, y) are located at the center of their respective micro-images, there are no extra pixels in the initial image block and initial receptive field that overlap with surrounding micro-images, and the initial image block and initial receptive field are directly output, resulting in the preprocessed image block and preprocessed receptive field. When the target pixel coordinates (x, y) gradually deviate from the center of the micro-image, the initial image block and initial receptive field overlap with surrounding micro-images. In this case, the pixels overlapping with the surrounding micro-images are cropped as extra pixels, and the corresponding pixels in the micro-image containing the target pixel are used to fill the cropped gaps, resulting in the preprocessed image block and preprocessed receptive field.

[0023] The adaptive block selection process for cropping and filling image blocks is shown in the attached figure. Figure 3 As shown in the figure. When the target pixel coordinates are (x1, y1), it deviates from the center of the micro-image. Pixels overlapping with the surrounding micro-images are cropped as extra pixels, and the missing parts after cropping are filled with pixels at the corresponding positions in the micro-image where the target pixel is located, thus completing the reorganization of the pixel block, as shown in the attached figure. Figure 3As shown in (a), when the target pixel coordinates are (x2, y2), it is located at the center of the micro-image. At this time, there are no extra pixels in the image block, and the image block is directly input without any additional processing, as shown in (a). Figure 3 As shown in (b).

[0024] To avoid the receptive field acquiring too much information from different micro-images, the adaptive block selection process will process the receptive field in the same way.

[0025] The motion and color value feature extraction process is shown in the appendix. Figure 4 As shown, two preprocessed receptive fields are input into the deep convolutional neural network used in this process. The preprocessed receptive fields are input as a single input from three channels: red, green, and blue. Convolutional layers and downconvolutional layers are alternated. A series of convolutional layers simultaneously extract motion and pixel features at different scales from the two receptive fields. Batch Normalization is used for regularization in each convolutional layer, with Leaky ReLU as the activation function. Downconvolutional layers are used during feature extraction intervals to reduce feature loss while downsampling. The final output one-dimensional vector is passed through a spatial softmax layer to ensure that all values ​​in the vector are non-negative. After processing, the one-dimensional vector is reconstructed into an interpolation kernel K that is twice the size of the image patch.

[0026] The process of densifying the micro-image array sequence involves convolving an interpolation kernel K with two preprocessed image blocks to obtain the color value at the target pixel coordinates (x, y). This process iterates through the entire micro-image array, arranging the color values ​​at the target pixel coordinates (x, y) obtained by convolving the interpolation kernel K with the corresponding preprocessed image block according to their coordinates, thus obtaining intermediate frames of the micro-image array. All intermediate frames of the micro-image array are added to the original sparse micro-image array sequence in chronological order, and the process is iterated to finally generate a dense micro-image array sequence.

[0027] The high frame rate micro-image array three-dimensional reconstruction process uses a dense micro-image array sequence to reconstruct four-dimensional light field angle information and spatial information with the help of an integrated imaging display device.

[0028] This invention proposes a lightweight, high-quality integrated imaging 3D video generation method. Based on the internal optical logic structure of a micro-image array, an adaptive block selection algorithm is designed and combined with a convolutional neural network. While strictly adhering to the internal optical logic structure of the micro-image array, this method fully exploits the motion and pixel features inherent in the micro-image array sequence, accurately generating intermediate frames of the micro-image array without reducing its resolution. Simultaneously, an iterative method is used to insert the intermediate frames generated in each round into the micro-image array sequence in chronological order, and new intermediate frames are generated again using this method. Ultimately, this achieves the transformation from a sparse micro-image array sequence to a dense micro-image array sequence, thus realizing lightweight, high-quality integrated imaging 3D video generation.

Claims

1. A lightweight, high-quality integrated imaging 3D video generation method, characterized in that... By incorporating the optical logic structure within the micro-image array to improve the accuracy of 3D video generation, the input sparse micro-image array sequence is transformed into a dense micro-image array sequence output, achieving lightweight, high-quality integrated imaging 3D video generation. The method includes six processes: acquisition of four-dimensional light field data, adaptive block selection, motion and color value feature extraction, target pixel filling, micro-image array sequence densification, and high frame rate 3D reconstruction. First, the four-dimensional light field is sparsely sampled in the time dimension. Subsequently, in the two adjacent frames of the acquired sparse micro-image array sequence, the receptive field and image patch ranges are defined with the target pixel as the center. An adaptive block selection algorithm is used to crop and fill the receptive field and image patch. This algorithm defines the initial image patch and initial receptive field range with the target pixel coordinates (x, y) as the center, determines the size of the initial image patch and initial receptive field based on the size of the micro-image, and crops and fills the initial image patch and initial receptive field based on whether the target pixel coordinates (x, y) are located at the center of its micro-image. When the target pixel coordinates (x, y) are located at the center of its micro-image, there are no extra pixels overlapping with surrounding micro-images in the initial image patch and initial receptive field, and the initial image patch and initial receptive field are directly output, resulting in the preprocessed image patch and preprocessed receptive field. When the target pixel coordinates (x, y) are located at the center of its micro-image, there are no extra pixels overlapping with surrounding micro-images in the initial image patch and initial receptive field, and the initial image patch and initial receptive field are directly output, resulting in the preprocessed image patch and preprocessed receptive field. When the target pixel deviates from the center of its micro-image, the pixels that overlap with the surrounding micro-images in the initial image block and the initial receptive field are cropped as additional pixels. The missing parts after cropping are filled with pixels at the corresponding positions in the micro-image where the target pixel is located, resulting in a preprocessed image block and a preprocessed receptive field. Then, a motion and color value feature extraction network is used to extract features from the preprocessed receptive field to obtain an interpolation kernel K. The interpolation kernel K implicitly contains the position information and color value information of the target pixel. The interpolation kernel K is convolved with the preprocessed image block to obtain the pixel information of the target pixel position. The process traverses the entire micro-image array and combines all the obtained pixel information to synthesize the intermediate frame of the micro-image array. In order to further improve the frame rate, the micro-image array sequence, including the newly generated micro-image array, will repeat the above steps again and iterate continuously to generate a dense micro-image array sequence. Finally, the four-dimensional light field angle information and spatial information are reconstructed with the help of an integrated imaging display device.

2. The lightweight, high-quality integrated imaging 3D video generation method according to claim 1, characterized in that, A real camera array is established to sparsely acquire real four-dimensional light field information over time, or the trajectory of an object is specified in computer graphics software, and a virtual camera array is established to acquire virtual four-dimensional light field information. Then, a micro-image array is generated through a pixel mapping algorithm.

3. The lightweight, high-quality integrated imaging 3D video generation method according to claim 1, characterized in that, The interpolation kernel K is convolved with two preprocessed image blocks to obtain the color value at the target pixel coordinates (x, y). This process traverses the entire micro-image array and arranges the color values ​​at the target pixel coordinates (x, y) obtained by convolving the interpolation kernel K with the corresponding preprocessed image block according to the coordinates to obtain the intermediate frames of the micro-image array. All the generated intermediate frames of the micro-image array are added to the original sparse micro-image array sequence in chronological order and iterated to finally generate a dense micro-image array sequence.

4. The lightweight, high-quality integrated imaging 3D video generation method according to claim 1, characterized in that, A dense array of micro-images is used to reconstruct four-dimensional light field angle and spatial information using an integrated imaging display device.