A fast compression and decomposition method and system for high-dimensional light field video
By constructing a high-dimensional light field pre-decomposition unit and a reverse transmission unit and utilizing the correlation between video key frames and non-key frames, the problems of large memory usage and long time of traditional light field decomposition algorithms are solved, efficient light field video decomposition is achieved, and the decomposition efficiency and signal-to-noise ratio are improved.
Patent Information
- Application Number
- CN202411634124.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-11-15
AI Technical Summary
In existing technologies for high-dimensional light field display, traditional compressed light field decomposition algorithms require the calculation of a large number of light rays, resulting in large memory usage and long processing time, making it difficult to meet the requirements of fast light field display of high-resolution complex scenes.
By constructing a high-dimensional light field pre-decomposition unit and/or a high-dimensional light field reverse transmission unit, the correlation between video key frames and non-key frames is utilized, and the first optimized value of the video key frame nearest to the non-key frame is used as the second iterative initial value of the compressed light field, thereby reducing the number of iterations and improving the decomposition speed.
It achieves fast compression decomposition of high-dimensional light field videos, improves decomposition efficiency, can handle higher resolution and more complex objects, overcomes the limitations of traditional methods, and improves the signal-to-noise ratio.
Smart Images

Figure CN119484867B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for rapid compression and decomposition of high-dimensional light field videos, and belongs to the technical field of light field display. Background Art
[0002] Compressed light field display (CLD) is an important naked-eye 3D display technology. It utilizes multi-layer flat-panel displays to construct high-dimensional light field information for optical reconstruction of 3D images. CLD systems are attracting considerable attention due to their ability to provide binocular parallax information and continuous motion parallax information, along with their high resolution, simple structure, and excellent compatibility.
[0003] In compressed light field display technology, compressing and decomposing high-dimensional light field information into multi-layered images is a challenging task. High-dimensional light field compression decomposition algorithms require recording a large amount of light information and solving for the intersection positions and pixel values between the light rays and each display layer based on their brightness and direction. Because the number of input light equations far exceeds the number of pixels to be solved on each screen layer, traditional compressed light field technology typically formulates the problem as solving an overdetermined system of linear equations. These equations typically require optimization algorithms to solve.
[0004] Current optimization algorithms fall into four main categories: least squares, nonnegative tensor decomposition, tomography, and machine learning. However, these optimization algorithms typically require simultaneous calculations of all collected light rays during the light field decomposition process. Even a low-resolution scene can involve hundreds of millions of light rays. Simultaneously calculating these rays inevitably consumes significant memory and takes a significant amount of time, making it difficult to meet the requirements for fast light field display of high-resolution, complex scenes. The information disclosed in this background technology is provided solely for understanding the context of the present invention and may therefore include information that does not constitute prior art. Summary of the Invention
[0005] In response to the above problem or one of the above problems, an object of the present invention is to provide a method and system for rapid compression and decomposition of high-dimensional light field videos, which uses the first optimized value of the video key frame nearest to the video non-key frame as the second iterative initial value of the compressed light field, thereby utilizing the correlation between the video key frame and the non-key frame to quickly obtain the iterative initial value of the non-key frame, thereby significantly reducing the number of iterations of the light field video, improving the compression and decomposition speed of the dynamically compressed light field video information, and further significantly improving the compression and decomposition efficiency of the compressed light field video information, and achieving compression and decomposition of higher resolution and more complex objects.
[0006] In response to the above problem or one of the above problems, the second object of the present invention is to provide a method and system for rapid compression decomposition of high-dimensional light field videos. By constructing a high-dimensional light field pre-decomposition unit and / or a high-dimensional light field reverse transmission unit to calculate the initial value of the iteration, the method overcomes the limitations of traditional iterative optimization methods that easily converge to a local optimal solution and the signal-to-noise ratio of the displayed results is low, thereby achieving high-quality compression decomposition of high-dimensional light fields.
[0007] In response to the above problem or one of the above problems, the third object of the present invention is to provide a method and system for rapid compression and decomposition of high-dimensional light field videos. By constructing a high-dimensional light field pre-decomposition unit and / or a high-dimensional light field reverse transmission unit to calculate the initial value of the iteration, the problems of the traditional iterative optimization method with a large number of iterations and long optimization time are further solved, and the rapid compression and decomposition of the high-dimensional light field is achieved.
[0008] To achieve one of the above purposes, the first technical solution of the present invention is:
[0009] A method for rapid compression and decomposition of high-dimensional light field videos, comprising the following steps:
[0010] Step 1: Using a pre-built light field acquisition model, the high-dimensional light field information of the three-dimensional object is temporally and spatially sampled to obtain a high-dimensional light field information video of the three-dimensional object;
[0011] Step 2: Using a pre-built first frame extraction model, the high-dimensional light field information video is processed to extract the video key frame and obtain the first iterative initial value of the compressed light field of the three-dimensional object at the corresponding moment;
[0012] Step 3: Using a pre-built first optimization simulation model, performing iterative optimization using a first iteration initial value, and obtaining a first optimized value of the compressed light field at a time corresponding to a key frame of the video;
[0013] Step 4: Use the pre-built second frame extraction model to extract the video non-key frame, and use the first optimized value of the video key frame that is the nearest neighbor to the video non-key frame as the second iteration initial value of the compressed light field;
[0014] Step 5: Using the pre-built second optimization simulation model, performing iterative optimization using the second iteration initial value, and obtaining a second optimized value of the compressed light field at a corresponding moment of a non-key frame of the video;
[0015] Step six: Using the pre-built object generation model, the first optimization value and the second optimization value are combined into a compressed light field video of the three-dimensional object in a time series, thereby completing the rapid compression decomposition of the high-dimensional light field video.
[0016] After continuous exploration and experimentation, the present invention constructs a light field acquisition model, a first frame extraction model, a first optimization simulation model, a second frame extraction model, a second optimization simulation model, and an object generation model. The first optimization value of the video key frame that is the nearest neighbor to the video non-key frame is used as the second iteration initial value of the compressed light field. This allows the correlation between the video key frame and the non-key frame to be used to quickly obtain the iteration initial value of the non-key frame, thereby significantly reducing the number of iterations of the light field video and improving the compression decomposition speed of the dynamically compressed light field video information. This in turn significantly improves the compression decomposition efficiency of the compressed light field video information, and can achieve compression decomposition of higher resolution and more complex objects.
[0017] As preferred technical measures:
[0018] Step 1: Use a pre-built light field acquisition model to perform spatiotemporal sampling on the high-dimensional light field information of the three-dimensional object. The method for obtaining a high-dimensional light field information video of the three-dimensional object is as follows:
[0019] Get the dimension information (u, v) of the spatial light information variable;
[0020] According to the dimension information (u, v), set the position coordinates of each camera unit;
[0021] Establishing a camera array according to the position coordinates of each camera unit;
[0022] The dynamic three-dimensional object is captured by a camera array, and spatiotemporal sampling is performed to obtain color image information of the three-dimensional object;
[0023] Calculating and obtaining depth information of the three-dimensional object based on the color image information of the three-dimensional object;
[0024] According to the depth information, a high-dimensional light field information video is obtained.
[0025] As preferred technical measures:
[0026] Spatiotemporal sampling is the sampling of the 5-dimensional light field function that describes a 3D object;
[0027] The 5-dimensional light field function is L(u, v, s, t, T);
[0028] Among them, u, v, s, t are spatial light information variables, and T is the time information variable;
[0029] Or / and, the camera array includes a physical camera array or a virtual digital camera array.
[0030] As preferred technical measures:
[0031] Step 2: Using the pre-built first frame extraction model, the high-dimensional light field information video is processed to extract the video key frame and obtain the first iterative initial value of the compressed light field of the three-dimensional object at that moment. The method is as follows:
[0032] By using the mapping law between the decomposition results of the compressed light field and the three-dimensional object, a high-dimensional light field pre-decomposition unit is constructed to pre-decompose the high-dimensional light field information video to obtain the first iterative initial value of the compressed light field;
[0033] Alternatively, a high-dimensional light field reverse transmission unit is constructed by utilizing the mapping law between spatial light of the high-dimensional light field and the three-dimensional object, and the image on the compressed light field plane is refocused to obtain the first iterative initial value of the compressed light field;
[0034] Alternatively, a large-aperture camera is used to record video key frames, and images captured by focusing the camera on different depth planes are used as initial values for the first iteration, where the plane focused by the camera corresponds to the information display plane of the compressed light field.
[0035] As preferred technical measures:
[0036] The method of constructing a high-dimensional light field pre-decomposition unit is as follows:
[0037] Extract a key frame and obtain the RGBD information recorded by the camera array at the corresponding moment, which includes the RGB information and D information of the object point;
[0038] Obtaining coordinate information of array units in the camera array, obtaining position information and resolution information of multiple layers of the compressed light field, and establishing a world coordinate system;
[0039] Based on the world coordinate system, the RGB information of the object point is decomposed into the layer of the compressed light field adjacent to the object point according to the D information of the object point;
[0040] When there are two adjacent layers of compressed light fields around an object point, the decomposition ratio between the two adjacent layers of compressed light fields is inversely proportional to the nth power of the ratio of the distance between the object point and the layers to the spacing between the two layers, where n is greater than or equal to 1;
[0041] When there is only one adjacent compressed light field layer around the object point, the RGB information of the object point is decomposed only on this layer, and the decomposition ratio is inversely proportional to the ratio of the distance between the object point and the layer to the nth power of the distance between the two layers, where n is greater than or equal to 1;
[0042] The RGB values decomposed into each layer of the compressed light field are accumulated according to the weights and the pre-decomposition results are output;
[0043] The pre-decomposition result is used as the iterative initial value of the compressed light field of the three-dimensional object at that moment.
[0044] As preferred technical measures:
[0045] The method for constructing a high-dimensional light field reverse transmission unit is as follows:
[0046] Obtain a key frame, and based on the key frame, obtain the light field information L(u, v, s, t) recorded by the camera array at the corresponding moment;
[0047] Obtain the position and resolution information of multiple layers of the compressed light field and construct the compressed light field L'(u',v',s',t');
[0048] There is a certain distance between the light field information L and the compressed light field L', and a world coordinate system XYZ is established, where the X coordinate axis is parallel to the u coordinate axis, and the Y coordinate axis is parallel to the v coordinate axis;
[0049] Reversely transmit the light in the light field information L to each plane of the compressed light field L';
[0050] The pixel value on each plane of the compressed light field is equal to the weighted average of the brightness of all reverse-transmitting rays that intersect with each plane;
[0051] The weighted average value is used as the first iteration initial value of the compressed light field of the three-dimensional object at that moment.
[0052] As preferred technical measures:
[0053] Step 3: Using the pre-built first optimization simulation model and the first iteration initial value to perform iterative optimization, the method for obtaining the first optimized value of the compressed light field at the time corresponding to the video key frame is as follows:
[0054] Based on the characteristics of the initial value of the first iteration, select an optimization algorithm from the gradient descent algorithm, stochastic gradient descent algorithm, mini-batch stochastic gradient descent method, and momentum method;
[0055] Based on the optimization algorithm, an optimization iteration target is set, and an iterative solver is established to form an iterative solver based on non-negative matrix decomposition or an iterative solver based on high-dimensional linear least squares method, which is used to iteratively optimize the initial value of the first iteration;
[0056] The optimization iteration objectives include the signal-to-noise ratio of the image and the number of iterations;
[0057] Using an iterative solver, the optimized values of multiple layers of information of the compressed light field at the corresponding moment of the video key frame are obtained;
[0058] The optimized value is the optimized value of the layer's transmittance, reflectivity, brightness or polarization parameter;
[0059] A first optimized value of the compressed light field is established according to the optimized value of the transmittance, reflectivity, brightness or polarization parameter of the layer.
[0060] As preferred technical measures:
[0061] Step 5: Using the pre-built second optimization simulation model, performing iterative optimization using the second iteration initial value, and obtaining a second optimized value of the compressed light field at a time corresponding to a non-key frame of the video;
[0062] According to the characteristics of the initial value of the second iteration, select an optimization algorithm from the gradient descent algorithm, stochastic gradient descent algorithm, mini-batch stochastic gradient descent method, and momentum method;
[0063] Based on the optimization algorithm and setting the optimization iteration target, an iterative solver is established to form an iterative solver based on non-negative matrix decomposition or an iterative solver based on high-dimensional linear least squares method, which is used to iteratively optimize the second iteration initial value;
[0064] The optimization iteration objectives include the signal-to-noise ratio of the image and the number of iterations;
[0065] Using an iterative solver, the optimized values of multiple layers of information of the compressed light field corresponding to the non-key frame of the video are obtained;
[0066] The optimized value is the optimized value of the layer's transmittance, reflectivity, brightness or polarization parameter;
[0067] A second optimized value of the compressed light field is established according to the optimized value of the transmittance, reflectivity, brightness or polarization parameter of the layer.
[0068] To achieve one of the above purposes, the second technical solution of the present invention is:
[0069] A method for rapid compression and decomposition of high-dimensional light field videos comprises the following steps:
[0070] Step 1: Perform spatiotemporal sampling on the high-dimensional light field information of the dynamic three-dimensional object to obtain a high-dimensional light field information video of the three-dimensional object;
[0071] Step 2: extract the video key frame and obtain the iterative initial value of the compressed light field of the three-dimensional object at the corresponding moment;
[0072] Step 3: Input the initial value of the iteration into the iterative solver for iterative optimization to obtain the optimized value of the compressed light field of the three-dimensional object at the time corresponding to the key frame;
[0073] Step 4: extract a non-key frame from the video, use the optimized value of the compressed light field of the high-dimensional light field video frame at the nearest time to a non-key frame as the iterative initial value of the compressed light field of the non-key frame, input the initial value into the iterative solver for iterative optimization, and obtain the optimized value of the compressed light field of the three-dimensional object at the time corresponding to the non-key frame;
[0074] Step 5: Combining the optimized values of the compressed light field in steps 3 and 4 into a compressed light field video of a three-dimensional object in a time sequence for output and display.
[0075] After continuous exploration and experimentation, the present invention uses the first optimized value of the video key frame that is the nearest neighbor to the video non-key frame as the second iteration initial value of the compressed light field. This can utilize the correlation between the video key frame and the non-key frame to quickly obtain the iteration initial value of the non-key frame, thereby significantly reducing the number of iterations of the light field video, improving the compression decomposition speed of the dynamic compressed light field video information, and further significantly improving the compression decomposition efficiency of the compressed light field video information, and can achieve compression decomposition of higher resolution and more complex objects.
[0076] To achieve one of the above purposes, the third technical solution of the present invention is:
[0077] A fast compression and decomposition system for high-dimensional light field videos, comprising a light field acquisition module, a first frame extraction module, a first optimization and simulation module, a second frame extraction module, a second optimization and simulation module, and an object generation module;
[0078] The light field acquisition module performs spatiotemporal sampling on the high-dimensional light field information of the three-dimensional object to obtain the high-dimensional light field information video of the three-dimensional object;
[0079] A first frame extraction module processes the high-dimensional light field information video, extracts the video key frames, and obtains the first iterative initial value of the compressed light field of the three-dimensional object;
[0080] A first optimization simulation module performs iterative optimization using a first iterative initial value to obtain a first optimized value of the compressed light field at a time corresponding to a key frame of the video;
[0081] A second frame extraction module extracts a non-key frame of the video and uses the first optimized value of the key frame of the video that is the nearest neighbor to the non-key frame of the video as the second iterative initial value of the compressed light field;
[0082] A second optimization simulation module performs iterative optimization using the second iterative initial value to obtain a second optimized value of the compressed light field of the non-key frame of the video;
[0083] The object generation module combines the first optimization value and the second optimization value into a compressed light field video of a three-dimensional object in a time sequence, thereby completing the rapid compression decomposition of the high-dimensional light field video.
[0084] After continuous exploration and experimentation, the present invention sets a light field acquisition module, a first frame extraction module, a first optimization simulation module, a second frame extraction module, a second optimization simulation module, and an object generation module, and uses the first optimization value of the video key frame that is closest to the video non-key frame as the second iteration initial value of the compressed light field. In this way, the correlation between the video key frame and the non-key frame can be used to quickly obtain the iteration initial value of the non-key frame, thereby significantly reducing the number of iterations of the light field video, improving the compression decomposition speed of the dynamic compressed light field video information, and thus significantly improving the compression decomposition efficiency of the compressed light field video information, and achieving higher resolution and compression decomposition of more complex objects.
[0085] To achieve one of the above purposes, the fourth technical solution of the present invention is:
[0086] An electronic device comprising:
[0087] one or more processors;
[0088] a storage device for storing one or more programs;
[0089] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned method for fast compression and decomposition of high-dimensional light field videos.
[0090] To achieve one of the above objectives, the fifth technical solution of the present invention is:
[0091] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned method for rapid compression and decomposition of a high-dimensional light field video.
[0092] Compared with the existing technical solutions, the present invention has the following beneficial effects:
[0093] After continuous exploration and experimentation, the present invention uses the first optimized value of the video key frame that is the nearest neighbor to the video non-key frame as the second iteration initial value of the compressed light field. This can utilize the correlation between the video key frame and the non-key frame to quickly obtain the iteration initial value of the non-key frame, thereby significantly reducing the number of iterations of the light field video, improving the compression decomposition speed of the dynamic compressed light field video information, and further significantly improving the compression decomposition efficiency of the compressed light field video information, and can achieve compression decomposition of higher resolution and more complex objects.
[0094] Furthermore, the present invention overcomes the limitations of traditional iterative optimization methods, which easily converge to local optimal solutions and have low signal-to-noise ratios of displayed results, by constructing a high-dimensional light field pre-decomposition unit and / or a high-dimensional light field reverse transmission unit to calculate the iterative initial value, thereby achieving high-quality compressed decomposition of high-dimensional light fields.
[0095] The present invention further solves the problems of large number of iterations and long optimization time in traditional iterative optimization methods by constructing a high-dimensional light field pre-decomposition unit and / or a high-dimensional light field reverse transmission unit to calculate the iteration initial value, and realizes the rapid compression decomposition of the high-dimensional light field. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] Figure 1 A schematic diagram of a process for performing rapid compression decomposition using a light field pre-decomposition method according to the present invention;
[0097] Figure 2 A schematic diagram of the light field pre-decomposition method of the present invention;
[0098] Figure 3 A schematic diagram of a process for performing rapid compression and decomposition using a reverse transmission method according to the present invention;
[0099] Figure 4 A schematic diagram of the reverse transmission method of the present invention;
[0100] Figure 5 A schematic diagram of a process for rapid compression decomposition using a camera multi-plane focusing method according to the present invention;
[0101] Figure 6 A schematic diagram of the multi-plane focusing method of the camera of the present invention. DETAILED DESCRIPTION
[0102] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0103] Rather, the present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention as defined by the claims. Furthermore, to facilitate a better understanding of the present invention, certain specific details are described in detail below in the detailed description of the present invention. Those skilled in the art will be able to fully understand the present invention without these details.
[0104] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention pertains. The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. As used herein, the term "or / and" includes any and all combinations of one or more of the associated listed items.
[0105] The first specific embodiment of the method for rapid compression and decomposition of high-dimensional light field video of the present invention is as follows:
[0106] A method for rapid compression and decomposition of high-dimensional light field videos, comprising the following steps:
[0107] Step 1: Using a pre-built light field acquisition model, the high-dimensional light field information of the three-dimensional object is temporally and spatially sampled to obtain a high-dimensional light field information video of the three-dimensional object;
[0108] Step 2: Using a pre-built first frame extraction model, the high-dimensional light field information video is processed to extract the video key frame and obtain the first iterative initial value of the compressed light field of the three-dimensional object at the corresponding moment;
[0109] Step 3: Using a pre-built first optimization simulation model, performing iterative optimization using a first iteration initial value, and obtaining a first optimized value of the compressed light field at a time corresponding to a key frame of the video;
[0110] Step 4: Use the pre-built second frame extraction model to extract the video non-key frame, and use the first optimized value of the video key frame that is the nearest neighbor to the video non-key frame as the second iteration initial value of the compressed light field;
[0111] Step 5: Using the pre-built second optimization simulation model, performing iterative optimization using the second iteration initial value, and obtaining a second optimized value of the compressed light field at a corresponding moment of a non-key frame of the video;
[0112] Step six: Using the pre-built object generation model, the first optimization value and the second optimization value are combined into a compressed light field video of the three-dimensional object in a time series, thereby completing the rapid compression decomposition of the high-dimensional light field video.
[0113] A second specific embodiment of the method for rapid compression and decomposition of high-dimensional light field video of the present invention:
[0114] A method for rapid compression and decomposition of high-dimensional light field videos comprises the following steps:
[0115] Step 1: Perform spatiotemporal sampling on the high-dimensional light field information of the dynamic three-dimensional object to obtain a high-dimensional light field information video of the three-dimensional object;
[0116] Step 2: extract the video key frame and obtain the iterative initial value of the compressed light field of the three-dimensional object at that moment;
[0117] Step 3: Input the initial iteration value into the iterative solver for iterative optimization to obtain the optimized value of the compressed light field of the three-dimensional object at the time corresponding to the key frame.
[0118] Step 4: Extract a non-key frame from the video, use the optimized value of the compressed light field of the high-dimensional light field video frame at the nearest time to a non-key frame as the iterative initial value of the compressed light field of the frame, input the initial value into the iterative solver for iterative optimization, and obtain the optimized value of the compressed light field of the three-dimensional object at the time corresponding to the non-key frame.
[0119] Step 5: Combining the optimized values of the compressed light field in steps 3 and 4 into a compressed light field video of a three-dimensional object in a time sequence for output and display.
[0120] In this embodiment, the spatiotemporal sampling described in step 1 is the sampling of a 5-dimensional light field function describing a dynamic 3D object. The 5-dimensional light field function is described as L(u, v, s, t, T). This includes the sampling of the spatial light information variables (u, v, s, t) and the sampling of the time information variable T.
[0121] The spatiotemporal sampling process can be achieved by capturing video of a dynamic three-dimensional object using a camera array. The position coordinates of each camera unit in the camera array correspond to the (u, v) dimensions of the spatial light information variable. The pixel coordinates of the video captured by the camera array correspond to the (s, t) dimensions. The capture time of all cameras in the camera array is synchronized, corresponding to the variable T.
[0122] The camera array includes a physical camera array and a virtual digital camera array. The cameras in the array preferably record color image information and depth information of a three-dimensional object, namely, RGBD information. The cameras in the array record the color image information of the three-dimensional object and obtain the depth information of the three-dimensional object through calculation.
[0123] In this embodiment, the method for obtaining the iterative initial value of the compressed light field of the three-dimensional object at this moment described in step 2 includes: (1) a high-dimensional light field pre-decomposition method, (2) a high-dimensional light field reverse transmission method, and (3) a camera multi-plane focusing method.
[0124] (1) The high-dimensional light field pre-decomposition method includes the following contents:
[0125] The high-dimensional light field is pre-decomposed using the mapping law between the decomposition results of the compressed light field and the three-dimensional object, which includes the following steps:
[0126] Input the RGBD information recorded by the camera array at the time corresponding to a key frame.
[0127] Input the coordinate information of the array elements in the camera array.
[0128] Input position and resolution information for multiple layers of a compressed light field.
[0129] A world coordinate system is established, and the RGB information of the object point recorded by each camera is decomposed onto the compressed light field layer adjacent to the object point based on the object point's D information. When there are two adjacent compressed light field layers around the object point, the decomposition ratio between the two adjacent compressed light field layers is inversely proportional to the ratio of the distance from the object point to the layer to the distance between the two layers to the nth power, where n is greater than or equal to 1. When there is only one adjacent layer around the object point, the RGB information of the object point is decomposed only on this one layer, and its decomposition ratio is inversely proportional to the ratio of the distance from the object point to the layer to the distance between the two layers to the nth power, where n is greater than or equal to 1.
[0130] The RGB values decomposed into each compressed light field layer are summed up according to the weights and then the pre-decomposition result is output. The pre-decomposition result is the iterative initial value of the compressed light field of the three-dimensional object at that moment. (2) The reverse transmission method of the high-dimensional light field includes the following contents:
[0131] The image on the compressed light field plane is refocused using the mapping law between the spatial light of the high-dimensional light field and the three-dimensional object. The process includes the following:
[0132] Input the light field information L(u, v, s, t) recorded by the light field camera at the time corresponding to a key frame.
[0133] The position and resolution information of multiple layers of the compressed light field are input to construct the compressed light field L'(u',v',s',t').
[0134] Set the distance between light fields L and L' and establish a world coordinate system (XYZ), where the X axis is parallel to the u axis and the Y axis is parallel to the v axis. Reverse-transmit light rays from light field L to the planes of the compressed light field L'. The pixel value on each plane of the compressed light field is equal to the weighted average of the brightness of all reverse-transmitted rays that intersect each plane. This weighted average becomes the initial value of the iteration of the compressed light field of the 3D object at that moment.
[0135] (3) The camera multi-plane focusing method includes the following contents:
[0136] Images captured by a large-aperture camera focused on different depth planes are used as initial values for iteration, where the plane on which the camera is focused corresponds to the information display plane of the compressed light field. In this embodiment, the iterative solver described in step 3 includes an iterative solver based on non-negative matrix factorization or an iterative solver based on high-dimensional linear least squares. The solver can employ an optimization algorithm, including a gradient descent algorithm, a stochastic gradient descent algorithm, a mini-batch stochastic gradient descent method, a momentum method, and the like.
[0137] In this embodiment, the optimized values of the compressed light field described in steps 3 and 4 include optimized values of information about multiple layers of the compressed light field. The physical meaning of the optimized values of the layer information can be optimized values of transmittance, reflectivity, brightness, or polarization parameters of the layer. The compressed light field can have at least two layers.
[0138] In this embodiment, the optimized values of the compressed light field in step 3 and step 4 are determined by a set optimization iteration target, which includes the signal-to-noise ratio of the image and the number of iterations.
[0139] In this embodiment, the compressed light field video described in step 5 includes both the video generated in real time through the process of steps 1 to 4 and the video file generated offline through the process of steps 1 to 4. The present invention utilizes pre-decomposition and / or reverse propagation techniques to calculate initial iteration values, overcoming the limitations of traditional iterative optimization methods, which tend to converge to a local optimal solution and exhibit a low signal-to-noise ratio (SNR) in the displayed results. This achieves high-quality compressed decomposition of high-dimensional light fields.
[0140] The present invention adopts pre-decomposition technology and / or reverse transmission technology to calculate the initial value of iteration, solves the problems of large number of iterations and long optimization time in traditional iterative optimization methods, and realizes rapid compression decomposition of high-dimensional light fields.
[0141] The present invention optimizes the initial value by utilizing the correlation of the compressed light field between non-key frames, greatly reduces the number of iterations of the light field video, and improves the compression decomposition speed of the dynamic compressed light field video information.
[0142] The present invention significantly improves the compression decomposition efficiency of compressed light field video information, and can achieve compression decomposition of higher resolution and more complex objects.
[0143] See also Figure 1 、 Figure 2 , the third specific embodiment of the method for rapid compression and decomposition of high-dimensional light field video of the present invention:
[0144] A fast compression decomposition method for high-dimensional light field videos, including the following contents:
[0145] First, a camera array is used to capture the 5-dimensional light field function L(u, v, s, t, T) of a dynamic 3D object to complete the spatiotemporal sampling of the high-dimensional light field. Keyframes are then set for the high-dimensional light field video, and the keyframes and non-keyframes are processed separately. A light field pre-decomposition method is used to determine the iterative initial values of the decompressed light field for the keyframes. This initial value is then used with an iterative solver based on non-negative matrix factorization for iterative optimization, using a stochastic gradient descent algorithm to obtain the optimized value of the compressed light field for the keyframes. For non-keyframes, the optimized value of the compressed light field of their nearest neighboring keyframes is used as the initial value. This initial value is then used with an iterative solver based on non-negative matrix factorization for iterative optimization, using the momentum method, to obtain the optimized value of the compressed light field for the non-keyframes. Finally, the optimized values of the compressed light fields of the keyframes and non-keyframes are combined in a time series to form a compressed light field video of the 3D object for output and display.
[0146] like Figure 2 The diagram shows the light field pre-decomposition method in this embodiment. First, a world coordinate system XYZ is established, where the dotted hexagon represents the three-dimensional scene to be imaged. The light field camera is located at the Z=0 plane, and each camera unit x on this plane is i , i∈(1,……,n) records the RGBD information of a certain perspective of the three-dimensional scene at that moment. The compressed light field is composed of two display layers, where the first display layer is located in the Z=l plane, and the coordinates on this plane are represented by (u',v'), and the second display layer is located in the Z=l+h plane, and the coordinates on this plane are represented by (s',t'). Taking the two-dimensional light field coordinates as an example, for the pre-decomposed values f(u' i ) and f(s' i ) are:
[0147]
[0148] Among them I xi (u i ') represents the xth i The camera records the first display layer u' i The brightness value of the object point corresponding to the coordinate, z xi (u i ') represents the depth value of the object point. xi (s i ') represents the xth i The camera records the second display layer s' i The brightness value of the object point corresponding to the coordinate, z xi (s i ') represents the depth value of the object point.
[0149] Taking the two-dimensional light field coordinates as an example, in another implementation of this embodiment, the pre-decomposed values f(u' i ) and f(s' i ) are defined as:
[0150]
[0151] See also Figure 3 、 Figure 4 , the fourth specific embodiment of the method for rapid compression and decomposition of high-dimensional light field video of the present invention:
[0152] A fast compression decomposition method for high-dimensional light field videos, including the following contents:
[0153] First, a light field camera is used to capture the 5-dimensional light field function L(u, v, s, t, T) of a dynamic 3D object to complete the spatiotemporal sampling of the high-dimensional light field. Keyframes are then set for the high-dimensional light field video, and the keyframes and non-keyframes are processed separately. For the keyframes, the high-dimensional light field backpropagation method is used to determine the iterative initial values of their decompressed light fields. This initial value is then used with an iterative solver based on a high-dimensional linear least squares method for iterative optimization, employing a gradient descent algorithm to obtain the optimized value of the keyframe's compressed light field. For non-keyframes, the optimized value of the compressed light field of their nearest neighboring keyframes is used as the initial value. This initial value is then used with an iterative solver based on a high-dimensional linear least squares method for iterative optimization, employing a mini-batch stochastic gradient descent method to obtain the optimized value of the non-keyframe's compressed light field. Finally, the optimized values of the compressed light fields of the keyframes and non-keyframes are combined in a temporal sequence to form a compressed light field video of the 3D object for output and display.
[0154] like Figure 4 The figure shows a schematic diagram of the reverse transmission method of the high-dimensional light field in this embodiment. First, a world coordinate system XYZ is established. The light field information recorded by the light field camera corresponding to the key frame at a certain moment is L(u, v, s, t). The compressed light field to be constructed is L'(u', v', s', t'). The distances Z1 and Z2 between the light fields L and L' are set, and the light in the light field L is reversely transmitted to each plane of the compressed light field L' through the light transmission matrix. The pixel value on each plane of the compressed light field is equal to the weighted average of the brightness of all reversely transmitted light rays intersecting with each plane. The weighted average is the iterative initial value of the compressed light field of the three-dimensional object at that moment:
[0155]
[0156] Among them L i (s,t) and L i(u, v) is the brightness of the light recorded by the i-th camera that is transmitted to the same pixel on the (s, t) or (u, v) plane.
[0157] See also Figure 5 、 Figure 6 , a fifth specific embodiment of the method for rapid compression and decomposition of high-dimensional light field video of the present invention:
[0158] A fast compression decomposition method for high-dimensional light field videos, including the following contents:
[0159] First, a camera array is used to capture the 5-dimensional light field function L(u, v, s, t, T) of a dynamic 3D object, completing the spatiotemporal sampling of the high-dimensional light field. Keyframes are then set for the high-dimensional light field video, and the keyframes and non-keyframes are processed separately. For the keyframes, the camera multi-plane focusing method is used to determine the iterative initial values of the decompressed light field. This initial value is then used with an iterative solver based on non-negative matrix factorization for iterative optimization, using a gradient descent algorithm to obtain the optimized value of the compressed light field for the keyframes. For non-keyframes, the optimized value of the compressed light field of their nearest neighboring keyframes is used as the initial value. This initial value is then used with an iterative solver based on non-negative matrix factorization for iterative optimization, using a mini-batch stochastic gradient descent algorithm to obtain the optimized value of the compressed light field for the non-keyframes. Finally, the optimized values of the compressed light fields of the keyframes and non-keyframes are combined in a time series to form a compressed light field video of the 3D object for output and display.
[0160] like Figure 6 The figure shows a schematic diagram of the multi-plane focusing method of the camera in this embodiment. Figure 6 As shown, the camera takes a picture of a 3D image. Focus the camera on the information plane (s', t') of the compressed light field and take a picture, which serves as the initial value for the information plane (s', t'). Focus the camera on the information plane (u', v') of the compressed light field and take another picture, which serves as the initial value for the information plane (u', v').
[0161] An embodiment of a device applying the method of the present invention:
[0162] An electronic device comprising:
[0163] one or more processors;
[0164] a storage device for storing one or more programs;
[0165] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned method for fast compression and decomposition of high-dimensional light field videos.
[0166] A computer medium embodiment of the method of the present invention:
[0167] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned method for rapid compression and decomposition of a high-dimensional light field video.
[0168] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, and computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, optical storage, etc.) containing computer-usable program code.
[0169] The present application is described in terms of flowcharts or / and block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process or / and block in the flowchart or / and block diagram and the combination of the processes or / and blocks in the flowchart or / and block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0170] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0172] The model in this application is an object that objectively describes the morphological structure with the help of physical or virtual representation. The object is not equal to the physical body and is not limited to physical and virtual. It can be a data processing function, software program, processing mode, usage method, operation method, workflow, application process, electronic hardware, circuit module, processing system, system imitation or simulation object.
[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field can still modify or replace the specific implementation methods of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be included in the scope of protection of the claims of the present invention.
Claims
1. A fast compression decomposition method for high-dimensional light field video, characterized by: The following steps are involved: Step 1: Using a pre-built light field acquisition model, the high-dimensional light field information of the three-dimensional object is temporally and spatially sampled to obtain a high-dimensional light field information video of the three-dimensional object; Step 2: Using a pre-built first frame extraction model, the high-dimensional light field information video is processed to extract the video key frame and obtain the first iterative initial value of the compressed light field of the three-dimensional object at the corresponding moment; Step 3: Using a pre-built first optimization simulation model, performing iterative optimization using a first iteration initial value, and obtaining a first optimized value of the compressed light field at a time corresponding to a key frame of the video; Step 4: Use the pre-built second frame extraction model to extract the video non-key frame, and use the first optimized value of the video key frame that is the nearest neighbor to the video non-key frame as the second iteration initial value of the compressed light field; Step 5: Using the pre-built second optimization simulation model, performing iterative optimization using the second iteration initial value, and obtaining a second optimized value of the compressed light field at a corresponding moment of a non-key frame of the video; Step six: Using the pre-built object generation model, the first optimization value and the second optimization value are combined into a compressed light field video of the three-dimensional object in a time series, thereby completing the rapid compression decomposition of the high-dimensional light field video.
2. The method for rapid compression and decomposition of high-dimensional light field video according to claim 1, wherein: Step 1: Use a pre-built light field acquisition model to perform spatiotemporal sampling on the high-dimensional light field information of the three-dimensional object. The method for obtaining a high-dimensional light field information video of the three-dimensional object is as follows: Get the dimension information (u, v) of the spatial light information variable; According to the dimension information (u, v), set the position coordinates of each camera unit; Establishing a camera array according to the position coordinates of each camera unit; The dynamic three-dimensional object is captured by a camera array, and spatiotemporal sampling is performed to obtain color image information of the three-dimensional object; Calculating and obtaining depth information of the three-dimensional object based on the color image information of the three-dimensional object; According to the depth information, a high-dimensional light field information video is obtained.
3. The method for rapid compression and decomposition of high-dimensional light field video according to claim 2, wherein: Spatiotemporal sampling is the sampling of the 5-dimensional light field function that describes a 3D object; The 5-dimensional light field function is L(u, v, s, t, T); Among them, u, v, s, t are spatial light information variables, and T is the time information variable; Or / and, the camera array includes a physical camera array or a virtual digital camera array.
4. The method for rapid compression and decomposition of high-dimensional light field video according to claim 1, wherein: Step 2: Using the pre-built first frame extraction model, the high-dimensional light field information video is processed to extract the video key frame and obtain the first iterative initial value of the compressed light field of the three-dimensional object at that moment. The method is as follows: By using the mapping law between the decomposition results of the compressed light field and the three-dimensional object, a high-dimensional light field pre-decomposition unit is constructed to pre-decompose the high-dimensional light field information video to obtain the first iterative initial value of the compressed light field; Alternatively, a high-dimensional light field reverse transmission unit is constructed by utilizing the mapping law between spatial light of the high-dimensional light field and the three-dimensional object, and the image on the compressed light field plane is refocused to obtain the first iterative initial value of the compressed light field; Alternatively, a large-aperture camera is used to record video key frames, and images captured by focusing the camera on different depth planes are used as initial values for the first iteration, where the plane focused by the camera corresponds to the information display plane of the compressed light field.
5. The method for rapid compression and decomposition of high-dimensional light field video according to claim 4, characterized in that: The method of constructing a high-dimensional light field pre-decomposition unit is as follows: Extract a key frame and obtain the RGBD information recorded by the camera array at the corresponding moment, which includes the RGB information and D information of the object point; Obtaining coordinate information of array units in the camera array, obtaining position information and resolution information of multiple layers of the compressed light field, and establishing a world coordinate system; Based on the world coordinate system, the RGB information of the object point is decomposed into the layer of the compressed light field adjacent to the object point according to the D information of the object point; When there are two adjacent layers of compressed light fields around an object point, the decomposition ratio between the two adjacent layers of compressed light fields is inversely proportional to the nth power of the ratio of the distance between the object point and the layers to the spacing between the two layers, where n is greater than or equal to 1; When there is only one adjacent compressed light field layer around the object point, the RGB information of the object point is decomposed only on this layer, and the decomposition ratio is inversely proportional to the ratio of the distance between the object point and the layer to the nth power of the distance between the two layers, where n is greater than or equal to 1; The RGB values decomposed into each layer of the compressed light field are accumulated according to the weights and the pre-decomposition results are output; The pre-decomposition result is used as the iterative initial value of the compressed light field of the three-dimensional object at that moment.
6. The method for rapid compression and decomposition of high-dimensional light field video according to claim 4, characterized in that: The method for constructing a high-dimensional light field reverse transmission unit is as follows: Obtain a key frame, and based on the key frame, obtain the light field information L(u, v, s, t) recorded by the camera array at the corresponding moment; Obtain the position and resolution information of multiple layers of the compressed light field and construct the compressed light field L'(u',v',s',t'); There is a certain distance between the light field information L and the compressed light field L', and the world coordinate system XYZ is established. The X coordinate axis is parallel to the u coordinate axis, and the Y coordinate axis is parallel to the v coordinate axis; Reversely transmit the light in the light field information L to each plane of the compressed light field L'; The pixel value on each plane of the compressed light field is equal to the weighted average of the brightness of all reverse-transmitting rays that intersect with each plane; The weighted average value is used as the first iteration initial value of the compressed light field of the three-dimensional object at that moment.
7. The method for rapid compression and decomposition of high-dimensional light field video according to claim 1, wherein: Step 3: Using the pre-built first optimization simulation model and the first iteration initial value to perform iterative optimization, the method for obtaining the first optimized value of the compressed light field at the time corresponding to the video key frame is as follows: Based on the characteristics of the initial value of the first iteration, select an optimization algorithm from the gradient descent algorithm, stochastic gradient descent algorithm, mini-batch stochastic gradient descent method, and momentum method; Based on the optimization algorithm, an optimization iteration target is set, and an iterative solver is established to form an iterative solver based on non-negative matrix decomposition or an iterative solver based on high-dimensional linear least squares method, which is used to iteratively optimize the initial value of the first iteration; The optimization iteration objectives include the signal-to-noise ratio of the image and the number of iterations; Using an iterative solver, the optimized values of multiple layers of information of the compressed light field at the corresponding moment of the video key frame are obtained; The optimized value is the optimized value of the layer's transmittance, reflectivity, brightness or polarization parameter; A first optimized value of the compressed light field is established according to the optimized value of the transmittance, reflectivity, brightness or polarization parameter of the layer.
8. The method for rapid compression and decomposition of high-dimensional light field video according to claim 1, wherein: Step 5: Using the pre-built second optimization simulation model, performing iterative optimization using the second iteration initial value, and obtaining a second optimized value of the compressed light field at a time corresponding to a non-key frame of the video; According to the characteristics of the initial value of the second iteration, select an optimization algorithm from the gradient descent algorithm, stochastic gradient descent algorithm, mini-batch stochastic gradient descent method, and momentum method; Based on the optimization algorithm and setting the optimization iteration target, an iterative solver is established to form an iterative solver based on non-negative matrix decomposition or an iterative solver based on high-dimensional linear least squares method, which is used to iteratively optimize the second iteration initial value; The optimization iteration objectives include the signal-to-noise ratio of the image and the number of iterations; Using an iterative solver, the optimized values of multiple layers of information of the compressed light field corresponding to the non-key frame of the video are obtained; The optimized value is the optimized value of the layer's transmittance, reflectivity, brightness or polarization parameter; A second optimized value of the compressed light field is established according to the optimized value of the transmittance, reflectivity, brightness or polarization parameter of the layer.
9. A method for rapid compression and decomposition of high-dimensional light field videos, characterized by: The steps include: Step 1: Perform spatiotemporal sampling on the high-dimensional light field information of the dynamic three-dimensional object to obtain a high-dimensional light field information video of the three-dimensional object; Step 2: extract the video key frame and obtain the iterative initial value of the compressed light field of the three-dimensional object at the corresponding moment; Step 3: Input the initial value of the iteration into the iterative solver for iterative optimization to obtain the optimized value of the compressed light field of the three-dimensional object at the time corresponding to the key frame; Step 4: extract a non-key frame from the video, use the optimized value of the compressed light field of the high-dimensional light field video frame at the nearest time to a non-key frame as the iterative initial value of the compressed light field of the non-key frame, input the initial value into the iterative solver for iterative optimization, and obtain the optimized value of the compressed light field of the three-dimensional object at the time corresponding to the non-key frame; Step 5: Combining the optimized values of the compressed light field in steps 3 and 4 into a compressed light field video of a three-dimensional object in a time sequence for output and display.
10. A fast compression and decomposition system for high-dimensional light field videos, characterized by: It includes a light field acquisition module, a first frame extraction module, a first optimization simulation module, a second frame extraction module, a second optimization simulation module and an object generation module; The light field acquisition module performs spatiotemporal sampling on the high-dimensional light field information of the three-dimensional object to obtain the high-dimensional light field information video of the three-dimensional object; A first frame extraction module processes the high-dimensional light field information video, extracts the video key frames, and obtains the first iterative initial value of the compressed light field of the three-dimensional object; A first optimization simulation module performs iterative optimization using a first iterative initial value to obtain a first optimized value of the compressed light field at a time corresponding to a key frame of the video; A second frame extraction module extracts a non-key frame of the video and uses the first optimized value of the key frame of the video that is the nearest neighbor to the non-key frame of the video as the second iterative initial value of the compressed light field; A second optimization simulation module performs iterative optimization using the second iterative initial value to obtain a second optimized value of the compressed light field of the non-key frame of the video; The object generation module combines the first optimization value and the second optimization value into a compressed light field video of a three-dimensional object in a time sequence, thereby completing the rapid compression decomposition of the high-dimensional light field video.
Citation Information
Patent Citations
Real-time compression and reconstruction method of Video-SAR (Synthetic Aperture Radar) image
CN105741333A
Multi-view video compression method and device based on light field, equipment and medium
CN111757125A