Anti-compression multi-view 3D light field coding method and device

By acquiring multi-view image and building mapping relationships, and using image interpolation algorithm for scaling and sampling, the color distortion and details loss caused by video color format conversion in multi-view 3D light field display is solved, and high-quality multi-view 3D light field display is achieved.

CN120343221APending Publication Date: 2025-07-18BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478288.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In multi-view 3D light field display, video color format conversion results in color distortion and loss of details in high-precision 3D coded images, affecting playback fluency and display quality.

Method used

By obtaining multi-view images with different parallax information, determining the number of views and constructing the mapping relationship between sub-pixels and sub-pixels of multi-view image in the composite graph, scaling and sampling is used for image interpolation algorithm, and accurately encoding mapping relationship is established to ensure the accurate rendering of colors and details of the image after format conversion.

Benefits of technology

It effectively solves the problem of color distortion and detail loss caused by video color format conversion, ensures that the color and details of the image are accurately presented after the format conversion, and achieves high-quality multi-viewpoint 3D light field display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343221A_ABST
    Figure CN120343221A_ABST
Patent Text Reader

Abstract

The invention discloses an anti-compression multi-view 3D light field coding method and device, and relates to the technical field of three-dimensional light field display, and the method comprises the steps: obtaining a multi-view image with different parallax information; determining the number of viewpoints, and constructing a mapping relation between sub-pixels in the composite image and sub-pixels of the multi-viewpoint image; scaling each multi-view image by using an image interpolation algorithm to enable the resolution of the scaled multi-view image to be consistent with the resolution of the composite image to be generated; and according to the mapping relation, sampling the scaled multi-view image to obtain a composite image, and completing multi-view 3D light field coding. According to different YUV sampling modes, a precise coding mapping relation is established through a specific multi-view 3D light field coding and format conversion mode, the problem that high-precision 3D coded images are displayed abnormally due to traditional video color format conversion can be effectively solved, and it is ensured that colors and details of the images are displayed accurately after format conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional light field display, and particularly to an anti-compression multi-viewpoint 3D light field encoding method and device. Background Art

[0002] With the continuous development of three-dimensional display technology, multi-viewpoint 3D light field display has received extensive attention in many fields such as film and television entertainment, virtual reality (VR), augmented reality (AR), medical surgical simulation, and industrial design because it can provide a more realistic and immersive viewing experience. In a multi-viewpoint 3D light field display system, a large number of images with different parallax information need to be processed to achieve a multi-angle stereoscopic visual effect.

[0003] However, there are many problems in the current video processing and playback process. In terms of color format conversion, common video color formats include RGB and YUV, etc. During video playback, due to video color format conversion, differences in sampling methods and data representations of different formats, it is easy to cause problems such as color distortion and detail loss in high-precision 3D encoded images, and thus they cannot be normally displayed. This seriously affects the playback smoothness and display quality of 3D videos, making users unable to obtain a good viewing experience. Summary of the Invention

[0004] Aiming at the above deficiencies in the prior art, the anti-compression multi-viewpoint 3D light field encoding method and device provided by the present invention solve the problems that high-precision 3D encoded images are prone to color distortion, detail loss, etc. due to video color format conversion, and thus cannot be normally displayed.

[0005] In order to achieve the above invention purpose, the technical solution adopted by the present invention is as follows:

[0006] Provide an anti-compression multi-viewpoint 3D light field encoding method, which includes:

[0007] Obtain multi-viewpoint images with different parallax information;

[0008] Determine the number of viewpoints, and construct a mapping relationship between sub-pixels in the synthesized image and sub-pixels in the multi-viewpoint images;

[0009] Use an image interpolation algorithm to scale each multi-viewpoint image so that the resolution of the scaled multi-viewpoint images is the same as the resolution of the synthesized image to be generated;

[0010] According to the mapping relationship, sample the scaled multi-viewpoint images to obtain a synthesized image, and complete the multi-viewpoint 3D light field encoding.

[0011] Further, the method for obtaining multi-viewpoint images with different parallax information includes:

[0012] Collect multiple images with different disparities of a target scene from different perspectives using a single or multiple cameras, obtaining multi-viewpoint images with different disparity information; among them, when a single or multiple cameras perform continuous acquisition from different perspectives, multi-viewpoint images with different disparity information corresponding to consecutive frames are obtained.

[0013] Furthermore, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images includes:

[0014] For YUV4:4:4 sampling, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images is code(i,j,k) = view[n](i,j,k), that is, the width and height of the view size remain unchanged; where code(i,j,k) is the k-th sub-pixel in the pixel at the i-th row and j-th column in the composite image, and view[n](i,j,k) is the k-th sub-pixel in the pixel at the i-th row and j-th column in the image of the n-th viewpoint; n is the number of the viewpoint.

[0015] For YUV4:2:2 sampling, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images is That is, the horizontal resolution of the view size of the U and V components is doubled, the width of the view is twice the original width, the height remains unchanged, and the view size of the Y component remains unchanged; represents rounding up;

[0016] For YUV4:4:0 sampling, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images is That is, the vertical resolution of the view size of the U and V components is doubled, the height of the view is twice the original height, the width remains unchanged, and the view size of the Y component remains unchanged;

[0017] For YUV4:2:0 sampling, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images is That is, the view size of the U and V components will be doubled in both the horizontal and vertical directions, the width and height of the view become twice the original width and original height, and the view size of the Y component remains unchanged;

[0018] For YUV4:1:1 sampling, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images is That is, the view size of the U and V components is four times the original width in the horizontal direction, the height remains unchanged, and the view size of the Y component remains unchanged.

[0019] Furthermore, the image interpolation algorithms include the nearest neighbor interpolation algorithm, the bilinear interpolation algorithm, and the bicubic interpolation algorithm.

[0020] Further, the horizontal distance from the left edge of any sub-pixel (i, j, k) in the synthesized image to the left edge of the leftmost grating unit is D p , and the distance from the left edge of this sub-pixel to the left edge of the grating unit corresponding to this sub-pixel is A p , then there are the following constraints:

[0021] D p = 3×(j - 1)+3×(i - 1)×tanθ+(k - 1)

[0022] A p = D p mod P e

[0023] where θ is the tilt angle of the cylindrical lens deviating from the vertical direction; tan represents the tangent function; P e is the width of the sub-pixel area of the synthesized image covered by the grating unit in the horizontal direction; mod represents the remainder function; the cylindrical lens is a component in the display, and a single cylindrical lens includes several grating units.

[0024] Provided is a device based on an anti-compression multi-viewpoint 3D light field encoding method, which includes:

[0025] A multi-viewpoint image acquisition module, configured to acquire multi-viewpoint images with different parallax information;

[0026] A mapping relationship construction module, configured to determine the number of viewpoints and construct the mapping relationship between the sub-pixels in the synthesized image and the sub-pixels in the multi-viewpoint images;

[0027] An image scaling module, configured to scale each multi-viewpoint image by using an image interpolation algorithm so that the resolution of the scaled multi-viewpoint images is consistent with the resolution of the to-be-generated synthesized image;

[0028] A 3D light field encoding module, configured to sample the scaled multi-viewpoint images according to the mapping relationship to obtain a synthesized image and complete the multi-viewpoint 3D light field encoding.

[0029] Further, it further includes:

[0030] A display, configured to load the synthesized image for stereoscopic display.

[0031] Further, the display includes a display unit and a cylindrical lens;

[0032] The display unit, configured to load the synthesized image;

[0033] The cylindrical lens, configured to perform light control and project each sub-pixel on the display unit into a volume pixel, thereby realizing stereoscopic display.

[0034] Provided is a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute an anti-compression multi-viewpoint 3D light field encoding method.

[0035] Provided is a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor is caused to execute an anti-compression multi-viewpoint 3D light field encoding method.

[0036] The beneficial effects of the present invention are as follows: According to different YUV sampling methods, the present invention establishes an accurate encoding mapping relationship through specific multi-viewpoint 3D light field encoding and format conversion methods, which can effectively solve the problem that the display of high-precision 3D encoded images is often abnormal due to traditional video color format conversion, and ensure that the color and details of the image are accurately presented after format conversion. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a schematic flowchart of the method;

[0038] Figure 2 is a schematic diagram of the constraint relationship between D p and A p in the embodiment;

[0039] Figure 3 is a schematic diagram of the geometric relationship in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0041] As Figure 1 shown, the anti-compression multi-viewpoint 3D light field encoding method includes:

[0042] S1. Obtain multi-viewpoint images with different parallax information;

[0043] S2. Determine the number of viewpoints and construct a mapping relationship between the sub-pixels in the synthesized image and the sub-pixels in the multi-viewpoint images;

[0044] S3. Use an image interpolation algorithm to scale each multi-viewpoint image so that the resolution of the scaled multi-viewpoint images is the same as the resolution of the synthesized image to be generated;

[0045] S4. According to the mapping relationship, sample the scaled multi-viewpoint images to obtain a synthesized image, and complete the multi-viewpoint 3D light field encoding.

[0046] A method for obtaining multi-viewpoint images with different parallax information includes:

[0047] In this embodiment, a single or multiple cameras are used to collect multiple images with different parallaxes of a target scene from different viewpoints, obtaining multi-viewpoint images with different parallax information; wherein, when a single or multiple cameras perform continuous collection from different viewpoints, multi-viewpoint images with different parallax information for consecutive frames are correspondingly obtained.

[0048] In this embodiment, according to geometric optics knowledge, based on the formula the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images can be determined, including:

[0049] For YUV4:4:4 sampling, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images is code(i,j,k) = view[n](i,j,k), that is, the width and height of the view size remain unchanged; where code(i,j,k) is the k-th sub-pixel in the pixel at the i-th row and j-th column in the composite image, and view[n](i,j,k) is the k-th sub-pixel in the pixel at the i-th row and j-th column in the image of the n-th viewpoint; n is the number of the viewpoint; N is the total number of viewpoints;

[0050] For YUV4:2:2 sampling, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images is that is, the horizontal resolution of the view size of the U and V components doubles, the width of the view is twice the original width, the height remains unchanged, and the view size of the Y component remains unchanged; represents rounding up;

[0051] For YUV4:4:0 sampling, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images is that is, the vertical resolution of the view size of the U and V components doubles, the height of the view is twice the original height, the width remains unchanged, and the view size of the Y component remains unchanged;

[0052] For YUV4:2:0 sampling, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images is that is, the view size of the U and V components doubles in both the horizontal and vertical directions, the width and height of the view both become twice the original width and original height, and the view size of the Y component remains unchanged;

[0053] For YUV4:1:1 sampling, the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images is that is, the view size of the U and V components is four times the original width in the horizontal direction, the height remains unchanged, and the view size of the Y component remains unchanged.

[0054] In this embodiment, the image interpolation algorithms include the nearest neighbor interpolation algorithm, the bilinear interpolation algorithm, and the bicubic interpolation algorithm. One of these algorithms can be selected for use.

[0055] In this embodiment, as Figure 2 and Figure 3 shown, the horizontal distance from the left edge of any sub-pixel (i, j, k) in the composite image to the left edge of the leftmost grating unit is D p , and the distance from the left edge of this sub-pixel to the left edge of the grating unit corresponding to this sub-pixel is A p . Then the following constraints exist:

[0056] D p = 3×(j - 1) + 3×(i - 1)×tanθ + (k - 1)

[0057] A p = D p mod P e

[0058] where θ is the tilt angle of the cylindrical lens deviating from the vertical direction; tan represents the tangent function; P e is the width of the sub-pixel area of the composite image covered by the grating unit in the horizontal direction; mod represents the remainder function; the cylindrical lens is a component in the display, and a single cylindrical lens includes several grating units.

[0059] Correspondingly, this embodiment provides a device based on an anti-compression multi-viewpoint 3D light field encoding method, which includes:

[0060] A multi-viewpoint image acquisition module for acquiring multi-viewpoint images with different parallax information;

[0061] A mapping relationship construction module for determining the number of viewpoints and constructing the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the multi-viewpoint images;

[0062] An image scaling module for scaling each multi-viewpoint image using an image interpolation algorithm so that the resolution of the scaled multi-viewpoint images is consistent with the resolution of the composite image to be generated;

[0063] A 3D light field encoding module for sampling the scaled multi-viewpoint images according to the mapping relationship to obtain a composite image and complete the multi-viewpoint 3D light field encoding.

[0064] The device further includes:

[0065] A display for loading the composite image for stereoscopic display.

[0066] where the display includes a display unit (LED display panel) and a cylindrical lens;

[0067] A display unit for loading a composite image;

[0068] A lenticular lens for controlling light rays to project each sub-pixel on the display unit into a volume pixel, thereby realizing stereoscopic display.

[0069] This embodiment also provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes an anti-compression multi-viewpoint 3D light field encoding method.

[0070] This embodiment also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by the processor, the processor executes an anti-compression multi-viewpoint 3D light field encoding method.

[0071] In an embodiment of the present invention, the resolution of the used display is 7680×4320, the screen size is 65 inches, the viewing angle range is 100 degrees, and a single grating unit covers 24.5 sub-pixels, i.e., P e = 24.5 and the tilt angle of the grating unit (i.e., the tilt angle of the lenticular lens deviating from the vertical direction) θ = arctan(-1 / 6).

[0072] Use an array of 100 cameras arranged horizontally to simultaneously capture the same scene from different angles to obtain a multi-viewpoint image sequence with different parallax information (resolution 3840×2160), and these images will be used for subsequent multi-viewpoint 3D light field encoding.

[0073] Determine the number of viewpoints: In this embodiment, the number of viewpoints is determined to be N = 100.

[0074] According to the viewpoint construction principle, determine the mapping relationship between the sub-pixels in the composite image and the sub-pixels in the parallax map:

[0075] According to the formula A p = D p mod P e , calculate the viewpoint n corresponding to the sub-pixel (i, j, k).

[0076] Exemplarily, for the sub-pixel (2, 3, 2), D p = 6.5, P e = 24.5, A p = 6.5, from the formula we get That is, the sub-pixel (2, 3, 2) corresponds to the 27th viewpoint.

[0077] Determine the anti-compression coding mapping relationship: Taking YUV4:2:0 sampling as an example, the coding mapping relationship is During conversion, each pixel in the RGB video frame is processed according to the coding mapping relationship. For the sub-pixel position (2, 3, 2), the corresponding pixel position in the YUV4:2:0 format according to the mapping relationship is (1, 2, 2), that is, code(2, 3, 2) = view

[27] (1, 2, 2) indicates that the second sub-pixel in the pixel at the second row and third column in the composite image has a corresponding relationship with the second sub-pixel in the pixel at the first row and second column in the disparity map of the 27th viewpoint. For an RGB video frame with a resolution of 3840×2160, during the conversion process, the view size of the Y component remains unchanged at 3840×2160; for the U and V components, their view sizes double in both the horizontal and vertical directions, that is, the width becomes 7680 and the height becomes 4320.

[0078] Use the bicubic interpolation algorithm to scale each disparity image so that it has the same resolution of 7680×4320 as the composite image. Finally, according to the mapping relationship, sample each disparity image to obtain the composite image.

[0079] Load the obtained video frame onto a display with a resolution of 7680×4320 and a screen size of 65 inches for stereoscopic display. Finally, the video plays smoothly and is displayed with high quality within a 100-degree viewing angle, without perceiving color distortion, realizing an anti-compression multi-viewpoint 3D light field coding method.

Claims

1. A compression-resistant multi-viewpoint 3D light field encoding method, characterized in that Including: Obtain multi-viewpoint images with different parallax information; Determine the number of viewpoints and construct the mapping relationship between sub-pixels in the synthesized image and sub-pixels in the multi-viewpoint images; Use an image interpolation algorithm to scale each multi-viewpoint image so that the resolution of the scaled multi-viewpoint image is the same as the resolution of the synthesized image to be generated; According to the mapping relationship, sample the scaled multi-viewpoint images to obtain a synthesized image, completing the multi-viewpoint 3D light field encoding.

2. The method according to claim 1, wherein The method for obtaining multi-viewpoint images with different parallax information includes: Use a single or multiple cameras to collect multiple images with different parallaxes of the target scene from different perspectives to obtain multi-viewpoint images with different parallax information; wherein, when a single or multiple cameras perform continuous collection from different perspectives, corresponding multi-viewpoint images with different parallax information for consecutive frames are obtained.

3. The method according to claim 1, wherein The mapping relationship between sub-pixels in the synthesized image and sub-pixels in the multi-viewpoint images includes: For YUV4:4:4 sampling, the mapping relationship between sub-pixels in the synthesized image and sub-pixels in the multi-viewpoint images is code(i,j,k) = view[n](i,j,k), that is, the width and height of the view size remain unchanged; where code(i,j,k) is the k-th sub-pixel in the pixel at the i-th row and j-th column in the synthesized image, and view[n](i,j,k) is the k-th sub-pixel in the pixel at the i-th row and j-th column in the image of the n-th viewpoint; n is the number of the viewpoint. For YUV4:2:2 sampling, the mapping relationship between the sub-pixels in the synthesized image and the sub-pixels in the multi-view image is That is, the horizontal resolution of the view size of the U and V components is doubled, the width of the view is twice the original width, the height remains unchanged, and the view size of the Y component remains unchanged; Indicates rounding up; For YUV4:4:0 sampling, the mapping relationship between the sub-pixels in the synthesized image and the sub-pixels in the multi-view image is That is, the view sizes of the U and V components are doubled in the vertical resolution, the height of the view is twice the original height, the width remains unchanged, and the view size of the Y component remains unchanged; For YUV4:2:0 sampling, the mapping relationship between the sub-pixels in the synthesized image and the sub-pixels in the multi-view image is That is, the view sizes of the U and V components increase by a factor of two in both the horizontal and vertical directions, and the width and height of the view become twice the original width and height, while the view size of the Y component remains unchanged; For YUV4:1:1 sampling, the mapping relationship between the sub-pixels in the synthesized image and the sub-pixels in the multi-view image is That is, the view sizes of the U and V components are four times the original width in the horizontal direction and remain unchanged in the vertical direction, and the view size of the Y component remains unchanged.

4. The method according to claim 1, wherein The image interpolation algorithm includes the nearest neighbor interpolation algorithm, the bilinear interpolation algorithm, and the bicubic interpolation algorithm.

5. The method according to claim 3, wherein The horizontal distance from the left edge of any sub-pixel (i, j, k) in the composite image to the left edge of the leftmost grating unit is D p , and the distance from the left edge of this sub-pixel to the left edge of the grating unit corresponding to this sub-pixel is A p , then there are the following constraints: D p = 3×(j - 1)+3×(i - 1)×tanθ+(k - 1) A p = D p mod P e where θ is the tilt angle of the cylindrical lens deviating from the vertical direction; tan represents the tangent function; P e is the width of the synthesized sub-pixel area covered by the grating unit in the horizontal direction; mod represents the remainder function; the cylindrical lens is a component in the display, and a single cylindrical lens includes a number of grating units.

6. An apparatus based on the method according to any one of claims 1 to 5, characterized in that Including: A multi-viewpoint image acquisition module for obtaining multi-viewpoint images with different parallax information; A mapping relationship construction module for determining the number of viewpoints and constructing the mapping relationship between sub-pixels in the synthesized image and sub-pixels in the multi-viewpoint images; An image scaling module for using an image interpolation algorithm to scale each multi-viewpoint image so that the resolution of the scaled multi-viewpoint image is the same as the resolution of the synthesized image to be generated; A 3D light field encoding module for sampling the scaled multi-viewpoint images according to the mapping relationship to obtain a synthesized image, completing the multi-viewpoint 3D light field encoding.

7. The device according to claim 6, characterized in that, Also including: A display for loading the synthesized image for stereoscopic display.

8. The device according to claim 7, characterized in that, The display includes a display unit and lenticular lenses; The display unit for loading the synthesized image; The lenticular lenses for performing light control to project each sub-pixel on the display unit into a volume pixel, thereby realizing stereoscopic display.

9. A computer device, characterized in that, Including a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method according to any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that, Stores a computer program, and when the computer program is executed by the processor, the processor executes the method according to any one of claims 1 to 5.