Image processing method and apparatus, device and computer-readable storage medium
By acquiring and transmitting pixel information and depth parameters of the original image, a target image with a three-dimensional effect is generated, which solves the problems of poor image processing flexibility and low transmission efficiency in the existing technology and achieves more efficient image processing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BOE TECHNOLOGY GROUP CO LTD
- Filing Date
- 2025-01-02
- Publication Date
- 2026-07-09
AI Technical Summary
In existing technologies, image processing is less flexible, and the amount of content sent from the sending end to the receiving end is large, resulting in high transmission bandwidth and low transmission efficiency, which in turn reduces the efficiency of image processing.
The system acquires pixel information and depth parameters of multiple pixels in the original image, generates image information based on this information, and sends the image information to the receiving end, enabling the receiving end to generate a target image with a three-dimensional effect. The number of reference images is determined based on the number of viewpoints supported by the receiving end's display screen.
This reduces the amount of content sent by the sender, lowers the bandwidth requirements, improves transmission efficiency, and enhances the flexibility and efficiency of image processing.
Smart Images

Figure CN2025070277_09072026_PF_FP_ABST
Abstract
Description
Image processing methods, apparatus, devices and computer-readable storage media Technical Field
[0001] This application relates to the field of computer technology, and in particular to an image processing method, apparatus, device, and computer-readable storage medium. Background Technology
[0002] With the continuous development of computer technology, image processing methods are becoming increasingly diversified. For example, 3D (three-dimensional) processing can be applied to images to give them a 3D effect.
[0003] In related technologies, the number of viewpoints supported by the display screen corresponding to the receiving end is determined, a number of original images of the viewpoints are acquired, and the pixel information of each pixel in each original image is determined. The pixel information of each pixel in each original image is sent to the receiving end so that the receiving end can generate a target image with a 3D effect based on the pixel information of each pixel in each original image.
[0004] However, in the aforementioned image processing methods, the content sent by the sending end needs to be determined based on the number of viewpoints supported by the receiving end's display screen, resulting in poor flexibility in image processing. Furthermore, the sending end transmits pixel information for each pixel in the original image to the receiving end, leading to high bandwidth requirements and low transmission efficiency, which in turn reduces the efficiency of image processing.
[0005] Public content
[0006] This application provides an image processing method, apparatus, device, and computer-readable storage medium, which can be used to solve the problems in related technologies such as poor flexibility in image processing, large amounts of content sent from the sending end to the receiving end, resulting in high bandwidth requirements and low transmission efficiency, thereby reducing the efficiency of image processing. The technical solution is as follows:
[0007] In a first aspect, embodiments of this application provide an image processing method, the method comprising:
[0008] Obtain pixel information of multiple pixels in the original image, whereby the pixel information of any one pixel is used to describe the brightness, chroma, and saturation of that pixel.
[0009] Determine the depth parameters of each pixel, whereby the depth parameters of any pixel are used to indicate the distance between that pixel and the image acquisition device that acquired the original image.
[0010] Image information is obtained based on the pixel information and depth parameters of the multiple pixels;
[0011] The image information is sent to the receiving end, and the image information is used by the receiving end to generate the original image and the reference image corresponding to the original image. A target image with a three-dimensional effect is generated based on the original image and the reference image. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end.
[0012] Secondly, embodiments of this application provide an image processing method, the method comprising:
[0013] The image information sent by the receiving end is obtained based on the pixel information and depth parameters of multiple pixels in the original image. The pixel information of any pixel is used to describe the brightness, chroma and saturation of the pixel, and the depth parameter of any pixel is used to indicate the distance between the pixel and the image acquisition device that acquired the original image.
[0014] Based on the image information, determine the pixel information and depth parameters of each pixel;
[0015] Based on the pixel information and depth parameters of each pixel, the original image and the reference image corresponding to the original image are generated. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end.
[0016] A target image with a three-dimensional effect is generated based on the original image and the reference image.
[0017] Thirdly, embodiments of this application provide an image processing apparatus, the apparatus comprising:
[0018] The acquisition module is used to acquire pixel information of multiple pixels in the original image, and the pixel information of any pixel is used to describe the brightness, chroma and saturation of the pixel.
[0019] The determination module is used to determine the depth parameters of each pixel, and the depth parameters of any pixel are used to indicate the distance between the pixel and the image acquisition device that acquired the original image.
[0020] The acquisition module is also used to acquire image information based on the pixel information and depth parameters of the plurality of pixels;
[0021] The sending module is used to send the image information to the receiving end. The image information is used by the receiving end to generate the original image and the reference image corresponding to the original image, and to generate a target image with a three-dimensional effect based on the original image and the reference image. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end.
[0022] In one possible implementation, when the depth information of each pixel is located within a reference interval, the depth parameter of each pixel is the depth information of each pixel, the reference interval is determined based on the number of bits of the depth channel used in the reference sampling method, and the depth information of any pixel is used to indicate the actual distance between any pixel and the image acquisition device.
[0023] In one possible implementation, if there are pixels whose depth information is outside the reference interval, the depth parameter of each pixel is the depth order of each pixel, and the depth order of any pixel is located in the reference interval. The reference interval is determined based on the number of bits of the depth channel used in the reference sampling method, and the depth information of any pixel is used to indicate the actual distance between the pixel and the image acquisition device.
[0024] In one possible implementation, the determining module is further configured to determine the depth level of each pixel based on the number of bits of the depth channel used in the reference sampling method, the maximum depth information in the depth information of each pixel, and the depth information of each pixel.
[0025] In one possible implementation, the image information includes sampling information of the original image; the acquisition module is used to sample the pixel information and depth parameters of each pixel according to a reference sampling method to obtain the sampling information of the original image, and the sampling information of the original image is used by the receiving end to determine the pixel information and depth parameters of each pixel included in the original image.
[0026] In one possible implementation, the sending module is further configured to send supplementary enhancement information to the receiving end, the supplementary enhancement information indicating the reference sampling method, the reference sampling method being used by the receiving end to determine the pixel information and depth parameters of each pixel.
[0027] In one possible implementation, the sending module is configured to encode the image information to obtain an encoded packet corresponding to the original image, wherein the encoded packet corresponding to the original image includes the image information; encapsulate the encoded packet corresponding to the original image to obtain an encapsulated packet corresponding to the original image; and send the encapsulated packet corresponding to the original image to the receiving end.
[0028] In one possible implementation, the depth parameter of each pixel is the depth order of each pixel. The sending module is used to encode the image information and the depth step size of the original image to obtain an encoded packet corresponding to the original image. The encoded packet corresponding to the original image also includes the depth step size of the original image. The depth step size of the original image is used to determine the depth information of each pixel included in the original image.
[0029] In one possible implementation, the determining module is further configured to determine the depth step size of the original image based on the maximum depth information in the depth information of each pixel in the original image and the number of bits of the depth channel used in the reference sampling method.
[0030] In one possible implementation, the determining module is further configured to determine the depth standard deviation of the original image based on the depth information of each pixel included in the original image; and to determine the depth step size of the original image based on the depth standard deviation of the original image and the number of bits of the depth channel used in the reference sampling method.
[0031] Fourthly, embodiments of this application provide an image processing apparatus, the apparatus comprising:
[0032] The receiving module is used to receive image information sent by the sending end. The image information is obtained based on the pixel information and depth parameters of multiple pixels included in the original image. The pixel information of any pixel is used to describe the brightness, chroma and saturation of the pixel. The depth parameter of any pixel is used to indicate the distance between the pixel and the image acquisition device that acquired the original image.
[0033] The determining module is used to determine the pixel information and depth parameters of each pixel point based on the image information;
[0034] The generation module is used to generate the original image and the reference image corresponding to the original image based on the pixel information and depth parameters of each pixel. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end.
[0035] The generation module is further configured to generate a target image with a three-dimensional effect based on the original image and the reference image.
[0036] In one possible implementation, the receiving module is configured to receive the encapsulation packet corresponding to the original image sent by the sending end; decapsulate the encapsulation packet corresponding to the original image to obtain the encoding packet corresponding to the original image; and decode the encoding packet corresponding to the original image to obtain the image information.
[0037] In one possible implementation, the depth parameter of each pixel is the depth order of each pixel, and the encoding packet also includes the depth step size of the original image;
[0038] The receiving module is further configured to decode the encoded packet corresponding to the original image to obtain the depth step of the original image;
[0039] The generation module is used to determine the depth information of each pixel based on the depth step size of the original image and the depth order of each pixel; and to generate the original image and a reference image corresponding to the original image based on the pixel information and depth information of each pixel.
[0040] In one possible implementation, the generation module is configured to generate the original image and a first intermediate image corresponding to the original image based on the pixel information and depth information of each pixel, wherein the size of the first intermediate image is larger than the size of the original image; crop the first intermediate image to obtain a second intermediate image, wherein the size of the second intermediate image is the same as the size of the original image; fill the blank areas in the second intermediate image to obtain a third intermediate image; and determine a reference image corresponding to the original image based on the third intermediate image.
[0041] In one possible implementation, the generation module is used to use the third intermediate image as a reference image corresponding to the original image; or, to repair the hole regions in the third intermediate image and use the repaired image as a reference image corresponding to the original image.
[0042] In one possible implementation, the generation module is configured to determine the reflection vector corresponding to the first intermediate image, the reflection vector corresponding to the first intermediate image being determined based on the position of the first intermediate image relative to the original image; determine the offset position information of each pixel based on the position information, depth information, and the reflection vector corresponding to the first intermediate image; and generate the first intermediate image corresponding to the original image based on the pixel information, depth information, and the offset position information of each pixel.
[0043] In one possible implementation, the reflection vector corresponding to the first intermediate image is a two-dimensional vector, and the position information of any pixel includes the horizontal and vertical coordinates of the pixel.
[0044] The generation module is configured to, for any pixel among the pixels, determine the offset horizontal coordinate of the pixel after the offset based on the horizontal coordinate of the pixel, the depth information of the pixel, and the value of the first dimension in the reflection vector corresponding to the first intermediate image; determine the offset vertical coordinate of the pixel after the offset based on the vertical coordinate of the pixel, the depth information of the pixel, and the value of the second dimension in the reflection vector corresponding to the first intermediate image; and determine the offset position information of the pixel after the offset based on the offset horizontal coordinate and the offset vertical coordinate.
[0045] Fifthly, embodiments of this application provide a computer device, the computer device including a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to enable the computer device to implement any of the image processing methods described above.
[0046] In a sixth aspect, a computer-readable storage medium is also provided, wherein at least one piece of program code is stored therein, the at least one piece of program code being loaded and executed by a processor to enable a computer to implement any of the image processing methods described above.
[0047] In a seventh aspect, a computer program or computer program product is also provided, wherein the computer program or computer program product stores at least one computer instruction, the at least one computer instruction being loaded and executed by a processor to enable the computer to implement any of the above-described image processing methods.
[0048] The technical solution provided in this application has at least the following beneficial effects:
[0049] The technical solution provided in this application only requires the sending end to obtain the pixel information and depth parameters of each pixel in the original image, and then send the image information determined based on the pixel information and depth parameters of each pixel to the receiving end. This enables the receiving end to generate the original image and a corresponding reference image based on the image information, and then generate a target image with a 3D effect based on the original image and the reference image. Since the sending end only needs to send the information of the original image itself, the amount of content sent by the sending end is small, reducing the bandwidth required to transmit content to the receiving end, thereby improving the content transmission efficiency and thus improving the efficiency of image processing.
[0050] Moreover, the content sent by the sending end does not need to be determined based on the number of viewpoints supported by the receiving end's display screen. In this way, regardless of the number of viewpoints supported by the receiving end's display screen, only the information of the original image itself is sent to the receiving end. This decouples the sending end and the receiving end, thereby improving the flexibility of image processing. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 is a schematic diagram of the implementation environment of an image processing method provided in an embodiment of this application;
[0053] Figure 2 is a flowchart of an image processing method provided in an embodiment of this application;
[0054] Figure 3 is a schematic diagram illustrating the determination of depth information of any pixel in an original image according to an embodiment of this application;
[0055] Figure 4 is an architecture diagram of a deep learning model provided in an embodiment of this application;
[0056] Figure 5 is a schematic diagram of YUVD image acquisition provided in an embodiment of this application;
[0057] Figure 6 is a schematic diagram of a YUV image and a YUVD image corresponding to the YUV image provided in an embodiment of this application;
[0058] Figure 7 is a flowchart of an encoder processing according to an embodiment of this application;
[0059] Figure 8 is a flowchart of an image processing method provided in an embodiment of this application;
[0060] Figure 9 is a schematic diagram of an original image and a first intermediate image corresponding to the original image provided in an embodiment of this application;
[0061] Figure 10 is a schematic diagram of determining a second intermediate image according to an embodiment of this application;
[0062] Figure 11 is a schematic diagram of determining a third intermediate image according to an embodiment of this application;
[0063] Figure 12 is a schematic diagram of an in-screen / out-of-screen effect provided in an embodiment of this application;
[0064] Figure 13 shows an RTC receiver report message format provided in an embodiment of this application.
[0065] Figure 14 is a schematic diagram showing a reference image corresponding to an original image provided in an embodiment of this application;
[0066] Figure 15 is a schematic diagram of a cavity repair provided in an embodiment of this application;
[0067] Figure 16 is a schematic diagram of a reference image corresponding to an original image provided in an embodiment of this application;
[0068] Figure 17 is a decoding flowchart of a decoder provided in an embodiment of this application;
[0069] Figure 18 is a schematic diagram of an image processing method provided in an embodiment of this application;
[0070] Figure 19 is a schematic diagram of the structure of an image processing device provided in an embodiment of this application;
[0071] Figure 20 is a schematic diagram of the structure of an image processing device provided in an embodiment of this application;
[0072] Figure 21 is a schematic diagram of the structure of a server provided in an embodiment of this application;
[0073] Figure 22 is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0074] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0075] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0076] Figure 1 is a schematic diagram of the implementation environment of an image processing method provided in an embodiment of this application. As shown in Figure 1, the implementation environment includes a transmitting end 101 and a receiving end 102. The image processing method provided in this embodiment is implemented through the interaction between the transmitting end 101 and the receiving end 102.
[0077] Optionally, the transmitter 101 can be a terminal device, which can be any electronic device that allows human-computer interaction with the user through one or more methods such as a keyboard, touchpad, remote control, voice interaction, or handwriting device. Examples include PCs (Personal Computers), mobile phones, smartphones, PDAs (Personal Digital Assistants), wearable devices, PPCs (Pocket PCs), tablets, smart car systems, smart TVs, smart speakers, and smartwatches. The receiver 102 follows the same principle.
[0078] A terminal device can refer to one of multiple terminal devices; this embodiment uses only one terminal device as an example. Those skilled in the art will understand that the number of terminal devices can be more or less. For example, there may be only one terminal device, or there may be dozens or hundreds, or even more. This application embodiment does not limit the number or type of terminal devices.
[0079] When the sending end 101 is a server, the server can be a single server, a server cluster consisting of multiple servers, or any one of a cloud computing platform and a virtualization center; this embodiment of the application does not limit this. The receiving end 102 is similar.
[0080] Those skilled in the art should understand that the above-described terminal devices and servers are merely illustrative examples. Other existing or future terminal devices or servers that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.
[0081] This application provides an image processing method that can be applied to the implementation environment shown in FIG1. Taking the flowchart of an image processing method provided by this application embodiment shown in FIG2 as an example, the method can be executed by the sending end 101 in FIG1. As shown in FIG2, the method includes the following steps 201 to 204.
[0082] In step 201, pixel information of multiple pixels in the original image is obtained, and the pixel information of any pixel is used to describe the brightness, chroma and saturation of any pixel.
[0083] In an exemplary embodiment of this application, the original image includes multiple pixels that correspond to color information, and the pixel information of each pixel is determined based on the color information of each pixel.
[0084] In one possible implementation, the color information of any pixel is used to describe the color of that pixel. Optionally, the color information of any pixel is the RGB (red-green-blue) information of that pixel. That is, the color information of any pixel includes the red component (R), green component (G), and blue component (B) of that pixel. The color of any pixel can be obtained from the red component, green component, and blue component of that pixel.
[0085] Optionally, the color information of any pixel is the RGBA (red-green-blue-alpha) information of that pixel. That is, the color information of any pixel includes the red component (R), green component (G), blue component (B), and alpha component (A). The color of any pixel can be obtained by using the red component, green component, blue component, and alpha component of any pixel.
[0086] In one possible implementation, the pixel information of any pixel is its YUV (Luminance-Chrominance-Chroma) information. That is, the pixel information of any pixel includes its luminance component (Y), chrominance component (U), and chrominance component (V). The luminance component describes luminance, the chrominance component describes chrominance, and the chrominance component describes chromaticity.
[0087] Optionally, the process of determining the pixel information of each pixel based on the color information of each pixel includes: for any pixel among the pixels, determining the pixel information of any pixel based on the color information of that pixel.
[0088] In one possible implementation, the color information of any pixel includes the red, green, and blue components of that pixel, and the pixel information of any pixel includes the luminance, chroma, and saturation components of that pixel. The process of determining the pixel information of any pixel based on its color information includes: determining the luminance, chroma, and saturation components of any pixel based on its red, green, and blue components.
[0089] In some embodiments, the luminance component, chroma component, and density component of any pixel are determined according to the following formula (1) based on the red component, green component, and blue component of any pixel.
[0090] In the above formula (1), R is the red component of any pixel, G is the green component of any pixel, B is the blue component of any pixel, Y is the luminance component of any pixel, U is the chroma component of any pixel, and V is the density component of any pixel.
[0091] In another possible implementation, the color information of any pixel includes the red, green, blue, and alpha components, and the pixel information of any pixel includes the luminance, chroma, and saturation components. The process of determining the pixel information of any pixel based on its color information includes: determining the median red component of any pixel based on its red and alpha components; determining the median green component of any pixel based on its green and alpha components; determining the median blue component of any pixel based on its blue and alpha components; and determining the luminance, chroma, and saturation components of any pixel based on its median red, median green, and median blue components.
[0092] The red component of any pixel is determined according to the red component and transparency component of that pixel, using the following formula (2): RED=R*A+C R *(1-A) Formula (2)
[0093] In the above formula (2), RED is the central red component of any pixel, R is the red component of any pixel, A is the transparency component of any pixel, and C is the lightness component of any pixel. R This refers to the red component of the background color.
[0094] Based on the green component and transparency component of any pixel, the intermediate green component of any pixel is determined according to the following formula (3): GREEN=G*A+C G *(1-A) Formula (3)
[0095] In the above formula (3), GREEN is the central green component of any pixel, G is the green component of any pixel, A is the transparency component of any pixel, and C is the green component of any pixel. G This is the green component of the background color.
[0096] Based on the blue component and transparency component of any pixel, the intermediate blue component of any pixel is determined according to the following formula (4): BLUE=B*A+C B *(1-A) Formula (4)
[0097] In the above formula (4), BLUE is the central blue component of any pixel, B is the blue component of any pixel, A is the transparency component of any pixel, and C is the lightness component of any pixel. B It represents the blue component of the background color.
[0098] It should be noted that the background color can be any color, and this application embodiment does not limit it. For example, if the background color is the color corresponding to (32, 128, 247), then the red component of the background color is 32, the green component is 128, and the blue component is 247. As another example, if the background color is white, and white is the color corresponding to (255, 255, 255), then the red, green, and blue components of the background color are all 255.
[0099] For example, the background color is the color corresponding to (255, 255, 255), and the color information of any pixel is (32, 128, 247, 0.4). According to the above formula (2), the central red component of any pixel is determined as RED = 32 * 0.4 + 255 * (1 - 0.4) = 166. According to the above formula (3), the central green component of any pixel is determined as GREEN = 128 * 0.4 + 255 * (1 - 0.4) = 204. According to the above formula (4), the central blue component of any pixel is determined as BLUE = 247 * 0.4 + 255 * (1 - 0.4) = 252.
[0100] Based on the central red component, central green component, and central blue component of any pixel, determine the luminance component, chroma component, and saturation component of any pixel according to the following formula (5).
[0101] In the above formula (5), RED is the central red component of any pixel, GREEN is the central green component of any pixel, BLUE is the central blue component of any pixel, Y is the luminance component of any pixel, U is the chroma component of any pixel, and V is the density component of any pixel.
[0102] For example, the central red component of any pixel is 166, the central green component is 204, and the central blue component is 252. The luminance component of any pixel is determined to be 198.11, the chroma component is 154.52098, and the saturation component is 99.83156 by the above formula (5).
[0103] In step 202, the depth parameters of each pixel are determined. The depth parameters of any pixel are used to indicate the distance between any pixel and the image acquisition device that acquired the original image.
[0104] It should be noted that the execution order of steps 201 and 202 is not limited in this embodiment. Step 201 can be executed first to determine the pixel information of each pixel, and then step 202 can be executed to determine the depth parameters of each pixel; or step 202 can be executed first to determine the depth parameters of each pixel, and then step 201 can be executed to determine the pixel information of each pixel; or steps 201 and 202 can be executed simultaneously, that is, the pixel information and depth parameters of each pixel can be determined simultaneously.
[0105] In one possible implementation, the process of determining the depth parameters of each pixel includes: determining the depth information of each pixel, whereby the depth information of any pixel indicates the actual distance between that pixel and the image acquisition device that acquired the original image; and determining the depth parameters of each pixel based on the depth information of each pixel.
[0106] Optionally, when the depth information of all pixels is within the reference interval, the depth parameter of each pixel is the depth information of that pixel. If there are pixels whose depth information is outside the reference interval, the depth parameter of each pixel is the depth order of that pixel. The depth order of any pixel is within the reference interval, which is determined based on the number of bits in the depth channel used in the reference sampling method. The depth order of any pixel is used to indicate the relative distance between that pixel and the image acquisition device that acquired the original image.
[0107] For example, if the reference sampling method is yuvd4204p and the depth channel of the reference sampling method has 8 bits, then the reference range is 0 to 2. 8 -1 means the reference range is 0 to 255. For example, if the reference sampling method is yuvd4204p9be and the depth channel has 9 bits, then the reference range is 0 to 255. 9 -1, meaning the reference range is 0 to 511. Of course, the reference sampling method can be other methods, and this application embodiment does not limit this.
[0108] Table 1 below is an exemplary table of reference sampling methods provided in the embodiments of this application.
[0109] Table 1
[0110] As shown in Table 1 above, when the reference sampling method is YUVD4204P, the number of channels is 4, the number of bits per pixel is 20, and the number of bits per channel is 8. That is, the Y channel (luminance channel), U channel (chroma channel), V channel (saturation channel), and D channel (depth channel) have 8 bits each. For other reference sampling methods, the corresponding number of channels, number of bits per pixel, and number of bits per channel are shown in Table 1 above, and will not be repeated here.
[0111] In one possible implementation, the process of determining the depth information of each pixel in the original image differs depending on the image acquisition device used to acquire the original image. This application's embodiments only illustrate the process of determining the depth information of each pixel in the original image using a binocular camera and a monocular camera as examples.
[0112] In one possible implementation, when the image acquisition device for acquiring the original image is a binocular camera, the depth information of each pixel in the original image is determined according to the following formula (6).
[0113] In the above formula (6), Z represents the depth information of any pixel in the original image, f is the focal length of the binocular camera, T is the distance between the two cameras in the binocular camera, and X... R X is the x-coordinate of any pixel in the original image within the imaging plane of the right camera of the binocular camera system. L The x-coordinate of any pixel in the original image within the imaging plane of the left camera of the binocular camera.
[0114] Figure 3 is a schematic diagram illustrating the determination of depth information of any pixel in an original image according to an embodiment of this application. Wherein, p is any pixel in the original image, ol is the left camera of the binocular camera, or is the right camera of the binocular camera, pl is the point of any pixel p in the imaging plane of the left camera of the binocular camera, pr is the point of any pixel p in the imaging plane of the right camera of the binocular camera, xl is the abscissa of any pixel p in the imaging plane of the left camera of the binocular camera, xr is the abscissa of any pixel p in the imaging plane of the right camera of the binocular camera, f is the focal length of the binocular camera, and Z is the depth information of any pixel p.
[0115] In another possible implementation, when the image acquisition device for capturing the original image is a monocular camera, a deep learning model is obtained through deep learning training or by convolution. Based on the pixel information of each pixel in the original image, a YUV image (an image format) corresponding to the original image is generated. The YUV image corresponding to the original image is then input into the deep learning model to obtain a YUVD image (an image format). The YUVD image includes the depth information of each pixel in the original image. This approach requires a certain amount of hardware computing power. Commonly used open-source models include Intel's OpenVino (a semi-open-source toolkit developed by Intel specifically for optimizing and deploying artificial intelligence inference), DPT (Dense Prediction Transformers, a deep learning model based on the transformer architecture), MegaDepth (a project for learning single-view depth prediction from images), etc. Of course, you can also train your own deep learning model to estimate depth information. Figure 4 is an architecture diagram of a deep learning model provided in an embodiment of this application. The architecture of a deep learning model can be divided into two main parts: a front-end convolutional neural network (F-MF) and a multi-scale fusion module with continuous conditional random fields (CRFs). The front-end F-MF consists of multiple convolutional layers used to extract features from the input image (image r), resulting in feature maps (s1, s2, s3, s4, s5). These feature maps contain information about the input image at different levels of abstraction. For example, s1 contains basic features such as edges and textures, while s5 contains semantic features. The multi-scale fusion module (C-MF) is used to fuse feature maps from different scales to obtain a YUVD image (image d).
[0116] Figure 5 is a schematic diagram illustrating the acquisition of a YUVD image according to an embodiment of this application. In Figure 5, the YUV image of the original image is input into a deep learning model to obtain the YUVD image corresponding to the original image.
[0117] Optionally, the model can be a cuDNN model (a neural network model provided by NVIDIA). The cuDNN model can work well with GPU hardware, maximizing its parallel computing power. PyTorch (an open-source deep learning framework) is used to load models such as MiDas (Monocular Depth Estimation with a Multi-Scale Deep Network, an advanced depth estimation model), and YUV images can be obtained by inputting YUV images.
[0118] Figure 6 is a schematic diagram of a YUV image and its corresponding YUVD image provided in an embodiment of this application. In the figure, 602 is the YUVD image corresponding to 601, 604 is the YUVD image corresponding to 603, 606 is the YUVD image corresponding to 605, and 608 is the YUVD image corresponding to 607.
[0119] In one possible implementation, depth information can also be calculated using focal length. By adjusting the focal length and phase distance of the image acquisition device to obtain the clearest image, the object distance (i.e., depth information) can be determined. The corresponding formula is 1 / u + 1 / v = 1 / f, where u is the object distance, v is the phase distance, and f is the focal length. Alternatively, depth information can be calculated using photosensitive distance measurement or based on structured light pulses. The embodiments in this application will not be elaborated further here.
[0120] After determining the depth information of each pixel, the depth parameters of each pixel are determined based on the depth information and the reference range. If the depth information of all pixels falls within the reference range, the depth parameters of each pixel are determined as its depth information. If there are pixels whose depth information falls outside the reference range, the depth order of each pixel is determined based on its depth information and the bit depth of the depth channel used in the reference sampling method. If the depth order of any pixel falls within the reference range, the depth parameters of that pixel are determined as its depth order.
[0121] The process of determining the depth order of each pixel based on the depth information of each pixel and the number of bits of the depth channel used in the reference sampling method includes: determining the depth order of each pixel based on the number of bits of the depth channel used in the reference sampling method, the maximum depth information in the depth information of each pixel, and the depth information of each pixel.
[0122] Optionally, the process of determining the depth level of each pixel based on the number of bits of the depth channel used in the reference sampling method, the maximum depth information in the depth information of each pixel, and the depth information of each pixel includes: determining a reference value based on the number of bits of the depth channel used in the reference sampling method; and determining the depth level of each pixel based on the maximum depth information in the depth information of each pixel, the reference value, and the depth information of each pixel.
[0123] Where the depth channel bit depth used in the reference sampling method is N, then the reference value is 2. N N is an integer greater than or equal to 1. For example, if the depth channel bit depth used in the reference sampling method is 8 bits, then the reference value is 256.
[0124] Based on the maximum depth information, reference value, and depth information of each pixel, the depth order of each pixel is determined according to the following formula (7).
[0125] In the above formula (7), D(x, y) is the depth order of pixel (x, y), Z(x, y) is the depth information of pixel (x, y), and N is the number of bits of the depth channel used in the reference sampling method. N For reference value, Z max This represents the maximum depth information among all pixels.
[0126] In step 203, image information is obtained based on pixel information and depth parameters of multiple pixels.
[0127] In one possible implementation, the image information includes sampling information of the original image. After determining the depth parameters of each pixel in step 202 above, sampling is performed on the pixel information and depth parameters of each pixel according to a reference sampling method to obtain the sampling information of the original image. The sampling information of the original image is used by the receiving end to determine the pixel information and depth parameters of each pixel included in the original image.
[0128] The reference sampling method is any one of the sampling methods in Table 1 above.
[0129] In one possible implementation, the pixel information and depth parameters of each pixel are sampled according to the reference sampling method to obtain the sampled information; the sampled information is then stored according to the reference storage mode to obtain the sampled information of the original image.
[0130] The reference storage modes include planar mode and packed mode. Planar mode stores different channels in separate memory addresses. That is, it stores the Y, U, V, and D channel information from the sampled information in different locations. Packed mode arranges each channel of each pixel sequentially in an interleaved manner. That is, it mixes the Y, U, V, and D channel information from the sampled information together.
[0131] For example, the original image includes eight pixels, each pixel corresponding to a pixel information and a depth parameter. The pixel information includes a luminance component, a chrominance component, and a saturation component. Optionally, through steps 201 and 202 above, it is determined that the luminance component of the first pixel is Y1, the chrominance component is U1, the saturation component is V1, and the depth parameter is D1; the luminance component of the second pixel is Y2, the chrominance component is U2, the saturation component is V2, and the depth parameter is D2; the luminance component of the third pixel is Y3, the chrominance component is U3, the saturation component is V3, and the depth parameter is D3; and the luminance component of the fourth pixel is Y4, the chrominance component is U4, and the saturation component is D1. The luminance component of the fifth pixel is Y5, the chroma component is U5, the saturation component is V5, and the depth parameter is D5; the luminance component of the sixth pixel is Y6, the chroma component is U6, the saturation component is V6, and the depth parameter is D6; the luminance component of the seventh pixel is Y7, the chroma component is U7, the saturation component is V7, and the depth parameter is D7; the luminance component of the eighth pixel is Y8, the chroma component is U8, the saturation component is V8, and the depth parameter is D8. Taking the reference sampling method yuvd4202p as an example, sampling is performed on the pixel information and depth parameters of each pixel to obtain the following sampled information: Y1, Y2, Y3, Y4, Y5, Y6, Y7, Y8, U1, U3, V5, V7, D1, D2, D3, D4, D5, D6, D7, D8. That is, each pixel corresponds to one luminance component and one depth component. The first four pixels share U1 and V5, and the last four pixels share U3 and V7. Taking the planar reference storage mode as an example, storing the sampled information yields the original image's sampled information as Y1Y2Y3Y4Y5Y6Y7Y8U1U3V5V7D1D2D3D4D5D6D7D8. Taking the packed reference storage mode as an example, storing the sampled information yields the original image's sampled information as Y1Y2Y3Y4Y5Y6Y7Y8U1V5U3V7D1D2D3D4D5D6D7D8.
[0132] Optionally, when the sampling ratio of the luminance component, chroma component, saturation component, and depth component is 1:1:1:1, the amount of data in the sampled information of the original image consists of the pixel information of multiple pixels and the data corresponding to the depth parameters. When the sampling ratio of the luminance component, chroma component, saturation component, and depth component is not 1:1:1:1, the amount of data in the sampled information of the original image is less than the amount of data corresponding to the pixel information of multiple pixels and the depth parameters. This reduces the amount of data sent to the receiving end, thereby improving transmission efficiency.
[0133] Taking the YUVD4204P sampling method as an example, the sampling ratio of the luminance component, chroma component, saturation component, and depth component is 1:1 / 4:1 / 4:1. According to YUVD4204P's sampling method, each pixel samples one luminance component and one depth parameter, and every four pixels samples one chroma component and one saturation component. The amount of content occupied by every four pixels in the sampled information is 4 + 1 + 1 + 4 = 10 bytes. Therefore, the amount of content occupied by one pixel is 10 / 4 = 2.5 bytes. Taking the reference sampling format yuvd4201p as an example, the sampling ratio of the luminance component, chroma component, saturation component, and depth component is 1:1 / 4:1 / 4:1 / 4. According to yuvd4201p, when sampling the pixel information and depth parameters of each pixel, one luminance component is sampled for each pixel, and one chroma component, one saturation component, and one depth parameter are sampled for every four pixels. The amount of content occupied by every four pixels in the sampled information is 4+1+1+1=7. Therefore, the amount of content occupied by one pixel is 7 / 4=1.75 bytes.
[0134] In step 204, image information is sent to the receiving end. The image information is used by the receiving end to generate an original image and a reference image corresponding to the original image. A target image with a three-dimensional stereoscopic effect is generated based on the original image and the reference image. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end.
[0135] In one possible implementation, before sending image information to the receiving end, supplemental enhancement information (SEI) needs to be sent first. This SEI indicates the reference sampling method, which is used by the receiving end to determine the pixel information and depth parameters of each pixel. In other words, after receiving the image information, the receiving end determines the pixel information and depth parameters of each pixel based on the reference sampling method and the image information.
[0136] Optionally, a Session Description Protocol (SDP) is sent to the receiving end, which includes supplementary enhancement information.
[0137] The following is an example of a session description protocol sent to a receiving end according to an embodiment of this application.
[0138] v = 0 (version number)
[0139] o = -0 0 IN IP4 127.0.0.1 (Owner / Creator and Session Identifier)
[0140] s = No Name(session name)
[0141] t = 0 0 (Session activity time, representing start and end; 0 indicates no limit)
[0142] a = tool:libavformat 60.3.100 (attribute, tool here means to generate tool)
[0143] m = video 0 RTP / AVP 96 (m represents media, video indicates video content, RTP packet format, payload_type = 96)
[0144] b = AS:2288 (b represents the bit rate, and 2288 indicates a maximum of 2288kbps)
[0145] a = rtpmap:96 H266 / 90000 (rtp mapping, 96 corresponds to the payload_type above, encoded as h266 / vvc, frequency 90000)
[0146] a = fmtp:96 sprop-vps = QAEMAf / / AUAAAAMAKAAA..; sprop-sps = QgEBAUA AAAMAKAAAAwaaawCWoACACAGwfEZZVSkIRkX9DAQEAAADAAQAAA MAY…; sprop-pps = RAHArf…; sprop-sei = VaBXE… (The VPS, SPS, PPS, Sei, etc. corresponding to the stream with payload_type = 96 are base64 encoded and need no explanation)
[0147] a = control:streamed = 0 (flow control information)
[0148] Optionally, based on the depth parameters of each pixel as the depth order of each pixel, the supplementary enhancement information also includes the depth step size of the original image. The depth step size of the original image is used by the receiver to determine the depth information of each pixel based on the depth order of each pixel and the depth step size of the original image. The method for determining the depth step size of the original image will be explained later. The supplementary enhancement information also includes depth order units, which are the units of the depth parameters of each pixel. For example, if the depth order unit is centimeters, then a depth parameter of 40 for any pixel means that the depth parameter of any pixel is 40 centimeters. The supplementary enhancement information also includes the depth order median, depth order variance, and the number of bits in the depth channels. By sending SEI to the receiver, the loss of critical information can be prevented, and the SEI identifier allows the decoder and demultiplexer at the receiver to make appropriate preparations or backups.
[0149] The following is an example of supplementary and enhanced information provided in an embodiment of this application.
[0150] In one possible implementation, the process of sending image information to the receiving end includes: encoding the image information to obtain an encoded packet (NAL packet) corresponding to the original image, the encoded packet of the original image including the image information; encapsulating the encoded packet of the original image to obtain an encapsulated packet (RTP packet) corresponding to the original image; and sending the encapsulated packet of the original image to the receiving end.
[0151] In this process, image information is encoded by an encoder to obtain the encoded packet corresponding to the original image. Optionally, the encoder can be a VVC (Versatile Video Coding) encoder, an h264 encoder (a video codec), or a hevc (High Efficiency Video Coding) encoder. This application embodiment does not limit the encoder. For example, the encoder is a VVC encoder.
[0152] In one possible implementation, when the depth parameter of each pixel is the depth order of each pixel, the encoded image information and the depth step size of the original image are used to obtain the encoded packet corresponding to the original image. The encoded packet corresponding to the original image also includes the depth step size of the original image, which is used to determine the depth information of each pixel included in the original image.
[0153] This application provides two methods for determining the depth step size.
[0154] The first method determines the depth step size of the original image based on the maximum depth information in the depth information of each pixel in the original image and the number of bits in the depth channel used by the reference sampling method.
[0155] Optionally, the depth step size of the original image is determined according to the following formula (8) based on the maximum depth information in the depth information of each pixel in the original image and the number of bits of the depth channel used in the reference sampling method.
[0156] In the above formula (8), P is the depth step size of the original image, D0 is the maximum depth information in the depth information of each pixel, and N is the number of bits of the depth channel used in the reference sampling method.
[0157] The second method involves determining the depth standard deviation of the original image based on the depth information of each pixel in the original image; and determining the depth step size of the original image based on the depth standard deviation of the original image and the number of bits in the depth channel used in the reference sampling method.
[0158] In one possible implementation, the depth step size of the original image is determined according to the following formula (9) based on the depth standard deviation of the original image and the number of bits of the depth channel used in the reference sampling method.
[0159] In the above formula (9), P is the depth step size of the original image, σ is the depth standard deviation of the original image, and N is the number of bits of the depth channel used in the reference sampling method.
[0160] In one possible implementation, since the reference sampling method includes a depth channel (d), the encoder needs to add an additional channel type (chroma_format) configuration. The encoder includes metadata, which comprises three parameter sets: SPS (Sequence Parameter Set), PPS (Picture Parameter Set), and VPS (Video Parameter Set). Each parameter set requires additional entries.
[0161] The following additional items need to be added to SPS:
[0162] sps_depth_channel_flag (used to indicate whether a depth channel exists, occupying 1 bit)
[0163] sps_depth_bit_depth_minus8 (used to identify the number of bits in the depth channel, such as 8 / 10, occupying 32 bits)
[0164] sps_chroma_format_idc (used to identify the reference sampling mode, occupying 32 bits)
[0165] sps_vui_parameters.vui_viewport_params_present_flag (used to indicate whether there is an optional viewport range, occupying 1 bit)
[0166] sps_vui_parameters.vui_viewport_min_x (This identifies the minimum x-axis coordinate of the optional viewpoint, occupying 16 bits)
[0167] sps_vui_parameters.vui_viewport_min_y (This identifies the minimum y-axis coordinate of the optional viewpoint, occupying 16 bits)
[0168] sps_vui_parameters.vui_viewport_max_x (This identifies the maximum x-axis coordinate of the optional viewpoint, occupying 16 bits)
[0169] sps_vui_parameters.vui_viewport_max_y (This identifies the maximum y-coordinate of the optional viewport, occupying 16 bits)
[0170] sps_vui_parameters.vui_depth_capture_mode (This identifies the depth information capture mode and occupies 8 bits)
[0171] sps_vui_parameters.vui_depth_unit (used to identify depth units, occupying 8 bits)
[0172] sps_vui_parameters.vui_depth_stride (used to identify the depth stride, occupying 32 bits)
[0173] The possible values for sps_chroma_format_idc are shown in Table 2 below.
[0174] Table 2
[0175] In Table 2 above, when sps_chroma_format_idc = 1, it indicates that the reference sampling method is YUV420P, meaning that the sampling rate of the chroma and saturation components is one-quarter that of the luminance component. When sps_chroma_format_idc = 4, it indicates that the reference sampling method is YUVD4204P, meaning that the sampling rate of the chroma and saturation components is one-quarter that of the luminance component, and the sampling rate of the depth component is the same as that of the luminance component. When sps_chroma_format_idc is any other value, it indicates that the reference sampling method is another, which will not be described in detail in this embodiment.
[0176] The values for sps_vui_parameters.vui_depth_capture_mode can be shown in Table 3 below.
[0177] Table 3
[0178] In Table 3 above, sps_vui_parameters.vui_depth_capture_mode = 0 indicates that the depth information acquisition mode is through dual / multi-view cameras. When sps_vui_parameters.vui_depth_capture_mode is other, it indicates that the depth information acquisition mode is other, which will not be elaborated here.
[0179] The possible values for sps_vui_parameters.vui_depth_unit are shown in Table 4 below.
[0180] Table 4
[0181] In Table 4 above, when sps_vui_parameters.vui_depth_unit = 0, it means that the depth unit is meters; when sps_vui_parameters.vui_depth_unit is other, it means that the depth unit is other. The embodiments of this application will not be described in detail here.
[0182] The following is an example of an SPS data organization logic provided in an embodiment of this application. The numerical values are examples and are for reference only.
[0183] if(sps_depth_channel_flag){
[0184] sps_depth_bit_depth_minus8 = 0; (The depth channel has 8 bits)
[0185] sps_chroma_format_idc = 4; (The fourth sampling method is used as a reference)
[0186] sps_vui_parameters.vui_viewport_params_present_flag = 1; (Optional viewpoint range)
[0187] sps_vui_parameters.vui_viewport_min_x = 0; (Optional: minimum x-axis coordinate of the viewpoint is 0)
[0188] sps_vui_parameters.vui_viewport_min_y = 0; (Optional: minimum y-axis coordinate of the viewpoint is 0)
[0189] sps_vui_parameters.vui_viewport_max_x = 1920; (Optional maximum x-axis coordinate of the viewpoint is 1920)
[0190] sps_vui_parameters.vui_viewport_max_y = 1080; (Optional: Maximum y-coordinate of the viewpoint is 1080)
[0191] sps_vui_parameters.vui_depth_capture_mode = 0; (The depth information acquisition mode is set to the 0th mode, which means that the depth information is determined by dual / multi-view cameras.)
[0192] sps_vui_parameters.vui_depth_unit = 3; (The depth unit is the third type, which is feet.)
[0193] sps_vui_parameters.vui_depth_stride = 10; (depth stride is 10)
[0194] }
[0195] In one possible implementation, PPS would require the following additional items.
[0196] pps_depth_enabled_flag (used to indicate whether a depth channel is enabled, occupying 1 bit)
[0197] pps_depth_mean (used to identify the depth median, occupying 8 bits)
[0198] pps_depth_deviation (used to identify deep-order variance, occupying 8 bits)
[0199] pps_depth_roi_enabled_flag (uses 1 bit to indicate whether depth-based ROI encoding is enabled)
[0200] pps_depth_roi_qp (Identifies depth-based ROI encoding quantization parameters, occupying 8 bits)
[0201] The following is a PPS data organization logic provided in an embodiment of this application, wherein the numerical values are examples and for reference only.
[0202] if(pps_depth_enabled_flag){
[0203] pps_depth_mean = 5; (The depth median is 5)
[0204] pps_depth_deviation = 3; (depth-order variance is 3)
[0205] pps_depth_roi_enabled_flag = 1; (Enables depth-based ROI encoding)
[0206] pps_depth_roi_qp = 60; (The quantization parameter for depth-based ROI encoding is 60)
[0207] }
[0208] In one possible implementation, the VPS would require the following additional features.
[0209] vps_num_layers_in_vps_minus1 (Optional, adds depth layers)
[0210] vps_depth_layer_id (optional, depth layer ID)
[0211] The following is a VPS data organization logic provided in an embodiment of this application, wherein the values are examples and for reference only.
[0212] vps_depth_enabled_flag = 1;
[0213] if(vps_depth_enabled_flag){
[0214] vps_num_layers_in_vps_minus1 = 1; (Adds 2 depth layers, one of which is a residual layer)
[0215] vps_depth_layer_id = 3; (depth layer ID is 3)
[0216] }
[0217] if(vps_depth_residual_flag){(Does it have a residual?)
[0218] vps_depth_residual_layer_id = 9; (Residual layer ID)
[0219] }
[0220] It should be noted that if the receiving end requires additional parameters when generating the reference image, these parameters can be added to the SPS and SEI. This will not be elaborated further in the embodiments of this application.
[0221] In one possible implementation, the process of sending image information to the receiving end includes sending image information to the receiving end via RTSP (Real-Time Streaming Protocol).
[0222] In some embodiments, the transmitting end can also acquire the original image, input the original image into the encoder, and the encoder processes the original image to obtain an encoded packet. The encoder's processing flow is shown in Figure 7. The original image is the input to the encoder, representing the image to be encoded. The original image can be in RGB format, RGBD (Red-Green-Blue-Depth) format, YUV format, RGBA format, or YUVD (Luminosity-Chromaticity-Saturation-Depth) format. The encoding controller is responsible for the management and control of the entire encoding process, including setting encoding control parameters and controlling the encoding flow. The depth estimation module is responsible for estimating the depth information of the original image. The frame prediction module performs intra / inter-frame prediction based on the depth information and encoding control parameters to generate a prediction frame; the loop filtering module filters the prediction frame to reduce block artifacts and noise, and calculates the image residual between the original image and the filtered image. The quantization module quantizes the image residual to obtain a quantization residual for further data compression. The entropy coding module entropy codes the quantization residual to obtain an encoded packet, which is the output of the encoding process.
[0223] Optionally, after sending image information to the receiving end, the receiving end receives the image information, determines the pixel information and depth parameters of each pixel based on the image information, and generates an original image and a reference image corresponding to the original image based on the pixel information and depth parameters of each pixel. When generating the reference image corresponding to the original image, the receiving end needs to first generate a first intermediate image, the size of which is larger than the size of the original image; crop the first intermediate image to obtain a second intermediate image, the size of which is the size of the original image, and the second intermediate image is different from the first intermediate image, including blank areas; fill the blank areas in the second intermediate image to obtain a third intermediate image; and then determine the reference image based on the third intermediate image.
[0224] The process of filling blank areas in the second intermediate image to obtain the third intermediate image at the receiving end includes: the receiving end sending the location information of the blank areas in the second intermediate image to the sending end; the sending end receiving the location information of the blank areas, calculating the edge image based on the location information, and sending the edge image to the receiving end in the form of a residual; the receiving end receiving the edge image and filling the blank areas with the edge image to obtain the third intermediate image. The location information of the blank areas includes the minimum x-coordinate, maximum x-coordinate, minimum y-coordinate, and maximum y-coordinate of the blank areas. The residual is the difference taken from the macroblock during frame prediction in the coding stage, which is then entropy-coded and entered into the next IDR frame (Instant Decoding Refresh Frame).
[0225] Optionally, when sending the image information of the next frame, the sending end also sends the edge image to the receiving end. For example, a layer parameter is added to the SPS of the encoded packet of the next frame. The layer parameter is used to indicate the edge image, so that the decoding end can perform edge compensation. The following is an example of a layer parameter provided by an embodiment of this application.
[0226] sps_edge_residual_block_flag = 1; (Image with edges)
[0227] if(sps_edge_residual_block_flag){(if there is an edge image){
[0228] sps_edge_residual_layerid = 8; (The layer ID where the edge image is located = 8)
[0229] }
[0230] In one possible implementation, the receiving end can adjust the depth step size of the original image and send the adjusted depth step size to the sending end. Upon receiving the adjusted depth step size, the sending end adjusts the depth step size of the original image accordingly. Optionally, the receiving end can also adjust the bit depth of the depth channel used in the reference sampling method and send the adjusted bit depth to the sending end. Upon receiving the adjusted bit depth, the sending end adjusts the bit depth of the depth channel used in the reference sampling method accordingly.
[0231] In one possible implementation, a region in the original image is designated as the Region of Interest (ROI), such as a region with depth information within the range of u ± σ (where u is the median depth value and σ is the variance depth value). The quantization parameter of this region can be reduced to enhance its image quality and improve the user's viewing experience. Optionally, the receiving end can send a quantization parameter acquisition request to the sending end. After receiving the request, the sending end sends the quantization parameters to the receiving end, allowing the receiving end to adjust the quantization parameters of the ROI region accordingly to enhance its image quality.
[0232] The above method requires the sending end to acquire only the pixel information and depth parameters of each pixel in the original image, and then send image information determined based on the pixel information and depth parameters to the receiving end. This allows the receiving end to generate the original image and its corresponding reference image based on the image information, and then generate a target image with a 3D effect based on the original image and the reference image. Since the sending end only needs to send information from the original image itself, the amount of content sent is less, reducing the bandwidth required to transmit content to the receiving end, thereby improving the content transmission efficiency and ultimately enhancing the efficiency of image processing.
[0233] Moreover, the content sent by the sending end does not need to be determined based on the number of viewpoints supported by the receiving end's display screen. In this way, regardless of the number of viewpoints supported by the receiving end's display screen, only the information of the original image itself is sent to the receiving end. This decouples the sending end and the receiving end, thereby improving the flexibility of image processing.
[0234] Figure 8 is a flowchart of an image processing method provided in an embodiment of this application. This method can be applied to the implementation environment shown in Figure 1 above, and can be executed by the receiving end 102 in Figure 1. As shown in Figure 8, the method includes the following steps 801 to 804.
[0235] In step 801, image information sent by the sending end is received.
[0236] The image information is derived from the pixel information and depth parameters of each pixel in the original image. The pixel information describes the brightness, chroma, and saturation of that pixel, while the depth parameter indicates the distance between that pixel and the image acquisition device used to acquire the original image. The depth parameter can be either the pixel's depth information or its depth level. The depth information indicates the actual distance between the pixel and the image acquisition device, while the depth level indicates the relative distance between them.
[0237] In one possible implementation, the receiving end sends a packet corresponding to the original image, unpacks the packet to obtain an encoded packet, and then decodes the encoded packet to obtain the image information.
[0238] For example, the received image information is Y1Y2Y3Y4Y5Y6Y7Y8U1U3V5V7D1D2D3D4D5D6D7D8.
[0239] Optionally, decoding the encoded packet corresponding to the original image can also yield sequence parameter sets, image parameter sets, and video parameter sets.
[0240] In step 802, the pixel information and depth parameters of each pixel are determined based on the image information.
[0241] In one possible implementation, the receiver receives supplemental enhancement information sent by the sender. This supplemental enhancement information indicates a reference sampling method, which is used to determine the pixel information and depth parameters of each pixel. Optionally, the receiver receives a session description protocol sent by the sender, which includes the supplemental enhancement information. The reference sampling method is also included in the sequence parameter set after decoding the encoded packets corresponding to the original image.
[0242] The process of determining the pixel information and depth parameters of each pixel based on the image information includes: determining the pixel information and depth parameters of each pixel based on the reference sampling method and the image information.
[0243] For example, the image information is Y1Y2Y3Y4Y5Y6Y7Y8U1U3V5V7D1D2D3D4D5D6D7D8, and the reference sampling method is yuvd4202p. Since the image information includes 8 luminance components, it can be determined that the original image includes 8 pixels. The reference sampling method yuvd4202p indicates that the sampling ratio of the luminance component, chroma component, saturation component, and depth parameter is 1:1 / 4:1 / 4:1. The following parameters are determined: For the first pixel, the luminance component is Y1, the chroma component is U1, the saturation component is V5, and the depth parameter is D1; for the second pixel, the luminance component is Y2, the chroma component is U1, the saturation component is V5, and the depth parameter is D2; for the third pixel, the luminance component is Y3, the chroma component is U3, the saturation component is V7, and the depth parameter is D3; for the fourth pixel, the luminance component is Y4, the chroma component is U3, the saturation component is V7, and the depth parameter is D4; for the fifth pixel, the luminance component is Y5, the chroma component is U1, the saturation component is V5, and the depth parameter is D5; for the sixth pixel, the luminance component is Y6, the chroma component is U1, the saturation component is V5, and the depth parameter is D6; for the seventh pixel, the luminance component is Y7, the chroma component is U3, the saturation component is V7, and the depth parameter is D7; and for the eighth pixel, the luminance component is Y8, the chroma component is U3, the saturation component is V7, and the depth parameter is D7.
[0244] In step 803, an original image and a reference image corresponding to the original image are generated based on the pixel information and depth parameters of each pixel. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end.
[0245] The receiving end's display screen can be a 3D display connected to the receiving end. The 3D display supports multiple viewpoints. The sum of the number of reference images and the original image is the number of viewpoints supported by the receiving end's display screen. That is, the number of reference images is the difference between the number of viewpoints supported by the receiving end's display screen and 1. For example, if the receiving end's display screen supports 2 viewpoints (i.e., a 2-viewpoint display screen), the number of reference images is 1. As another example, if the receiving end's display screen supports 4 viewpoints (i.e., a 4-viewpoint display screen), the number of reference images is 3. As yet another example, if the receiving end's display screen supports 9 viewpoints (i.e., a 9-viewpoint display screen), the number of reference images is 8. And as yet another example, if the receiving end's display screen supports 16 viewpoints (i.e., a 16-viewpoint display screen), the number of reference images is 15.
[0246] In one possible implementation, the depth information of each pixel is determined based on the depth parameters of each pixel, and the depth information of any pixel is used to indicate the actual distance between any pixel and the image acquisition device that acquires the original image; based on the pixel information and depth information of each pixel, the original image and the reference image corresponding to the original image are generated.
[0247] Specifically, based on the depth parameters of each pixel, the depth information of each pixel is obtained. Then, based on the pixel information and depth parameters of each pixel, the original image and its corresponding reference image are generated. Based on the depth parameters of each pixel, the depth order of each pixel is obtained. The encoding packet also includes the depth step size of the original image. Decoding the encoding packet corresponding to the original image yields the depth step size of the original image. Based on the depth step size of the original image and the depth order of each pixel, the depth information of each pixel is determined. Finally, based on the pixel information and depth information of each pixel, the original image and its corresponding reference image are generated.
[0248] Optionally, the process of determining the depth information of each pixel based on the depth step size of the original image and the depth order of each pixel includes: using the product between the depth step size of the original image and the depth order of each pixel as the depth information of each pixel.
[0249] For example, if the depth step of the original image is 10 and the depth order of any pixel is 10, then the depth information of any pixel is 10*10=100.
[0250] In one possible implementation, the process of generating an original image and a corresponding reference image based on the pixel information and depth information of each pixel includes: generating an original image and a first intermediate image corresponding to the original image based on the pixel information and depth information of each pixel, wherein the size of the first intermediate image is larger than the size of the original image; cropping the first intermediate image to obtain a second intermediate image, wherein the size of the second intermediate image is the same as the size of the original image and the second intermediate image is different from the first intermediate image; filling the blank areas in the second intermediate image to obtain a third intermediate image; and determining the reference image corresponding to the original image based on the third intermediate image.
[0251] The original image can have one or more reference images, the number of which is determined by the number of viewpoints supported by the display screen. The generation process for each reference image corresponding to the original image is consistent; the above describes the generation process for any reference image corresponding to the original image.
[0252] In one possible implementation, each pixel corresponds to a location information, and the original image is generated based on the location information, pixel information, and depth information of each pixel. Optionally, the location information, pixel information, and depth information of each pixel are input into an image generation model to obtain the original image. The image generation model can be any model capable of generating a corresponding image based on the location information, pixel information, and depth information of pixels; this embodiment does not limit its scope.
[0253] In one possible implementation, each pixel corresponds to a position information, and a first intermediate image corresponding to the original image is generated based on the position information, pixel information, and depth information of each pixel.
[0254] Optionally, the process of generating a first intermediate image corresponding to the original image based on the position information, pixel information, and depth information of each pixel includes: determining the reflection vector corresponding to the first intermediate image, wherein the reflection vector corresponding to the first intermediate image is determined based on the position of the first intermediate image relative to the original image; determining the offset position information of each pixel based on the position information, depth information, and reflection vector corresponding to the first intermediate image; and generating the first intermediate image corresponding to the original image based on the pixel information, depth information, and offset position information of each pixel.
[0255] The reflection vector of the first intermediate image varies depending on its position relative to the original image. For example, if the first intermediate image is to the left of the original image, its reflection vector is (1, 0); if it is below the original image, its reflection vector is (0, -1); if it is to the lower left of the original image, its reflection vector is (1, -1); if it is to the right of the original image, its reflection vector is (-1, 0); if it is above the original image, its reflection vector is (0, 1); if it is to the upper right of the original image, its reflection vector is (-1, 1); if it is to the lower right of the original image, its reflection vector is (-1, -1); and if it is to the upper left of the original image, its reflection vector is (1, 1). When the first intermediate image is in a different position than the original image, the reflection vector corresponding to the first intermediate image is different, which will not be described in detail here.
[0256] Optionally, the reflection vector corresponding to the first intermediate image is a two-dimensional vector, and the position information of any pixel includes the x-coordinate and y-coordinate of any pixel. The process of determining the offset position information of each pixel based on its position information, depth information, and the reflection vector corresponding to the first intermediate image includes: for any pixel, determining the offset x-coordinate based on its x-coordinate, depth information, and the value of the first dimension of the reflection vector corresponding to the first intermediate image; determining the offset y-coordinate based on its y-coordinate, depth information, and the value of the second dimension of the reflection vector corresponding to the first intermediate image; and finally, determining the offset position information of any pixel based on its offset x-coordinate and y-coordinate.
[0257] The horizontal coordinate of any pixel after offset is determined according to the following formula (10), based on the horizontal coordinate of any pixel, the depth information of any pixel, and the value of the first dimension in the reflection vector corresponding to the first intermediate image.
[0258] In the above formula (10), X2 is the x-coordinate of any pixel after offset, A xZ represents the value of the first dimension in the reflection vector corresponding to the first intermediate image, Z represents the depth information of any pixel, X1 represents the x-coordinate of any pixel, f represents the focal length of the image acquisition device that acquired the original image, and T represents the distance between the two cameras when the image acquisition device that acquired the original image is a binocular camera. When the image acquisition device that acquired the original image is another device, T represents other values. For example, when the image acquisition device that acquired the original image is another device, T is set based on experience or adjusted according to the implementation environment. This application embodiment does not limit this.
[0259] Based on the ordinate of any pixel, the depth information of any pixel, and the value of the second dimension in the reflection vector corresponding to the first intermediate image, the ordinate of any pixel after offset is determined according to the following formula (11).
[0260] In the above formula (11), Y2 is the ordinate of any pixel after offset, and A y Z represents the value of the second dimension in the reflection vector corresponding to the first intermediate image, Z represents the depth information of any pixel, Y1 represents the ordinate of any pixel, f represents the focal length of the image acquisition device that acquired the original image, and T represents the distance between the two cameras when the image acquisition device that acquired the original image is a binocular camera. When the image acquisition device that acquired the original image is another device, T represents other values. For example, when the image acquisition device that acquired the original image is another device, T is set based on experience or adjusted according to the implementation environment. This application embodiment does not limit this.
[0261] In one possible implementation, the process of determining the position information of any pixel after its offset, based on the offset x-coordinate and offset y-coordinate of any pixel, includes: using the position information composed of the offset x-coordinate and offset y-coordinate of any pixel as the offset position information of any pixel. For example, the offset x-coordinate of any pixel is X2, the offset y-coordinate is Y2, and the offset position information of any pixel is (X2, Y2).
[0262] The process of generating a first intermediate image corresponding to the original image based on the pixel information, depth information, and offset position information of each pixel includes: using the pixel information and depth information of each pixel as the offset position information and depth information of each pixel; and generating the first intermediate image corresponding to the original image based on the offset position information, the pixel information of the offset position information, and the depth information. The process of generating the first intermediate image corresponding to the original image based on the offset position information, the pixel information of the offset position information, and the depth information is similar to the process of generating the original image described above, and will not be repeated here.
[0263] Figure 9 is a schematic diagram of an original image and a corresponding first intermediate image provided in an embodiment of this application. 901 is the original image, and 902 is the first intermediate image. The size of the first intermediate image 902 is larger than the size of the original image 901.
[0264] In one possible implementation, the encoding packet corresponding to the original image also includes a sequence parameter set. This sequence parameter set includes the dimensions of the original image. For example, the sequence parameter set might include the minimum x-coordinate of the selectable viewpoint being 0, the maximum x-coordinate being 1920, the minimum y-coordinate being 0, and the maximum y-coordinate being 1280, meaning the original image size is 1920*1280. Parsing the encoding packet corresponding to the original image further reveals the dimensions of the original image. Based on these dimensions, the first intermediate image is cropped to obtain a second intermediate image. The second intermediate image has the same dimensions as the original image, but is different from the original image.
[0265] Optionally, if the first intermediate image is the image on the left side of the original image, then a portion of the right side of the first intermediate image is cropped to obtain the second intermediate image. If the first intermediate image is the image on the right side of the original image, then a portion of the left side of the first intermediate image is cropped to obtain the second intermediate image. If the first intermediate image is the image on the top side of the original image, then a portion of the bottom side of the first intermediate image is cropped to obtain the second intermediate image. If the first intermediate image is the image on the bottom side of the original image, then a portion of the top side of the first intermediate image is cropped to obtain the second intermediate image. If the first intermediate image is the image on the top left side of the original image, then a portion of the right and bottom sides of the first intermediate image is cropped to obtain the second intermediate image. If the first intermediate image is the image on the bottom left side of the original image, then a portion of the right and top sides of the first intermediate image is cropped to obtain the second intermediate image. If the first intermediate image is the image on the top right side of the original image, then a portion of the left and bottom sides of the first intermediate image is cropped to obtain the second intermediate image. Since the first intermediate image is the image on the lower right side of the original image, the left and upper parts of the first intermediate image are cropped to obtain the second intermediate image.
[0266] Figure 10 is a schematic diagram of determining a second intermediate image according to an embodiment of this application. Since the first intermediate image 1001 is the left side of the original image, a portion 1002 of the right side of the first intermediate image 1001 is cropped to obtain the second intermediate image 1003. The second intermediate image 1003 has the same size as the first intermediate image 1001, but the second intermediate image 1003 is different from the first intermediate image 1001.
[0267] Since the first intermediate image is an image offset from the original image, and the second intermediate image is an image cropped from a portion of the first intermediate image, there are blank areas in the second intermediate image, which need to be filled to obtain the third intermediate image.
[0268] Optionally, the process of filling blank areas to obtain a third intermediate image is not limited in this embodiment. Optionally, the receiving end performs edge prediction using a convolutional neural network model to obtain the filling content for the blank areas, and fills the blank areas with the filling content to obtain the third intermediate image. Alternatively, the receiving end sends the location information of the blank areas to the sending end; the sending end receives the location information of the blank areas, calculates an edge image based on the location information of the blank areas, and sends the edge image to the receiving end; the receiving end receives the edge image and fills the blank areas with the edge image to obtain the third intermediate image. The location information of the blank areas includes the minimum x-coordinate, maximum x-coordinate, minimum y-coordinate, and maximum y-coordinate of the blank areas. Figure 11 is a schematic diagram of determining a third intermediate image provided by an embodiment of this application. 1101 is the second intermediate image, 1102 is the filling content for the blank areas, and the filling content 1102 is filled into the blank areas of the second intermediate image 1101 to obtain the third intermediate image 1103.
[0269] Optionally, the receiving end sends the location information of the blank area to the sending end via RTCP (Real-Time Transport Control Protocol) messages. The RTCP message is called an rtcp receiver report (a type of RTCP message).
[0270] Optionally, since the depth step size of the original image is used to determine the depth information of each pixel, and the depth information of each pixel affects the in-screen and out-of-screen effects of the target image, the receiving end can also adjust the depth step size of the original image to improve the 3D stereoscopic effect and in-screen and out-of-screen effects of the generated target image. Specifically, an out-of-screen effect occurs when an object in the image appears closer to the viewer than the actual distance to the display screen. This can be achieved by reducing the object's depth information (i.e., reducing the depth step size of the original image), making the object appear closer to the viewer. Conversely, an in-screen effect occurs when an object in the image appears farther away than the actual distance to the viewer. This can be achieved by increasing the object's depth information (i.e., increasing the depth step size of the original image), making the object appear farther away.
[0271] Figure 12 is a schematic diagram of an in-screen / out-of-screen effect provided in an embodiment of this application. The screen baseline is a horizontal dashed line representing the distance from the viewer's eye to the display screen, serving as a reference line for the 3D visual effect. The depth axis is a dashed line perpendicular to the screen baseline, representing the foreground / background position of an object in 3D space. 1201 represents the display screen. The foreground B is an object in the original image closest to the observer. Since the foreground B is in front of the screen baseline, meaning it appears closer than the actual distance from the viewer's eye to the display screen, an out-of-screen effect occurs. A smaller depth step size for the foreground B, due to its smaller depth information and more forward depth position (i.e., closer to the observer), produces a better out-of-screen effect than the foreground B. The background A is an object in the original image farther from the observer. Since the background A is behind the screen baseline, meaning it appears farther than the actual distance from the viewer's eye to the screen, an in-screen effect occurs. Increasing the depth step size will result in a better screen entry effect for background A, as it has more depth information and a more rearward depth position (i.e., further away from the observer).
[0272] The receiver can also adjust the depth step size of the original image and send the adjusted depth step size to the transmitter, so that the transmitter can adjust the depth step size of the original image accordingly. The receiver can also adjust the bit depth of the depth channels in the reference sampling mode and send the adjusted depth step size to the transmitter, so that the transmitter can adjust the bit depth of the depth channels in the reference sampling mode accordingly. The receiver can also adjust the depth unit and send the adjusted depth unit to the transmitter, so that the transmitter can adjust the depth unit accordingly.
[0273] Figure 13 illustrates an RTP receiver report message format provided in an embodiment of this application. The headers SSRC and above are general RTP headers and will not be discussed further here. The fields following SSRC are explained as follows: DeepBitDepthMinus8 changes the depth level of the encoded depth map (changing the number of bits in the depth channel used by the reference sampling method); ViewPortMinX is the minimum possible X-coordinate of the composite viewpoint (the minimum X-coordinate of the blank area in the composite map); ViewPortMinY is the minimum possible Y-coordinate of the composite viewpoint (the minimum Y-coordinate of the blank area in the composite map); ViewPortMaxX is the maximum possible X-coordinate of the composite viewpoint (the maximum X-coordinate of the blank area in the composite map); ViewPortMaxY is the maximum possible Y-coordinate of the composite viewpoint (the maximum Y-coordinate of the blank area in the composite map); R indicates whether depth ROI encoding is enabled; RoiQP refers to the quantization parameters of the ROI encoded image; C indicates whether residual layer encoding is enabled; Unit is the depth unit determined by the receiver; and Stride is the depth step size determined by the receiver.
[0274] When R is 1, it indicates that RoiQP is active. RoiQP is a 7-bit parameter, so it can take values from 0 to 127. Since the highest bit is usually used as the sign bit, but in this case it's an unsigned integer, the maximum value is 128. The quantization parameter range defined in VVC encoding is 0-63, meaning that RoiQP's value range completely covers the quantization parameter range in VVC encoding, and it also provides additional values for expansion.
[0275] In one possible implementation, the process of determining the reference image corresponding to the original image based on the third intermediate image includes: using the third intermediate image as the reference image corresponding to the original image; or, repairing the hole regions in the third intermediate image and using the repaired image as the reference image corresponding to the original image. Here, the hole regions in the third intermediate image refer to the regions composed of pixels where the luminance, chroma, and saturation components are all zero.
[0276] Optionally, the third intermediate image contains hollow regions, as shown in Figure 11, where 1103 and 1104 are hollow regions. Therefore, the hollow regions in the third intermediate image can be repaired, and the repaired image can be used as the reference image corresponding to the original image, so that the reference image corresponding to the original image does not contain hollow regions.
[0277] This application does not limit the process of repairing the hole regions in the third intermediate image. Optionally, the hole regions in the third intermediate image can be repaired using the DIBR (Depth Image Based Rendering, a computer graphics technique) algorithm. Alternatively, the hole regions in the third intermediate image can be repaired using the Crimnisi (an algorithm for image inpainting) algorithm.
[0278] Crimnisi is an image inpainting algorithm. It selects a pixel with the highest depth information at the edge of a hole region, constructs an n*n pixel block centered on this pixel, searches for the most similar sample block in the intact region of the third intermediate image, and updates the pixel block with the found sample block to complete the inpainting. This process is repeated iteratively until the hole region is completely repaired, yielding a reference image corresponding to the original image. Here, n is a value greater than 0, and the intact region of the third intermediate image is the region in the third intermediate image excluding the hole region. Figure 14 is a schematic diagram showing a reference image corresponding to the original image provided in an embodiment of this application. The reference image does not contain hole regions.
[0279] When repairing the hole areas in the third intermediate image, the principle of depth processing is used. High-depth pixels are covered by low-depth pixels. First, the pixels with the farthest depth are repaired, and then the pixels with the closest depth are repaired. This allows the foreground repair image to cover the background repair image, resulting in a better effect for the repaired reference image.
[0280] Figure 15 is a schematic diagram of a cavity repair method provided in an embodiment of this application. It includes images from different perspectives (left view, right view, top view, bottom view, middle view, and lower right view). Each image is divided into three layers: foreground, middle ground, and background. The dashed arrows and numbers in the images indicate the processing flow and order between different layers, and the dashed areas represent cavity areas in the images. In the left view, the foreground, middle ground, and background are arranged sequentially, and none of them have cavity areas, so no crimnisi processing is required. In the right view, the foreground, middle ground, and background are arranged sequentially, and each corresponds to a cavity area, requiring crimnisi processing. The processing order is: first process the background, then the middle ground, and finally the foreground. In the top view, the foreground, middle ground, and background are arranged sequentially, and none of them have cavity areas, so no crimnisi processing is required. In the bottom view, the foreground, middle ground, and background are arranged sequentially, each corresponding to a hole area. These require crimnisi processing, with the processing order being: background first, then middle ground, and finally foreground. In the middle view, the foreground, middle ground, and background are arranged sequentially, and none have hole areas, so crimnisi processing is not needed. In the bottom right view, the foreground, middle ground, and background are arranged sequentially, each corresponding to a hole area. These require crimnisi processing, with the processing order being: background first, then middle ground, and finally foreground.
[0281] In one possible implementation, a reference image corresponding to the original image is generated according to the number of viewpoints supported by the display screen of the receiving end, following the process described above. Figure 16 is a schematic diagram of a reference image corresponding to an original image provided in an embodiment of this application. When the number of viewpoints supported by the display screen of the receiving end is 2, the reference image corresponding to the original image 1601 is 1602. When the number of viewpoints supported by the display screen of the receiving end is 4, the reference images corresponding to the original image 1601 are 1603, 1604, and 1605. When the number of viewpoints supported by the display screen of the receiving end is 9, the reference images corresponding to the original image 1601 are 1606, 1607, 1608, 1609, 1610, 1611, 1612, and 1613.
[0282] In one possible implementation, the receiver includes a decoder. After receiving the encapsulated packet corresponding to the original image sent by the sender, the receiver unpacks the encapsulated packet corresponding to the original image to obtain the encoded packet corresponding to the original image. The encoded packet corresponding to the original image is then input into the decoder to obtain the original image and the reference image corresponding to the original image.
[0283] Figure 17 is a decoding flowchart of a decoder provided in an embodiment of this application. The process includes: inputting the encoded packet corresponding to the original image into the decoder. The decoder includes a decoding controller, an entropy encoding module, an inverse transform / quantization module, a frame prediction module, an inner loop filtering module, and a post-filtering module. The decoding controller is responsible for the management and control of the entire decoding process. The decoding controller processes the encoded packet to obtain encoded data and a parameter set. The encoded data is sent to the entropy encoding module for entropy encoding to obtain the sampling information of the original image; the sampling information and parameter set of the original image are sent to the inverse transform / quantization module for inverse transform / quantization processing to obtain the pixel information and depth information of each pixel in the original image; the pixel information and depth information of each pixel in the original image are sent to the frame prediction module to obtain a predicted frame; the predicted frame is input to the inner loop filtering module for filtering processing to obtain the original image. A first intermediate image corresponding to the original image is generated through the frame prediction module, and the first intermediate image corresponding to the original image is cropped and padded to obtain a third intermediate image corresponding to the original image. The third intermediate image corresponding to the original image is input into the post-filtering processing module to repair the hole regions in the third intermediate image corresponding to the original image, thereby obtaining the reference image corresponding to the original image.
[0284] In step 804, a target image with a three-dimensional effect is generated based on the original image and the reference image.
[0285] In one possible implementation, after generating the original image and the reference image in step 803 above, the process of generating a target image with a three-dimensional effect based on the original image and the reference image includes: performing an interlacing operation on the original image and the reference image to obtain a target image with a three-dimensional effect.
[0286] Optionally, the interleaving algorithm corresponding to the display screen of the receiving end is determined, and the original image and the reference image are interleaved according to the interleaving algorithm corresponding to the display screen of the receiving end to obtain a target image with a three-dimensional stereoscopic effect. The interleaving algorithm varies depending on the implementation of the three-dimensional stereoscopic effect polarization process of each manufacturer.
[0287] In some embodiments, the image parameter set included in the encoding packet corresponding to the original image includes deep-order median, deep-order variance, enabled depth-based ROI coding quantization parameters, and depth-based ROI coding quantization parameters. Enabling depth-based ROI coding quantization parameters means that the receiving end can enable depth-based ROI coding quantization parameters. The receiving end can determine the region of interest (ROI) of the original image based on the deep-order median and deep-order variance; modifying the quantization parameters of the ROI to depth-based ROI coding quantization parameters reduces the quantization parameters of the ROI, thereby improving the image quality of the ROI and enhancing the user's viewing experience of the original image.
[0288] Optionally, if the image parameter set included in the encoding packet corresponding to the original image does not include depth-based ROI encoding quantization parameters, the receiving end can send a request to the sending end to enable depth-based ROI encoding quantization parameters in order to obtain the depth-based ROI encoding quantization parameters sent by the sending end to the receiving end, and then modify the quantization parameters of the user's region of interest to depth-based ROI encoding quantization parameters.
[0289] The receiving end sends a request to the sending end via the RTCP protocol to enable depth-based ROI encoding quantization parameters.
[0290] The above method requires the sending end to acquire only the pixel information and depth parameters of each pixel in the original image, and then send image information determined based on the pixel information and depth parameters to the receiving end. This allows the receiving end to generate the original image and its corresponding reference image based on the image information, and then generate a target image with a 3D effect based on the original image and the reference image. Since the sending end only needs to send information from the original image itself, the amount of content sent is less, reducing the bandwidth required to transmit content to the receiving end, thereby improving the content transmission efficiency and ultimately enhancing the efficiency of image processing.
[0291] Moreover, the content sent by the sending end does not need to be determined based on the number of viewpoints supported by the receiving end's display screen. In this way, regardless of the number of viewpoints supported by the receiving end's display screen, only the information of the original image itself is sent to the receiving end. This decouples the sending end and the receiving end, thereby improving the flexibility of image processing.
[0292] Figure 18 is a schematic diagram of an image processing method provided in an embodiment of this application. The executing entities include a sending end and a receiving end. The sending end includes an image acquisition device and an encoder, and the receiving end includes a decoder and a display screen.
[0293] Image acquisition devices are used to acquire raw images;
[0294] The encoder determines the pixel information and depth parameters of each pixel in the original image; based on the pixel information and depth parameters, it obtains the encoded packet. This way, the encoder only needs to encode one original image, resulting in a smaller encoded packet, which facilitates packet transmission and improves transmission efficiency.
[0295] The sending end encapsulates the encoded packet to obtain the encapsulated packet, and then sends the encapsulated packet to the receiving end.
[0296] The receiving end receives the encapsulated packet, decapsulates it, and obtains the encoded packet.
[0297] The decoder parses the encoded packet to obtain the pixel information and depth parameters of each pixel; based on the pixel information and depth parameters of each pixel, it generates the original image and a corresponding reference image. The number of reference images is determined based on the number of viewpoints supported by the display screen.
[0298] The display screen shows a target image with a three-dimensional effect based on the original image and its corresponding reference image. The display screen can also adjust the depth step size to adjust the in-screen and out-of-screen effects of the target image.
[0299] Figure 19 is a schematic diagram of an image processing apparatus provided in an embodiment of this application. As shown in Figure 19, the apparatus includes:
[0300] The acquisition module 1901 is used to acquire pixel information of multiple pixels in the original image, whereby the pixel information of any pixel is used to describe the brightness, chroma, and saturation of any pixel.
[0301] The determination module 1902 is used to determine the depth parameters of each pixel. The depth parameters of any pixel are used to indicate the distance between any pixel and the image acquisition device that acquires the original image.
[0302] The acquisition module 1901 is also used to acquire image information based on pixel information and depth parameters of multiple pixels;
[0303] The transmitting module 1903 is used to transmit image information to the receiving end. The image information is used by the receiving end to generate an original image and a reference image corresponding to the original image. Based on the original image and the reference image, a target image with a three-dimensional stereoscopic effect is generated. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end.
[0304] In one possible implementation, when the depth information of each pixel is located within the reference interval, the depth parameter of each pixel is the depth information of each pixel. The reference interval is determined based on the number of bits of the depth channel used in the reference sampling method. The depth information of any pixel is used to indicate the actual distance between any pixel and the image acquisition device.
[0305] In one possible implementation, when there are pixels whose depth information is outside the reference interval, the depth parameter of each pixel is the depth order of each pixel. The depth order of any pixel is located within the reference interval, which is determined based on the number of bits of the depth channel used in the reference sampling method. The depth information of any pixel is used to indicate the actual distance between any pixel and the image acquisition device.
[0306] In one possible implementation, the determining module 1902 is further configured to determine the depth level of each pixel based on the number of bits of the depth channel used in the reference sampling method, the maximum depth information in the depth information of each pixel, and the depth information of each pixel.
[0307] In one possible implementation, the image information includes sampling information of the original image;
[0308] The acquisition module 1901 is used to sample the pixel information and depth parameters of each pixel according to the reference sampling method to obtain the sampling information of the original image. The sampling information of the original image is used by the receiving end to determine the pixel information and depth parameters of each pixel included in the original image.
[0309] In one possible implementation, the transmitting module 1903 is further configured to transmit supplementary enhancement information to the receiving end, the supplementary enhancement information indicating a reference sampling method, the reference sampling method being used by the receiving end to determine the pixel information and depth parameters of each pixel.
[0310] In one possible implementation, the sending module 1903 is used to encode image information to obtain an encoded packet corresponding to the original image, the encoded packet of the original image including image information; encapsulate the encoded packet of the original image to obtain an encapsulated packet corresponding to the original image; and send the encapsulated packet corresponding to the original image to the receiving end.
[0311] In one possible implementation, the depth parameter of each pixel is the depth order of that pixel.
[0312] The transmitting module 1903 is used to encode image information and the depth step size of the original image to obtain the encoded packet corresponding to the original image. The encoded packet corresponding to the original image also includes the depth step size of the original image. The depth step size of the original image is used to determine the depth information of each pixel in the original image.
[0313] In one possible implementation, the determining module 1902 is further configured to determine the depth step size of the original image based on the maximum depth information in the depth information of each pixel included in the original image and the number of bits of the depth channel used in the reference sampling method.
[0314] In one possible implementation, the determining module 1902 is further configured to determine the depth standard deviation of the original image based on the depth information of each pixel included in the original image; and to determine the depth step size of the original image based on the depth standard deviation of the original image and the number of bits of the depth channel used in the reference sampling method.
[0315] Figure 20 is a schematic diagram of an image processing apparatus provided in an embodiment of this application. As shown in Figure 20, the apparatus includes:
[0316] The receiving module 2001 is used to receive image information sent by the sending end. The image information is obtained based on the pixel information and depth parameters of multiple pixels included in the original image. The pixel information of any pixel is used for the brightness, chroma and saturation of any pixel. The depth parameter of any pixel is used to indicate the distance between any pixel and the image acquisition device that acquired the original image.
[0317] The determination module 2002 is used to determine the pixel information and depth parameters of each pixel based on the image information;
[0318] The generation module 2003 is used to generate an original image and a reference image corresponding to the original image based on the pixel information and depth parameters of each pixel. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end.
[0319] The generation module 2003 is also used to generate a target image with a three-dimensional effect based on the original image and the reference image.
[0320] In one possible implementation, the receiving module 2001 is used to receive the encapsulation packet corresponding to the original image sent by the sending end; to unpack the encapsulation packet corresponding to the original image to obtain the encoded packet corresponding to the original image; and to decode the encoded packet corresponding to the original image to obtain image information.
[0321] In one possible implementation, the depth parameter of each pixel is the depth order of each pixel, and the encoding packet also includes the depth step of the original image;
[0322] The receiving module 2001 is also used to decode the encoded packet corresponding to the original image to obtain the depth step of the original image;
[0323] The generation module 2003 is used to determine the depth information of each pixel based on the depth step size and depth order of the original image; and to generate the original image and the reference image corresponding to the original image based on the pixel information and depth information of each pixel.
[0324] In one possible implementation, the generation module 2003 is used to generate an original image and a first intermediate image corresponding to the original image based on the pixel information and depth information of each pixel. The size of the first intermediate image is larger than that of the original image. The first intermediate image is cropped to obtain a second intermediate image, the size of which is the same as that of the original image. If the second intermediate image includes blank areas, the blank areas are filled to obtain a third intermediate image. The reference image corresponding to the original image is determined based on the third intermediate image.
[0325] In one possible implementation, the generation module 2003 is used to use the third intermediate image as a reference image corresponding to the original image; or, to repair the hole regions in the third intermediate image and use the repaired image as a reference image corresponding to the original image.
[0326] In one possible implementation, the generation module 2003 is used to determine the reflection vector corresponding to the first intermediate image, the reflection vector corresponding to the first intermediate image being determined based on the position of the first intermediate image relative to the original image; determine the offset position information of each pixel based on the position information, depth information and the reflection vector corresponding to the first intermediate image; and generate the first intermediate image corresponding to the original image based on the pixel information, depth information and the offset position information of each pixel.
[0327] In one possible implementation, the reflection vector corresponding to the first intermediate image is a two-dimensional vector, and the position information of any pixel includes the horizontal and vertical coordinates of any pixel.
[0328] The generation module 2003 is used to determine the offset horizontal coordinate of any pixel point based on the horizontal coordinate of the pixel point, the depth information of the pixel point, and the value of the first dimension in the reflection vector corresponding to the first intermediate image; determine the offset vertical coordinate of any pixel point based on the vertical coordinate of the pixel point, the depth information of the pixel point, and the value of the second dimension in the reflection vector corresponding to the first intermediate image; and determine the offset position information of any pixel point based on the offset horizontal coordinate and the offset vertical coordinate of the pixel point.
[0329] It should be understood that the above-described apparatus is only illustrated by the division of the functional modules described above when implementing its functions. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0330] Figure 21 is a schematic diagram of the server structure provided in an embodiment of this application. The server 2100 can vary considerably due to different configurations or performance. It may include one or more processors (Central Processing Units, CPUs) 2101 and one or more memories 2102. The one or more memories 2102 store at least one line of program code, which is loaded and executed by the one or more processors 2101 to implement the image processing methods provided in the above-described method embodiments. Of course, the server 2100 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 2100 may also include other components for implementing device functions, which will not be elaborated here.
[0331] Figure 22 shows a structural block diagram of a terminal device 2200 provided in an exemplary embodiment of this application. The terminal device 2200 can be any electronic device product capable of human-computer interaction with a user through one or more methods such as a keyboard, touchpad, remote control, voice interaction, or handwriting device. Examples include PCs (Personal Computers), mobile phones, smartphones, PDAs (Personal Digital Assistants), wearable devices, PPCs (Pocket PCs), tablet computers, smart car systems, smart TVs, smart speakers, and smartwatches.
[0332] Typically, terminal device 2200 includes a processor 2201 and a memory 2202.
[0333] Processor 2201 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 2201 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 2201 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 2201 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 2201 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0334] The memory 2202 may include one or more computer-readable storage media, which may be non-transitory. The memory 2202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 2202 are used to store at least one instruction, which is executed by the processor 2201 to implement the image processing method provided in the method embodiments of this application.
[0335] In some embodiments, the terminal device 2200 may also optionally include a peripheral device interface 2203 and at least one peripheral device. The processor 2201, memory 2202, and peripheral device interface 2203 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 2203 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: radio frequency circuitry 2204, display screen 2205, camera assembly 2206, audio circuitry 2207, and power supply 2208.
[0336] Peripheral device interface 2203 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 2201 and memory 2202. In some embodiments, processor 2201, memory 2202 and peripheral device interface 2203 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 2201, memory 2202 and peripheral device interface 2203 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0337] The radio frequency (RF) circuit 2204 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 2204 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 2204 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 2204 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 2204 can communicate with other terminal devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 2204 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0338] Display screen 2205 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 2205 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 2201 for processing. In this case, display screen 2205 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 2205, disposed on the front panel of terminal device 2200; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal device 2200 or in a folded design; in still other embodiments, display screen 2205 may be a flexible display screen, disposed on a curved or folded surface of terminal device 2200. Furthermore, display screen 2205 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 2205 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0339] The camera assembly 2206 is used to acquire images or videos. Optionally, the camera assembly 2206 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal device 2200, and the rear-facing camera is located on the back of the terminal device 2200. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 2206 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0340] The audio circuit 2207 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 2201 for processing, or to the radio frequency circuit 2204 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal device 2200. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 2201 or the radio frequency circuit 2204 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 2207 may also include a headphone jack.
[0341] Power supply 2208 is used to power the various components in terminal device 2200. Power supply 2208 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 2208 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0342] In some embodiments, the terminal device 2200 further includes one or more sensors 2209. The one or more sensors 2209 include, but are not limited to: an acceleration sensor 2210, a gyroscope sensor 2211, a pressure sensor 2212, an optical sensor 2213, and a proximity sensor 2214.
[0343] Accelerometer 2210 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal device 2200. For example, accelerometer 2210 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 2201 can control display screen 2205 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 2210. Accelerometer 2210 can also be used for games or for acquiring user motion data.
[0344] The gyroscope sensor 2211 can detect the orientation and rotation angle of the terminal device 2200. The gyroscope sensor 2211 can work in conjunction with the accelerometer sensor 2210 to collect the user's 3D movements on the terminal device 2200. Based on the data collected by the gyroscope sensor 2211, the processor 2201 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0345] The pressure sensor 2212 can be disposed on the side bezel of the terminal device 2200 and / or on the lower layer of the display screen 2205. When the pressure sensor 2212 is disposed on the side bezel of the terminal device 2200, it can detect the user's grip signal on the terminal device 2200, and the processor 2201 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 2212. When the pressure sensor 2212 is disposed on the lower layer of the display screen 2205, the processor 2201 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 2205. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0346] Optical sensor 2213 is used to collect ambient light intensity. In one embodiment, processor 2201 can control the display brightness of display screen 2205 based on the ambient light intensity collected by optical sensor 2213. Specifically, when the ambient light intensity is high, the display brightness of display screen 2205 is increased; when the ambient light intensity is low, the display brightness of display screen 2205 is decreased. In another embodiment, processor 2201 can also dynamically adjust the shooting parameters of camera assembly 2206 based on the ambient light intensity collected by optical sensor 2213.
[0347] The proximity sensor 2214, also known as a distance sensor, is typically installed on the front panel of the terminal device 2200. The proximity sensor 2214 is used to detect the distance between the user and the front of the terminal device 2200. In one embodiment, when the proximity sensor 2214 detects that the distance between the user and the front of the terminal device 2200 is gradually decreasing, the processor 2201 controls the display screen 2205 to switch from a screen-on state to a screen-off state; when the proximity sensor 2214 detects that the distance between the user and the front of the terminal device 2200 is gradually increasing, the processor 2201 controls the display screen 2205 to switch from a screen-off state to a screen-on state.
[0348] Those skilled in the art will understand that the structure shown in FIG22 does not constitute a limitation on the terminal device 2200, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0349] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one piece of program code that is loaded and executed by a processor to enable a computer to implement any of the above-described image processing methods.
[0350] Optionally, the aforementioned computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0351] In an exemplary embodiment, a computer program or computer program product is also provided, which stores at least one computer instruction that is loaded and executed by a processor to enable the computer to implement any of the above-described image processing methods.
[0352] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the original images involved in this application were obtained with full authorization.
[0353] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0354] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. An image processing method, characterized in that, The method includes: Obtain pixel information of multiple pixels in the original image, whereby the pixel information of any one pixel is used to describe the brightness, chroma, and saturation of that pixel. Determine the depth parameters of each pixel, whereby the depth parameters of any pixel are used to indicate the distance between that pixel and the image acquisition device that acquired the original image. Image information is obtained based on the pixel information and depth parameters of the multiple pixels; The image information is sent to the receiving end, and the image information is used by the receiving end to generate the original image and the reference image corresponding to the original image. A target image with a three-dimensional effect is generated based on the original image and the reference image. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end.
2. The method according to claim 1, characterized in that, When the depth information of each pixel is within the reference interval, the depth parameter of each pixel is the depth information of each pixel. The reference interval is determined based on the number of bits of the depth channel used in the reference sampling method. The depth information of any pixel is used to indicate the actual distance between any pixel and the image acquisition device.
3. The method according to claim 1, characterized in that, In the case where there are pixels whose depth information is outside the reference interval, the depth parameter of each pixel is the depth order of each pixel. The depth order of any pixel is located within the reference interval, which is determined based on the number of bits of the depth channel used in the reference sampling method. The depth information of any pixel is used to indicate the actual distance between any pixel and the image acquisition device.
4. The method according to claim 3, characterized in that, The method further includes: The depth level of each pixel is determined based on the number of bits in the depth channel used in the reference sampling method, the maximum depth information in the depth information of each pixel, and the depth information of each pixel.
5. The method according to any one of claims 1 to 4, characterized in that, The image information includes sampling information of the original image; the acquisition of image information based on pixel information and depth parameters of the multiple pixels includes: According to the reference sampling method, the pixel information and depth parameters of each pixel are sampled to obtain the sampling information of the original image. The sampling information of the original image is used by the receiving end to determine the pixel information and depth parameters of each pixel included in the original image.
6. The method according to claim 5, characterized in that, The method further includes: Supplementary enhancement information is sent to the receiving end, the supplementary enhancement information indicating the reference sampling method, the reference sampling method being used by the receiving end to determine the pixel information and depth parameters of each pixel.
7. The method according to any one of claims 1 to 4, characterized in that, Sending the image information to the receiving end includes: The image information is encoded to obtain an encoded packet corresponding to the original image, wherein the encoded packet corresponding to the original image includes the image information; Encapsulate the encoded packet corresponding to the original image to obtain the encapsulated packet corresponding to the original image; The encapsulation packet corresponding to the original image is sent to the receiving end.
8. The method according to claim 7, characterized in that, The depth parameter of each pixel is the depth order of each pixel. Encoding the image information to obtain the encoded packet corresponding to the original image includes: The image information and the depth step size of the original image are encoded to obtain the encoding packet corresponding to the original image. The encoding packet corresponding to the original image also includes the depth step size of the original image. The depth step size of the original image is used to determine the depth information of each pixel in the original image.
9. The method according to claim 8, characterized in that, The method further includes: The depth step size of the original image is determined based on the maximum depth information in the depth information of each pixel in the original image and the number of bits in the depth channel used in the reference sampling method.
10. The method according to claim 8, characterized in that, The method further includes: Based on the depth information of each pixel in the original image, determine the depth standard deviation of the original image; The depth step size of the original image is determined based on the depth standard deviation of the original image and the number of bits in the depth channel used in the reference sampling method.
11. An image processing method, characterized in that, The method includes: The image information sent by the receiving end is obtained based on the pixel information and depth parameters of multiple pixels in the original image. The pixel information of any pixel is used to describe the brightness, chroma and saturation of the pixel, and the depth parameter of any pixel is used to indicate the distance between the pixel and the image acquisition device that acquired the original image. Based on the image information, determine the pixel information and depth parameters of each pixel; Based on the pixel information and depth parameters of each pixel, the original image and the reference image corresponding to the original image are generated. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end. A target image with a three-dimensional effect is generated based on the original image and the reference image.
12. The method according to claim 11, characterized in that, The image information sent by the receiving and sending end includes: Receive the encapsulation packet corresponding to the original image sent by the sending end; The encapsulation packet corresponding to the original image is unpacked to obtain the encoded packet corresponding to the original image; The image information is obtained by decoding the encoded packet corresponding to the original image.
13. The method according to claim 12, characterized in that, The depth parameter of each pixel is the depth order of each pixel, and the encoded packet also includes the depth step size of the original image; the method further includes: Decode the encoded packet corresponding to the original image to obtain the depth step size of the original image; The step of generating the original image and the corresponding reference image based on the pixel information and depth parameters of each pixel includes: The depth information of each pixel is determined based on the depth step size of the original image and the depth order of each pixel. Based on the pixel information and depth information of each pixel, the original image and the corresponding reference image are generated.
14. The method according to claim 13, characterized in that, The step of generating the original image and a reference image corresponding to the original image based on the pixel information and depth information of each pixel includes: Based on the pixel information and depth information of each pixel, the original image and a first intermediate image corresponding to the original image are generated, wherein the size of the first intermediate image is larger than the size of the original image. The first intermediate image is cropped to obtain a second intermediate image, the size of which is the same as the size of the original image; If the second intermediate image includes blank areas, the blank areas are filled to obtain a third intermediate image; Based on the third intermediate image, a reference image corresponding to the original image is determined.
15. The method according to claim 14, characterized in that, The step of determining the reference image corresponding to the original image based on the third intermediate image includes: The third intermediate image is used as the reference image corresponding to the original image; Alternatively, the hollow areas in the third intermediate image can be repaired, and the repaired image can be used as the reference image corresponding to the original image.
16. The method according to claim 14, characterized in that, Based on the pixel information and depth information of each pixel, a first intermediate image corresponding to the original image is generated, including: Determine the reflection vector corresponding to the first intermediate image, which is determined based on the position of the first intermediate image relative to the original image; Based on the position information and depth information of each pixel and the reflection vector corresponding to the first intermediate image, the offset position information of each pixel is determined; Based on the pixel information, depth information, and offset position information of each pixel, a first intermediate image corresponding to the original image is generated.
17. The method according to claim 16, characterized in that, The reflection vector corresponding to the first intermediate image is a two-dimensional vector, and the position information of any pixel includes the horizontal and vertical coordinates of any pixel. The step of determining the offset position information of each pixel based on the position information, depth information, and reflection vector corresponding to the first intermediate image includes: For any pixel among the pixels, the offset horizontal coordinate of any pixel is determined based on the horizontal coordinate of any pixel, the depth information of any pixel, and the value of the first dimension in the reflection vector corresponding to the first intermediate image. The offset ordinate of any pixel is determined based on the ordinate of any pixel, the depth information of any pixel, and the value of the second dimension in the reflection vector corresponding to the first intermediate image. The position information of any pixel after offset is determined based on the horizontal coordinate and the vertical coordinate of any pixel after offset.
18. An image processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire pixel information of multiple pixels in the original image, and the pixel information of any pixel is used to describe the brightness, chroma and saturation of the pixel. The determination module is used to determine the depth parameters of each pixel, and the depth parameters of any pixel are used to indicate the distance between the pixel and the image acquisition device that acquired the original image. The acquisition module is also used to acquire image information based on the pixel information and depth parameters of the plurality of pixels; The sending module is used to send the image information to the receiving end. The image information is used by the receiving end to generate the original image and the reference image corresponding to the original image, and to generate a target image with a three-dimensional effect based on the original image and the reference image. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end.
19. An image processing apparatus, characterized in that, The device includes: The receiving module is used to receive image information sent by the sending end. The image information is obtained based on the pixel information and depth parameters of multiple pixels included in the original image. The pixel information of any pixel is used to describe the brightness, chroma and saturation of the pixel. The depth parameter of any pixel is used to indicate the distance between the pixel and the image acquisition device that acquired the original image. The determining module is used to determine the pixel information and depth parameters of each pixel point based on the image information; The generation module is used to generate the original image and the reference image corresponding to the original image based on the pixel information and depth parameters of each pixel. The number of reference images is determined based on the number of viewpoints supported by the display screen of the receiving end. The generation module is further configured to generate a target image with a three-dimensional effect based on the original image and the reference image.
20. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to enable the computer device to implement the image processing method as described in any one of claims 1 to 17.
21. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to enable the computer to implement the image processing method as described in any one of claims 1 to 17.
22. A computer program product, characterized in that, The computer program product stores at least one computer instruction, which is loaded and executed by a processor to enable the computer to implement the image processing method as described in any one of claims 1 to 17.