A perceptual weighting-based point cloud rate-distortion encoding method and related device

CN117376568BActive Publication Date: 2026-09-25SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310183247.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-22
Publication Date
2026-09-25
Estimated Expiration
2043-02-22

AI Technical Summary

Technical Problem

然而,现有的V-PCC模型均未考虑动态点云的视觉感知质量,无法利用动态点云的感知冗余,进而影响编码效率

Benefits of technology

[0036]有益效果:与现有技术相比,本申请提供了一种基于感知加权的点云率失真编码方法及相关装置,方法包括获取待编码点云对应的投影图像以及所述投影模块对应的若干参考投影图像;获取投影图像及各参考投影图像的各编码块的梯度权重及结构相似性失真度,并基于各编码块的梯度权重及结构相似性失真度计算投影图像的感知失真度;根据投影图像的感知失真度确定所述待编码点云对应的率失真成本,并基于所述率失真成本编码所述投影图像。本申请通过投影生成二维的投影图像,然后计算以编码块为单位确定每个投影图像对应的感知失真度,并基于感知损失度进行编码决策,提高了点云编码与视觉感知质量的匹配性,从而提高了编码效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117376568B_ABST
    Figure CN117376568B_ABST
Patent Text Reader

Abstract

The application discloses a point cloud rate-distortion coding method based on perception weighting and a related device. The method comprises the following steps: acquiring a projection image corresponding to a point cloud to be coded and a plurality of reference projection images corresponding to a projection module; acquiring gradient weights and structural similarity distortion degrees of each coding block of the projection image and each reference projection image, and calculating a perception distortion degree of the projection image based on the gradient weights and the structural similarity distortion degrees of each coding block; determining a rate-distortion cost corresponding to the point cloud to be coded according to the perception distortion degree of the projection image, and coding the projection image based on the rate-distortion cost. The application generates a two-dimensional projection image through projection, then calculates a perception distortion degree corresponding to each projection image in a coding block unit, and makes a coding decision based on the perception distortion degree, thereby improving the matching of point cloud coding and visual perception quality, and improving the coding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a point cloud rate-distortion coding method and related apparatus based on perceptual weighting. Background Technology

[0002] In recent years, three-dimensional technologies such as Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) have become popular in many applications due to their ability to provide users with unique six degrees of freedom (6DoF) interaction, realistic and immersive 3D visual experiences, including 3D movie viewing, heritage preservation, navigation, immersive telephony, and remote surgery. Among these, Dynamic Point Cloud (DPC) has become one of the mainstream forms of expression in emerging immersive VR, AR, and MR media due to its realistic representation capabilities.

[0003] Dynamic point clouds (DPCs) represent 3D scenes using a large number of unstructured, high-dimensional points. A DPC is a series of time-continuous point clouds that not only reflects motion and temporal changes in light, but each point also includes geometric components for identifying its position in 3D space, as well as photometric information reflecting light and object properties, such as RGB color, reflection, and transparency. However, due to the large number of high-dimensional points, dynamic point clouds generate a massive amount of data, requiring enormous storage space and network bandwidth for transmission.

[0004] To address this issue, dynamic point clouds need to be compressed before storage and transmission to reduce the required storage space and bandwidth. Currently, the most commonly used compression technique for dynamic point clouds is Video Point Cloud Compression (V-PCC). V-PCC measures distortion by the signal difference between the original and distorted point clouds, such as point-to-point (D1), point-to-area (D2), and point-to-mesh (P2mesh). However, existing V-PCC models do not consider the visual perceptual quality of dynamic point clouds and cannot utilize their perceptual redundancy, thus affecting coding efficiency.

[0005] Therefore, the existing technology still needs to be improved and enhanced. Summary of the Invention

[0006] The technical problem to be solved by this application is to provide a point cloud rate-distortion coding method and related apparatus based on perceptual weighting, which addresses the shortcomings of the existing technology.

[0007] To address the aforementioned technical problems, a first aspect of this application provides a point cloud rate-distortion coding method based on perceptual weighting, the method comprising:

[0008] Obtain the projection image corresponding to the point cloud to be encoded and several reference projection images corresponding to the projection module, wherein the projection image and several reference projection images each include a geometric projection image and a texture projection image;

[0009] The gradient weights and structural similarity distortion of each coded block of the projected image and each reference projected image are obtained, and the perceptual distortion of the projected image is calculated based on the gradient weights and structural similarity distortion of each coded block.

[0010] The rate-distortion cost corresponding to the point cloud to be encoded is determined based on the perceptual distortion of the projected image, and the projected image is encoded based on the rate-distortion cost.

[0011] The perceptually weighted point cloud rate-distortion coding method wherein the reference projection image is obtained by downsampling the projection image, and the projection image and several reference projection images have different image scales.

[0012] The perceptually weighted point cloud rate-distortion coding method, wherein the process of obtaining the gradient weights specifically includes:

[0013] Calculate the gradient value of each pixel in the coding block, and calculate the gradient weight of the coding block based on the gradient value of each pixel.

[0014] The perceptually weighted point cloud rate-distortion coding method, wherein determining the rate-distortion cost corresponding to the point cloud to be encoded based on the perceptual distortion of the projected image specifically includes:

[0015] The structural similarity distortion in the perceptual distortion is converted into mean square error distortion to obtain the perceptual coefficients corresponding to the projected image;

[0016] The perceptual Lagrange multiplier corresponding to the projected image is determined based on the perceptual coefficient;

[0017] The rate-distortion cost corresponding to the point cloud to be encoded is determined based on the perceived distortion and the perceived Lagrange multiplier.

[0018] The perceptually weighted point cloud rate-distortion coding method, wherein the correspondence between the structural similarity distortion and the mean square error distortion is as follows:

[0019]

[0020]

[0021] in, This represents the structural similarity distortion of the coded block i of the k-th target projection image. Φ represents the mean square error distortion of the coded block i of the k-th target projection image.k,i blk represents the perceptual coefficient of the coded block i of the k-th target projection image. n ρ represents the nth sub-image block of coded block i. n E represents the linear model parameters of the nth sub-image block of coded block i, where N represents the number of sub-image blocks. i The pixel content weight is represented by the k-th target projection image, which is a projection image in the projection image set formed by the projection image and several reference projection images.

[0022] The perceptually weighted point cloud rate-distortion coding method, wherein when the perceptual Lagrange multiplier is used for coding, determining the perceptual Lagrange multiplier corresponding to the projected image based on the perceptual coefficients specifically includes:

[0023] The first target perceptual coefficient is calculated based on the perceptual coefficient of each coded block of the projected image, and the ratio of the first target perceptual coefficient to the perceptual coefficient of the coded block is calculated to obtain the first perceptual coefficient ratio.

[0024] Calculate the product of the first perceptual coefficient ratio and the mean square error Lagrange multiplier of the projected image to obtain the perceptual Lagrange multiplier corresponding to the coding block.

[0025] The perceptually weighted point cloud rate-distortion coding method, wherein when the perceptual Lagrange multiplier is used for pattern decision, determining the perceptual Lagrange multiplier corresponding to the projected image based on the perceptual coefficients specifically includes:

[0026] The second target perceptual coefficient is calculated based on the perceptual coefficient of each coded block of the projected image, and the ratio of the second target perceptual coefficient to the square root of the perceptual coefficient of the coded block is calculated to obtain the second perceptual coefficient ratio.

[0027] The product of the second perceptual coefficient ratio and the square root of the mean square error Lagrange multiplier of the projected image is calculated to obtain the perceptual Lagrange multiplier corresponding to the coding block.

[0028] A second aspect of this application provides a point cloud rate-distortion coding system based on perceptual weighting, the system comprising:

[0029] The acquisition module is used to acquire the projection image corresponding to the point cloud to be encoded and a plurality of reference projection images corresponding to the projection module, wherein the projection image and the plurality of reference projection images each include a geometric projection image and a texture projection image.

[0030] The calculation module is used to obtain the gradient weights and structural similarity distortion of each coded block of the projected image and each reference projected image, and to calculate the perceptual distortion of the projected image based on the gradient weights and structural similarity distortion of each coded block.

[0031] The encoding module is used to determine the rate-distortion cost corresponding to the point cloud to be encoded based on the perceptual distortion of the projected image, and to encode the projected image based on the rate-distortion cost.

[0032] A third aspect of this application provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps in the perceptually weighted point cloud rate-distortion coding method described above.

[0033] A fourth aspect of this application provides a terminal device, which includes: a processor, a memory, and a communication bus; the memory stores a computer-readable program that can be executed by the processor;

[0034] The communication bus enables communication between the processor and the memory;

[0035] When the processor executes the computer-readable program, it implements the steps in any of the above-described perceptually weighted point cloud rate-distortion coding methods.

[0036] Beneficial Effects: Compared with existing technologies, this application provides a point cloud rate-distortion coding method and related apparatus based on perceptual weighting. The method includes acquiring a projection image corresponding to the point cloud to be encoded and several reference projection images corresponding to the projection module; acquiring the gradient weights and structural similarity distortion of each coding block of the projection image and each reference projection image, and calculating the perceptual distortion of the projection image based on the gradient weights and structural similarity distortion of each coding block; determining the rate-distortion cost corresponding to the point cloud to be encoded based on the perceptual distortion of the projection image, and encoding the projection image based on the rate-distortion cost. This application generates a two-dimensional projection image through projection, then calculates the perceptual distortion corresponding to each projection image on a per-coding-block basis, and makes coding decisions based on the perceptual loss, thereby improving the matching between point cloud coding and visual perception quality and thus improving coding efficiency. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 A flowchart of the point cloud rate-distortion coding method based on perceptual weighting provided in this application.

[0039] Figure 2The flowchart illustrates the principle of the point cloud rate-distortion coding method based on perceptual weighting provided in this application.

[0040] Figure 3 This is a scatter plot showing the correlation between the texture projection maps of the approximate model and the original model.

[0041] Figure 4 This is a scatter plot showing the correlation between the approximate model and the original model on their geometric projections.

[0042] Figure 5 The structural principle diagram of the point cloud rate-distortion coding method based on perceptual weighting provided in this application is shown.

[0043] Figure 6 A schematic diagram of the terminal device provided in this application. Detailed Implementation

[0044] This application provides a point cloud rate-distortion coding method and related apparatus based on perceptual weighting. To make the purpose, technical solution, and effects of this application clearer and more explicit, the following detailed description is provided with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.

[0045] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0046] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0047] It should be understood that the sequence number and size of each step in this embodiment do not imply the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application embodiment.

[0048] Research has shown that in recent years, 3D technologies such as Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) have become increasingly popular in many applications due to their ability to provide users with unique six degrees of freedom (6DoF) interaction, realistic and immersive 3D visual experiences, including 3D movie viewing, heritage preservation, navigation, immersive telephony, and remote surgery. Among these, Dynamic Point Cloud (DPC), with its realistic representation capabilities, has become one of the mainstream forms of expression in emerging immersive VR, AR, and MR media.

[0049] Dynamic point clouds (DPCs) represent 3D scenes using a large number of unstructured, high-dimensional points. A DPC is a series of time-continuous point clouds that not only reflects motion and temporal changes in light, but each point also includes geometric components for identifying its position in 3D space, as well as photometric information reflecting light and object properties, such as RGB color, reflection, and transparency. However, due to the large number of high-dimensional points, dynamic point clouds generate a massive amount of data, requiring enormous storage space and network bandwidth for transmission.

[0050] To address this issue, dynamic point clouds need to be compressed before storage and transmission to reduce the required storage space and bandwidth. Currently, the most commonly used compression technique for dynamic point clouds is Video Point Cloud Compression (V-PCC). V-PCC measures distortion by the signal difference between the original and distorted point clouds, such as point-to-point (D1), point-to-area (D2), and point-to-mesh (P2mesh). However, existing V-PCC models do not consider the visual perceptual quality of dynamic point clouds and cannot utilize their perceptual redundancy, thus affecting coding efficiency.

[0051] Therefore, combining visual perception with coding has always been a hot topic in the fields of video and image processing. In existing research, Jiang et al. proposed a new perceptual coding scheme conforming to the H.265 / HEVC standard by using spatiotemporal saliency. Wu et al. proposed a rate-distortion (RD) model based on perceptual weighted average squared error (PWMSE) and derived the Lagrange multiplier for the rate-distortion optimization (RDO) process based on equivalent distortion. Based on WS-PSNR, Li et al. optimized the RDO task using a 360-degree video evaluation metric in the spherical domain. Li et al. proposed an RDO method for spherical domain videos to address the problem of inaccurate calculation of planar distortion in two-dimensional images. The basic idea of ​​these methods is to use the HVS model or its approximations as the distortion in the coding module and select the optimal coding mode or parameters with the minimization of perceptual distortion as a constraint. However, these methods are designed for traditional 2D or 360-degree video coding and cannot be used for dynamic point clouds.

[0052] To utilize visual redundancy in point clouds, Li et al. employed an occupancy map-based RDO method. This method designed an occupancy map-guided video compression framework, leveraging the occupancy map to improve DPC compression performance. Since current RDO and geometric distortion methods are inconsistent with the evaluation standards D1 and D2, Xiong et al. proposed an EPM-based RDO method. This method first describes the relationship between existing distortion models and geometric quality measurements, then estimates the D1 and D2 normals by estimating the CU normals, thereby modifying the 3D geometric distance. However, the reconstructed point cloud quality in these rate-distortion coding schemes is still measured using D1 and D2, failing to accurately reflect human perception. Therefore, existing point cloud coding schemes do not fully consider the perceptual characteristics of point clouds and thus do not utilize their perceptual redundancy.

[0053] To address the aforementioned issues, this application embodiment acquires a projected image corresponding to the point cloud to be encoded and several reference projected images corresponding to the projection module; acquires the gradient weights and structural similarity distortion of each coding block in the projected image and each reference projected image, and calculates the perceptual distortion of the projected image based on the gradient weights and structural similarity distortion of each coding block; determines the rate-distortion cost corresponding to the point cloud to be encoded based on the perceptual distortion of the projected image, and encodes the projected image based on the rate-distortion cost. This application generates a two-dimensional projected image through projection, then calculates the perceptual distortion corresponding to each projected image on a per-coding-block basis, and makes encoding decisions based on the perceptual loss, thereby improving the matching between point cloud encoding and visual perception quality and thus improving encoding efficiency.

[0054] The application content will be further explained below with reference to the accompanying drawings and the description of the embodiments.

[0055] This embodiment provides a point cloud rate-distortion coding method based on perceptual weighting, such as... Figure 1 and Figure 2 As shown, the method includes:

[0056] S10. Obtain the projection image corresponding to the point cloud to be encoded and several reference projection images corresponding to the projection module.

[0057] Specifically, the projected image and the plurality of reference projected images each include a geometric projection map and a texture projection map. The projected image is formed by patching and packing the point cloud to be encoded. Both the geometric projection map and the texture projection map are two-dimensional images, and they have the same image scale. Each of the plurality of reference projected images is an image group, which includes a reference geometric projection map and a reference texture projection map. The image scales of the geometric projection maps in each reference projected image are different from each other and are not equal to the image scale of the geometric projection map; similarly, the image scales of the texture projection maps in each reference projected image are different from each other and are not equal to the image scale of the texture projection map.

[0058] The reference projection image is obtained by downsampling the projection image. Specifically, this means the reference geometric projection image is obtained by downsampling the geometric projection image, and the reference texture projection image is obtained by downsampling the geometric projection image. In one implementation, several reference projection images can be downsampled using a Gaussian pyramid. This involves inputting the geometric and texture projection images (including the projection image) into a Gaussian pyramid, obtaining the output terms of each layer of the pyramid, and thus obtaining several reference projection images. It can be understood that through each layer of the Gaussian pyramid, the resolution of the output term of that layer is reduced to half of the input term. For example, let k represent different scales: k=1 represents the image scale corresponding to the projection image, k=2 indicates that the projection image has been downsampled once, and the image scale of the resulting reference projection image is half of the image scale of the projection image, and so on, to obtain several projection images.

[0059] Of course, in practical applications, after generating the projected image through projection, the acquisition of several reference projected images can also be achieved in other ways. For example, several projected images can be sorted according to their image scale from largest to smallest to form a first reference projected image, a second reference projected image, ..., an Nth reference projected image. The first reference projected image is obtained by downsampling the projected image by a factor of 2, the second reference projected image is obtained by downsampling the projected image by a factor of 4, and so on, until the Nth reference projected image is obtained. Furthermore, the correspondence between the image scales of the Nth reference projected images can also be different. For example, the image scale corresponding to the projected image can be three times that of the first reference projected image, and the image scale corresponding to the first reference projected image can be three times that of the second reference projected image, and so on.

[0060] S20. Obtain the gradient weights and structural similarity distortion of each coded block of the projected image and each reference projected image, and calculate the perceptual distortion of the projected image based on the gradient weights and structural similarity distortion of each coded block.

[0061] Specifically, the gradient weights are determined based on the gradients of each pixel in the coding block, and the structural similarity distortion is used to reflect the distortion of the corresponding distortion block of the coding block. The distortion block is determined based on the geometric reconstruction map and texture reconstruction map formed by the reconstructed point cloud corresponding to the point cloud to be encoded. That is, when acquiring the projected image and several reference projected images, the reconstructed projected image of the reconstructed point cloud corresponding to the point cloud to be encoded is acquired simultaneously. The reconstructed projected image includes both a geometric reconstruction map and a texture reconstruction map. Then, a downsampling operation is performed on the reconstructed projected image to obtain several reference reconstructed projected images. The process of determining the reconstructed projected image and several reference reconstructed projected images is the same as that of determining the projected image and several reference projected images, so it will not be described in detail here. It is only noted that the projected image corresponds to the reconstructed projected image, and the several reference projected images correspond one-to-one with the several reference reconstructed projected images. Furthermore, the image scale of the reference projected image is the same as the image scale of its corresponding reference reconstructed projected image; that is, the image scale of the reference geometric projection map in the reference projected image is equal to the image scale of the reference geometric reconstruction map in its corresponding reference reconstructed projected image, and the image scale of the reference texture projection map in the reference projected image is equal to the image scale of the reference texture reconstruction map in its corresponding reference reconstructed projected image.

[0062] In one implementation, the process of obtaining the gradient weights specifically includes:

[0063] Calculate the gradient value of each pixel in the coding block, and calculate the gradient weight of the coding block based on the gradient value of each pixel.

[0064] Specifically, the gradient value of each pixel in the coded block includes the horizontal gradient and the vertical gradient, and the formulas for calculating the horizontal gradient and the vertical gradient are as follows:

[0065]

[0066]

[0067] Among them, F H I represents the gradient operator in the horizontal direction. i Let represent the i-th coded block, with a size of M×N, and j represent the j-th pixel in the coded block.

[0068] After obtaining the pixel gradient values ​​of each pixel, the gradient value of the coding block is calculated based on the pixel gradient values ​​of each pixel in the coding block. The formula for calculating the candidate gradient value of the coding block can be:

[0069]

[0070] in, This represents the candidate gradient value.

[0071] After obtaining the candidate gradient values ​​for each coding block, a normalization operation is performed on each candidate gradient value to obtain the gradient value for each coding block. The gradient value of a coding block is used to reflect the importance of that coding block relative to all coding blocks.

[0072] The formula for calculating the gradient value is as follows:

[0073]

[0074] Among them, W i This represents the gradient value of the coded block, and Normal(·) represents the normalization operation, which is calculated as follows:

[0075]

[0076] The C(·) operation constrains the calculated maximum and minimum values ​​to the range [0,1], i.e., values ​​less than 0 are set to 0, and values ​​greater than 1 are set to 1. x x represents i The mean, σ x x represents i The standard deviation.

[0077] Furthermore, the structural similarity distortion is the structural similarity distortion between the coded block and its corresponding reconstructed coded block. Having obtained the gradient weights and structural similarity distortion of each coded block, the perceptual distortion of the projected image can be determined based on these values. Specifically, the texture image block distortion of the i-th texture distortion image block in the k-th target projected image of the projected image set, formed by the projected image and several reference projected images, can be expressed as:

[0078]

[0079] Among them, W k,i This represents the gradient weight of the i-th coded block in the k-th image group. This represents the SSIM distortion of the i-th coded block relative to the i-th reconstructed coded block in the k-th image group.

[0080] Of course, it is worth noting that when calculating the perceptual distortion of the coded block, the geometric projection map and texture projection map in the projected image calculate their respective perceptual distortion, and the calculation process for both is the same. Therefore, this embodiment uses the projected image to illustrate the process of determining the perceptual distortion.

[0081] S30. Determine the rate-distortion cost corresponding to the point cloud to be encoded based on the perceptual distortion of the projected image, and encode the projected image based on the rate-distortion cost.

[0082] Specifically, the objective function for rate-distortion cost is:

[0083]

[0084] Where J represents the total rate-distortion (RD) cost, D represents the distortion difference between the point cloud to be encoded and the reconstructed point cloud corresponding to the unit to be encoded, and R represents the encoded bits. λ represents the Lagrange multiplier used to balance distortion and bit rate.

[0085] Based on this, the objective function for determining the rate-distortion cost corresponding to the point cloud to be encoded according to the perceptual distortion of the projected image can be expressed as:

[0086]

[0087] Among them, D PPCM Indicates the degree of perceived distortion.

[0088] Since the V-PCC standard uses HEVC to encode geometric video and texture video separately, the compression of geometric video and texture video is treated as two independent processes. Furthermore, the distortion of texture video depends only on the texture video coding parameters, and the distortion of geometric video depends only on the geometric video coding parameters. Therefore, the objective functions for the rate-distortion cost of texture projection images and the rate-distortion cost of geometric projection images are expressed as follows:

[0089]

[0090]

[0091] Among them, J G J represents the independent constraints of the texture projection image. T Represents the independent constraints of the geometrically projected image, R T For a geometric video encoder, R is a constant. G For texture video encoders.

[0092] Furthermore, since the encoding decision process for texture projection images is the same as that for geometric projection images, only the encoding decision process will be described here, without explaining the texture projection images and geometric projection images. The encoding decision process for texture projection images and geometric projection images can be adopted as follows.

[0093] In one implementation, determining the bitrate distortion cost corresponding to the point cloud to be encoded based on the perceptual distortion of the projected image specifically includes:

[0094] The structural similarity distortion in the perceptual distortion is converted into mean square error distortion to obtain the perceptual coefficients corresponding to the projected image;

[0095] The perceptual Lagrange multiplier corresponding to the projected image is determined based on the perceptual coefficient;

[0096] The rate-distortion cost corresponding to the point cloud to be encoded is determined based on the perceived distortion degree and the perceived Lagrange multiplier.

[0097] Specifically, under high-resolution quantization approximation, the encoding process will preserve as much luminance information as possible. Furthermore, for different video content, even at high QP settings, The Pearson correlation coefficient (PCC) between i and μ also exceeds 0.99. Based on this, i can be... th The SSIM value of the pixel is rewritten as:

[0098]

[0099]

[0100] in, E represents the pixel-level squared error. i The pixel content weights are represented by M, the number of coding blocks is represented by C2, and ε is the C2 constant in SSIM. t This represents the filter coefficient at position t in the current coding block. and This represents the variance at point it before and after encoding of the coded block.

[0101] Furthermore, to reduce computational complexity, the SSIM distortion of the coded block is set to the average of 16 sub-blocks, where n represents the number of sub-image blocks in the coded block, i.e.:

[0102]

[0103] The coding units in HEVC are known to have the following relationship.

[0104]

[0105] Where, ρ n These are linear model parameters related to the image content, where the result of the encoded co-occurrence block of the previous frame is used as the reference image content (except for the first frame), Q n It is the quantization step size of the sub-block, and is usually applied to all sub-blocks with the same value.

[0106] For the coded block, D MSE It can be represented as:

[0107]

[0108]

[0109]

[0110] in, The structural similarity distortion of the coded block i of the k-th projected image is represented by . Φ represents the mean square error distortion of the coded block i of the k-th projected image. k,i blk represents the perceptual coefficient of the coded block i of the k-th projected image. n ρ represents the nth sub-image block of coded block i. n E represents the linear model parameters of the nth sub-image block of coded block i, where N represents the number of sub-image blocks. i Represents the pixel content weight. The k-th projected image is a projected image in a set of projected images formed by the projected image and several reference projected images.

[0111] Based on this, the image MSE numerical loss caused by video compression can be approximated as the SSIM loss of the neighborhood image after appropriate transformation, i.e., D PPCM It can be represented as:

[0112]

[0113] Furthermore, after obtaining the perceptual coefficients, the perceptual Lagrange multiplier corresponding to the projected image is determined. The process for determining the perceptual Lagrange multiplier can be as follows:

[0114] When the distortion metric D based on PPCM is used PPCM When applying the RD objective function in V-PCC, it is necessary to adjust the Lagrange multipliers. and Adjustments are made to reduce perceptual distortion in channel φ (where a channel includes texture and geometry channels, denoted by φ here). and bit rate R φ To find the optimal trade-off between these factors. Rate-distortion cost based on MSE. The calculation for each coded block i is as follows:

[0115]

[0116] in, and λ MSE In HEVC, this refers to distortion, bit rate, and Lagrange multiplier based on MSE, and bit rate. and distortion The relationship between them can be modeled as

[0117]

[0118] in, It is the variance of the coding residual after inter-frame or intra-frame prediction, where α is a scaling constant. right Find the partial derivatives, and when the partial derivatives are zero, we can obtain:

[0119]

[0120] Solving the above equation yields the optimal solution. and for:

[0121]

[0122] The total bitrate of a video is the sum of the number of bits in all blocks, which can be calculated as follows:

[0123]

[0124] Where M is the number of coded blocks in one frame / video.

[0125] Similarly, the objective function based on rate coding distortion is applied to D. MSE Taking the partial derivatives, and when the partial derivatives are zero, we can obtain:

[0126]

[0127] Solving the above equation yields the following result:

[0128]

[0129] The total number of bits for the encoding channel φ is calculated as follows:

[0130]

[0131] Based on this, the sensing Lagrange parameters can be determined based on the total number of bits in the encoding channel φ.

[0132] In one implementation, when the perceptual Lagrange multiplier is used for encoding, determining the perceptual Lagrange multiplier corresponding to the projected image based on the perceptual coefficients specifically includes:

[0133] The first target perceptual coefficient is calculated based on the perceptual coefficient of each coded block of the projected image, and the ratio of the first target perceptual coefficient to the perceptual coefficient of the coded block is calculated to obtain the first perceptual coefficient ratio.

[0134] Calculate the product of the first perceptual coefficient ratio and the mean square error Lagrange multiplier of the projected image to obtain the perceptual Lagrange multiplier corresponding to the coding block.

[0135] Specifically, in the V-PCC model, attribute and geometric video are encoded using an MSE-based encoder. The total bitrate R of MSE-based V-PCC is... MSE Compared with the total bit rate R of V-PCC based on perceived distortion φ The same applies to different point cloud encoding channels φ, i.e., R. MSE =R φ Therefore, we can obtain λ. φ and λ MSE The relationship between them is:

[0136]

[0137] Therefore, based on formulas (1) and (2), the RD cost of the pattern decision is obtained as follows:

[0138]

[0139] Therefore, the perceived Lagrange coefficient is:

[0140]

[0141] In one implementation, when the perceptual Lagrange multiplier is used for pattern decision-making, determining the perceptual Lagrange multiplier corresponding to the projected image based on the perceptual coefficients specifically includes:

[0142] The second target perceptual coefficient is calculated based on the perceptual coefficient of each coded block of the projected image, and the ratio of the second target perceptual coefficient to the square root of the perceptual coefficient of the coded block is calculated to obtain the second perceptual coefficient ratio.

[0143] The product of the second perceptual coefficient ratio and the square root of the mean absolute error Lagrange multiplier of the projected image is calculated to obtain the perceptual Lagrange multiplier corresponding to the coding block.

[0144] In traditional MSE-based video encoders, SAD / MAD is used as the distortion term in rate-distortion optimization to avoid squaring operations and maintain low complexity. The RD cost is calculated as follows:

[0145]

[0146] in and These represent the distortion, bit rate, and Lagrange multiplier based on MAD in HEVC. The Lagrange multiplier for ME is also shown. Similarly, based on The relationship has

[0147] Therefore, the RD cost is:

[0148]

[0149] Therefore, the perceptual Lagrange multiplier based on perceptual distortion Updated to:

[0150]

[0151] To further illustrate the encoding process of this embodiment, the V-PCC encoding standard maps the geometric and texture features of the 3D point cloud into two video sequences, denoted as the geometric projection map sequence and the texture projection map sequence, respectively. Then, the video sequence of the dynamic point cloud is compressed using an existing HEVC / VVC or other video encoder. During the encoding process, some metadata, such as placeholder maps and auxiliary block information, is also generated to describe the two video sequences. The encoding process is the same as the existing process and will not be described in detail here.

[0152] In V-PCC, a point cloud projection patch is a collection of information, including the 3D bounding box of the point cloud, related geometric and texture information, and atlas information required for 3D reconstruction. The point cloud projection patch partitioning process projects each frame of the point cloud from 3D space onto a given 2D plane, minimizing the number of smoothly bounded patch blocks to reduce reconstruction error. The point cloud projection process can begin by calculating the normal vector of each point in the input point cloud, then clustering the entire point cloud according to the normal vectors and projecting it onto the six faces of a cube to form initial clusters. Next, based on the normal vectors and the indices of the nearest neighbors, the clustering index of each point is iteratively updated to achieve finer patch partitioning. Finally, connected component extraction is applied to extract connected patches and merge them into larger patch blocks, resulting in the final patch set. Thus, by packaging and recombining the point cloud projection patch set on a 2D plane, the geometric and texture information in 3D space is converted into 2D images, enabling the generation and filling of geometric and texture images to obtain geometric projection map sequences and texture projection map sequences.

[0153] The following explanation assumes that the geometric projection map sequence and texture projection map sequence have already been obtained. The encoding process of the geometric projection map sequence and texture projection map sequence includes:

[0154] a) Perform geometric projection analysis to obtain the perceptual coefficients of each coded block in the geometric projection map. In addition to the first frame reference frame, the code is obtained based on the coded co-occurring blocks of the current coded block and the previous frame. ρ in n parameter;

[0155] b) Determine the perception Lagrange multiplier based on the perception coefficient;

[0156] c) Encode the current geometric projection frame; if the geometric projection images corresponding to all video frames have been encoded, proceed to step d) for texture encoding; otherwise, go to step a) for encoding the next frame.

[0157] d) Perform texture projection map analysis to obtain the perceptual coefficients of each coded block in the geometric projection map. In addition to the first frame reference frame, the code is obtained based on the coded co-occurring blocks of the current coded block and the previous frame. ρ in n parameter;

[0158] e) Determine the perception Lagrange multiplier based on the perception coefficient;

[0159] f) Encode the current texture projection map, then proceed to step d) to encode the texture projection map of the next frame, until all texture projection maps have been encoded.

[0160] In summary, this embodiment provides a point cloud rate-distortion coding method based on perceptual weighting. The method includes acquiring a projected image corresponding to the point cloud to be encoded and several reference projected images corresponding to the projection module; acquiring the gradient weights and structural similarity distortion of each coding block in the projected image and each reference projected image, and calculating the perceptual distortion of the projected image based on the gradient weights and structural similarity distortion of each coding block; determining the rate-distortion cost corresponding to the point cloud to be encoded based on the perceptual distortion of the projected image, and encoding the projected image based on the rate-distortion cost. This application generates a two-dimensional projected image through projection, then calculates the perceptual distortion corresponding to each projected image on a block-by-block basis, and makes coding decisions based on the perceptual loss, improving the matching between point cloud coding and visual perception quality, thereby improving coding efficiency.

[0161] Furthermore, to illustrate the accuracy of this embodiment, statistical analysis was performed on the point clouds of Longdress and Andrew. Each point cloud was compressed with 17 quantization parameter (QP) pairs, ranging from 20 to 32 and 27 to 42 respectively. Each point cloud was projected onto a two-dimensional image and decomposed into 400 blocks, i.e., 400 × 17 sampling blocks. Figure 3 and Figure 4 As shown, R is used 2 To calculate and The approximate model accuracy between the two is expressed as the average R-value across the attribute and geometric channels. 2 The approximations are 0.9732 and 0.9556, respectively, and are accurate.

[0162] To verify the proposed coding efficiency, the coding method was implemented on the V-PCC reference software TMC2-10.0 and the corresponding HEVC reference software HM16.20-SCM8.8. The state-of-the-art V-PCC (TMC2-10.0 + HM16.20-SCM8.8) was used as the anchor method for comparison. Furthermore, two state-of-the-art V-PCC RDO methods, occupancy-map based RDO (denoted as "OC-RDO") and EPM-based RDO (denoted as "EPM-RDO"), were also used as benchmarks for comparison. Since the proposed coding method can be applied to intra-frame and inter-frame coding of V-PCC, V-PCC was configured as a random access coding scheme (RA) for geometry and texture coding. Five bitrate points from low (r1) to high (r5) conforming to the V-PCC Common Test Conditions (CTC) were selected for testing. For a fair comparison, GraphSIM (independent of the proposed PPCM) was used to measure the perceptual quality of the compressed point clouds of different coding schemes. Then, the Bjontegaard Delta Bit Rate (BDBR) was used to compare the RD performance of the proposed scheme with that of the benchmark scheme.

[0163] Table 1 shows the RD comparisons across the six tested DPCs, where the visual quality of the compressed video was measured using GraphSIM, R... T The total bit rate (Mbps) is represented. It can be seen that OC-RDO achieves a bit rate reduction of 0.34% to 17.05% compared to V-PCC, with an average of 9.61%. EPM-RDO achieves a BDBR saving of -4.08% to 5.61%, with an average of 1.17%. As for the proposed PWRDO, it achieves a BDBR reduction of 7.09% to 22.79%, an average reduction of 13.52% compared to V-PCC, significantly exceeding the comparison schemes of OC-RDO and EPM-RDO. To evaluate the perceptual coding efficiency of the proposed PWRDO, five different perceptual PCQA metrics, including D1, GraphSIM, MPED, SIAT_PCQA, and the proposed PPCM, were used to measure the visual quality of compressed DPC in different schemes. Then, the BDBR, which includes the total number of coded bits including geometric bits, attribute bits, and metadata bits, was calculated for each PCQA. A negative BDBR indicates bit savings, while a positive BDBR indicates a decrease in coding efficiency compared to the anchor point. Table 1 shows the RD comparison between the proposed PWRDO and the benchmark RDO scheme under five different perceptual PCQA metrics. These comparison results demonstrate that the proposed PWRDO consistently achieves high coding gains across all PCQA metrics, proving the effectiveness of the proposed PWRDO.

[0164] Table 1. Gains of the proposed method and two point cloud coding rate-distortion optimization methods on five evaluation metrics.

[0165]

[0166]

[0167] Based on the above-described perceptually weighted point cloud rate-distortion coding method, this embodiment provides a perceptually weighted point cloud rate-distortion coding system, such as... Figure 5 As shown, the system includes:

[0168] The acquisition module 100 is used to acquire the projection image corresponding to the point cloud to be encoded and a plurality of reference projection images corresponding to the projection module, wherein the projection image and the plurality of reference projection images each include a geometric projection image and a texture projection image.

[0169] The calculation module 200 is used to obtain the gradient weights and structural similarity distortion of each coding block of the projected image and each reference projected image, and to calculate the perceptual distortion of the projected image based on the gradient weights and structural similarity distortion of each coding block.

[0170] The encoding module 300 is used to determine the rate-distortion cost corresponding to the point cloud to be encoded based on the perceptual distortion of the projected image, and to encode the projected image based on the rate-distortion cost.

[0171] Based on the above-described perceptually weighted point cloud rate distortion coding method, this embodiment provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps in the perceptually weighted point cloud rate distortion coding method described in the above embodiment.

[0172] Based on the aforementioned perceptually weighted point cloud rate-distortion coding method, this application also provides a terminal device, such as... Figure 6 As shown, it includes at least one processor 20; a display screen 21; and a memory 22, and may also include a communications interface 23 and a bus 24. The processor 20, display screen 21, memory 22, and communications interface 23 can communicate with each other via the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setup mode. The communications interface 23 can transmit information. The processor 20 can invoke logical instructions in the memory 22 to execute the methods described in the above embodiments.

[0173] Furthermore, the logical instructions in the aforementioned memory 22 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0174] The memory 22, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of this disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, thereby implementing the methods in the above embodiments.

[0175] The memory 22 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 22 may include high-speed random access memory (RAM) and non-volatile memory. Examples include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, as well as transient storage media.

[0176] Furthermore, the specific process of loading and executing multiple instruction processors in the aforementioned storage medium and terminal device has been described in detail in the above method, and will not be repeated here.

[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A point cloud rate-distortion coding method based on perceptual weighting, characterized in that, The method includes: Obtain the projection image corresponding to the point cloud to be encoded and several reference projection images corresponding to the projection image, wherein the projection image and several reference projection images each include a geometric projection image and a texture projection image; The gradient weights and structural similarity distortion of each coded block in the projected image and each reference projected image are obtained. The perceptual distortion of the projected image is calculated based on the gradient weights and structural similarity distortion of each coded block. The structural similarity distortion is used to reflect the distortion of the corresponding distortion block. The distortion block is determined by the geometric reconstruction map and texture reconstruction map formed based on the reconstructed point cloud corresponding to the point cloud to be encoded. The perceptual distortion is obtained by weighting the structural similarity distortion of each coded block by using the gradient weights of each coded block as weight coefficients. The rate-distortion cost corresponding to the point cloud to be encoded is determined based on the perceptual distortion of the projected image, and the projected image is encoded based on the rate-distortion cost.

2. The point cloud rate-distortion coding method based on perceptual weighting according to claim 1, characterized in that, The reference projection image is obtained by downsampling the projection image, and the projection image and several reference projection images have different image scales.

3. The point cloud rate-distortion coding method based on perceptual weighting according to claim 1, characterized in that, The process of obtaining the gradient weights specifically includes: Calculate the gradient value of each pixel in the coding block, and calculate the gradient weight of the coding block based on the gradient value of each pixel.

4. The point cloud rate-distortion coding method based on perceptual weighting according to claim 1, characterized in that, The step of determining the rate-distortion cost corresponding to the point cloud to be encoded based on the perceptual distortion of the projected image specifically includes: The structural similarity distortion in the perceptual distortion is converted into mean square error distortion to obtain the perceptual coefficients corresponding to the projected image; The perceptual Lagrange multiplier corresponding to the projected image is determined based on the perceptual coefficient; The rate-distortion cost corresponding to the point cloud to be encoded is determined based on the perceived distortion and the perceived Lagrange multiplier.

5. The point cloud rate-distortion coding method based on perceptual weighting according to claim 4, characterized in that, The correspondence between the structural similarity distortion and the mean square error distortion is as follows: in, Indicates the first Encoded blocks of a target projection image Structural similarity distortion Indicates the first Encoded blocks of a target projection image Mean squared error distortion Indicates the first Encoded blocks of a target projection image Perception coefficient, Indicates a coded block The Sub-image blocks, Indicates a coded block The Linear model parameters for each sub-image patch Indicates the number of sub-image patches. Represents the pixel content weight, the first A target projection image is a projection image in a set of projection images formed by the projection image and several reference projection images.

6. The point cloud rate-distortion coding method based on perceptual weighting according to claim 4, characterized in that, When the perceptual Lagrange multiplier is used for encoding, determining the perceptual Lagrange multiplier corresponding to the projected image based on the perceptual coefficients specifically includes: The first target perceptual coefficient is calculated based on the perceptual coefficient of each coded block of the projected image, and the ratio of the first target perceptual coefficient to the perceptual coefficient of the coded block is calculated to obtain the first perceptual coefficient ratio. Calculate the product of the first perceptual coefficient ratio and the mean square error Lagrange multiplier of the projected image to obtain the perceptual Lagrange multiplier corresponding to the coding block.

7. The point cloud rate-distortion coding method based on perceptual weighting according to claim 4, characterized in that, When the perceptual Lagrange multiplier is used for pattern decision-making, determining the perceptual Lagrange multiplier corresponding to the projected image based on the perceptual coefficients specifically includes: The second target perceptual coefficient is calculated based on the perceptual coefficient of each coded block of the projected image, and the ratio of the second target perceptual coefficient to the square root of the perceptual coefficient of the coded block is calculated to obtain the second perceptual coefficient ratio. The product of the second perceptual coefficient ratio and the square root of the mean square error Lagrange multiplier of the projected image is calculated to obtain the perceptual Lagrange multiplier corresponding to the coding block.

8. A point cloud rate-distortion coding system based on perceptual weighting, characterized in that, The system includes: The acquisition module is used to acquire the projection image corresponding to the point cloud to be encoded and several reference projection images corresponding to the projection image, wherein the projection image and several reference projection images each include a geometric projection image and a texture projection image. The calculation module is used to obtain the gradient weights and structural similarity distortion of each coded block of the projected image and each reference projected image, and to calculate the perceptual distortion of the projected image based on the gradient weights and structural similarity distortion of each coded block. The structural similarity distortion is used to reflect the distortion of the corresponding distortion block of the coded block. The distortion block is determined by the geometric reconstruction map and texture reconstruction map formed based on the reconstructed point cloud corresponding to the point cloud to be encoded. The perceptual distortion is obtained by weighting the structural similarity distortion of each coded block by using the gradient weights of each coded block as weight coefficients. The encoding module is used to determine the rate-distortion cost corresponding to the point cloud to be encoded based on the perceptual distortion of the projected image, and to encode the projected image based on the rate-distortion cost.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps in the perceptually weighted point cloud rate-distortion coding method as described in any one of claims 1-7.

10. A terminal device, characterized in that, include: Processor, memory, and communication bus; the memory stores a computer-readable program that can be executed by the processor; The communication bus enables communication between the processor and the memory; When the processor executes the computer-readable program, it implements the steps in the perceptually weighted point cloud rate-distortion coding method as described in any one of claims 1-7.