Point cloud geometric information motion estimation method, system, device and storage medium
By using a multi-scale binary prediction residual calculation method in point cloud inter-frame prediction, the problem of inaccurate calculation of motion estimation residual bits in the prior art is solved, the inter-frame prediction accuracy and coding efficiency are improved, and the bit rate saving is achieved.
Patent Information
- Application Number
- CN202110811925.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-07-19
AI Technical Summary
In the prior art, in point cloud inter-frame prediction, there is a problem of inaccurate calculation of residual bit counts during motion estimation, resulting in low inter-frame prediction accuracy and no saving in bit rate.
Using the G-PCC encoding framework based on octree, a linear relationship between the residual bits and the multi-scale binary prediction residual is established through multi-scale binary prediction residual calculation, thereby minimizing the overall number of bits including residual bits and motion vector bits to obtain the optimal motion vector.
The accuracy of inter prediction and point cloud encoding efficiency are improved, the bit rate is saved, and more accurate residual bit prediction is achieved.
Smart Images

Figure CN115643415B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of point cloud compression coding, and in particular to a point cloud geometric information motion estimation method, system, device and storage medium. Background Art
[0002] Point cloud is a collection of points in space, which can be used to represent the surface of a 3D object or 3D scene. Each point in the point cloud contains geometric information (spatial position x, y, z of each point) and attribute information (color or reflectance of each point, etc.). Typical point cloud application scenarios include: autonomous driving, urban planning, archaeological heritage protection, medical imaging, surveying and mapping, etc.
[0003] Point cloud compression is a key technology that enables typical point cloud applications. It usually includes encoding both point cloud geometry information and attribute information. In most point cloud compression frameworks, geometry information and attribute information are encoded using different encoding methods, and geometry information is encoded first, and then attribute information is encoded based on geometry information. The most typical point cloud geometry information encoding algorithm uses an octree to divide the space. According to different entropy coding methods, it includes many variants, one of which is adopted by the latest MPEG G-PCC standard as one of its methods for compressing geometry information, namely, octree-based G-PCC geometry information encoding.
[0004] The octree-based G-PCC geometric information coding scheme is as follows: Assume {X i =(x i ,y i ,z i )} i=1…N It is the set of geometric information of all points in the input point cloud. The octree-based G-PCC first quantizes the input point cloud based on the following formula to generate a quantized point cloud
[0005]
[0006] Where X shift =(x min ,y min ,z min ) is usually the minimum value of all point coordinates in the input point cloud. s is the quantization step size, which is equal to 0.001953125 to 0.25 under the general test conditions of G-PCC, and is used to quantize the input point cloud into different qualities. Points with different coordinates in the input point cloud may be quantized to the same coordinates, and points with the same geometric coordinates can be removed. At the decoding end, The reconstructed point cloud is generated based on the following inverse quantization formula
[0007]
[0008] Then, the octree-based G-PCC uses the following steps to losslessly encode the quantized point cloud. 1) Create a bounding box containing the quantized point cloud, which is a 2 n ×2 n ×2 n , where n is the smallest integer that satisfies the following formula to ensure that the cube completely contains the quantized point cloud:
[0009]
[0010] 2) Recursively divide the cube containing the quantized point cloud into sub-cubes based on the octree. Figure 1 As shown, starting from the largest cube containing the quantized point cloud, the cube containing the point is recursively divided into 8 sub-cubes. Each octree partition generates a byte to represent whether the 8 sub-cubes contain points in the point cloud. A bit of 1 indicates that the corresponding sub-cube contains points in the point cloud, and a bit of 0 indicates that the corresponding sub-cube does not contain points in the point cloud. It should be noted that the sub-cubes containing points need to be further divided to the deepest depth. In the octree-based G-PCC, each byte is formed into a binary string using breadth-first search, and then written into the bitstream using binary context adaptive arithmetic coding, which uses whether the neighboring sub-cubes of the parent node contain points as context. At the decoding end, a bounding box containing the quantized point cloud is also first established, that is, a size of 2 n ×2 n ×2 n The cube is then recursively divided based on the bytes decoded from the bitstream until the deepest depth is reached to complete the reconstruction of the point cloud.
[0011] In the octree-based G-PCC coding framework, the loss of point cloud quality is determined by the quantization process shown in formula (1), and the process of recursively dividing the quantized point cloud by the octree is lossless.
[0012] In the above scheme, whether the spatially adjacent sub-cubes contain points is used as the context of entropy coding to remove the correlation within the point cloud frame. In addition, inter-frame prediction technology has also been proposed recently, which uses the correlation between adjacent point cloud frames to improve coding efficiency. The concept of point cloud frame is similar to that of video frame, that is, video frames are images collected at different times, and point cloud frames are static point clouds collected at different times.
[0013] Inter-frame prediction technology is an inter-frame prediction technology compatible with the G-PCC standard, which has not yet been adopted by the G-PCC standard. One of the core technologies of inter-frame prediction is motion compensation. The basic idea of motion compensation is to find the sub-cube in the reference frame corresponding to the current sub-cube based on the motion vector, and then use the octree partition of the sub-cube in the reference frame to predict the octree partition of the current sub-cube, and finally encode the predicted residual. The principle is as follows:
[0014] Assume that the coordinates of the upper left corner of the current cube are (x0, y0, z0), and the motion vector of the current cube is MV = (MV x ,MV y ,MV z ), the upper left corner coordinates (x1, y1, z1) of the corresponding cube of the current cube in the reference frame can be obtained by the following formula:
[0015] (x1,y1,z1)=(x0,y0,z0)+(MV x ,MV y ,MV z )(4)
[0016] The coordinates of the upper left corner of the cube in the reference frame combined with the side length of the cube determine the cube of the reference frame, and the coordinates of the points contained in the cube can also be obtained. Octree partitioning of the cube can determine whether each sub-cube contains a point, and the byte representation o1 of the octree partitioning of the cube can be obtained. Then the byte representation can be used to predict the octree byte representation o0 of the current cube using the following formula:
[0017]
[0018] Since the byte representation of the octree division is binary, the XOR operation can be used The XOR residual r0 is written into the bitstream in the same way as the octree partitioning. It should be noted that the reference cube found by a cube through the motion vector is used not only to predict the octree partitioning of the current cube, but also to predict the octree partitioning of all its subcubes.
[0019] Another core technology of inter-frame prediction is motion estimation, which aims to find the motion vector MV of the current cube. Since the point cloud is losslessly compressed after quantization, the optimization goal of motion estimation does not include distortion, but minimizes the residual bit R. res and motion vector bit R MV The total number of bits:
[0020]
[0021] Among them, RMV It can be obtained directly by entropy coding the motion vector, so the main question is how to calculate R res .
[0022] As described in the motion compensation process: the current cube finds the reference cube through the motion vector, which is used not only to predict the octree partitioning of the current cube, but also to predict the octree partitioning of all its subcubes. Therefore, R res It also contains the residual bits of the current cube and all its sub-cubes. For example, for a 128×128×128 cube, it is necessary to accumulate the residual bits of all cubes of sizes from 128×128×128 to 1×1×1. However, since the octree uses breadth-first search, when performing motion estimation on the current 128×128×128 cube, the context for entropy coding of its sub-cubes cannot be obtained, so the residual bits of all sub-cubes cannot be accurately calculated.
[0023] In the prior art, the residual bit R is estimated based on the prediction distortion using the following formula: res :
[0024]
[0025] Among them, N c is the number of points contained in the current cube; λ is usually set to 0.5; is the distance between the i-th point in the current cube and the nearest point in the reference cube, which can be calculated using the following formula:
[0026]
[0027] in, and are the coordinates of the i-th and j-th points of the current cube and the reference cube, N r is the number of points contained in the reference cube.
[0028] However, the prior art estimates the number of residual bits based on the prediction distortion, so there is Figure 2 For ease of understanding, Figure 2 Use a square instead of a cube. Figure 2 In the figure, the large square and the small square represent the current cube and its subcube respectively, and the thick dots represent the points in the point cloud. 1) The same prediction distortion does not mean the same number of residual bits. Figure 2As shown in part (a) of , the two points in the current cube and the two points in the reference cube have the same distance. However, the two points in the current cube belong to different sub-cubes, while the two points in the reference cube belong to the same sub-cube, so the same distance has a completely different impact on the number of bits. 2) The distance between the points in the current cube and the points in the reference cube (the points in the upper left sub-cube) only considers the distribution of the points in the current cube, not the distribution of the points in the reference cube. Figure 2 As shown in part (b) of , the formula shown in (7) can only reflect the distortion of the points in the upper left sub-cube of the reference cube, but cannot reflect the distortion of the points in the remaining three sub-cubes. Summary of the invention
[0029] The purpose of the present invention is to provide a point cloud geometric information motion estimation method, system, device and storage medium, which can accurately predict residual bits, thereby improving inter-frame prediction accuracy and saving bit rate.
[0030] The objective of the present invention is achieved through the following technical solutions:
[0031] A point cloud geometric information motion estimation method, comprising:
[0032] Based on the octree-based G-PCC coding framework, a cubic cube containing the quantized point cloud is constructed;
[0033] Set multiple scales according to the size of the cube. For each scale, use the binary representation of the current cube and the corresponding reference cube to calculate the corresponding prediction residual; where 1 in the binary representation represents the point in the point cloud, and 0 represents the point in the point cloud is not included;
[0034] A linear relationship between the multi-scale prediction residual and the residual bit is established, and the optimal motion vector is obtained by minimizing the overall number of bits including the residual bit and the motion vector bit.
[0035] A point cloud geometric information motion estimation system, comprising:
[0036] Point cloud cube construction unit, used in the octree-based G-PCC coding framework to construct a cubic cube containing the quantized point cloud;
[0037] A multi-scale residual calculation unit is used to set multiple scales according to the size of the cube, and for each scale, the corresponding prediction residual is calculated using the binary representation of the current cube and the corresponding reference cube; wherein 1 in the binary representation represents that a point in the point cloud is included, and 0 represents that a point in the point cloud is not included;
[0038] The optimal motion vector calculation unit is used to establish a linear relationship between the residual bits and the multi-scale prediction residuals, and obtain the optimal motion vector by minimizing the overall number of bits including the residual bits and the motion vector bits.
[0039] A processing device, comprising: one or more processors; a memory for storing one or more programs;
[0040] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0041] A readable storage medium stores a computer program, which implements the above method when the computer program is executed by a processor.
[0042] It can be seen from the technical solution provided by the present invention that the present invention establishes a multi-scale binary prediction residual and residual bit R res The linear relationship between res , and finally a better point cloud geometric information motion estimation scheme is obtained. Compared with the prior art, the motion estimation scheme proposed in the present invention can more accurately predict R res , thereby improving the inter-frame prediction accuracy and point cloud coding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0044] Figure 1 A schematic diagram of an octree recursively partitioning a cube and its corresponding binary representation provided as the background technology of the present invention;
[0045] Figure 2 A schematic diagram of problems existing in the prior art provided as the background technology of the present invention;
[0046] Figure 3 A flowchart of a method for estimating motion of point cloud geometric information provided by an embodiment of the present invention;
[0047] Figure 4 A schematic diagram of multi-scale binary prediction residual calculation provided by an embodiment of the present invention;
[0048] Figure 5 A schematic diagram of a fast search path provided by an embodiment of the present invention;
[0049] Figure 6A schematic diagram of a point cloud geometric information motion estimation system provided by an embodiment of the present invention;
[0050] Figure 7 A schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.
[0052] First, the terms that may be used in this article are explained as follows:
[0053] The terms "include", "comprises", "contains", "has" or other descriptions with similar semantics should be interpreted as non-exclusive inclusion. For example, including certain technical feature elements (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or products, etc.) should be interpreted as including not only certain technical feature elements explicitly listed, but also other technical feature elements known in the art that are not explicitly listed.
[0054] The following is a detailed description of a point cloud geometric information motion estimation method provided by the present invention. The contents not described in detail in the embodiments of the present invention belong to the prior art known to professional and technical personnel in the field. If no specific conditions are specified in the embodiments of the present invention, the conventional conditions in the field or the conditions recommended by the manufacturer shall be followed.
[0055] like Figure 3 As shown, a point cloud geometric information motion estimation method includes the following steps:
[0056] Step 1: Based on the octree-based G-PCC coding framework, a cubic cube containing the quantized point cloud is constructed.
[0057] This step can be implemented by referring to the solution introduced in the background technology, and will not be described in detail here.
[0058] Step 2: Set multiple scales according to the size of the cube. For each scale, use the binary representation of the current cube and the corresponding reference cube to calculate the corresponding prediction residual; where 1 in the binary representation represents the inclusion of a point in the point cloud, and 0 represents the exclusion of a point in the point cloud.
[0059] In the embodiment of the present invention, the maximum divisible depth K of the octree is determined according to the size of the cube, and the set scale needs to be smaller than the divisible depth K of the octree, and 2 to K scales are all acceptable. For example, for a 128x128x128 block, the maximum divisible depth of the octree is 7, so a maximum of 7 scales are set. In comparison, setting 7 scales can achieve the best effect. Of course, the present invention can still be implemented with less than 7 scales and the effect is better than the existing solution introduced in the background technology; therefore, the specific scale data can be set by the user according to the divisible depth of the octree combined with actual needs.
[0060] Figure 4 A schematic diagram of multi-scale binary prediction residual calculation is given. Figure 4 Similarly, squares are used instead of cubes. The largest square represents the octree node (cube), the small squares represent its child nodes (sub-cubes), and the small circles represent the points in the point cloud. Figure 4 In , two scales are taken as an example to calculate the multi-scale binary prediction residuals.
[0061] For each scale, the binary representations o0 and o1 of the current cube and the reference cube can be obtained by the method described above; as mentioned above, 1 in the binary representation represents the point included in the point cloud, and 0 represents the point not included in the point cloud; then, o0 and o1 are substituted into the above formula (5) to obtain the prediction residual r0. After performing the above calculations at each scale, multi-scale binary prediction residuals can be obtained. It should be noted that since only non-empty nodes (that is, nodes containing points) need to be further divided, when calculating the residual of scale 2, only two nodes are further calculated for binary prediction residuals.
[0062] Step 3: Establish a linear relationship between the multi-scale prediction residual and the residual bit, and obtain the optimal motion vector by minimizing the overall number of bits including the residual bit and the motion vector bit.
[0063] After obtaining the multi-scale binary prediction residual, the present invention further establishes the number of 1s and 0s in the multi-scale binary prediction residual and the residual bit R res The linear relationship between:
[0064] R res =a1×N1+b1×N0+c1(9)
[0065] Among them, R res represents the residual bits, a1, b1 and c1 are the linear model parameters, and N1 and N0 represent the number of 1s and 0s in the multi-scale binary prediction residuals respectively (summing the number of 1s and 0s in the binary prediction residuals in all scales).
[0066] Substituting the above formula (9) into the above formula (6), a more accurate motion estimation criterion is obtained as shown in the following formula:
[0067]
[0068] Since the linear model parameter c1 is a constant and does not affect the solution of the motion vector, c1 is omitted in the above formula (10).
[0069] The above-mentioned point cloud geometric information motion estimation method provided by the embodiment of the present invention can be used in combination with any motion estimation search path, and the motion estimation search path can be a full search path or any fast search path. In addition, the cube can be a cube of fixed size or variable size; if a cube of variable size is used, the optimal motion vector is calculated for each size of the cube, and then the final motion vector and the corresponding cube division method are determined according to the number of bits of each optimal motion vector; after obtaining the optimal motion vector, motion compensation is performed to obtain the residual, and then the optimal motion vector and the residual are encoded to complete the encoding of the point cloud.
[0070] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, a point cloud geometric information motion estimation method provided by an embodiment of the present invention is described in detail below with reference to a specific embodiment.
[0071] Embodiment 1
[0072] This embodiment combines the proposed motion estimation criterion with a variable cube size and a fast search path. The variable cube sizes used in this embodiment are 32×32×32, 64×64×64 and 128×128×128. It should be noted that the motion estimation criterion proposed in the present invention can support larger or smaller variable cube sizes. The supported cube size refers to transmitting a motion vector for a cube of corresponding size: the smaller the cube, the more accurate the motion estimation, that is, the fewer bits required for the residual and the more bits required for the motion vector; the larger the cube, the less accurate the motion estimation, that is, the more bits required for the residual and the fewer bits required for the motion vector. The use of different block sizes is to find the best balance between the number of bits for the residual and the motion vector. In addition, the fast search path used in this embodiment is as follows Figure 5 As shown, the motion vector consuming the least total number of bits is searched at 18 points around the center point (i.e., the darkest black point in the center of the cube), and the initial search step is 8. If the point consuming the least number of bits is not the center point, further search is performed with the point consuming the least number of bits as the center point; if the point consuming the least number of bits is the center point, the search step is reduced to half of the original until the search step is 1. The main steps are as follows:
[0073] Step 1: If there is a point within the preset search range of the reference point cloud frame, first calculate the number of bits consumed by the motion vector and the number of 0s and 1s in the corresponding multi-scale residual when the motion vector is (0,0,0) in a 128×128×128 cube, and then estimate the total number of bits of the motion vector (0,0,0) based on formula (10); if there is no point within the preset search range in the reference, inter-frame prediction is not used.
[0074] Step 2: Use Figure 5 The search path shown calculates the number of bits consumed by the motion vector of each non-center point and the number of 0s and 1s in the corresponding multi-scale residual, then estimates the total number of bits for each motion vector based on formula (10), and finally determines the optimal motion vector of a cube of size 128×128×128 by comparing the number of bits.
[0075] Step 3: Use depth-first search to traverse the octree of the 128×128×128 cube until it reaches a 32×32×32 cube. Use operations similar to the first two steps to find the optimal motion vectors for the 64×64×64 and 32×32×32 cubes.
[0076] Step 4: Compare the total bit consumption of 8 cubes of 32×32×32 and 1 cube of 64×64×64 to determine whether each cube of 64×64×64 needs to be divided; compare the total bit consumption of 8 cubes of 64×64×64 and 1 cube of 128×128×128 to determine whether each cube of 128×128×128 needs to be divided. Finally, the optimal division and motion vector are obtained.
[0077] Step 5: Use the optimal partitioning and motion vector to perform motion compensation to obtain the predicted value, then calculate the residual, and finally encode the motion vector and residual to complete the encoding of the point cloud.
[0078] Embodiment 2
[0079] This embodiment combines the proposed motion estimation criterion with the variable cube size and global search. The difference between this embodiment and the first embodiment is that a global search is used instead of a certain fast search method; the main steps are as follows:
[0080] Step 1: If there is a point within the preset search range of the reference point cloud frame, first calculate the number of bits consumed by the motion vector when the fixed-size cube is at the motion vector (0,0,0) and the number of 0s and 1s in the corresponding multi-scale residual, and then estimate the total number of bits of the motion vector (0,0,0) based on formula (10); if there is no point within the preset search range in the reference, inter-frame prediction is not used;
[0081] Step 2: Use Figure 5 The search path shown calculates the number of bits consumed by the motion vector of each non-center point and the number of 0s and 1s in the corresponding multi-scale residual, and then estimates the total number of bits of each motion vector based on formula (10), and finally determines the optimal motion vector of a cube of size 128×128×128 by comparing the number of bits;
[0082] Step 3: Use the optimal motion vector to perform motion compensation to obtain the predicted value, then calculate the residual, and finally encode the motion vector and residual to complete the encoding of the point cloud.
[0083] Step 1: If there is a point within the preset search range of the reference point cloud frame, first calculate the number of bits consumed by the motion vector and the number of 0s and 1s in the corresponding multi-scale residual when the motion vector is (0,0,0) in a 128×128×128 cube, and then estimate the total number of bits of the motion vector (0,0,0) based on formula (10); if there is no point within the preset search range in the reference, inter-frame prediction is not used.
[0084] Step 2: Use global search to calculate the number of bits consumed by each motion vector within the preset search range and the number of 0s and 1s in the corresponding multi-scale residuals. Then, estimate the total number of bits for each motion vector based on formula (10). Finally, determine the optimal motion vector of a 128×128×128 cube by comparing the number of bits.
[0085] Step 3: Use depth-first search to traverse the octree of the 128×128×128 cube until it reaches a 32×32×32 cube. Use operations similar to the first two steps to find the optimal motion vectors for the 64×64×64 and 32×32×32 cubes.
[0086] Step 4: Compare the total bit consumption of 8 cubes of 32×32×32 and 1 cube of 64×64×64 to determine whether each cube of 64×64×64 needs to be divided; compare the total bit consumption of 8 cubes of 64×64×64 and 1 cube of 128×128×128 to determine whether each cube of 128×128×128 needs to be divided. Finally, the optimal division and motion vector are obtained.
[0087] Step 5: Use the optimal partitioning and motion vector to perform motion compensation to obtain the predicted value, then calculate the residual, and finally encode the motion vector and residual to complete the encoding of the point cloud.
[0088] Embodiment 3
[0089] This embodiment combines the proposed motion estimation criteria with a fixed cube size and a fast search path. The difference between this embodiment and the first embodiment is that the fixed cube size is used instead of a variable cube size. The fixed cube size can be 128×128×128, 64×64×64, 32×32×32 and other sizes. The main steps are as follows:
[0090] Step 1: If there is a point within the preset search range of the reference point cloud frame, first calculate the number of bits consumed by the motion vector when the fixed-size cube is at the motion vector (0,0,0) and the number of 0s and 1s in the corresponding multi-scale residual, and then estimate the total number of bits of the motion vector (0,0,0) based on formula (10); if there is no point within the preset search range in the reference, inter-frame prediction is not used;
[0091] Step 2: Use Figure 5 The search path shown calculates the number of bits consumed by the motion vector of each non-center point and the number of 0s and 1s in the corresponding multi-scale residual, and then estimates the total number of bits of each motion vector based on formula (10), and finally determines the optimal motion vector of a cube of size 128×128×128 by comparing the number of bits;
[0092] Step 3: Use the optimal motion vector to perform motion compensation to obtain the predicted value, then calculate the residual, and finally encode the motion vector and residual to complete the encoding of the point cloud.
[0093] Embodiment 4
[0094] This embodiment combines the proposed motion estimation criterion with a fixed cube size and a global search. The difference between this embodiment and the second embodiment is that the cube size is fixed instead of a variable cube size.
[0095] The specific steps of this embodiment are as follows:
[0096] Step 1: If there is a point within the preset search range of the reference point cloud frame, first calculate the number of bits consumed by the motion vector when the fixed-size cube is at the motion vector (0,0,0) and the number of 0s and 1s in the corresponding multi-scale residual, and then estimate the total number of bits of the motion vector (0,0,0) based on formula (10); if there is no point within the preset search range in the reference, inter-frame prediction is not used;
[0097] Step 2: Use the global search path to calculate the number of bits consumed by the motion vector of each non-center point and the number of 0s and 1s in the corresponding multi-scale residual, then estimate the total number of bits of each motion vector based on formula (10), and finally determine the optimal motion vector of the 128×128×128 cube by comparing the number of bits;
[0098] Step 3: Use the optimal motion vector to perform motion compensation to obtain the predicted value, then calculate the residual, and finally encode the motion vector and residual to complete the encoding of the point cloud.
[0099] The above is the main principle and working process of the above method in the embodiment of the present invention. In order to illustrate the technical effect brought by the above method, it is explained through performance comparison experiments.
[0100] Taking the above-mentioned embodiment 1 as an example, a1 and b1 are set to 0.317 and 0.011, and the performance is compared with the motion estimation criterion based on prediction distortion described in the background technology under the inter-frame prediction configuration. At this time, the standard point cloud test conditions of MPEG from low bit rate (r1) to high bit rate (r6) are tested, and D1-PSNR and D2-PSNR are used to measure the reconstruction quality of the geometric information of the point cloud, and BD-rate is used to measure the coding gain of D1 and D2. Table 1 shows the test results using a typical point cloud sequence.
[0101]
[0102] Table 1 Typical point cloud sequence test results
[0103] It can be seen from Table 1 that the scheme proposed in the present invention saves 1.25% of the bit rate under D1-PSNR and D2-PSNR compared with the existing method, thereby improving the coding efficiency.
[0104] Another embodiment of the present invention further provides a point cloud geometric information motion estimation system, which is mainly used to implement the method provided in the above embodiment, such as Figure 6 As shown, the system mainly includes:
[0105] Point cloud cube construction unit, used in the octree-based G-PCC coding framework to construct a cubic cube containing the quantized point cloud;
[0106] A multi-scale residual calculation unit is used to set multiple scales according to the size of the cube, and for each scale, the corresponding prediction residual is calculated using the binary representation of the current cube and the corresponding reference cube; wherein 1 in the binary representation represents that a point in the point cloud is included, and 0 represents that a point in the point cloud is not included;
[0107] The optimal motion vector calculation unit is used to establish a linear relationship between the residual bits and the multi-scale prediction residuals, and obtain the optimal motion vector by minimizing the overall number of bits including the residual bits and the motion vector bits.
[0108] It should be noted that the specific technical details involved in each unit have been described in detail in the previous method embodiment, so they will not be repeated here.
[0109] Another embodiment of the present invention further provides a processing device, such as Figure 7 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the methods provided in the aforementioned embodiments.
[0110] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0111] In the embodiment of the present invention, the specific types of the memory, input device and output device are not limited; for example:
[0112] The input device may be a touch screen, an image acquisition device, a physical button or a mouse, etc.;
[0113] The output device may be a display terminal;
[0114] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.
[0115] Another embodiment of the present invention further provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.
[0116] In the embodiment of the present invention, the readable storage medium is a computer-readable storage medium and can be set in the aforementioned processing device, for example, as a memory in the processing device. In addition, the readable storage medium can also be a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk, etc., which can store program codes.
[0117] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for motion estimation of point cloud geometric information, characterized in that: include: Based on the octree-based G-PCC coding framework, a cube containing the quantized point cloud is constructed; Set multiple scales according to the size of the cube. For each scale, use the binary representation of the current cube and the corresponding reference cube to calculate the corresponding prediction residual; where 1 in the binary representation represents the point in the point cloud, and 0 represents the point in the point cloud is not included; A linear relationship between the prediction residual and the residual bit is established through multi-scale prediction residual, and the optimal motion vector is obtained by minimizing the overall number of bits including the residual bit and the motion vector bit; The linear relationship between the multi-scale prediction residual and the residual bit is expressed as: R res =a1×N1+b1×N0+c1 Among them, R res represents the residual bits, a1, b1 and c1 are the linear model parameters, N1 and N0 represent the number of 1s and 0s in the multi-scale binary prediction residual respectively; The optimal motion vector obtained by minimizing the number of overall bits including residual bits and motion vector bits is expressed as: In the above process, the linear model parameter c1 is omitted; R MV Indicates the motion vector MV bit.
2. A point cloud geometric information motion estimation method according to claim 1, characterized in that: The setting of multiple scales according to the size of the cube includes: determining a maximum divisible depth K of the octree according to the size of the cube, and setting a number of scales of 2 to K.
3. The method for estimating motion of point cloud geometric information according to claim 1, characterized in that: The formula for calculating the corresponding prediction residual using the binary representation of the current cube and the corresponding reference cube includes: Among them, r0 represents the prediction residual, o0 and o1 are the binary representation of the current cube and the binary representation of the reference cube respectively. It is an XOR operation.
4. A point cloud geometric information motion estimation method according to any one of claims 1 to 3, characterized in that: The cube is a cube of fixed size or variable size; if a cube of variable size is used, the optimal motion vector is calculated for each size of the cube, and then the final motion vector and the corresponding cube division method are determined according to the number of bits of each optimal motion vector.
5. A point cloud geometric information motion estimation method according to any one of claims 1 to 3, characterized in that: The method also includes: using it in combination with any motion estimation search path, after obtaining the optimal motion vector, performing motion compensation to obtain a residual, and then encoding the optimal motion vector and the residual to complete the encoding of the point cloud.
6. A point cloud geometric information motion estimation system, characterized in that: include: Point cloud cube construction unit, used in the octree-based G-PCC encoding framework to construct a cube containing the quantized point cloud; A multi-scale residual calculation unit is used to set multiple scales according to the size of the cube, and for each scale, the corresponding prediction residual is calculated using the binary representation of the current cube and the corresponding reference cube; wherein 1 in the binary representation represents that a point in the point cloud is included, and 0 represents that a point in the point cloud is not included; An optimal motion vector calculation unit, used to establish a linear relationship between the prediction residual and the residual bit through multi-scale prediction residual, and obtain the optimal motion vector by minimizing the overall number of bits including the residual bit and the motion vector bit; The linear relationship between the multi-scale prediction residual and the residual bit is expressed as: R res =a1×N1+b1×N0+c1 Among them, R res represents the residual bits, a1, b1 and c1 are the linear model parameters, N1 and N0 represent the number of 1s and 0s in the multi-scale binary prediction residual respectively; The optimal motion vector obtained by minimizing the number of overall bits including residual bits and motion vector bits is expressed as: In the above process, the linear model parameter c1 is omitted; R MV Indicates the motion vector MV bit.
7. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
8. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.