Data coding method and device, data decoding method and device, storage medium and computer equipment

By encoded point group division and context point group construction, and using entropy encoder to generate attribute code streams, the problem of low compression efficiency of point cloud data is solved and more efficient storage and transmission is achieved.

CN120017836AActive Publication Date: 2025-05-16PENG CHENG LAB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510109384.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Due to its irregular structure and high correlation, point cloud data is difficult to effectively compress, resulting in large storage space and low transmission efficiency.

Method used

By encoded point group division and context point group construction on point cloud data, an entropy encoder is used to generate attribute code streams based on the probability distribution of attribute values ​​of point data, reducing redundancy and improving compression efficiency.

Benefits of technology

On the basis of retaining massive amounts of original attribute information of point cloud data, reduce or avoid redundancy, reduce storage space usage, and improve transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017836A_ABST
    Figure CN120017836A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data coding method and device, a data decoding method and device, a storage medium and computer equipment, and the method comprises the steps: obtaining point cloud data, carrying out the coding point group division of the point data, and obtaining a plurality of coding point groups; a set of context points is then constructed for each set of coding points. For the first coding point group with the serial number of 1, coding the attribute value of the point data in the first coding point group to obtain a first sub-attribute code stream; and for a second coding point group with a non-one serial number, determining attribute value probability distribution according to the point data of the second coding point group and the point data in the context point group, and inputting the attribute value probability distribution into the entropy encoder to obtain a second sub-attribute code stream. And finally, integrating the first sub-attribute code stream and the second sub-attribute code stream into an attribute code stream of the point cloud data, thereby reducing or avoiding redundancy, reducing storage space occupation and improving data transmission efficiency. On the basis of retaining original attribute information of massive point cloud data, redundancy is reduced or avoided, storage space occupation is reduced, and transmission efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a data encoding method, decoding method, device, storage medium and computer equipment. Background Art

[0002] Point cloud is a commonly used form of 3D data representation, which is widely used in computer vision, autonomous driving, robotics and other fields. Point cloud representation retains the original geometric information in 3D space, so it plays an important role in scene understanding related applications. In recent years, with the development of deep learning technology, point cloud processing methods have also made significant progress.

[0003] In related technologies, point clouds require a lot of storage space and have high transmission costs. Compared with traditional image and video data, point clouds present highly irregular structures of different densities and scales. General attribute compression schemes cannot effectively capture the correlation between points, resulting in redundancy in attribute information of compressed point cloud data, which occupies a large storage space and has low transmission efficiency due to the existence of redundancy. Summary of the invention

[0004] The main purpose of this application is to provide a data encoding method, decoding method, device, storage medium and computer equipment, aiming to provide a new attribute compression scheme, effectively capture the correlation between points, reduce or avoid redundancy on the basis of retaining the original attribute information of massive point cloud data, reduce storage space occupancy, and improve transmission efficiency. The technical solution is as follows:

[0005] In a first aspect, an embodiment of the present application provides a data encoding method, including:

[0006] Acquire point cloud data, and divide multiple point data in the point cloud data into coded point groups to obtain multiple coded point groups;

[0007] Based on each of the coding point groups, context point groups are constructed for multiple point data in the point cloud data to obtain context point groups corresponding to each of the constructed coding point groups;

[0008] Encode the attribute value of the point data in the first code point group with a sequence number of 1 to obtain a first sub-attribute code stream;

[0009] For each second code point group whose serial number is not one, determining a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group;

[0010] Each of the attribute value probability distributions is input into an entropy encoder to obtain a second sub-attribute code stream corresponding to each of the second coding point groups, and to obtain an attribute code stream of the point cloud data, wherein the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream.

[0011] In a second aspect, an embodiment of the present application provides a data decoding method, including:

[0012] Get the attribute code stream of point cloud data;

[0013] Dividing the plurality of point data in the point cloud data into decoding point groups to obtain a plurality of decoding point groups;

[0014] Constructing context point groups for a plurality of point data in the point cloud data based on each of the decoded point groups, and obtaining context point groups corresponding to each of the constructed decoded point groups;

[0015] Decoding the attribute value of the point data in the first code point group with sequence number 1 in the attribute code stream to obtain a reconstructed value of the point data in the first code point group;

[0016] For each second code point group whose serial number is not one, determining a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group;

[0017] Each of the attribute value probability distributions and the corresponding attribute code stream is input into an entropy decoder to obtain an attribute reconstruction value of the point data in the second code point group corresponding to each of the second code point groups.

[0018] In a third aspect, an embodiment of the present application provides a data encoding device, including:

[0019] A first division unit is used to obtain point cloud data, and divide a plurality of point data in the point cloud data into coded point groups to obtain a plurality of coded point groups;

[0020] A first construction unit is used to construct context point groups for multiple point data in the point cloud data based on each of the coded point groups, so as to obtain context point groups corresponding to each of the constructed coded point groups;

[0021] An encoding unit, used for encoding the attribute value of the point data in the first code point group with a sequence number of one to obtain a first sub-attribute code stream;

[0022] A first determining unit is used to determine, for each second code point group whose serial number is not one, a probability distribution of an attribute value of each second code point group based on point data in each second code point group and point data in a corresponding context point group;

[0023] The first input unit is used to input each of the attribute value probability distributions into an entropy encoder to obtain a second sub-attribute code stream corresponding to each of the second coding point groups, and obtain an attribute code stream of the point cloud data, wherein the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream.

[0024] In a fourth aspect, an embodiment of the present application provides a data decoding device, including:

[0025] An acquisition unit, used to acquire the attribute code stream of point cloud data;

[0026] A second division unit is used to divide the multiple point data in the point cloud data into decoding point groups to obtain multiple decoding point groups;

[0027] A second construction unit is used to construct context point groups for multiple point data in the point cloud data based on each of the decoding point groups, so as to obtain context point groups corresponding to each of the constructed decoding point groups;

[0028] A decoding unit, used for decoding the attribute value of the point data in the first code point group with sequence number 1 from the attribute code stream to obtain a reconstructed value of the point data in the first code point group;

[0029] A second determining unit is used to determine, for each second code point group whose serial number is not one, a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group;

[0030] The second input unit is used to input each of the attribute value probability distributions and the corresponding attribute code stream into an entropy decoder to obtain an attribute reconstruction value of the point data in the second code point group corresponding to each of the second code point groups.

[0031] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a plurality of instructions suitable for a processor to load to execute any of the above data encoding methods or data decoding methods.

[0032] In a sixth aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above data encoding methods or data decoding methods when executing the computer program.

[0033] In an embodiment of the present application, by acquiring point cloud data, multiple point data in the point cloud data are divided into coding point groups to obtain multiple coding point groups; context point groups are constructed for multiple point data in the point cloud data based on each coding point group to obtain context point groups corresponding to each constructed coding point group; the attribute values ​​of the point data in the first coding point group with a sequence number of one are encoded to obtain a first sub-attribute code stream; for each second coding point group with a sequence number not of one, the attribute value probability distribution of each second coding point group is determined based on the point data in each second coding point group and the point data in the corresponding context point group; each attribute value probability distribution is input into an entropy encoder to obtain a second sub-attribute code stream corresponding to each second coding point group, and obtain an attribute code stream of the point cloud data, the attribute code stream including the first sub-attribute code stream and the second sub-attribute code stream. In this way, by dividing the acquired point cloud data into coding point groups, the point cloud data is divided into multiple coding point groups. This grouping method can put points with similar features or close spatial positions in the same group, which is convenient for subsequent processing. When constructing context point groups based on each coded point group, it can better consider the local information around each point, effectively capture the correlation between point data, reduce or avoid redundancy while retaining the original attribute information of massive point cloud data, reduce storage space occupancy, and improve transmission efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0035] Figure 1 A flowchart of a data encoding method provided in an embodiment of the present application.

[0036] Figure 2 A schematic diagram of a coding point group and a corresponding context point group provided in an embodiment of the present application.

[0037] Figure 3 A schematic diagram of the structure of the attention network provided in an embodiment of the present application.

[0038] Figure 4 A flowchart of a data decoding method provided in an embodiment of the present application.

[0039] Figure 5 A schematic diagram of the structure of a data encoding device provided in an embodiment of the present application.

[0040] Figure 6A schematic diagram of the structure of a data decoding device provided in an embodiment of the present application.

[0041] Figure 7 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0043] It should be noted that in some processes described in the specification, claims and the above-mentioned drawings, multiple steps appearing in a specific order are included, but it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second" or "target" in this document are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0044] Before further describing the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations:

[0045] Point cloud data is a data format used to represent the surface information of objects or scenes in three-dimensional space.

[0046] The data structure of point cloud data includes:

[0047] 1. Spatial coordinate information: The most basic is the three-dimensional spatial coordinates containing a large number of points, usually expressed in the form of (X, Y, Z). These coordinate points can depict the geometric shape of an object or scene in three-dimensional space. For example, when scanning a building, the position coordinates of each point on the outer surface of the building are obtained by a laser scanning device. These coordinates are combined to outline the general outline of the building. For some application scenarios, higher-dimensional spatial representations may also be involved. For example, in geographic information system (GIS) applications, coordinate information under the geographic spatial reference coordinate system may also be included for precise positioning.

[0048] 2. Attribute information: In addition to coordinates, point cloud data can also contain other attributes for each point. Among them, the color (RGB) attribute is a common one, which can make the point cloud present a more realistic appearance when visualized. For example, in a 3D reconstruction scene, after assigning color attributes to the point cloud, people can more intuitively distinguish the different parts of the object. The normal vector is also an important attribute information. It is used to indicate the direction of the point cloud surface at that point, which is very critical for tasks such as lighting calculation and surface reconstruction. For example, in computer graphics, the reflection and scattering effects of light on the surface of an object can be accurately simulated based on the normal vector of the point cloud. The reflection intensity attribute records the intensity of the reflection signal of each point, which is more common in acquisition methods such as laser scanning. It is related to factors such as the material and roughness of the object surface. For example, the reflection intensity of a smooth metal surface is higher, while the reflection intensity of a rough cloth surface is lower.

[0049] At present, considering the coding performance and computational efficiency, the traditional point cloud attribute coding (PCAC) scheme is the Moving Picture Experts Group Geometry-based Point Cloud Compression (MPEG G-PCC) reference software TMC13, in which the attribute coding tool is Pred-Lifting (PLT). PLT is an enhanced lifting framework based on the level of detail (LoD) structure. The LoD structure used in G-PCC relies on Euclidean distance calculation, resulting in high computational complexity, and the point-wise autoregressive coding process limits the efficiency of parallel computing. In addition, the manually designed method has several disadvantages: it relies on manually designed graphics and transformation matrices, which cannot fully capture complex and diverse geometries; in addition, it makes a strong assumption on the high correlation between geometry and attributes, which may not always hold true in the real world.

[0050] An end-to-end cloud attribute combination framework is adopted based on deep learning methods, which extends the multi-scale structure of Sparse Point Cloud Attribute Coding (SparsePCGC). It uses sparse CNN to estimate the parameters of Laplace distribution to derive attribute probabilities. However, existing deep learning based PCAC methods have several limitations. One of the key issues is their poor generalization ability to point clouds with different attributes and geometric scales and different densities. In addition, most of these methods rely on autoregressive background models, so they have high time complexity.

[0051] In order to solve the above problems, the embodiment of the present application proposes to obtain point cloud data, divide multiple point data in the point cloud data into coding point groups, and divide multiple coding point groups; construct context point groups for multiple point data in the point cloud data based on each coding point group, and obtain the context point group corresponding to each constructed coding point group; encode the attribute value of the point data in the first coding point group with a sequence number of one to obtain a first sub-attribute code stream; for each second coding point group with a sequence number not one, determine the attribute value probability distribution of each second coding point group based on the point data in each second coding point group and the point data in the corresponding context point group; input each attribute value probability distribution into an entropy encoder to obtain a second sub-attribute code stream corresponding to each second coding point group, and obtain the attribute code stream of the point cloud data, and the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream. In this way, the point cloud data is divided into multiple coding point groups by dividing the obtained point cloud data into coding point groups. This grouping method can put points with similar features or close spatial positions in the same group, which is convenient for subsequent processing. When constructing context point groups based on each coded point group, the local information around each point can be better considered, the correlation between point data and point data can be effectively captured, redundancy can be reduced or avoided on the basis of retaining the original attribute information of massive point cloud data, storage space usage can be reduced, and transmission efficiency can be improved. Please continue to refer to the following specific embodiments for details.

[0052] See also Figure 1 , Figure 1 A schematic diagram of a data encoding method provided in an embodiment of the present application. The data encoding method includes:

[0053] In step 201, point cloud data is acquired, and multiple point data in the point cloud data are divided into coded point groups to obtain multiple coded point groups.

[0054] Among them, the point cloud data is obtained by slicing all the collected point data. Slice partitioning usually refers to dividing a data set (such as an array, matrix, point cloud, etc.) into multiple smaller parts according to certain rules. These smaller parts are called "slices".

[0055] For example, spatial division in point cloud data encoding: In point cloud data encoding, "slice division" may be to segment the point cloud according to the spatial position. For example, a point cloud in a three-dimensional space is divided into multiple horizontal "slices" according to the coordinate axis direction (such as the Z axis direction). Each "slice" contains points within a specific height range, so that point clouds at different height layers can be processed separately. For example, in terrain mapping, the terrain features at different altitudes can be analyzed separately.

[0056] Specifically, after obtaining a plurality of slices, a level of detail structure (LoD) is quickly generated for the point cloud data in each slice, wherein each level of detail structure corresponds to a coding point group and a context point group.

[0057] The multiple point data in the point cloud data are divided into coding point groups to obtain multiple coding point groups. For example, the coding point groups R1, R2, ..., R L .

[0058] In some implementations, dividing the plurality of point data in the point cloud data into coded point groups to obtain a plurality of coded point groups includes:

[0059] (1) respectively obtaining a corresponding sorting code for each point data in each of the point cloud data, and sorting each point data in each of the point cloud data in ascending order of the sorting code to obtain a candidate point data set, wherein the sorting code is a Morton code or a Hilbert code;

[0060] (2) obtaining the number of preset point data for each code point group to be written;

[0061] (3) According to the number of real-time point data in the candidate point data set and the number of each preset point data, multiple point data in the point cloud data are divided into coded point groups to obtain multiple coded point groups.

[0062] Among them, the sorting code of each point data can be a Morton code or a Hilbert code. When the sorting code is a Hilbert code, it is determined by the Hilbert curve. In the point cloud data processing, the points in the three-dimensional space are indexed according to the order of the Hilbert curve. First, the spatial range of the three-dimensional point cloud data needs to be divided into a three-dimensional grid (similar to the small cube units in the three-dimensional space), and then these grid units are traversed according to the rules of the Hilbert curve. Each unit is assigned an index value. When the points in the point cloud data fall into a certain grid unit, these points inherit the index of the unit, thereby realizing the indexing of all points. In this way, the points in the point cloud data have a certain order in indexing, and this order reflects the relative position relationship of the points in the three-dimensional space.

[0063] After obtaining the sorting code of each point data, each point data is sorted according to the sorting code to obtain a candidate point data set.

[0064] For example, a point cloud containing N point data is defined as P = {p1, ..., p N}, a list of unselected points, denoted by P ns , P ns This is the candidate point dataset. The initial setting is P ns=P. In the process of dividing the code point group, P ns The point data in is arranged in the order of their Hilbert sorting codes. ns When it becomes empty, the build is complete.

[0065] For each LoD layer corresponding to the code point group to be written, a corresponding preset point data quantity is pre-set, specifically {n1,…,n L For example, the number of preset point data of the coding point group to be written corresponding to the first LoD layer is n1=16, the number of preset point data of the coding point group to be written corresponding to the second LoD layer is n2=32, etc., where n l =2×n l-1 , that is, the number of preset point data of the coding point group to be written corresponding to the current LoD layer is twice the number of preset point data of the coding point group to be written corresponding to the previous LoD layer. In the division process, according to the number of real-time point data in the candidate point data set and the number of each preset point data, the point data to be written into each coding point group is determined from the candidate point data set in turn to obtain multiple coding point groups.

[0066] In some implementations, dividing the plurality of point data in the point cloud data into coded point groups according to the number of real-time point data in the candidate point data set and the number of each of the preset point data to obtain a plurality of coded point groups includes:

[0067] (1.1) determining the number of preset point data of the target coding point group to be written, and determining the number of real-time point data in the candidate point data set;

[0068] (1.2) when the target code point group is not the last code point group to be added, dividing the point data of the candidate point data set according to the preset point data quantity and the real-time point data quantity to obtain a plurality of candidate point data subsets;

[0069] (1.3) extracting one point data from each of the candidate point data subsets in turn and putting it into the target code point group to obtain a preliminarily divided code point group;

[0070] (1.4) when the next code point group to be added is not the last code point group to be added, the next code point group to be added is determined as the target code point group, and the steps of determining the number of preset point data of the target code point group and determining the number of real-time point data in the candidate point data set are returned to be executed until the target code point group is the last code point group to be added, thereby obtaining a plurality of preliminarily divided code point groups;

[0071] (1.5) putting the remaining point data in the candidate point data set into the last code point group to be added to obtain a preliminary divided code point group;

[0072] (1.6) When the number of target point data of each of the initially divided coding point groups is less than or equal to the first point number threshold, each of the initially divided coding point groups is determined as a divided coding point group.

[0073] The division process of the code point group is as follows:

[0074] Determine the target code point group R to be written l The number of preset point data n l , and determine the candidate point dataset P ns The number of real-time point data; if R l Not the last code point group to be added R L , then according to the preset point data quantity R l And the number of real-time point data |P ns |, for the candidate point dataset P ns The point data is divided to obtain multiple candidate point data subsets; one point data is uniformly selected and extracted from each candidate point data subset and put into the target code point group in turn to obtain the initially divided code point group; after the current target code point group to be written is written, if the next code point group to be added R l+1 Not the last code point group to be added R L When the next code point group to be added is determined as the target code point group, the cycle is repeated. It should be noted that the candidate point data set P ns The candidate point data in the will gradually decrease with each extraction, making the number of real-time point data |P ns | Gradually decrease, the end condition of the loop is that the target code point group is the last code point group to be added. In this way, the code point group corresponding to the added code point group corresponding to the first LoD layer to the preliminary division of the code point group corresponding to the L-1th LoD layer is obtained. For the code point group R to be added corresponding to the Lth LoD layer L , put the point data in the candidate point dataset into the last coded point group to be added, that is, R L = Current P ns , get the initial divided code point group R L .

[0075] Among them, set the first point quantity threshold M r To limit the maximum number of point data of the coding point group corresponding to each LoD layer. If the number of target point data of each initially divided coding point group is less than or equal to the first point number threshold, it means that the initially divided coding point group does not exceed the first point number threshold, and each initially divided coding point group is determined as a divided coding point group.

[0076] In some embodiments, the method further comprises:

[0077] (1.1) when there is a code point group to be divided whose corresponding target point data quantity is greater than a first point quantity threshold in each of the initially divided code point groups, dividing the code point group to be divided according to the first point quantity threshold to obtain divided code point groups;

[0078] (1.2) Determine the other preliminarily divided code point groups except the code point group to be divided as the divided code point groups.

[0079] Among them, for each preliminary divided code point group R l , l∈1,…,L, obtain the number of target point data of each initially divided coding point group|R l |, if the number of target point data |R l |Less than or equal to the first point quantity threshold M r , then the initially divided coding point group whose corresponding target point data quantity is less than or equal to the first point quantity threshold is determined as the divided coding point group; if the target point data quantity|R l |Greater than the first point number threshold M r , that is, from the multiple initially divided coding point groups, select the coding point groups to be divided whose corresponding target point data quantity is greater than the first point quantity threshold, and then according to the first point quantity threshold M r The code point groups to be divided are divided to obtain divided code point groups.

[0080] For example, the number of target point data |R4| of the initially divided coding point group R4 is 8, and the first point number threshold M r is 4, then the initially divided code point group R4 is divided into 4, and R 4,1 and R 4,2 There are two code point groups, and the number of point data in each code point group is 4.

[0081] Based on the target point data quantity of each preliminarily divided code point group and the first point quantity threshold, at least one preliminarily divided code point group is divided to obtain divided code point groups.

[0082] In some implementations, dividing the point data of the candidate point data set according to the number of preset point data and the number of real-time point data to obtain a plurality of candidate point data subsets includes:

[0083] (1.1) Calculating the ratio of the number of real-time point data to the number of preset point data to obtain the number of divisions;

[0084] (1.2) Dividing the point data of the candidate point data set according to the number of divisions to obtain multiple candidate point data subsets.

[0085] The specific method of dividing the point data of the candidate point data set to obtain multiple candidate point data subsets is as follows:

[0086] Calculate the number of real-time point data |P ns | and the number of preset point data R l The ratio of l , that is, the interval length in the candidate point data set, and the number of divisions K is obtained l After that, according to the number of divisions K l For the candidate point dataset P ns The point data is evenly divided to obtain multiple candidate point data subsets.

[0087] In step 202, context point groups are constructed for a plurality of point data in the point cloud data based on each of the coded point groups, to obtain context point groups corresponding to each of the constructed coded point groups.

[0088] After the coding point group corresponding to each LoD layer is determined in step 201, context point groups are constructed for multiple point data in the point cloud data according to each coding point group to obtain context point groups corresponding to each constructed coding point group.

[0089] In some implementations, constructing a context point group for a plurality of point data in the point cloud data based on each of the coded point groups to obtain a context point group corresponding to each of the constructed coded point groups includes:

[0090] (1) obtaining the serial number of each context point group, and obtaining the serial number of each of the encoding point groups;

[0091] (2) Set the context point group with sequence number 1 to empty;

[0092] (3) For each context point group whose sequence number is not one, obtain the corresponding target sequence number;

[0093] (4) obtaining the total number of point data of each code point group with a sequence number from 1 to the target sequence number;

[0094] (5) comparing the total number with a second point number threshold to obtain a comparison result;

[0095] (6) determining target point data of the context point group of the target sequence number based on the comparison result;

[0096] (7) Writing the target point data into the context point group of the target sequence number to obtain the context point group corresponding to each of the constructed encoding point groups.

[0097] Since each LoD layer corresponds to a code point group and a context point group, the code point group and the context point group have correspondence. Therefore, the sequence number of each context point group is obtained, and the sequence number of each code point group is obtained; for the context point group C1 with a sequence number of 1, it is regarded as an empty set; for the context point groups C2, ..., C L , determine the corresponding target sequence number, obtain the total number of point data for each coded point group with sequence numbers from 1 to the target sequence number; set the second point quantity threshold M C To limit the size of each context point group, the total number is compared with the second point number threshold to obtain a comparison result, and the target point data of the context point group of the target sequence number is determined according to the comparison result; the target point data is written into the context point group of the target sequence number to obtain the context point group corresponding to each of the constructed coding point groups.

[0098] In some implementations, determining the target point data of the context point group of the target sequence number based on the comparison result includes:

[0099] (1.1) if the comparison result indicates that the total number is less than or equal to the second point number threshold, the point data of each coded point group with a sequence number from 1 to the target sequence number is determined as the target point data of the context point group of the target sequence number;

[0100] (1.2) If the comparison result indicates that the total number is greater than the second point number threshold, then obtaining the average index value of each point data in the encoding point group of the target sequence number;

[0101] (1.3) Calculate the index distance between the index value of each point data and the average index value;

[0102] (1.4) Arrange each point data in the order of small to large index distance to obtain a point data sequence;

[0103] (1.5) Point data with the second point quantity threshold are screened out from the point data sequence in order from the front to the back, and the screened point data are determined as target point data of the context point group with the target sequence number.

[0104] Among them, if the comparison result represents that the total number is less than or equal to the second point number threshold, it means that the number of points in the context point group to be used as the target sequence number does not exceed the second point number threshold, then the point data of each encoding point group with sequence numbers from one to the target sequence number is determined as the target point data of the context point group with the target sequence number.

[0105] If the comparison result indicates that the total number is greater than the second point quantity threshold, it means that the number of points in the context point group to be used as the target sequence number exceeds the second point quantity threshold, then the average index value of each point data in the coded point group of the target sequence number is obtained, and then the index distance between the index value of each point data and the average index value is calculated; each point data is arranged in order of index distance from small to large to obtain a point data sequence; point data of the second point quantity threshold are filtered out from the point data sequence in order from front to back, and the filtered point data are determined as the target point data of the context point group of the target sequence number.

[0106] For example, if the target number is l, it will be used as context point group C. l The number of points exceeds the second point number threshold M C When calculating the code point group R with sequence number l l The average index value of each point data in the , and according to the index distance between each point data and the average index value, select the nearest M C The context point group C l The number of target points.

[0107] The specific method for determining the context point group can refer to the following formula:

[0108]

[0109] For details, please refer to Figure 2 , Figure 2 A schematic diagram of a coding point group and a corresponding context point group provided in an embodiment of the present application.

[0110] The point cloud data includes 15 points, arranged in Hilbert order, and the Hilbert index value is the subscript number. The LoD structure has 4 layers, and the number of points in the 4th layer is greater than M r =4, so the refinement is divided into two sub-layers. After constructing the coding point group, construct the corresponding context point group. The initialized C 4,2 The number of points contained is greater than M C =8, calculate R 4,2 The average Hilbert index size is 9.25, and then the 8 nearest points are selected according to the index distance to obtain C 4,2 .

[0111] In step 203, the attribute value of the point data in the first code point group with sequence number 1 is encoded to obtain a first sub-attribute code stream.

[0112] Since the context point group C1 corresponding to the first LoD layer is an empty set, the attributes of the corresponding code point group R1 are directly saved, and the attribute values ​​of the point data in the first code point group with sequence number 1 are encoded to obtain the first sub-attribute code stream.

[0113] In step 204, for each second code point group whose serial number is not one, the attribute value probability distribution of each second code point group is determined based on the point data in each second code point group and the point data in the corresponding context point group.

[0114] Wherein, for each second code point group R2,…,R whose serial number is not one L , based on each second code point group R l Point data in, corresponding context point group C l The point data in each second code point group R l The probability distribution of attribute values.

[0115] In some implementations, determining the attribute value probability distribution of each second coded point group based on the point data in each second coded point group and the point data in the corresponding context point group includes:

[0116] (1) for each first point data in each of the second coding point groups, a first preset number of candidate point data with a smaller distance are screened out from the corresponding target context point group based on a preset proximity algorithm;

[0117] (2) sorting each candidate point data according to the Euclidean distance to obtain a corresponding sub-context point group;

[0118] (3) determining a preliminary estimated value of the first point data based on the coordinates of the first point data and the attribute values ​​and coordinates of a first preset number of candidate point data in the corresponding sub-context point group;

[0119] (4) determining the point data that is in front of a second preset number in the corresponding sub-context point group as adjacent point data to form an adjacent point data set;

[0120] (5) selecting a third preset number of second point data that are close to each of the adjacent point data from the sub-context point group, and obtaining a second point data set corresponding to each of the adjacent point data;

[0121] (6) inputting the attribute value and coordinates of each of the second point data in the second point data set into the first attention network to obtain an output local area feature;

[0122] (7) inputting the local area feature and the adjacent point data in the adjacent point data set into a second attention network to obtain a feature vector of the first point data as an output;

[0123] (8) inputting the feature vector into a multi-layer perceptron to obtain a predicted value of the first point data and a corresponding scale parameter;

[0124] (9) calculating the sum of the preliminary estimated value of the first point data and the corresponding predicted value to obtain the distribution parameter of the first point data;

[0125] (10) performing attribute value probability distribution calculation on each first point data in the second coding point group based on the distribution parameter of the first point data and the corresponding scale parameter to obtain the attribute value probability distribution of each first point data;

[0126] (11) Performing a multiplication calculation on the attribute value probability distribution of each of the first point data in the second code point group to obtain the attribute value probability distribution of each of the second code point groups.

[0127] Among them, for the Lod layer with l>1, for each corresponding second code point group R l Each first point data p in m (subscript l omitted), use the k-nearest neighbor (KNN) algorithm to select the corresponding context point group C l Find the first preset number K of candidate point data closest to each other, and sort them by Euclidean distance to form a sub-context point group S m .

[0128] After getting the sub-context point group S m Then, a preliminary prediction method based on interpolation is used to estimate p m The attribute value, that is, the initial estimate The specific calculation method of the preliminary estimated value is to determine the preliminary estimated value of the first point data based on the coordinates of the first point data and the coordinates of the first preset number of candidate point data in the corresponding sub-context point group.

[0129] Among them, the preliminary estimate The determination method can refer to the following formula:

[0130]

[0131] Among them, p j Subcontext point group S m The first preset number (for example, the first three) of candidate point data in x j is the candidate point data p j The corresponding attribute value. mj is the attribute value x j The corresponding weight can be determined by referring to the following formula:

[0132]

[0133] Among them, z j is the candidate point data pj The corresponding coordinate, z m is the first point data p m The coordinates of .

[0134] After that, the subcontext point group S m The first K1 (second preset number) point data in the m The data set of adjacent points is called S′ mk ={p mk}, where k = 1,…,K1. For each adjacent point data p mk , find it in the context point subgroup S m The K2 (third preset number) point data closer to the inner side, that is, the second point data, form each adjacent point data p nk The corresponding second point data set S″ mk ,Right now Among them, t=1,…,K2.

[0135] Specifically, after obtaining the second point data set S″ mk Then, each second point data in the second point data set The attribute value of and coordinates Input to the first attention network. The first attention network firstly pays attention to each second point data Coordinates And the corresponding attribute values Perform normalization to obtain normalized coordinates And the normalized attribute value

[0136] Normalized coordinates The determination method can refer to the following formula:

[0137]

[0138] in, The first and second data points The corresponding coordinates, The second point data currently being calculated The corresponding coordinates, To calculate the coordinates corresponding to each second point data from 2 to K2 and the first second point data The corresponding coordinates The maximum distance value among the distances is calculated by ratio, so as to obtain the data of each second point The normalized coordinates of The first attention network is mainly used to determine each second point data Corresponding sub-local region features And the corresponding attention score And for multiple second point data The corresponding sub-local area features and the corresponding attention scores are weighted and summed to obtain the final output local area feature f mk .

[0139] Normalized attribute values The determination method can refer to the following formula:

[0140]

[0141] Among them, MAX attri is the maximum attribute value, The second data point For color attributes, the compression calculation is performed in the YCoCg color space, with the maximum brightness value being 56 and the maximum chromaticity value being 512. The purpose of coordinate normalization is to ensure that S″ mk The point position in the neighborhood is based on the adjacent point data p mk As the center, the role of attribute normalization is to take the first point data p m From another perspective, attribute value normalization can be regarded as a kind of residual learning, which can explicitly model the attribute residual. Experimental results confirm that the loss converges faster.

[0142] Among them, the attention score It can be determined according to the following formula:

[0143]

[0144] This formula is the attention mechanism proposed in PTv2 that includes position embedding and subtraction relationship, where: The first and second data points The corresponding sub-local region feature, δ mul () is used to normalize the coordinates Perform some kind of multiplication-related transformation operation. This can be a linear or nonlinear operation of scaling, weighting, or other multiplication of the input coordinates, ψ key () and ψ query ()These two functions are usually used to convert the input feature vector into the "key" and "query" vectors for attention calculation. Their input is the feature vector, such as And convert it into a form suitable for calculating correlation or matching. bias () is used to input Perform bias-related operations to further adjust the coordinate information or add bias information to affect the final attention score.

[0145] Specifically, after obtaining the local area feature f mk After that, the local region feature f mk And the coordinates of each second point data Input to the second attention network, similarly, the coordinates are normalized so that the first point data p m If the features of the first stage of the first attention network have been centered on the first point data p m centered, so no additional attribute value normalization is required, and the query value in the attention score formula should be set to zero. Finally, the first point data p is obtained m The eigenvector f m .

[0146] In order to estimate each first point data p m The location of the Laplace distribution (μ m ), that is, the predicted value of the first point data and the scale parameter (σ m ), the features obtained from the attention module are input into a multi-layer perceptron (MLP). Due to residual learning, it is necessary to add the predicted values ​​back, as shown below:

[0147] (μ′ m ,σ m )=MLP(f m ),

[0148]

[0149] Among them, μ′ m is the predicted value of the first point data, μ m This is the distribution parameter of the first point data.

[0150] Specifically, the probability distribution P is simulated by a hierarchical attention network θ (X) is used to estimate the true distribution P(X). The model adopts a layer-by-layer autoregressive encoding process. Compared with the point-by-point encoding method, it is proposed to use a shared context point group to achieve parallel processing of points in the same Lod level. The mathematical expression is:

[0151]

[0152] Among them, x i is the point data currently being processed i The corresponding attribute value, C l and Represents the spatial coordinates of the context point group and the target point group, x lm and z lm They represent the attributes and spatial coordinates of the mth point in the lth target group respectively.

[0153] When l is greater than 1, the Laplace distribution L is used to model the probability distribution of the attribute, attribute x i By parameter μ i and σ i Estimation, these two parameters come from the network gθ(), that is, the attention network including the first attention network and the second attention network. The estimated value of the lth layer is expressed as:

[0154]

[0155] Among them, the attribute value of each point data is in a specific range The integral calculation is performed on , and the results are multiplied to obtain the probability distribution of the attribute value of the coding point group corresponding to the lth layer. In this way, the probability distribution of the attribute value of the coding point group corresponding to each Lod layer is calculated in turn.

[0156] During the training process, the loss function is the total number of bits used to encode all attribute values. For details, refer to the following formula:

[0157]

[0158] Among them, this formula is a loss function used to measure model performance. It is usually used when training deep learning models to quantify the difference between the model's predicted results and the actual results. In this specific formula, it is based on the probability distribution P θ (x i ) or P θ (x lm ) to calculate the loss. This loss function is used to optimize the model's parameters so that the probability distribution predicted by the model is as close to the true distribution as possible. In this process, minimizing the loss function can help the model fit the data better. bits(x i ) represents the encoding x i The number of bits required. In information theory, according to the concept of information entropy, the number of bits required to encode an event is related to the probability of the event occurring. The greater the probability, the fewer bits are required, and vice versa.

[0159] See also Figure 3 , Figure 3 A schematic diagram of the structure of the attention network provided in the embodiment of the present application, where each second point data Corresponding attribute value Input to the attribute value normalization module to obtain the normalized attribute value Normalize attribute values Input to the encoder to obtain the corresponding sub-local area features The first attention network of the attention mechanism calculates the features of each sub-local area Attention score After weighted summation, the final output local area feature f is obtained mk ; Each second point data The corresponding coordinates Input to the coordinate normalization module to obtain the normalized coordinates The local area feature f mk And the normalized coordinates Input into the second attention network of the attention mechanism to obtain the final first point data p m The eigenvector f m . The feature vector f m Input to the multi-layer perceptron to get μ′ m and the scale parameter σ m , combined with the initial estimate Calculate the first point data p m The distribution parameter μ m .

[0160] In step 205, each of the attribute value probability distributions is input into an entropy encoder to obtain a second sub-attribute code stream corresponding to each of the second coding point groups, and to obtain an attribute code stream of the point cloud data, wherein the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream.

[0161] Among them, the attribute value probability distribution of each second coding point group is used as an input parameter of an entropy encoder (such as a commonly used arithmetic entropy encoder, an adaptive entropy encoder, etc.), and the entropy encoder encodes the attribute value according to its own encoding rules based on the probability distribution.

[0162] For example, the arithmetic entropy encoder constructs probability intervals based on the probability distribution, maps the attribute values ​​to the corresponding intervals, and finally outputs the second sub-attribute code stream corresponding to each second coding point group by continuously subdividing the intervals and generating corresponding codes. The first sub-attribute code stream and all the second sub-attribute code streams are then integrated together to obtain the attribute code stream of the complete point cloud data.

[0163] By utilizing the characteristic of the entropy encoder that encodes based on the probability distribution of attribute values, efficient compression of the second coded point group is achieved to reduce the amount of data. Finally, all encoding results are integrated to form an attribute code stream to complete the encoding and compression of the attribute part of the entire point cloud data, thereby achieving the purpose of reducing data storage space occupancy and improving transmission efficiency while retaining the original attribute information.

[0164] See also Figure 4 , Figure 4 A flowchart of a data decoding method provided in an embodiment of the present application.

[0165] In step S301, the attribute code stream of the point cloud data is obtained.

[0166] When the encoded attribute code stream of the point cloud data is obtained, the coordinates of each point data in the point cloud data are also obtained for subsequent attribute value decoding.

[0167] In step S302, a plurality of point data in the point cloud data are divided into decoding point groups to obtain a plurality of decoding point groups.

[0168] The multiple point data in the multiple point cloud data are divided into coding point groups in the same way as the coding process to obtain multiple coding point groups.

[0169] In step S303, context point groups are constructed for multiple point data in the point cloud data based on each of the decoded point groups to obtain context point groups corresponding to each of the constructed decoded point groups;

[0170] In the same manner as the encoding process, context point groups are constructed for multiple point data in the point cloud data based on each coded point group to obtain context point groups corresponding to each of the constructed coded point groups. It should be noted that before decoding the attribute value code stream, the coded point group and the context point group only have coordinates but no attribute values, and the attribute values ​​are updated during the decoding process.

[0171] In step S304, the attribute value of the point data in the first code point group with sequence number 1 is decoded from the attribute code stream to obtain the reconstructed value of the point data in the first code point group.

[0172] The attribute value of the midpoint of the first coding point group R1 is decoded from the attribute code stream, and the attribute value of the midpoint of C1 is updated according to the attribute value of the midpoint of R1.

[0173] In step S305, for each second code point group whose serial number is not one, the attribute value probability distribution of each second code point group is determined based on the point data in each second code point group and the point data in the corresponding context point group.

[0174] Among them, for each second code point group whose sequence number is not one, that is, R2,…,R L , and perform attribute decoding in turn. Attention network structure, input R l Coordinates of the midpoint, C l The spatial coordinates and attribute information of the midpoint are obtained l Probability distribution of attribute values ​​at the midpoint.

[0175] In step S306, each of the attribute value probability distributions and the corresponding attribute code stream is input into an entropy decoder to obtain an attribute reconstruction value of the point data in the second code point group corresponding to each of the second code point groups.

[0176] Among them, the attribute value probability distribution and the corresponding attribute code stream are input into the entropy decoder to obtain the decompressed R l The midpoint attribute is reconstructed. C is also updated l+1 The midpoint attribute reconstruction value is finally obtained. L The attribute reconstruction value of the midpoint.

[0177] As can be seen from the above, the embodiment of the present application obtains point cloud data, divides multiple point data in the point cloud data into coding point groups, and divides multiple coding point groups; constructs context point groups for multiple point data in the point cloud data based on each coding point group, and obtains the context point group corresponding to each constructed coding point group; encodes the attribute value of the point data in the first coding point group with a sequence number of one to obtain a first sub-attribute code stream; for each second coding point group with a sequence number not being one, based on the point data in each second coding point group and the point data in the corresponding context point group, the attribute value probability distribution of each second coding point group is determined; each attribute value probability distribution is input into the entropy encoder to obtain the second sub-attribute code stream corresponding to each second coding point group, and obtain the attribute code stream of the point cloud data, and the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream. In this way, the point cloud data is divided into multiple coding point groups by dividing the acquired point cloud data into coding point groups. This grouping method can put points with similar features or close spatial positions in the same group, which is convenient for subsequent processing. When constructing context point groups based on each coded point group, it can better consider the local information around each point, effectively capture the correlation between point data, reduce or avoid redundancy while retaining the original attribute information of massive point cloud data, reduce storage space occupancy, and improve transmission efficiency.

[0178] The specific implementation of the above steps can be found in the previous embodiments, which will not be described in detail here.

[0179] In order to better implement the data transmission method provided in the embodiment of the present application, the embodiment of the present application also provides a device based on the above data transmission method. The meanings of the terms are the same as those in the above data transmission method, and the specific implementation details can refer to the description in the method embodiment.

[0180] See also Figure 5 , Figure 5 The data transmission method device may include a first division unit 601, a first construction unit 602, an encoding unit 603, a first determination unit 604, and a first input unit 605.

[0181] The first division unit 601 is used to obtain point cloud data, and divide multiple point data in the point cloud data into coded point groups to obtain multiple coded point groups;

[0182] A first construction unit 602 is used to construct context point groups for multiple point data in the point cloud data based on each of the coded point groups, to obtain context point groups corresponding to each of the constructed coded point groups;

[0183] The encoding unit 603 is used to encode the attribute value of the point data in the first code point group with a sequence number of 1 to obtain a first sub-attribute code stream;

[0184] A first determining unit 604 is configured to determine, for each second code point group whose sequence number is not one, a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group;

[0185] The first input unit 605 is used to input each of the attribute value probability distributions into the entropy encoder to obtain a second sub-attribute code stream corresponding to each of the second coding point groups, and obtain an attribute code stream of the point cloud data, wherein the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream.

[0186] In some embodiments, the first dividing unit 601 includes:

[0187] A first sorting subunit is used to respectively obtain a sorting code corresponding to each point data in each of the point cloud data, and sort each point data in each of the point cloud data in ascending order of the sorting code to obtain a candidate point data set, wherein the sorting code is a Morton code or a Hilbert code;

[0188] A first acquisition subunit is used to acquire the number of preset point data of each code point group to be written;

[0189] The division subunit is used to divide the multiple point data in the point cloud data into coded point groups according to the number of real-time point data in the candidate point data set and the number of each preset point data, so as to divide the multiple coded point groups.

[0190] In some embodiments, the sub-units are divided to:

[0191] Determine the number of preset point data of the target coded point group to be written, and determine the number of real-time point data in the candidate point data set;

[0192] When the target code point group is not the last code point group to be added, dividing the point data of the candidate point data set according to the preset point data quantity and the real-time point data quantity to obtain a plurality of candidate point data subsets;

[0193] Extracting one point data from each of the candidate point data subsets in turn and putting it into the target code point group to obtain a preliminarily divided code point group;

[0194] When the next code point group to be added is not the last code point group to be added, the next code point group to be added is determined as the target code point group, and the steps of determining the number of preset point data of the target code point group and determining the number of real-time point data in the candidate point data set are returned to be executed until the target code point group is the last code point group to be added, thereby obtaining a plurality of preliminarily divided code point groups;

[0195] Put the remaining point data in the candidate point data set into the last code point group to be added to obtain the initially divided code point group;

[0196] When the number of target point data of each of the preliminarily divided code point groups is less than or equal to the first point number threshold, each of the preliminarily divided code point groups is determined as a divided code point group.

[0197] In some embodiments, the sub-units are further used to:

[0198] When there is a code point group to be divided whose corresponding target point data quantity is greater than a first point quantity threshold in each of the initially divided code point groups, dividing the code point group to be divided according to the first point quantity threshold to obtain divided code point groups;

[0199] The other preliminarily divided code point groups except the code point group to be divided are determined as divided code point groups.

[0200] In some embodiments, the sub-units are divided to:

[0201] Calculate the ratio of the real-time point data quantity to the preset point data quantity to obtain the division quantity;

[0202] The point data of the candidate point data set is divided according to the number of divisions to obtain a plurality of candidate point data subsets.

[0203] In some embodiments, the first building unit 602 includes:

[0204] A second acquisition subunit, used to acquire the serial number of each context point group, and to acquire the serial number of each encoding point group;

[0205] A setting subunit is used to set the context point group with sequence number one to be empty;

[0206] A third acquisition subunit is used to acquire a corresponding target sequence number for each context point group whose sequence number is not one;

[0207] A fourth acquisition subunit is used to acquire the total number of point data of each coded point group having a sequence number from 1 to the target sequence number;

[0208] A comparison subunit, used for comparing the total number with a second point number threshold to obtain a comparison result;

[0209] A first determining subunit, configured to determine target point data of a context point group of a target sequence number based on the comparison result;

[0210] The writing subunit is used to write the target point data into the context point group of the target sequence number to obtain the context point group corresponding to each of the constructed coding point groups.

[0211] In some embodiments, the alignment subunit is used to:

[0212] If the comparison result indicates that the total number is less than or equal to the second point number threshold, then determining the point data of each coded point group with a sequence number from 1 to the target sequence number as the target point data of the context point group of the target sequence number;

[0213] If the comparison result indicates that the total number is greater than the second point number threshold, then obtaining an average index value of each point data in the encoding point group of the target sequence number;

[0214] Calculate the index distance between the index value of each point data and the average index value;

[0215] Arrange each point data in the order of small to large index distance to obtain a point data sequence;

[0216] Point data with the second point quantity threshold are screened out from the point data sequence in a front-to-back order, and the screened point data are determined as target point data of the context point group with a target sequence number.

[0217] In some embodiments, the first determining unit 604 includes:

[0218] A first screening subunit is configured to screen out a first preset number of candidate point data with a smaller distance from a corresponding target context point group based on a preset proximity algorithm for each first point data in each of the second coding point groups;

[0219] A second sorting subunit is used to sort each candidate point data according to the Euclidean distance to obtain a corresponding sub-context point group;

[0220] a second determining subunit, configured to determine a preliminary estimated value of the first point data based on the coordinates of the first point data and attribute values ​​and coordinates of a first preset number of candidate point data in a corresponding sub-context point group;

[0221] A third determining subunit is used to determine the point data that is in front of a second preset number in the corresponding sub-context point group as adjacent point data to form an adjacent point data set;

[0222] A second screening subunit is used to screen out a third preset number of second point data that are close to each of the adjacent point data from the sub-context point group, to obtain a second point data set corresponding to each of the adjacent point data;

[0223] A first input subunit, used for inputting the attribute value and coordinates of each second point data in the second point data set into the first attention network to obtain an output local area feature;

[0224] A second input subunit is used to input the local area feature and the coordinates of each second point data into a second attention network to obtain an output feature vector of the first point data;

[0225] A third input subunit, used to input the feature vector into a multi-layer perceptron to obtain a predicted value of the first point data and a corresponding scale parameter;

[0226] A first calculation subunit is used to calculate the sum of the preliminary estimated value of the first point data and the corresponding predicted value to obtain the distribution parameter of the first point data;

[0227] A second calculation subunit is used to perform attribute value probability distribution calculation on each first point data in the second coding point group based on the distribution parameters of the first point data and the corresponding scale parameters to obtain the attribute value probability distribution of each first point data;

[0228] The third calculation subunit is used to perform a multiplication calculation on the attribute value probability distribution of each of the first point data in the second code point group to obtain the attribute value probability distribution of each of the second code point groups.

[0229] The specific implementation of each of the above units can be found in the previous embodiments, which will not be described in detail here.

[0230] As can be seen from the above, the embodiment of the present application obtains point cloud data through the first division unit 601, divides multiple point data in the point cloud data into coded point groups, and divides multiple coded point groups; the first construction unit 602 constructs context point groups for multiple point data in the point cloud data based on each coded point group, and obtains the context point group corresponding to each constructed coded point group; the encoding unit 603 encodes the attribute value of the point data in the first coded point group with a sequence number of one to obtain a first sub-attribute code stream; the first determination unit 604 determines the attribute value probability distribution of each second coded point group based on the point data in each second coded point group and the point data in the corresponding context point group for each second coded point group with a sequence number other than one; the first input unit 605 inputs each attribute value probability distribution into the entropy encoder to obtain the second sub-attribute code stream corresponding to each second coded point group, and obtains the attribute code stream of the point cloud data, and the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream. In this way, by dividing the acquired point cloud data into coded point groups, the point cloud data is divided into multiple coded point groups. This grouping method can put points with similar features or close spatial positions in the same group for easy subsequent processing. When constructing context point groups based on each coded point group, it can better consider the local information around each point, effectively capture the correlation between point data, reduce or avoid redundancy on the basis of retaining the original attribute information of massive point cloud data, reduce storage space occupancy, and improve transmission efficiency.

[0231] See also Figure 6 , Figure 6 The data transmission method device may include an acquisition unit 701, a second division unit 702, a second construction unit 703, a decoding unit 704, a second determination unit 705, and a second input unit 706.

[0232] An acquisition unit 701 is used to acquire an attribute code stream of point cloud data;

[0233] The second division unit 702 is used to divide the multiple point data in the point cloud data into decoding point groups to obtain multiple decoding point groups;

[0234] A second construction unit 703 is used to construct context point groups for multiple point data in the point cloud data based on each of the decoding point groups, to obtain a context point group corresponding to each of the constructed decoding point groups;

[0235] A decoding unit 704 is used to decode the attribute value of the point data in the first code point group with a sequence number of 1 from the attribute code stream to obtain a reconstructed value of the point data in the first code point group;

[0236] A second determining unit 705 is configured to determine, for each second code point group whose sequence number is not one, a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group;

[0237] The second input unit 706 is used to input each of the attribute value probability distributions and the corresponding attribute code stream into an entropy decoder to obtain an attribute reconstruction value of the point data in the second code point group corresponding to each of the second code point groups.

[0238] The specific implementation of each of the above units can be found in the previous embodiments, which will not be described in detail here.

[0239] Reference Figure 7 , Figure 7 The block diagram of the structure of part of the computer device 110 for implementing the embodiment of the present disclosure. The computer device 110 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 622 (for example, one or more processors) and memory 632, and one or more storage media 630 (for example, one or more mass storage devices) storing application programs 642 or data 644. Among them, the memory 632 and the storage medium 630 can be temporary storage or permanent storage. The program stored in the storage medium 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 600. Furthermore, the central processing unit 622 can be configured to communicate with the storage medium 630 and execute a series of instruction operations in the storage medium 630 on the server 600.

[0240] The computer device 110 may also include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input and output interfaces 658, and / or one or more operating systems 641, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0241] The central processor 622 in the computer device 110 may be used to execute the data encoding method of the embodiment of the present disclosure, for example:

[0242] Acquire point cloud data, and divide multiple point data in the point cloud data into coded point groups to obtain multiple coded point groups;

[0243] Based on each of the coding point groups, context point groups are constructed for multiple point data in the point cloud data to obtain context point groups corresponding to each of the constructed coding point groups;

[0244] Encode the attribute value of the point data in the first code point group with a sequence number of 1 to obtain a first sub-attribute code stream;

[0245] For each second code point group whose serial number is not one, determining a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group;

[0246] Each of the attribute value probability distributions is input into an entropy encoder to obtain a second sub-attribute code stream corresponding to each of the second coding point groups, and to obtain an attribute code stream of the point cloud data, wherein the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream.

[0247] The central processor 622 in the computer device 110 may be used to execute the data decoding method of the embodiment of the present disclosure, for example:

[0248] Get the attribute code stream of point cloud data;

[0249] Dividing the plurality of point data in the point cloud data into decoding point groups to obtain a plurality of decoding point groups;

[0250] Constructing context point groups for a plurality of point data in the point cloud data based on each of the decoded point groups, and obtaining context point groups corresponding to each of the constructed decoded point groups;

[0251] Decoding the attribute value of the point data in the first code point group with sequence number 1 in the attribute code stream to obtain a reconstructed value of the point data in the first code point group;

[0252] For each second code point group whose serial number is not one, determining a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group;

[0253] Each of the attribute value probability distributions and the corresponding attribute code stream is input into an entropy decoder to obtain an attribute reconstruction value of the point data in the second code point group corresponding to each of the second code point groups.

[0254] The embodiments of the present disclosure also provide a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the data encoding methods of the aforementioned embodiments.

[0255] The present disclosure also provides a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that when the computer device is an encoding device, the computer program is executed to implement the above-mentioned data encoding method. For example:

[0256] Acquire point cloud data, and divide multiple point data in the point cloud data into coded point groups to obtain multiple coded point groups;

[0257] Based on each of the coding point groups, context point groups are constructed for multiple point data in the point cloud data to obtain context point groups corresponding to each of the constructed coding point groups;

[0258] Encode the attribute value of the point data in the first code point group with a sequence number of 1 to obtain a first sub-attribute code stream;

[0259] For each second code point group whose serial number is not one, determining a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group;

[0260] Each of the attribute value probability distributions is input into an entropy encoder to obtain a second sub-attribute code stream corresponding to each of the second coding point groups, and to obtain an attribute code stream of the point cloud data, wherein the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream.

[0261] The processor of the computer device reads and executes the computer program, so that when the computer device is a decoding device, the above-mentioned data decoding method is implemented. For example:

[0262] Get the attribute code stream of point cloud data;

[0263] Dividing the plurality of point data in the point cloud data into decoding point groups to obtain a plurality of decoding point groups;

[0264] Constructing context point groups for a plurality of point data in the point cloud data based on each of the decoded point groups, and obtaining context point groups corresponding to each of the constructed decoded point groups;

[0265] Decoding the attribute value of the point data in the first code point group with sequence number 1 in the attribute code stream to obtain a reconstructed value of the point data in the first code point group;

[0266] For each second code point group whose serial number is not one, determining a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group;

[0267] Each of the attribute value probability distributions and the corresponding attribute code stream is input into an entropy decoder to obtain an attribute reconstruction value of the point data in the second code point group corresponding to each of the second code point groups.

[0268] In addition, the terms "comprises" and "includes" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product or apparatus.

[0269] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0270] It should be understood that in the description of the embodiments of the present application, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to not include the number, and above, below, within, etc. are understood to include the number.

[0271] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0272] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0273] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0274] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store program codes.

[0275] It should also be understood that the various implementations provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0276] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0277] The above is a specific description of the implementation method of the present application, but the present application is not limited to the above-mentioned implementation method. Technical personnel familiar with the field can also make various equivalent modifications or substitutions without violating the spirit of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.

Claims

1. A data encoding method, characterized in that: include: Acquire point cloud data, and divide multiple point data in the point cloud data into coded point groups to obtain multiple coded point groups; Based on each of the coding point groups, context point groups are constructed for multiple point data in the point cloud data to obtain context point groups corresponding to each of the constructed coding point groups; Encode the attribute value of the point data in the first code point group with a sequence number of 1 to obtain a first sub-attribute code stream; For each second code point group whose serial number is not one, determining a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group; Each of the attribute value probability distributions is input into an entropy encoder to obtain a second sub-attribute code stream corresponding to each of the second coding point groups, and to obtain an attribute code stream of the point cloud data, wherein the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream.

2. The data encoding method according to claim 1, characterized in that: The step of dividing the plurality of point data in the point cloud data into coded point groups to obtain a plurality of coded point groups includes: Respectively obtain the corresponding sorting code of each point data in each of the point cloud data, and sort each point data in each of the point cloud data in ascending order of the sorting code to obtain a candidate point data set; Get the number of preset point data for each code point group to be written; According to the number of real-time point data in the candidate point data set and the number of each preset point data, multiple point data in the point cloud data are divided into coded point groups to obtain multiple coded point groups.

3. The data encoding method according to claim 2, characterized in that: The step of dividing the plurality of point data in the point cloud data into coded point groups according to the number of real-time point data in the candidate point data set and the number of each of the preset point data to obtain a plurality of coded point groups includes: Determine the number of preset point data of the target coded point group to be written, and determine the number of real-time point data in the candidate point data set; When the target code point group is not the last code point group to be added, dividing the point data of the candidate point data set according to the preset point data quantity and the real-time point data quantity to obtain a plurality of candidate point data subsets; Extracting one point data from each of the candidate point data subsets in turn and putting it into the target code point group to obtain a preliminarily divided code point group; When the next code point group to be added is not the last code point group to be added, the next code point group to be added is determined as the target code point group, and the steps of determining the number of preset point data of the target code point group and determining the number of real-time point data in the candidate point data set are returned to be executed until the target code point group is the last code point group to be added, thereby obtaining a plurality of preliminarily divided code point groups; Put the remaining point data in the candidate point data set into the last code point group to be added to obtain the initially divided code point group; When the number of target point data of each of the preliminarily divided code point groups is less than or equal to the first point number threshold, each of the preliminarily divided code point groups is determined as a divided code point group.

4. The data encoding method according to claim 3, characterized in that: The method further comprises: When there is a code point group to be divided whose corresponding target point data quantity is greater than a first point quantity threshold in each of the initially divided code point groups, dividing the code point group to be divided according to the first point quantity threshold to obtain divided code point groups; The other preliminarily divided code point groups except the code point group to be divided are determined as divided code point groups.

5. The data encoding method according to claim 3, characterized in that: The point data of the candidate point data set is divided according to the number of preset point data and the number of real-time point data to obtain multiple candidate point data subsets, including: Calculate the ratio of the real-time point data quantity to the preset point data quantity to obtain the division quantity; The point data of the candidate point data set is divided according to the number of divisions to obtain a plurality of candidate point data subsets.

6. The data encoding method according to claim 1, characterized in that: The step of constructing a context point group for a plurality of point data in the point cloud data based on each of the coded point groups to obtain a context point group corresponding to each of the constructed coded point groups includes: Obtaining the serial number of each context point group, and obtaining the serial number of each of the encoding point groups; Set the context point group with sequence number one to empty; For each context point group whose sequence number is not one, obtain the corresponding target sequence number; Obtain the total number of point data of each coded point group with sequence numbers from 1 to the target sequence number; Comparing the total number with a second point number threshold to obtain a comparison result; Based on the comparison result, determining the target point data of the context point group of the target sequence number; The target point data is written into the context point group of the target sequence number to obtain the context point group corresponding to each of the constructed encoding point groups.

7. The data encoding method according to claim 6, characterized in that: The step of determining the target point data of the context point group of the target sequence number based on the comparison result includes: If the comparison result indicates that the total number is less than or equal to the second point number threshold, the point data of each coded point group with a sequence number from 1 to the target sequence number is determined as the target point data of the context point group of the target sequence number; If the comparison result indicates that the total number is greater than the second point number threshold, then obtaining an average index value of each point data in the encoding point group of the target sequence number; Calculate the index distance between the index value of each point data and the average index value; Arrange each point data in the order of small to large index distance to obtain a point data sequence; Point data with the second point quantity threshold are screened out from the point data sequence in a front-to-back order, and the screened point data are determined as target point data of the context point group with a target sequence number.

8. The data encoding method according to claim 1, characterized in that: The determining, based on the point data in each of the second coded point groups and the point data in the corresponding context point group, the attribute value probability distribution of each of the second coded point groups comprises: For each first point data in each of the second coding point groups, a first preset number of candidate point data with a smaller distance are screened out from the corresponding target context point group based on a preset proximity algorithm; Sort each candidate point data according to the Euclidean distance to obtain a corresponding sub-context point group; Determine a preliminary estimated value of the first point data based on the coordinates of the first point data and the attribute values ​​and coordinates of a first preset number of candidate point data in the corresponding sub-context point group; Determine the point data that is in front of a second preset number in the corresponding sub-context point group as adjacent point data to form an adjacent point data set; Filtering out a third preset number of second point data that are close to each of the adjacent point data from the sub-context point group, to obtain a second point data set corresponding to each of the adjacent point data; Inputting the attribute value and coordinates of each of the second point data in the second point data set into the first attention network to obtain output local area features; Inputting the local area features and the coordinates of each of the second point data into a second attention network to obtain an output feature vector of the first point data; Inputting the feature vector into a multi-layer perceptron to obtain a predicted value of the first point data and a corresponding scale parameter; Calculate the sum of the preliminary estimated value of the first point data and the corresponding predicted value to obtain the distribution parameter of the first point data; Based on the distribution parameters of the first point data and the corresponding scale parameters, performing attribute value probability distribution calculation on each first point data in the second coding point group to obtain the attribute value probability distribution of each first point data; A multiplication calculation is performed on the attribute value probability distribution of each of the first point data in the second code point group to obtain the attribute value probability distribution of each of the second code point groups.

9. A data decoding method, characterized in that: include: Get the attribute code stream of point cloud data; Dividing the plurality of point data in the point cloud data into decoding point groups to obtain a plurality of decoding point groups; Constructing context point groups for a plurality of point data in the point cloud data based on each of the decoded point groups, and obtaining context point groups corresponding to each of the constructed decoded point groups; Decoding the attribute value of the point data in the first code point group with sequence number 1 in the attribute code stream to obtain a reconstructed value of the point data in the first code point group; For each second code point group whose serial number is not one, determining a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group; Each of the attribute value probability distributions and the corresponding attribute code stream is input into an entropy decoder to obtain an attribute reconstruction value of the point data in the second code point group corresponding to each of the second code point groups.

10. A data encoding device, comprising: A first division unit is used to obtain point cloud data, and divide a plurality of point data in the point cloud data into coded point groups to obtain a plurality of coded point groups; A first construction unit is used to construct context point groups for multiple point data in the point cloud data based on each of the coded point groups, so as to obtain context point groups corresponding to each of the constructed coded point groups; An encoding unit, used for encoding the attribute value of the point data in the first code point group with a sequence number of one to obtain a first sub-attribute code stream; A first determining unit is used to determine, for each second code point group whose serial number is not one, a probability distribution of an attribute value of each second code point group based on point data in each second code point group and point data in a corresponding context point group; The first input unit is used to input each of the attribute value probability distributions into an entropy encoder to obtain a second sub-attribute code stream corresponding to each of the second coding point groups, and obtain an attribute code stream of the point cloud data, wherein the attribute code stream includes the first sub-attribute code stream and the second sub-attribute code stream.

11. A data decoding device, comprising: An acquisition unit, used to acquire the attribute code stream of point cloud data; A second division unit is used to divide the multiple point data in the point cloud data into decoding point groups to obtain multiple decoding point groups; A second construction unit is used to construct context point groups for multiple point data in the point cloud data based on each of the decoding point groups, so as to obtain context point groups corresponding to each of the constructed decoding point groups; A decoding unit, used for decoding the attribute value of the point data in the first code point group with sequence number 1 from the attribute code stream to obtain a reconstructed value of the point data in the first code point group; A second determining unit is used to determine, for each second code point group whose serial number is not one, a probability distribution of attribute values ​​of each second code point group based on point data in each second code point group and point data in a corresponding context point group; The second input unit is used to input each of the attribute value probability distributions and the corresponding attribute code stream into an entropy decoder to obtain an attribute reconstruction value of the point data in the second code point group corresponding to each of the second code point groups.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the data encoding method according to any one of claims 1 to 8 or the data decoding method according to claim 9.

13. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it is implemented to execute the data encoding method according to any one of claims 1 to 8 or the data decoding method according to claim 9.

Citation Information

Patent Citations

  • Method and device for encoding and decoding data, and system

    CN111641826A

  • Point cloud attribute coding method, point cloud attribute decoding method and storage medium

    CN115278269A

  • Coding and decoding method and device for point cloud coordinate conversion residual error

    CN115412716A