Three-dimensional reconstruction method and apparatus, and electronic device
By employing a dual-centroid technique in triangulation encoding, the vertex position information associated with each centroid is obtained, thus solving the problem of point cloud distribution distortion within nodes and improving the accuracy and realism of 3D reconstruction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2026-01-13
- Publication Date
- 2026-07-23
AI Technical Summary
In existing triangular representation encoding techniques, using a single centroid to characterize the point cloud distribution inside a node leads to distortion in point cloud reconstruction.
By employing the dual centroid technique, the position information of each centroid is determined by acquiring the vertex position information associated with the two centroids, so as to more finely characterize the point cloud distribution inside the node.
It improves the accuracy and realism of 3D reconstruction and reduces distortion in point cloud reconstruction.
Smart Images

Figure CN2026072230_23072026_PF_FP_ABST
Abstract
Description
3D reconstruction methods, devices and electronic equipment
[0001] Cross-references to related applications
[0002] This application is based on Chinese Patent Application No. 202510058260.2, filed on January 14, 2025, and the priority of that Chinese Patent Application is incorporated herein by reference in its entirety. Technical Field
[0003] This application relates to the field of encoding and decoding technology, and more specifically, to a three-dimensional reconstruction method, apparatus, and electronic device. Background Technology
[0004] Trisoup encoding is a lossy geometric encoding technique that uses nodes of a certain size as units and constructs triangular patches within each node to fit a real surface. Each triangular patch consists of two vertices and a centroid, with the vertices located on the edges of the node, and the point cloud distribution inside the node being entirely represented by the centroid. The challenge lies in obtaining the centroid of each node so that it can better characterize the point cloud distribution within the node, thus avoiding distortion during point cloud reconstruction. Summary of the Invention
[0005] This application provides a three-dimensional reconstruction method, apparatus, and electronic device that can solve the problem that using a single centroid is insufficient to characterize the point cloud distribution inside a node, resulting in distortion during point cloud reconstruction.
[0006] Firstly, a three-dimensional reconstruction method is provided, executed by the encoding end, which includes:
[0007] When it is determined that the current node satisfies the first dual-centroid technical condition, obtain the vertices associated with each centroid among the two centroids;
[0008] Based on the position information of the vertices associated with each of the two centroids, determine the position information of each centroid in the two centroids;
[0009] Three-dimensional reconstruction is performed based on the position information of each of the two centroids.
[0010] Secondly, a three-dimensional reconstruction method is provided, executed by the decoding end, which includes:
[0011] When it is determined that the current node satisfies the second dual centroid technical condition, obtain the vertices associated with each centroid in the two centroids;
[0012] Based on the position information of the vertices associated with each of the two centroids, determine the position information of each centroid in the two centroids;
[0013] Three-dimensional reconstruction is performed based on the position information of each of the two centroids.
[0014] Thirdly, a three-dimensional reconstruction device is provided for use at the encoding end, the method comprising:
[0015] The first vertex acquisition module is used to acquire the vertex associated with each of the two centroids when it is determined that the current node satisfies the first dual centroid technical condition;
[0016] The first centroid position information determination module is used to determine the position information of each centroid in the two centroids based on the position information of the vertex associated with each centroid in the two centroids;
[0017] The first three-dimensional reconstruction module is used to perform three-dimensional reconstruction based on the position information of each of the two centroids.
[0018] Fourthly, a three-dimensional reconstruction device is provided for use at the decoding end, the method comprising:
[0019] The second vertex acquisition module is used to acquire the vertex associated with each of the two centroids when it is determined that the current node satisfies the second dual centroid technical condition;
[0020] The second centroid position information determination module is used to determine the position information of each centroid in the two centroids based on the position information of the vertex associated with each centroid in the two centroids;
[0021] The second 3D reconstruction module is used to perform 3D reconstruction based on the position information of each of the two centroids.
[0022] Fifthly, an electronic device is provided, the terminal including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect, or implementing the steps of the method as described in the second aspect.
[0023] A sixth aspect provides an electronic device comprising: a memory configured to store point cloud data, and processing circuitry configured to implement the steps of the method described in the first aspect, or the steps of the method described in the second aspect.
[0024] In a seventh aspect, an electronic device is provided, including a processor and a communication interface, wherein the processor is configured to run a program or instructions to implement the steps in the method embodiments of the first aspect, or to implement the steps in the method of the second aspect, and the communication interface is configured to receive an original point cloud, or to transmit a point cloud reconstructed using the method of the first aspect or the method of the second aspect.
[0025] Eighthly, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect, or implement the steps of the method described in the second aspect.
[0026] A ninth aspect provides an encoding / decoding system, comprising: an encoding end device and a decoding end device, wherein the encoding end device is configured to perform the steps of the method described in the first aspect, and the decoding end device is configured to perform the steps of the method described in the second aspect.
[0027] In a tenth aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0028] Eleventhly, a computer program / program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the steps of the method as described in the first aspect, or to implement the steps of the method as described in the second aspect.
[0029] In this embodiment, when the encoding or decoding end determines that the current node satisfies the first dual-centroid technical condition, the vertices associated with the two centroids are obtained. Then, based on the position information of the vertices associated with each centroid in the two centroids, the position information of each centroid in the two centroids is determined. Finally, 3D reconstruction is performed based on the position information of each centroid in the two centroids. This dual-centroid approach can more meticulously depict the point cloud distribution within a node, effectively avoiding the distortion problem that occurs when using a single centroid for point cloud reconstruction, thereby improving the accuracy and realism of 3D reconstruction. Attached Figure Description
[0030] Figure 1 is a schematic diagram of an encoding / decoding system that can be used in an embodiment of this application;
[0031] Figure 2a is a flowchart of an encoding process performed by an encoder that can be used in an embodiment of this application;
[0032] Figure 2b is a flowchart of an encoding process performed by another encoder available in an embodiment of this application;
[0033] Figure 3a is a flowchart of a decoding process performed by a decoder available in an embodiment of this application;
[0034] Figure 3b is a flowchart of a decoding process performed by another decoder available in an embodiment of this application;
[0035] Figure 4 is a schematic diagram of a node centroid offset in an embodiment of this application;
[0036] Figure 5A is a schematic diagram of the original point cloud distribution within a node in an embodiment of this application;
[0037] Figure 5B is a schematic diagram of the construction of triangular facets using relevant technologies for the nodes shown in Figure 5A;
[0038] Figure 5C is a schematic diagram of reconstructing the point cloud using relevant technologies for the nodes shown in Figure 5A;
[0039] Figure 6 is a schematic diagram of a three-dimensional reconstruction process in an embodiment of this application;
[0040] Figure 7A is a schematic diagram of constructing a triangular facet using the node shown in Figure 5A according to an embodiment of this application;
[0041] Figure 7B is a schematic diagram of the point cloud reconstruction of the node shown in Figure 5A using the embodiments of this application;
[0042] Figure 8 is a schematic diagram of a node division process based on the main axis direction in an embodiment of this application;
[0043] Figure 9 is a schematic diagram of another three-dimensional reconstruction process in an embodiment of this application;
[0044] Figure 10 is a schematic diagram of another three-dimensional reconstruction process in an embodiment of this application;
[0045] Figure 11 is a schematic diagram of another node division process based on the main axis direction in an embodiment of this application;
[0046] Figure 12 is a schematic diagram of another three-dimensional reconstruction process in an embodiment of this application;
[0047] Figure 13 is a schematic diagram of a three-dimensional reconstruction device according to an embodiment of this application;
[0048] Figure 14 is a schematic diagram of another three-dimensional reconstruction device in an embodiment of this application;
[0049] Figure 15 is a schematic diagram of the structure of an electronic device according to an embodiment of this application;
[0050] Figure 16 is a schematic diagram of the structure of another electronic device in an embodiment of this application. Detailed Implementation
[0051] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0052] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0053] Before introducing the technical solutions provided in the embodiments of this application, the meanings of some terms will be explained first.
[0054] Point cloud: A point cloud is a set of discrete points in space that are randomly distributed and represent the spatial structure and surface properties of a three-dimensional object or scene. Point clouds can be classified into different categories according to different classification criteria. For example, according to the method of acquiring the point cloud, it can be divided into dense point clouds and sparse point clouds; or according to the temporal type of the point cloud, it can be divided into static point clouds and dynamic point clouds.
[0055] Point cloud data: Point cloud data is composed of the geometric coordinates and attribute information of each point. Geometric coordinate information, also known as 3D position information, refers to the spatial coordinates (x, y, z) of a point in the point cloud. This can include the coordinate values of the point along each coordinate axis of a 3D coordinate system, such as the coordinate value x along the X-axis, the coordinate value y along the Y-axis, and the coordinate value z along the Z-axis. The attribute information of a point in the point cloud can include at least one of the following: color information, material information, and laser reflection intensity information (also known as reflectivity). Typically, each point in the point cloud has the same number of attribute information. For example, each point in the point cloud can have both color information and laser reflection intensity information, or it can have color information, material information, and laser reflection intensity information.
[0056] Point cloud compression (PCC) refers to the process of encoding the geometric coordinates and attribute information of each point in a point cloud to obtain a compressed bitstream. Point cloud compression includes two main processes: geometric coordinate information encoding and attribute information encoding. Currently, point cloud compression frameworks that can compress point clouds include the Geometry Point Cloud Compression (G-PCC) or Video Point Cloud Compression (V-PCC) framework provided by the Moving Picture Experts Group (MPEG), or the AVS-PCC framework provided by the Audio Video Standard (AVS).
[0057] Point cloud decoding: Point cloud decoding refers to decoding the compressed bitstream obtained from point cloud encoding to reconstruct the point cloud. More specifically, it refers to the process of reconstructing the geometric coordinates and attribute information of each point in the point cloud based on the geometric bitstream and attribute bitstream in the compressed bitstream. After obtaining the compressed bitstream at the decoding end, for the geometric bitstream, entropy decoding is first performed to obtain the quantized information of each point in the point cloud, and then inverse quantization is performed to reconstruct the geometric coordinates of each point in the point cloud. For the attribute bitstream, entropy decoding is first performed to obtain the quantized attribute residual information or quantized transform coefficients of each point in the point cloud; then, inverse quantization is performed on the quantized attribute residual information to obtain the reconstructed residual information, and inverse quantization is performed on the quantized transform coefficients to obtain the reconstructed transform coefficients. The reconstructed transform coefficients are then inversely transformed to obtain the reconstructed residual information. Based on the reconstructed residual information of each point in the point cloud, the attribute information of each point in the point cloud can be reconstructed. The reconstructed attribute information of each point in the point cloud is then matched one-to-one with the reconstructed geometric coordinate information in sequence to reconstruct the point cloud.
[0058] Figure 1 is a schematic diagram of the encoding / decoding system 10 provided in an embodiment of this application. The technical solution of this application embodiment relates to encoding / decoding (CODEC) point cloud data (including encoding or decoding).
[0059] As shown in Figure 1, the encoding / decoding system 10 includes a source device 100, which provides encoded point cloud data to be decoded and displayed by the destination device 110. Specifically, the source device 100 provides the point cloud data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of the following: desktop computer, laptop computer, tablet computer, set-top box, mobile phone, wearable device (e.g., smartwatch or wearable camera), television, camera, display device, in-vehicle device, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, digital media player, video game console, video conferencing equipment, video streaming equipment, broadcast receiver equipment, broadcast transmitter equipment, spacecraft, aircraft, robot, satellite, etc.
[0060] In the example of Figure 1, source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. Destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. Source device 100 represents an example of an encoding device, while destination device 110 represents an example of a decoding device. In other examples, source device 100 and destination device 110 may not include some of the components shown in Figure 1, or they may include components other than those shown in Figure 1. For example, source device 100 may acquire point cloud data through an external capture device. Similarly, destination device 110 may interface with an external display device instead of including an integrated display device. Furthermore, memory 102 and memory 113 may be external memories.
[0061] Although Figure 1 illustrates the source device 100 and the destination device 110 as separate devices, in some examples, they may be integrated into a single device. In such embodiments, the same hardware or software, separate hardware or software, or any combination thereof may be used to implement the functionality corresponding to the source device 100 and the functionality corresponding to the destination device 110.
[0062] In some examples, source device 100 and destination device 110 can perform unidirectional or bidirectional data transmission. In the case of bidirectional data transmission, source device 100 and destination device 110 can operate in a substantially symmetrical manner, i.e., each of source device 100 and destination device 110 includes an encoder and a decoder.
[0063] Data source 101 represents the source of point cloud data (i.e., raw, unencoded point cloud data) and provides the point cloud data to encoder 200, which encodes the point cloud data. Source device 100 may include capture devices (e.g., camera devices, sensing devices, or scanning devices), archives containing previously captured point cloud data, or feed interfaces for receiving point cloud data from data content providers. Camera devices may include ordinary cameras, stereo cameras, and light field cameras; sensing devices may include laser devices, radar devices, etc.; and scanning devices may include 3D laser scanning devices, etc. Point cloud data can be obtained by capturing real-world visual scenes using capture devices. Alternatively, data source 101 may generate computer graphics-based data as source data, or combine real-time data, archived data, and computer-generated data. For example, the data source may generate point cloud data based on virtual objects (e.g., virtual 3D objects and virtual 3D scenes obtained through 3D modeling).
[0064] Encoder 200 encodes captured, pre-captured, or computer-generated data. Encoder 200 can rearrange point cloud data from the received order (sometimes referred to as the "display order") according to the encoded order. Encoder 200 can generate a bitstream including the encoded point cloud data. Source device 100 can then output the encoded point cloud data to communication medium 120 via output interface 104 for reception or retrieval, for example, by input interface 111 of destination device 110.
[0065] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general-purpose memory. In some examples, memory 102 may store raw data from data source 101, and memory 113 may store decoded point cloud data from decoder 300. Additionally or alternatively, memories 102 and 113 may respectively store software instructions executable by, for example, encoder 200 and decoder 300. Although memories 102 and 113 are shown separately from encoder 200 and decoder 300 in this example, it should be understood that encoder 200 and decoder 300 may also include internal memory for functionally similar or equivalent purposes. If encoder 200 and decoder 300 are deployed on the same hardware device, memories 102 and 113 may be the same memory. Furthermore, memories 102 and 113 may store, for example, encoded point cloud data output from encoder 200 and input to decoder 300. In some examples, portions of memories 102 and 113 may be allocated as one or more point cloud buffers, for example, to store raw, decoded, or encoded point cloud data.
[0066] In some examples, source device 100 can output encoded data from output interface 104 to memory 113. Similarly, destination device 110 can access encoded data from memory 113 via input interface 111. Memory 113 or memory 102 can include any of a variety of distributed or locally accessed data storage media, such as hard drives, Blu-ray discs, digital versatile discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded point cloud data.
[0067] Output interface 104 may include any type of medium or device capable of transmitting encoded point cloud data from source device 100 to destination device 110. For example, output interface 104 may include a transmitter or transceiver, such as an antenna, configured to transmit encoded point cloud data directly from source device 100 to destination device 110 in real time. The encoded point cloud data may be modulated according to the communication standards of a wireless communication protocol and transmitted to destination device 110.
[0068] Communication medium 120 may include transient media, such as wireless broadcasting or wired network transmission. For example, communication medium 120 may include radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). Communication medium 120 may form part of a packet-based network (such as a local area network, a wide area network, or a global network such as the Internet). Communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, flash drive, compact disk, digital point cloud disk, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded point cloud data.
[0069] In some implementations, the communication medium 120 may include a router, switch, base station, or any other device that can be used to facilitate communication from source device 100 to destination device 110. For example, a server (not shown) may receive encoded point cloud data from source device 100 and provide it to destination device 110, for example, via network transmission. The server may include (e.g., a web server for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or File Delivery Over Unidirectional Transport (FLUTE) protocol), a Content Delivery Network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or Evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached Storage (NAS) device, etc. The server can implement one or more HTTP streaming protocols, such as MPEG Media Transport (MMT), Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), or Real Time Streaming Protocol (RTSP).
[0070] Destination device 110 can access encoded point cloud data from a server, for example, via a wireless channel (e.g., WiFi connection) or a wired connection (e.g., Digital subscriber line (DSL), cable modem, etc.) for accessing encoded point cloud data stored on the server.
[0071] Output interface 104 and input interface 111 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to the IEEE 802.11 or IEEE 802.15 standard (e.g., ZigBee™ transmission mode), Bluetooth standard, or other physical components. In an example where output interface 104 and input interface 111 include wireless components, output interface 104 and input interface 111 can be configured to operate according to Wireless Fidelity (WIFI), Ethernet, cellular networks (such as 4G (4G)... th Generation 4G mobile communication networks, Long Term Evolution (LTE), Advanced LTE, 5G (5G) th Generation 5G mobile communication network, sixth generation (6G) th Data, such as encoded point cloud data, is transmitted using Generation 6G mobile communication networks.
[0072] The technology provided in this application can be applied to support one or more of the following application scenarios: machine-perceived point clouds, which can be used in autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, disaster relief robots, and other scenarios; human-perceived point clouds, which can be used in point cloud application scenarios such as digital cultural heritage, free-viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.
[0073] The input interface 111 of the destination device 110 receives an encoded bitstream from the communication medium 120. The encoded bitstream may include high-level syntax elements and encoded data units (e.g., sequences, image groups, images, slices, blocks, etc.), where the high-level syntax elements are used to decode the encoded data units to obtain decoded point cloud data. The display device 114 displays the decoded point cloud data to the user. The display device 114 may include a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices. In some examples, the destination device 110 may not have a display device 114; for example, if the decoded point cloud data is used to determine the location of a physical object, the display device 114 may be replaced by a processor.
[0074] The encoder 200 and decoder 300 can be implemented as one or more of various processing circuits, which may include microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. When the technology is implemented wholly or partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided in the embodiments of this application.
[0075] The basic principles of the encoder 200 and decoder 300 provided in this application embodiment are introduced below, taking the G-PCC and AVS-PCC codec frameworks as examples.
[0076] The encoding and decoding frameworks of G-PCC and AVS-PCC are largely the same. Figure 2a shows the encoding flowchart executed by the encoder based on the AVS-PCC encoding framework, and Figure 2b shows the encoding flowchart executed by the encoder based on the MPEG G-PCC encoding framework. The encoder mentioned above can be the encoder 200 shown in Figure 1. The above encoding frameworks can generally be divided into a geometric coordinate information encoding process and an attribute information encoding process. In the geometric information encoding process, the geometric coordinate information of each point in the point cloud is encoded to obtain a geometric bitstream; in the attribute information encoding process, the attribute information of each point in the point cloud is encoded to obtain an attribute bitstream; the geometric bitstream and the attribute bitstream together constitute the compressed bitstream of the point cloud.
[0077] For the geometric information encoding process, the encoding flow executed by encoder 200 is as follows:
[0078] 1. Pre-processing: This can include coordinate transformation and voxelization. Through scaling and translation operations, pre-processing converts the point cloud data in 3D space into integer form and moves its smallest geometric position to the origin. In some examples, encoder 200 may not perform pre-processing.
[0079] 2. Geometric Coding: For the AVS-PCC coding framework, geometric coding includes two modes: octree-based geometric coding and prediction tree-based geometric coding. For the G-PCC coding framework, geometric coding includes three modes: octree-based geometric coding, trisoup-based geometric coding, and prediction tree-based prediction coding. Among them:
[0080] Octree-based geometric encoding: An octree is a tree-like data structure that uniformly divides a predefined bounding box in three-dimensional space, with each node having eight child nodes. By using "1" and "0" to indicate whether each child node of the octree is occupied, occupancy code information is obtained as the bitstream of point cloud geometric information.
[0081] Geometric coding based on prediction trees: A prediction tree is generated using a prediction strategy. Starting from the root node of the prediction tree, each node is traversed, and the residual coordinate value corresponding to each traversed node is encoded.
[0082] Geometric encoding based on triangulation: The point cloud is divided into blocks of a certain size, and the intersection points (called vertices) of the point cloud surface at the edges of the blocks are located. Geometric information is compressed by encoding whether there are intersection points on the edges of the blocks and the positions of the intersection points.
[0083] 3. Geometric Entropy Encoding: This method uses statistical compression encoding on the occupancy code information of the octree, the prediction residual information of the prediction tree, and the vertex information of the triangular representation, finally outputting a binary (0 or 1) compressed bitstream. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to represent the same signal. A commonly used statistical coding method is Content Adaptive Binary Arithmetic Coding (CABAC).
[0084] 4. Geometric Reconstruction: Decoding and reconstructing the geometric information after geometric encoding.
[0085] For the attribute information encoding process, the encoding flow executed by encoder 200 is as follows:
[0086] 1. Color Transformation: Apply transformations to change the color information of an attribute to a different domain. For example, color information can be transformed from the RGB color space to the YCbCr color space.
[0087] 2. Attribute Recoloring: In lossy encoding, after encoding the geometric coordinate information, the encoding end needs to decode and reconstruct the geometric information, that is, restore the geometric information of each point in the point cloud. Attribute information corresponding to one or more neighboring points in the original point cloud is used as the attribute information for the reconstructed point.
[0088] In some examples, encoder 200 may not perform color transformation or attribute recoloring.
[0089] 3. Attribute information processing: In AVS-PCC, attribute information processing can include three modes: prediction coding, transformation coding, and prediction & transformation coding. These three coding modes can be used under different conditions.
[0090] Predictive coding refers to determining the neighboring points of the point to be coded as prediction points among the already coded points based on information such as distance or spatial relationships. Based on set criteria, the predicted attribute information of the point to be coded is calculated according to the attribute information of the prediction points. The difference between the actual attribute information and the predicted attribute information of the point to be coded is calculated as attribute residual information. This attribute residual information is then quantized, transformed (optional), and entropy encoded.
[0091] Transform coding refers to using transformation methods such as Discrete Cosine Transform (DCT) and Haar Transform (Haar) to group and transform attribute information, quantize the transformation coefficients, obtain attribute reconstruction information through inverse quantization and inverse transformation, calculate the difference between the real attribute information and the attribute reconstruction information to obtain attribute residual information and quantize it, and then entropy-encode the quantized transformation coefficients and attribute residuals.
[0092] Predictive transform coding refers to using the attribute residual information obtained from prediction to perform transformation, and then quantizing and entropy coding the transform coefficients.
[0093] In MPEG G-PCC, attribute information processing can include three modes: Prediction Transform coding, Lifting Transform coding, and Region Adaptive Hierarchical Transform (RAHT) coding. These three coding modes can be used under different conditions.
[0094] Predictive transform coding refers to dividing the point cloud into multiple different levels of detail (LoD) based on distance-selected subsets of points, achieving a multi-quality, hierarchical point cloud representation from coarse to fine. Bottom-up prediction is possible between adjacent layers, where neighboring points in the coarse layer predict the attribute information of points introduced in the fine layer, obtaining the corresponding attribute residual information. The points at the lowest level are encoded as reference information.
[0095] Lift transform coding refers to introducing a weight update strategy for neighboring points on the basis of LoD neighboring layer prediction, and finally obtaining the predicted attribute information of each point and the corresponding attribute residual information.
[0096] Hierarchical region adaptive transform coding refers to the process of transforming attribute information into the transform domain, which is called the transform coefficient.
[0097] 4. Attribute Quantization: The fineness of quantization is usually determined by the quantization parameters. The transformation coefficients or attribute residuals obtained from attribute information processing are quantized, and the quantized results are entropy-coded. For example, in predictive transform coding and boost transform coding, entropy coding is performed on the quantized attribute residuals; in RAHT, entropy coding is performed on the quantized transform coefficients.
[0098] 5. Entropy Coding: The quantized attribute residual information and / or transform coefficients are generally compressed using run-length coding and arithmetic coding. The corresponding coding mode, quantization parameters, and other information are also encoded using an entropy encoder.
[0099] The encoder 200 encodes the geometric coordinate information of each point in the point cloud to obtain a geometric bitstream, and encodes the attribute information of each point in the point cloud to obtain an attribute bitstream. The encoder 200 can transmit the encoded geometric bitstream and attribute bitstream together to the decoder 300.
[0100] Figure 3a shows a decoding flowchart executed by the decoder in the AVS-PCC-based decoding framework, and Figure 3b shows a decoding flowchart executed by the decoder in the MPEG G-PCC-based decoding framework. The decoder can be the decoder 300 shown in Figure 1. After receiving the compressed bitstream (i.e., attribute bitstream and geometric bitstream) transmitted by the encoder 200, the decoder 300 decodes the geometric bitstream to reconstruct the geometric coordinate information of each point in the point cloud, and decodes the attribute bitstream to reconstruct the attribute information of each point in the point cloud.
[0101] The decoding process performed by decoder 300 is as follows:
[0102] 1. Entropy Decoding: Perform entropy decoding on the geometric bitstream and attribute bitstream respectively to obtain geometric syntax elements and attribute syntax elements.
[0103] 2. Geometric Decoding: For the AVS-PCC coding framework, geometric decoding includes two modes: octree-based geometric decoding and prediction tree-based geometric decoding. For the G-PCC coding framework, geometric decoding includes three modes: octree-based geometric decoding, trisoup-based geometric decoding, and prediction tree-based prediction decoding.
[0104] Octree-based geometric decoding: reconstructing the octree based on the geometric syntax elements obtained from parsing the geometric bitstream.
[0105] Geometric Decoding Based on Prediction Trees: Reconstructing the prediction tree based on the geometric syntax elements obtained from parsing the geometric bitstream.
[0106] Geometric Decoding Based on Triangle Representation: Reconstructing the triangular model based on the geometric syntax elements obtained from parsing the geometric bitstream.
[0107] 3. Geometric Reconstruction: Perform reconstruction to obtain the geometric coordinate information of the points in the point cloud.
[0108] 4. Inverse coordinate transformation: Perform an inverse transformation on the reconstructed geometric coordinate information to convert the reconstructed coordinates (positions) of points in the point cloud from the transformation domain back to the initial domain.
[0109] 5. Dequantization: Dequantizes attribute syntax elements.
[0110] 6. Attribute Information Processing: In AVS-PCC, attribute information processing determines the color information of points in the point cloud by predicting or predicting the transformation of the inverse-quantized prediction residual or prediction residual transformation coefficients, or by transforming the transformation coefficients of the inverse-quantized transformation.
[0111] In MPEG G-PCC, attribute information processing determines the color information of points in the point cloud by using RAHT to invert the attribute information, or by using LOD and inverse boosting to determine the color information of points in the point cloud.
[0112] 7. Inverse Color Transformation: Transforms color information from the YCbCr color space to the RGB color space. In some examples, the inverse color transformation operation may not be necessary.
[0113] The three-dimensional reconstruction method provided in the embodiments of this application is described below with reference to the accompanying drawings. The three-dimensional reconstruction method provided in the embodiments of this application can be executed by an encoding end, such as the encoder 200 shown in Figure 1. The encoding end can be implemented by software, hardware, or a combination thereof. When it is implemented by hardware, the encoding end can be referred to as an encoding end device.
[0114] When reconstructing point clouds using Trisoup encoding, the point cloud distribution within a node is entirely represented by its centroid. One way to characterize the point cloud distribution within a node using its centroid is to use the node's original centroid. Alternatively, centroid offset can be considered to further optimize the characterization of the point cloud distribution within the node. This centroid offset refers to calculating the offset based on the distribution of the surrounding original point cloud, using the original centroid and vertices as a basis. For example, triangular faces are constructed based on the original centroid and vertices, and the sum of the normal vectors of each facet is obtained. The offset direction is calculated based on the sum of the normal vectors, and then a weighted average is taken based on the projection lengths of the original point clouds within a certain vertical distance from this normal vector onto this normal vector to obtain the offset length, thus determining the offset position. Centroid offset can better fit the distribution of the original point cloud within a node. For example, as shown in Figure 4, the original centroid C is determined based on the vertices within the node. mean Then, based on the calculated offset length drift and offset direction, i.e. the direction of the normal vector n, the position of the centroid C after offset is determined, and then the point cloud is reconstructed using the centroid C.
[0115] Because the size of trisoup nodes varies depending on the bitrate, for example, the lower the bitrate, the larger the trisoup node, and the more complex the point cloud distribution inside the node. In this case, if a single centroid method is still used, it is insufficient to characterize the point cloud distribution inside the node. For example, when the original point cloud distribution inside the node is as shown in Figure 5A, there are two parts of the original point cloud in the node, and these two parts of the original point cloud are not continuous. Only one centroid is used to characterize the point cloud. The triangular facet constructed using one centroid can be shown in Figure 5B, and the final reconstructed point cloud can be shown in Figure 5C. As can be seen from Figure 5C, the reconstructed point cloud constructed using the existing technology will be distributed throughout the entire node, which is significantly different from the actual original point cloud distribution in Figure 5A. Therefore, the existing point cloud reconstruction method will lead to significant distortion in the reconstructed point cloud.
[0116] Therefore, this application provides a three-dimensional reconstruction method. When applied to the encoding end, as shown in Figure 6, the method includes the following steps:
[0117] Step 601: When it is determined that the current node satisfies the first dual-centroid technical condition, obtain the vertices associated with each centroid among the two centroids;
[0118] In this embodiment, the aforementioned node is a basic unit in a data structure used for point cloud encoding. It can be a trisoup node, and each trisoup node can be regarded as corresponding to a cube. The surface of the cube is approximately represented by a series of triangles. The aforementioned first dual centroid technique condition is a pre-set criterion used to determine whether the current node needs to adopt the dual centroid technique. The aforementioned centroid is different from the original centroid of the current node. The original centroid is the weighted average position of all vertices within the node and has not undergone centroid offset. In this embodiment, the position of each centroid in the two centroids is the weighted average position of some vertices associated with it. The aforementioned vertex refers to the intersection of the point cloud surface with the edge of the node, while the vertex associated with the centroid refers to the vertex used to determine the centroid of a certain region. This vertex can be an edge vertex, that is, a vertex located on the edge of the current node.
[0119] In some embodiments, the vertices within each node can be divided into two parts along the principal axis according to the coordinates of the original centroid, such as a first part and a second part. The vertices of the first part are used to calculate a centroid, and the vertices of the first part are the vertices associated with that centroid. The vertices of the second part are used to calculate another centroid, and the vertices of the second part are the vertices associated with the other centroid. Specifically, obtaining the vertices associated with each of the two centroids refers to obtaining the position information of the vertices associated with each of the two centroids, such as the coordinates of the vertices.
[0120] Step 602: Determine the position information of each centroid in the two centroids based on the position information of the vertices associated with each centroid in the two centroids;
[0121] Based on step 601, the position information of each centroid is determined by using the position information of the vertices associated with each centroid of the two centroids. For example, when the point cloud distribution within a node is uneven, the coordinates of the vertices associated with each centroid and their corresponding weights are used to perform a weighted average calculation to obtain the coordinates of the centroid.
[0122] Step 603: Perform three-dimensional reconstruction based on the position information of each of the two centroids.
[0123] Based on step 602, three-dimensional reconstruction, i.e. point cloud reconstruction, is completed at the encoding end according to the determined position information of each centroid.
[0124] In this embodiment, after determining that the current node meets the first dual-centroid technical condition, the vertices associated with each centroid of the two centroids are obtained. Then, the position information of each centroid is determined using the position information of the vertices. Finally, three-dimensional reconstruction is performed based on the determined position information of each centroid. This allows for a better depiction of the point cloud distribution within the node using dual centroids, effectively avoiding the distortion problem that exists when using a single centroid for point cloud reconstruction, thereby improving the accuracy and detail of the point cloud reconstruction. Still targeting the original point cloud shown in Figure 5A, point cloud reconstruction according to the scheme in this embodiment yields two centroids, such as centroid A and centroid B in Figure 7A. The triangular facets constructed using these two centroids are shown in Figure 7A, and the final reconstructed point cloud is shown in Figure 7B. This reconstructed point cloud is still distributed according to two regions, similar to the original point cloud distribution in Figure 5A, without significant distortion.
[0125] In this embodiment, when acquiring the vertices associated with each of the two centroids at the encoding end, it is necessary to first determine whether there are vertices on the edges of the current node, i.e., edge vertices. If edge vertices exist, their position information is quantized for transmission. A vertex presence flag can be used to indicate the presence of a vertex; for example, a vertex presence flag Flag0 = 1 indicates that the current node has vertices. Specifically, the non-repeating edges of the current node are first reordered lexicographically. Then, the vertex presence flag and the context information of the vertex position are determined using the node's edge neighbor information. Following the reordered order, an encoding tool, such as Dynamic OBUF, is used to encode the vertex presence flag and the quantized vertex position information into a bitstream for output. If the current node does not have edge vertices, only the vertex presence flag needs to be encoded into a bitstream and transmitted to the decoding end. Thus, determining the existence of vertices before performing the dual-centroid technique and encoding the relevant valid information into a bitstream improves the accuracy and efficiency of encoding, providing reliable data support for subsequent point cloud reconstruction.
[0126] In this embodiment, when determining whether the current node meets the first dual-centroid technical condition, the distribution of the point cloud around the original centroid of the current node can be judged first to determine whether the dual-centroid technique needs to be executed. This can be determined by whether the number of point clouds within a second preset distance of the original centroid of the current node is less than or equal to a first quantity. Specifically, if the number of point clouds within the second preset range of the original centroid is too small, it indicates that the point cloud distribution within the current node is relatively scattered and not concentrated around the original centroid. In this case, using a single original centroid to characterize the point cloud would lead to significant distortion in the reconstructed point cloud. Therefore, by dividing the point cloud within the node into two parts, and then calculating the corresponding centroid based on each part, the dual-centroid technique is implemented, resulting in better point cloud reconstruction. Thus, by comparing the number of point clouds within the second preset range of the original centroid of the current node with the first quantity, it is possible to more accurately determine whether the current node meets the first dual-centroid technical condition, thereby ensuring both coding efficiency and the accuracy and quality of the reconstructed point cloud. The aforementioned second preset range refers to a certain spatial region surrounding the original centroid, such as a spherical region with a radius of 2cm, a cubic region with a side length of 2cm, or other shaped regions surrounding the original centroid. It can be a preset value, or it can be obtained by the encoding end by adjusting the aforementioned preset value according to the adjustment parameters, or it can be determined according to the bitrate. That is, when the bitrate is different, the corresponding second preset range value will also be different.
[0127] In this embodiment, the aforementioned first quantity can be a preset value, or it can be obtained by adjusting the preset value according to adjustment parameters. The adjustment parameters include at least one of the following: average point cloud distance, vertex search distance, point cloud density parameter, and point cloud sparsity parameter. For example, the encoding end adjusts the preset value of the first quantity according to the point cloud density parameter. When the point cloud density parameter is true, the first quantity = preset value + n; when the point cloud density parameter is false, the first quantity = preset value. By comprehensively considering these adjustment parameters, the specific value of the first quantity can be determined more accurately, thereby ensuring that reasonable judgments and decisions can be made based on the actual situation when performing the dual-centroid technique.
[0128] The average point cloud distance can be calculated based on the positional information of each point cloud. The vertex search distance is used to determine the vertices of the current node, which can be calculated and quantified based on the positional information of the point clouds within a certain distance around the edge of the current node. The point cloud density parameter can be obtained by dividing the number of points in the current node by the spatial size of the current node. After the vertices in the current node are divided into two parts, the point cloud density parameter of each part needs to be recalculated. The point cloud sparsity parameter is determined based on the distribution and number of points in the current node and can be used as a quantitative indicator to evaluate the sparsity of the point cloud in the current node. The average point cloud distance and vertex search distance mentioned above are adjustment parameters already used in the prior art, while the point cloud density parameter and point cloud sparsity parameter need to be calculated based on the number of points in the current node and the size of the node space, or based on the number of points in each part of the current node and the spatial size of each part.
[0129] In this embodiment, a first identifier for the current node can also be generated. This first identifier is used to instruct the current node to perform the dual-centroid technique. For example, if it is determined that the current node meets the first dual-centroid technique condition, that is, the number of point clouds within a second preset distance of the original centroid of the current node is less than or equal to a first number, then the first identifier is set. Specifically, this can be achieved by setting the value of the identifier Flag1, i.e., Flag1 = 1 can be set as the first identifier. When the first dual-centroid technique condition is not met and the encoding end does not perform the dual-centroid technique, Flag1 = 0 is set. Subsequently, the first identifier Flag1 = 1 can be added to the bitstream, and Flag1 = 0, which indicates that the dual-centroid technique is not performed, can also be added to the bitstream, i.e., both can be encoded into a bitstream for transmission. The bitstream refers to a data stream composed of a string of binary numbers (0 and 1).
[0130] In some embodiments, qualification conditions for each node to perform the dual centroid technique can be set, and a qualification determination operation can be performed. Then, only when the current node meets the qualification conditions for performing the dual centroid technique is the above-mentioned point cloud distribution based on the original centroid of the current node used to determine whether the dual centroid technique needs to be performed. That is, determining whether the current node meets the first dual centroid technique condition also includes performing a qualification judgment on the current node to determine whether the current node meets the qualification conditions for performing the dual centroid technique. If the current node meets the qualification conditions for performing the dual centroid technique, and the point cloud distribution also meets the conditions for performing the dual centroid technique, it means that the current node can use the centroid position information determined by the position information of the vertices associated with each centroid to perform 3D reconstruction. Otherwise, 3D reconstruction is performed using existing methods, such as the above-mentioned 3D reconstruction method based on the original centroid.
[0131] In this scenario, the encoder can generate the first identifier and add it to the bitstream only when a node meets the eligibility criteria for executing the dual centroid technique. For nodes that do not meet the eligibility criteria, it is no longer necessary to determine whether the number of point clouds within the second preset distance of the current node's original centroid is less than or equal to the first number. Therefore, the first identifier Flag1=1 will not be generated and sent, nor will Flag1=0 be sent. Similarly, the decoder only needs to obtain the first identifier from the bitstream when the eligibility criteria for executing the dual centroid technique are met. When the eligibility criteria are not met, it is not necessary to obtain the first identifier from the bitstream. This reduces the amount of data transmitted between the encoder and decoder, and lowers encoding redundancy and encoding costs.
[0132] In this embodiment of the application, the qualification determination operation of the current node can be performed based on various information about the node, specifically including at least one of the following operations:
[0133] Determine that the size of the current node is greater than or equal to the first preset size;
[0134] Determine if the number of vertices in the current node is greater than or equal to the first preset number;
[0135] Determine if the vertex distribution information of the current node meets the preset distribution conditions;
[0136] The centroid offset information within the node is determined to meet the preset conditions.
[0137] Optionally, the size information of the current node includes the node width, which represents the size of the cube space represented by the current node. In this application, it can represent the size of the space occupied by the current node in the original point cloud. The node size information needs to be set according to the actual encoding requirements and scenario to achieve the best encoding effect and performance. The first preset size is used to filter nodes that can be subjected to the double centroid technique to better meet the encoding requirements of point cloud reconstruction. For example, the node width in the first preset size information can be set to 8 to meet the encoding requirements of point cloud reconstruction. The first preset quantity is used to filter nodes suitable for the double centroid technique based on the number of vertices. That is, the double centroid technique can only be performed when the number of vertices in the node reaches at least the preset value. The setting of this parameter is used to ensure that the current node has enough vertices to achieve a high-quality point cloud reconstruction effect, thereby avoiding the distortion problem of point cloud reconstruction caused by too few vertices in the node.
[0138] The vertex distribution information of the current node described above describes the distribution of vertices within the current node. The preset distribution conditions are used to filter nodes eligible for the dual-centroid technique based on the vertex distribution within the current node. Setting this parameter ensures that the selected nodes have dual centroids and that the dual-centroid technique is applied. The centroid offset information within the node indicates whether the current node has a centroid offset; the preset conditions are used to filter nodes eligible for the dual-centroid technique based on the centroid offset information within the node.
[0139] In the embodiments of this application, the first preset size, the first preset quantity, the preset distribution conditions, and the preset conditions mentioned above can be preset, or they can be obtained by the encoding end adjusting the preset values or preset conditions according to the adjustment parameters, or they can be determined according to the bit rate. That is, when the bit rate is different, the corresponding second preset range value will also be different.
[0140] In some embodiments, when performing a qualification determination operation based on the size information of the current node, it can be determined by judging whether the size information of the current node reaches a first preset size. Specifically, it can be determined whether the node width of the current node is greater than or equal to a first preset value. If the node width of the current node is greater than or equal to the first preset value, it is determined that the current node meets the qualification conditions for performing the dual centroid technique, and point cloud reconstruction can be performed using dual centroids. Otherwise, it is considered that the current node does not meet the qualification conditions for the dual centroid technique, and point cloud reconstruction can be completed in another way, such as using the point cloud reconstruction method based on the original centroid described above. For example, when the first preset value th1 = 8 is set, if the node width of the current node is greater than or equal to 8, it is determined that it can perform the dual centroid technique. If the node width of the current node is less than 8, point cloud reconstruction can be completed in another way, such as using the point cloud reconstruction method based on the original centroid.
[0141] In some embodiments, when performing a qualification determination operation based on the vertex distribution information of the current node, it can be determined by judging whether the vertex distribution information of the current node meets a preset distribution condition. Specifically, it can be judged whether the vertices of the current node can be divided into two parts in the principal axis direction, and whether the number of vertices in each part is greater than or equal to a second preset number. In this case, the following situations are included:
[0142] (1) When it is determined that the vertices of the current node can be divided into two parts in the direction of the main axis, and the number of vertices in each part is greater than or equal to the second preset number, for example, when the current node is the node shown in Figure 5A, it is determined that it can perform the double centroid technique and calculate the centroid of each part.
[0143] (2) If the vertex of the current node cannot be divided into two parts in the direction of the main axis, other methods can be used, such as the method based on the original centroid, to complete the point cloud reconstruction.
[0144] (3) If the vertices of the current node can be divided into two parts in the direction of the main axis, but the number of vertices in at least one part is less than the second preset number, then the current node is considered not to meet the qualification conditions for performing the double centroid technique. In this case, other methods can be used, such as the original centroid method, to complete the point cloud reconstruction.
[0145] In this way, by performing the qualification determination operation based on the vertex distribution information of the current node, the point cloud distribution within the current node can be better characterized by the double centroid, thereby reducing the degree of distortion of the reconstructed point cloud.
[0146] In this embodiment of the application, when it is determined that the vertex of the current node can be divided into two parts along the principal axis direction, the principal axis direction can be determined first, and then the vertex can be divided according to the principal axis direction. Specifically, as shown in Figure 8, the following steps are included:
[0147] Step 701: Determine the main axis direction of the current node;
[0148] In this embodiment, the principal axis direction can be either the direction of the sum of the normal vectors of the triangular facets formed by the centroid and the vertices, or the direction of the coordinate axis components of the centroid. Specifically, the principal axis direction can be determined through the following process: First, calculate an original centroid using all vertices within the current node. Then, determine the arrangement order of the vertices along the x, y, and z directions respectively, and construct triangular facets together with the original centroid. Next, calculate the sum of the normal vectors of each triangular facet. Finally, select the direction with the largest sum of normal vector components as the principal axis direction.
[0149] Step 702: Determine the vertices in the main axis direction whose coordinate values are greater than the original centroid of the current node as the first part of vertices, and the vertices whose coordinate values are less than or equal to the original centroid of the current node as the second part of vertices. The distance between the first vertex in the first part of vertices and the second vertex in the second part of vertices is greater than a first preset distance. The first vertex is the vertex in the first part of vertices that is closest to the original centroid, and the second vertex is the vertex in the second part of vertices that is closest to the original centroid.
[0150] Based on step 701, along the main axis, the coordinates of the vertices and the original centroid are calculated according to the position information of all vertices within the current node and the original centroid. Then, these two types of coordinates are compared. If the coordinates of a vertex are greater than the coordinates of the original centroid, it is classified as a first part of vertices. If the coordinates of a vertex are less than or equal to the coordinates of the original centroid, it is classified as a second part of vertices. Furthermore, if the distance between the vertex closest to the original centroid in the first part and the vertex closest to the original centroid in the second part is greater than a first preset distance, then the vertex division is considered valid. The first preset distance is used to determine whether the vertices of the current node can be divided into two parts along the main axis.
[0151] In this embodiment, by comprehensively considering information such as the principal axis direction, vertex position, centroid position, and distance between vertices, it is possible to determine whether the current node's vertex can be divided into two parts along the principal axis direction. This improves the accuracy and reliability of the judgment, thereby effectively reducing the degree of distortion during point cloud reconstruction and improving the reconstruction effect of the point cloud.
[0152] In some embodiments, when performing a qualification determination operation based on the centroid offset information within the current node, it can be determined by judging whether the centroid offset information within the current node meets a preset condition. Specifically, it can be determined whether there is an offset within the current node's centroid, which includes the following situations:
[0153] (1) If there is no offset of the centroid within the current node, then the offset information of the centroid within the current node meets the preset conditions, and subsequent operations can be performed.
[0154] (2) If the centroid of the current node is offset but the centroid offset value is zero, then the centroid offset information of the current node meets the preset conditions and subsequent operations can be performed.
[0155] (3) If the centroid of the current node is offset and the centroid offset value is not zero, it is determined that the centroid offset information of the current node does not meet the preset conditions. At this time, other methods can be used to complete the point cloud reconstruction, such as using the above-mentioned method based on the original centroid to complete the point cloud reconstruction.
[0156] In this way, by determining whether the centroid offset information within a node meets the preset conditions, it is possible to determine whether the current node meets the eligibility criteria for executing the dual-centroid technique. This reduces unnecessary coding processes and improves the overall efficiency of the dual-centroid technique.
[0157] In this embodiment, after determining that the node meets the eligibility criteria for performing the dual-centroid technique based on the centroid offset information within the current node, the centroid offset information also needs to be encoded. The encoding includes three parts: whether the centroid offset value is zero, the sign bit of the centroid offset value, and the absolute value of the centroid offset value. The encoding of whether the centroid offset value is zero is performed first. The centroid offset value is calculated based on the position information of the vertices surrounding the original centroid within the current node.
[0158] In the embodiments of this application, when performing qualification determination on the current node, one type of information or a combination of multiple types of information of the node can be selected to perform qualification determination on the current node according to actual needs. In some embodiments, in order to improve the effect of point cloud reconstruction, the size information, number of vertices, distribution information of vertices, and centroid offset information of the current node can be comprehensively considered to perform qualification determination on the current node. If the current node meets the corresponding conditions in the above multiple qualification determination operations, it is considered that the node has the qualification conditions to perform the dual centroid technique.
[0159] In this embodiment of the application, when the encoding end performs 3D reconstruction based on the position information of each centroid in the two centroids, it also needs to combine the position information of the vertices associated with each centroid to perform 3D reconstruction collaboratively. Specifically, as shown in Figure 9, it includes the following steps:
[0160] Step 801: Based on the position information of each of the two centroids, construct triangular facets using the position information of the vertices associated with each centroid.
[0161] In this embodiment, the triangular facets are the basic units that constitute a three-dimensional surface model, used to approximate or fit the actual distribution of the point cloud; the positions of the vertices correspond to the positions of the centroids, that is, when the centroid A is located in the first part of the current node and the centroid B is located in the second part of the current node, the vertices in the first part and the centroid A are used to jointly construct the triangular facets of the first part, and the vertices in the second part and the centroid B are used to jointly construct the triangular facets of the second part of the current node.
[0162] In some embodiments, before constructing triangular patches using the centroid and its associated vertices at the encoding end, it is necessary to determine the face vertices of the current node. Then, the vertices and face vertices within each part of the current node are sorted, for example, by one of the x, y, and z coordinates of the vertices. Subsequently, the sorted edge vertices and face vertices are used with the centroid of that part to construct the triangular patches of that part. The aforementioned face vertices are some vertices within the current node. During the process of determining the face vertices, a face vertex existence flag is generated. This existence flag is used to indicate whether the current node has face vertices. Specifically, if the value of the face vertex existence flag is 1, it means that the current node has face vertices; if the value of the face vertex existence flag is 0, it means that the current node does not have face vertices. The aforementioned face vertex existence flag is also encoded into a bit stream for transmission.
[0163] Step 802: Reconstruct the point cloud based on the triangular facets.
[0164] Based on step 801, ray tracing sampling is performed on the constructed triangular facets to obtain the reconstructed point cloud. This ray tracing sampling is based on the principle of ray tracing: light rays originate from a source, travel in a straight line, and intersect with objects in the scene. At the intersection, the light rays undergo reflection, refraction, or absorption depending on the object's material properties and lighting conditions. By recording these interactions, the geometric shape, surface features, and lighting information of objects in the scene can be obtained. Applying ray tracing sampling technology to point cloud reconstruction can improve the accuracy and realism of the reconstructed point cloud.
[0165] In this embodiment, before encoding begins at the encoding end, regardless of whether the current node meets the first dual-centrifugal technology condition, it is ensured that the dual-centrifugal technology is enabled. Then, the condition of whether the current node meets the first dual-centrifugal technology condition is determined. The second identifier of the current node indicates that the dual-centrifugal technology is enabled at the encoding end, meaning the encoding end uses dual-centrifugal technology. This second identifier also needs to be added to the bitstream for transmission, thereby instructing the decoding end to also use dual-centrifugal technology.
[0166] In this embodiment, the second identifier of the current node can also be added to the geometric parameter set for use by both the encoder and decoder. The geometric parameter set can be an existing set used to define Trisoup encoding parameters, which, in addition to the second identifier, may also include parameters such as the size of the current node, the size of the base node, and the incremental offset value. Adding the second identifier to the geometric parameter set requires minimal modification to existing protocols and maintains good compatibility with existing technologies.
[0167] Corresponding to the embodiments executed at the encoding end shown in Figures 6-9 above, this application also provides embodiments executed at the decoding end. Referring to Figure 10, this application provides a three-dimensional reconstruction method that can be applied at the decoding end. As shown in Figure 10, the three-dimensional reconstruction method includes the following steps:
[0168] Step 901: When it is determined that the current node satisfies the second dual centroid technical condition, obtain the vertices associated with each centroid in the two centroids;
[0169] In this embodiment, the aforementioned second dual-centroid technology condition is a pre-set criterion used to determine whether the current node needs to employ dual-centroid technology. The aforementioned vertex refers to the intersection of the point cloud surface with the node's edge, while the vertex associated with the centroid refers to the vertex used to determine the centroid of a certain region. This vertex can be an edge vertex, i.e., a vertex located on the edge of the current node. The aforementioned acquisition of the vertex associated with each of the two centroids specifically refers to acquiring the position information of the vertex associated with each of the two centroids, such as the vertex coordinates. This position information can be obtained by decoding the bitstream obtained from the decoding end. Specifically, it is necessary to first determine the neighbor information of each edge of the current node for subsequent entropy decoding, i.e., to decode the vertex existence flag Flag0 and quantized vertex position information of each edge of the current node based on the determined neighbor information of the node's edges. The aforementioned quantized vertex position information only exists when the vertex exists. Thus, determining the existence of vertices before performing dual-centroid technology and decoding the relevant valid information provides reliable data support for subsequent point cloud reconstruction.
[0170] Step 902: Determine the position information of each centroid in the two centroids based on the position information of the vertices associated with each centroid in the two centroids;
[0171] Based on step 901, the position information of each centroid is determined by using the position information of the vertices associated with each centroid of the two centroids. For example, when the point cloud distribution within a node is uneven, the coordinates of the vertices associated with each centroid and their corresponding weights are used to perform a weighted average calculation to obtain the coordinates of the centroid.
[0172] Step 903: Perform three-dimensional reconstruction based on the position information of each of the two centroids.
[0173] Based on step 902, the three-dimensional reconstruction, i.e. point cloud reconstruction, is completed at the decoding end according to the determined position information of each centroid.
[0174] In this embodiment, after the decoding end determines that the current node meets the second dual-centroid technical condition, the vertex associated with each centroid of the two centroids is obtained. Then, the position information of each centroid is determined using the position information of the aforementioned vertices. Finally, 3D reconstruction is performed based on the determined position information of each centroid. In this way, the dual centroids can better depict the distribution of the point cloud inside the node, thereby effectively avoiding the distortion problem that exists when using a single centroid for point cloud reconstruction, and thus improving the accuracy and detail representation of point cloud reconstruction.
[0175] In this embodiment, when determining whether the current node meets the second dual-centroid technical condition at the decoding end, the determination can be made using the first identifier of the current node obtained from the bitstream. For example, as described in the above embodiment, the first identifier can be Flag1 = 1. If received, it is determined that the current node meets the second dual-centroid technical condition, and point cloud reconstruction is performed using dual centroids. If Flag1 = 0 is received, it is determined that the current node does not meet the second dual-centroid technical condition, and the encoding end has not performed dual-centroid technology. In this case, other methods can be used for point cloud reconstruction, such as using a method based on the original centroids. The first identifier Flag1 = 1 is obtained by decoding the bitstream sent from the encoding end. Alternatively, Flag1 = 0 can also be obtained by decoding the bitstream sent from the encoding end.
[0176] In some embodiments, eligibility criteria for each node to perform the dual centroid technique can be defined, and an eligibility determination operation can be performed. Then, the operation of judging using the first identifier of the current node obtained from the bitstream is only performed if the current node meets the eligibility criteria for performing the dual centroid technique. That is, before using the first identifier of the current node obtained from the bitstream, the eligibility criteria for the current node are judged to determine whether the current node meets the eligibility criteria for performing the dual centroid technique. If the current node meets the eligibility criteria for performing the dual centroid technique, and the first identifier of the current node obtained from the bitstream satisfies the second dual centroid technique condition, it indicates that the current node can perform 3D reconstruction using the centroid position information determined by the position information of the vertices associated with each centroid. Otherwise, 3D reconstruction is performed using existing methods, such as the above-described 3D reconstruction method based on the original centroid, and it is not necessary to obtain the first identifier of the current node from the bitstream.
[0177] In this scenario, for the encoding end, the first identifier is generated and added to the bitstream only when a node meets the eligibility criteria for executing the dual-centroid technique. For nodes that do not meet the eligibility criteria, the first identifier Flag1=1 is not sent, nor is it necessary to send Flag1=0. Similarly, for the decoding end, the first identifier is only retrieved from the bitstream when a node meets the eligibility criteria. This embodiment reduces the amount of data transmitted between the encoding and decoding ends, and lowers encoding redundancy and cost.
[0178] In this embodiment of the application, the qualification determination operation of the current node can be performed based on various information of the current node, specifically including at least one of the following operations:
[0179] Determine that the size of the current node is greater than or equal to the first preset size;
[0180] Determine if the number of vertices in the current node is greater than or equal to the first preset number;
[0181] Determine if the vertex distribution information of the current node meets the preset distribution conditions;
[0182] The centroid offset information within the node is determined to meet the preset conditions.
[0183] Optionally, the size information of the current node includes the node width, which represents the size of the cube space represented by the current node. In this application, it can represent the size of the space occupied by the current node in the original point cloud. The node size information needs to be set according to the actual encoding requirements and scenario to achieve better point cloud reconstruction effect and performance. The first preset size is used to filter nodes that can be used for the dual centroid technique to better meet the needs of point cloud reconstruction. For example, the node width in the first preset size information can be set to 8 to meet the needs of point cloud reconstruction. The first preset quantity is used to filter nodes suitable for the dual centroid technique based on the number of vertices. That is, only when the number of vertices in a node reaches at least the preset value can dual centroid be used for point cloud reconstruction. The setting of this parameter is used to ensure that the current node has enough vertices to achieve a high-quality point cloud reconstruction effect, thereby avoiding the distortion problem of point cloud reconstruction caused by too few vertices in the node.
[0184] The vertex distribution information of the current node described above describes the distribution of vertices within the current node. The preset distribution conditions are used to filter nodes suitable for the dual-centroid technique based on the vertex distribution within the current node. Setting this parameter ensures that the selected nodes have dual centroids, and point cloud reconstruction is performed based on these centroids. The centroid offset information within the node indicates whether the current node has a centroid offset; the preset conditions are used to filter nodes suitable for the dual-centroid technique based on the centroid offset information within the node.
[0185] In the embodiments of this application, the first preset size, the first preset quantity, the preset distribution conditions, and the preset conditions mentioned above can be preset, or they can be obtained by the encoding end adjusting the preset values or preset conditions according to the adjustment parameters, or they can be determined according to the bit rate. That is, when the bit rate is different, the corresponding second preset range value will also be different.
[0186] In some embodiments, when performing a qualification determination operation based on the size information of the current node, it can be determined by judging whether the size information of the current node reaches a first preset size. Specifically, it can be determined whether the node width of the current node is greater than or equal to a first preset value. If the node width of the current node is greater than or equal to the first preset value, it is determined that the current node meets the qualification conditions for performing the dual centroid technique and can be used for point cloud reconstruction using dual centroids; otherwise, it is considered that the current node does not meet the qualification conditions for the dual centroid technique. In this case, other methods can be used for point cloud reconstruction, such as using the single centroid point cloud reconstruction method described above. For example, when the first preset value th1 = 8 is set, if the node width of the current node is greater than or equal to 8, it is determined that it can perform the dual centroid technique; if the node width of the current node is less than 8, point cloud reconstruction is performed using the method based on the original centroid.
[0187] In some embodiments, when performing a qualification determination operation based on the vertex distribution information of the current node, it can be determined by judging whether the vertex distribution information of the current node meets a preset distribution condition. Specifically, it can be judged whether the vertices of the current node can be divided into two parts in the principal axis direction, and whether the number of vertices in each part is greater than or equal to a second preset number. In this case, the following situations are included:
[0188] (1) When the vertex of the current node can be divided into two parts in the direction of the main axis, and the number of vertices in each part is greater than or equal to the second preset number, for example, when the current node is the node shown in Figure 5, it is determined that it can perform the double centroid technique and the centroid of each part is calculated.
[0189] (2) If the vertex of the current node cannot be divided into two parts in the direction of the main axis, other methods can be used, such as the single centroid method, to complete the point cloud reconstruction.
[0190] (3) If the vertices of the current node can be divided into two parts in the main axis direction, but the number of vertices in each part is less than the second preset number, then the current node is considered not to meet the qualification conditions for performing the dual centroid technique. In this case, other methods can be used, such as the original centroid method, to complete the point cloud reconstruction.
[0191] In this way, by performing qualification determination operations based on the vertex distribution information of the current node at the decoding end, the point cloud distribution within the current node can be better characterized by the dual centroids, thereby reducing the distortion of the reconstructed point cloud.
[0192] In this embodiment of the application, when it is determined that the vertex of the current node can be divided into two parts along the principal axis direction, the principal axis direction can be determined first, and then the vertex can be divided according to the principal axis direction. Specifically, as shown in Figure 11, the following steps are included:
[0193] Step 1001: Determine the main axis direction of the current node;
[0194] In this embodiment, the principal axis direction can be either the direction of the sum of the normal vectors of the triangular facets formed by the centroid and the vertices, or the direction of the coordinate axis components of the centroid. Specifically, the principal axis direction can be determined through the following process: First, calculate an original centroid using all vertices within the current node. Then, determine the arrangement order of the vertices along the x, y, and z directions respectively, and construct triangular facets together with the original centroid. Next, calculate the sum of the normal vectors of each triangular facet. Finally, select the direction with the largest sum of normal vector components as the principal axis direction.
[0195] Step 1002: Determine the vertices in the main axis direction whose coordinate values are greater than the coordinate values of the original centroid of the current node as the first part of vertices, and the vertices whose coordinate values are less than or equal to the coordinate values of the original centroid of the current node as the second part of vertices. The distance between the first vertex in the first part of vertices and the second vertex in the second part of vertices is greater than a first preset distance. The first vertex is the vertex in the first part of vertices that is closest to the original centroid, and the second vertex is the vertex in the second part of vertices that is closest to the original centroid.
[0196] Based on step 1001, along the main axis, the coordinates of the vertices and the original centroid are calculated according to the position information of all vertices within the current node and the original centroid. Then, these two types of coordinates are compared. If the coordinates of a vertex are greater than the coordinates of the original centroid, it is classified as a first part of vertices. If the coordinates of a vertex are less than or equal to the coordinates of the original centroid, it is classified as a second part of vertices. Furthermore, if the distance between the vertex closest to the original centroid in the first part and the vertex closest to the original centroid in the second part is greater than a first preset distance, then the vertex division is considered valid. The first preset distance is used to determine whether the vertices of the current node can be divided into two parts along the main axis.
[0197] In this embodiment, by comprehensively considering information such as the principal axis direction, vertex position, centroid position, and distance between vertices, it is possible to determine whether the current node's vertex can be divided into two parts along the principal axis direction. This improves the accuracy and reliability of the judgment, thereby effectively reducing the degree of distortion during point cloud reconstruction and improving the reconstruction effect of the point cloud.
[0198] In some embodiments, when performing a qualification determination operation based on the centroid offset information within the current node, it is necessary to determine whether the centroid offset information within the current node meets preset conditions. Specifically, it is necessary to determine whether there is an offset within the centroid of the current node. This includes the following situations:
[0199] (1) If there is no offset of the centroid within the current node, then the offset information of the centroid within the current node meets the preset conditions, and subsequent operations can be performed.
[0200] (2) If the centroid of the current node is offset but the centroid offset value is zero, then the centroid offset information of the current node meets the preset conditions and subsequent operations can be performed.
[0201] (3) If the centroid of the current node is offset and the centroid offset value is not zero, it is determined that the centroid offset information of the current node does not meet the preset conditions. At this time, other methods can be used to reconstruct the point cloud, such as using the original centroid to complete the point cloud reconstruction.
[0202] In this way, by determining whether the centroid offset information within a node meets the preset conditions, it is possible to determine whether the current node meets the eligibility criteria for performing the dual-centroid technique. This allows for a more accurate assessment of whether the current node is suitable for point cloud reconstruction using dual centroids, thereby reducing unnecessary resource waste.
[0203] In this embodiment, when performing a qualification determination operation on the current node, one type of information or a combination of multiple types of information can be used to perform the qualification determination operation according to actual needs. In some embodiments, in order to improve the point cloud reconstruction effect, the size information, number of vertices, distribution information of vertices, and centroid offset information within the current node can be comprehensively considered to perform the qualification determination operation on the current node. If the current node meets the corresponding qualification conditions in all of the above qualification determination operations, it is considered that the node has the qualification conditions for performing the dual centroid technique.
[0204] In this embodiment, the decoding end can determine the above-mentioned situations based on the bit stream sent by the encoding end. For example, by decoding the bit stream, the centroid offset value is obtained. If the centroid offset value is zero, it is determined that the current node meets the qualification conditions for performing the dual centroid technique, and point cloud reconstruction can be performed using dual centroids. If the centroid offset value is not zero, it is determined that the current node does not meet the qualification conditions for performing the dual centroid technique. In this case, point cloud reconstruction can be performed using other methods, such as single centroid point cloud reconstruction. In this embodiment, when the decoding end performs 3D reconstruction based on the position information of each centroid in the two centroids, it also needs to combine the position information of the vertices associated with each centroid to perform 3D reconstruction collaboratively. Specifically, as shown in Figure 12, it includes the following steps:
[0205] Step 1101: Based on the position information of each of the two centroids, construct triangular facets using the position information of the vertices associated with each centroid.
[0206] In this embodiment, the triangular facets are the basic units that constitute a three-dimensional surface model, used to approximate or fit the actual distribution of the point cloud; the positions of the vertices correspond to the positions of the centroids, that is, when the centroid A is located in the first part of the current node and the centroid B is located in the second part of the current node, the vertices in the first part and the centroid A are used to jointly construct the triangular facets of the first part, and the vertices in the second part and the centroid B are used to jointly construct the triangular facets of the second part of the current node.
[0207] In some embodiments, before constructing a triangular facet using the centroid and its associated vertices at the decoding end, it is necessary to determine the face vertices of the current node. Then, the vertices and face vertices within each part of the current node are sorted, for example, by one of the x, y, and z coordinates of the vertices. Subsequently, the sorted edge vertices and face vertices are used with the centroid of that part to construct the triangular facet of that part. The face vertices can be determined based on a face vertex presence identifier, which can be obtained by decoding the bitstream sent by the encoding end.
[0208] Step 1102: Reconstruct the point cloud based on the triangular facets.
[0209] Based on step 1101, ray tracing sampling is performed on the constructed triangular facets to obtain the reconstructed point cloud. This ray tracing sampling is based on the principle of ray tracing, where light rays originate from a source, travel along a straight line, and intersect with objects in the scene. At the intersection, the light rays undergo reflection, refraction, or absorption depending on the object's material properties and lighting conditions. By recording these interactions, the geometric shape, surface features, and lighting information of objects in the scene can be obtained. Applying ray tracing sampling technology to point cloud reconstruction can improve the accuracy and realism of the reconstructed point cloud.
[0210] In this embodiment, before using dual centroids for point cloud reconstruction at the decoding end, regardless of whether the current node meets the second dual centroid technology condition, it is necessary to ensure that the dual centroid technology is enabled before determining whether the current node meets the second dual centroid technology condition. The second identifier of the current node indicates that the dual centroid technology is enabled at the encoding end, meaning that the encoding end uses the dual centroid technology. This second identifier can be obtained from the bitstream sent by the encoding end, indicating that the decoding end also needs to use the dual centroid technology.
[0211] In this embodiment, the second identifier of the current node can also be added to the geometric parameter set for use by the decoding end. The geometric parameter set can be an existing set used to define Trisoup encoding-related parameters, which, in addition to the second identifier, may also include parameters such as the size of the current node, the size of the base node, and the incremental offset value. Adding the second identifier to the geometric parameter set requires minimal modification to existing protocols and maintains good compatibility with existing technologies.
[0212] In this embodiment, the aforementioned 3D reconstruction method can split a reconstructed surface that should not be continuous into two reconstructed surfaces that better conform to the original point cloud distribution, reducing the number of reconstructed point cloud points and lowering the bitrate. Under lossy encoding and decoding conditions, a comparative analysis of the aforementioned 3D reconstruction method at both the decoding and encoding ends is performed on the same dataset, compared with existing methods. The comparison results show that the 3D reconstruction method provided in this application has improved performance compared to existing methods. While the encoding and decoding time is slightly increased, it remains within a reasonable range. Specifically, under the evaluation measures of point-to-point and point-to-surface, the encoding and decoding time of the aforementioned 3D reconstruction method is basically unchanged compared to existing methods. Furthermore, the BD-rate parameter value indicates that, at the same quality, the aforementioned 3D reconstruction method can achieve comparable results to existing methods using fewer bits. The BD-rate parameter is a performance indicator used to represent the bitrate saved at the same quality.
[0213] Furthermore, to further verify the effectiveness of the 3D reconstruction method provided in this application, the above method was extended and applied by improving a simplified geometry-based point cloud compression method. The method was then validated on different types of data. The results show that the improved simplified geometry-based point cloud compression method provided in this application achieves a certain performance gain compared to the unimproved method, while maintaining essentially the same encoding and decoding time. Specifically, the performance was improved in both point-to-point and point-to-surface evaluation metrics. In other words, the encoding and decoding time remains essentially the same compared to the unimproved method. Furthermore, based on the BD-rate parameter value, it can achieve comparable results to the unimproved method using fewer bits at the same quality.
[0214] The three-dimensional reconstruction method provided in this application can be executed by a three-dimensional reconstruction device. This application uses a three-dimensional reconstruction device to perform the three-dimensional reconstruction method as an example to illustrate the three-dimensional reconstruction device provided in this application.
[0215] Specifically, referring to Figure 13, when the 3D reconstruction device is a device applied to the encoding end, the 3D reconstruction device 1200 includes a first vertex acquisition module 1201, used to acquire vertices associated with each centroid of the two centroids when it is determined that the current node satisfies the first dual centroid technical condition; a first centroid position information determination module 1202, used to determine the position information of each centroid of the two centroids based on the position information of the vertices associated with each centroid of the two centroids; and a first 3D reconstruction module 1203, used to perform 3D reconstruction based on the position information of each centroid of the two centroids.
[0216] In some embodiments, determining that the first dual-centroid technical condition is met includes:
[0217] The number of point clouds within a second preset distance from the original centroid of the current node is determined to be less than or equal to a first number.
[0218] In some embodiments, before determining that the number of point clouds within a second preset distance of the original centroid of the current node is less than or equal to a first number, the method further includes:
[0219] The first identifier addition module is used to generate a first identifier for the current node and add the first identifier of the current node to the bitstream. The first identifier is used to instruct the current node to perform the dual centroid technique.
[0220] In some embodiments, determining that the current node satisfies the first dual-centroid technical condition includes:
[0221] Perform a qualification check on the current node to determine if it meets the eligibility criteria for executing the dual-centroid technique.
[0222] In some embodiments, the qualification determination operation for the current node includes at least one of the following:
[0223] Determine that the size of the current node is greater than or equal to the first preset size;
[0224] Determine if the number of vertices in the current node is greater than or equal to the first preset number;
[0225] Determine if the vertex distribution information of the current node meets the preset distribution conditions;
[0226] The centroid offset information within the node is determined to meet the preset conditions.
[0227] In some embodiments, it also includes:
[0228] The second identifier addition module is used to encode the second identifier into the bitstream, and the second identifier is used to indicate the use of dual centroid technology.
[0229] By using the aforementioned 3D reconstruction device to reconstruct point clouds at the encoding end, the distribution of point clouds within the current node can be better fitted using dual centroids. This effectively avoids the problem of significant distortion when reconstructing point clouds based on a single centroid, thereby providing a more accurate 3D model for subsequent applications.
[0230] The three-dimensional reconstruction device provided in this application embodiment can realize the various processes implemented in the method embodiments of Figures 6 to 9 and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0231] Referring to Figure 14, when the 3D reconstruction device is applied to the decoding end, the 3D reconstruction device 1300 includes a second vertex acquisition module 1301, used to acquire vertices associated with each centroid in the two centroids when it is determined that the current node satisfies the second dual centroid technical condition; a second centroid position information determination module 1302, used to determine the position information of each centroid in the two centroids based on the position information of the vertices associated with each centroid in the two centroids; and a second 3D reconstruction module 1303, used to perform 3D reconstruction based on the position information of each centroid in the two centroids.
[0232] In some embodiments, determining that the second dual-centroid technical condition is met includes:
[0233] The first identifier of the current node is determined from the bitstream. This first identifier is used to instruct the current node to perform the dual centroid technique.
[0234] In some embodiments, before determining the first identifier of the current node from the bitstream, the process includes:
[0235] Perform a qualification check on the current node to determine if it meets the eligibility criteria for executing the dual-centroid technique.
[0236] In some embodiments, the qualification determination operation for the current node includes at least one of the following:
[0237] Determine that the size of the current node is greater than or equal to the first preset size;
[0238] Determine if the number of vertices in the current node is greater than or equal to the first preset number;
[0239] Determine if the vertex distribution information of the current node meets the preset distribution conditions;
[0240] The centroid offset information within the node is determined to meet the preset conditions.
[0241] In some embodiments, determining that the centroid offset information within a node satisfies a preset condition includes:
[0242] There is no centroid offset within the current node; or, the current node has a centroid offset, but the centroid offset value is zero.
[0243] In some embodiments, before determining that the current node satisfies the second dual-centroid technical condition, the method further includes:
[0244] A second identifier is obtained from the bitstream, which is used to indicate the use of dual centroid technology.
[0245] By using the aforementioned 3D reconstruction device to reconstruct point clouds at the decoding end, the distribution of point clouds within the current node can be better fitted using dual centroids. This effectively avoids the problem of significant distortion when reconstructing point clouds based on a single centroid, thereby providing a more accurate 3D model for subsequent applications.
[0246] The three-dimensional reconstruction device provided in this application embodiment can realize the various processes implemented in the method embodiments of Figures 10 to 12 and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0247] As shown in Figure 15, this application embodiment also provides an electronic device 1400, including a processor 1401 and a memory 1402. The memory 1402 stores programs or instructions that can run on the processor 1401. For example, when the electronic device 1400 is an encoding device, the program or instructions executed by the processor 1401 implement the various steps of any of the three-dimensional reconstruction method embodiments described in Figures 6-9, and achieve the same technical effect. When the electronic device 1400 is a decoding device, the program or instructions executed by the processor 1401 implement the various steps of any of the three-dimensional reconstruction method embodiments described in Figures 10-12, and achieve the same technical effect. To avoid repetition, these steps will not be repeated here. Optionally, the memory 1402 can be the memory 102 or memory 113 in the embodiment shown in Figure 1, and the processor 1401 can implement the functions of the encoder 200 or decoder 300 in the embodiments shown in Figures 1-3.
[0248] This application also provides an electronic device, including: a memory configured to store point cloud data; and a processing circuit configured to implement the various steps of the three-dimensional reconstruction method embodiments as described in any of Figures 6-9 or 10-12 above. Optionally, the memory may be memory 102 or memory 113 in the embodiment shown in Figure 1, and the processing circuit may implement the functions of encoder 200 or decoder 300 in the embodiments shown in Figures 1-3.
[0249] This application also provides an electronic device, including a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps in any of the method embodiments shown in Figures 6-9 or 10-12. This device embodiment corresponds to the above method embodiments, and all implementation processes and methods of the above method embodiments can be applied to this terminal embodiment and achieve the same technical effect.
[0250] The processor or processing circuit in the embodiments of this application may include general-purpose processors, special-purpose processors, etc., such as central processing units (CPUs), microprocessors, digital signal processors (DSPs), artificial intelligence (AI) processors, graphics processing units (GPUs), application-specific integrated circuits (ASICs), network processors (NPs), field-programmable gate arrays (FPGAs), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc. The communication interface in the embodiments of this application may include transceivers, pins, circuits, buses, etc.
[0251] The aforementioned electronic devices can be terminals or other devices besides terminals, such as servers, network attached storage (NAS), etc.
[0252] Among them, the terminal can also be called user equipment (UE), which can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, mixed reality (MR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipborne equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication functions, such as refrigerators, televisions, washing machines or furniture, etc.), game console, personal computer (PC), ATM or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the embodiments in this application do not limit the specific type of terminal.
[0253] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server. A cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.
[0254] For example, the aforementioned electronic device may include, but is not limited to, the type of source device 100 or destination device 110 shown in FIG1.
[0255] Taking an electronic device as an example, Figure 16 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of this application.
[0256] The terminal 1500 includes, but is not limited to, at least some of the following components: radio frequency unit 1501, network module 1502, audio output unit 1503, input unit 1504, sensor 1505, display unit 1506, user input unit 1507, interface unit 1508, memory 1509, and processor 1510.
[0257] Those skilled in the art will understand that the terminal 1500 may also include a power supply (such as a battery) for powering various components. The power supply can be logically connected to the processor 1510 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The terminal structure shown in Figure 16 does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0258] It should be understood that, in this embodiment, the input unit 1504 may include a graphics processor 15041 and a microphone 15042. The graphics processor 15041 processes image data of still images or videos obtained by an image acquisition device (such as a camera) in video acquisition mode or image acquisition mode, or it may process the obtained point cloud data. The display unit 1506 may include a display panel 15061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1507 includes at least one of a touch panel 15071 and other input devices 15072. The touch panel 15071 is also called a touch screen. The touch panel 15071 may include a touch detection device and a touch controller. Other input devices 15072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0259] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 1501 can transmit it to the processor 1510 for processing; in addition, the radio frequency unit 1501 can send uplink data to the network-side device. Typically, the radio frequency unit 1501 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0260] The memory 1509 can be used to store software programs or instructions, as well as various data. The memory 1509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1509 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0261] Processor 1510 may include one or more processing units; optionally, processor 1510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1510.
[0262] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant descriptions of the three-dimensional reconstruction methods described in any of Figures 6-9 or Figures 10-12 in the method embodiment, and achieve the same or corresponding technical effects. To avoid repetition, it will not be described again here.
[0263] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of any of the three-dimensional reconstruction method embodiments described in Figures 6-9 or 10-12 above, and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0264] The processor mentioned above is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.
[0265] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of any of the three-dimensional reconstruction method embodiments described in Figures 6-9 or 10-12 above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0266] It should be understood that the chips mentioned in the embodiments of this application may include system-on-a-chip (also known as system chip, chip system, or system-on-a-chip) or discrete display chips, etc.
[0267] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described three-dimensional reconstruction method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0268] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0269] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.), and the computer software product includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0270] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. A three-dimensional reconstruction method, applied at the encoding end, comprising: When it is determined that the current node satisfies the first dual-centroid technical condition, obtain the vertices associated with each centroid among the two centroids; Based on the position information of the vertices associated with each of the two centroids, determine the position information of each centroid in the two centroids; Three-dimensional reconstruction is performed based on the position information of each of the two centroids.
2. The method of claim 1, wherein, The determination that the first dual-centroid technical condition is met includes: The number of point clouds within a second preset distance from the original centroid of the current node is determined to be less than or equal to a first number.
3. The method of claim 2, wherein, The first quantity is a preset value, or the first quantity is obtained by adjusting the preset value according to the adjustment parameters.
4. The method of claim 3, wherein, The adjustment parameters include at least one of the following: average point cloud distance, vertex search distance, point cloud density parameter, and point cloud sparsity parameter.
5. The method of claim 2, wherein, Also includes: A first identifier for the current node is generated and added to the bitstream. The first identifier is used to instruct the current node to perform the dual centroid technique.
6. The method of any one of claims 2-5, wherein, Before determining that the number of point clouds within a second preset distance of the original centroid of the current node is less than or equal to the first number, the method further includes: Perform a qualification check on the current node to determine if it meets the eligibility criteria for executing the dual-centroid technique.
7. The method of claim 6, wherein, The qualification determination operation for the current node includes at least one of the following: Determine that the size of the current node is greater than or equal to the first preset size; Determine if the number of vertices in the current node is greater than or equal to the first preset number; Determine if the vertex distribution information of the current node meets the preset distribution conditions; The centroid offset information within the node is determined to meet the preset conditions.
8. The method of claim 7, wherein, Determining that the current node's size information reaches the first preset size includes: Determine that the width of the current node is greater than or equal to the first preset value.
9. The method of claim 7, wherein, The determination that the vertex distribution information of the current node satisfies the preset distribution conditions includes: Determine if the vertices of the current node can be divided into two parts along the main axis, and if the number of vertices in each part is greater than or equal to a second preset number.
10. The method of claim 9, wherein, The determination that the vertex of the current node can be divided into two parts along the principal axis includes: Determine the main axis direction of the current node; In the direction of the main axis, vertices whose coordinate values are greater than the coordinate values of the original centroid of the current node are identified as the first set of vertices, and vertices whose coordinate values are less than or equal to the coordinate values of the original centroid of the current node are identified as the second set of vertices. The distance between the first vertex in the first set of vertices and the second vertex in the second set of vertices is greater than a first preset distance. The first vertex is the vertex in the first set of vertices that is closest to the original centroid, and the second vertex is the vertex in the second set of vertices that is closest to the original centroid.
11. The method of claim 7, wherein, The determination of the centroid offset information within the node satisfies preset conditions, including: There is no centroid offset within the current node; or, the current node has a centroid offset, but the centroid offset value is zero.
12. The method of any one of claims 1-11, wherein, The three-dimensional reconstruction based on the position information of each of the two centroids includes: Based on the position information of each of the two centroids, triangular facets are constructed using the position information of the vertices associated with each centroid. Point cloud reconstruction is performed based on the triangular facets.
13. The method of any one of claims 1-12, wherein, Also includes: A second identifier is added to the bitstream, which is used to indicate the use of dual centroid technology.
14. The method of claim 13, wherein, Adding the second identifier to the bitstream includes: Add the second identifier to the geometry parameter set.
15. A three-dimensional reconstruction method, applied at the decoding end, comprising: When it is determined that the current node satisfies the second dual centroid technical condition, obtain the vertices associated with each centroid in the two centroids; Based on the position information of the vertices associated with each of the two centroids, determine the position information of each centroid in the two centroids; Three-dimensional reconstruction is performed based on the position information of each of the two centroids.
16. The method of claim 15, wherein, The determination that the current node satisfies the second dual centroid technical condition includes: The first identifier of the current node is determined from the bitstream. The first identifier is used to instruct the current node to perform the dual centroid technique.
17. The method of claim 16, wherein, Before determining the first identifier of the current node obtained from the bitstream, the method further includes: Perform a qualification check on the current node to determine if it meets the eligibility criteria for executing the dual-centroid technique.
18. The method of claim 17, wherein, The qualification determination operation for the current node includes at least one of the following: Determine that the size of the current node is greater than or equal to the first preset size; Determine if the number of vertices in the current node is greater than or equal to the first preset number; Determine if the vertex distribution information of the current node meets the preset distribution conditions; The centroid offset information within the node is determined to meet the preset conditions.
19. The method of claim 18, wherein, Determining that the size information of the current node is greater than or equal to the first preset size includes: Determine that the width of the current node is greater than or equal to the first preset value.
20. The method of claim 18, wherein, The determination that the vertex distribution information of the current node satisfies the preset distribution conditions includes: Determine if the vertices of the current node can be divided into two parts along the main axis, and if the number of vertices in each part is greater than or equal to a second preset number.
21. The method of claim 20, wherein, The determination that the vertex of the current node can be divided into two parts along the principal axis includes: Determine the main axis direction of the current node; In the direction of the main axis, vertices whose coordinate values are greater than the coordinate values of the original centroid of the current node are identified as the first set of vertices, and vertices whose coordinate values are less than or equal to the coordinate values of the original centroid of the current node are identified as the second set of vertices. The distance between the first vertex in the first set of vertices and the second vertex in the second set of vertices is greater than a first preset distance. The first vertex is the vertex in the first set of vertices that is closest to the original centroid, and the second vertex is the vertex in the second set of vertices that is closest to the original centroid.
22. The method of claim 18, wherein, The determination of the centroid offset information within the node satisfies preset conditions, including: There is no centroid offset within the current node; or, the current node has a centroid offset, but the centroid offset value is zero.
23. The method of any one of claims 15-22, wherein, The three-dimensional reconstruction based on the position information of each of the two centroids includes: Based on the position information of each of the two centroids, triangular facets are constructed using the position information of the vertices associated with each centroid. Point cloud reconstruction is performed based on the triangular facets.
24. The method of any one of claims 15-23, wherein, Before determining that the current node satisfies the second dual-centroid technical condition, the following steps are also included: A second identifier is obtained from the bitstream, which is used to indicate the use of dual centroid technology.
25. The method of claim 24, wherein, The determination of obtaining the second identifier from the bitstream includes: Determine that the second identifier is obtained from the set of geometric parameters.
26. A three-dimensional reconstruction device, applied at an encoding end, comprising: The first vertex acquisition module is used to acquire the vertex associated with each of the two centroids when it is determined that the current node satisfies the first dual centroid technical condition; The first centroid position information determination module is used to determine the position information of each centroid in the two centroids based on the position information of the vertex associated with each centroid in the two centroids; The first three-dimensional reconstruction module is used to perform three-dimensional reconstruction based on the position information of each of the two centroids.
27. The apparatus of claim 26, wherein, The determination that the first dual-centroid technical condition is met includes: The number of point clouds within a second preset distance from the original centroid of the current node is determined to be less than or equal to a first number.
28. The apparatus of claim 27, wherein, Also includes: The first identifier addition module is used to generate a first identifier for the current node and add the first identifier of the current node to the bitstream. The first identifier is used to instruct the current node to perform the dual centroid technique.
29. The apparatus of claim 27 or 28, wherein, Before determining that the number of point clouds within a second preset distance of the original centroid of the current node is less than or equal to a first number, the following steps are included: Perform a qualification check on the current node to determine if it meets the eligibility criteria for executing the dual-centroid technique.
30. The apparatus of claim 29, wherein, The qualification determination operation for the current node includes at least one of the following: Determine that the size of the current node is greater than or equal to the first preset size; Determine if the number of vertices in the current node is greater than or equal to the first preset number; Determine if the vertex distribution information of the current node meets the preset distribution conditions; The centroid offset information within the node is determined to meet the preset conditions.
31. The apparatus of any of claims 26-30, wherein, Also includes: The second identifier addition module is used to encode the second identifier into the bitstream, and the second identifier is used to indicate the use of dual centroid technology.
32. A three-dimensional reconstruction device, applied at a decoding end, comprising: The second vertex acquisition module is used to acquire the vertex associated with each of the two centroids when it is determined that the current node satisfies the second dual centroid technical condition. The second centroid position information determination module is used to determine the position information of each centroid in the two centroids based on the position information of the vertex associated with each centroid in the two centroids; The second 3D reconstruction module is used to perform 3D reconstruction based on the position information of each of the two centroids.
33. The apparatus of claim 32, wherein, The determination that the second dual-centroid technical condition is met includes: The first identifier of the current node is determined from the bitstream. The first identifier is used to instruct the current node to perform the dual centroid technique.
34. The apparatus of claim 33, wherein, Before determining the first identifier of the current node obtained from the bitstream, the method further includes: Perform a qualification check on the current node to determine if it meets the eligibility criteria for executing the dual-centroid technique.
35. The apparatus of claim 34, wherein, The qualification determination operation for the current node includes at least one of the following: Determine that the size of the current node is greater than or equal to the first preset size; Determine if the number of vertices in the current node is greater than or equal to the first preset number; Determine if the vertex distribution information of the current node meets the preset distribution conditions; The centroid offset information within the node is determined to meet the preset conditions.
36. The apparatus of claim 35, wherein, The determination of the centroid offset information within the node satisfies preset conditions, including: There is no centroid offset within the current node; or, the current node has a centroid offset, but the centroid offset value is zero.
37. The apparatus of any of claims 32-36, wherein, Before determining that the current node satisfies the second dual-centroid technical condition, the following steps are also included: A second identifier is obtained from the bitstream, which is used to indicate the use of dual centroid technology.
38. An electronic device comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the three-dimensional reconstruction method as claimed in any one of claims 1 to 14, or implementing the steps of the three-dimensional reconstruction method as claimed in claims 15 to 25.
39. A readable storage medium storing a program or instructions that, when executed by a processor, implement the three-dimensional reconstruction method as claimed in any one of claims 1 to 14, or implement the steps of the three-dimensional reconstruction method as claimed in any one of claims 15 to 25.