Three-dimensional scene reconstruction method and device, equipment and storage medium

Through compression coding and weighted prediction technology, the storage demand for anchor feature attributes is reduced, and the problem of large storage occupancy in the anchor-based 3DGS representation method is solved, achieving more efficient storage and transmission.

CN120070731AActive Publication Date: 2025-05-30PENG CHENG LAB

Patent Information

Application Number
CN202510012498.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-30
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

In the structured 3DGS representation method based on anchor points, the storage of anchor feature attributes takes up a large amount, resulting in waste of storage space and transmission bandwidth.

Method used

By obtaining the viewing image of the target scene and sparse point clouds, the initial three-dimensional representation model based on the anchor is initialized, and the anchor attributes are compressed and encoded using the preset entropy model. Then, entropy decoding and weighted prediction are performed to reduce redundancy of feature attributes and reduce storage requirements through channel dimension grouping.

Benefits of technology

It effectively reduces the storage usage of anchor feature attributes, reduces transmission bandwidth, and improves the rate distortion performance of the three-dimensional representation model, while maintaining higher quality reconstruction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070731A_ABST
    Figure CN120070731A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a three-dimensional scene reconstruction method and device, equipment and a storage medium, and relates to the technical field of computer vision. The method comprises the following steps: acquiring an anchor-point-based initial three-dimensional representation model corresponding to a target scene, for each anchor point, carrying out compressed encoding on the attribute of the anchor point to obtain an encoding code stream corresponding to the target scene, carrying out entropy decoding on the encoding code stream to obtain a decoding feature attribute, a decoding size attribute and a decoding offset attribute, and carrying out decoding on the decoding feature attribute, the decoding size attribute and the decoding offset attribute; and performing weighted prediction according to the decoding feature attributes and the feature attribute mean value to obtain weighted feature attributes, performing channel dimension grouping on the weighted feature attributes to obtain rendering feature attributes, and finally generating a three-dimensional representation model of the target scene based on the rendering feature attributes, the decoding size attributes and the decoding offset attributes. The feature attributes of the anchor points are simplified through the weighted prediction process and the channel dimension grouping process, the storage space is reduced, and the rate distortion performance of the initial three-dimensional representation model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to a three-dimensional scene reconstruction method, apparatus, device, and storage medium. Background Art

[0002] As a three-dimensional scene representation method, the 3D Gaussian Splatting (3DGS) technology has surpassed the previous generation of scene representation technologies such as point clouds and Neural Radiance Fields (NeRF) in terms of training speed, rendering speed, and synthesis quality. 3DGS represents a three-dimensional scene as a set of Gaussian ellipsoids with local scene attributes (such as opacity, color, size, etc.), and synthesizes virtual views with the help of a fast and differentiable rasterization pipeline. However, 3DGS relies on numerous Gaussian ellipsoids to achieve high-quality view synthesis results, resulting in a large amount of data processed, complex calculation processes, and high storage and transmission costs.

[0003] In the related art, Scaffold-GS proposed an anchor-based structured 3DGS representation method. This method uses anchors to cluster Gaussian ellipsoids with similar attributes in a local area, and predicts the attributes of Gaussian points based on the attributes of the anchors such as offset and features, thereby achieving more efficient rendering. Among them, Scaffold-GS reduces the model size to a certain extent by clustering Gaussian ellipsoids into anchor attributes, but the storage requirements for the relevant features of its anchors are still large. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a three-dimensional scene reconstruction method, apparatus, device, and storage medium to reduce the storage occupancy of anchor features in the anchor-based structured 3DGS representation process.

[0005] To achieve the above objective, the first aspect of the embodiments of this application proposes a three-dimensional scene reconstruction method, including:

[0006] Obtain perspective images corresponding to a target scene from multiple perspectives respectively, and obtain a sparse point cloud based on the perspective images. Initialize an anchor-based initial three-dimensional representation model according to the sparse point cloud. The initial three-dimensional representation model includes multiple anchors and the anchor attributes corresponding to each anchor. The anchor attributes include feature attributes, size attributes, and position offsets;

[0007] For each anchor, use a preset entropy model to perform compression encoding on the anchor attributes to obtain a coded bitstream corresponding to the target scene;

[0008] Perform entropy decoding on the coded bitstream to obtain decoded feature attributes, decoded size attributes, and decoded offset attributes;

[0009] Perform weighted prediction based on the decoded feature attributes and the mean of the feature attributes predicted based on the preset entropy model to obtain weighted feature attributes, and group the weighted feature attributes in the channel dimension to obtain rendering feature attributes;

[0010] Generate a three-dimensional representation model of the target scene based on the rendering feature attributes, the decoded size attributes, and the decoded offset attributes.

[0011] In some embodiments, the preset entropy model is a context model, including a binary hash table and a first multi-layer perceptron. The method of using the preset entropy model to perform compression encoding on the anchor attributes to obtain the encoded bitstream corresponding to the target scene includes:

[0012] Perform an interpolation operation in the binary hash table using the coordinate values of the anchor to obtain context features;

[0013] Input the context features into the first multi-layer perceptron for data prediction to obtain probability distribution parameters and quantization step sizes;

[0014] Quantize the anchor attributes using the quantization step sizes to obtain quantized attributes, and perform entropy encoding on the quantized attributes using the probability distribution parameters to obtain the encoded bitstream.

[0015] In some embodiments, the method of performing entropy decoding on the encoded bitstream to obtain decoded feature attributes, decoded size attributes, and decoded offset attributes includes:

[0016] Perform entropy decoding on the encoded bitstream using the probability distribution parameters to obtain decoded quantization attributes, where the decoded quantization attributes include quantized feature attributes, quantized size attributes, and quantized position offsets;

[0017] Obtain the quantization step sizes corresponding to each of the decoded quantization attributes, multiply the quantized feature attributes by the quantization step sizes to obtain the decoded feature attributes, multiply the quantized size attributes by the quantization step sizes to obtain the decoded size attributes, and multiply the quantized position offsets by the quantization step sizes to obtain the decoded offset attributes.

[0018] In some embodiments, the method of performing weighted prediction based on the mean of the feature attributes and the decoded feature attributes to obtain weighted feature attributes includes:

[0019] Input the context features into a second multi-layer perceptron for data prediction to obtain feature channel fusion weights;

[0020] Perform weighted prediction based on the feature channel fusion weights, the mean of the feature attributes, and the decoded feature attributes to obtain weighted feature attributes.

[0021] In some embodiments, inputting the context features into a second multi-layer perceptron for data prediction to obtain the feature channel fusion weights includes:

[0022] Inputting the context features into a second multi-layer perceptron for data prediction to obtain a first intermediate value;

[0023] Inputting the first intermediate value into a preset activation function for data calculation to obtain a second intermediate value;

[0024] Obtaining a difference from the second intermediate value to obtain the feature channel fusion weights.

[0025] In some embodiments, performing weighted prediction based on the feature channel fusion weights, the feature attribute mean, and the decoded feature attributes to obtain weighted feature attributes includes:

[0026] Obtaining the product of the feature channel fusion weights and the decoded feature attributes to obtain a third intermediate value;

[0027] Obtaining a difference from the feature channel fusion weights to obtain a weight intermediate value, multiplying the weight intermediate value by the feature attribute mean to obtain a fourth intermediate value;

[0028] Accumulating the third intermediate value and the fourth intermediate value to obtain the weighted feature attributes.

[0029] In some embodiments, performing channel dimension grouping on the weighted feature attributes to obtain rendered feature attributes includes:

[0030] Based on a preset grouping method, grouping the weighted feature attributes into a first channel feature and a second channel feature in the channel dimension;

[0031] Adding the first channel feature and the second channel feature to obtain a fifth intermediate value, and selecting one of the first channel feature and the second channel feature as the residual feature;

[0032] Concatenating the non-residual feature and the fifth intermediate value to obtain the rendered feature attributes.

[0033] To achieve the above object, a second aspect of the embodiments of the present application proposes a three-dimensional scene reconstruction device, including:

[0034] A point cloud processing module: configured to obtain perspective images respectively corresponding to a target scene from multiple perspectives, and based on the perspective images, obtain a sparse point cloud, and initialize an anchor-based initial three-dimensional representation model according to the sparse point cloud, where the initial three-dimensional representation model includes a plurality of anchors and anchor attributes corresponding to each anchor, and the anchor attributes include feature attributes, size attributes, and position offsets;

[0035] Entropy encoding module: For each of the anchor points, perform compression encoding on the anchor point attributes to obtain the encoded bitstream corresponding to the target scene;

[0036] Entropy decoding module: For performing entropy decoding on the encoded bitstream to obtain decoded feature attributes, decoded size attributes, and decoded offset attributes;

[0037] Redundancy processing module: For estimating the mean of the feature attributes using a preset entropy model, performing weighted prediction based on the decoded feature attributes and the mean of the feature attributes to obtain weighted feature attributes, and performing channel dimension grouping on the weighted feature attributes to obtain rendered feature attributes;

[0038] Rendering and reconstruction module: For generating a three-dimensional representation model of the target scene based on the rendered feature attributes, the decoded size attributes, and the decoded offset attributes.

[0039] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0040] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a storage medium, the storage medium is a storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0041] The 3D scene reconstruction method, device, equipment, and storage medium proposed in the embodiments of this application obtain perspective images corresponding to a target scene from multiple perspectives, and obtain a sparse point cloud based on the perspective images. An initial 3D representation model based on anchor points is initialized according to the sparse point cloud. The initial 3D representation model includes multiple anchor points and the anchor point attributes corresponding to each anchor point. The anchor point attributes include feature attributes, size attributes, and position offsets. For each anchor point, the anchor point attributes are compressed and encoded to obtain a coded bitstream corresponding to the target scene. Then, the coded bitstream is entropy decoded to obtain decoded feature attributes, decoded size attributes, and decoded offset attributes. Next, a weighted prediction is performed based on the decoded feature attributes and the mean of the feature attributes predicted by a preset upper model to obtain weighted feature attributes, and the weighted feature attributes are grouped by channel dimension to obtain rendered feature attributes. Finally, a 3D representation model of the target scene is generated based on the rendered feature attributes, decoded size attributes, and decoded offset attributes. In the embodiments of this application, the feature attributes of the anchor points are reasonably simplified. On the one hand, although the feature attributes of the anchor points contain rich scene information, some information is repeated with the mean of the feature attributes predicted by the preset entropy model. Therefore, the mean of the feature attributes is introduced through the weighted prediction process, and the feature attributes of the anchor points only need to represent the missing and inaccurate scene information in the mean of the feature attributes, thereby effectively reducing the scene information contained in the feature attributes. On the other hand, by analyzing the similarity between channels, it is found that the similarity between all channels is very high. Therefore, if divided into two groups, the similarity between the two groups will also be very high. Therefore, the small-storage-cost residual is used to model the slight difference between the two groups. That is to say, in the related art, two groups of features need to be stored, while in the embodiments of this application, only one group of features and the residual need to be stored, and the other group of features can be represented by combining the stored group of features with the residual. This can reduce the channel redundancy of the feature attributes while retaining the key channel information, and further reduce the storage occupancy of the feature attributes of the anchor points. Through the optimization of the anchor point attributes, it is possible to reduce the transmission bandwidth, reduce the storage space, and improve the rate-distortion performance of the 3D representation model on the basis of obtaining higher-quality reconstruction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a flowchart of the 3D scene reconstruction method provided by the embodiments of this application.

[0043] Figure 2 is a flowchart of using a preset entropy model to compress and encode the anchor point attributes to obtain a coded bitstream corresponding to the target scene provided by the embodiments of this application.

[0044] Figure 3 is a flowchart of entropy decoding the coded bitstream to obtain decoded feature attributes, decoded size attributes, and decoded offset attributes provided by the embodiments of this application.

[0045] Figure 4 It is a schematic diagram of redundancy analysis of feature attributes provided by an embodiment of the present application.

[0046] Figure 5 It is a flowchart of weighted prediction to obtain weighted feature attributes based on the mean of feature attributes and decoded feature attributes provided by an embodiment of the present application.

[0047] Figure 6 It is a flowchart of inputting context features into a second multi-layer perceptron for data prediction to obtain feature channel fusion weights provided by an embodiment of the present application.

[0048] Figure 7 It is a flowchart of weighted prediction to obtain weighted feature attributes based on feature channel fusion weights, the mean of feature attributes, and decoded feature attributes provided by an embodiment of the present application.

[0049] Figure 8 It is a schematic diagram of channel correlation analysis of feature attributes provided by an embodiment of the present application.

[0050] Figure 9 It is a flowchart of grouping the weighted feature attributes by channel dimension to obtain rendered feature attributes provided by an embodiment of the present application.

[0051] Figure 10 It is a schematic diagram of the overall process provided by an embodiment of the present application.

[0052] Figure 11 It is a structural block diagram of a three-dimensional scene reconstruction device provided by another embodiment of the present application.

[0053] Figure 12 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0054] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0055] It should be noted that although functional module grouping is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module grouping in the device or the order in the flowchart.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0057] First, some terms involved in this application are analyzed:

[0058] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.

[0059] With the rapid development of technologies such as Virtual Reality (VR), Augmented Reality (AR), and Digital Twin, three-dimensional scene reconstruction technology has become an important research direction. Traditional three-dimensional reconstruction methods mostly rely on depth sensors (such as LiDAR) to extract scene geometric information, and usually have problems such as high cost, high computational complexity, and low reconstruction accuracy. In recent years, in order to achieve more efficient and higher-precision three-dimensional reconstruction, methods based on deep learning and view synthesis have attracted much attention in the academic and industrial fields.

[0060] As a three-dimensional scene representation method, 3D Gaussian Splatting (3DGS) has surpassed previous-generation scene representation technologies such as point clouds and Neural Radiance Fields (NeRF) in terms of training speed, rendering speed, and synthesis quality. 3DGS represents a three-dimensional scene as a set of Gaussian ellipsoids with local scene attributes (such as opacity, color, size, etc.), and uses a fast and differentiable rasterization pipeline to synthesize virtual views. However, 3DGS relies on a large number of Gaussian ellipsoids to achieve high-quality view synthesis results, resulting in a large amount of data processed, a complex calculation process, and high storage and transmission costs for the model.

[0061] In related technologies, Scaffold-GS proposed an anchor-based structured 3DGS representation method. This method uses anchors to cluster Gaussian ellipsoids with similar attributes in a local area, and predicts the attributes of Gaussian points based on the attributes such as offsets and features of the anchors, thereby achieving more efficient rendering. Among them, Scaffold-GS reduces the model size to a certain extent by clustering Gaussian ellipsoids into anchor attributes, but the storage space required for the related features of its anchors is still large.

[0062] Based on this, the embodiments of the present application provide a three-dimensional scene reconstruction method, apparatus, device, and storage medium, which reasonably simplify the feature attributes of anchor points. On the one hand, although the feature attributes of anchor points contain rich scene information, some information is repeated with the mean value of the feature attributes predicted by the preset entropy model. Therefore, by introducing the mean value of the feature attributes through a weighted prediction process, the feature attributes of the anchor points only need to represent the missing and inaccurate scene information in the mean value of the feature attributes, thereby effectively reducing the scene information contained in the feature attributes. On the other hand, by analyzing the similarity between channels, it is found that the similarity between all channels is very high. Therefore, if divided into two groups, the similarity between the two groups will also be very high. Therefore, the small-storage-cost residual is used to model the micro differences between the two groups. That is to say, in the related art, two groups of features need to be stored, while in the embodiments of the present application, only one group of features and the residual need to be stored, and the other group of features can be represented by combining the stored group of features with the residual. This can reduce the channel redundancy of the feature attributes while retaining the key channel information, further reducing the storage occupancy of the feature attributes of the anchor points. By optimizing the anchor point attributes, it is possible to reduce the transmission bandwidth, reduce the storage space, and improve the rate-distortion performance of the three-dimensional representation model on the basis of obtaining a higher-quality reconstruction result.

[0063] The embodiments of the present application provide a three-dimensional scene reconstruction method, apparatus, device, and storage medium, which will be specifically described through the following embodiments. First, the three-dimensional scene reconstruction method in the embodiments of the present application will be described.

[0064] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0065] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0066] The 3D scene reconstruction method provided by the embodiments of this application relates to the field of computer vision technology. The 3D scene reconstruction method provided by the embodiments of this application can be applied to a terminal, or to a server, or can be a computer program running on a terminal or a server. For example, the computer program can be a native program or software module in an operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in the operating system to run, such as a client that supports 3D scene generation, that is, a program that only needs to be downloaded into a browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module or plug-in. Among them, the terminal communicates with the server through a network. The 3D scene reconstruction method can be executed by the terminal or the server, or by the terminal and the server in cooperation.

[0067] In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, or a smart watch, etc. In addition, the terminal can also be a smart vehicle-mounted device. The smart vehicle-mounted device applies the 3D scene reconstruction method of this embodiment to provide relevant services and improve the driving experience. The server can be an independent server, or can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; it can also be a service node in a blockchain system, and the service nodes in the blockchain system form a peer-to-peer (Peer To Peer, P2P) network, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP) protocol. The connection between the terminal and the server can be made through communication connection methods such as Bluetooth, Universal Serial Bus (USB), or network, and this embodiment does not make any restrictions here.

[0068] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0069] It should be noted that in each specific embodiment of this application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.

[0070] The three-dimensional scene reconstruction method in the embodiments of this application will be described below.

[0071] Figure 1 is an optional flowchart of the three-dimensional scene reconstruction method provided by the embodiments of this application. Figure 1 The method in can include but is not limited to steps 110 to 150. At the same time, it can be understood that this embodiment does not specifically limit the order of steps 110 to 150 in Figure 1 and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added.

[0072] Step 110: Obtain perspective images respectively corresponding to a target scene from multiple perspectives, obtain a sparse point cloud based on the perspective images, and initialize an anchor-based initial three-dimensional representation model according to the sparse point cloud.

[0073] In one embodiment, the target scene is the scene that needs to be three-dimensionally modeled. For the target scene, a 360-degree panoramic surround layout can be adopted to deploy image acquisition devices to achieve synchronous acquisition. This surround layout can cover the target scene in all directions, ensuring that no angular information is missed, so as to obtain data of the target scene from all directions. Different image acquisition devices correspond to different perspectives, similar to observing the target scene from different positions, and there are differences in the pictures seen from different positions. At each perspective, the corresponding perspective image is obtained by means of an image acquisition device.

[0074] In one embodiment, the Colmap calibration method is used to calibrate the perspective images corresponding to all perspectives collected at the same time, to obtain the sparse point cloud representing the target scene, and the internal and external parameters corresponding to each image acquisition device.

[0075] Specifically, in this embodiment, the Colmap calibration method is used to analyze the feature information in each perspective image. Such feature information includes corner points, edges, and other significant visual features in the perspective image, etc. Subsequently, with the help of a specific feature extraction algorithm, these feature points are accurately identified. Then, matching operations are carried out according to the corresponding relationships of these feature points between different perspective images. For example, for the corner point of a building in a perspective image, the corresponding corner point of the same building needs to be found in other perspective images. In the matching process, factors such as the position of the feature point and the gray-scale change of the surrounding pixels are comprehensively considered to ensure the accuracy of the matching.

[0076] After the feature matching is completed, the internal and external parameters corresponding to each image acquisition device are calculated. Among them, the internal parameters reflect the internal imaging characteristics of the camera itself, such as the focal length of the camera, the position of the principal point, and possible lens distortion parameters, etc.; the external parameters reflect the position and attitude information of the camera in space, such as the rotation angle and translation vector of the camera relative to the target scene, etc.

[0077] At the same time, using the information of the already matched feature points, the sparse point cloud representing the target scene is obtained. This sparse point cloud is a discretized representation form of the target scene in three-dimensional space, composed of numerous feature points. These feature points are relatively sparsely distributed, and they outline the general outlines of the main objects in the target scene and key features such as spatial position relationships.

[0078] In one embodiment, in the Scaffold-GS scenario, the obtained sparse point cloud is initialized to obtain an initial three-dimensional representation model based on anchor points, and the initial three-dimensional representation model is composed of multiple different three-dimensional Gaussian distributions, each of which can be regarded as a three-dimensional ellipsoid, and multiple three-dimensional ellipsoids are used to fit the sparse point cloud. The anchor point is a reference point for organizing and processing Gaussian ellipsoids, similar to a "center" or "benchmark", which plays a role in grouping and managing Gaussian ellipsoids in a local area. Among them, in the initial three-dimensional representation model, each anchor point corresponds to one or more three-dimensional Gaussian distributions, and a three-dimensional Gaussian distribution can be represented by a Gaussian point at its center.

[0079] For the case of only one 3D Gaussian distribution, the 3D Gaussian distribution can usually describe the relatively simple and concentrated data distribution near the anchor point by virtue of its specific distribution characteristics. For example, for the anchor point corresponding to a relatively isolated and regular-shaped small object, only one 3D Gaussian distribution is sufficient to reflect its position in space and the data density within a certain range.

[0080] When an anchor point corresponds to more than one three-dimensional Gaussian distribution, it means that the data distribution in the area where the anchor point is located is relatively complex and needs to be jointly characterized by multiple three-dimensional Gaussian distributions. For example, in a scene covering a complex mechanical structure, the key node of a large mechanical device is regarded as an anchor point, and the parts around it have different shapes and complex spatial layouts. In this case, multiple three-dimensional Gaussian distributions are required, each of which has different central positions, covariances and other Gaussian parameters. Through the cooperation and synergy of these Gaussian points, the complex spatial form around the anchor point, the density changes of data, and the relationship between different parts can be accurately presented.

[0081] In one embodiment, each anchor point also has a corresponding anchor attribute a. Specifically, the anchor point attribute includes a feature attribute f, a size attribute l, and a position offset o, which is expressed as a∈{f, l, o}.

[0082] The feature attribute f reflects some inherent characteristics of the anchor point, such as the color and geometry of the object in the target scene. For example, in a three-dimensional scene with multiple objects, different anchor points often have different feature attributes.

[0083] The size attribute l is related to the size of the initial 3D representation model associated with the anchor point. In the 3D space, the initial 3D representation model can be used to present the shape and size of the object. For a larger object, the size attribute value of the corresponding anchor point may be relatively large; while for a smaller object, the size attribute value is relatively small.

[0084] In addition, since the anchor contains one or more three-dimensional Gaussian distributions, the position offset attribute o is used to represent the position offset of the three-dimensional Gaussian distributions within the anchor. With the position offset, the positional relationship between each three-dimensional Gaussian distribution and the anchor can be determined, and thus the coordinate information of the center position of each three-dimensional Gaussian distribution can be uniquely determined based on the coordinate value of the anchor and the position offset. Once the coordinate value and the attribute-related information of the anchor are determined, an initial three-dimensional representation model based on the anchor can be constructed, and the three-dimensional Gaussian distributions around each anchor can be used to describe the probability distribution of the scene in this local area.

[0085] Step 120: For each anchor, use a preset entropy model to perform compression encoding on the anchor attributes to obtain a coded bitstream corresponding to the target scene.

[0086] In one embodiment, the preset entropy model is used to compress the anchor attributes in the three-dimensional scene representation model. The preset entropy model can be a context model, a superprior model, etc. For example, the preset entropy model can be a context model, and this context model includes a binary hash table and a first multi-layer perceptron. Refer to Figure 2 , Figure 2 is a flowchart of using a preset entropy model to perform compression encoding on anchor attributes to obtain a coded bitstream corresponding to the target scene provided by an embodiment of the present application, including:

[0087] Step 210: Use the coordinate value of the anchor to perform an interpolation operation in the binary hash table to obtain context features.

[0088] In some embodiments, in order to achieve the best rate-distortion performance, a binary hash table is pre-constructed to store the context information corresponding to each anchor. For example, in the target scene, the distribution of objects around the anchor, the spatial relationship with adjacent feature points, and the possible semantic information, etc. By pre-constructing this binary hash table, these rich context information can be stored in a structured form, facilitating subsequent retrieval and use. In addition, the binary hash table can be optimized during the training process. During training, based on a large number of data samples and corresponding annotation information, the way of storing information, association rules, and related parameters in the hash table are continuously adjusted. For example, when processing data containing various different types of scenes (such as indoor scenes, outdoor scenes, etc.), by allowing the hash table to come into contact with various anchors and their actual context situations, it can learn how to store and reflect context information more accurately in different scenes, thereby improving the accuracy and effectiveness of subsequent information acquisition using it.

[0089] Next, input the coordinate value corresponding to the anchor in the three-dimensional space into the binary hash table, perform a mapping operation on the coordinate value, and the context feature f corresponding to the anchor can be obtained c, where the coordinate values of the anchor points determine their specific positions in the three-dimensional space of the target scene. Taking this position information as input is like providing an "index" to a binary hash table to clarify the information to be searched. Since the context information stored in the hash table may be discrete, interpolation operations are used to perform numerical calculations based on the existing stored data in the hash table, and the accurate context features corresponding to the anchor points are obtained according to the relevant information stored around.

[0090] Step 220: Input the context features into the first multi-layer perceptron for data prediction to obtain probability distribution parameters and quantization step sizes.

[0091] In one embodiment, the obtained context features are input into the first multi-layer perceptron to perform data prediction on each attribute of the anchor point, obtaining a quantization step size and probability distribution parameters corresponding to each attribute. Among them, the probability distribution parameters include an attribute mean and an attribute standard deviation. The attribute mean is used to reflect the average level of the anchor point attribute as a whole, while the attribute standard deviation reflects the degree of dispersion of the attribute data relative to the mean, that is, the fluctuation of the data.

[0092] The above process is expressed as:

[0093]

[0094] where, μ a represents the attribute mean corresponding to attribute a, σ a represents the attribute standard deviation corresponding to attribute a, q a represents the quantization step size corresponding to attribute a, and MLP1 represents the first multi-layer perceptron.

[0095] As can be seen from the above, through the first multi-layer perceptron, the following can be obtained: the feature attribute mean μ f corresponding to the feature attribute, the feature attribute standard deviation σ f and the quantization step size q f of the feature attribute; the size attribute mean μ l corresponding to the size attribute, the size attribute standard deviation σ l and the quantization step size q l of the size attribute; the position offset mean μ o corresponding to the position offset, the position offset standard deviation σ o and the quantization step size of the position offset q o .

[0096] Step 230: Quantize the anchor point attributes using the quantization step size to obtain quantized attributes, and perform entropy coding on the quantized attributes using the probability distribution parameters to obtain an encoded bitstream.

[0097] In one embodiment, for each different anchor point attribute, it can be first quantized using a quantization step size to obtain a corresponding quantized attribute, and then the quantized attribute related to the anchor point attribute is entropy-encoded using the corresponding attribute mean and attribute standard deviation to obtain an encoded bitstream. The process of quantizing the anchor point attribute using the quantization step size is expressed as:

[0098]

[0099] where a represents the specific value corresponding to the anchor point attribute a.

[0100] In one embodiment, there are multiple encoding methods for entropy encoding, such as arithmetic encoding, Huffman tree encoding, etc. The specific method of entropy encoding in the embodiments of the present application is not limited. Taking arithmetic encoding as an example, as an entropy encoding technology based on probability statistics, the data of the entire anchor point attribute is regarded as a whole symbol sequence for processing. When performing arithmetic encoding on the anchor point attribute, the probability distribution corresponding to each parameter and its value in the anchor point attribute is first analyzed. Subsequently, based on this probability information, all the data in the anchor point attribute is gradually mapped to the corresponding encoding interval, and by continuously subdividing and determining the encoding interval, a uniquely corresponding encoded bitstream is finally generated. After obtaining the encoded bitstream, subsequent decoding operations are performed on it.

[0101] In one embodiment, the attribute mean μ a corresponding to the probability distribution parameter a and the attribute standard deviation σ can also be used to reflect the bit rate loss during the encoding process. For example, when the bit rate loss is proportional to the calculated information entropy, the bit rate loss value can be obtained based on the information entropy, and then the total bit rate loss value is calculated according to the bit rate loss values of all attributes. Based on this total bit rate loss value, the overall parameters such as the attribute value of the anchor point and the preset entropy model are adjusted, thereby controlling the bit consumption generated by the information related to the attributes of the anchor point during the compression process.

[0102] It can be understood that the encoding process in the embodiments of the present application has two advantages. On the one hand, it can compress the storage space of the anchor point attribute. After encoding, the attributes of the anchor point exist in a more compact encoded bitstream form, greatly saving storage resources, which is very beneficial for long-term storage of a large amount of anchor point attribute data or storing relevant data on devices with limited storage resources. On the other hand, if there is a transmission requirement, considering that the data volume of the encoded bitstream is relatively small, the transmission rate can also be improved to ensure that the data can quickly and stably reach the receiving end from the sending end.

[0103] In one embodiment, the encoding process and the corresponding decoding process in the embodiments of the present application are performed within the same processing device. If the resources of the processing device are limited, encoding and decoding operations are performed in the device to reduce the amount of data to be processed, relieve the computing burden of the device, so that the device can complete subsequent processing tasks with limited resources. In addition, in application scenarios with high real-time requirements, performing encoding and decoding operations within the same device can also reduce the time consumed in intermediate links such as data transmission, quickly achieve the conversion from the original data to the final reconstruction result, thereby ensuring the processing efficiency.

[0104] In addition, the encoding process and the corresponding decoding process can also be carried out in a distributed manner. First, the encoding operation is performed at the encoding end, then the corresponding encoded bitstream is transmitted, and then the decoding operation is implemented on the corresponding processing device at the decoding end. In the case of limited transmission bandwidth, transmitting the encoded bitstream can effectively reduce the amount of data transmitted and improve the transmission efficiency. In this way, the entire reconstruction process will not be affected by delays in the transmission link, thereby improving the overall reconstruction efficiency. It can be understood that the execution locations of the encoding process and the corresponding decoding process are not limited in this embodiment.

[0105] Step 130: Perform entropy decoding on the encoded bitstream to obtain decoded feature attributes, decoded size attributes, and decoded offset attributes.

[0106] In one embodiment, referring to Figure 3 , Figure 3 is the flowchart of performing entropy decoding on the encoded bitstream provided by the embodiments of the present application to obtain decoded feature attributes, decoded size attributes, and decoded offset attributes, which specifically includes the following steps:

[0107] Step 310: Perform entropy decoding on the encoded bitstream using probability distribution parameters to obtain decoded quantization attributes.

[0108] In one embodiment, entropy decoding is performed on the encoded bitstream using an entropy decoding method corresponding to entropy encoding to obtain decoded quantization attributes, where the decoded quantization attributes include quantization feature attributes, quantization size attributes, and quantization position offsets.

[0109] Step 320: Obtain the quantization step size corresponding to each decoded quantization attribute, multiply the quantization feature attribute by the quantization step size to obtain the decoded feature attribute, multiply the quantization size attribute by the quantization step size to obtain the decoded size attribute, and multiply the quantization position offset by the quantization step size to obtain the decoded offset attribute.

[0110] In one embodiment, the quantization step size can be obtained according to the previous calculation, so the decoded feature attribute is expressed as:

[0111]

[0112] The decoded size attribute is represented as:

[0113]

[0114] The decoded offset attribute is represented as:

[0115]

[0116] Thus, the decoded attribute information is obtained.

[0117] In one embodiment, the above decoded attribute information is different during the training process and the inference process. During the training process, mean noise is added to the attribute to simulate the decoding step, which is represented as:

[0118]

[0119] Wherein, represents the attribute information of the decoded encoded attribute a' during the training process, represents a uniform distribution, a′ represents the value of the encoded attribute a' during the training process, and q a ′ represents the quantization step of the encoded attribute a' during the training process.

[0120] Step 140: Perform weighted prediction based on the decoded feature attribute and the mean value of the feature attributes predicted by the preset entropy model to obtain a weighted feature attribute, and group the weighted feature attribute in the channel dimension to obtain a rendering feature attribute.

[0121] In one embodiment, the embodiments of the present application select feature attributes for redundancy analysis. Different from the related art that directly uses the decoded feature attributes obtained by decoding for subsequent rendering and reconstruction processes, instead, they are simplified and compressed, thereby reducing the storage requirements of the anchor points. Although the feature attributes of the anchor points contain rich scene information, some information is repeated with the mean value of the feature attributes predicted by the preset entropy model. Therefore, by introducing the mean value of the feature attributes through the weighted prediction process, the feature attributes of the anchor points only need to represent the missing and inaccurate scene information in the mean value of the feature attributes, thereby effectively reducing the scene information contained in the feature attributes.

[0122] In one embodiment, referring to Figure 4 , Figure 4 is a schematic diagram of the redundancy analysis of the feature attributes provided by the embodiments of the present application. From Figure 4It can be seen that the similarity between the conventional rendering image obtained from the complete anchor point attributes and the mean rendering image corresponding to the mean of the feature attributes is relatively high. From this, it can be known that using the mean of the feature attributes to replace the feature attributes in the anchor point to participate in the rendering process can also obtain good scene geometric structures and colors. In the related art, by focusing on the role of the mean of the feature attributes in calculating the bitrate loss, the fact that the mean of the feature attributes carries rich scene information that can be used to assist reconstruction is ignored. Therefore, in the embodiments of the present application, weighted prediction is combined with the mean of the feature attributes to reduce data redundancy.

[0123] In one embodiment, referring to Figure 5 , Figure 5 is a flowchart of weighted prediction based on the mean of the feature attributes and the decoded feature attributes provided by the embodiments of the present application to obtain weighted feature attributes, which specifically includes the following steps:

[0124] Step 510: Input the context features into the second multi-layer perceptron for data prediction to obtain the feature channel fusion weights.

[0125] In one embodiment, the feature channel fusion weights are used to assign corresponding weights according to the importance degrees of different feature attributes to the reconstruction results, so as to adaptively aggregate the mean of the feature attributes μ f and the decoded feature attributes to guide the model to use the scene information in the mean for reconstruction and reduce the amount of information to be represented by the decoded feature attributes . Referring to Figure 6 , Figure 6 is a flowchart of inputting the context features into the second multi-layer perceptron for data prediction to obtain the feature channel fusion weights provided by the embodiments of the present application, which specifically includes the following steps:

[0126] Step 610: Input the context features into the second multi-layer perceptron for data prediction to obtain a first intermediate value.

[0127] In one embodiment, the first intermediate value is expressed as:

[0128] MLP w (f c )

[0129] where MLP w represents the second multi-layer perceptron.

[0130] Step 620: Input the first intermediate value into a preset activation function for data calculation to obtain a second intermediate value.

[0131] In one embodiment, the preset activation function is the sigmoid activation function, and the second intermediate value is expressed as:

[0132] Sigmoid(MLP w (fc ))

[0133] Step 630: Obtain the difference from the second intermediate value to get the feature channel fusion weight.

[0134] In one embodiment, the feature channel fusion weight w is expressed as:

[0135] w = 1 - Sigmoid(MLP w (f c ))

[0136] Step 520: Perform weighted prediction based on the feature channel fusion weight, the feature attribute mean, and the decoded feature attribute to obtain the weighted feature attribute.

[0137] In one embodiment, referring to Figure 7 , Figure 7 is the flowchart of performing weighted prediction based on the feature channel fusion weight, the feature attribute mean, and the decoded feature attribute provided by the embodiment of the present application to obtain the weighted feature attribute, which specifically includes the following steps:

[0138] Step 710: Obtain the product of the feature channel fusion weight and the decoded feature attribute to get the third intermediate value.

[0139] In one embodiment, the third intermediate value is expressed as:

[0140]

[0141] Step 720: Obtain the difference from the feature channel fusion weight to get the weight intermediate value, and multiply the weight intermediate value by the feature attribute mean to get the fourth intermediate value.

[0142] In one embodiment, the fourth intermediate value is expressed as:

[0143] (1 - w)·μ f

[0144] where (1 - w) represents the weight intermediate value.

[0145] Step 730: Accumulate the third intermediate value and the fourth intermediate value to get the weighted feature attribute.

[0146] In one embodiment, the weighted feature attribute is expressed as:

[0147]

[0148] In one embodiment, after obtaining the weighted feature attributes, by analyzing the similarity between channels, it is found that the similarity between all channels is very high. Therefore, if divided into two groups, the similarity between the two groups will also be very high. Thus, the small-storage-cost residual is used to model the slight difference between the two groups. That is to say, in the related art, two groups of features need to be stored, while in the embodiment of the present application, only one group of features and the residual need to be stored, and the other group of features can be represented by combining the stored group of features with the residual. This can reduce the channel redundancy of the feature attributes, while retaining the key channel information, and further reduce the storage occupancy of the feature attributes of the anchor point. Refer to Figure 8 , Figure 8 FIG. Figure 8 is a schematic diagram of channel correlation analysis of the feature attributes provided by the embodiment of the present application. In the figure, 4 channels are taken as an example for illustration. By visualizing the channels and using different channels for rendering, channel rendering image 1, channel rendering image 2, channel rendering image 3, and channel rendering image 4 are obtained respectively. It can be seen that the similarity of these four rendering images is extremely high. Thus, it can be seen that there is a high similarity between the channels of the feature attributes of the anchor point. Therefore, the embodiment of the present application uses the method of grouping in the channel dimension to remove the redundancy of the channels.

[0149] In one embodiment, refer to Figure 9 , Figure 9 FIG. Figure 9 is a flowchart of grouping the weighted feature attributes in the channel dimension to obtain the rendered feature attributes provided by the embodiment of the present application, which specifically includes the following steps:

[0150] Step 910: Group the weighted feature attributes into a first channel feature and a second channel feature in the channel dimension based on a preset grouping method.

[0151] In one embodiment, the preset grouping method may be uniform grouping, that is, randomly dividing into two equal parts, or other grouping methods, which are not limited in this embodiment. Assuming that the channel dimension is D, the preset grouping method includes but is not limited to:

[0152] 1) Let [0, 1,..., D / 2 - 1] be the first group and [D / 2, D / 2 + 1,..., D - 1] be the second group; 2) Let [0, 2,..., D] be the first group and [1, 3,..., D - 1] be the second group, etc. Group the weighted feature attributes into a first channel feature and a second channel feature in the channel dimension according to the preset grouping method, which is expressed as:

[0153]

[0154] where D represents the channel dimension, represents the first channel feature, represents the second channel feature.

[0155] Step 920: Add the first-channel feature and the second-channel feature to obtain a fifth intermediate value, and select one of the first-channel feature and the second-channel feature as the residual feature.

[0156] In one embodiment, the fifth intermediate value is expressed as:

[0157]

[0158] At this time, in order to reduce the redundancy between these two sets of features, that is, the redundancy in the channel dimension of the anchor features, it is necessary to select one of the first-channel feature and the second-channel feature as the residual feature to reduce the amount of information while not damaging the representation ability of the anchor features. Since the rendering results of different channels are highly similar, this selection can be made randomly.

[0159] Step 930: Concatenate the non-residual feature and the fifth intermediate value to obtain the rendering feature attribute.

[0160] In one embodiment, if the second-channel feature is selected as the residual feature, at this time the non-residual feature is the first-channel feature This way based on residual prediction can explicitly utilize the correlation between feature channels, reduce the redundancy between the first-channel feature and the second-channel feature, reduce the amount of information that needs to be represented in the second-channel feature, and thus reduce the computational overhead.

[0161] After concatenating the non-residual feature and the fifth intermediate value, the rendering feature attribute is obtained Expressed as:

[0162]

[0163] Next, use the rendering feature attribute to replace the decoded feature attribute obtained originally and participate in the subsequent rendering process.

[0164] Step 150: Generate a three-dimensional representation model of the target scene based on the rendering feature attribute, the decoded size attribute, and the decoded offset attribute.

[0165] In one embodiment, according to the above process, the rendering feature attribute, the decoded size attribute, and the decoded offset attribute form a new three-dimensional representation model, which is used as the reconstruction input data. Among them, the decoded size attribute and the decoded offset attribute are the decoded attribute data, and the reconstruction input data is expressed as

[0166] Next, similar to the Scaffold-GS method, first obtain the decoded feature attributes, size attributes, and position offsets corresponding to each anchor point from the reconstructed input data, and then predict the Gaussian distribution attributes corresponding to all the three-dimensional Gaussian distributions contained in the anchor points. For example, predict features such as the relevant category, material, and general shape of the three-dimensional Gaussian distribution based on the feature attributes; clarify the distribution range of the three-dimensional Gaussian distribution in space with the help of the size attributes; accurately determine the position of the center of the three-dimensional Gaussian distribution relative to the anchor point through the position offset.

[0167] After obtaining the Gaussian distribution attributes, it enters the view synthesis stage. In this stage, by combining the internal and external parameters corresponding to the image acquisition device and the Gaussian distribution attributes, the appearance of the target scene from different perspectives can be determined. Since the Gaussian distribution attributes cover multiple aspects of information such as material, spatial distribution, and relative position, when observing the target scene from different perspectives, these attributes will work together, just like individual "building blocks", and are combined according to their respective characteristics and positional relationships to present the specific appearance of the scene from different perspectives, ultimately achieving the three-dimensional reconstruction of the target scene.

[0168] In one embodiment, refer to Figure 10 , Figure 10 which is the overall process schematic diagram provided by the embodiment of the present application. In Figure 10 , first generate an initial three-dimensional representation model corresponding to multiple anchor points according to the sparse point cloud. Each anchor point corresponds to a voxel, and each voxel contains one or more Gaussian distributions. Among them, in addition to having coordinate values, the anchor point also includes a feature attribute f, a size attribute l, and a position offset o.

[0169] Next, enter the entropy modeling process, which includes an encoding and a decoding process. Use the coordinate values of the anchor points to perform an interpolation operation in the binary hash table to obtain the context feature f c , and input the context feature f c into the first multi-layer perceptron MLP1 to perform data prediction on each anchor point attribute of the anchor point, and obtain the attribute mean μ a , the attribute standard deviation σ a and the corresponding quantization step q a .

[0170] Then, according to the obtained feature attribute mean μ f and the decoded feature attribute , perform a mean-based weighted prediction process. First, input the context feature f c into the second multi-layer perceptron MLP w to perform data prediction and obtain the feature channel fusion weight w. Then, based on the feature channel fusion weight w, the feature attribute mean μ f and the decoded feature attribute Perform weighted prediction to obtain weighted feature attributes

[0171] Next, enter the channel dimension grouping process. Based on a preset grouping method, group the weighted feature attributes into a first channel feature and a second channel feature in the channel dimension. Then, add the first channel feature and the second channel feature to obtain a fifth intermediate value. Select one of the first channel feature and the second channel feature as the residual feature, and splice the non-residual feature and the fifth intermediate value to obtain the rendered feature attributes

[0172] Finally, enter the rendering process. Combining with the three-dimensional Gaussian splash technology, generate the Gaussian point attributes corresponding to the anchor points based on the rendered feature attributes, the decoded size attributes, and the decoded offset attributes, and generate a three-dimensional representation model of the target scene according to all the Gaussian point attributes

[0173] The three-dimensional scene reconstruction method provided by the embodiments of this application first uses a weighted prediction method based on the mean to adaptively use the mean of the probability distribution predicted by the entropy model for the anchor point features for reconstruction, reducing the amount of scene information that needs to be represented in the anchor point features and reducing the storage overhead without damaging the reconstruction quality of the scene representation model. In addition, a cross-channel residual prediction process is performed. The anchor point features are divided into two groups according to the channel dimension, and one of the groups is regarded as the residual of the other group to make full use of the similarity between the channels of the anchor point feature attributes to reduce the redundancy between the two groups of features, thereby reducing the storage requirements of the three-dimensional scene representation model

[0174] In one embodiment, for rendering performance verification, the embodiments of this application carried out a comparative analysis with the three-dimensional static scene compression methods in the related art for the two datasets of BungeeNeRF and Mip-NeRF360. Among them, the selected three-dimensional static scene compression methods cover 3DGS, Scaffold-GS, and HAC (Hash-grid Assisted). The selected analysis metrics are: Peak Signal to Noise Ratio (PSNR), Structural Similarity Index (SSIM), Learned Perceptual Image Patch Similarity (LPIPS), and Storage Capacity Storage (MB). The specific comparative analysis results are presented in Tables 1 and 2 below

[0175] The results show that, compared with the best-performing HAC in the related art, the embodiments of the present application can reduce the model storage cost while basically not damaging the rendering quality of the model, and even achieve better rendering quality in the case of the BungeeNeRF dataset. This fully demonstrates the strong superiority of the 3D scene reconstruction method involved in the embodiments of the present application.

[0176] Method PSNR (dB) SSIM LPIPS Storage (MB) 3DGS 24.87 0.841 0.205 1616 Scaffold - GS 26.62 0.865 0.241 183.0 HAC 26.48 0.845 0.250 18.49 Embodiment of the present application 26.58 0.850 0.245 15.15

[0177] Table 1 Analysis results on the BungeeNeRF dataset

[0178] Method PSNR (dB) SSIM LPIPS Storage (MB) 3DGS 27.49 0.813 0.222 744.7 Scaffold - GS 27.50 0.806 0.252 253.9 HAC 27.53 0.807 0.238 15.26 Embodiment of the present application 27.53 0.806 0.240 12.82

[0179] Table 2 Analysis results on the Mip-NeRF 360 dataset

[0180] The technical solution provided by the embodiments of the present application obtains perspective images corresponding to a target scene from multiple perspectives respectively, obtains a sparse point cloud based on the perspective images, initializes an anchor-based initial three-dimensional representation model according to the sparse point cloud. The initial three-dimensional representation model includes multiple anchors and anchor attributes corresponding to each anchor. The anchor attributes include feature attributes, size attributes, and position offsets. For each anchor, the anchor attributes are compressed and encoded to obtain a coded bitstream corresponding to the target scene. Then, the coded bitstream is entropy decoded to obtain decoded feature attributes, decoded size attributes, and decoded offset attributes. Next, a weighted prediction is performed based on the decoded feature attributes and the mean of the feature attributes predicted by using a preset upper model to obtain weighted feature attributes, and the weighted feature attributes are grouped in the channel dimension to obtain rendered feature attributes. Finally, a three-dimensional representation model of the target scene is generated based on the rendered feature attributes, the decoded size attributes, and the decoded offset attributes. In the embodiments of the present application, the feature attributes of the anchors are reasonably simplified. On the one hand, although the feature attributes of the anchors contain rich scene information, some information is redundant with the mean of the feature attributes predicted by the preset entropy model. Therefore, the mean of the feature attributes is introduced through the weighted prediction process, so that the feature attributes of the anchors only need to represent the missing and inaccurate scene information in the mean of the feature attributes, thereby effectively reducing the scene information contained in the feature attributes. On the other hand, by analyzing the similarity between channels, it is found that the similarity between all channels is very high. Therefore, if divided into two groups, the similarity between the two groups will also be very high. Therefore, the small-storage-cost residual is used to model the slight difference between the two groups. That is to say, in the related art, two groups of features need to be stored, while in the embodiments of the present application, only one group of features and the residual need to be stored, and the other group of features can be represented by combining the stored group of features with the residual. This can reduce the channel redundancy of the feature attributes, while retaining the key channel information, and further reduce the storage occupancy of the feature attributes of the anchors. Through the optimization of the anchor attributes, it is possible to reduce the transmission bandwidth, reduce the storage space, and improve the rate-distortion performance of the three-dimensional representation model on the basis of obtaining a higher-quality reconstruction result.

[0181] The embodiments of the present application further provide a three-dimensional scene reconstruction device, which can implement the above three-dimensional scene reconstruction method. Refer to Figure 11 , the device includes:

[0182] The point cloud processing module 1110 is configured to obtain perspective images corresponding to a target scene from multiple perspectives respectively, obtain a sparse point cloud based on the perspective images, and initialize an anchor-based initial three-dimensional representation model according to the sparse point cloud. The initial three-dimensional representation model includes multiple anchors and anchor attributes corresponding to each anchor. The anchor attributes include feature attributes, size attributes, and position offsets.

[0183] The entropy encoding module 1120 is configured to, for each anchor, compress and encode the anchor attributes to obtain a coded bitstream corresponding to the target scene.

[0184] Entropy decoding module 1130: It is used to perform entropy decoding on the encoded bitstream to obtain decoded feature attributes, decoded size attributes, and decoded offset attributes.

[0185] Redundancy processing module 1140: It is used to estimate the mean of feature attributes using a preset entropy model, perform weighted prediction based on the decoded feature attributes and the mean of feature attributes to obtain weighted feature attributes, and group the weighted feature attributes in the channel dimension to obtain rendered feature attributes.

[0186] Rendering and reconstruction module 1150: It is used to generate a three-dimensional representation model of the target scene based on the rendered feature attributes, decoded size attributes, and decoded offset attributes.

[0187] The specific implementation manner of the three-dimensional scene reconstruction device in this embodiment is basically the same as that of the above three-dimensional scene reconstruction method, and will not be elaborated here.

[0188] This application embodiment also provides an electronic device, including:

[0189] At least one memory;

[0190] At least one processor;

[0191] At least one program;

[0192] The program is stored in the memory, and the processor executes the at least one program to implement the three-dimensional scene reconstruction method described above in this application. This electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0193] Please refer to Figure 12 , Figure 12 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0194] A processor 1201, which can be implemented in ways such as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in this application embodiment;

[0195] The memory 1202 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1202 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1202, and the processor 1201 is used to call and execute the three-dimensional scene reconstruction method of the embodiments of this application;

[0196] The input / output interface 1203 is used to implement information input and output;

[0197] The communication interface 1204 is used to implement communication and interaction between this device and other devices. It can implement communication through a wired method (such as USB, network cable, etc.), or can also implement communication through a wireless method (such as mobile network, WIFI, Bluetooth, etc.); and the bus 1205 is used to transmit information between various components of the device (such as the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204);

[0198] Among them, the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204 are communicatively connected to each other inside the device through the bus 1205.

[0199] The embodiments of this application also provide a storage medium. The storage medium is a storage medium that stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned three-dimensional scene reconstruction method.

[0200] As a non-transitory storage medium, the memory can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0201] The 3D scene reconstruction method, device, equipment, and storage medium proposed in the embodiments of this application obtain perspective images corresponding to a target scene from multiple perspectives, and obtain a sparse point cloud based on the perspective images. An initial 3D representation model based on anchor points is initialized according to the sparse point cloud. The initial 3D representation model includes multiple anchor points and the anchor point attributes corresponding to each anchor point. The anchor point attributes include feature attributes, size attributes, and position offsets. For each anchor point, the anchor point attributes are compressed and encoded to obtain a coded bitstream corresponding to the target scene. Then, the coded bitstream is entropy decoded to obtain decoded feature attributes, decoded size attributes, and decoded offset attributes. Next, weighted prediction is performed based on the decoded feature attributes and the mean of the feature attributes predicted by using a preset upper model to obtain weighted feature attributes, and the weighted feature attributes are grouped in the channel dimension to obtain rendered feature attributes. Finally, a 3D representation model of the target scene is generated based on the rendered feature attributes, decoded size attributes, and decoded offset attributes. In the embodiments of this application, the feature attributes of the anchor points are reasonably simplified. On the one hand, although the feature attributes of the anchor points contain rich scene information, some of the information is redundant with the mean of the feature attributes predicted by the preset entropy model. Therefore, the mean of the feature attributes is introduced through the weighted prediction process, and the feature attributes of the anchor points only need to represent the missing and inaccurate scene information in the mean of the feature attributes, thereby effectively reducing the scene information contained in the feature attributes. On the other hand, by analyzing the similarity between channels, it is found that the similarity between all channels is very high. Therefore, if divided into two groups, the similarity between the two groups will also be very high. Therefore, the small storage cost residual is used to model the micro differences between the two groups. That is to say, in the related art, two groups of features need to be stored, while in the embodiments of this application, only one group of features and the residual need to be stored, and the other group of features can be represented by combining the stored group of features with the residual. This can reduce the channel redundancy of the feature attributes, while retaining the key channel information, and further reduce the storage occupancy of the feature attributes of the anchor points. Through the optimization of the anchor point attributes, it is possible to reduce the transmission bandwidth, reduce the storage space, and improve the rate-distortion performance of the 3D representation model on the basis of obtaining a higher-quality reconstruction result.

[0202] The embodiments described in the embodiments of this application are for more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems. Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0203] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0204] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0205] It should be understood that in this application, the terms "first", "second", "third", "fourth", etc. (if any) in the specification and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0206] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one)" or a similar expression below refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0207] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the grouping of the above-mentioned units is only a logical function grouping. In actual implementation, there may be other grouping methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0208] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0209] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0210] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store programs.

[0211] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.

Claims

1. A three-dimensional scene reconstruction method, characterized in that: include: Obtaining perspective images corresponding to the target scene at multiple perspectives, and obtaining a sparse point cloud based on the perspective images, and initializing an initial three-dimensional representation model based on anchor points according to the sparse point cloud, wherein the initial three-dimensional representation model includes multiple anchor points and anchor point attributes corresponding to each anchor point, and the anchor point attributes include feature attributes, size attributes, and position offsets; For each of the anchor points, compress and encode the anchor point attributes using a preset entropy model to obtain an encoded bitstream corresponding to the target scene; Performing entropy decoding on the encoded bitstream to obtain a decoding feature attribute, a decoding size attribute, and a decoding offset attribute; Performing weighted prediction according to the decoded feature attribute and the feature attribute mean value predicted based on the preset entropy model to obtain a weighted feature attribute, and grouping the weighted feature attribute according to the channel dimension to obtain a rendering feature attribute; A three-dimensional representation model of the target scene is generated based on the rendering feature attribute, the decoded size attribute, and the decoded offset attribute.

2. The three-dimensional scene reconstruction method according to claim 1, characterized in that: The preset entropy model is a context model, including a binary hash table and a first multi-layer perceptron. The preset entropy model is used to compress and encode the anchor point attributes to obtain the encoding stream corresponding to the target scene, including: Using the coordinate value of the anchor point to perform an interpolation operation in the binary hash table to obtain a context feature; Inputting the context features into the first multi-layer perceptron for data prediction to obtain probability distribution parameters and quantization step length; The anchor point attribute is quantized using the quantization step to obtain a quantized attribute, and the quantized attribute is entropy encoded using the probability distribution parameter to obtain the encoded bitstream.

3. The three-dimensional scene reconstruction method according to claim 2, characterized in that: The entropy decoding of the encoded code stream to obtain a decoding feature attribute, a decoding size attribute and a decoding offset attribute includes: Performing entropy decoding on the coded bitstream using the probability distribution parameters to obtain decoded quantization attributes, wherein the decoded quantization attributes include quantization feature attributes, quantization size attributes, and quantization position offsets; Obtain a quantization step size corresponding to each of the decoding quantization attributes, multiply the quantization feature attribute by the quantization step size to obtain the decoding feature attribute, multiply the quantization size attribute by the quantization step size to obtain the decoding size attribute, and multiply the quantization position offset by the quantization step size to obtain the decoding offset attribute.

4. The three-dimensional scene reconstruction method according to claim 2, characterized in that: The step of performing weighted prediction according to the feature attribute mean and the decoded feature attribute to obtain a weighted feature attribute includes: Inputting the context features into a second multi-layer perceptron for data prediction to obtain feature channel fusion weights; A weighted prediction is performed based on the feature channel fusion weight, the feature attribute mean and the decoded feature attribute to obtain a weighted feature attribute.

5. The three-dimensional scene reconstruction method according to claim 4, characterized in that: The step of inputting the context features into a second multi-layer perceptron for data prediction to obtain feature channel fusion weights includes: Inputting the context feature into a second multi-layer perceptron to perform data prediction to obtain a first intermediate value; Inputting the first intermediate value into a preset activation function to perform data calculation to obtain a second intermediate value; Obtain a difference between the first and the second intermediate value to obtain the feature channel fusion weight.

6. The three-dimensional scene reconstruction method according to claim 4, characterized in that: The step of performing weighted prediction based on the feature channel fusion weight, the feature attribute mean and the decoded feature attribute to obtain the weighted feature attribute includes: Obtaining the product of the feature channel fusion weight and the decoded feature attribute to obtain a third intermediate value; Obtaining a difference between the first and second feature channel fusion weights to obtain a weight intermediate value, and multiplying the weight intermediate value by the feature attribute mean to obtain a fourth intermediate value; The third intermediate value and the fourth intermediate value are accumulated to obtain the weighted feature attribute.

7. The three-dimensional scene reconstruction method according to claim 1, characterized in that: The step of grouping the weighted feature attributes by channel dimension to obtain rendering feature attributes includes: Grouping the weighted feature attributes into first channel features and second channel features in a channel dimension based on a preset grouping method; Adding the first channel feature and the second channel feature to obtain a fifth intermediate value, and selecting one of the first channel feature and the second channel feature as a residual feature; The rendering feature attribute is obtained by concatenating the non-residual feature and the fifth intermediate value.

8. A three-dimensional scene reconstruction device, characterized in that: include: Point cloud processing module: used to obtain the view images corresponding to the target scene under multiple view angles, and obtain the sparse point cloud based on the view images, and initialize the initial three-dimensional representation model based on the anchor point according to the sparse point cloud, wherein the initial three-dimensional representation model includes multiple anchor points and anchor point attributes corresponding to each anchor point, and the anchor point attributes include feature attributes, size attributes and position offset; Entropy coding module: used for compressing and coding the anchor point attributes for each anchor point to obtain a coding stream corresponding to the target scene; Entropy decoding module: used to perform entropy decoding on the coded bitstream to obtain decoding feature attributes, decoding size attributes and decoding offset attributes; Redundancy processing module: used to estimate the feature attribute mean using a preset entropy model, perform weighted prediction according to the decoded feature attribute and the feature attribute mean to obtain weighted feature attributes, and group the weighted feature attributes according to channel dimensions to obtain rendering feature attributes; Rendering and reconstruction module: used to generate a three-dimensional representation model of the target scene based on the rendering feature attribute, the decoding size attribute and the decoding offset attribute.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the three-dimensional scene reconstruction method according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the three-dimensional scene reconstruction method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Detection method and device, electronic equipment and computer readable medium

    CN116773534A

  • Three-dimensional volume video coding compression method

    CN119052510A

  • System and method for compressing and / or reconstructing medical image

    US20240428927A1

  • Point cloud encoding / decoding method and apparatus, device, and storage medium

    WO2024221458A1

Cited By

  • Light field image coding method

    CN121217935A

  • Three-dimensional scene reconstruction method and apparatus, and device and storage medium

    WO2026144340A1