3D scene implicit reconstruction method and device based on adaptive sparse coding guidance
Through the 3D scene implicit reconstruction method based on adaptive sparse encoding guidance, the problem in the prior art that dependence on truth data and global coding cannot take into account details and fixed local coding is difficult to adaptively allocate resources, and high-precision three-dimensional reconstruction and detail restoration are achieved, which significantly improves the expression ability and reconstruction accuracy of complex areas.
Patent Information
- Application Number
- CN202411884131.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Existing implicit neural representation technology relies on true data when processing non-completely closed models, resulting in high data processing complexity and computational cost; global coding is difficult to take into account complex details, and fixed local coding is difficult to adaptively allocate resources, resulting in limited performance in high-resolution scenarios.
The implicit reconstruction method of 3D scenes based on adaptive sparse encoding guidance is adopted. By obtaining the surface point cloud data in the point cloud data, local coding points with important geometric features are selected, the distance weight between the query point and the coded point is calculated, and the new query point features are fused to generate new query point features. Through network training and loss optimization, the coded point distribution is dynamically adjusted, and the Marching Cubes algorithm is finally used for three-dimensional model reconstruction.
High-precision three-dimensional reconstruction and detail restoration are achieved, taking into account the accuracy of the global structure and the fineness of local details. Through the adaptive mechanism, the model's expression ability and reconstruction accuracy of complex areas are significantly improved.
Smart Images

Figure CN119339023B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method and a device for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance, and belongs to the technical field of three-dimensional reconstruction. Background Art
[0002] 3D reconstruction is widely used in fields such as virtual reality (VR), real-life 3D, medical imaging, and autonomous driving. As an emerging technology, Implicit Neural Representations (INRs) have shown great potential in high-quality 3D reconstruction in recent years. INRs parameterize continuous functions through neural networks and can efficiently model complex geometric shapes and functional features. Compared with traditional discrete representation methods (such as point clouds, voxels, and grids), it has the advantages of high storage efficiency and can accurately capture high-resolution geometric details.
[0003] However, existing implicit neural representation techniques still have the following shortcomings:
[0004] 1. High reliance on true data. Most existing implicit neural representation methods rely on pre-computed true data. However, when dealing with models that are not completely closed, generating true data faces great challenges and often requires additional model completion operations. This not only increases the complexity of early data processing, but also significantly increases computational costs and time consumption.
[0005] Second, global coding cannot take into account complex details. Although the global coding strategy can effectively describe the overall geometric structure, it performs poorly when dealing with areas with high curvature or dense details. This results in the inability to fully capture detailed information and makes it difficult to achieve fine restoration of complex areas. In addition, the coding capacity of simple areas is redundant, further resulting in a waste of computing resources.
[0006] 3. Fixed local coding makes it difficult to adaptively allocate resources. When dealing with models with uneven shape complexity, fixed coding methods often generate redundant coding in simple areas, while in areas with high curvature or dense details, insufficient coding may lead to insufficient expression of details. This makes it difficult to achieve an effective balance between detail capture and computing resource allocation. Although the fixed local coding method enhances the attention to details, due to the lack of adaptive ability, a large amount of redundant coding is generated in simple areas, occupying additional storage and computing resources. In areas with rich details, the expression ability may be limited due to insufficient number of codes. In addition, fixed local coding usually requires high precision to meet modeling requirements, which significantly increases memory overhead and is not conducive to promotion in practical applications.
[0007] Whether it is global encoding or fixed local encoding, it is difficult to achieve a good balance between detail restoration of complex models and resource utilization, resulting in limited performance in high-resolution scenarios. Summary of the invention
[0008] In order to solve the above problems, the present invention proposes a 3D scene implicit reconstruction method and device based on adaptive sparse coding guidance, which can achieve high-precision three-dimensional reconstruction and detail restoration.
[0009] The technical solution adopted by the present invention to solve the technical problem is:
[0010] In a first aspect, an embodiment of the present invention provides a 3D scene implicit reconstruction method based on adaptive sparse coding guidance, comprising the following steps:
[0011] Acquire surface point cloud data in the point cloud data of the 3D scene to be processed, and select local coding points with important geometric features from the surface point cloud data based on curvature;
[0012] Select a query point, calculate the distance between the query point and the local code point, and use the distance value as the weight of the local code point to weight it;
[0013] The weighted coding point features are combined with the query points to generate new query point features, and the new query point features are used as network input for network training;
[0014] During network training, the loss is calculated to optimize network parameters;
[0015] Calculate the similarity of code points and dynamically adjust the code point distribution for network optimization;
[0016] Based on the optimized network and network parameters, the Marching Cubes (MC) algorithm is used to reconstruct the three-dimensional model.
[0017] As a possible implementation of this embodiment, the initial coding information of the local coding point is represented as two matrices: and ,in, represents the latent coding matrix, Represents the local coding point matrix, n represents the number of local coding points used, and m is the dimension of the coding.
[0018] As a possible implementation of this embodiment, the step of selecting a query point, calculating a distance value between the query point and a local code point, and weighting the local code point using the distance value as a weight of the local code point includes the following steps:
[0019] Set a set of query points in three-dimensional space , from the query point set Select query points, the query point set Each row Represents an independent query point;
[0020] Construct query points and code points The distance matrix between :
[0021] (1),
[0022] in, Represents the distance matrix The middle element, represents a local code point, Represents a set of query points The query point in , , represents the number of query points;
[0023] According to the distance matrix A weight matrix W is calculated:
[0024] (2),
[0025] (3);
[0026] Weight matrix and the updated latent coding matrix Perform matrix multiplication to obtain the weighted code point matrix :
[0027] (4),
[0028] in, .
[0029] As a possible implementation of this embodiment, the step of fusing the weighted code point features with the query points to generate new query point features, and using the new query point features as network input for network training includes the following steps:
[0030] Weighted code point matrix Merge with the query point matrix and pass it as input to the network for network training to generate the output vector , where each element Indicates the query point The predicted signed distance of .
[0031] As a possible implementation of this embodiment, during the network training process, the function for calculating the loss is as follows:
[0032] (5),
[0033] (6),
[0034] (7),
[0035] (8),
[0036] (9),
[0037] Where: represents the total loss function, and represents the weight coefficient of each loss, represents the SDF loss function, represents the SDF gradient loss function, represents the Eikonal loss function, represents the weighted coding loss function, and the query point of surface sampling is expressed as , Represents the normal vector.
[0038] As a possible implementation of this embodiment, the calculation of the similarity of code points and the dynamic adjustment of the code point distribution for network optimization include the following steps:
[0039] By code point As the center, select the surrounding Nearby points , and measure these neighbors by Euclidean distance Point and center point latent codes The similarity between:
[0040] (10);
[0041] Determine the mask value of each code point , when the similarity between two code points is lower than the similarity threshold When , the corresponding mask value is assigned is zero, indicating that the code point is redundant or unnecessary information; using the matrix represents the mask matrix, , each mask value in the mask matrix For each code point, ;
[0042] By expanding the mask matrix dimension, making it similar to the latent encoding matrix The dimensions are aligned and expanded by column assignment, so that the mask matrix after expansion Each column of is equal to the mask matrix before expansion The same, the expanded matrix and the latent coding matrix Perform a dot multiplication to update Dynamically remove redundant or similar code points:
[0043] (11),
[0044] in, represents the updated latent coding matrix.
[0045] As a possible implementation of this embodiment, the three-dimensional model reconstruction using the Marching Cubes algorithm based on the optimized network and network parameters includes the following steps:
[0046] Generate the model surface based on the optimized network and network parameters, perform uniform sampling in three-dimensional space, and predict the signed distance value (SDF) of each sampling point;
[0047] The Marching Cubes algorithm is used to process the sampling points and extract the model surface.
[0048] As a possible implementation of this embodiment, the use of the Marching Cubes algorithm to process the sampling points includes the following steps:
[0049] Traverse the cubic unit composed of sampling points and determine the triangle topology structure according to the relationship between the SDF value and the isosurface;
[0050] The model surface is extracted by interpolating the coordinates of the intersection points on the cube units and merging the triangular meshes.
[0051] In a second aspect, an embodiment of the present invention provides a 3D scene implicit reconstruction device based on adaptive sparse coding guidance, comprising:
[0052] A data acquisition module is used to obtain surface point cloud data from the point cloud data of the 3D scene to be processed, and select local coding points with important geometric features from the surface point cloud data based on curvature;
[0053] A weight calculation module is used to select a query point and calculate the distance value between the query point and the local code point, and use the distance value as the weight of the local code point to weight the local code point;
[0054] A feature fusion module is used to fuse the weighted coding point features with the query point to generate new query point features, and use the new query point features as network input for network training;
[0055] The network training module is used to calculate the loss and optimize the network parameters during the network training process;
[0056] Network optimization module, used to calculate the similarity of code points and dynamically adjust the distribution of code points for network optimization;
[0057] The model reconstruction module is used to reconstruct the three-dimensional model based on the optimized network and network parameters using the Marching Cubes algorithm.
[0058] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform any step of the above-mentioned implicit reconstruction method of 3D scenes guided by adaptive sparse coding.
[0059] In a fourth aspect, an embodiment of the present invention provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above-mentioned methods for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance are executed.
[0060] The beneficial effects of the technical solution of the embodiment of the present invention are as follows:
[0061] The present invention combines Eikonal term constraints with sparse coding strategies, taking into account both the accuracy of the global structure and the fineness of local details. The Eikonal term optimizes the reconstruction accuracy of the global model by constraining the gradient of the network; sparse coding focuses on the detailed expression of the local area to ensure the accurate capture of key features in complex geometric areas; the present invention combines the dual strategies so that the model can maintain global consistency and achieve significant improvement in detail restoration. The present invention achieves high-precision three-dimensional reconstruction and detail restoration.
[0062] To further optimize resource allocation, the present invention introduces an adaptive mechanism to dynamically adjust the number of codes according to the geometric complexity of the region, reduce redundant codes in low-complexity regions, and concentrate more computing resources on regions with rich details, effectively improving the model's ability to depict boundary features and surface changes. The resource allocation method of the present invention not only improves coding efficiency, but also achieves more accurate reconstruction in high-detail regions.
[0063] Unlike traditional methods, the present invention only relies on limited model surface point cloud data, and can achieve high-quality reconstruction without additional model completion operations or true value data support. Through reasonable coding allocation and detail capture strategies, the present invention demonstrates significant advantages in reconstruction accuracy and detail restoration capabilities, providing a new solution for the efficient reconstruction of complex three-dimensional scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a flow chart of a method for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance according to an exemplary embodiment;
[0065] Figure 2 is a structural schematic diagram of a 3D scene implicit reconstruction device based on adaptive sparse coding guidance according to an exemplary embodiment;
[0066] Figure 3 It is a specific implementation flow chart of a method for implementing implicit reconstruction of a 3D scene based on adaptive sparse coding guidance using the device of the present invention according to an exemplary embodiment;
[0067] Figure 4 is a flow chart of a Marching Cubes algorithm according to an exemplary embodiment. DETAILED DESCRIPTION
[0068] In order to more clearly illustrate the technical features of the solution of the present invention, the present invention is described in detail below through specific implementation methods and in conjunction with the accompanying drawings.
[0069] In response to the problems that implicit neural representation fails to capture sufficient details in 3D reconstruction tasks and is highly dependent on the true value of data, the present invention proposes a 3D reconstruction method based on implicit neural representation with adaptive sparse coding. High-precision reconstruction and detail restoration are achieved by combining Eikonal term constraints with adaptive sparse coding. It not only takes into account global consistency and local detail performance, but also optimizes resource allocation through an adaptive mechanism, thus breaking through the limitations of existing 3D reconstruction methods.
[0070] like Figure 1 As shown, an embodiment of the present invention provides a 3D scene implicit reconstruction method based on adaptive sparse coding guidance, comprising the following steps:
[0071] Acquire surface point cloud data in the point cloud data of the 3D scene to be processed, and select local coding points with important geometric features from the surface point cloud data based on curvature;
[0072] Select a query point, calculate the distance between the query point and the local code point, and use the distance value as the weight of the local code point to weight it;
[0073] The weighted coding point features are combined with the query points to generate new query point features, and the new query point features are used as network input for network training;
[0074] During network training, the loss is calculated to optimize network parameters;
[0075] Calculate the similarity of code points and dynamically adjust the code point distribution for network optimization;
[0076] Based on the optimized network and network parameters, the Marching Cubes algorithm is used to reconstruct the three-dimensional model.
[0077] As a possible implementation of this embodiment, the initial coding information of the local coding point is represented as two matrices: and ,in, represents the latent coding matrix, Represents the local coding point matrix, n represents the number of local coding points used, and m is the dimension of the coding.
[0078] As a possible implementation of this embodiment, the step of selecting a query point, calculating a distance value between the query point and a local code point, and weighting the local code point using the distance value as a weight of the local code point includes the following steps:
[0079] Set a set of query points in three-dimensional space , from the query point set Select query points, the query point set Each row Represents an independent query point;
[0080] Construct query points and code points The distance matrix between :
[0081] (1),
[0082] in, Represents the distance matrix The middle element, represents a local code point, Represents a set of query points The query point in , , represents the number of query points;
[0083] According to the distance matrix A weight matrix W is calculated:
[0084] (2),
[0085] (3);
[0086] Weight matrix and the updated latent coding matrix Perform matrix multiplication to obtain the weighted code point matrix :
[0087] (4),
[0088] in, .
[0089] As a possible implementation of this embodiment, the step of fusing the weighted code point features with the query points to generate new query point features, and using the new query point features as network input for network training includes the following steps:
[0090] Weighted code point matrix Merge with the query point matrix and pass it as input to the network for network training to generate the output vector , where each element Indicates the query point The predicted signed distance of .
[0091] As a possible implementation of this embodiment, during the network training process, the function for calculating the loss is as follows:
[0092] (5),
[0093] (6),
[0094] (7),
[0095] (8),
[0096] (9),
[0097] Where: represents the total loss function, and represents the weight coefficient of each loss, represents the SDF loss function, represents the SDF gradient loss function, represents the Eikonal loss function, represents the weighted coding loss function, Represents all query points, and the query points of surface sampling are represented as , Represents the normal vector.
[0098] As a possible implementation of this embodiment, the calculation of the similarity of code points and the dynamic adjustment of the code point distribution for network optimization include the following steps:
[0099] By code point As the center, select the surrounding Nearby points , and measure these neighbors by Euclidean distance Point and center point latent codes The similarity between:
[0100] (10);
[0101] Determine the mask value of each code point , when the similarity between two code points is lower than the similarity threshold When , the corresponding mask value is assigned is zero, indicating that the code point is redundant or unnecessary information; using the matrix represents the mask matrix, , each mask value in the mask matrix For each code point, ;
[0102] By expanding the mask matrix dimension, making it similar to the latent encoding matrix The dimensions are aligned and expanded by column assignment, so that the mask matrix after expansion Each column of is equal to the mask matrix before expansion The same, the expanded matrix and the latent coding matrix Perform a dot multiplication to update Dynamically remove redundant or similar code points:
[0103] (11),
[0104] in, represents the updated latent coding matrix.
[0105] As a possible implementation of this embodiment, the three-dimensional model reconstruction using the Marching Cubes algorithm based on the optimized network and network parameters includes the following steps:
[0106] Generate the model surface based on the optimized network and network parameters, perform uniform sampling in three-dimensional space, and predict the signed distance value (SDF) of each sampling point;
[0107] The Marching Cubes algorithm is used to process the sampling points and extract the model surface.
[0108] As a possible implementation of this embodiment, the use of the Marching Cubes algorithm to process the sampling points includes the following steps:
[0109] Traverse the cubic unit composed of sampling points, and determine the triangle topology structure according to the relationship between the SDF value and the isosurface. The isosurface is usually 0.
[0110] The model surface is extracted by interpolating the coordinates of the intersection points on the cube units and merging the triangular meshes.
[0111] like Figure 2 As shown, an embodiment of the present invention provides a 3D scene implicit reconstruction device based on adaptive sparse coding guidance, comprising:
[0112] A data acquisition module is used to obtain surface point cloud data from the point cloud data of the 3D scene to be processed, and select local coding points with important geometric features from the surface point cloud data based on curvature;
[0113] A weight calculation module is used to select a query point and calculate the distance value between the query point and the local code point, and use the distance value as the weight of the local code point to weight the local code point;
[0114] A feature fusion module is used to fuse the weighted coding point features with the query point to generate new query point features, and use the new query point features as network input for network training;
[0115] The network training module is used to calculate the loss and optimize the network parameters during the network training process;
[0116] Network optimization module, used to calculate the similarity of code points and dynamically adjust the distribution of code points for network optimization;
[0117] The model reconstruction module is used to reconstruct the three-dimensional model based on the optimized network and network parameters using the Marching Cubes algorithm.
[0118] like Figure 3 As shown, the specific process of using the device of the present invention to perform implicit reconstruction of a 3D scene based on adaptive sparse coding guidance is as follows.
[0119] Step 1: Input surface point cloud data and select local coding points with important geometric features based on curvature. Areas with higher curvature will be sampled with more coding points to ensure that complex areas have higher resolution coding.
[0120] Based on the curvature, local coding points P with important geometric features are selected. This sampling strategy will sample more coding points in areas with higher curvature, thereby ensuring that complex areas have higher coding density. This not only helps the network to identify and focus on detail areas more quickly during initialization, but also enhances its ability to adapt to shape complexity. The initial coding information is represented as two matrices: and . Where n represents the number of code points used and m is the dimension of the code.
[0121] Step 2: Calculate the distance weight between the query point and the encoding point. For each query point, calculate the distance between it and the encoding point, use the distance value as the weight, and give a larger weight to the closer encoding point. This weight is used to guide the feature fusion of the query point to better capture the local geometric features.
[0122] Set a set of query points in three-dimensional space , where each line represents an independent query point. Next, this paper constructs query points and encoding points The distance matrix between The elements of this matrix are The geometric distance between each query point and the corresponding encoding point is quantified by the following formula, as shown in formula (1):
[0123] (1).
[0124] Then according to A weight matrix W is calculated, and the calculation formula is shown in equations (2) and (3):
[0125] (2),
[0126] (3).
[0127] The invention aims to ensure that the local coding weight far from the query point is small, so the cubic inverse of the distance is taken as the weight and normalized. The weight matrix is used to quantify the distance relationship between the query point and the coding point, reflecting the contribution of the coding point to the query point. and the updated latent coding matrix Perform matrix multiplication to obtain the weighted code point matrix , the formula is as formula (4):
[0128] (4),
[0129] The weight matrix is based on the geometric distance from each query point to all code points. The distance is used as a weight factor to weight the coding information of the code points, giving code points with closer distances greater influence.
[0130] Step 3: The weighted coded point features are fused with the query points to generate new query point features, which are used as the input of the network to ensure that the query points can convey the geometric information of the local area.
[0131] Weighted coded information Merge with the query point feature and pass it to the subsequent network as input to finally generate the output vector , where each element Indicates the query point This method enables the model to dynamically adjust the degree of attention to local areas according to the distance. Since the encoding points are selected at locations with higher curvature, the weight distribution is more inclined to these complex areas, thereby highlighting the information of key areas. This not only improves the model's ability to express complex areas, but also enhances the accurate capture and reconstruction performance of the overall spatial structure.
[0132] Step 4: Network training and loss calculation. During the network training process, the SDF value, normal vector loss, and Eikonal loss are jointly optimized to ensure that the network learns the accurate distance field and normal vector direction.
[0133] The final loss is defined as follows:
[0134] (5),
[0135] (6),
[0136] (7),
[0137] (8),
[0138] (9),
[0139] Where: represents all query points, and the query points of surface sampling are represented by , Represents the normal vector. The query points are mainly obtained by extracting surface points, introducing noise processing in the point cloud, and implementing uniform sampling in space, thus ensuring its wide distribution and representativeness.
[0140] Step 5: Dynamically adjust the distribution of code points. During the training process, the similarity of code points is calculated, and the number of code points with high similarity is gradually reduced to adapt to the changes in regional complexity.
[0141] Step 5: In order to achieve adaptive screening and optimization of code points, the mask matrix is introduced , where each mask value Corresponding to each code point. First, the code point As the center, select the surrounding Nearby points , and measure these neighbors by Euclidean distance Point and center point latent codes The similarity between them is calculated by the formula (10):
[0142] (10);
[0143] According to the preset similarity threshold , determine the mask value of each code point , when the similarity between two code points is lower than the threshold, the corresponding mask value is assigned to zero ( ), indicating that the code point is redundant or unnecessary information. Finally, by expanding the mask matrix Update the latent coding matrix By dynamically removing redundant or similar coding points, the updated latent coding matrix is obtained. The calculation formula is shown in formula (11):
[0144] ,
[0145] Through this masking mechanism, when the similarity of some coding points is lower than the preset threshold, they will be adaptively adjusted through masking operations to guide the network to focus more on learning key coding points. The dynamic adjustment of the mask matrix enables the model to flexibly allocate coding points according to their importance, thereby achieving adaptive elimination of invalid or redundant coding points and optimizing the overall expression ability of the model.
[0146] Step 6: After multiple rounds of iterative training, the Marching Cubes algorithm is used to reconstruct the model in three dimensions.
[0147] Based on the training network to generate the model surface, we first uniformly sample points in the three-dimensional space and predict the signed distance value (SDF) of each sampling point. Then, Figure 4 As shown in the figure, the Marching Cubes algorithm is used to process these points: the cubic unit composed of the sampling points is traversed, and the triangle topology structure is determined by looking up the table according to the relationship between the SDF value and the isosurface (usually 0), the intersection coordinates are calculated by interpolation, and the triangle meshes are generated and merged to extract the model surface.
[0148] Compared with the prior art, the present invention has the following characteristics:
[0149] 1) A reconstruction method that takes into account both global and local details is adopted. By introducing Eikonal term constraints, the global model is finely optimized, which enhances the consistency of the global structure. The sparse coding strategy focuses on local complex areas and key geometric details to ensure the integrity of the detailed expression. This dual strategy breaks through the limitation of traditional methods that are difficult to balance between global consistency and local refinement.
[0150] 2) Adopting an adaptive coding resource allocation mechanism: The present invention significantly reduces coding redundancy in low-complexity areas by dynamically adjusting the number of codes and allocating computing resources according to the geometric complexity of the area.
[0151] 3) High-quality 3D reconstruction based on limited point cloud data: The present invention does not need to rely on data truth or additional model completion operations, and can achieve high-quality 3D reconstruction based only on limited surface point clouds.
[0152] An electronic device provided by an embodiment of the present invention includes a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform any step of the above-mentioned implicit reconstruction method of a 3D scene guided by adaptive sparse coding.
[0153] Specifically, the above-mentioned memory and processor can be general-purpose memory and processor, which are not specifically limited here. When the processor runs the computer program stored in the memory, the above-mentioned implicit reconstruction method of 3D scene guided by adaptive sparse coding can be executed.
[0154] Those skilled in the art will appreciate that the structure of the electronic device does not limit the electronic device and may include more or fewer components than shown in the figure, or combine or split certain components, or arrange the components differently.
[0155] In some embodiments, the electronic device may also include a touch screen that can be used to display a graphical user interface (e.g., a startup interface of an application) and receive user operations on the graphical user interface (e.g., startup operations on an application). The specific touch screen may include a display panel and a touch panel. The display panel may be configured in the form of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc. The touch panel may collect the user's contact or non-contact operations on or near it, and generate pre-set operation instructions, for example, the user uses any suitable object such as a finger, stylus, or accessories on or near the touch panel. In addition, the touch panel may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and posture, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into information that the processor can process, and then sends it to the processor, and can receive and execute commands from the processor. In addition, the touch panel can be implemented by various types such as resistive, capacitive, infrared and surface acoustic wave, and any technology developed in the future can also be used to implement the touch panel. Further, the touch panel can cover the display panel, and the user can operate on or near the touch panel covered on the display panel according to the graphical user interface displayed on the display panel. After the touch panel detects the operation on or near it, it is transmitted to the processor to determine the user input, and then the processor provides corresponding visual output on the display panel in response to the user input. In addition, the touch panel and the display panel can be implemented as two independent components or integrated.
[0156] Corresponding to the method for starting the above-mentioned application, an embodiment of the present invention further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned implicit reconstruction methods of 3D scenes guided by adaptive sparse coding are executed.
[0157] The application startup device provided in the embodiment of the present application can be specific hardware on the device or software or firmware installed on the device. The device provided in the embodiment of the present application, its implementation principle and the technical effect produced are the same as those in the aforementioned method embodiment. For the sake of brief description, the parts not mentioned in the device embodiment can refer to the corresponding contents in the aforementioned method embodiment. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiment, and will not be repeated here.
[0158] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0159] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0160] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0161] In addition, each functional module in the embodiments provided in the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0162] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0163] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0164] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A 3D scene implicit reconstruction method based on adaptive sparse coding guidance, characterized in that: The steps include: Acquire surface point cloud data in the point cloud data of the 3D scene to be processed, and select local coding points with important geometric features from the surface point cloud data based on curvature; Select a query point, calculate the distance between the query point and the local code point, and use the distance value as the weight of the local code point to weight it; The weighted coding point features are combined with the query points to generate new query point features, and the new query point features are used as network input for network training; During network training, the loss is calculated to optimize network parameters; Calculate the similarity of code points and dynamically adjust the code point distribution for network optimization; Based on the optimized network and network parameters, the Marching Cubes algorithm is used to reconstruct the three-dimensional model.
2. The method for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance according to claim 1, characterized in that: The initial coding information of the local coding point is represented by two matrices: and ,in, represents the latent coding matrix, Represents the local coding point matrix, n represents the number of local coding points used, and m is the dimension of the coding.
3. The method for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance according to claim 2, characterized in that: The step of selecting a query point, calculating the distance between the query point and the local code point, and weighting the local code point using the distance as the weight of the local code point comprises the following steps: Set a set of query points in three-dimensional space , from the query point set Select query points, the query point set Each row Represents an independent query point; Construct query points and code points The distance matrix between : (1), in, Represents the distance matrix Medium element, represents a local code point, Represents a set of query points The query point in , , represents the number of query points; According to the distance matrix A weight matrix W is calculated: (2), (3); Weight matrix and the latent coding matrix Perform matrix multiplication to obtain the weighted code point matrix : (4), in, .
4. The method for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance according to claim 3, characterized in that: The step of fusing the weighted code point features with the query points to generate new query point features, and using the new query point features as network input for network training includes the following steps: Weighted code point matrix Merge with the query point matrix and pass it as input to the network for network training to generate the output vector , where each element Indicates the query point The predicted signed distance of .
5. The method for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance according to claim 4, characterized in that: During network training, the function for calculating loss is as follows: (5), (6), (7), (8), (9), Where: represents the total loss function, and represents the weight coefficient of each loss, represents the SDF loss function, represents the SDF gradient loss function, represents the Eikonal loss function, represents the weighted coding loss function, and the number of query points sampled on the surface is expressed as , Represents the normal vector.
6. The method for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance according to claim 5, characterized in that: The method of calculating the similarity of code points and dynamically adjusting the distribution of code points for network optimization includes the following steps: By code point As the center, select the surrounding Nearby points , and measure the encoding of these adjacent points by Euclidean distance and the center point latent code The similarity between: (10); Determine the mask value of each code point , when the similarity between two code points is lower than the similarity threshold When , the corresponding mask value is assigned is zero, indicating that the code point is redundant or unnecessary information; using the matrix represents the mask matrix, , each mask value in the mask matrix For each code point, ; By expanding the mask matrix dimension, making it similar to the latent encoding matrix The dimensions are aligned and expanded by column assignment, so that the mask matrix after expansion Each column of is equal to the mask matrix before expansion The same, the expanded matrix and the latent coding matrix Perform point multiplication for updating and dynamically remove redundant or similar encoding points: (11), in, represents the updated latent coding matrix.
7. The method for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance according to any one of claims 1 to 6, characterized in that: The method of reconstructing a three-dimensional model based on the optimized network and network parameters using the Marching Cubes algorithm includes the following steps: Generate the model surface based on the optimized network and network parameters, perform uniform sampling in three-dimensional space, and predict the signed distance value of each sampling point; The Marching Cubes algorithm is used to process the sampling points and extract the model surface.
8. A 3D scene implicit reconstruction device based on adaptive sparse coding guidance, characterized in that: include: A data acquisition module is used to obtain surface point cloud data from the point cloud data of the 3D scene to be processed, and select local coding points with important geometric features from the surface point cloud data based on curvature; A weight calculation module is used to select a query point and calculate the distance value between the query point and the local code point, and use the distance value as the weight of the local code point to weight the local code point; A feature fusion module is used to fuse the weighted coding point features with the query point to generate new query point features, and use the new query point features as network input for network training; The network training module is used to calculate the loss and optimize the network parameters during the network training process; Network optimization module, used to calculate the similarity of code points and dynamically adjust the distribution of code points for network optimization; The model reconstruction module is used to reconstruct the three-dimensional model based on the optimized network and network parameters using the Marching Cubes algorithm.
9. An electronic device, characterized in that: The electronic device comprises a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the method for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, executes the steps of the method for implicitly reconstructing a 3D scene based on adaptive sparse coding guidance as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Implicit three-dimensional scene characterization method based on multilayer dynamic feature point cloud
CN115512077A
Composite three-dimensional curved surface reconstruction method
CN118537506A