A 3D reconstruction method for coarse aggregate based on implicit neural model

Through the implicit neural model and multi-view image acquisition system, the data incomplete, high equipment cost and operational complexity of the three-dimensional reconstruction of coarse aggregate particles in the prior art are solved, and high-precision and efficient three-dimensional reconstruction are achieved, and high-quality visual effects are generated.

CN120259564BActive Publication Date: 2025-08-26CHANGAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510737898.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-26
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The prior art has problems such as incomplete data acquisition, high equipment cost, complex operation, low scanning efficiency, environmental radiation risks and complex data processing in the three-dimensional reconstruction of coarse aggregate particles, making it difficult to achieve high-precision and efficient three-dimensional reconstruction.

Method used

Using an implicit neural model method, a geometric and texture model is established by constructing an unobstructed multi-view image acquisition system, combining multi-view image acquisition and implicit neural representation, and using spherical harmonic function coding and volume rendering technology to achieve high-precision reconstruction of three-dimensional contour images.

Benefits of technology

It improves reconstruction accuracy and efficiency, reduces computational complexity, enhances model adaptability and visual effects, avoids occlusion problems and complex operations of traditional methods, and generates high-quality three-dimensional reconstruction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259564B_ABST
    Figure CN120259564B_ABST
Patent Text Reader

Abstract

The present invention relates to the fields of computer vision and three-dimensional modeling technology, and in particular to a method for three-dimensional reconstruction of coarse aggregate based on an implicit neural model. An unobstructed multi-view image acquisition system is constructed to collect three-dimensional contour image data of coarse aggregate particles; spatial sampling points associated with the coarse aggregate particles in the three-dimensional contour image are obtained; a geometric model and a texture model are respectively established based on the implicit neural model; the spatial sampling points are input into the geometric model to obtain the signed distance function value of the spatial sampling points; the input vector of the texture model is obtained according to the signed distance function value combined with spherical harmonic function encoding and input into the texture model to obtain the RGB value of the spatial sampling point; a volume rendering model is used to perform image synthesis based on the signed distance function value and the RGB value of the spatial sampling point to obtain the RGB value of the pixel corresponding to the spatial sampling point in the three-dimensional contour image. The present invention can achieve high-precision reconstruction of the three-dimensional geometric morphology of coarse aggregate particles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and three-dimensional modeling technology, and in particular to a coarse aggregate three-dimensional reconstruction method based on an implicit neural model. Background Art

[0002] As global infrastructure construction continues to advance, highway transportation, as a key area of ​​national infrastructure development, plays a fundamental and leading role in supporting national land development, logistics and transportation, social interaction, and international exchanges. Asphalt pavement has become the preferred form of highway construction in my country due to its ease of construction and maintenance, comfortable driving, low noise, and excellent drainage performance. Aggregate, as an indispensable component of pavement, serves as both a skeleton and filler, directly impacting pavement performance. In the wave of digital transformation, digital measurement and analysis of the geometric characteristics of coarse aggregate particles are of great significance for improving the quality and efficiency of pavement design, construction and maintenance. Digital technology can measure and analyze the shape, size, surface texture and other characteristics of coarse aggregate particles. These characteristics can be used to improve the quality and service life of pavement during the design, construction and maintenance stages of pavement. In addition, the study of the geometric characteristics and particle size distribution of coarse aggregate particles is of great significance in many fields such as materials science, civil engineering and geology, including mechanical properties, concrete and asphalt mixture design, particle stacking and filling efficiency, soil and rock engineering, classification and property evaluation of geological materials, environmental impact assessment, material processing and manufacturing, scientific research and education, and technical specifications and standards formulation.

[0003] Existing technologies for digitally measuring and analyzing the geometric morphology of coarse aggregate particles rely primarily on a variety of 3D scanning and imaging techniques. These include structured light 3D scanning, binocular stereo imaging, and computed tomography (CT). Structured light 3D scanning acquires 3D information about an object by projecting a specific light pattern and analyzing its deformation. Binocular stereo imaging uses two or more cameras to capture images from different angles and reconstructs the 3D shape through parallax calculations. CT scanning uses X-rays to penetrate an object and acquire data from multiple angles to reconstruct its internal structure. While these technologies have achieved some success in acquiring 3D data on coarse aggregate particles, they each have limitations.

[0004] (1) Data acquisition integrity and accuracy issues: Due to light obstruction and viewing angle limitations, existing technologies may not be able to fully capture all the geometric features of coarse aggregate particles, resulting in incomplete data and inaccurate shape reconstruction;

[0005] (2) Equipment cost and operation complexity: High-precision 3D scanning equipment such as CT scanning and some laser scanning equipment are expensive and complex to operate, which limits their widespread promotion in industrial applications;

[0006] (3) Scanning efficiency and real-time issues: When processing a large amount of coarse aggregate particles, the existing technology may not be able to meet the real-time or near-real-time requirements for scanning and data processing, affecting production efficiency;

[0007] (4) Environmental and health safety issues: CT scanning involves radiation exposure, which poses potential risks to operators and the environment;

[0008] (5) Complexity of data processing and analysis: The three-dimensional data obtained from existing technologies require complex post-processing and analysis to extract the geometric features of the particles, which increases the difficulty of applying the technology;

[0009] (6) Dependence on specific conditions: Some technologies, such as binocular stereo imaging, have certain requirements for surface texture. For coarse aggregate particles with unclear surface features, it may not be possible to obtain accurate three-dimensional information. Summary of the Invention

[0010] In order to solve the problems existing in the prior art, the present invention provides a method for three-dimensional reconstruction of coarse aggregate based on an implicit neural model. The method collects three-dimensional contour image data of coarse aggregate particles by constructing an unobstructed multi-view image acquisition system; obtains spatial sampling points of coarse aggregate particles in the three-dimensional contour image; establishes a geometric model and a texture model based on the implicit neural model; inputs the spatial sampling points into the geometric model to obtain the signed distance function value of the spatial sampling points; obtains the input vector of the texture model according to the signed distance function value combined with spherical harmonic function encoding and inputs it into the texture model to obtain the RGB value of the spatial sampling point; uses a volume rendering model to synthesize the signed distance function value and RGB value of the spatial sampling point to obtain the RGB value of the corresponding pixel of the spatial sampling point in the three-dimensional contour image. The present invention can achieve high-precision reconstruction of the three-dimensional geometric morphology of coarse aggregate particles.

[0011] The present invention adopts the following technical solution: a coarse aggregate 3D reconstruction method based on implicit neural model,

[0012] Construct an unobstructed multi-view image acquisition system to collect 3D contour image data of coarse aggregate particles;

[0013] Obtaining spatial sampling points of coarse aggregate particles in a three-dimensional contour image;

[0014] Based on the implicit neural model, the geometric model and texture model are established respectively;

[0015] Inputting the spatial sampling points into the geometric model to obtain signed distance function values ​​of the spatial sampling points;

[0016] Acquire an input vector of a texture model according to the signed distance function value combined with spherical harmonics encoding;

[0017] Input the input vector into the texture model to obtain the RGB value of the spatial sampling point;

[0018] The signed distance function value and RGB value of the spatial sampling point are synthesized using a volume rendering model to obtain the final RGB value of the corresponding pixel of the spatial sampling point in the three-dimensional contour image.

[0019] Furthermore, the unobstructed multi-view image acquisition system comprises: a coarse aggregate falling control device, an optical fiber falling detection sensor and a multi-view image acquisition system body;

[0020] The multi-view image acquisition system mainly includes: multiple industrial cameras distributed around, COB lighting sources, uniform light balls, multi-channel 10G switches and a control computer.

[0021] Furthermore, the method for collecting three-dimensional contour image data of coarse aggregate particles is specifically as follows:

[0022] Setting a plurality of different exposure times to collect three-dimensional contour image data of the coarse aggregate particles;

[0023] The Laplace operator is used to evaluate the blurriness of the 3D contour image data corresponding to different exposure times and determine the optimal exposure time.

[0024] Obtain three-dimensional contour image data of coarse aggregate particles collected at the optimal exposure time.

[0025] Furthermore, the spatial sampling points of the coarse aggregate particles in the three-dimensional contour image are obtained, specifically:

[0026] The bounding box size of the 3D reconstruction is established according to the maximum size of the coarse aggregate particles;

[0027] Sending rays to the bounding box using an unobstructed multi-view image acquisition system;

[0028] Ray marching technology is used to select coordinate points on the ray to obtain the spatial sampling points of coarse aggregate particles in the three-dimensional contour image.

[0029] Furthermore, the geometric model is established by: using a multi-resolution hash coding technique combined with a multi-layer perceptron to construct the geometric model;

[0030] The texture model is established by using two multi-layer perceptrons to construct the texture model; each multi-layer perceptron contains 64 hidden neurons.

[0031] Furthermore, the signed distance function value of the spatial sampling point is obtained, specifically:

[0032] The geometric model uses a multi-resolution hash coding technique to input code to the spatial sampling points to obtain a coding vector;

[0033] Concatenate the encoding vector with the coordinates of the spatial sampling points to obtain a 35-dimensional input vector;

[0034] The 35-dimensional input vector is subjected to feature extraction by a multi-layer perceptron to obtain the signed distance function value of the spatial sampling point.

[0035] Furthermore, the input vector of the texture model is obtained according to the signed distance function value combined with the spherical harmonic function encoding, specifically:

[0036] Obtain its eigenvector and the normal vector of the spatial sampling point according to the signed distance function value output by the geometric model;

[0037] The 35-dimensional input vector of the geometric model is encoded using a 4th-order spherical harmonic function to obtain a 16-dimensional vector;

[0038] The normal vector, the eigenvector and the 16-dimensional vector are fused to obtain an input vector of a texture model.

[0039] Furthermore, before the volume rendering model is used to synthesize the signed distance function value and the RGB value of the spatial sampling point, the method further includes: converting the signed distance function value of the spatial sampling point into a volume density, which is expressed as:

[0040] ;

[0041] in, Represents spatial sampling points The signed distance function value at represents the signed distance function, is the offset parameter, is a learnable scale parameter that determines the sensitivity of volume density to changes in the signed distance function value.

[0042] Furthermore, when the geometric model and texture model are established based on the implicit neural model, the following steps are also included:

[0043] The total loss function of the implicit neural model is set to the weighted sum of RGB loss, mask loss, Eikonal loss and curvature loss;

[0044] The expression of the RGB loss is:

[0045] ;

[0046] in, is the RGB loss, represents a rendered image, represents the real image, Represents the difference in pixel value between the rendered image and the real image at the spatial sampling point i, is the 2-norm;

[0047] The expression of the mask loss is:

[0048] ;

[0049] in, is the mask loss, is the predicted opacity value at spatial sampling point i, is the true binary mask at the spatial sampling point i;

[0050] The expression of the Eikonal loss is:

[0051] ;

[0052] in, For Eikonal losses, At the spatial sampling point The gradient of the signed distance function at , represents the gradient operator;

[0053] The curvature loss is expressed as:

[0054] ;

[0055] in, is the curvature loss, At the spatial sampling point The mean curvature of the signed distance function at .

[0056] The beneficial effects of the present invention are as follows: by integrating unobstructed multi-view image acquisition technology and a 3D contour refinement reconstruction method based on implicit neural representation, the present invention achieves high-precision reconstruction of the 3D geometric form of coarse aggregate particles, achieving significant improvements and enhancements in the following aspects:

[0057] 1. Improved reconstruction accuracy and efficiency: The implicit neural representation method used in this paper can accurately restore the three-dimensional contours of coarse aggregate particles from multi-view images, avoiding the occlusion problem and tedious camera calibration process in traditional methods, and significantly improving the reconstruction accuracy and efficiency;

[0058] 2. Reduced computational complexity: The present invention reduces the number of parameters of the implicit neural network model by adopting multi-resolution hash coding and linear interpolation technology, while maintaining the reconstruction quality and reducing computational complexity;

[0059] 3. Enhanced model adaptability: The method of the present invention can encode sampling points at different scales, capturing local details and global structural features of the particle surface, and enhancing the model's adaptability to particles of different sizes;

[0060] 4. Improved visual effects and realism: Neural rendering technology simulates the interaction between light and the surface of objects through a deep learning model to generate high-quality visual effects and learns the reflective properties of particle surfaces, enhancing the realism of the reconstruction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0062] Figure 1 Schematic diagram of a process for 3D reconstruction of coarse aggregate based on an implicit neural model according to an embodiment of the present invention;

[0063] Figure 2 This is a schematic diagram of an overview of the main body of a spherical cavity structure multi-view image acquisition system according to an embodiment of the present invention;

[0064] Figure 3 A schematic diagram of camera layering in a multi-view image acquisition system for a spherical cavity structure according to an embodiment of the present invention;

[0065] Figure 4 A schematic structural diagram of a geometric model and a texture model according to an embodiment of the present invention;

[0066] Figure 5 A schematic diagram of the principle of volume rendering according to an embodiment of the present invention;

[0067] Figure 6 A schematic diagram of a typical loss curve during training of a 3D reconstruction network model according to an embodiment of the present invention;

[0068] Figure 7 A three-dimensional reconstruction error heat map according to an embodiment of the present invention;

[0069] Figure 8 A schematic diagram comparing the results of three-dimensional reconstruction using different methods according to an embodiment of the present invention;

[0070] Figure 9 A heat map of mesh error of three-dimensional reconstruction of coarse aggregate particles according to an embodiment of the present invention;

[0071] Figure 10Schematic diagram comparing quantified results of three-dimensional reconstruction accuracy using different methods according to an embodiment of the present invention. DETAILED DESCRIPTION

[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0073] A flow chart of a method for 3D reconstruction of coarse aggregate based on an implicit neural model according to an embodiment of the present invention is shown in FIG. Figure 1 Shown, including:

[0074] Construct an unobstructed multi-view image acquisition system to collect 3D contour image data of coarse aggregate particles;

[0075] The unobstructed multi-view image acquisition system constructed in the embodiment of the present invention includes a coarse aggregate falling control device, an optical fiber falling detection sensor and a spherical cavity structure multi-view image acquisition system body. The system body overview is shown in FIG. Figure 2 As shown in the figure, it consists of multiple industrial cameras distributed in a circle, COB lighting sources, uniform light balls, multi-channel 10G switches and control computers. Its working process is as follows: the drop control device releases coarse aggregate particles from the top of the equipment, allowing them to fall freely along the central axis of the equipment; when the particles pass through the top window of the imaging device, the fiber optic sensor generates a trigger signal and sends it synchronously to each industrial camera. After receiving the trigger signal, the camera starts the delayed trigger function and sets the delay time according to the time required for the particles to fall freely. When the particles fall exactly to the center of the device, all industrial cameras are exposed at the same time to capture images of the coarse aggregate particles from different perspectives.

[0076] The number of cameras and their viewing angle distribution in the spherical cavity structure multi-view imaging system of the embodiment of the present invention determine the completeness of the representation of the target's complete contour information by the collected image samples. The specific camera layering diagram is shown in FIG. Figure 3 As shown in Figure 2, the multi-view image acquisition device designed in this embodiment of the present invention is equipped with 16 industrial cameras to balance 3D reconstruction accuracy and cost. The cameras are distributed at four "latitudes" with equal angular differences on the surface of a sphere. Each layer has four cameras spaced at equal angles, and the cameras in adjacent layers are staggered at 45°. All 16 cameras maintain a viewpoint facing the center of the sphere.

[0077] The core function of the optical imaging system is that multiple cameras simultaneously capture the coarse aggregate particles falling to the center of the device, and the two-dimensional contours of the coarse aggregate particles in each captured perspective image are complete and the details are clear. The system further selected 16 JHEM131GM global exposure black and white industrial cameras with the same parameters, equipped with high-definition distortion-free lenses to meet the requirements of high-quality imaging of fast-moving target objects in a specific field of view; illumination compensation design is a key factor in achieving high-quality and robust machine vision applications. The system proposed in the embodiment of the present invention combines backlighting and diffuse lighting. By installing a high-brightness COBLED light source board on the inner wall of the carbon fiber board and embedding a white acrylic uniform light ball in the device, the uniformity of the illumination is optimized to ensure high-quality multi-perspective images within a short exposure time.

[0078] Each industrial camera in the system is equipped with a GigE Gigabit Ethernet interface and connected to an H3C S1226FX switch via a star topology to reduce data transmission delays and network bottlenecks. The switch is connected to the server via two 10G fiber optic uplink data ports and adopts QoS policies to prioritize different types of data streams, ensuring that critical image data is transmitted first, thereby improving the overall transmission efficiency and stability of the system.

[0079] When acquiring the three-dimensional contour image of the coarse aggregate particles, the free-fall motion process of the coarse aggregate particles is analyzed to determine the motion parameters of the particles in the device. By calculating the total height, required time, and instantaneous velocity of the particles falling from the drop control device to the center of the device, a particle motion model is established. This model, combined with the physical scale corresponding to the unit pixel on the camera imaging target surface, is used to calculate the time required for the particles to move per unit pixel on the imaging target surface, providing a theoretical basis for the subsequent optimization of the exposure time. The exposure time of all industrial cameras in the multi-view image capture system is further set to multiple different values ​​between 15us and 200us. Multi-view imaging is performed on the same standard sphere, and the Laplace operator is used to calculate the edge response of each pixel in the image. The optimal exposure time is determined by comparing the blur of the multi-view images collected at different exposure times. The Laplace operator calculation formula for the image is shown below:

[0080] ;

[0081] in, Indicates that the image is The pixel value at and Represent the second-order partial derivatives of the image in the x and y directions, respectively. is the Laplace response of the image. For clear edges, the Laplace response has a larger absolute value because the grayscale value of the image changes dramatically at the edge. Therefore, the fuzziness evaluation index B is defined as the variance of the Laplace response of the image. The calculation formula is shown below:

[0082] ;

[0083] in, is the total number of pixels in the image, represents the coordinates of the i-th pixel, represents the Laplace operator, Representing an image The Laplace response value at pixel i, is an image The Laplace response mean of the pixel points in the image is calculated, and the blur of the multi-view images collected at different exposure times is compared. By reasonably setting the parameters of the multi-view image capture system, the target edges in the collected images are ensured to be as clear as possible.

[0084] Based on the blur evaluation results obtained in the above steps, a scatter plot of exposure time and image blur was plotted. The trend of image blur changing with exposure time was analyzed. It was determined that the target edges in the captured image were clearest at an exposure time of 150 μs. Taking into account the displacement of particles during exposure and the influence of the amount of light entering the lens, the performance of the system imaging components and the illumination compensation device were comprehensively evaluated. The optimal exposure time was determined to be 150 μs. This ensures that the target edges in the captured image are as clear as possible while avoiding overexposure and dark images.

[0085] Obtaining spatial sampling points of coarse aggregate particles in a three-dimensional contour image;

[0086] After obtaining the three-dimensional contour image, the embodiment of the present invention sets the bounding box size of the three-dimensional reconstructed scene according to the maximum size of the coarse aggregate particle sample. After the initial setting is completed, a ray is emitted from the camera principal point in the spherical cavity structure multi-view imaging system, passing through specific pixels in the image, and directed into the scene bounding box. Through the ray marching technology, a series of spatial sampling points are selected on the ray according to the established strategy.

[0087] Based on the implicit neural model, the geometric model and texture model are established respectively;

[0088] The structural diagram of the geometric model and texture model in the embodiment of the present invention is as follows Figure 4As shown in the figure, the geometric model is first initialized. The model consists of a multi-layer perceptron (MLP) and is used to predict the signed distance function (SDF) value of the sampling point in the three-dimensional space. The model input is the coordinate value of the spatial sampling point. The input is encoded through multi-resolution hash coding technology, and the encoded feature vector is spliced ​​with the original coordinate to construct an input vector with a dimension of 35 to enhance the model's perception of the spatial structure. The structural block diagram is shown in the figure. Figure 4 On the left, the SDF is implicitly expressed as follows:

[0089] ;

[0090] in, represents three-dimensional Euclidean space, is a point in three-dimensional space, represents the signed distance function, the set Contains all points that make the signed distance function value 0, that is, the surface of the object, that is, the points that satisfy the above formula The SDF value quantifies the spatial distance between a point and the surface of the object. Points inside the surface have negative values, while points outside the surface have positive values.

[0091] The texture model consists of a two-layer MLP, each layer contains 64 hidden neurons, which are used to predict the RGB color value of the sampling point. The input data of this part consists of three key parts: the normal vector (Normal) of the sampling point predicted by the geometric network, the 13-dimensional feature vector output by the geometric network, and the use of the fourth-order spherical harmonic function (Spherical Harmonics) to encode the input 2-dimensional observation angle to generate a 16-dimensional vector; the fusion of these three parts forms a 32-dimensional input vector, which provides a multi-dimensional feature space for the model, thereby improving the realism and dynamic range of the rendering results, enabling the texture model to accurately predict the color value. The texture model model is as follows: Figure 4 Shown in the right half.

[0092] Inputting the spatial sampling points into the geometric model to obtain signed distance function values ​​of the spatial sampling points;

[0093] Specifically, the embodiment of the present invention inputs the collected multi-view image data into the implicit neural representation model to perform implicit encoding of the three-dimensional geometric structure. For each coarse aggregate particle captured in the image, a spatial sampling point in the three-dimensional space is It is input into a multi-layer perceptron (MLP), which maps these coordinate points to the corresponding signed distance function (SDF) values ​​through MLP learning, expressed as:

[0094] ;

[0095] in, is the spatial sampling point The signed distance function value to the particle surface is negative if the point is inside the surface and positive if it is outside the surface.

[0096] The input vector of the texture model is obtained according to the signed distance function value combined with the spherical harmonic function encoding; the input vector is input into the texture model to obtain the RGB value of the spatial sampling point;

[0097] The embodiment of the present invention utilizes the output of the geometric model and the view direction encoded by the spherical harmonic function to input them into the texture MLP to predict the RGB value of the current spatial sampling point. The expression is as follows:

[0098] ;

[0099] in, represents texture MLP, is the RGB value of the predicted spatial sampling point, is the viewing direction information encoded by spherical harmonics, is the output of the geometric model, i.e., the spatial sampling point of value.

[0100] The signed distance function value and RGB value of the spatial sampling point are synthesized using a volume rendering model to obtain the final RGB value of the corresponding pixel of the spatial sampling point in the three-dimensional contour image.

[0101] Before synthesis through the volume rendering model, the SDF value output by the geometric model is converted to a volume density value for volume rendering. The conversion process is shown in the following formula:

[0102] ;

[0103] in, is the sigmoid function, represents the volume density of spatial sampling point i, Represents spatial sampling points The corresponding SDF value is, Indicates the spatial sampling points The corresponding SDF value.

[0104] The embodiment of the present invention corrects the deviations in traditional methods through this conversion method, thereby achieving more accurate surface reconstruction. By using SDF to represent the object surface and applying advanced rendering technology, it performs well in reconstructing detailed three-dimensional surfaces from two-dimensional images.

[0105] The embodiment of the present invention further uses volume rendering technology to combine the SDF and RGB values ​​of all sampling points along a ray to render the final RGB value of the corresponding pixel in the synthesized image. Volume rendering uses image rendering technology to synthesize a new image with implicitly represented three-dimensional geometric structure information of particulate matter at a specified camera perspective in a scene. Then, the difference between the rendered image and the input multi-view image is quantified and compared, and a loss function is constructed. Finally, through optimization methods such as backpropagation, supervised training of the three-dimensional geometric structure of particulate matter is achieved without relying on strong supervisory information such as scene depth maps. When the volume rendering model is applied to three-dimensional reconstruction or new-view image synthesis application scenarios, the view position is usually used as the starting point to construct a ray passing through the image pixel point. When the ray passes through the three-dimensional scene space pre-defined according to the target object or scene scale, a number of sampling points are selected on the ray according to specific rules. The associated data at all sampling points are accumulated to obtain the RGB color of the pixel point in the synthesized image corresponding to the ray. In this way, volume rendering associates the pixels in the input multi-view image with the pixels in the synthesized image. A schematic diagram of the principle of volume rendering is provided in the embodiment of the present invention as follows: Figure 5 Specifically, the volume rendering formula is as follows:

[0106] ;

[0107] in, It's a pixel The final RGB value of From camera to pixel The ray path of is the color value of the point on the ray path, is the density value at a point on the ray path, representing the attenuation of light.

[0108] The embodiment of the present invention further constructs a loss function to measure the difference between the synthesized image and the input multi-view image, and performs a backpropagation algorithm on the entire model to update the trainable parameters in the model.

[0109] Therefore, the embodiment of the present invention learns and reconstructs high-precision three-dimensional geometric structures from multi-view images through an implicit neural model, and can also predict the texture information of the surface of coarse aggregate particles.

[0110] In another specific embodiment of the present invention:

[0111] When encoding the spatial sampling points output geometric model, the multi-resolution hash coding technology is used to encode the sampling points in space, specifically:

[0112] First, the continuous 3D scene space is divided into three-dimensional grids of varying resolutions, from coarse to fine. Each vertex of the grid cube is associated with a fixed-length eigenvalue. When a sampling point falls into a grid cube, the eigenvectors associated with the cube mesh vertex are processed using trilinear interpolation based on the distance between the sampling point and the cube vertex, resulting in the eigenvector of the sampling point at the corresponding resolution.

[0113] The multi-resolution hash-encoded feature vector is concatenated with the original coordinates to construct an input vector with a dimension of 35 to enhance the model's perception of spatial structure. The hash encoding expression is shown in the following formula:

[0114] ;

[0115] in, Indicates bitwise exclusive OR operation. Bitwise exclusive OR operation is to perform exclusive OR operation on the corresponding bits of two binary numbers. The same is 0, and different is 1. Represents the unique large prime numbers corresponding to the three dimensions of the integer coordinates of the input mesh vertices, The components representing the three-dimensional coordinate values ​​of the input mesh vertices, the superscript d represents the dimension, and the subscript represents the index of the data point, The modulo operation is to take the remainder after dividing the result of the XOR operation by T. T is a positive integer, which is usually used to limit the hash value to a specific range, that is, between 0 and T-1. This formula is implemented by performing an XOR operation on the linear congruential (pseudo-random) permutation result of each dimension, thereby removing the influence of the dimension on the hash value. In order to achieve (pseudo) independence, only two of the three dimensions of the input mesh vertex integer coordinates need to be permuted. Therefore, in the embodiment of the present invention, , as well as To improve cache coherence.

[0116] The occupancy grid estimator is used to select sampling points for rays. The occupancy grid estimator is based on caching the scene density in a binary voxel grid, allowing to effectively remove empty areas when the ray passes through a predefined grid with a set step size.

[0117] During the ray sampling process, the distribution of sampling points is adjusted according to the occupancy status of the voxels. The proportion of sampling points in occupied grids is significantly higher than that in unoccupied grids, thereby ensuring that more sampling points are concentrated near the object surface, which has an important contribution to the rendering results. In the embodiment of the present invention, the sampling points in the occupied grids may account for 70% to 80% of the total number of ray sampling points, while the sampling points in the unoccupied grids account for only 20% to 30%.

[0118] The embodiment of the present invention uses this multi-resolution hash coding and importance strategy spatial sampling point optimization method to effectively improve the efficiency of neural implicit representation three-dimensional reconstruction model training and reasoning, reduce the probability of weakly correlated sampling points being extracted, and thus improve rendering efficiency and accuracy.

[0119] In another specific embodiment of the present invention:

[0120] The total loss function of the implicit neural model is designed as a weighted sum of RGB loss, mask loss, Eikonal loss, and curvature loss;

[0121] The expression of RGB loss is:

[0122] ;

[0123] in, is the RGB loss, represents a rendered image, represents the real image, Represents the difference in pixel value between the rendered image and the real image at the spatial sampling point i, is the 2-norm;

[0124] The expression of mask loss is:

[0125] ;

[0126] in, is the mask loss, is the predicted opacity value at spatial sampling point i, is the true binary mask at the spatial sampling point i;

[0127] The expression of Eikonal loss is:

[0128] ;

[0129] in, For Eikonal loss, At the spatial sampling point The gradient of the signed distance function at , represents the gradient operator;

[0130] The curvature loss is expressed as:

[0131] ;

[0132] in, is the curvature loss, At the spatial sampling point The mean curvature of the signed distance function at .

[0133] For curvature loss, embodiments of the present invention estimate the surface normal and curvature of the SDF by using numerical gradients, which estimate the gradient by evaluating the function at nearby points and calculating the rate of change:

[0134] ;

[0135] Among them, the formula 、 、 is a unit vector pointing in the positive direction of the x, y, and z axes; 、 、 is the step size along the x, y, and z axes.

[0136] During the optimization process, the step size is initialized according to the size of the coarsest hash grid and gradually reduced to match the different hash grid sizes during the optimization process, so that numerical gradients are applied to calculate the surface normal and curvature of the SDF; through these steps, the loss function design and end-to-end joint training can ensure that the model maintains geometric accuracy and surface smoothness while generating visually realistic 3D reconstructions.

[0137] In an experimental embodiment of the present invention:

[0138] The 3D reconstruction network model proposed in this experiment uses the PyTorch Lightning framework, which separates research and engineering code, simplifies data processing, and improves code readability. Multi-resolution hash coding is implemented using NVIDIA's tiny-cuda-nn library, and volume rendering is performed using the NerfAcc toolbox, optimizing model training and inference efficiency. Experiments were conducted on a Linux system.

[0139] The typical loss curve when training the implicit neural representation 3D reconstruction network model is as follows: Figure 6 As shown in the figure, the loss value of the model drops rapidly after only 200 training iterations. When the total number of training iterations is set to 2000, the model loss value and the peak signal-to-noise ratio of the validation image tend to be stable.

[0140] In the experimental embodiment of the present invention, the parameters of 16 cameras in the unobstructed multi-view imaging system were optimized. The initial camera parameters were obtained based on the design blueprint and further optimized during model training. By using the camera pose as a trainable variable, a complete training was completed, resulting in more accurate camera positions and poses. In subsequent training, these optimized parameters were directly used without the need for retraining.

[0141] The experimental embodiment of the present invention further reconstructed the standard parts in three dimensions through the constructed implicit neural representation method and analyzed its accuracy. The experimental embodiment used a sphere with a diameter of 15 mm and a cylinder of the same size as the standard parts. The comparison of the reconstruction results with the CAD model showed that the maximum surface error of the sphere was 0.401 mm and the average error was 0.097 mm; the maximum surface error of the cylinder was 0.281 mm and the average error was 0.156 mm. The error heat map is shown in Figure 2. Figure 7 shown.

[0142] The experimental example of the present invention further selected particles of five different shapes as experimental samples, including round, irregular, angular, flake, and elongated shapes. First, a high-precision commercial 3D scanner, namely the AutoScan Inspec system manufactured by SHINING 3D, with an accuracy of ≤10 microns, was used to scan and reconstruct the particles. The process of scanning one particle usually takes about 4 minutes. Figure 8 The results of three-dimensional reconstruction using different methods are shown. The first row is used as a reference, showing the results from a high-precision 3D scanner. Rows 2 to 4 show the results of reconstruction using three different methods. Among them, row 2 shows the particle surface reconstructed by the space carving algorithm; row 3 shows the particle surface reconstructed by NeRF; row 4 shows the particle surface reconstructed using the method proposed in this invention, including surface texture.

[0143] The results show that the surface reconstructed using the contour method exhibits noticeable grid-like artifacts due to its voxel-based storage mechanism, resulting in a less than smooth surface. Furthermore, concave regions on the particle surface are flattened, resulting in poor reconstruction of fine surface details. The surface reconstructed using NeRF displays richer details than the contour method, but its use of density for geometric representation results in a surface filled with pits of varying depths, causing significant errors. The proposed method achieves particle reconstruction results that exhibit both fine surface detail and smoothness, eliminating pits similar to the NeRF method, resulting in optimal reconstruction performance. Figure 9 The reconstruction results obtained by the experimental embodiment of the present invention are shown. The particles are colored according to the distance value from the aligned reference surface. The reconstruction results show that the maximum surface distance is less than 0.828 mm and the average distance is less than 75 μm. Figure 10 The quantitative results of particle reconstruction using the three methods are presented. The results show that the average distance between the particle reconstruction method proposed in this invention and the reference value is the smallest compared with the other two methods.

[0144] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A 3D reconstruction method for coarse aggregate based on an implicit neural model, characterized in that: include: Construct an unobstructed multi-view image acquisition system to collect 3D contour image data of coarse aggregate particles; Obtaining spatial sampling points of coarse aggregate particles in a three-dimensional contour image; Based on the implicit neural model, the geometric model and texture model are established respectively; The geometric model is established by: using a multi-resolution hash coding technique combined with a multi-layer perceptron to construct the geometric model; The texture model is established by: using two multi-layer perceptrons to construct the texture model; each multi-layer perceptron contains 64 hidden neurons; The spatial sampling points are input into the geometric model to obtain the signed distance function values ​​of the spatial sampling points, specifically: The spatial sampling points are input encoded using multi-resolution hash coding technology to obtain a coding vector; Concatenate the encoding vector with the coordinates of the spatial sampling points to obtain a 35-dimensional input vector; The 35-dimensional input vector is subjected to feature extraction by a multi-layer perceptron to obtain the signed distance function value of the spatial sampling point; The input vector of the texture model is obtained according to the signed distance function value combined with the spherical harmonic function encoding, specifically: Obtain its eigenvector and the normal vector of the spatial sampling point according to the signed distance function value output by the geometric model; The 35-dimensional input vector of the geometric model is encoded using a 4th-order spherical harmonic function to obtain a 16-dimensional vector; Fusing the normal vector, the eigenvector and the 16-dimensional vector to obtain an input vector of a texture model; Input the input vector into the texture model to obtain the RGB value of the spatial sampling point; The signed distance function value and RGB value of the spatial sampling point are synthesized using a volume rendering model to obtain the final RGB value of the corresponding pixel of the spatial sampling point in the three-dimensional contour image.

2. The method for 3D reconstruction of coarse aggregate based on an implicit neural model according to claim 1, characterized in that: The unobstructed multi-view image acquisition system comprises: a coarse aggregate falling control device, an optical fiber falling detection sensor and a spherical cavity structure multi-view image acquisition system body; The main body of the spherical cavity structure multi-view image acquisition system includes: multiple industrial cameras distributed around, COB lighting sources, uniform light balls, multi-channel 10G switches and a control computer.

3. The method for 3D reconstruction of coarse aggregate based on an implicit neural model according to claim 1, characterized in that: The specific method for collecting three-dimensional contour image data of coarse aggregate particles is as follows: Setting a plurality of different exposure times to collect three-dimensional contour image data of the coarse aggregate particles; The Laplace operator is used to evaluate the blurriness of the 3D contour image data corresponding to different exposure times and determine the optimal exposure time. Obtain three-dimensional contour image data of coarse aggregate particles collected at the optimal exposure time.

4. The method for 3D reconstruction of coarse aggregate based on an implicit neural model according to claim 1, characterized in that: Obtain the spatial sampling points of coarse aggregate particles in the 3D contour image, specifically: The bounding box size of the 3D reconstruction is established according to the maximum size of the coarse aggregate particles; Sending rays to the bounding box using an unobstructed multi-view image acquisition system; Ray marching technology is used to select coordinate points on the ray to obtain the spatial sampling points of coarse aggregate particles in the three-dimensional contour image.

5. The method for 3D reconstruction of coarse aggregate based on an implicit neural model according to claim 1, characterized in that: Before the volume rendering model is used to synthesize the signed distance function value and RGB value of the spatial sampling point, the signed distance function value of the spatial sampling point is converted into a volume density, which is expressed as: ; in, Represents spatial sampling points The signed distance function value at represents the signed distance function, is the offset parameter, is a learnable scale parameter that determines how sensitive the volume density is to changes in the signed distance function value.

6. The method for 3D reconstruction of coarse aggregate based on an implicit neural model according to claim 1, characterized in that: When establishing the geometric model and texture model based on the implicit neural model, it also includes: The total loss function of the implicit neural model is set to the weighted sum of RGB loss, mask loss, Eikonal loss and curvature loss; The expression of the RGB loss is: ; in, is the RGB loss, represents a rendered image, represents the real image, Represents the difference in pixel value between the rendered image and the real image at the spatial sampling point i, is the 2-norm; The expression of the mask loss is: ; in, is the mask loss, is the predicted opacity value at spatial sampling point i, is the true binary mask at the spatial sampling point i; The expression of the Eikonal loss is: ; in, For Eikonal losses, At the spatial sampling point The gradient of the signed distance function at , represents the gradient operator; The curvature loss is expressed as: ; in, is the curvature loss, At the spatial sampling point The mean curvature of the signed distance function at .

Citation Information

Patent Citations

  • Method for obtaining tetrahedral grid from object three-dimensional image

    CN101436303A

  • Reconstruction model geometry and texture optimization method based on adaptive mesh subdivision

    CN116863101A