Coarse aggregate three-dimensional reconstruction method based on implicit neural model
Through the implicit neural model and multi-view image acquisition system, the problems of incomplete data, high equipment cost and complex operation of coarse aggregate particles in the prior art are solved, and the three-dimensional reconstruction effect with high precision and low complexity are achieved.
Patent Information
- Application Number
- CN202510737898.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing three-dimensional scanning technology has problems such as incomplete data collection, high equipment cost, complex operation, low scanning efficiency, environmental radiation risks and complex data processing when reconstructing coarse aggregate particles, resulting in insufficient reconstruction accuracy and efficiency.
Using an implicit neural model method, a geometric and texture model is established by constructing an occlusion-free multi-view image acquisition system, combining multi-resolution hash coding and spherical harmonic function coding, and three-dimensional reconstruction is carried out using volume rendering technology to achieve high-precision reconstruction of coarse aggregate particles.
It improves reconstruction accuracy and efficiency, reduces calculation complexity, enhances model adaptability, improves visual effects and realism, and avoids occlusion problems and cumbersome operations in traditional methods.
Smart Images

Figure CN120259564A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and three-dimensional modeling technology, and particularly relates to a three-dimensional reconstruction method for coarse aggregates based on an implicit neural model. Background Art
[0002] With the continuous advancement of global infrastructure construction, highway transportation, as a key area of national infrastructure construction, plays a fundamental and leading role in supporting aspects such as national territorial space development, logistics transportation, social interaction, and international exchanges. Asphalt pavement has become the preferred form of highway construction in China due to its convenient construction and maintenance, comfortable driving, low noise, and good drainage performance. Aggregates, as an indispensable and important component in the pavement, play a role of skeleton and filling in the pavement and have a direct impact on the pavement performance. In the wave of digital transformation, the digital measurement and analysis of the geometric shape characteristics of coarse aggregate particles are of great significance for improving the quality and efficiency of pavement design, construction, and maintenance; digital technology can measure and analyze the characteristics of coarse aggregate particles such as shape, size, and surface texture, and these characteristics can be used to improve the pavement quality and service life in stages such as pavement design, construction, and maintenance. In addition, the research on the geometric shape characteristics and particle size distribution of coarse aggregate particles is of great significance in many fields such as materials science, civil engineering, and geology, including mechanical properties, concrete and asphalt mixture design, particle packing and filling efficiency, soil and rock engineering, classification and property evaluation of geological materials, environmental impact assessment, material processing and manufacturing, scientific research and education, and formulation of technical specifications and standards.
[0003] In the field of digital measurement and analysis of the geometric shape characteristics of coarse aggregate particles, the existing technologies mainly rely on a variety of three-dimensional scanning and imaging technologies. These technologies include structured light 3D scanning, binocular stereo imaging, and computed tomography (CT), etc.; structured light 3D scanning obtains the three-dimensional information of an object by projecting a specific light pattern and analyzing its deformation; binocular stereo imaging uses two or more cameras to capture images from different angles and reconstructs the three-dimensional shape through parallax calculation; CT scanning penetrates an object with X-rays and obtains data from multiple angles to reconstruct the internal structure of the object; these technologies have achieved certain results in obtaining the three-dimensional data of coarse aggregate particles, but each has its limitations: (1) Problems of integrity and accuracy of data acquisition: Due to light occlusion and perspective limitations, the existing technologies may not be able to completely capture all the geometric features of coarse aggregate particles, resulting in incomplete data and inaccurate shape reconstruction; (2) Equipment cost and operation complexity: High-precision three-dimensional scanning equipment such as CT scanning and some laser scanning equipment is costly and complex to operate, which limits its wide promotion in industrial applications; (3) Scanning efficiency and real-time issues: In the existing technology, when dealing with a large number of coarse aggregate particles, the scanning and data processing speeds may not meet the real-time or near-real-time requirements, affecting production efficiency; (4) Environmental and health safety issues: CT scanning involves radiation exposure, posing potential risks to operators and the environment; (5) Complexity of data processing and analysis: The three-dimensional data obtained from the existing technology requires complex post-processing and analysis to extract the geometric shape characteristics of the particles, which increases the difficulty of technology application; (6) Dependence on specific conditions: Some technologies such as binocular stereo imaging have certain requirements for surface texture. For coarse aggregate particles with unclear surface features, accurate three-dimensional information may not be obtained. Summary of the Invention
[0004] To solve the problems existing in the existing technology, the present invention provides a three-dimensional reconstruction method for coarse aggregates based on an implicit neural model. This method acquires the three-dimensional contour image data of coarse aggregate particles by constructing an unobstructed multi-view image acquisition system; obtains the spatial sampling points of the coarse aggregate particles in the three-dimensional contour image; respectively establishes a geometric model and a texture model based on the implicit neural model; inputs the spatial sampling points into the geometric model to obtain the signed distance function values of the spatial sampling points; combines the signed distance function values with spherical harmonic function encoding to obtain the input vector of the texture model and inputs it into the texture model to obtain the RGB values of the spatial sampling points; synthesizes according to the signed distance function values and RGB values of the spatial sampling points using a volume rendering model to obtain the RGB values of the corresponding pixels of the spatial sampling points in the three-dimensional contour image. The present invention can achieve high-precision reconstruction of the three-dimensional geometric shape of coarse aggregate particles.
[0005] The present invention adopts the following technical solution, a three-dimensional reconstruction method for coarse aggregates based on an implicit neural model, Construct an unobstructed multi-view image acquisition system to acquire the three-dimensional contour image data of coarse aggregate particles; Obtain the spatial sampling points of the coarse aggregate particles in the three-dimensional contour image; Respectively establish a geometric model and a texture model based on the implicit neural model; Input the spatial sampling points into the geometric model to obtain the signed distance function values of the spatial sampling points; Combine the signed distance function values with spherical harmonic function encoding to obtain the input vector of the texture model; Input the input vector into the texture model to obtain the RGB values of the spatial sampling points; Synthesize according to the signed distance function values and RGB values of the spatial sampling points using a volume rendering model to obtain the final RGB values of the corresponding pixels of the spatial sampling points in the three-dimensional contour image.
[0006] Furthermore, the unobstructed multi-view image acquisition system includes: a coarse aggregate blanking control device, an optical fiber blanking detection sensor, and a multi-view image acquisition system main body; The multi-view image acquisition system main body includes: multiple industrial cameras distributed in a ring, a COB illumination light source, a diffuser sphere, a multi-channel 10 Gigabit switch, and a control computer.
[0007] Furthermore, the method for collecting three-dimensional contour image data of coarse aggregate particles is specifically as follows: Set multiple different exposure times to collect three-dimensional contour image data of the coarse aggregate particles; Use the Laplace operator to evaluate the blurriness of the three-dimensional contour image data corresponding to different exposure times, and determine the optimal exposure time; Obtain the three-dimensional contour image data of the coarse aggregate particles collected at the optimal exposure time.
[0008] Furthermore, obtaining the spatial sampling points of the coarse aggregate particles in the three-dimensional contour image is specifically as follows: Establish the size of the bounding box for three-dimensional reconstruction according to the maximum size of the coarse aggregate particles; Emit rays from the unobstructed multi-view image acquisition system to the bounding box; Use the ray marching technique to select coordinate points on the rays to obtain the spatial sampling points of the coarse aggregate particles in the three-dimensional contour image.
[0009] Furthermore, the geometric model is established by: using the multi-resolution hashing encoding technique combined with a multi-layer perceptron to construct the geometric model; The texture network is established by: using two multi-layer perceptrons to construct the texture network; each multi-layer perceptron contains 64 hidden neurons.
[0010] Furthermore, obtaining the signed distance function value of the spatial sampling points is specifically as follows: The geometric model uses the multi-resolution hashing encoding technique to encode the spatial sampling points to obtain an encoded vector; Concatenate the encoded vector with the coordinates of the spatial sampling points to obtain a 35-dimensional input vector; The 35-dimensional input vector undergoes feature extraction through a multi-layer perceptron to obtain the signed distance function value of the spatial sampling points.
[0011] Furthermore, according to the signed distance function value combined with spherical harmonic function encoding to obtain the input vector of the texture model, specifically as follows: Obtain its feature vector and the normal vector of the spatial sampling points according to the signed distance function value output by the geometric model; Encode the 35-dimensional input vector of the geometric model using a fourth-order spherical harmonic function to obtain a 16-dimensional vector; Fuse the normal vector, the feature vector, and the 16-dimensional vector to obtain the input vector of the texture model.
[0012] Further, before performing synthesis using the volume rendering model based on the signed distance function value and the RGB value of the spatial sampling points, it further includes: converting the signed distance function value of the spatial sampling points into a volume density, expressed as: ; Wherein, represents the signed distance function value at the spatial sampling point , represents the signed distance function, is an offset parameter, is a learnable scale parameter that determines the sensitivity of the volume density to changes in the signed distance function value.
[0013] Further, when establishing the geometric model and the texture model based on the implicit neural model respectively, it further includes: Setting the total loss function of the implicit neural model as the weighted sum of the RGB loss, the mask loss, the Eikonal loss, and the curvature loss; The expression of the RGB loss is: ; Wherein, is the RGB loss, represents the rendered image, represents the real image, represents the difference in pixel values between the rendered image and the real image at the spatial sampling point i, is the 2-norm; The expression of the mask loss is: ; Wherein, is the mask loss, is the predicted opacity value at the spatial sampling point i, is the real binary mask at the spatial sampling point i; The expression of the Eikonal loss is: ; Wherein, is the Eikonal loss, is the gradient of the signed distance function at the spatial sampling point , represents the gradient operator; The curvature loss is expressed as: ; Among them, is the curvature loss, which is the mean curvature of the signed distance function at the spatial sampling point .
[0014] The beneficial effects of the present invention are as follows: By integrating the unobstructed multi-view image acquisition technology and the three-dimensional contour refinement reconstruction method based on implicit neural representation, the present invention realizes the high-precision reconstruction of the three-dimensional geometric shape of coarse aggregate particles, and has made remarkable improvements and enhancements in the following aspects: 1. Improved reconstruction accuracy and efficiency: The implicit neural representation method adopted by the present invention can accurately recover the three-dimensional contour of coarse aggregate particles from multi-view images, avoiding the occlusion problem and the cumbersome camera calibration process in traditional methods, and significantly improving the reconstruction accuracy and efficiency; 2. Reduced computational complexity: By adopting multi-resolution hash encoding and linear interpolation techniques, the present invention reduces the number of parameters of the implicit neural network model while maintaining the reconstruction quality, thereby reducing the computational complexity; 3. Enhanced model adaptability: The method of the present invention can encode sampling points at different scales, capture the local details and global structural features of the particle surface, and enhance the adaptability of the model to particles of different sizes; 4. Improved visual effect and realism: The neural rendering technology simulates the interaction between light and the object surface through a deep learning model, generates high-quality visual effects, and learns the reflection characteristics of the particle surface, enhancing the realism of the reconstruction result. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is a schematic flow chart of a three-dimensional reconstruction method of coarse aggregate based on an implicit neural model according to an embodiment of the present invention; Figure 2 is a schematic overview diagram of the main body of a multi-view image acquisition system with a spherical cavity structure according to an embodiment of the present invention; Figure 3 is a schematic diagram of the camera stratification in a multi-view image acquisition system with a spherical cavity structure according to an embodiment of the present invention; Figure 4 is a schematic diagram of the structure of a geometric model and a texture model according to an embodiment of the present invention; Figure 5 Schematic diagram of the principle of volume rendering according to an embodiment of the present invention; Figure 6 Schematic diagram of a typical loss curve during the training of a 3D reconstruction network model according to an embodiment of the present invention; Figure 7 3D reconstruction error heat map according to an embodiment of the present invention; Figure 8 Schematic diagram for comparing the results of 3D reconstruction by different methods according to an embodiment of the present invention; Figure 9 3D reconstruction grid error heat map of coarse aggregate particles according to an embodiment of the present invention; Figure 10 Schematic diagram for comparing the quantification results of 3D reconstruction accuracy by different methods according to an embodiment of the present invention. Detailed implementation manners
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0018] A schematic flow chart of a method for 3D reconstruction of coarse aggregates based on an implicit neural model according to an embodiment of the present invention is as Figure 1 shown and includes: Construct an unobstructed multi-view image acquisition system to collect 3D contour image data of coarse aggregate particles; The unobstructed multi-view image acquisition system constructed in the embodiment of the present invention includes a coarse aggregate blanking control device, an optical fiber blanking detection sensor, and a spherical cavity structure multi-view image acquisition system main body. The overview diagram of the system main body is as Figure 2 shown, which is composed of multiple industrial cameras distributed around, a COB illumination light source, a diffuser sphere, a multi-channel 10G switch, and a control computer. Its working process is as follows: the blanking control device releases coarse aggregate particles from the top of the device and makes them freely fall along the central axis of the device; when the particles pass through the top window of the imaging device, the optical fiber sensor generates a trigger signal and synchronously sends it to each industrial camera. After receiving the trigger signal, the camera starts the delayed trigger function and sets the delay duration according to the time required for the particles to freely fall. When the particles just fall to the center of the device, all industrial cameras are exposed simultaneously to capture images of the coarse aggregate particles from different perspectives.
[0019] In the spherical cavity structure multi-view imaging system of the embodiment of the present invention, the number of cameras included and their view angle distributions determine the representational completeness of the collected image samples for the complete contour information of the target. The specific camera layer diagram is asFigure 3 As shown in the figure. The multi-view image acquisition device designed in the embodiment of the present invention is equipped with 16 industrial cameras to balance the accuracy and cost of three-dimensional reconstruction; the cameras are distributed on the surface of the sphere at 4 "latitudes" with equal angular differences, 4 cameras are arranged at equal angular intervals on each layer, and the cameras between adjacent layers are staggered at an angle of 45°, and all 16 cameras maintain a front view of the center of the sphere.
[0020] The core function of the optical imaging system is that multiple cameras simultaneously capture the coarse aggregate particles falling to the center of the device, and the two-dimensional contours of the coarse aggregate particles in each captured view image are complete and the details are clear. The system further selects 16 JHEM131GM global exposure black-and-white industrial cameras with the same parameters and installs high-definition distortion-free lenses to meet the high-quality imaging of fast-moving target objects in a specific field of view; the light compensation design is a key factor for realizing high-quality and robust machine vision applications. The system proposed in the embodiment of the present invention combines two methods of backlight illumination and diffused illumination. By installing a high-brightness COB LED light source board on the inner wall of the carbon fiber board and embedding a white acrylic diffuser sphere in the device, the uniformity of the illumination is optimized to ensure high-quality multi-view images are obtained within a short exposure time.
[0021] Each industrial camera in the system is equipped with 1 GigE Gigabit Ethernet interface and is connected to an H3C S1226FX switch through a star topology to reduce data transmission latency and network bottlenecks. The switch is connected to the server through 2 10G fiber uplink data ports and adopts a QoS strategy to manage the priorities of different types of data streams to ensure the priority transmission of key image data and improve the overall transmission efficiency and stability of the system.
[0022] When obtaining the three-dimensional contour image of the coarse aggregate particles, analyze the free-fall motion process of the coarse aggregate particles to determine the motion parameters of the particles in the device. By calculating the total height, the required time, and the instantaneous velocity when the particles fall from the blanking control device to the center of the device, establish a particle motion model. Using this model, combined with the physical scale corresponding to a unit pixel on the imaging target surface of the camera, calculate the time required for the particles to move a unit pixel on the imaging target surface, providing a theoretical basis for the subsequent optimization of the exposure time; further, sequentially set the exposure times of all industrial cameras in the multi-view image capture system to multiple different values between 15 us and 200 us, perform multi-view imaging on the same standard sphere, calculate the edge response of each pixel point in the image using the Laplace operator, and determine the optimal exposure time by comparing the blurriness of the multi-view images collected at different exposure times. The calculation formula of the Laplace operator of the image is shown as follows: ; Where Indicates the pixel value of the image at The pixel value at the position, And Respectively represent the second-order partial derivatives of the image in the x and y directions. Is the Laplacian response of the image. For clear edges, the Laplacian response has a large absolute value because the image gray value changes violently at the edge. Therefore, the blurriness evaluation index B is defined as the variance of the Laplacian response of the image, and the calculation formula is as shown in the following formula: ; Among them, Is the total number of pixels in the image, Represents the coordinates of the i-th pixel point, Represents the Laplacian operator, Represents the image The Laplacian response value at the pixel point i, Is the image The mean value of the Laplacian response of the pixel points in. Compare the blurriness of multi-view images collected at different exposure times, and ensure that the target edges in the collected images are as clear as possible by reasonably setting the parameters of the multi-view image capture system.
[0023] Further, according to the blurriness evaluation results obtained in the above steps, draw a scatter plot of the exposure time and the image blurriness, and analyze the trend of the image blurriness changing with the exposure time. It is determined that when the exposure time is 150 us, the target edges in the collected images are the clearest. Considering the displacement of the particles during the exposure and the influence of the lens light input, comprehensively evaluate the performance of the system imaging components and the light compensation device, and determine the optimal exposure time to be 150 us to ensure that the target edges in the collected images are as clear as possible, while avoiding problems such as overexposure and underexposed images.
[0024] Obtain the spatial sampling points of the coarse aggregate particles in the three-dimensional contour image; In the embodiment of the present invention, after obtaining the three-dimensional contour image, the size of the bounding box of the three-dimensional reconstruction scene is set according to the maximum size of the coarse aggregate particle sample. After the initial setting is completed, a ray is emitted from the camera principal point in the multi-view imaging system of the spherical cavity structure through a specific pixel in the image and shoots into the scene bounding box. Through the ray tracing technology, a series of spatial sampling points are selected on the ray according to the established strategy.
[0025] Based on the implicit neural model, establish a geometric model and a texture model respectively; In the embodiment of the present invention, the structural schematic diagrams of the geometric model and the texture model are as Figure 4As shown, first, the geometric model is initialized. This model consists of a multi-layer perceptron (MLP) and is used to predict the signed distance function (SDF) values of sampling points in three-dimensional space. The input of the model is the coordinate values of the spatial sampling points, and input encoding is performed through multi-resolution hash encoding technology. The encoded feature vector is concatenated with the original coordinates to construct an input vector with a dimension of 35, so as to enhance the model's perception ability of the spatial structure. The structural block diagram is as shown in Figure 4 the left half. The SDF is implicitly shown as follows: ; where represents three-dimensional Euclidean space, is a point in three-dimensional space, represents the signed distance function, and the set contains all points for which the signed distance function value is 0, that is, the surface of the object, namely the points that satisfy the above formula constitute the surface of the object. The SDF value quantifies the spatial distance between the point and the object surface. Points inside the surface are negative, and points outside the surface are positive.
[0026] The texture model consists of a two-layer MLP, with each layer containing 64 hidden neurons, and is used to predict the RGB color values of sampling points. The input data of this part consists of three key parts: the normal vector (Normal) of the sampling points predicted by the geometric network, the 13-dimensional feature vector output by the geometric network, and the 16-dimensional vector generated by encoding the input 2D observation view using the 4th-order spherical harmonics; the fusion of these three parts forms a 32-dimensional input vector, providing a multi-dimensional feature space for the model, thereby improving the realism and dynamic range of the rendering results, enabling the texture model to accurately predict color values. The texture network model is as shown in Figure 4 the right half.
[0027] Input the spatial sampling points into the geometric model to obtain the signed distance function values of the spatial sampling points; Specifically, in the embodiments of the present invention, the multi-view image data collected is input into the implicit neural representation model for implicit encoding of the three-dimensional geometric structure. For each coarse aggregate particle captured in the image, a spatial sampling point in three-dimensional space is input into a multi-layer perceptron (MLP), and through the MLP learning, these coordinate points are mapped to the corresponding signed distance function (SDF) values. The expression is: ; where is the spatial sampling point The signed distance function value to the particle surface, which is negative if the point is inside the surface and positive if it is outside the surface.
[0028] Obtain the input vector of the texture model according to the signed distance function value combined with spherical harmonic function encoding; input the input vector into the texture model to obtain the RGB values of the spatial sampling points; In the embodiment of the present invention, the output of the geometric model and the viewing direction encoded by the spherical harmonic function are jointly input into the texture MLP to predict the RGB values of the current spatial sampling points. The expression is as follows: ; Where represents the texture MLP, is the predicted RGB value of the spatial sampling point, is the viewing direction information encoded by the spherical harmonic function, is the output of the geometric model, that is, the value of the spatial sampling point .
[0029] Integrate according to the signed distance function value and RGB value of the spatial sampling point using the volume rendering model to obtain the final RGB value of the pixel corresponding to the spatial sampling point in the three-dimensional contour image.
[0030] Before integrating through the volume rendering model, convert the SDF value output by the geometric model into a volume density value for use in volume rendering. The conversion process is shown as follows: ; Where is the sigmod function, represents the volume density of the spatial sampling point i, represents the spatial sampling point corresponding SDF value, represents the th spatial sampling point corresponding SDF value.
[0031] In the embodiment of the present invention, through this conversion method, the deviation in the traditional method is corrected, and thus more accurate surface reconstruction is achieved. By using SDF to represent the object surface and applying advanced rendering techniques, it performs excellently in reconstructing detailed three-dimensional surfaces from two-dimensional images.
[0032] In an embodiment of the present invention, further through volume rendering technology, the SDF and RGB values of all sampling points along a certain ray are combined to render the final RGB value of the corresponding pixel in the synthesized image; volume rendering is to use image rendering technology to synthesize new images of the three-dimensional geometric structure information of implicitly represented particulate matter from a specified camera perspective in a scene. Subsequently, the difference between the rendered image and the input multi-view image is quantitatively compared to construct a loss function. Finally, through optimization methods such as backpropagation, supervised training of the three-dimensional geometric structure of particulate matter is realized without relying on strong supervision information such as scene depth maps. When applying the volume rendering model to three-dimensional reconstruction or new view image synthesis application scenarios, usually starting from the view position, rays passing through the image pixel points are constructed. When the rays pass through the three-dimensional scene space predefined according to the scale of the target object or scene, several sampling points are selected on the rays according to specific rules, and the data associated with all sampling points are accumulated, that is, the RGB color of the pixel point in the synthesized image corresponding to the ray is obtained. In this way, volume rendering associates the pixel points in the input multi-view image with the pixel points in the synthesized image. In an embodiment of the present invention, a schematic diagram of the principle of volume rendering is as shown in Figure 5 shown. Specifically, the formula for volume rendering is as shown in the following formula: ; where, is the final RGB value of pixel , is the ray path from the camera to pixel , is the color value of the point on the ray path, is the density value of the point on the ray path, indicating the attenuation of light.
[0033] In an embodiment of the present invention, a loss function is further constructed to measure the difference between the synthesized image and the input multi-view image, and the backpropagation algorithm is executed on the entire model to update the trainable parameters in the model.
[0034] Thus, in an embodiment of the present invention, a high-precision three-dimensional geometric structure is learned and reconstructed from multi-view images through an implicit neural model, and at the same time, the texture information on the surface of coarse aggregate particles can be predicted.
[0035] In another specific embodiment of the present invention: When encoding the spatial sampling point output geometric model, a multi-resolution hash encoding technology is used to encode the sampling points in the space. Specifically: First, divide the continuous three-dimensional scene space into stereo grids with different resolutions from coarse to fine, and associate eigenvalue vectors of a fixed length with the vertices of each grid cube. When a sampling point falls into a certain grid cube, according to the distance relationship between the sampling point and the cube vertices, use the trilinear interpolation method to process the eigenvector associated with the cube grid vertices, so as to obtain the eigenvector of the sampling point at the corresponding resolution.
[0036] Concatenate the eigenvector after multi-resolution hash encoding with the original coordinates to construct an input vector with a dimension of 35 to enhance the model's perception ability of the spatial structure. The expression of the hash encoding is as follows: ; where, represents the bitwise XOR operation. The bitwise XOR operation performs an exclusive OR operation on the corresponding bits of two binary numbers. If they are the same, the result is 0; if they are different, the result is 1. represents three unique large prime numbers corresponding to the three dimensions of the integer coordinates of the input grid vertices respectively. represents the component of the integer three-dimensional coordinate value of the input grid vertex. The superscript d represents the dimension, and the subscript represents the index of the data point. is the modulo operation, which takes the remainder after dividing the result of the XOR operation by T. T is a positive integer, usually used to limit the hash value within a specific range, that is, between 0 and T−1. This formula is implemented by performing an XOR operation on the results of the linear congruence (pseudo-random) permutations of each dimension, so as to eliminate the influence of the dimension on the hash value. To achieve (pseudo) independence, only two of the three dimensions of the integer coordinates of the input grid vertices need to be permuted. Therefore, in the embodiments of the present invention, , and are selected to improve cache coherence.
[0037] Use an Occupancy Grid Estimator to select sampling points for the ray. The Occupancy Grid Estimator caches the scene density in a binary voxel grid, allowing for effective elimination of empty regions when the ray passes through a predefined grid at a set step size.
[0038] During the ray sampling process, adjust the distribution of sampling points according to the occupancy status of the voxels. The proportion of sampling points in the occupied grid is significantly higher than that in the unoccupied grid, so as to ensure that more sampling points are concentrated near the object surface and make an important contribution to the rendering result. In the embodiments of the present invention, the sampling points in the occupied grid may account for 70% - 80% of the total number of ray sampling points, while the sampling points in the unoccupied grid only account for 20% - 30%.
[0039] Through this method for optimizing spatial sampling points with multi-resolution hash encoding and importance strategy in the embodiments of the present invention, the efficiency of training and inference of the neural implicit representation 3D reconstruction model can be effectively improved, the probability of extracting weakly relevant sampling points can be reduced, thereby improving the rendering efficiency and accuracy.
[0040] In another specific embodiment of the present invention: The total loss function of the implicit neural model is designed as a weighted sum of RGB loss, mask loss, Eikonal loss, and curvature loss; The expression of the RGB loss is: ; Where, is the RGB loss, represents the rendered image, represents the ground truth image, represents the difference in pixel values between the rendered image and the ground truth image at the spatial sampling point i, is the 2-norm; The expression of the mask loss is: ; Where, is the mask loss, is the predicted opacity value at the spatial sampling point i, is the ground truth binary mask at the spatial sampling point i; The expression of the Eikonal loss is: ; Where, is the Eikonal loss, is the gradient of the signed distance function at the spatial sampling point , represents the gradient operator; The curvature loss is expressed as: ; Where, is the curvature loss, is the mean curvature of the signed distance function at the spatial sampling point .
[0041] For the curvature loss, in the embodiments of the present invention, the surface normal and curvature of the SDF are estimated by using numerical gradients, and the numerical gradients estimate the gradients by evaluating the function at nearby points and calculating the rate of change: ; Where, in the formula , , is a unit vector pointing in the positive directions of the x, y, and z axes; , , are the step sizes along the x, y, and z axes.
[0042] During the optimization process, the step sizes are initialized according to the size of the coarsest hash grid and gradually decreased during the optimization process to match different hash grid sizes, so as to apply numerical gradients to calculate the surface normal and curvature of the SDF; through these steps, the loss function design and end-to-end joint training can ensure that the model maintains the geometric accuracy and surface smoothness while generating visually realistic 3D reconstructions.
[0043] In an experimental embodiment of the present invention: The 3D reconstruction network model proposed in this experiment adopts the Pytorch Lightning framework, which separates research and engineering code, simplifies the data processing flow, and improves the readability of the code. The multi-resolution hash encoding is implemented using NVIDIA's tiny-cuda-nn library, and volume rendering is performed through the NerfAcc toolbox, optimizing the model training and inference efficiency. The experiment is carried out on a Linux system.
[0044] When training the implicit neural representation 3D reconstruction network model by inputting 16 multi-view images of a single coarse aggregate particle, i.e., camera pose information, a typical loss curve is as Figure 6 shown. The loss value of the model drops rapidly after only 200 iterations of training. When the total number of training iterations is set to 2000, the model loss value and the peak signal-to-noise ratio of the validation images tend to stabilize.
[0045] In the experimental embodiment of the present invention, the parameters of 16 cameras in the unoccluded multi-view imaging system are optimized. The initial camera parameters are obtained based on the design blueprint and further optimized during model training. By taking the camera pose as a trainable variable, a complete training is completed, and more accurate camera positions and poses are obtained. In subsequent training, these optimized parameters are directly used without retraining.
[0046] In the experimental embodiment of the present invention, 3D reconstructions of standard parts are further performed through the constructed implicit neural representation method, and their accuracies are analyzed. The experimental embodiment uses a sphere with a diameter of 15 mm and a cylinder of the same size as standard parts. The comparison between the reconstruction results and the CAD model shows that the maximum surface error of the sphere is 0.401 mm and the average error is 0.097 mm; the maximum surface error of the cylinder is 0.281 mm and the average error is 0.156 mm. The error heat map is as Figure 7 shown.
[0047] In the experimental embodiment of the present invention, five different-shaped particles were further selected as experimental samples, including circular, irregular, angular, flaky, and elongated. First, a high-precision commercial 3D scanner, namely the AutoScan Inspec system manufactured by SHINING 3D with a precision of ≤10 microns, was used to scan and reconstruct the particles. The process of scanning one particle usually takes about 4 minutes. Figure 8 The results of 3D reconstruction using different methods are shown. The first row serves as a reference and shows the results from the high-precision 3D scanner. The second to fourth rows show the results reconstructed using three different methods. Among them, the second row shows the particle surface reconstructed by the spatial carving algorithm; the third row shows the particle surface reconstructed by NeRF; the fourth row shows the particle surface reconstructed using the method proposed in the present invention, including surface texture.
[0048] It can be observed from the results that the surface reconstructed using the contour method shows obvious grid-like artifacts due to the voxel-based storage mechanism, resulting in a non-smooth surface. In addition, the concave regions on the particle surface are flattened, leading to poor reconstruction of fine surface details. The surface reconstructed using NeRF shows richer details compared to the contour method, but the mechanism of using density for geometric representation causes the surface to be filled with pits of different depths, resulting in significant errors. The method proposed in the present invention achieves the particle reconstruction result, demonstrating the fine surface details of the particles and showing smoothness, eliminating the pits similar to the NeRF method, thus achieving the best reconstruction performance. Figure 9 The reconstruction results obtained from the experimental embodiment of the present invention are shown, colored according to the distance value from the aligned reference surface. The reconstruction result of the particle shows that the maximum surface distance is less than 0.828 mm and the average distance is less than 75 microns. At the same time, Figure 10 The quantitative results of particle reconstruction by the three methods are shown. The results indicate that, compared with the other two methods, the method proposed in the present invention has the smallest average distance from the reference value in particle reconstruction.
[0049] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A three-dimensional reconstruction method for coarse aggregates based on an implicit neural model, characterized in that, Including: Construct an unobstructed multi-view image acquisition system to collect three-dimensional contour image data of coarse aggregate particles; Obtain the spatial sampling points of the coarse aggregate particles in the three-dimensional contour image; Based on the implicit neural model, establish a geometric model and a texture model respectively; Input the spatial sampling points into the geometric model to obtain the signed distance function values of the spatial sampling points; According to the signed distance function values, combine with spherical harmonic function encoding to obtain the input vector of the texture model; Input the input vector into the texture model to obtain the RGB values of the spatial sampling points; According to the signed distance function values and RGB values of the spatial sampling points, use the volume rendering model for synthesis to obtain the final RGB values of the corresponding pixels of the spatial sampling points in the three-dimensional contour image.
2. The three-dimensional reconstruction method of coarse aggregate based on an implicit neural model according to claim 1, wherein: The unobstructed multi-view image acquisition system includes: a coarse aggregate blanking control device, an optical fiber blanking detection sensor, and a spherical cavity structure multi-view image acquisition system main body; The spherical cavity structure multi-view image acquisition system main body includes: multiple industrial cameras distributed around, a COB lighting source, a diffuser sphere, a multi-channel 10 Gigabit switch, and a control computer.
3. The three-dimensional reconstruction method of coarse aggregate based on an implicit neural model according to claim 1, characterized in that: The method for collecting three-dimensional contour image data of coarse aggregate particles is specifically as follows: Set multiple different exposure times to collect three-dimensional contour image data of the coarse aggregate particles; Use the Laplace operator to evaluate the blurriness of the three-dimensional contour image data corresponding to different exposure times, and determine the optimal exposure time; Obtain the three-dimensional contour image data of the coarse aggregate particles collected at the optimal exposure time.
4. A three-dimensional reconstruction method of coarse aggregate based on an implicit neural model according to claim 1, characterized in that: Obtaining the spatial sampling points of the coarse aggregate particles in the three-dimensional contour image is specifically as follows: Establish the size of the bounding box for three-dimensional reconstruction according to the maximum size of the coarse aggregate particles; Emit rays from the unobstructed multi-view image acquisition system to the bounding box; Use the ray marching technique to select coordinate points on the rays to obtain the spatial sampling points of the coarse aggregate particles in the three-dimensional contour image.
5. A three-dimensional reconstruction method of coarse aggregate based on an implicit neural model according to claim 1, characterized in that: The establishment method of the geometric model is: using the multi-resolution hash encoding technique combined with a multi-layer perceptron to construct the geometric model; the establishment method of the texture network is: using two multi-layer perceptrons to construct the texture network; each multi-layer perceptron contains 64 hidden neurons.
6. The three-dimensional reconstruction method of coarse aggregate based on an implicit neural model according to claim 5, characterized in that: Obtaining the signed distance function values of the spatial sampling points is specifically as follows: Use the multi-resolution hash encoding technique to encode the spatial sampling points to obtain an encoded vector; Concatenate the encoded vector with the coordinates of the spatial sampling points to obtain a 35-dimensional input vector; The 35-dimensional input vector undergoes feature extraction through a multi-layer perceptron to obtain the signed distance function values of the spatial sampling points.
7. A three-dimensional reconstruction method of coarse aggregate based on an implicit neural model according to claim 6, characterized in that: According to the signed distance function values, combining with spherical harmonic function encoding to obtain the input vector of the texture model is specifically as follows: Obtain its feature vector and the normal vector of the spatial sampling points according to the signed distance function values output by the geometric model; Use the 4th-order spherical harmonic function to encode the 35-dimensional input vector of the geometric model to obtain a 16-dimensional vector; Fuse the normal vector, feature vector, and the 16-dimensional vector to obtain the input vector of the texture model.
8. A three-dimensional reconstruction method of coarse aggregate based on an implicit neural model according to claim 1, characterized in that: Before synthesizing according to the signed distance function value and RGB value of the spatial sampling points using a volume rendering model, it further includes: converting the signed distance function value of the spatial sampling points into a volume density, expressed as: ; Among them, represents the value of the signed distance function at the spatial sampling point, represents the signed distance function, is the offset parameter, is a learnable scale parameter, which determines the sensitivity of the volume density to the change of the signed distance function value.
9. A three-dimensional reconstruction method of coarse aggregate based on an implicit neural model according to claim 1, characterized in that: When respectively establishing a geometric model and a texture model based on an implicit neural model, it further includes: Setting the total loss function of the implicit neural model as a weighted sum of RGB loss, mask loss, Eikonal loss, and curvature loss; The expression of the RGB loss is: ; Among them, is the RGB loss, represents the rendered image, represents the ground truth image, represents the difference in pixel values between the rendered image and the ground truth image at the spatial sampling point i, is the L2 norm; The expression of the mask loss is: ; Among them, is the mask loss, is the predicted opacity value at the spatial sampling point i, is the true binary mask at the spatial sampling point i; The expression of the Eikonal loss is: ; Among them, is the Eikonal loss, which is the gradient of the signed distance function at the spatial sampling points and represents the gradient operator; The curvature loss is expressed as: ; Among them, is the curvature loss, which is the mean curvature of the signed distance function at the spatial sampling point.
Citation Information
Patent Citations
Method for obtaining tetrahedral grid from object three-dimensional image
CN101436303A
Multi-view three-dimensional reconstruction method based on implicit neural representation
CN115761178A
Reconstruction model geometry and texture optimization method based on adaptive mesh subdivision
CN116863101A
Nerve implicit SLAM method based on voxel tetrahedron coding
CN118470241A
Indoor three-dimensional scene reconstruction method and system based on implicit coding and geometric prior
CN119251402A