Trusted twin modeling method and storage medium of neural radiation field based on evidence fusion
Through the evidence fusion of neural radiation field models, the uncertainty of spatial points is quantified and the training process is optimized, which solves the problems of large computational overhead and model uncertainty in three-dimensional reconstruction, and realizes efficient and reliable complex surface defect detection and high-precision modeling.
Patent Information
- Application Number
- CN202510148862.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing three-dimensional reconstruction technology has high computational overhead and long training time in complex scenarios, high dependence on input image quality, and poor reconstruction effect when facing high noise or sparse data, and the model uncertainty is not fully considered, resulting in inaccurate reconstruction results.
Using a trusted twin modeling method of neural radiation field based on evidence fusion, the uncertainty of spatial points is quantified by introducing evidence deep learning modules and normal inverse gamma distribution, combined with adaptive resampling strategies and loss weighting, the training process is optimized to improve the robustness and accuracy of the model in complex environments.
It improves the accuracy and robustness of 3D reconstruction, reduces computational overhead, and can provide reliable predictions in complex surface defect detection and high-precision modeling tasks, solving the problem of missing sample size and difficulty in quantifying uncertainty.
Smart Images

Figure CN119648923B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method, device and storage medium for credible twin modeling of neural radiation fields based on evidence fusion. Background Art
[0002] With the rapid development of computer vision and deep learning technologies, 3D reconstruction has been widely applied in various fields. Traditional 3D reconstruction methods mostly rely on multi-view stereo (MVS) or hardware devices such as structured light and laser scanning to perform 3D geometric modeling of objects. However, these traditional methods often fail to efficiently restore high-quality 3D reconstructions when faced with complex surface geometries and objects with rich details. This is particularly true when textures are unclear or lighting conditions vary significantly, resulting in poor reconstruction results.
[0003] In recent years, the emergence of Neural Radiance Field (NeRF) technology has significantly advanced the field of 3D reconstruction. NeRF utilizes a multi-layer perceptron (MLP) to infer 3D scene details and lighting information from multi-view 2D images. By modeling the scene's radiance field, NeRF is able to generate highly realistic 3D images and demonstrates significant advantages when handling complex surfaces and details. However, NeRF models still face challenges in complex scenes, such as high computational overhead, long training time, and dependence on input image quality. Furthermore, reconstruction performance can be unsatisfactory when dealing with highly noisy or sparse data.
[0004] In deep learning applications, especially in the field of 3D reconstruction, model uncertainty is often not fully considered. When outputting predictions, the model may be overconfident in some unknown or difficult-to-model scenarios, resulting in inaccurate reconstructions. To address this issue, Deep Evidential Regression (DER) has been proposed as a novel regression method designed to improve the model's robustness in complex environments by modeling the uncertainty of predictions. DER combines the advantages of deep learning and Bayesian reasoning, providing not only accurate predictions but also a reliability metric, effectively quantifying the model's uncertainty when exposed to diverse data (Regolini, R., et al., "Deep Evidential Regression," NeurIPS, 2019).
[0005] In NeRF applications, deep evidence-based regression (DER) methods provide an effective approach to addressing uncertainty in 3D reconstruction. By incorporating an evidence-based regression model into NeRF, a reliability metric can be added to the output 3D scene, reducing overconfident predictions and improving reconstruction accuracy, particularly when input data is incomplete or noisy. Recent research has demonstrated that combining NeRF with DER can significantly improve 3D reconstruction quality, particularly in surface defect detection and complex object modeling (Zhang, J., et al., "Multivariate Deep Evidential Regression," ICML, 2020).
[0006] In addition to deep evidence regression, traditional uncertainty modeling methods such as MC-Dropout (Gal, Y., "Uncertainty in Deep Learning," Ph.D. dissertation, University of Cambridge, 2016) and deep ensembles (Lakshminarayanan, B., et al., "Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles," Advances in Neural Information Processing Systems, 2017) have also improved the robustness of 3D reconstruction to some extent. However, these methods are often accompanied by high computational costs and complex implementations, particularly requiring multiple sampling or training of multiple models. In contrast, 3D reconstruction techniques combined with deep evidence regression can not only effectively reduce these computational costs but also provide more accurate and reliable predictions.
[0007] With the combination of NeRF technology and deep evidence regression methods, these technologies have broader application prospects in industrial inspection, virtual reality, medical imaging and other fields, especially in the tasks of three-dimensional reconstruction, defect detection and high-precision modeling of complex surfaces, which have important practical significance. Summary of the Invention
[0008] The present invention proposes a neural radiation field trusted twin modeling method, device and storage medium based on evidence fusion to solve the problems of missing sample size and difficulty in quantifying the uncertainty of supplementary samples in surface defect detection of industrial components based on computer vision.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] A neural radiation field trusted twin modeling method based on evidence fusion, which performs the following steps through computer equipment:
[0011] Step 1: Obtain multi-view image data for 3D scene reconstruction, including multi-view images and their corresponding camera internal and external parameters.
[0012] Step 2: Construct a neural radiation field model based on evidence-based deep learning to model the radiation field in three-dimensional space, and model the attributes of spatial points as volume density and color and their uncertainty parameters.
[0013] Step 2.1, neural radiation field input construction. As a 3D scene modeling method based on deep learning, neural radiation field has been widely used in recent years. By learning the radiation field representation of the 3D scene from image data and camera poses of different perspectives, the neural radiation field can efficiently perform 3D reconstruction and more accurately capture the lighting and geometric features in complex scenes. When processing complex and high-detail scenes, the neural radiation field can provide higher quality image reconstruction. and the corresponding camera pose are input into the neural radiance field model, where the image data is RGB three-channel, and the width and height are and , the input image data will be subdivided into coarse sampling and fine sampling in the input construction stage, which will be used as coarse network and fine network input respectively. The coarse network input is ,in is the number of coarse samples, and the fine network input is ,in is the number of fine samples;
[0014] Step 2.2: Evidence Deep Learning Module Construction. We embed the evidence deep learning module into the neural radiation field model to effectively model the prediction uncertainty of each spatial point and improve the stability and robustness of the model in areas with high uncertainty. We introduce the normal inverse gamma distribution (NIG) and assume that the color of the sampling point (correspond , are sampling points on the light) each obeys a mutually independent normal distribution, that is, , its mean and variance Controlled by the normal inverse gamma distribution parameters: mean , precision parameter , shape parameter , scale parameter ,These parameters are used to quantify the uncertainty of color prediction, and thus improve the confidence and stability of the model for predictions in different regions;
[0015] Step 2.3: Construct the Neural Radiance Field Network. The Neural Radiance Field Network consists of two identical components: a coarse network and a fine network. Each network is composed of two multi-layer perceptrons (MLPs). The first MLP extracts geometric information and uncertainty parameters from spatial coordinates to generate volume density, internal features, accuracy, shape, and scale parameters. The second MLP uses spatial coordinates and viewing direction, combined with internal features, to predict the color mean, capturing illumination characteristics and view-dependent color variations.
[0016] Step 3: A loss weighting based on data uncertainty is proposed. The parameters of the normal inverse gamma distribution are used to calculate the data uncertainty of each sampling point and weight the loss of each sample in the coarse network.
[0017] Step 3.1, in order to model uncertainty, we use the conjugate prior normal inverse gamma distribution in Bayesian statistics to simultaneously model the inherent uncertainty of the data and the uncertainty of the model parameters. In our model, data uncertainty represents the error in the color prediction of each spatial point due to data noise. For the first pixels and color channels , the data uncertainty and model uncertainty can be calculated as follows As shown:
[0018]
[0019]
[0020] in represents the scale parameter after volume rendering fusion, Represents the shape parameters after volume rendering fusion, Represents the accuracy parameter after volume rendering fusion. The greater the data uncertainty, the greater the prediction variance of the sampling point, and the model's prediction result for this point is more vague or unreliable;
[0021] Step 3.2, in our neural radiation field model, in order to reduce the negative impact of areas with large data uncertainty on the coarse network training, a loss weighting strategy is adopted. This strategy adaptively reduces the impact of the model on low uncertainty areas based on the data uncertainty of each pixel color, thereby improving the model's learning effect on high uncertainty areas. The weight of each spatial point is calculated based on the uncertainty of its color prediction and is normalized to avoid over-weighting. This ensures that after the coarse network loss is weighted, the model can more effectively learn areas with small data uncertainty and optimize the training process. The weighted loss of the coarse network is as follows: :
[0022]
[0023] in , , This is the first pixels, color channels The uncertainty weight of The data uncertainty of the color prediction for this sampling point.
[0024] In step 4, an adaptive resampling method is proposed. During the training process, the sampling density is increased for areas with higher uncertainty based on the model uncertainty estimation of spatial points, and the sampling density is appropriately reduced for areas with lower model uncertainty. This is used as the input of the fine network to improve training efficiency and rendering quality.
[0025] Step 4.1, based on the idea of layered sampling of neural radiation fields, perform preliminary sampling through the coarse network, and then use the fine network for fine sampling. First, use the calculation of each sampling point in the coarse network The volume density and weight of the coarse sampling points are , and evenly distributed in the range of light Each sampling point See the formula for the position of :
[0026]
[0027] Then use a thin network for fine sampling in key areas;
[0028] Input the sampling points into the coarse network to obtain the volume density and calculate the corresponding weight , which is defined as the formula :
[0029] in Indicates the The weight of the sampling points, It is The volume density of the sampling points, It is the total number of sampling points on the ray in coarse sampling. A higher volume density means that the area contributes more to the rendering of the model.
[0030] In the inference phase, the fine network is based on the weights of the coarse sampling points. Determine the position of the additional sampling point in each ray and perform further sampling based on this. The position of the fine sampling point is determined by the formula Sure:
[0031]
[0032] in Indicates the distance between two adjacent sampling points in coarse sampling;
[0033] Step 4.2: Generate preliminary volume density distribution and normal inverse gamma parameters based on the coarse network, combined with the model uncertainty described in step 3 , calculate the three-channel average model uncertainty , and introduced into the sampling weight calculation, such as formula :
[0034]
[0035] in is the adjustment coefficient, which is used to control the impact of uncertainty on the weight. The sequence number of the maximum sampling point of the model uncertainty.
[0036] Based on these weights, more sampling points are dynamically allocated to high uncertainty areas. During the training phase, in the fine network, the positions of the newly added sampling points are as follows: :
[0037]
[0038] During the training phase, an adaptive resampling strategy focuses resources on learning challenging areas, improving the ability to capture detail. During the inference phase, stratified sampling is used to balance efficiency and quality, focusing on refining key areas. This approach not only improves training efficiency and rendering accuracy, but also enhances the ability to reliably model complex surface defects.
[0039] Step 5: Fuse the normal inverse gamma distribution of each sampling point for volume rendering to generate a rendered image and uncertainty quantification map.
[0040] In step 5.1, for multiple sampling points along the ray, the color channel of each sampling point corresponds to the parameters of a normal inverse gamma distribution, including mean, precision, shape, and scale parameters, as well as weights related to uncertainty. To obtain the final pixel color and uncertainty quantification map, a weighted fusion method is used to calculate the expected value and uncertainty parameters of the pixel color. This method weights and sums the parameters of each sampling point, ultimately obtaining a comprehensive pixel color and uncertainty parameter, effectively improving rendering accuracy and uncertainty quantification capabilities.
[0041] These parameters are calculated by weighted fusion and numbered as The expected pixel color value and uncertainty parameters , as shown in the formula :
[0042]
[0043]
[0044]
[0045]
[0046] in , Indicates the number of sampling points on the ray, Indicates the The volume density of the sampling points, Represents the distance between two adjacent sampling points;
[0047] In step 5.2, combine the data uncertainty and model uncertainty definitions in step 3.1 to generate an uncertainty quantification diagram.
[0048] Step 6: Combine the differentiable rendering norm loss and the uncertainty loss function to optimize the global radiation field.
[0049] In step 6.1, a new regression loss function is introduced to use the input samples to learn the normal inverse gamma distribution parameter optimization model; for a batch size of For each ray in, calculate each color channel Color Normal Inverse Gamma Negative Log-Likelihood Loss , see formula :
[0050]
[0051] in, , Indicates the true color of light, Respectively Root ray color channel The mean, precision, shape and scale parameters of is the standard loss coefficient;
[0052] This loss function minimizes the error between the predicted color mean and the true color by maximizing the log-likelihood of the normal inverse gamma distribution of the color channel, taking into account the confidence of each prediction, and is applied to both coarse and fine network optimization. Combined with the loss weighting algorithm described in step 3.2, the coarse network loss can be rewritten as formula :
[0053]
[0054] In step 6.2, in order to prevent the fine network model from over-relying on uncertainty parameters, a norm loss is designed to ensure that the color estimate takes uncertainty into account. Compared with the estimate without considering the uncertainty The difference between them is as small as possible, see formula :
[0055]
[0056] in , is the weight coefficient of the normative loss;
[0057] Step 6.3, in summary, we propose a complex shape trusted twin modeling method for evidence fusion neural radiation field. The overall loss is shown in formula (14):
[0058]
[0059] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the above method.
[0060] On the other hand, the present invention further discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0061] It can be seen from the above technical solution that the present invention relates to the field of evidence deep learning and three-dimensional reconstruction technology, and discloses a credible twin modeling method for surface defects of complex-shaped products based on evidence fusion neural radiation field. The method includes: acquiring multi-view image data for three-dimensional scene reconstruction; constructing a neural radiation field model based on evidence deep learning, and modeling the volume density, color and uncertainty of spatial points through normal inverse gamma distribution; training the model using a weighted loss function of data uncertainty; increasing the sampling density of high-uncertainty areas through an adaptive resampling strategy during the training process to improve rendering quality; applying the trained model to a defect sample test data set; generating a three-dimensional reconstructed image and an uncertainty quantification map through the model, performing a credibility assessment on the defect area, and generating the final detection result by combining the rendered image and the uncertainty quantification map. The present invention effectively solves the problem of poor reconstruction quality caused by noise and data sparsity in traditional three-dimensional reconstruction technology by introducing a fusion method of evidence deep learning and neural radiation field, and provides a new solution to the common problem of missing sample size in surface defect detection of industrial products, namely, supplementing samples through high-precision and high-credibility evidence fusion neural radiation field twin reconstruction to improve the reliability and robustness of detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is the overall flow chart of the present invention;
[0063] Figure 2This is the structural diagram of the neural radiation field model constructed based on evidence-based deep learning proposed by the present invention;
[0064] Figure 3 It is the adaptive resampling module proposed by the present invention;
[0065] Figure 4 It is a credible modeling and visualization effect of complex surface defects. DETAILED DESCRIPTION
[0066] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.
[0067] like Figure 1 As shown,
[0068] For the purpose of formula unification, the following definitions are given for the subscripts and subscripts used in the formulas of this invention. Represents RGB color channel, the value is ; Represents the first sampling points, the values range from 1 to A natural number, the value in fine sampling ranges from 1 to natural numbers; Represents a training batch The A ray of light.
[0069] like Figure 1 As shown, the specific implementation steps of the credible twin modeling method for complex-shaped product surface defects based on evidence fusion neural radiation field of the present invention are as follows:
[0070] Step 1: Obtain training data and perform coarse sampling of light. First, obtain RGB image data of the target object through a multi-view camera to ensure that there is no blind spot in the view angle coverage, and calibrate the internal and external parameters of the camera (such as focal length, position, and posture) to provide high-quality data for model input. The camera view angle collects scene images in a diverse layout with a resolution of The pixel color is taken as the true value, and then the light corresponding to each pixel is placed in the direction for coarse sampling. The light starts from the center of the camera and the direction is (p is the pixel coordinate, o is the camera position), in the depth range Internal Generation uniform sampling points, Represent the nearest and farthest distances respectively, then the depth value of each sampling point can be obtained by the formula gives:
[0071]
[0072] The sampling point coordinates are calculated as These evenly distributed coarse sampling points provide the necessary basis for subsequent radiation field modeling, and the sampling distribution is further optimized in combination with uncertainty.
[0073] Step 2, such as Figure 2 As shown, an evidence fusion neural radiation field model is constructed.
[0074] Specifically, it is divided into the following steps:
[0075] Step 2.1, model input construction. The model input construction phase is based on pixels and light, combining multi-view image data with camera viewing direction to generate sampling points in three-dimensional space. The coarse network input is ,in is the number of coarse samples; the fine network input is ,in is the number of fine samples;
[0076] Step 2.2, embedding the evidence deep learning module. The evidence deep learning module is embedded in the neural radiation field model to model the uncertainty of each spatial point, assuming that the color of the sampling point Each obeys a mutually independent normal distribution, that is, , its mean and variance Controlled by the normal inverse gamma distribution parameters: mean , precision parameter , shape parameter , scale parameter , then its probability density function can be defined as ;
[0077] The prediction of the color of the sampling point follows the normal inverse gamma distribution , these parameters are used to quantify the uncertainty of the model, and the output parameters are as follows:
[0078] Color mean : represents the mean of color prediction;
[0079] Accuracy parameters : used to reflect the model's confidence in color prediction;
[0080] Shape parameters : Indicates the prediction stability of the model for this spatial point;
[0081] scale parameter : Used to control the range of color distribution and reflect the discrete degree of predicted distribution;
[0082] Step 2.3: Build the neural radiation field network structure. The neural radiation field model consists of a coarse network and a fine network, which are used for global feature extraction and local detail optimization, respectively. The structures of the two networks are exactly the same, both consisting of two multi-layer perceptrons (MLPs). The network input and output are constructed as follows:
[0083] The first MLP is mainly responsible for extracting geometric information and its uncertainty parameters from spatial coordinates to generate volume density , internal implicit features , precision parameters , shape parameters , scale parameters , whose parameters are :
[0084]
[0085] Color mean Through the second multi-layer perceptron. This MLP uses spatial coordinates and viewing direction , combined with internal implicit features , predict the color mean , which captures the scene's lighting characteristics and view-dependent color variations. Its parameters are :
[0086]
[0087] in , Represent the NIG estimated means of the red, green, and blue channels, respectively.
[0088] Step 3, use the proposed adaptive resampling method to obtain fine sampling points, such as Figure 3 As shown, train the fine network.
[0089] Step 3.1, stratified sampling. After the sampling points obtained by coarse sampling are input into the coarse network, the key areas can be sampled preliminarily, and then fine sampling can be performed using the fine network.
[0090] Input the sampling points into the coarse network to obtain the volume density and calculate the corresponding weight , which is defined as the formula :
[0091]
[0092] in Indicates the The weight of the sampling points, It is The volume density of the sampling points, It is the total number of sampling points on the ray in coarse sampling. A higher volume density means that the area contributes more to the rendering of the model.
[0093] Step 3.2, adaptive resampling. Adaptively adjust the sampling density to more accurately model high uncertainty areas, and dynamically allocate sampling points based on model uncertainty. The basis for adaptation is the initial volume density distribution of sampling points generated by the coarse network and the model uncertainty, where the model uncertainty can be obtained by the normal inverse gamma parameter: , and use it to calculate the sampling weight, as shown in the formula :
[0094]
[0095] in is the adjustment coefficient, which is used to control the impact of uncertainty on the weight.
[0096] Based on these weights, more sampling points are dynamically allocated to high uncertainty areas. In the thin network, the locations of the newly added sampling points are as follows: :
[0097]
[0098] in Represents the distance between two adjacent sampling points in the coarse sampling. During the training phase, the resampling strategy concentrates resources on learning difficult areas to improve the ability to capture details.
[0099] In step 4, the new view and the uncertainty view are synthesized using a volume rendering method fused with normal inverse gamma distribution.
[0100] Step 4.1, for multiple sampling points along the ray , each sampling point Color channels All correspond to the parameters of the normal inverse gamma (NIG) distribution: mean , precision ,shape ,scale , and the weights associated with the uncertainty ;
[0101] In order to obtain the final pixel color and uncertainty quantization map, we calculate these parameters by weighted fusion and obtain the number The expected pixel color value and uncertainty parameters , as shown in the formula :
[0102]
[0103]
[0104]
[0105]
[0106] in , Indicates the number of sampling points on the ray, Indicates the The volume density of the sampling points, Indicates the distance between two adjacent sampling points.
[0107] Step 4.2, uncertainty quantification map generation, combined with the data uncertainty and model uncertainty definitions in step 3.1, the pixel data uncertainty can be derived , model uncertainty ,These uncertainty quantification results help to evaluate the prediction confidence of the model in different regions.
[0108] Step 5: Train the model using the proposed loss.
[0109] Step 5.1 introduces a new regression loss function, which uses the input samples to learn the normal inverse gamma distribution parameters for a batch size of For each ray in, calculate each color channel Color Normal Inverse Gamma Negative Log-Likelihood Loss , see formula :
[0110]
[0111] Step 5.2, calculate the coarse network loss weight. To reduce the negative impact of areas with large data uncertainty on the training process in the coarse network, a loss weighting strategy is adopted. Based on the data uncertainty of each predicted pixel color, the influence of areas with low model accuracy on the training model is adaptively reduced. For the color of the predicted pixel, the data uncertainty weight can be expressed as: , in order to avoid excessive weight, we use the normalization method to obtain We can get the weighted loss of the coarse network, see formula :
[0112]
[0113] Step 5.3, Differentiable Rendering Norm Loss, in order to prevent the fine network model from over-relying on uncertainty parameters, a norm loss is designed to ensure that the color estimation value taking uncertainty into account Compared with the estimate without considering the uncertainty The difference between them is as small as possible, as shown in the formula :
[0114]
[0115] in , is the weight coefficient of the normative loss;
[0116] Step 5.4, model overall loss. In summary, we propose a complex shape credible twin modeling method for evidence fusion neural radiation field. The overall loss is shown in formula (38):
[0117]
[0118] Step 6: Get test data and perform coarse sampling on the light corresponding to the pixel color to be predicted. The specific content is the same as step 1.
[0119] Step 7: Use the resampling strategy to perform fine sampling. The coarse sampling points are input into the coarse network to obtain the volume density and calculate the corresponding weights. , which is defined as the formula :
[0120]
[0121] in Indicates the The weight of the sampling points, It is The volume density of the sampling points, It is the total number of sampling points on the ray in coarse sampling. A higher volume density means that the area contributes more to the rendering of the model.
[0122] The fine network is based on the weights of the coarse sampling points Determine the position of the additional sampling point in each ray and perform further sampling based on this. The position of the fine sampling point is determined by the formula Sure:
[0123]
[0124] in Indicates the distance between two adjacent sampling points in coarse sampling;
[0125] In step 8, the fine sampling results are input into the trained fine network to obtain the volume density, color, and uncertainty information of each point in space. Through the volume rendering module fused with the normal inverse gamma distribution described in step 4, new views and uncertainty views can be rendered. The model performance can be evaluated by combining the real images of the test data.
[0126] Table 1
[0127]
[0128] like Figure 4 As shown in Table 1, in order to verify the effectiveness of this scheme in the task of trustworthy modeling of complex surface defects, the present invention constructs an experimental scenario based on the llff public dataset (17 training images and 3 test images), and performs model evaluation indicators through model training and output results:
[0129] The present invention selects the 400k-step model fully fitting training condition and uses PSNR (peak signal-to-noise ratio), SSIM, and LPIPS as the measures of the difference between the predicted image and the real image.
[0130] In another aspect, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0131] On the other hand, the present invention further discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0132] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any of the neural radiation field trusted twin modeling methods based on evidence fusion in the above-mentioned embodiments.
[0133] It is understandable that the system, device and storage medium provided in the embodiments of the present invention correspond to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts of the above methods.
[0134] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk, tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0135] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0136] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.
[0137] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A neural radiation field trusted twin modeling method based on evidence fusion, characterized by: The following steps are included: Step 1: Acquire multi-view image data for 3D scene reconstruction; Step 2: Construct a neural radiation field model and embed it into an evidence deep learning framework to model the radiation field in three-dimensional space. The attributes of spatial points are modeled as volume density and color and their uncertainty parameters. Step 3: Weight the loss based on data uncertainty. Use the parameters of the normal inverse gamma distribution to calculate the data uncertainty of each sampling point and weight the loss of each sample in the coarse network. Step 4: Adaptive resampling. Based on the model uncertainty estimation, an adaptive resampling strategy is proposed during the training process to increase the sampling density in areas with high model uncertainty and reduce the sampling density in areas with low model uncertainty. Step 5: Based on the volume rendering method of neural radiation field, the volume rendering method of normal inverse gamma distribution is integrated, and the uncertainty parameters of all sampling points are integrated to generate a rendered image and uncertainty quantification map; Step 6: credible evidence regression, combining differentiable rendering norm loss and uncertainty loss function to optimize the global radiation field; Step 2 includes the following steps: Step 2.1, neural radiation field input construction: The input multi-view image data The corresponding camera pose is input into the neural radiation field model, where the image data is RGB three-channel, with width and height respectively W0 and H0. The input image data is subdivided into coarse sampling and fine sampling during the input construction phase: Coarse sampling: where N coarse is the number of coarse samples, x coarse Represents the three-dimensional coordinates of the coarse sampling points; Fine sampling: where N Fine is the number of fine sampling, x fine Represents the three-dimensional coordinates of the coarse sampling points; The network receives the spatial coordinates x = (x, y, z) and the viewing direction As input; Step 2.2, Evidence Deep Learning Module Construction: The evidence deep learning module is embedded in the neural radiance field model to model the uncertainty of each spatial point, assuming that the color C of the sampling point i,j , corresponding to i∈{R,G,B}, j is the sampling point on the light, each obeys an independent normal distribution, that is, Its mean μ i,j and variance Controlled by the normal inverse gamma distribution parameters: mean Accuracy parameters Shape parameters scale parameter Then its probability density function is defined as The fused evidence model will output the following parameters: Color mean δ j (x,d): Output by the network, j∈{R,G,B}, represents the mean of color prediction; Precision parameter v j (x): used to reflect the confidence of the model in color prediction, defined as the inverse of the accuracy, where τ 2 is the variance of the predicted color; Shape parameter α j (x): represents the prediction stability of the model at that spatial point; a higher α value indicates that the model has higher stability at that point; scale parameter β j (x): used to control the range of color distribution, combined with α to reflect the discrete degree of the predicted distribution; At each sampling point, the color prediction follows the normal inverse gamma distribution NIG(δ,γ,α,β), and the uncertainty of the model is quantified by these parameters; Step 2.3, neural radiation field network construction: The neural radiation field network structure consists of two parts: a coarse network and a fine network. The structures of the two are exactly the same, both consisting of two multi-layer perceptrons (MLPs). The input and output of the network are constructed as follows: The first MLP is responsible for extracting geometric information and its uncertainty parameters from the spatial coordinates, generating volume density σ(x), internal implicit features f, accuracy parameter υ, shape parameter α, and scale parameter β. Its parameter is θ1: The color mean δ passes through a second multi-layer perceptron; this MLP uses the spatial coordinate x and the viewing direction d, combined with the internal implicit feature f, to predict the color mean δ, capturing the scene's lighting characteristics and color changes related to viewing angles; its parameter is θ2: where δ=[δ R ,δ G ,δ B ], R, G, and B represent the NIG estimated means of the red, green, and blue channels, respectively.
2. The method for credible twin modeling of neural radiation fields based on evidence fusion according to claim 1 is characterized by: Step 3 includes the following steps: Step 3.1, calculate data uncertainty. Data uncertainty represents the error caused by data noise in the color prediction of each spatial point. For the k-th pixel and color channel i∈{R,G,B} in the training batch, calculate data uncertainty and model uncertainty: where β i,j Represents the scale parameter after volume rendering fusion, α i,j Represents the shape parameters after volume rendering fusion, v i,k Represents the accuracy parameter after volume rendering fusion. The greater the data uncertainty, the greater the prediction variance of the sampling point, and the model's prediction result for this point is more vague or unreliable; Step 3.2: Calculate the weighted coarse network loss. This strategy uses a weighted loss strategy to adaptively reduce the impact of regions with low model accuracy on the training model based on the data uncertainty of each predicted pixel color, thereby improving the model's learning effect in regions with low data uncertainty. For predicting the color of a pixel, the data uncertainty weight is expressed as: In order to avoid excessive weight, the normalization method is used to obtain w i,j For size N batch The uncertainty weight of the k-th pixel and color channel i in the training batch, U Aleatoric The data uncertainty of the color prediction for this sampling point; The weighted loss of the coarse network is obtained as shown in formula (1):
3. The method for credible twin modeling of neural radiation fields based on evidence fusion according to claim 2 is characterized by: Step 4 includes: Step 4.1, hierarchical sampling, first perform preliminary sampling through the coarse network to obtain the key areas, and then use the fine network for fine sampling; First, use the coarse network to perform preliminary sampling of the light and calculate the volume density and weight of each sampling point j; the number of sampling points is N coarse , and evenly distributed in the interval of light [t n ,t f ] on, t n ,t f Represents the closest and farthest distances between the sampling point and the optical center coordinates of the light, so each sampling point The position of is given by formula (2): The sampling points are input into the coarse network to obtain the volume density and the corresponding weight w can be calculated, which is defined as formula (3): where w j represents the weight of the jth sampling point, σ j is the volume density of the jth sampling point, N coarse It is the total number of sampling points on the ray in coarse sampling. A higher volume density means that the area contributes more to the rendering of the model. The fine network is based on the weight w of the coarse sampling point j Determine the position of the additional sampling point in each ray and perform further sampling based on this. The position of the fine sampling point is determined by formula (4): in Indicates the distance between two adjacent sampling points in coarse sampling; Step 4.2, adaptive resampling, adaptively adjusts the sampling density to more accurately model high uncertainty areas and dynamically allocates sampling points based on model uncertainty; First, a preliminary volume density distribution and normal inverse gamma parameters are generated by the coarse network, combined with the model uncertainty U described in step 3 Epistemic Dynamically adjust the fine sampling points; introduce model uncertainty into the sampling weight calculation, as shown in formula (5): Where λ is the adjustment coefficient, which is used to control the impact of uncertainty on the weight, and n is the number of the sampling point with the maximum uncertainty of the model; Based on these weights, more sampling points are dynamically allocated to high uncertainty areas. In the thin network, the locations of the newly added sampling points are as shown in formula (6): During the training phase, the resampling strategy concentrates resources on learning difficult areas and improves the ability to capture details; during the inference phase, stratified sampling is used to balance efficiency and quality, focusing on refining key areas.
4. The method for credible twin modeling of neural radiation fields based on evidence fusion according to claim 3 is characterized by: Step 5 includes the following steps: Step 5.1, weighted fusion of sampling points. For multiple sampling points j∈{1,2,...,N} along the ray, the color channel i∈{R,G,B} of each sampling point j corresponds to the parameters of the normal inverse gamma distribution: mean δ i,j , precision v i,j , shape α i,j , scale β i,j , and the weight w associated with the uncertainty j ; These parameters are calculated by weighted fusion to obtain the expected color value δ of the pixel numbered k k,i and uncertainty parameter v k,i ,α k,i ,β k,i , as shown in formulas (7)(8)(9)(10): in N represents the number of sampling points on the ray, σ j represents the volume density of the jth sampling point, Δ j =t j+1 -t j Represents the distance between two adjacent sampling points; Step 5.2: Generate uncertainty quantification map. Combine the data uncertainty and model uncertainty definitions in step 3.1 to derive the pixel data uncertainty. Model uncertainty These uncertainty quantifications help assess the confidence in the model's predictions in different regions.
5. The method for credible twin modeling of neural radiation fields based on evidence fusion according to claim 4 is characterized by: Step 6 includes the following steps: In step 6.1, a new regression loss function is introduced to learn the parameters of the normal inverse gamma distribution using the input samples; For a batch size of N batch For each ray in, calculate the color normal inverse gamma negative log-likelihood loss L for each color channel i∈{R,G,B} fine,NIG , see formula (11): Among them, Ω k,i =2β k,i (1+v k,i ), C k,i Indicates the true color of light, δ k,i 、v k,i , α k,i , β k,i are the mean, precision, shape and scale parameters of the k-th ray color channel i, λ N is the standard loss coefficient; This loss function minimizes the error between the predicted color mean and the true color by maximizing the log-likelihood of the normal inverse gamma distribution of the color channel, taking into account the confidence of each prediction, and is applied to both coarse and fine network optimization. Combined with the loss weighting algorithm described in step 3.2, the coarse network loss is Rewritten as formula (12): In step 6.2, a regularization loss is designed to ensure that the color estimate δ is considered with uncertainty. k,i Compared with the estimate without considering the uncertainty The difference between them should be as small as possible, as shown in formula (13): in λ norm is the weight coefficient of the normative loss; In step 6.3, the overall loss of the complex shape credible twin modeling method of the evidence fusion neural radiation field is shown in formula (14):
6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
New view angle synthesis method based on point features and neural radiation field
CN118135363A