Scene uncertainty quantification method and system, terminal and storage medium

By constructing a prediction model and a refinement model, combining the view generation model and significance binary graph, the uncertainty value in the three-dimensional reconstruction process is solved, and the problem of uncertainty calculations in the existing technology is affected by environmental factors, achieving more accurate uncertainty analysis.

CN120088494APending Publication Date: 2025-06-03GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411989031.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

When calculating uncertainty in the three-dimensional reconstruction process, the prior art is susceptible to environmental factors, resulting in inaccurate uncertainty results.

Method used

By constructing a prediction model and refinement model, multiple target feature maps of the target image are obtained, and the uncertainty value of the pixel value is calculated through the view generation model, and the final uncertainty value is calculated by combining the significance binary graph.

Benefits of technology

It improves the accuracy of uncertainty calculations, reduces the impact of environmental factors on the results, and provides a more reliable three-dimensional reconstruction uncertainty analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088494A_ABST
    Figure CN120088494A_ABST
Patent Text Reader

Abstract

The invention discloses a scene uncertainty quantification method and system, a terminal and a storage medium. The method comprises the following steps: constructing a plurality of saliency binary images of a target object through a saliency check model; defining a deformation value according to vector displacement of pixel values in the plurality of feature maps of the target object so as to train a view generation model; generating a model for the target feature map view, outputting a predicted pixel value color, comparing the predicted pixel value color with a real pixel value color to obtain a ray residual error, and calculating an uncertainty value of the target feature map according to the ray residual error and a Jacobian matrix; and calculating a final uncertainty value of each target feature map according to the saliency binary image and the uncertainty value. The invention introduces a saliency target detection method for identifying a target object and generating a binary image, quantifies an uncertainty value by using a plug-and-play probability method, and calculates a spatial uncertainty field based on the target object by combining the binary image and the uncertainty value of the existing scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of scene parameter quantization, and particularly to a quantization method, system, terminal and computer-readable storage medium for scene uncertainty. Background Art

[0002] Three-dimensional reconstruction refers to depicting a real scene as a mathematical model that conforms to computer logic expression, which can assist research such as cultural relic protection, game development, and architectural design.

[0003] Currently, for the process of three-dimensional reconstruction, uncertainty is mainly incorporated into NeRF (Neural Radiance Fields), and it is used to iteratively generate and plan the next best view, or by identifying the regions in NeRF that can be spatially perturbed with the least impact on the reconstruction loss, a scene uncertainty field is established.

[0004] However, the existing uncertainty calculation methods cannot well measure the uncertainty of the reconstruction subject, making the planning results greatly affected by background factors, resulting in large errors in the uncertainty results.

[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0006] The main purpose of the present invention is to provide a quantization method, system, terminal and computer-readable storage medium for scene uncertainty, aiming to solve the problem that the calculation process of the uncertainty result in the existing technology is easily affected by the environment, resulting in inaccurate uncertainty results.

[0007] To achieve the above object, the present invention provides a quantization method for scene uncertainty, and the quantization method for scene uncertainty includes the following steps:

[0008] Construct a prediction model and a refinement model, input the target image input by the user into the prediction model, output a plurality of target feature maps, and input all the target feature maps into the refinement model for bilinear interpolation processing to output a plurality of saliency binary maps;

[0009] Obtain the vector displacement of the pixel values in each target feature map, define the deformation value of each vector displacement, and use all the deformation values to train the pre-constructed initial view generation model to obtain a view generation model;

[0010] Input the multiple target feature maps into the view generation model respectively, output the corresponding predicted pixel value colors, compare each predicted pixel value color with the true pixel value color of the corresponding target feature map, obtain multiple ray residual errors, and calculate the uncertainty values of the pixel values in the multiple target feature maps according to the multiple ray residual errors and the Jacobian matrix;

[0011] Calculate the final uncertainty value of the pixel value in each target feature map according to the saliency binary map and the uncertainty value of each target feature map.

[0012] Optionally, in the method for quantifying scene uncertainty, wherein the prediction model and the refinement model are constructed, the target image input by the user is input into the prediction model, multiple target feature maps are output, and all the target feature maps are respectively input into the refinement model for bilinear interpolation processing, and multiple saliency binary maps are output, specifically including:

[0013] Construct an encoder and a decoder, add a bridging stage between the encoder and the decoder to obtain a prediction model, and construct a residual encoder and a residual decoder, add a bridging stage between the residual encoder and the residual decoder to obtain a refinement model;

[0014] Input the target image input by the user into the prediction model, and use the encoder of the prediction model to sample the target image according to a preset pooling window to obtain the feature maps corresponding to multiple stages in the encoder and the pixel values in each feature map;

[0015] Stitch the multiple feature maps in the current stage with the multiple feature maps in the previous stage, and use the decoder of the prediction model to perform multi-channel output on the stitched content in different stages to obtain multiple target feature maps in the final stage;

[0016] Input all the target feature maps into the refinement model, use the residual encoder of the refinement model to downsample all the target feature maps according to the pooling window, and use the residual decoder to upsample all the target feature maps by bilinear interpolation to output the saliency binary map of each target feature map.

[0017] Optionally, in the method for quantifying scene uncertainty, wherein the encoder and the decoder are constructed, a bridging stage is added between the encoder and the decoder to obtain a prediction model, and a residual encoder and a residual decoder are constructed, and a bridging stage is added between the residual encoder and the residual decoder to obtain a refinement model, specifically including:

[0018] In the first first preset number of stages of the encoder, a plurality of convolutional layers are respectively constructed, and in the second second preset number of stages of the encoder, a plurality of basic residual blocks are respectively constructed;

[0019] In all stages of the decoder, a plurality of convolutional layers, batch normalization layers, and rectified linear units are respectively constructed, and a bridging stage is used to connect the encoder and the decoder to obtain a prediction model;

[0020] In each stage of the residual encoder and the residual decoder, a convolutional layer, a batch normalization layer, and a rectified linear unit are respectively constructed, and a bridging stage is used to connect the residual encoder and the residual decoder to obtain a refinement model.

[0021] Optionally, in the method for quantifying scene uncertainty, wherein, obtaining the vector displacement of the pixel values in each of the target feature maps, defining the deformation value of each of the vector displacements, and training the pre-constructed initial view generation model using all the deformation values to obtain a view generation model, specifically including:

[0022] Taking the maximum eigenvalue in each of the target feature maps as the pixel value, introducing each of the pixel values into the deformation field to store at the grid vertices, and taking the vector displacement of each grid vertex as the vector displacement of the corresponding pixel value;

[0023] By the trilinear interpolation method, defining the deformation value D of each of the pixel values θ(x) :

[0024] D θ(x) = Trilinear(x, θ);

[0025] wherein, θ represents the deformation parameter of the deformation field, x represents the vector displacement, and Trilinear represents the trilinear interpolation method;

[0026] Constructing an initial view generation model, and reparameterizing the initial view generation model using all the deformation values to obtain a view generation model.

[0027] Optionally, in the method for quantifying scene uncertainty, wherein, inputting the plurality of target feature maps into the view generation model respectively, outputting the corresponding predicted pixel value colors, comparing each of the predicted pixel value colors with the true pixel value colors of the corresponding target feature maps, obtaining a plurality of ray residual errors, and calculating the uncertainty values of the pixel values in the plurality of target feature maps according to the plurality of ray residual errors and the Jacobian matrix, specifically including:

[0028] Inputting the plurality of target feature maps into the view generation model respectively, and outputting the predicted pixel value colors of the pixel values corresponding to each of the target feature maps;

[0029] Obtain the true pixel value color corresponding to each of the target feature maps, and calculate the degree of consistency between each predicted pixel value color and the true pixel value color of the corresponding target feature map, to obtain the ray residual error ∈ corresponding to each target feature map θ (r):

[0030]

[0031] where r represents the light ray of the camera, represents the predicted pixel value color, represents the true pixel value color, and n represents the number of training samples;

[0032] Define the Jacobian matrix J θ (r):

[0033]

[0034] where the Jacobian matrix represents the first-order partial derivative of the predicted pixel color with respect to the parameter θ;

[0035] According to the filtered ray residual error and the Jacobian matrix, calculate the diagonal entry approximate covariance matrix ∑ for each target feature map:

[0036]

[0037] where diag() is used to construct a matrix, R represents the parameters of the view generation model, T represents the transpose, λ represents the regularization parameter, and I represents the training image;

[0038] According to each diagonal entry approximate covariance matrix, screen the grid vertices corresponding to each ray residual error to obtain the final target feature map after screening;

[0039] Through the trilinear interpolation method, obtain the uncertainty value of the pixel value in each final target feature map.

[0040] Optionally, for the method for quantifying the scene uncertainty, where the uncertainty value of the pixel value in each final target feature map is obtained through the trilinear interpolation method, specifically includes:

[0041] Obtain the variance vector of the diagonal entry approximate covariance matrix of each final target feature map;

[0042] Define the uncertainty field U(x) through the trilinear interpolation method:

[0043] U(x) = Trilinear(x, σ);

[0044] where σ represents the variance vector;

[0045] Determine the uncertainty values of the pixel values in each of the final target feature maps in the uncertainty field according to the variance vectors of each of the final target feature maps.

[0046] Optionally, for the method for quantifying the scene uncertainty, wherein calculating the final uncertainty values of the pixel values in each of the target feature maps according to the significance binary maps and the uncertainty values of each of the target feature maps specifically includes:

[0047] Calculate the significance value of each of the final target feature maps according to the significance binary maps corresponding to the pixel values of all the final target feature maps;

[0048] Calculate the final uncertainty values of the pixel values in each of the final target feature maps according to each of the significance values and the uncertainty values corresponding to each of the final target feature maps.

[0049] In addition, to achieve the above object, the present invention further provides a system for quantifying scene uncertainty, wherein the system for quantifying scene uncertainty includes:

[0050] A binary map construction module, configured to construct a prediction model and a refinement model, input a target image input by a user into the prediction model, output a plurality of target feature maps, and input all the target feature maps into the refinement model respectively for bilinear interpolation processing, and output a plurality of significance binary maps;

[0051] A view generation model construction module, configured to obtain the vector displacements of the pixel values in each of the target feature maps, define the deformation values of each of the vector displacements, and train the constructed initial view generation model by using all the deformation values to obtain a view generation model;

[0052] An uncertainty value calculation module, configured to input the plurality of target feature maps into the view generation model respectively, output the corresponding predicted pixel value colors, compare each of the predicted pixel value colors with the true pixel value colors of the corresponding target feature maps, obtain a plurality of ray residual errors, and calculate the uncertainty values of the pixel values in the plurality of target feature maps according to the plurality of ray residual errors and the Jacobian matrix;

[0053] A final uncertainty value calculation module, configured to calculate the final uncertainty values of the pixel values in each of the target feature maps according to the significance binary maps and the uncertainty values of each of the target feature maps.

[0054] In addition, to achieve the above object, the present invention further provides a terminal, wherein the terminal includes: a memory, a processor, and a quantization program of scene uncertainty stored on the memory and executable on the processor. When the quantization program of scene uncertainty is executed by the processor, the steps of the quantization method of scene uncertainty as described above are implemented.

[0055] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a quantization program of scene uncertainty. When the quantization program of scene uncertainty is executed by a processor, the steps of the quantization method of scene uncertainty as described above are implemented.

[0056] In the present invention, a prediction model and a refinement model are constructed. The target image input by the user is input into the prediction model, and a plurality of target feature maps are output. All the target feature maps are respectively input into the refinement model for bilinear interpolation processing, and a plurality of saliency binary maps are output; the vector displacement of the pixel values in each target feature map is obtained, the deformation value of each vector displacement is defined, and the constructed initial view generation model is trained using all the deformation values to obtain a view generation model; the plurality of target feature maps are respectively input into the view generation model, and the corresponding predicted pixel value colors are output. The predicted pixel value color of each one is compared with the true pixel value color of the corresponding target feature map to obtain a plurality of ray residual errors. According to the plurality of ray residual errors and the Jacobian matrix, the uncertainty values of the pixel values in the plurality of target feature maps are calculated; according to the saliency binary map and the uncertainty value of each target feature map, the final uncertainty value of the pixel value in each target feature map is calculated. The present invention introduces a saliency object detection method for identifying target objects, generating binary maps, and using a plug-and-play probability method to quantify uncertainty values, combining the binary maps with the uncertainty values of the existing scene to calculate a spatial uncertainty field based on the target object. Description of the Drawings

[0057] Figure 1 is a flowchart of a preferred embodiment of the quantization method of scene uncertainty of the present invention;

[0058] Figure 2 is a schematic diagram of view extraction of a preferred embodiment of the quantization method of scene uncertainty of the present invention;

[0059] Figure 3 is a structural diagram of a preferred embodiment of the quantization system of scene uncertainty of the present invention;

[0060] Figure 4 is a schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. Detailed Embodiments

[0061] To make the objectives, technical solutions and advantages of the present invention more clear and definite, the present invention will be further described in detail below with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not intended to limit the present invention.

[0062] The method for quantifying scene uncertainty according to a preferred embodiment of the present invention is as Figure 1 shown, and the method for quantifying scene uncertainty includes the following steps:

[0063] Step S10: Construct a prediction model and a refinement model, input the target image entered by the user into the prediction model, output multiple target feature maps, and input all the target feature maps into the refinement model for bilinear interpolation processing to output multiple saliency binary maps.

[0064] Among them, the target image input into the model can select the initial view by uniformly and randomly sampling the hemisphere, and this process can be realized by the following formula:

[0065] φ = 2πu;

[0067] θ = cos -1 (1 - v);

[0068] Among them, φ represents the position of the view, both u and v represent constants, and the range is between 0 - 1 (0 or 1 can be taken), and θ represents the deformation parameter of the deformation field; after determining the position of the view, the camera view can be directed towards the center of the object, so as to randomly generate the initial viewpoints, and the candidate views (i.e., target feature maps) can be obtained according to these initial viewpoints.

[0069] Specifically, construct an encoder and a decoder, add a bridging stage between the encoder and the decoder to obtain a prediction model, and construct a residual encoder and a residual decoder, add a bridging stage between the residual encoder and the residual decoder to obtain a refinement model; input the target image entered by the user into the prediction model, use the encoder of the prediction model to sample the target image according to the preset pooling window to obtain the feature maps corresponding to multiple stages in the encoder and the pixel values in each feature map; splice the multiple feature maps in the current stage with the multiple feature maps in the previous stage, and use the decoder of the prediction model to perform multi-channel output on the spliced content at different stages to obtain multiple target feature maps in the final stage; input all the target feature maps into the refinement model, use the residual encoder of the refinement model to downsample all the target feature maps according to the pooling window, and use the residual decoder to upsample all the target feature maps by bilinear interpolation to output the saliency binary map of each target feature map.

[0070] Further, in the first first preset number of stages of the encoder, a plurality of convolutional layers are respectively constructed, and in the second second preset number of stages of the encoder, a plurality of basic residual blocks are respectively constructed; in all stages of the decoder, a plurality of convolutional layers, batch normalization layers, and rectified linear units are respectively constructed, and a bridging stage is used to connect the encoder and the decoder to obtain a prediction model; in each stage of the residual encoder and the residual decoder, a convolutional layer, a batch normalization layer, and a rectified linear unit are respectively constructed, and a bridging stage is used to connect the residual encoder and the residual decoder to obtain a refinement model.

[0071] Among them, through a saliency object detection network, the main object in the image, that is, the building to be reconstructed, is identified to generate a binary map; it mainly includes a prediction module and a refinement module.

[0072] Among them, the input convolutional layer of the encoder part of the prediction module has 64 convolutional filters with a size of 3×3 and a stride of 1, and there is no pooling operation, which makes the feature maps in the first two stages have the same spatial resolution as the input image, better retaining the original structure and local features in the image, thus ensuring the integrity of the result; the first four stages adopt the ResNet-34 (ResNet, Residual Network) structure. The residual connections in ResNet-34 help to solve the problem of gradient disappearance when the network depth increases, enabling the network to better learn deep feature representations. On the basis of maintaining the early high-resolution feature maps, combining the structural advantages of ResNet-34 can obtain global context information while not losing local details; three basic residual blocks (basic res-block) are added in the last two stages of the prediction module, and downsampling is performed through a non-overlapping max pooling layer to gradually expand the receptive field and learn features at different scales.

[0073] Among them, the decoder part of the prediction module is roughly symmetric to the encoder part, and each stage consists of three convolutional layers, batch normalization, and a ReLU (Rectified Linear Unit) activation function. A bridging stage is added between the encoder and the decoder, which consists of three convolutional layers with 512 dilation convolution (dilation = 2) filters for capturing global information.

[0074] Further, after inputting the target feature map into the prediction model, in the prediction model, according to the preset pooling window, the target image is sampled (i.e., downsampled using a non-overlapping max pooling layer). For example, if the pooling window size is selected as 2*2, then the stride is also 2, ensuring that the windows do not overlap. Starting from the upper left corner of the input feature map, a region of size 2*2 is selected as the processing unit, and the maximum value within each region is used as the element value at the corresponding position of the downsampled feature map. Then, the pooling window is moved to the right or down according to the stride to select the next region. If the size of the input feature map is H*W (H is the height and W is the width), after non-overlapping max pooling downsampling with a stride of k, the size of the output feature map becomes (H / k)*(W / k). For each stage, the encoder part collects the pixel values in the feature map, and the feature map collected in the current stage is concatenated with the feature map collected in the previous stage to obtain the target feature map for inputting to the decoder. Among them, through the multi-channel outputs of the bridging stage and each decoder stage, seven saliency maps are obtained through a common 3×3 convolutional layer, bilinear upsampling, and sigmoid function (S-shaped growth function), where the last one has the highest accuracy, as Figure 2 shown, and is passed as the final output of the prediction module to the refinement module.

[0075] Among them, the refinement module adopts a multi-scale residual refinement module based on a residual encoder-decoder architecture, including an input layer, a residual encoder, a bridge, a residual decoder, and an output layer; both the residual encoder and the residual decoder have four stages, with only one convolutional layer in each stage, and each layer has 64 filters of size 3×3, followed by batch normalization and ReLU activation function. The bridging stage also has one convolutional layer. The residual encoder uses non-overlapping max pooling for downsampling, and the residual decoder uses bilinear interpolation for upsampling to output the saliency binary map of each target feature map.

[0076] Further, before inputting the target feature map, the target feature map needs to be adjusted to the preset format size, and the finally obtained saliency binary map is also of the same preset format size. Through the saliency object detection network (i.e., the prediction model and the refinement model), the salient object region can be effectively segmented, and the fine structure with clear boundaries can be accurately predicted, so as to accurately locate the main body to be reconstructed.

[0077] Step S20: Obtain the vector displacement of the pixel values in each of the target feature maps, define the deformation value of each vector displacement, and use all the deformation values to train the constructed initial view generation model to obtain the view generation model.

[0078] Among them, according to the pixels in each target feature map, the uncertainty of the scene can be calculated by constructing an initial NeRF model (Representing Scenes as Neural Radiance Fields, view generation model).

[0079] Specifically, take the maximum eigenvalue in each of the said target feature maps as the pixel value, introduce each of the pixel values into the deformation field and store them in the grid vertices, and take the vector displacement of each grid vertex as the vector displacement of the corresponding pixel value; by the trilinear interpolation method, define the deformation value of each of the pixel values

[0080]

[0081] Among them, θ represents the deformation parameter of the deformation field, x represents the vector displacement, and Trilinear represents the trilinear interpolation method; construct an initial view generation model, and use all the deformation values to reparameterize the initial view generation model to obtain the view generation model.

[0082] Among them, the reparameterization of the initial view generation model can be achieved by introducing a deformation field (D:R D →R D )(where the R on the left side of the arrow D represents the domain (input space) of the deformation field, and the R on the right side D represents the range (output space) of the deformation field, which means that the deformation field maps the coordinates in the D-dimensional space to the space of the same dimension to obtain the deformed coordinates), select the vector displacement stored on the grid vertices with length M as a spatially meaningful parameterization form, so that θ can be represented as a matrix and define the deformation value for each spatial coordinate (i.e., pixel value) by trilinear interpolation Thus, the reparameterization of the initial view generation model is realized.

[0083] Step S30: Input the multiple target feature maps into the view generation model respectively, output the corresponding predicted pixel value colors, compare each predicted pixel value color with the true pixel value color of the corresponding target feature map to obtain multiple ray residual errors, and calculate the uncertainty values of the pixel values in the multiple target feature maps according to the multiple ray residual errors and the Jacobian matrix.

[0084] Among them, the target feature map is input into the view generation model for rendering. During the rendering process, the model calculates the propagation and interaction of light in the scene according to the position and direction of the viewing angle, determines information such as the color and brightness of each pixel, and finally generates a complete candidate view image. For example, during the calculation process, the model calculates the radiance value of each point in the scene observed from this viewing angle based on the input candidate viewing angle coordinates and combines the scene representation it has learned, and then synthesizes the image content of the candidate view. For each candidate view with calculation uncertainty, since the uncertainty is calculated based on the light beams emitted from each pixel point of the image, that is, for a set U of n candidate views n ∈H r ×W r ×scale, where H r ×W r are the pixels of the image, and scale is the ratio for compressing the image to prevent the pixels from being too large. The uncertainty estimation for a certain viewing angle can be calculated using a simple function:

[0085]

[0086] Among them, f i represents the uncertainty estimation of a certain viewing angle, X i represents the candidate view in U n ; Select the top k candidate views with large fi as the new views to be added. If these views are input into the above-mentioned constructed saliency object detection network, the saliency binary map of the corresponding target feature map is obtained, so as to obtain the saliency value of each pixel in each map.

[0087] Specifically, input multiple said target feature maps into said view generation model respectively, and output the predicted pixel value color of the pixel value corresponding to each said target feature map; obtain the true pixel value color corresponding to each said target feature map, and calculate the consistency degree between each said predicted pixel value color and the true pixel value color corresponding to the said target feature map respectively, to obtain the ray residual error ∈ θ (r):

[0088]

[0089] Among them, r represents the light ray of the camera, represents the predicted pixel value color, represents the true pixel value color, and n represents the number of training samples; define the Jacobian matrix J θ (r):

[0090]

[0091] Among them, the Jacobian matrix represents the first-order partial derivative of the predicted pixel color with respect to the parameter θ; according to the filtered ray residual error and the Jacobian matrix, the diagonal entry approximate covariance matrix ∑ of each target feature map is calculated:

[0092]

[0093] Among them, diag() is used to construct a matrix, R represents the parameter of the view generation model, T represents the transpose, λ represents the regularization parameter, and I represents the training image; according to each diagonal entry approximate covariance matrix, the grid vertices corresponding to each ray residual error are filtered to obtain the final target feature map after filtering; through the trilinear interpolation method, the uncertainty value of the pixel value in each final target feature map is obtained.

[0094] Among them, after reparameterizing the initial view generation model by introducing the deformation field, the predicted pixel color calculated according to the light ray can, under the condition of the given deformation field parameter (i.e., the parameter θ), pass through obeys a normal distribution with a mean of and a variance of which can be used to measure the consistency degree between the predicted pixel color output by the model and the real pixel color. This process can be expressed in the form of likelihood ( represents the predicted pixel color, and N represents the normal distribution). In order to more effectively estimate the deformation field parameter θ, a regularization independent Gaussian prior (θ ∼ N(0, λ -1 )) is placed on it; among them, λ is the regularization parameter, which determines the variance λ -1 of the prior distribution, thereby controlling the degree of constraint on the parameter θ.

[0095] Furthermore, by combining the prior distribution with the likelihood form mentioned above, the posterior distribution is constructed using Bayes' theorem The posterior distribution combines the prior knowledge (the assumption about the parameter θ when there is no observed data) and the likelihood information obtained based on the observed data ( represents the set of training images), reflecting the uncertainty distribution of the parameter θ after considering the data and the prior assumption.

[0096] The negative log-likelihood of the deformation field Among them, E n represents the expectation operation on all images in the set of training images , represents the expectation operation on the light ray r from the nth training image I n . is the predicted pixel color With the true pixel color The mean squared error on ray r, which is used to measure the difference between the model prediction and the true value, and the model's fit to the data is optimized by minimizing it. And λ|θ| 2 is the regularization term, and |θ| 2 is the square of the L2 norm of the parameter θ, which penalizes the magnitude of the parameter to prevent overfitting of the model and ensure that the model has good generalization ability.

[0097] Furthermore, according to the Fisher information (the Fisher information of the deformation field with deformation parameter θ), H(θ) can be defined (where H(θ) represents the Hessian matrix of θ (the Hessian matrix, a square matrix composed of the second-order partial derivatives of a multivariate function). By representing the Fisher information as the negative value of the Hessian matrix, the properties of the Hessian matrix can be used to approximately calculate the posterior distribution of the parameter, thereby realizing the quantification of the model uncertainty.

[0098] The relationship between it and the fractional variance of the parameter can also be determined through the Fisher information, and then the ray residual error ∈ θ (r) and the Jacobian matrix J θ (r) are defined. Among them, the Jacobian matrix represents the matrix composed of the first-order partial derivatives of the predicted pixel color with respect to the parameter θ, which is used to describe the rate of change of the model output predicted pixel color with respect to the parameter θ, and we get Then combined with By sampling and approximating the expectation for multiple rays (for example, R rays), an approximate expression of H(θ) is obtained:

[0099] Furthermore, obtain the variance vector of the approximate covariance matrix of the diagonal entries of each of the final target feature maps; define the uncertainty field U(x) through the trilinear interpolation method:

[0100] U(x) = Trilinear(x, σ);

[0101] where σ represents the variance vector; according to the variance vector of each of the final target feature maps, determine the uncertainty value of the pixel values in each of the final target feature maps in the uncertainty field.

[0102] Among them, since the grid vertices corresponding to each vector entry of the deformation field, that is, the influence of each vector entry on the coordinates is only limited to the interior of the local grid cell containing the grid vertex. When calculating H(θ), due to this limitation, only a few parameters related to the grid vertex will participate in the calculation of the second-order partial derivative, making a large number of elements in H(θ) equal to 0, showing sparsity. Therefore, through approximation processing (such as only considering the diagonal terms in this embodiment), the number of related parameters is reduced.

[0103] Among them, only consider the diagonal entries of H(θ) to approximate the covariance matrix: consider the (root) diagonal entries of ∑ to obtain the variance vector σ, and its norm is a positive scalar (σ = |σ| 2 ), measure the local spatial uncertainty of the radiation field at each grid vertex, and define the spatial uncertainty field through trilinear interpolation (U(x) = Trilinear(x, σ)), so as to obtain the uncertainty value U corresponding to each pixel position. i .

[0104] Step S40: Calculate the final uncertainty value of the pixel value in each of the target feature maps according to the significance binary map and the uncertainty value of each of the target feature maps.

[0105] Specifically, calculate the significance value of each of the final target feature maps according to the significance binary maps corresponding to the pixel values of all the final target feature maps; calculate the final uncertainty value of the pixel value in each of the final target feature maps according to each of the significance values and the uncertainty value corresponding to each of the final target feature maps.

[0106] Among them, for the significance binary map of each target feature map, the significance value of each final target feature map (the result after screening the target feature map) can be calculated, and then according to the uncertainty value U corresponding to each pixel position therein i , the final uncertainty value of each pixel in the final target feature map can be obtained:

[0107] U′ i = U i * S i ;

[0108] Among them, U′ i represents the final uncertainty value of pixel i, and S represents the significance value of pixel i.

[0109] The present invention introduces a significance object detection method for identifying target objects, generating binary maps, and using a plug-and-play probability method to quantify uncertainty values, combining the binary maps with the uncertainty values of the existing scene, and calculating a spatial uncertainty field based on the target object.

[0110] Furthermore, asFigure 3 As shown, based on the above-mentioned method for quantifying scene uncertainty, the present invention also correspondingly provides a system for quantifying scene uncertainty, wherein the system for quantifying scene uncertainty includes:

[0111] A binary map construction module 51, configured to construct a prediction model and a refinement model, input a target image input by a user into the prediction model, output a plurality of target feature maps, and input all the target feature maps into the refinement model respectively for bilinear interpolation processing to output a plurality of saliency binary maps;

[0112] A view generation model construction module 52, configured to obtain the vector displacement of pixel values in each target feature map, define the deformation value of each vector displacement, and train the constructed initial view generation model by using all the deformation values to obtain a view generation model;

[0113] An uncertainty value calculation module 53, configured to input a plurality of the target feature maps into the view generation model respectively, output the corresponding predicted pixel value colors, compare each predicted pixel value color with the true pixel value color of the corresponding target feature map to obtain a plurality of ray residual errors, and calculate the uncertainty values of the pixel values in the plurality of target feature maps according to the plurality of ray residual errors and the Jacobian matrix;

[0114] A final uncertainty value calculation module 54, configured to calculate the final uncertainty values of the pixel values in each target feature map according to the saliency binary map and the uncertainty value of each target feature map.

[0115] Furthermore, as Figure 4 shown, based on the above-mentioned method and system for quantifying scene uncertainty, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 4 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively.

[0116] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as the hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in some other embodiments, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes of the installed terminal, etc. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a quantization program 40 for scene uncertainty is stored on the memory 20, and this quantization program 40 for scene uncertainty can be executed by the processor 10, thereby implementing the quantization method for scene uncertainty in the present application.

[0117] The processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips in some embodiments, and is used to run the program codes stored in the memory 20 or process data, such as executing the quantization method for scene uncertainty, etc.

[0118] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and for displaying a visual user interface. Components of the terminal communicate with each other through a system bus.

[0119] In one embodiment, when the processor 10 executes the quantization program 40 for scene uncertainty in the memory 20, the following steps are implemented:

[0120] Construct a prediction model and a refinement model, input a target image input by a user into the prediction model, output a plurality of target feature maps, input all the target feature maps into the refinement model for bilinear interpolation processing, and output a plurality of saliency binary maps;

[0121] Obtain the vector displacement of pixel values in each target feature map, define the deformation value of each vector displacement, and use all the deformation values to train the constructed initial view generation model to obtain a view generation model;

[0122] Input the multiple target feature maps into the view generation model respectively, output the corresponding predicted pixel value colors, compare each predicted pixel value color with the true pixel value color of the corresponding target feature map, obtain multiple ray residual errors, and calculate the uncertainty values of the pixel values in the multiple target feature maps according to the multiple ray residual errors and the Jacobian matrix;

[0123] Calculate the final uncertainty value of the pixel value in each target feature map according to the saliency binary map and the uncertainty value of each target feature map.

[0124] Among them, the construction of the prediction model and the refinement model, input the target image input by the user into the prediction model, output multiple target feature maps, and input all the target feature maps into the refinement model for bilinear interpolation processing, output multiple saliency binary maps, specifically including:

[0125] Construct an encoder and a decoder, add a bridging stage between the encoder and the decoder to obtain a prediction model, and construct a residual encoder and a residual decoder, add a bridging stage between the residual encoder and the residual decoder to obtain a refinement model;

[0126] Input the target image input by the user into the prediction model, use the encoder of the prediction model, sample the target image according to a preset pooling window, and obtain the feature maps corresponding to multiple stages in the encoder and the pixel values in each feature map;

[0127] Stitch the multiple feature maps in the current stage with the multiple feature maps in the previous stage, and use the decoder of the prediction model to perform multi-channel output on the stitched content in different stages to obtain multiple target feature maps in the final stage;

[0128] Input all the target feature maps into the refinement model, use the residual encoder of the refinement model to downsample all the target feature maps according to the pooling window, and use the residual decoder to upsample all the target feature maps by bilinear interpolation to output the saliency binary map of each target feature map.

[0129] Among them, the construction of the encoder and the decoder, adding a bridging stage between the encoder and the decoder to obtain a prediction model, and constructing a residual encoder and a residual decoder, adding a bridging stage between the residual encoder and the residual decoder to obtain a refinement model, specifically including:

[0130] In the first preset number of stages in front of the encoder, construct multiple convolutional layers respectively, and in the second preset number of stages behind the encoder, construct multiple basic residual blocks respectively;

[0131] Construct multiple convolutional layers, batch normalization layers, and rectified linear units respectively in all stages of the decoder, and connect the encoder and the decoder using a bridging stage to obtain a prediction model;

[0132] In each stage of the residual encoder and the residual decoder, construct a convolutional layer, a batch normalization layer, and a rectified linear unit respectively, and connect the residual encoder and the residual decoder using a bridging stage to obtain a refinement model.

[0133] Among them, obtaining the vector displacement of the pixel values in each of the target feature maps, defining the deformation value of each vector displacement, and training the constructed initial view generation model using all the deformation values to obtain a view generation model specifically includes:

[0134] Taking the maximum eigenvalue in each of the target feature maps as the pixel value, introducing each pixel value into the deformation field to store at the grid vertices, and taking the vector displacement of each grid vertex as the vector displacement of the corresponding pixel value;

[0135] Defining the deformation value D of each pixel value through trilinear interpolation method θ(x) :

[0136] D θ(x) = Trilinear(x,θ);

[0137] Among them, θ represents the deformation parameter of the deformation field, x represents the vector displacement, and Trilinear represents the trilinear interpolation method;

[0138] Construct an initial view generation model, and reparameterize the initial view generation model using all the deformation values to obtain a view generation model.

[0139] Among them, inputting the multiple target feature maps into the view generation model respectively, outputting the predicted pixel value colors corresponding thereto, comparing each predicted pixel value color with the true pixel value color of the corresponding target feature map, obtaining multiple ray residual errors, and calculating the uncertainty values of the pixel values in the multiple target feature maps according to the multiple ray residual errors and the Jacobian matrix specifically includes:

[0140] Inputting the multiple target feature maps into the view generation model respectively, and outputting the predicted pixel value colors of the pixel values corresponding to each target feature map;

[0141] Obtaining the true pixel value colors corresponding to each target feature map, and calculating the degree of consistency between each predicted pixel value color and the true pixel value color of the corresponding target feature map respectively to obtain the ray residual error ∈ θ (r):

[0142]

[0143] Among them, r represents the light of the camera, represents the predicted pixel value color, represents the true pixel value color, and n represents the number of training samples;

[0144] Define the Jacobian matrix J θ (r):

[0145]

[0146] Among them, the Jacobian matrix represents the first-order partial derivative of the predicted pixel color with respect to the parameter θ;

[0147] According to the filtered ray residual error and the Jacobian matrix, calculate the diagonal entry approximate covariance matrix ∑ of each of the target feature maps:

[0148]

[0149] Among them, diag() is used to construct a matrix, R represents the parameters of the view generation model, T represents the transpose, λ represents the regularization parameter, and I represents the training image;

[0150] According to each of the diagonal entry approximate covariance matrices, screen the grid vertices corresponding to each of the ray residual errors to obtain the final target feature map after screening;

[0151] By means of trilinear interpolation, obtain the uncertainty value of the pixel value in each of the final target feature maps.

[0152] Among them, the method of obtaining the uncertainty value of the pixel value in each of the final target feature maps by means of trilinear interpolation specifically includes:

[0153] Obtain the variance vector of the diagonal entry approximate covariance matrix of each of the final target feature maps;

[0154] Define the uncertainty field U(x) by means of trilinear interpolation:

[0155] U(x) = Trilinear(x, σ);

[0156] Among them, σ represents the variance vector;

[0157] According to the variance vector of each of the final target feature maps, determine the uncertainty value of the pixel value in each of the final target feature maps in the uncertainty field.

[0158] Among them, calculating the final uncertainty value of the pixel value in each of the target feature maps according to the saliency binary map and the uncertainty value of each of the target feature maps specifically includes:

[0159] Calculating the saliency value of each of the final target feature maps according to the saliency binary maps corresponding to the pixel values of all the final target feature maps;

[0160] Calculating the final uncertainty value of the pixel value in each of the final target feature maps according to each of the saliency values and the uncertainty value corresponding to each of the final target feature maps.

[0161] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a quantization program for scene uncertainty, and when the quantization program for scene uncertainty is executed by a processor, the steps of the quantization method for scene uncertainty as described above are implemented.

[0162] In summary, the present invention provides a quantization method for scene uncertainty and related devices. The method includes: constructing a prediction model and a refinement model, inputting a target image input by a user into the prediction model, outputting a plurality of target feature maps, respectively inputting all the target feature maps into the refinement model for bilinear interpolation processing, and outputting a plurality of saliency binary maps; obtaining the vector displacement of the pixel value in each of the target feature maps, defining the deformation value of each of the vector displacements, training the constructed initial view generation model by using all the deformation values to obtain a view generation model; respectively inputting the plurality of target feature maps into the view generation model, outputting the corresponding predicted pixel value colors, comparing each of the predicted pixel value colors with the true pixel value color of the corresponding target feature map to obtain a plurality of ray residual errors, calculating the uncertainty values of the pixel values in the plurality of target feature maps according to the plurality of ray residual errors and the Jacobian matrix; calculating the final uncertainty value of the pixel value in each of the target feature maps according to the saliency binary map and the uncertainty value of each of the target feature maps. The present invention introduces a saliency object detection method for identifying target objects, generating binary maps, and using a plug-and-play probability method to quantify uncertainty values, combining the binary maps with the uncertainty values of the existing scenes to calculate a spatial uncertainty field based on the target objects.

[0163] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal including such element.

[0164] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above-described embodiment methods can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer. When the program is executed, it can include the processes of the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0165] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for quantifying scene uncertainty, characterized in that: The method for quantifying the scene uncertainty includes: Constructing a prediction model and a refinement model, inputting a target image input by a user into the prediction model, outputting a plurality of target feature maps, inputting all the target feature maps into the refinement model for bilinear interpolation processing, and outputting a plurality of saliency binary maps; Obtaining a vector displacement of a pixel value in each of the target feature maps, defining a deformation value of each of the vector displacements, and using all of the deformation values ​​to train the constructed initial view generation model to obtain a view generation model; Inputting a plurality of the target feature maps into the view generation model respectively, outputting corresponding predicted pixel value colors, comparing each predicted pixel value color with a true pixel value color corresponding to the target feature map to obtain a plurality of ray residual errors, and calculating uncertainty values ​​of pixel values ​​in the plurality of target feature maps according to the plurality of ray residual errors and the Jacobian matrix; According to the saliency binary map and the uncertainty value of each of the target feature maps, a final uncertainty value of the pixel value in each of the target feature maps is calculated.

2. The method for quantifying scene uncertainty according to claim 1, characterized in that: The construction of the prediction model and the refinement model, inputting the target image input by the user into the prediction model, outputting a plurality of target feature maps, inputting all the target feature maps into the refinement model for bilinear interpolation processing, and outputting a plurality of saliency binary maps, specifically includes: constructing an encoder and a decoder, adding a bridge stage between the encoder and the decoder to obtain a prediction model, and constructing a residual encoder and a residual decoder, adding a bridge stage between the residual encoder and the residual decoder to obtain a refinement model; Inputting a target image input by a user into the prediction model, sampling the target image using the encoder of the prediction model according to a preset pooling window, and obtaining feature maps corresponding to multiple stages in the encoder and pixel values ​​in each feature map; Splicing multiple feature maps of the current stage with multiple feature maps of the previous stage, and using the decoder of the prediction model to perform multi-channel output on the spliced ​​content of different stages to obtain multiple target feature maps of the final stage; All the target feature maps are input into the refinement model, all the target feature maps are downsampled according to the pooling window using the residual encoder of the refinement model, and all the target feature maps are upsampled by bilinear interpolation using the residual decoder, and a saliency binary map of each target feature map is output.

3. The method for quantifying scene uncertainty according to claim 2, characterized in that: The constructing of the encoder and the decoder, adding a bridge stage between the encoder and the decoder to obtain a prediction model, and constructing a residual encoder and a residual decoder, adding a bridge stage between the residual encoder and the residual decoder to obtain a refinement model, specifically includes: In the first preset number of stages of the encoder, a plurality of convolutional layers are respectively constructed, and in the second preset number of stages of the encoder, a plurality of basic residual blocks are respectively constructed; Constructing a plurality of convolutional layers, batch normalization layers, and linear rectification functions in all stages of the decoder, respectively, and connecting the encoder and the decoder using a bridge stage to obtain a prediction model; In each stage of the residual encoder and the residual decoder, a convolution layer, a batch normalization layer and a linear rectification function are respectively constructed, and the residual encoder and the residual decoder are connected using a bridge stage to obtain a refined model.

4. The method for quantifying scene uncertainty according to claim 1, characterized in that: The obtaining of the vector displacement of the pixel value in each of the target feature maps, defining the deformation value of each of the vector displacements, and using all the deformation values ​​to train the constructed initial view generation model to obtain the view generation model specifically includes: Taking the maximum eigenvalue in each of the target feature maps as a pixel value, introducing each of the pixel values ​​into a deformation field to be stored in a mesh vertex, and taking the vector displacement of each mesh vertex as a vector displacement of a corresponding pixel value; By trilinear interpolation method, the deformation value of each pixel value is defined Among them, θ represents the deformation parameter of the deformation field, x represents the vector displacement, and Trilinear represents the trilinear interpolation method; An initial view generation model is constructed, and all the deformation values ​​are used to re-parameterize the initial view generation model to obtain a view generation model.

5. The method for quantifying scene uncertainty according to claim 4, characterized in that: The step of inputting the plurality of target feature maps into the view generation model respectively, outputting the corresponding predicted pixel value color, comparing each predicted pixel value color with the actual pixel value color corresponding to the target feature map, obtaining a plurality of ray residual errors, and calculating the uncertainty values ​​of the pixel values ​​in the plurality of target feature maps according to the plurality of ray residual errors and the Jacobian matrix specifically includes: Inputting the plurality of target feature maps into the view generation model respectively, and outputting the predicted pixel value color of the pixel value corresponding to each target feature map; Obtain the true pixel value color corresponding to each target feature map, and calculate the consistency between each predicted pixel value color and the true pixel value color corresponding to the target feature map, and obtain the ray residual error ∈ corresponding to each target feature map θ (r): Among them, r represents the light of the camera, Represents the predicted pixel value color, Represents the true pixel value color, n represents the number of training samples; Define the Jacobian matrix J θ (r): Among them, the Jacobian matrix represents the first-order partial derivative of the predicted pixel color with respect to the parameter θ; According to the screened ray residual error and Jacobian matrix, the approximate covariance matrix ∑ of the diagonal entries of each target feature map is calculated: Among them, diag() is used to construct the matrix, R represents the parameters of the view generation model, T represents the transpose, λ represents the regularization parameter, and I represents the training image; According to each of the diagonal entries of the approximate covariance matrix, the mesh vertices corresponding to each of the ray residual errors are screened to obtain a screened final target feature map; The uncertainty value of the pixel value in each of the final target feature maps is obtained by a trilinear interpolation method.

6. The method for quantifying scene uncertainty according to claim 5, characterized in that: The step of obtaining the uncertainty value of the pixel value in each of the final target feature maps by the trilinear interpolation method specifically includes: Obtaining the variance vector of the diagonal entry approximation covariance matrix of each of the final target feature maps; The uncertainty field U(x) is defined by the trilinear interpolation method: U(x)=Trilinear(x,σ); Where σ represents the variance vector; According to the variance vector of each of the final target feature maps, an uncertainty value of a pixel value in each of the final target feature maps is determined in the uncertainty field.

7. The method for quantifying scene uncertainty according to claim 6, characterized in that: Calculating a final uncertainty value of a pixel value in each target feature map according to the saliency binary map and the uncertainty value of each target feature map specifically includes: Calculating the significance value of each of the final target feature maps according to the significance binary maps corresponding to the pixel values ​​of all the final target feature maps; According to each of the saliency values ​​and the uncertainty value corresponding to each of the final target feature maps, a final uncertainty value of the pixel value in each of the final target feature maps is calculated.

8. A system for quantifying scene uncertainty, characterized in that: The scene uncertainty quantification system includes: A binary image construction module is used to construct a prediction model and a refinement model, input a target image input by a user into the prediction model, output a plurality of target feature maps, input all the target feature maps into the refinement model for bilinear interpolation processing, and output a plurality of saliency binary maps; A view generation model construction module, used to obtain a vector displacement of a pixel value in each of the target feature maps, define a deformation value of each of the vector displacements, and use all the deformation values ​​to train the constructed initial view generation model to obtain a view generation model; an uncertainty value calculation module, used to input the plurality of target feature maps into the view generation model respectively, output the corresponding predicted pixel value color, compare each predicted pixel value color with the real pixel value color corresponding to the target feature map, obtain a plurality of ray residual errors, and calculate the uncertainty values ​​of the pixel values ​​in the plurality of target feature maps according to the plurality of ray residual errors and the Jacobian matrix; The final uncertainty value calculation module is used to calculate the final uncertainty value of the pixel value in each target feature map according to the saliency binary map and the uncertainty value of each target feature map.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a scene uncertainty quantification program stored in the memory and executable on the processor. When the scene uncertainty quantification program is executed by the processor, the steps of the scene uncertainty quantification method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a scene uncertainty quantification program, and when the scene uncertainty quantification program is executed by a processor, the steps of the scene uncertainty quantification method according to any one of claims 1 to 7 are implemented.