Language radiation field modeling optimization method based on semantic uncertainty

By introducing a semantic uncertainty prediction network into the language radiation field model, the reconstruction quality reduction and artifact problems caused by inconsistency in perspective are solved, the query similarity score is improved, and the higher quality three-dimensional reconstruction effect is achieved.

CN120337694APending Publication Date: 2025-07-18BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510174625.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing language radiation field models during training result in the reduction of reconstruction quality due to inconsistent image viewing angles, especially in areas with sharp changes in semantics, and the query similarity score is not high.

Method used

A semantic uncertainty prediction network is introduced, semantic uncertainty on the light is estimated through a multi-layer perceptron network, and the loss function is adjusted during the training and inference stages to reduce the impact of the region of sharply changing semantics.

Benefits of technology

It effectively reduces the artifacts during the reconstruction process, improves the similarity score of the query results, and improves the reconstruction quality of the model in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337694A_ABST
    Figure CN120337694A_ABST
Patent Text Reader

Abstract

The invention relates to a language radiation field modeling optimization method based on semantic uncertainty. A lightweight prediction model is provided for the situation that a language radiation field of semantic model distillation under complex changes of a reconstruction scene can change sharply. At present, research on a language radiation field rarely pays attention to reconstruction quality reduction and distribution of query similarity scores caused by changes of a semantic level, and generation of artifacts can be slowed down and similarity scores can be improved while basic query and reconstruction functions are guaranteed. According to the method, a semantic uncertainty pre-estimation model is designed, and during model training and reasoning, the pre-estimated uncertainty is applied, so that the quality is improved; a mixed loss function applying uncertainty is also designed, and uncertain semantics are embedded into training and reasoning. Experiments prove that artifacts are reduced on four datasets of NeRFStudio, the similarity score is improved, and the average level is superior to that of an existing language radiation field query method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision, computer graphics, and natural language processing, and involves an optimization method for language radiance field modeling based on semantic uncertainty ( Figure 1 ). Specifically, by adding an estimation of semantic uncertainty during the training of the language radiance field, applying this value to improve the rendering quality of the reconstruction, and increasing the confidence during inference. Ultimately, while ensuring the reconstruction quality and query accuracy, artifacts during reconstruction are significantly reduced and the similarity score of the query results is improved, making it more suitable for actual application scenarios. Background Art

[0002] 3D reconstruction is a technology that uses computer vision techniques to recover the 3D structure from multiple images of the same scene, and it is a subtopic under the cross - field of computer vision and computer graphics. Due to its application value in virtual reality, cultural relic protection, and the gaming field, it has become a research hotspot in recent years. In 1996, the Structure - from - Motion method was proposed, which first reconstructed the 3D structure by matching feature points between images. Until 2020, 3D reconstruction mainly relied on traditional image - processing methods to match feature points in images to recover the structure. However, since these feature - matching methods were all artificially constructed, it was difficult to capture high - dimensional and high - frequency features, and the reconstruction quality was difficult to reach photo - realistic level.

[0003] Since 2020, deep - learning techniques have shone in the field of 3D reconstruction. With the improvement of computer hardware computing power, the number of model parameters has become larger and the calculations have become more complex. Related scholars have started to study implicit scene representation methods, gradually bringing the reconstruction quality close to the photo - realistic level. Neural Radiance Fields (NeRF) mainly estimates the density and color values on the ray through a neural network, and then uses traditional volume rendering to complete the rendering. So far, this technology has gradually been extended to using language radiance fields to express semantic information on the ray to achieve multi - modal applications of text and images. The present invention mainly focuses on modeling the language radiance field combined with semantic uncertainty to improve its reconstruction quality and query quality. The background introduction related to the language radiance field is as follows.

[0004] (1) 3D Reconstruction Method Based on Neural Radiance Field

[0005] In 2020, Google proposed Neural Radiance Fields. Different from traditional methods that use complex artificial features and mathematical formulas to match feature points between views, Neural Radiance Fields directly estimate the density and color information of each point on the ray through a neural network. That is, given the point coordinate triple (x, y, z) and the direction polar angle and azimuth angle (θ, φ), the neural network estimates the voxel density σ and color c of the point in the three-dimensional scene represented by this five-tuple, and the formal expression is shown in Equation (1) (Reference Document 1: Mildenhall B, Srinivasan P P, Tancik M, et al. Nerf: Representing scenes as neural radiance fields for view synthesis[J]. Communications of the ACM, 2021, 65(1): 99-106.)

[0006] σ, c = F(x, y, z, θ, φ) (1)

[0007] When the color and density information of each point on the ray are known, voxel rendering can be performed. Voxel rendering was proposed by Drebin et al. and is shown in Equations (2) and (3). (Reference Document 2: Drebin R A, Carpenter L, Hanrahan P. Volume rendering[J]. ACM Siggraph Computer Graphics, 1988, 22(4): 65-74.)

[0008]

[0009] Among them, tn and tf respectively represent the starting point and ending point parameters of the ray, σ t represents the density of the ray at time t, and c t represents the color of the ray at time t. At the same time, according to the absorption rate of color by the material in physics, voxel rendering calculates the transmittance of the ray through T(t) based on the density of each point on the current ray. T(t) can also be understood as the proportion of the ray absorbed at time t. The final color value is calculated by integrating the product of the transmittance, density, and color.

[0010] To apply the voxel rendering algorithm to a neural network model in a discrete state, Andrea et al. modified the above voxel rendering formula into a discrete differentiable rendering equation, as shown in Equation (4). (Reference Document 3: D Tagliasacchi A, Mildenhall B. Volume rendering digest (for nerf) [J]. arXiv preprint arXiv:2209.02417, 2022.)

[0011]

[0012] In Equation (4), the continuous integral equation is decomposed into a discrete summation, where the three-term product corresponds to the three terms in Equation (2), and N represents the N sampling points on the ray for upsampling and rendering. C(t N+1 ) can be recursively obtained from the previous point n, achieving discrete differentiable rendering.

[0013] The content in the previous part of this section encompasses the core rendering formula of the neural radiance field. To better complete the reconstruction, the model also includes two optimizations. One optimization is the dual network to accelerate training, which uses two NeRF networks with similar structures for training simultaneously, and makes the density information output by the first network serve as the basis for importance sampling of the ray at the same time, while the second network is responsible for estimating the density and color information; the other optimization is position encoding. To resist the learning bias of the neural network for low-frequency features, the model uses the encoding method in Equation (5) to perform high-dimensional expansion on the coordinate p, encoding the position p into 2L trigonometric functions, expanding to L different frequencies, enabling the model to capture subtle changes in position during training.

[0014] γ(p) = (sin(2 0 πp), cos(2 0 πp), …, sin(2 L-1 πp), cos(2 L-1 πp) (5)

[0015] Regarding the dataset, any method can be used to obtain image data with camera positions and orientations. Finally, for each physically acquired or estimated view, the overall objective function is optimized, as shown in Equation (6), where C + and C f respectively represent the color values output by the dual network, and r refers to the ray The optimization objective is the mean squared error between the real image color C(r) and the outputs of the dual network and .

[0016]

[0017] After that, numerous studies have been based on NeRF for modification to adapt to different vertical scenarios and optimize performance. Among them, aiming at non-static objects and environmental changes that may exist in the original images, Radwan et al. proposed an outdoor-available NeRF without constraints (Reference Document 4: Martin-Brualla R, Radwan N, Sajjadi M S M, et al. Nerf in the wild: Neural radiance fields for unconstrained photo collections[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2021: 7210-7219.). The biggest optimization is to change the reconstructed image into S stati+ and S dynami+ , estimating the static and dynamic parts through two networks. During training, S stati+ and S dynami+ are put together for modeling to calculate the overall loss; during prediction, only S stati+ is used for voxel rendering, cleverly avoiding the interference of S dynami+ to the image. Its architecture is as shown in Figure 2 . The overall model will also estimate uncertainty through an additional network. This is also the main basis of the present invention. The present invention assumes that from a semantic perspective, there are also some positions in the scene where there are obvious semantic sharp changes, resulting in a decrease in the overall reconstruction quality. Therefore, the present invention proposes a semantic uncertainty estimation network to mitigate the impact of the semantic high-frequency change region on the performance of the overall model.

[0018] (2) Text-image contrast pre-training

[0019] The similarity between cross-modal contrast images and texts in multi-modal has always been a hot and difficult topic under the cross-disciplinary subject of natural language processing and computer vision. Before the text-image contrast pre-training model (CLIP, Contrastive Language-Image Pre-training) was proposed by OpenAI (Reference Document 4: Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C] / / International conference on machine learning. PMLR, 2021: 8748-8763.), the applications for large-scale texts and images mainly included "text-to-text" GPT-3 (Generative Pre-trained Transformer 3, a large-scale pre-trained attention model) and "image-to-text" supervised training tasks centered around the ImageNet dataset. However, the two tasks were relatively independent, belonging to different modalities and unable to be used together. The CLIP model was the first to model texts and images in a shared space. As Figure 3 shown, the text is input into a text encoder centered around the attention mechanism, and the image is input into an image encoder centered around the convolutional neural network. During training, the CLIP model adopted a contrastive training method. By selecting multiple groups of text-image samples as positive samples and all other combinations as negative samples in a simple training strategy, it could well obtain high-cosine-similarity expressions of similar samples in the shared space.

[0020] The overall optimization objective function of CLIP is shown in Equation (7). Sim i,j represents the cosine similarity between the i-th and j-th embeddings output by the encoder, and Sim j,i represents the cosine similarity between the j-th and i-th. Due to the cross-modal requirements of CLIP, it is necessary to ensure the consistency of similarity in both the image-to-text and text-to-image directions. Therefore, the optimization objective includes two different directions.

[0021]

[0022] The pre-trained CLIP model can be used as a scoring tool for comparing the similarity of image texts and can be embedded into other models out of the box. It is a fundamental model and pioneering work for many multi-modal models.

[0023] (3) Modeling method of language radiation field

[0024] In 2022, Boyi et al. proposed Lseg (Language-Driven Semantic Segmentation). Lseg adopts the CLIP architecture and became the most advanced approach in the field of image semantic style at that time. It encodes text and images through a unified shared space and performs pixel-by-pixel supervised learning downstream of the shared space. After the neural radiance field method became the most advanced method in multiple metrics in the industry, more and more scholars have tried to embed information into the radiance field, and the language radiance field is the most successful attempt at present. (Reference document 5, Kerr J, Kim C M, Goldberg K, et al. Lerf: Language embedded radiance fields[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023:19729-19739.)

[0025] Volume rendering outputs density σ and color c through a neural network and maps the three-dimensional scene to the coordinates on the two-dimensional plane at a specified viewing angle through integration. In the neural network, the output of the language radiance field does not need to be calculated strictly in accordance with the physical definition as in Equation (2). LeRF obtains the semantic information of the pixel points defined above through Equation (8), where r(t) and s(t) respectively correspond to the two-dimensional plane of the ray and the current semantic interval, and they are weighted by the semantic weight w(t) after being calculated jointly through the semantic encoding function F lang after calculation.

[0026]

[0027] This is the same as the discrete volume rendering proposed by NeRF, and the main core lies in training F lang this neural network. Based on NeRF rendering, for each sampling point of each ray in each image, N images are sampled, and the image S img ∈(S min , S max ), where S min and S max are the preset minimum and maximum values of the image size. In terms of the training strategy, the N images have different sizes, and each image will obtain a shared space embedding produced by a CLIP model and optimize the objective function Equation (9), where λ lang is the weight and φ lang is the final semantic result rendered by Equation (8).

[0028]

[0029] In practical applications, the DINO (self-distillation with no labels) network is also used to calculate DINO features to ensure that the edges of the query are clearer.

[0030] Compared with NeRF that can directly render images, LeRF queries require users to manually input a query, and the query is also embedded into the shared space through CLIP. query , and then the image with the highest similarity score is calculated through Equation (10). Here, Canno refers to the Canonical Phrase, which is a preset hyperparameter. Initially set as the large classifications of objects such as Texture, Things, Stuff, and Texture, it can be understood as the pre-classification of the category of the user's query, thus endowing the query with the ability of abstract / concrete composite understanding.

[0031]

[0032] Since the LeRF network combines semantic embedding and reconstruction tasks, it extends the previous two-dimensional semantic task to three dimensions, making its query accuracy in three-dimensional scenes much higher than that of traditional methods that perform semantic segmentation after rendering images in three-dimensional scenes to two-dimensional planes, thus expanding the applicable scenarios of semantic segmentation. Summary of the Invention

[0033] The present invention aims to overcome the influence on the quality of language radiation field reconstruction caused by inconsistent image perspectives during NeRF training, and corrects the semantic deviation during training and inference through the proposed semantic uncertainty measurement algorithm, thereby further improving the quality of scene reconstruction under the language radiation field.

[0034] The present invention mainly studies the language radiation field from the influence of perspective differences. As Figure 4 shown, for the same object, due to the lack of training images, the rendering results vary greatly at different perspectives, and even include areas not covered in the image ( Figure 4 the blurred part in the rightmost perspective). As can be seen from the previous introduction, the objective function of NeRF can be abbreviated as The optimization granularity of the model is the ray r. The optimization between two rays is relatively independent, and in the LeRF scenario, the overall reconstruction loss calculation changes from L rgb to L rgb +L lang . The sharp change in semantics will dominate the optimization direction of the overall loss, bringing a very large deviation, thereby reducing the ability of the model.

[0035] Based on the above considerations, the present invention proposes an effective semantic uncertainty measurement method to enable the model to learn some regions with sharp semantic changes, and on this basis, a semantic uncertainty network is proposed, which has two different uses. One is to reduce the impact of regions with sharp semantic changes on the overall loss during training; the other is to further increase the probability of positive benefits during inference. This has certain practical value for subsequent research and development.

[0036] Next, the main content of the present invention will be introduced in detail, specifically including the following steps:

[0037] Step 1: Model semantic uncertainty

[0038] After LeRF introduces L lang loss into the overall loss, according to the characteristics of neural network training, it is equivalent to expanding the gradient to L′ rgb +L′ lang . When L′ lang oscillates within a range and is difficult to converge, then L′ rgb will also increase the bias accordingly. The present invention defines semantic uncertainty μ lang as the semantic stability of the light ray upsampling points. When the uncertainty of S img with multiple sizes on the light ray changes greatly, its impact on the overall Loss is ultimately reduced.

[0039] Uncertainty modeling is divided into two types. One is from the perspective of loss and is defined from the posterior perspective. The present invention reduces L lang to F(L lang , μ lang ), and the final overall loss is L rgb +F(L lang , μ lang ) to reduce the impact of high-uncertainty regions on the loss; the second is from the output of CLIP and is defined from the prior perspective. Assuming that N embeddings are obtained from N S img of the CLIP model Obtain the prior uncertainty measurement as shown in equations (11) and (12).

[0040]

[0041] The advantage of this method is to introduce the concept of uncertainty and involve it in the calculation of the model loss function to offset the loss bias in the semantic oscillation region.

[0042] Step 2: Design a semantic uncertainty renderer

[0043] To date, all methods of using radiation fields require rendering with light and multiple sampling points thereon, including but not limited to rendering methods such as mean, accumulation, maximum and minimum values, and regularization values. After trying various renderers, the formula that is most stable in performance and most consistent with the LeRF logic (Equation (8)) is shown in (13).

[0044]

[0045] To prevent the uncertainty from being too large or too small, in actual applications, it is also necessary to perform clipping, limit its threshold to [0.1, 1.0], which is also more in line with the physical definition.

[0046] Step 3: Design the uncertainty network

[0047] Under the concepts of the first two steps, the third step of the present invention designs a specific semantic uncertainty estimation network. It is divided into the following three aspects:

[0048] 2.1 Design of the network structure

[0049] The schematic diagram of the network structure of the present invention is as shown in Figure 5 The main objective of this network is to offset the reconstruction bias introduced by LeRF, that is, to determine which regions have extremely large semantic changes by estimating uncertainty. This network is a multi-layer perceptron network that receives the current view as input. The input dimension is the same as the input dimensions of the CLIP and DINO components in LeRF, and produces a single output result μ.

[0050] This network is based on the posterior modeling method of Step 1. K rays (by default, K is 4096) are input in each batch, and P sampling points (by default, P is 25) are generated using importance sampling on each ray. They are fed into the uncertainty estimation model in parallel to produce K*P estimated values. The estimated values will be stored in a temporary dictionary together with the color and density outputs of NeRF and the original semantic outputs of LeRF, parallelizing the entire network training process through the temporary dictionary. See Figure 5 .

[0051] The parameters of the model specifically use a multi-layer perceptron (MLP) with an output dimension of 1 and an input dimension of ∑|φ lang |, with a total of 3 layers and 32 parameters, activated by the ReLu function and output by the Sigmoid function. In the experiment, the present invention hopes to ensure both the effect and performance, so a single-dimensional output is adopted, and each network has only ∑|φ lang |*32*3 parameters.

[0052] 2.2 Network training strategy

[0053] In the previous section, the structure of the model was mainly introduced. In the training strategy, the most important thing is the trade-off between the loss function, dimension, and accuracy to achieve the optimal solution of performance and performance.

[0054] To train the semantic uncertainty model, the present invention applies the estimated uncertainty value to the overall loss calculation function.

[0055] Semantic loss

[0056] Semantic loss refers to whether the features learned by the semantic radiation field in the scene reconstructed by the present invention are close to those learned by the CLIP-supervised labels after integration. This is a result-based distillation that distills the CLIP embedding into the LeRF network. Here, let φ learn be the image semantics rendered by the present invention using Equation (13), and φ +lip be the label semantics generated by estimating the image through the CLIP model. λ1 is the weight hyperparameter of the semantic loss. As shown in Equation (14), it is the Huber loss function used to balance between the mean square error and the mean error and smooth the overall gradient.

[0057]

[0058] Unlabeled self-distillation loss

[0059] The CLIP model contains rich semantic information. However, it is a general model and not a model for semantic segmentation. Whether the input is blurred or has clear boundaries, CLIP can relatively accurately give its semantic embedding in the shared space. However, it does not itself contain the determination of object boundaries. Therefore, it is also necessary to introduce the DINO network to supervise the LeRF model to correctly learn the edges of objects, where D represents the features output by the DINO model.

[0060]

[0061] Uncertainty loss

[0062] The present invention innovates the disadvantages introduced by the previous fixed distillation from CLIP by introducing uncertainty loss in the model training strategy. The specific objective function is Equation (16), where λ2, λ lang , λ dino are three configurable hyperparameters used to adjust the numerical scaling of the overall and local uncertainty losses.

[0063]

[0064] Uncertainty is a term used to reduce the overall loss. To prevent the model from being lazy and choosing to explain the scene with uncertainty, it is also necessary to introduce the regularization term Equation (17) through the preset hyperparameter λ regand logarithmic functions to ensure that the prediction is within a reasonable threshold.

[0065]

[0066] Reconstruction loss

[0067] Finally, there is also a loss L of NeRF itself as shown in Equation (6), rgb which will not be elaborated here.

[0068] In summary, the ultimate optimization goal of the present invention is Equation (18).

[0069]

[0070] The main contribution of the present invention is to add a network and loss related to uncertainty, and provide a feasible training scheme to model the sharp changes in semantics, effectively reducing the artifact phenomenon and improving the numerical similarity of queries.

[0071] In the overall training, in order to complete the training task faster, parallelization, quantization and other schemes are also applied, which have been explained in the above design part and have strong engineering implementation value.

[0072] 2.3 Network application strategy

[0073] Previously, it was mainly elaborated that the present invention balances the overall loss through uncertainty, and the inference part was not mentioned. The LeRF model uses Equation (10) for inference. In the present invention, by adding uncertainty to the inference formula, the value range of the similarity score can be significantly improved. The higher the similarity score, the more the model can distinguish positive and negative examples. It provides more room for future research and implementation.

[0074] The present invention scales the uncertainty of the query language through Equation (19). In the calculation logic of the query positive samples of LeRF, for all words (Phrases), the difference is calculated once for each word, and the S i,j represents the similarity between the i-th word and the j-th word.

[0075]

[0076] When uncertain semantics are not introduced, the formula will calculate the similarity of all words as a positive sample once with all other words, and finally find a decisive word that has the farthest vector space distance from other words.

[0077] In summary, the present invention mainly models uncertain semantics through the assumption that semantics will change sharply, and during training, the posterior uncertain semantics are used to enable the model to find which points have a greater impact on the model loss. The uncertain semantics play a role in both the training and inference stages, effectively reducing the reconstruction artifacts and the quality of inference queries. The experimental results prove that these optimizations enhance the basic capabilities of the model without reducing the quality of model reconstruction, and have practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 It is the overall structure diagram of the uncertainty network proposed by the present invention, which is introduced in detail in step three.

[0079] Figure 2 It is the structure diagram of the NeRF model available outdoors without constraints.

[0080] Figure 3 It is the diagram of the CLIP model training method.

[0081] Figure 4 It is the schematic diagram of the area where semantics change sharply, and the three figures represent three perspectives respectively.

[0082] Figure 5 It is the specific design scheme of the network structure.

[0083] Figure 6 It is the experimental diagram for eliminating artifact areas.

[0084] Figure 7 It is the visualization diagram of the uncertainty in the artifact area.

[0085] Figure 8 It is the visualization diagram of the uncertainty in the normal area.

[0086] Figure 9 It is the training data related to railing shadows in the Egypt dataset.

[0087] Figure 10 It is the comparison diagram of the railing shadow rendering results. The left figure is the method of the present invention, and the right figure is the LeRF method.

[0088] Figure 11 It is the comparison diagram of the query result scores (Egypt dataset).

[0089] Figure 12 It is the comparison diagram of the query result scores (Poster dataset). DETAILED DESCRIPTION OF THE EMBODIMENTS

[0090] Next, the technical solutions, experimental methods, and test results of the present invention will be further described in detail in combination with the drawings and specific experimental embodiments.

[0091] The present invention is an interdisciplinary topic in graphics, computer vision, and natural language processing, and proposes an optimization method based on semantic uncertainty modeling, which mainly includes the following steps.

[0092] Step 1: Construct a dataset. First, obtain the original photo information, take as many photos as possible around the scene to be reconstructed, and then measure the pose of the camera when each photo is taken by any means (such as COLMAP).

[0093] Step 2: Construct a neural network to implement the corresponding loss function. Download the weights of the CLIP model and the DINO model. You can observe the training situation by pre-training LeRF and adjust the parameters of uncertainty modeling by combining manual or automatic means.

[0094] Step 3: Conduct tests based on the training results. Write some words to be queried, and subjectively judge the quality of the effect according to the visualized radiance field rendering map; objectively judge the performance according to the offline evaluation metrics.

[0095] The experimental situation and conclusions of the present invention are specifically described below.

[0096] All experiments were conducted on an RTX4090 graphics card, and exhaustive tests were carried out on four different datasets.

[0097] (1) NeRF reconstruction quality metrics

[0098] Regarding 3D scene reconstruction, there are usually three major metrics, PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), and LPIPS (Learned Perceptual Image Patch Similarity), which correspond to equations (20) - (22) respectively. PSNR is the quotient of the maximum value of the image pixels and the mean squared error, measuring the absolute value, and the larger the better; in SSIM, μ and σ represent the mean and standard deviation respectively, and σ xy represents the covariance, measuring the pixel statistical characteristics of the image, and the larger the better; in LPIPS, the high-order feature differences of each pixel of the image of H*W pixels on the neural network are mainly perceived through the neural network, and the smaller the better.

[0099]

[0100] The present invention compared four datasets through experiments, namely the relatively simple indoor scene Poster, the empty outdoor scene Person, the complex outdoor scene Egypt, and the complex indoor scene Kitchen (the above datasets were all obtained from the official website of NeRFStudio). The hyperparameters selected for this experiment are: the semantic uncertainty weight is set to -0.06, the CLIP loss weight is set to 0.01, and the DINO loss weight is set to 0.1.

[0101] Table 1 Evaluation of rendering indicators of three methods on four datasets

[0102]

[0103] NeRF-related indicators are shown in Table 1. NeRFacto is a baseline version officially provided by NeRFStudio that integrates various optimizations of NeRF. It can be considered as the best NeRF in the past two years. The present invention is basically the same as the previous one or slightly improved in terms of NeRF-related indicators. On the complex outdoor dataset Egypt, it greatly exceeds the LeRF method. Since image quality is a relatively subjective attribute, and the current existing indicators can only evaluate a single feature in the image, they are not very comprehensive, and specific details need to be compared one by one.

[0104] (2) Artifact detail optimization

[0105] Artifacts in this field usually refer to the appearance of unrealistic images in the picture due to various reasons such as unreasonable reconstruction objective function and data quality, usually manifested as a shadow in an area. In the LeRF method, since the semantic loss distilled from CLIP is added to the overall loss, the optimization direction of the overall loss is affected by the high-frequency changes in semantics, and artifacts are generated under complex lighting conditions.

[0106] The most direct contribution of the present invention is to reduce artifacts after reconstruction. By rendering LeRF and the video rendering results of the present invention rotating around the scene, the uncertainty semantics proposed by the present invention significantly reduces the appearance of artifacts under two perspectives after comparison.

[0107] like Figure 6 As shown in the figure, the red frame area is the artifact area. It can be seen that the language radiation field optimized by the present invention is smoother in the corresponding area. On both sides of the figure, the left is the original LeRF reconstruction result, and the right is the result of the present invention. The dark part on the left side appears near the wall, which is not intuitive, while the artifact on the right side is greatly alleviated.

[0108] In order to theoretically prove that the remodeled area and the artifact area overlap, the present invention uses formula (13) to render the uncertainty layer, and the darker the color, the higher the uncertainty. The uncertainty layer rendering result is shown in Figure 7 As shown, it can be seen that the uncertainty layer fits the artifacts relatively closely, and the present invention plays a key role in correctly estimating the areas with quality defects in these reconstructions.

[0109] At the same time, in other non-obvious artifact areas such as Figure 8As shown, the left side is the visualization of the uncertainty region, where the darker the color, the higher the estimated uncertainty value; the right side is the reconstruction result region. The present invention also has a reasonable modeling for uncertainty from other perspectives and has a good estimation effect even in the case of non-artifacts. These regions with high uncertainty are mainly areas near the edges of some objects that do not have actual semantics.

[0110] For some geometric structures, an example will be given here using the shadow of a railing in the Egypt dataset. All the training data of the railing shadow is as Figure 9 shown. In this scenario, the training data for the shadow at the railing is extremely scarce, and the 4 pictures are all taken from the same orientation. The neural radiance field simply cannot be fully trained. Especially after introducing semantic estimation, it becomes even more difficult to optimize the overall objective function of the model. The semantic uncertainty estimation method of the present invention can reduce the impact of sparse data on the overall loss during training. As Figure 10 shown, where the left figure is the rendering result of the present invention, and the right figure is the rendering result of LeRF. The method of the present invention has a relatively obvious improvement in the rendering of the geometric structure of the shadow compared to LeRF. To sum up, for a scene like Egypt that is difficult to fully train, including close views, distant views, and obstacles, the present invention has a significant improvement in rendering metrics. For other scenes with complete training data and a relatively simple overall situation, the main contribution of the present invention is to reduce artifacts.

[0111] (3) Improvement in the similarity score of query results

[0112] As Figure 11 and Figure 12 shown respectively, in this experiment, the shades of colors in the same scene and from the same perspective were compared. Through the same Query for querying, the modified version better balanced the distribution of positive samples, reduced the impact of uncertainty on the classification function (Softmax) in the query statement, and successfully increased the threshold of its similarity score.

[0113] This threshold has no direct relation to whether an object can be detected, but the higher the score, the higher the credibility. By giving a more accurate score to the downstream, it can promote the development of downstream applications. For example, more confident positive samples can be found through the threshold.

[0114] (4) Performance metrics

[0115] The present invention evaluated the rendering performance and training performance respectively.

[0116] Table 2 Performance evaluation results of LeRF and the method proposed in the present invention

[0117]

[0118] In this experiment, each model ran for 5000 training steps and 200 warm-up steps, for a total of 5200 epochs. As shown in Table 2, the addition of the uncertain semantics module in the present invention only slightly increases the training time consumption and memory usage compared with LeRF, and does not significantly increase the training overhead, having practical application value.

[0119] In summary, the present invention proposes an optimization method for language radiation field modeling based on semantic uncertainty. Through a lightweight module, it uses the same input as the model for predicting semantic parts, outputs a single uncertain semantic value, and automatically finds the uncertain semantic values at each position during the neural network gradient optimization process, overcoming the problem of artifacts generated in the region of sharp semantic jitter. Most researchers in the current world are conducting in-depth research on influencing factors such as query accuracy and reconstruction of transient objects. The present invention makes an attempt on the high-frequency changes of semantics within the region, which can provide some references for the future development of language radiation fields.

Claims

1. An optimization method for language radiation field modeling based on semantic uncertainty, characterized in that: Through a lightweight semantic uncertainty estimation model, estimate the semantic uncertainty of each sampling point on the ray. By integrating the semantic uncertainty as a term of the loss into the overall loss for training, the uncertainty semantic value after training will be used as a scaling factor for the semantic loss during training and as a balancing factor for optimizing the semantic similarity score during inference; S1. Design of semantic uncertainty estimation network structure: The semantic uncertainty estimation network is an additional lightweight network embedded in the language radiation field. It shares the same input weights as the training of the language radiation field, and the output is a numerical value, which represents the degree of semantic uncertainty of the current input; First, the network selects N sampling points in a batch, where N = P * K * S. P represents P rays in a batch. Each ray selects K different sampling points according to the overall training needs, and S different image sizes are selected for each sampling point. Then, the N different sampled images are input into the image encoder to obtain N different image embeddings; Use a neural network to fit the N different embeddings into the voxel rendering formula of the three-dimensional scene. During training, the uncertainty value estimated by the semantic uncertainty estimation module is applied as a scaling factor to the loss function calculation module. The loss scaled by this value is added to the losses required for other reconstructions, and the iteration is carried out continuously until the model weights converge; S2. Semantic query strategy based on uncertainty: During the inference process of the language radiation field, through the user input query, after being uniformly calculated by the text encoder to obtain the embedding in the image-text shared space, it is compared with the embedding of the image during the training of the language radiation field. The image with the highest score is selected as the output similarity score through the classification function Softmax; the uncertainty score estimated by the uncertainty estimation network can be used as a distribution balance tool on the user interaction layer to limit the similarity score of the semantic region with high uncertainty and maintain the score of the semantic region with high certainty; S3. Network joint training strategy: In the specific training strategy, full consideration is given to parallel computing. The N images are simultaneously fed into the training network of the language radiation field and the semantic uncertainty estimation network, which mainly includes the following losses; Semantic loss, used as the output φ of the image-text shared space encoding model after pre-training completion when reconstructing the scene lang and φ after voxel rendering of the network language radiation field operator learn The difference, where λ1 is a hyperparameter controlling the overall semantic loss; Without a labeled self-distillation loss, the image-text shared space encoding model does not focus on calculating the similarity between images and text images. Therefore, the DINO model, which is good at dealing with boundaries, is introduced, and the output is labeled as D lang , and the distilled language radiation field model learns the boundary information D of the image learn , where λ2 controls the overall loss value; The semantic uncertainty loss. Through the trained semantic uncertainty estimation module, its value can be used as the loss to balance the whole, where λ2, λ lang and λ dino are used to balance the individual and the whole loss respectively, and eps is a decimal number to prevent overflow; Semantic uncertainty regularization term. To prevent the abuse of the uncertainty interpretation scenario, in practice, a regularization term for it also needs to be introduced to balance the use and overuse of uncertainty, λ reg is a hyperparameter for controlling the overall loss; The final loss L is the sum of all losses, and the overall optimal value is found through joint backpropagation of gradients. Note that when calculating the final loss, it is default that all calculations are based on N as the total number, and the training of multiple rays, multiple sampling points, and multiple perspectives is completed in parallel; 2. A method for optimizing the modeling of a language radiation field based on semantic uncertainty as described in claim 1, proposes a solution method based on the optimization direction of the loss function in the high-frequency change region of semantics, characterized in that: Scale the original calculation process of the language embedding loss with the same weight appropriately according to the estimated semantic uncertainty, and also use this value to guide the discrimination of positive and negative examples during the final inference. The implementation steps are as follows: For a group of images that have not been reconstructed, first estimate or accurately obtain the position and angle of the camera in the three-dimensional space when each picture was taken by any method; Secondly, when reconstructing the neural radiance field and semantic radiance field for this set of images, calculate the reconstruction loss, semantic loss, unsupervised self-distillation loss, and semantic uncertainty loss, and update each loss while the neural network continuously seeks the optimal overall loss. The uncertainty value is used to affect the loss related to semantics, slowing down the side effect of introducing semantic loss on the overall loss; Finally, when the model training is completed, it can interact with the user. After encoding the query text input by the user into the shared embedding space of the image, execute the same process as during training to calculate semantic uncertainty and similarity, and use this value to guide the positive and negative examples of the query text in the image and perform rendering.