A neural radiation field plant rendering method and device integrating a large language model

By combining large language models and neural radiation fields, using pseudo-labels and uncertainty estimation to optimize the training of neural radiation fields, the problems of large data demand and insufficient rendering quality in fruit tree rendering technology are solved, and efficient and robust three-dimensional fruit tree rendering is achieved, supporting the visual analysis of fruit trees.

CN120411325BActive Publication Date: 2025-09-02JIANGXI AGRICULTURAL UNIVERSITY +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510914606.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-02
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

The existing fruit tree rendering technology has problems such as large data demand, high computational complexity, long training time and insufficient rendering quality, which affects the efficient visualization and analysis of fruit tree scenes.

Method used

Combining the large language model and neural radiation field, multi-view plant images are collected through drones, pseudo-labels are generated using large language models, segmentation model fine-tuning is carried out, uncertainty estimation loss and active learning strategies are introduced, and the training process of neural radiation field is optimized to generate high-quality three-dimensional fruit tree rendering images.

Benefits of technology

It realizes the improvement of the quality and efficiency of fruit tree rendering under a limited input budget, enhances the robustness and generalization ability of the model, and quickly generates high-quality three-dimensional fruit tree rendering images, supporting the analysis of key factors such as canopy morphology analysis, growth monitoring and maturity evaluation of fruit trees.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411325B_ABST
    Figure CN120411325B_ABST
Patent Text Reader

Abstract

The present invention discloses a neural radiation field plant rendering method and device that integrates a large language model. The method collects a set of plant images with text prompts, uses the large language model to extract semantic information from the text prompts corresponding to the plant images, and generates pseudo-labels. The pseudo-labels are used to fine-tune the segmentation model based on the geometric correspondence between multi-view plant images and the self-supervised relationship within a single plant image. The position parameters of the plant images are input into the neural network model to predict new image parameters, and then the rendered image is generated using the volume rendering formula. The uncertainty estimation loss is added while the target text prompt is given. The segmentation model generates a two-dimensional segmentation map on each input view to predict the supervision signal for learning and training, and outputs a three-dimensional plant rendered image. The method of the present invention performs excellently in complex scene reconstruction performance with occlusion and small targets, providing a new solution for the efficient reconstruction of complex fruit tree scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional reconstruction of crops, and relates to a neural radiation field plant rendering method and device integrating a large language model. Background Art

[0002] Fruit tree visualization technology is used in various agricultural applications, including crown morphology analysis, growth monitoring, fruit counting and distribution analysis, and maturity assessment. By modeling factors such as crown volume, density, and branch distribution, realistic fruit tree scenes can be generated. The purpose of fruit tree rendering is to visualize fruit trees to help agricultural professionals, farmers, and policymakers better understand and analyze their characteristics, problems, and potential. In the context of digital agriculture, fruit tree visualization allows observation of the impact of environmental conditions on crop growth morphology, analysis of environmental factors affecting crop growth, and process analysis of crop growth. This can improve the crop growing environment, optimize land use, and increase agricultural production and income. This holds significant theoretical and practical value in botany and even agronomy.

[0003] Existing fruit tree rendering technology is implemented using tools such as computer graphics, geographic information systems, and remote sensing data. By using point cloud data collected through lidar or satellite remote sensing data, accurate rendering models can be generated. Unlike traditional 3D reconstruction methods, Neural Radiance Fields (NeRF) does not require large discrete point clouds or sparse depth images as input. Instead, it learns the 3D representation and volume density distribution of the scene from continuous observations from the camera's perspective. Using a deep neural network as a model, it can learn the mapping relationship from the input image to the implicit representation of the scene in an end-to-end manner. However, due to the large amount of training data, high computational complexity, rendered image generation and hyperparameter tuning, and training data limitations, NeRF's training time is long and the rendering quality is insufficient. Summary of the Invention

[0004] To address the limitations of traditional crop rendering methods in the aforementioned technical background, such as large data requirements and poor performance, the present invention provides a neural radiation field plant rendering method that integrates a large language model. This method, combined with deep learning to process real-time plant images, can rapidly render high-quality plant images. This method, for example, can be used for fruit tree rendering, facilitating analysis of key factors such as crown morphology, growth monitoring, fruit counting and distribution analysis, and maturity assessment, thereby helping to uncover patterns in increasing crop yields and incomes.

[0005] The present invention is implemented through the following technical solution: a neural radiation field plant rendering method integrating a large language model, comprising the following steps:

[0006] Step 1: Use a drone to capture multi-view plant images and preprocess them to calculate the position parameters of the plant images. Then, collect a set of plant images with text prompts and use a large language model to extract semantic information from the text prompts corresponding to the plant images to generate pseudo labels.

[0007] Step 2: Use the pseudo labels generated by the large language model to fine-tune the segmentation model based on the geometric correspondence between multi-view plant images and the self-supervised relationship within a single plant image;

[0008] Step 3: Input the position parameters of the plant image into the neural network model of the neural radiation field to predict the new image parameters, and then generate the rendered image using the volume rendering formula;

[0009] Step 4: Use active learning strategies to estimate the uncertainty of the image parameters of the rendered image, select new samples that can bring the most information gain for learning, and expand the training set with newly captured rendered images;

[0010] Step 5: Add uncertainty estimation loss and give the target text prompt at the same time. Use the segmentation model to generate a two-dimensional segmentation map prediction supervision signal on each input view for learning and training, and output a three-dimensional plant rendering image.

[0011] Further preferably, in step 2, the segmentation model is fine-tuned based on the geometric correspondence between the multi-view plant images, and the loss function is as follows:

[0012] ;

[0013] in: is the geometric correspondence supervision loss, is the pseudo label of the i-th sampling point in different magnified views, is the predicted value of the segmentation probability of the i-th sampling point in the zoomed-out view.

[0014] Further preferably, in step 2, within a single plant image, the segmentation of the global image is supervised by cropping the segmentation result of the local area:

[0015] ;

[0016] in: To trim the supervision loss, is the pseudo label of the i-th sampling point in the cropped local area, It is the predicted value of the segmentation probability of the i-th sampling point in the corresponding cropped local area of ​​the global image.

[0017] Further preferably, in step 4, when estimating uncertainty, firstly establish the output branch variance of the additional multi-layer perceptron;

[0018] By calculating the uncertainty reduction of all sampling points in the scene, we can evaluate the effect of the new perspective on the overall cognition of the model, and use the new perspective with more information gain as the Defined as:

[0019] ;

[0020] in, is the prior variance, is the uncertainty of the current model, is the posterior variance, which is the uncertainty after assuming new data, is the jth ray, For light The depth value of the i-th sampling point, is a candidate new dataset, corresponding to pixel light under a new perspective, The number of 3D space points sampled on each ray, i=1 means The information gain is accumulated from the first sampling point on sampling points.

[0021] The prior mean of the radiation color after obtaining a new perspective with more information gain The posterior update of :

[0022] ;

[0023] in, Indicates along the ray The kth sampling point when performing volume rendering The depth value of for The prior mean of the radiation color at , Update the weights for Bayesian, balancing the weights of new observations and prior estimates, is the light from the newly captured perspective image The corresponding real color, is the ray r divided by The contribution of other sampling points to the rendered color, is the kth sampling point Volume rendering weights;

[0024] The loss function for introducing uncertainty estimation is:

[0025] ;

[0026] in, Estimating losses for uncertainty, is the number of rays, is the jth ray The corresponding rendering color, is the jth ray The corresponding real color, represents the uncertainty variance of the radiation color, Represents light Render color after volume rendering The variance of To control the strength parameter of the regularization, Represents light The i-th sampling point The body density.

[0027] Further preferably, in step 5, the volume probability of a given target text prompt is learned, and according to the target text prompt given at the beginning of the three-dimensional reconstruction training, a segmentation model is used on each input view to generate a two-dimensional segmentation map prediction supervision signal under the constraint of the target text prompt, and the binary cross entropy loss is used to optimize the semantic volume probability, and it is compared with the two-dimensional segmentation map on the sampling ray. At the same time, uncertainty estimation loss is added through active learning during training to determine the perspective with more information gain, further optimize the three-dimensional reconstruction training process, and finally generate a three-dimensional fruit tree rendering image.

[0028] It is further preferred that during the training process, the uncertainty estimation loss and the image loss calculated by variance are combined to calculate the total rendering loss and the neural network model of the neural radiation field is back-propagated, so as to optimize the geometric shape and appearance texture of the three-dimensional model, thereby generating a three-dimensional fruit tree rendering image with better rendering effect and quality.

[0029] The present invention also provides a neural radiation field plant rendering device integrated with a large language model, comprising:

[0030] A data acquisition and transmission module, which includes a drone camera device and a wireless local area network. The drone camera device is used to collect images of the plants, and the wireless local area network is used to transmit data to a computer;

[0031] The image data processing module is deployed on a computer and is used to execute the various steps of the aforementioned neural radiation field plant rendering method integrating a large language model.

[0032] The present invention successfully applies large language models and neural radiation fields to the three-dimensional reconstruction task of complex fruit tree scenes, achieving efficient three-dimensional reconstruction of fruit tree scenes. This method fully utilizes the excellent three-dimensional reconstruction performance of neural radiation fields, while combining large language models with uncertainty estimation. Through the large language model technology, the geometric correspondence between text metadata and perspective is used to guide the semantic areas in the image with semantics, thereby improving the reconstruction quality of the model in the semantic area. Uncertainty estimation is introduced to select samples with the highest information gain, thereby reducing the uncertainty of training. Under a limited input budget, the generalization ability of the model is improved, the robustness of the model is enhanced, and high-quality rendered images are quickly generated.

[0033] The innovations and advantages of the present invention are:

[0034] (1) The introduction of a large language model, through the rich prior knowledge and image understanding ability of the large model, establishes the correspondence between pseudo-labels and fruit tree scene views to learn image-level and pixel-level semantics. Through two types of geometric connections, the segmentation model is fine-tuned to improve the robustness of open vocabulary semantic segmentation to adapt to fruit tree scene segmentation and achieve the ability to perform multiple semantic concept positioning in a single view. During 3D reconstruction, a semantically adapted 2D segmentation map is provided for each input view for the specified target text prompt, increasing the semantic understanding of special areas to learn about new volume probabilities and improve the reconstruction quality of the model in semantic areas.

[0035] (2) Uncertainty estimation training optimization method, through active learning to supplement new samples from existing data, evaluate the new input, and select the samples with the highest information gain, thereby reducing the uncertainty of training. Rendering high-frequency textures with a limited input budget ensures robustness and provides an explanation of how NeRF understands the scene.

[0036] (3) While improving the rendering quality, the present invention achieves efficient training of the neural radiation field. During training, a large language model is used to provide supervision signals to obtain a more expressive segmentation model, and uncertainty estimation is then used to supplement new sample knowledge, thereby improving the reconstruction results of the neural radiation field. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0038] Figure 1 This is a flow chart of a neural radiation field plant rendering method integrating a large language model provided by an embodiment of the present invention;

[0039] Figure 2 This is a network diagram of a neural radiation field plant rendering method integrating a large language model provided by an embodiment of the present invention;

[0040] Figure 3 This is a schematic diagram of a neural radiation field plant rendering device that integrates a large language model, provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0042] like Figure 1 and Figure 2 As shown, a neural radiation field plant rendering method integrating a large language model includes the following steps:

[0043] Step 1: Use drone photography to capture multi-view fruit tree images and pre-process them to calculate the position parameters of the fruit tree images. , where x, y, z are three-dimensional coordinates, It is a two-dimensional azimuth perspective; a collection of fruit tree images with text prompts is collected, and a large language model is used to extract semantic information from the text prompts corresponding to the fruit tree images to generate pseudo labels;

[0044] Step 2: Use the pseudo-labels generated by the large language model to fine-tune the segmentation model based on the geometric correspondence between the multi-view fruit tree images and the self-supervised relationship within a single fruit tree image;

[0045] Step 3: Input the position parameters of the fruit tree image into the neural network model of the neural radiation field to predict the new image parameters , RGB is color, is the volume density, and then the rendered image is generated through the volume rendering formula;

[0046] Step 4: Use active learning strategies to estimate the uncertainty of the image parameters of the rendered image, select new samples that can bring the most information gain for learning, and expand the training set with newly captured rendered images;

[0047] Step 5: Add uncertainty estimation loss and give the target text prompt at the same time. Use the segmentation model to generate a two-dimensional segmentation map prediction supervision signal on each input view for learning and training, and output a three-dimensional fruit tree rendering image.

[0048] In step 1 of this embodiment, the drone takes a picture from O 1, O2 and other multi-angle perspectives 1,The d2 isoview line collects fruit tree images in the fruit tree scene, and then uses colmap preprocessing to obtain the camera position parameters for taking the picture , where x, y, z are three-dimensional coordinates, It is a two-dimensional azimuth perspective, specifically 1, O2 etc. corresponds to d 1, d2 isoview , Three-dimensional coordinates.

[0049] In step 1, a collection of fruit tree images with text prompts can be obtained by collecting publicly available fruit tree images and corresponding text prompts from the internet, or by creating a custom collection of fruit tree images and manually adding text prompts to the fruit tree images. These text prompts include information such as the file name, title, hierarchy, and features. This information is described as a string to construct the prompt for the model. The user can enter a descriptive text prompt, such as "What architectural features of the crop are described in the following image? If unspecified, write "unknown"." Beam search decoding is then used to generate pseudo-labels, and these outputs are processed using standard text cleaning techniques. This allows the use of a large language model (LLM) to extract semantic information from the text prompts of the fruit tree images, and feature matching is used to establish correspondence between images, ensuring that the pseudo-labels are consistent in three-dimensional space.

[0050] In step 2, to adapt the segmentation model to fruit tree scene segmentation, the segmentation model CLIP is fine-tuned based on the geometric correspondence between multi-view fruit tree images. Through multi-view geometric consistency (feature matching to establish corresponding point pairs between images), pseudo labels are first used to form weak supervision so that the segmentation model calculates the similarity between image regions and pseudo labels, and then the high-response areas (zoomed-in views) of different images are segmented. The segmented zoomed-in views are then used as supervisory signals for training guidance of global (zoomed-out view) segmentation prediction:

[0051] ;

[0052] in: is the geometric correspondence supervision loss, is the pseudo label of the i-th sampling point in different magnified views, is the segmentation probability prediction value of the i-th sampling point in the reduced view, which enhances the semantic consistency across views.

[0053] In step 2, the self-supervised relationship within a single fruit tree image is used to fine-tune the segmentation model. Within a single fruit tree image, the segmentation results of the local area are cropped to supervise the segmentation of the global image:

[0054] ;

[0055] in: To trim the supervision loss, is the pseudo label of the i-th sampling point in the cropped local area, It is the segmentation probability prediction value of the i-th sampling point in the corresponding cropped local area of ​​the global image, which improves the model to maintain the consistency of local predictions in the global context.

[0056] In step 3 of this embodiment, refer to Figure 2 The neural network model is composed of a multi-layer perceptron, including multiple sets of linear layers (256 dimensions), which enables the neural network model to better fit complex data distribution and functional relationships. A ray is projected for each 3D point in the collected fruit tree image scene. The neural network model samples along each ray and uses the multi-layer perceptron to obtain the volume density of the 3D point based on the coordinates of any point in space. With color : , It is a multi-layer perceptron.

[0057] In step 3 of this embodiment, image generation is performed using a volume rendering formula. The volume rendering formula is used to generate an image based on the color and volume density of each sampling point. The volume rendering process can be implemented by referring to patent document CN117475067B.

[0058] In step 4 of this embodiment, when estimating uncertainty, first establish the additional multi-layer perceptron output branch variance:

[0059] ;

[0060] ;

[0061] Among them, f is the internal implicit feature, represents the variance of the neural radiance field on the target surface rendering pixel distribution, Subscript is the geometric characterization parameter, the subscript is an uncertain characterization parameter, the subscript is the color characterization parameter; Represents the sampling point coordinate position, represents the color mean, and d is the camera field of view direction.

[0062] For a given image with high resolution and width, the number of sampling = height × width is independent of the ray, and the position of the 3D space point is sampled from each ray. By calculating the uncertainty reduction of all sampled points in the scene, the effect of the new perspective on the overall cognition of the model is evaluated. And the new perspective with more information gain is used. Defined as:

[0063] ;

[0064] in, is the prior variance, is the uncertainty of the current model, is the posterior variance, which is the uncertainty after assuming new data, is the jth ray, For light The depth value of the i-th sampling point, is a candidate new dataset, corresponding to pixel light under a new perspective, The number of 3D space points sampled on each ray, i=1 means The information gain is accumulated from the first sampling point on sampling points.

[0065] The prior mean of the radiation color after obtaining a new perspective with more information gain The posterior update of :

[0066] ;

[0067] in, Indicates along the ray The kth sampling point when performing volume rendering The depth value (that is, the distance parameter in the light parameter), for The prior mean of the radiation color at , Update the weights for Bayesian, balancing the weights of new observations and prior estimates, is the light from the newly captured perspective image The corresponding real color, is the ray r divided by The contribution of other sampling points to the rendered color, is the kth sampling point Volume rendering weights, Indicates from Eliminate Get the kth sampling point The independent color influence of the point. Then through the weight Normalization ( The contribution to rendering is proportional to Decision). If There is a large difference from the current forecast and high uncertainty ( ), then a major update If the uncertainty is low ( ), then keep the prior estimate.

[0068] The loss function for introducing uncertainty estimation is:

[0069] ;

[0070] in, Estimating losses for uncertainty, is the number of rays, is the jth ray The corresponding rendering color, is the jth ray The corresponding real color, represents the uncertainty variance of the radiation color, Represents light Render color after volume rendering The variance of (weighted sum of the variances of all sampling points on the ray path), when the variance When the variance is large (high uncertainty), the penalty for the light error is reduced, allowing the model to maintain fuzzy predictions in unobserved areas. When the variance is small (low uncertainty), the color accuracy is constrained. is a variance regularization term that prevents the model from predicting high uncertainty for all regions. To penalize non-zero volume density, the model is prevented from predicting spurious geometry in empty regions. To control the strength parameter of the regularization, Represents light The i-th sampling point The volume density of

[0071] Reference Figure 2 In step 5, the volume probability of a given target text prompt is learned. According to the target text prompt given at the beginning of the 3D reconstruction training, a segmentation model is used on each input view to generate a two-dimensional segmentation map prediction supervision signal under the constraint of the target text prompt. The binary cross entropy loss is used to optimize the semantic volume probability and compare it with the two-dimensional segmentation map on the sampling ray. At the same time, uncertainty estimation loss is added through active learning during training to determine the perspective with more information gain, further optimize the 3D reconstruction training process, and finally generate a 3D fruit tree rendering image.

[0072] During the training process, the uncertainty estimation loss and the image loss calculated by variance are combined to calculate the total rendering loss and the neural network model of the neural radiation field is back-propagated to optimize the geometry and appearance texture of the three-dimensional model, thereby generating a three-dimensional fruit tree rendering image with better rendering effect and quality.

[0073] In the context of the rapid development of 3D technology, neural radiation field stands out. The use of neural network combined with volume rendering to generate plant models is both accurate and convenient. The present invention uses a very good method for 3D scene reconstruction. network, A neural network is used to learn the features of objects in an image and combine it with volume rendering to generate a new image. The method of the present invention (NeRF+large language model+uncertainty estimation) is based on combining NeRF with a large language model, and improves the original NeRF network model by using an uncertainty estimation training strategy, and back-propagation optimization model. The present invention selected three main evaluation indicators such as PSNR (peak signal-to-noise ratio), SSIM (structural similarity index‌), and LPIPS (learnable perceptual image block similarity). The data set images used by the present invention to test various improved networks include 87 training sets, 12 test sets, and 60 validation sets. Various improved networks were uniformly run for 100,000 rounds to obtain evaluation index data. The comparison data of the three evaluation indicators of the ablation experiment of various improved networks after running the data set are shown in Table 1 below. The present invention compares the original Uncertainty estimation is introduced into the network, and a supervision signal is provided by a large language model. In various experiments conducted, the combination of uncertainty estimation and supervision signal provided by a large language model has the best effect. Therefore, the present invention selects a network that combines uncertainty estimation and a large language model.

[0074] Table 1

[0075]

[0076] Table 2

[0077]

[0078] like Figure 3 As shown, an embodiment of the present invention provides a neural radiation field plant rendering device integrated with a large language model, comprising:

[0079] A data acquisition and transmission module, which includes a drone camera device and a wireless local area network. The drone camera device is used to collect images of fruit trees, and the wireless local area network is used to transmit data to a computer;

[0080] The image data processing module is deployed on a computer and is used to execute the various steps of the aforementioned neural radiation field plant rendering method integrating a large language model.

[0081] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A neural radiation field plant rendering method integrating a large language model, characterized by: The following steps are involved: Step 1: Use a drone to capture multi-view plant images and preprocess them to calculate the position parameters of the plant images. Then, collect a set of plant images with text prompts and use a large language model to extract semantic information from the text prompts corresponding to the plant images to generate pseudo labels. Step 2: Use the pseudo labels generated by the large language model to fine-tune the segmentation model based on the geometric correspondence between multi-view plant images and the self-supervised relationship within a single plant image; Step 3: Input the position parameters of the plant image into the neural network model of the neural radiation field to predict the new image parameters, and then generate the rendered image using the volume rendering formula; Step 4: Use active learning strategies to estimate the uncertainty of the image parameters of the rendered image, select new samples that can bring the most information gain for learning, and expand the training set with newly captured rendered images; In step 4, when estimating uncertainty, first establish the output branch variance of the additional multi-layer perceptron; By calculating the uncertainty reduction of all sampling points in the scene, we can evaluate the effect of the new perspective on the overall cognition of the model, and use the new perspective with more information gain as the Defined as: ; in, is the prior variance, is the posterior variance, is the jth ray, For light The depth value of the i-th sampling point, is a candidate new dataset, corresponding to pixel light under a new perspective, The number of 3D space points sampled on each ray, i=1 means The information gain is accumulated from the first sampling point on sampling points; The prior mean of the radiation color after obtaining a new perspective with more information gain The posterior update of : ; in, Indicates along the ray The kth sampling point when performing volume rendering The depth value of for The prior mean of the radiation color at , Update the weights for Bayesian, balancing the weights of new observations and prior estimates, is the light from the newly captured perspective image The corresponding real color, is the ray r divided by The contribution of other sampling points to the rendered color, is the kth sampling point Volume rendering weights; The loss function for introducing uncertainty estimation is: ; in, Estimating losses for uncertainty, is the number of rays, is the jth ray The corresponding rendering color, is the jth ray The corresponding real color, represents the uncertainty variance of the radiation color, Represents light Render color after volume rendering The variance of To control the strength parameter of the regularization, Represents light The i-th sampling point The volume density of Step 5: Add uncertainty estimation loss and give the target text prompt at the same time, learn the volume probability of the given target text prompt, and use the segmentation model to generate a two-dimensional segmentation map prediction supervision signal under the target text prompt constraint on each input view according to the target text prompt given at the beginning of the 3D reconstruction training. Use binary cross entropy loss to optimize the semantic volume probability and compare it with the two-dimensional segmentation map on the sampling ray. While training, add uncertainty estimation loss through active learning to determine the perspective with more information gain, further optimize the 3D reconstruction training process, and finally generate a 3D fruit tree rendering image.

2. The neural radiation field plant rendering method according to claim 1, characterized in that: In step 2, the segmentation model is fine-tuned based on the geometric correspondence between multi-view plant images. The loss function is as follows: ; in: is the geometric correspondence supervision loss, is the pseudo label of the i-th sampling point in different magnified views, is the predicted value of the segmentation probability of the i-th sampling point in the zoomed-out view.

3. The neural radiation field plant rendering method according to claim 1, characterized in that: In step 2, within a single plant image, the segmentation results of the local region are cropped to supervise the segmentation of the global image: ; in: To trim the supervision loss, is the pseudo label of the i-th sampling point in the cropped local area, is the predicted value of the segmentation probability of the i-th sampling point in the corresponding cropped local area of ​​the global image.

4. The neural radiation field plant rendering method according to claim 1, characterized in that: During training, the uncertainty estimation loss and the image loss calculated by variance are combined to calculate the total rendering loss and backpropagate the neural network model of the neural radiation field.

5. A neural radiation field plant rendering device integrating a large language model, characterized by: include: A data acquisition and transmission module, which includes a drone camera device and a wireless local area network. The drone camera device is used to collect images of the plants, and the wireless local area network is used to transmit data to a computer; An image data processing module is deployed on a computer, and is used to execute the various steps of the neural radiation field plant rendering method integrating a large language model as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Field rendering visualization fast rendering method and device based on neural radiation field

    CN117475067B

  • Field rendering visual rapid rendering method and device based on neural radiation field

    CN117475067A

  • Indoor three-dimensional object reconstruction method and apparatus, computer device and storage medium

    WO2024230151A1