Methods, apparatus, equipment and storage media for three-dimensional reconstruction of plants using multi-view fusion

By employing a multi-view fusion method for 3D plant reconstruction, utilizing 3D representation networks and self-supervised learning, and combining neural radiation fields and Markov random field optimization, the problem of low efficiency and accuracy in 3D plant reconstruction in greenhouse crop production is solved, achieving rapid and accurate 3D plant reconstruction.

CN119762686BActive Publication Date: 2026-01-06INTELLIGENT EQUIPMENT RESEARCH CENTER BEIJING ACADEMY OF AGRICULTURE AND FORESTRY SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510273853.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2026-01-06
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Existing methods for 3D reconstruction of plants are difficult to improve efficiency and accuracy in greenhouse crop production, especially in environments with long plant growth cycles, abundant branching, and complex and variable light conditions. Existing technologies are unable to meet the requirements for efficient and high-precision 3D reconstruction.

Method used

A multi-view fusion method for 3D reconstruction of plants is adopted. By training a 3D representation network based on unlabeled multi-view images of plants, combined with self-supervised learning and attention mechanisms, the 3D prior knowledge of plants is obtained. The 3D reconstruction of plants is then achieved through optimization using neural radiation fields and Markov random fields.

Benefits of technology

It achieves high efficiency and high precision in plant 3D reconstruction, and can quickly and accurately generate 3D reconstruction models of plants, thus improving the efficiency and accuracy of plant 3D reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762686B_ABST
    Figure CN119762686B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing, and provides a plant three-dimensional reconstruction method, device and equipment based on multi-view fusion and a storage medium.The method comprises the following steps: obtaining plant three-dimensional prior knowledge based on a three-dimensional representation network; the three-dimensional representation network is obtained by training based on unannotated plant multi-view images; performing feature extraction on multi-view time-series images of a plant to be monitored to obtain multi-view image fusion features; and performing three-dimensional reconstruction on the plant to be monitored based on the plant three-dimensional prior knowledge and the multi-view image fusion features to obtain a three-dimensional reconstruction model of the plant to be monitored.The three-dimensional representation network obtained by training based on unannotated plant multi-view images is used to obtain three-dimensional prior knowledge, and then an attention mechanism is used to fuse multi-view image features, and finally, the plant to be monitored is reconstructed based on the three-dimensional prior knowledge and the multi-view image fusion features, so that an accurate and fast three-dimensional reconstruction model of the plant to be monitored is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and in particular to a method, apparatus, device, and storage medium for three-dimensional reconstruction of plants using multi-view fusion. Background Technology

[0002] Three-dimensional reconstruction of plants is mainly used for plant phenotypic analysis. However, in greenhouse crop production, due to the long growth cycle of plants, extensive branching, and complex and variable light conditions in greenhouse environments, existing three-dimensional reconstruction methods (such as laser scanning and image matching) are difficult to improve the efficiency and accuracy of plant three-dimensional reconstruction. Summary of the Invention

[0003] This invention provides a method, apparatus, device, and storage medium for multi-view fusion of plant 3D reconstruction, which addresses the shortcomings of existing plant 3D reconstruction technologies in improving the efficiency and accuracy of plant 3D reconstruction, thereby achieving an improvement in the efficiency and accuracy of plant 3D reconstruction.

[0004] This invention provides a method for three-dimensional reconstruction of plants using multi-view fusion, comprising the following steps:

[0005] Three-dimensional prior knowledge of plants is obtained based on a three-dimensional representation network; the three-dimensional representation network is trained based on unlabeled multi-view images of plants.

[0006] Feature extraction was performed on multi-view time-series images of the plants to be monitored to obtain multi-view image fusion features;

[0007] Based on the plant's three-dimensional prior knowledge and the multi-view image fusion features, the plant to be monitored is reconstructed in three dimensions to obtain a three-dimensional reconstruction model of the plant to be monitored.

[0008] According to the multi-view fusion method for 3D reconstruction of plants provided by the present invention, the step of obtaining 3D prior knowledge of the plant based on a 3D representation network includes:

[0009] Based on the unlabeled multi-view images of the plants, the global image reconstruction error and image contrast loss are determined.

[0010] Based on the global image reconstruction error and the image contrast loss, a self-supervised learning objective is determined.

[0011] A three-dimensional representation network is trained based on the self-supervised learning objective.

[0012] According to the multi-view fusion method for three-dimensional reconstruction of plants provided by the present invention, the step of extracting features from the multi-view time-series images of the plant to be monitored to obtain multi-view image fusion features includes:

[0013] Feature extraction was performed on multi-view time-series images of the plants to be monitored to obtain the visual features of the images;

[0014] Multi-view feature tensors are obtained by stitching together feature maps from multiple perspectives; the feature maps are determined based on the visual features of the images.

[0015] The multi-view feature tensors are fused to obtain multi-view image fusion features.

[0016] According to the present invention, a method for three-dimensional reconstruction of plants based on multi-view fusion, wherein the three-dimensional reconstruction of the plant to be monitored is performed based on the prior three-dimensional knowledge of the plant and the multi-view image fusion features to obtain a three-dimensional reconstruction model of the plant to be monitored includes:

[0017] The continuous three-dimensional space corresponding to the plant to be monitored is divided to obtain a regular grid composed of multiple voxels;

[0018] Determine the occupying voxels in the regular grid;

[0019] Coordinates are extracted from the occupied voxels to obtain the voxelized three-dimensional information of the plant.

[0020] Based on the plant's three-dimensional prior knowledge, the multi-view image fusion features, and the plant's voxelized three-dimensional information, the plant to be monitored is reconstructed in three dimensions to obtain a three-dimensional reconstruction model of the plant to be monitored.

[0021] According to the present invention, a method for three-dimensional reconstruction of plants based on multi-view fusion is provided, wherein the plant to be monitored is reconstructed in three dimensions based on prior knowledge of the plant's three dimensions and the multi-view image fusion features to obtain a three-dimensional reconstruction model of the plant to be monitored, and the method further includes:

[0022] A temporal feature field is added to the first neural radiation field to obtain a second neural radiation field; the first neural radiation field is used to perform three-dimensional reconstruction of the plant to be monitored; the second neural radiation field is used to perform N-dimensional reconstruction of the plant to be monitored, where N>3; the temporal feature field is determined based on the temporal morphology of the plant to be monitored.

[0023] Determine the changes in the temporal pattern along the time dimension;

[0024] Based on the second neural radiation field and the change information, the plant to be monitored is reconstructed in N dimensions to obtain an N-dimensional reconstruction model of the plant to be monitored.

[0025] According to the present invention, a method for three-dimensional reconstruction of plants based on multi-view fusion is provided, wherein the plant to be monitored is reconstructed in three dimensions based on prior knowledge of the plant's three dimensions and the multi-view image fusion features to obtain a three-dimensional reconstruction model of the plant to be monitored, and the method further includes:

[0026] Acquire motion information of the image acquisition device; the image acquisition device is used to acquire multi-view time-series images of the plant to be monitored.

[0027] Obtain the posture information and real three-dimensional model of the plant to be monitored;

[0028] The image acquisition device is optimized based on the motion information, the posture information, and the real 3D model.

[0029] The present invention also provides a multi-view fusion plant three-dimensional reconstruction device, comprising the following modules:

[0030] A three-dimensional prior knowledge acquisition module for plants is used to obtain three-dimensional prior knowledge of plants based on a three-dimensional representation network; the three-dimensional representation network is trained based on unlabeled multi-view images of plants.

[0031] The multi-view image fusion feature acquisition module is used to extract features from multi-view time-series images of the plant to be monitored, and obtain multi-view image fusion features.

[0032] The plant 3D reconstruction module is used to perform 3D reconstruction of the plant to be monitored based on the plant's 3D prior knowledge and the multi-view image fusion features, so as to obtain the 3D reconstruction model of the plant to be monitored.

[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the multi-view fusion plant three-dimensional reconstruction method as described above.

[0034] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-view fusion method for three-dimensional reconstruction of plants as described above.

[0035] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the multi-view fusion method for three-dimensional reconstruction of plants as described above.

[0036] The present invention provides a method, apparatus, device, and storage medium for multi-view fusion of plant 3D reconstruction. It introduces self-supervised learning into plant 3D reconstruction, obtains 3D prior knowledge by training a 3D representation network through unlabeled multi-view images of the plant, and then fuses the features of the multi-view images with an attention mechanism. Finally, based on the trained 3D prior knowledge and the fused features of the multi-view images, the plant to be monitored is reconstructed in 3D, and the 3D reconstruction model of the plant to be monitored is obtained accurately and quickly. While improving the efficiency of plant 3D reconstruction, it also ensures the accuracy of plant 3D reconstruction. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0038] Figure 1 This is one of the flowcharts of the multi-view fusion plant three-dimensional reconstruction method provided by the present invention.

[0039] Figure 2 This is the second flowchart of the multi-view fusion method for three-dimensional reconstruction of plants provided by the present invention.

[0040] Figure 3 This is a schematic diagram of the structure of the multi-view fusion plant three-dimensional reconstruction device provided by the present invention.

[0041] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0043] The following is combined with Figures 1-4 The present invention describes a method, apparatus, device, and storage medium for three-dimensional reconstruction of plants using multi-view fusion.

[0044] Figure 1 This is one of the flowcharts illustrating the multi-view fusion method for three-dimensional reconstruction of plants provided by the present invention, such as... Figure 1 As shown, the method includes the following:

[0045] Step 100: Obtain three-dimensional prior knowledge of the plant based on a three-dimensional representation network; the three-dimensional representation network is trained based on unlabeled multi-view images of the plant.

[0046] This invention utilizes a mobile monitoring vehicle equipped with multiple cameras to capture continuous video of tomato plants from different perspectives through the movement of a robotic arm on the vehicle. A 3D representation network is pre-trained on a large number of unlabeled continuous views of tomato plants using self-supervised learning to acquire prior knowledge of the 3D scene.

[0047] Step 200: Extract features from the multi-view time-series images of the plant to be monitored to obtain multi-view image fusion features;

[0048] Multi-view time-series images of tomato plants are acquired by a mobile monitoring vehicle. The acquired multi-view video sequences (i.e., multi-view time-series images of the plants to be monitored) are input into an attention-based multi-view feature aggregation module (AMFAM) for adaptive weighted fusion of image features from different perspectives. This invention designs an attention-based multi-view feature aggregation module for adaptively weighted fusion of image features of tomato plants from different perspectives to obtain multi-view image fusion features.

[0049] Step 300: Based on the plant's three-dimensional prior knowledge and the multi-view image fusion features, perform three-dimensional reconstruction on the plant to be monitored to obtain a three-dimensional reconstruction model of the plant to be monitored.

[0050] This invention designs a NeRF-guided Rapid Modeling Algorithm. Guided by NeRF, the three-dimensional structure of tomato plants is rapidly recovered through Markov random field optimization, ensuring the accuracy of the three-dimensional structure while reducing computational overhead.

[0051] This embodiment introduces self-supervised learning into plant 3D reconstruction. A 3D representation network is trained using unlabeled multi-view images of the plant to obtain 3D prior knowledge. Then, an attention mechanism is used to fuse the features of the multi-view images. Finally, based on the trained 3D prior knowledge and the fused features of the multi-view images, the plant to be monitored is reconstructed in 3D. The 3D reconstruction model of the plant to be monitored is obtained accurately and quickly, which improves the efficiency of plant 3D reconstruction while ensuring the accuracy of plant 3D reconstruction.

[0052] Figure 2This is the second flowchart of the multi-view fusion method for three-dimensional reconstruction of plants provided by the present invention, as shown below. Figure 2 As shown, the method may further include:

[0053] Step 10: Based on the unlabeled multi-view images of the plants, determine the global image reconstruction error and image contrast loss;

[0054] Step 20: Determine the self-supervised learning objective based on the global image reconstruction error and the image contrast loss;

[0055] Step 30: Train a three-dimensional representation network based on the self-supervised learning objective.

[0056] Specifically, a mobile monitoring vehicle collects video (continuous-time images) of tomato plants from different perspectives to determine the location of the monitoring vehicle. The location coordinates of the cameras are The perspective direction is ,in, Indicates the horizontal direction angle; This represents the vertical angle. By precisely controlling the movement of the mobile monitoring vehicle, continuous acquisition of multi-view images of the tomato plant is achieved. Assume the robotic arm of the mobile monitoring vehicle is in... The position of the moment is Orientation is Then in At that moment, the first mobile monitoring vehicle The pose of each camera This can be represented as shown in Formula 1. Where, For the robotic arm in The rotation matrix at time step.

[0057] (1)

[0058] To obtain a comprehensive and high-quality time-series view of the tomato plant, the robotic arm's movements must meet the following conditions:

[0059] Horizontal angular coverage of each viewpoint Panorama, that is ;

[0060] The angle between adjacent viewpoints is not less than Generally take ,Right now ;in, For the first on the mobile monitoring vehicle The pose of each camera;

[0061] The robotic arm maintains a safe distance from the plant. , Generally taken ,Right now ,in These are the approximate center coordinates of the plant;

[0062] The movement speed of the robotic arm satisfy Among them, the maximum speed Generally taken .

[0063] Under the above constraints, the motion of the robotic arm can be planned using a gradient-based trajectory optimization algorithm, which can maximize the field of view coverage and spacing uniformity of the acquired view while avoiding collisions.

[0064] For unlabeled video data (i.e., unlabeled multi-view images of plants) collected by a mobile monitoring vehicle, this invention pre-trains a 3D representation network on the unlabeled multi-view images of plants using self-supervised learning. The training set of the unlabeled multi-view images of plants is determined as follows: Each video clip Include Frame image, i.e. One of the goals of self-supervised pre-training is to minimize the global image reconstruction error. As shown in Formula 2, where, A three-dimensional representation of the network; Networks for image reconstruction; This indicates the search for the minimum reconstruction error in the temporal neighborhood; Indicates the first The video clip is at a certain moment Images; Indicates the first The video clip is at a certain moment Images were used. PointNet++ (a deep learning model for point cloud segmentation and classification) and FoldingNet (an autoencoder for point clouds) were employed respectively to learn the 3D shape and appearance of tomato plants from point cloud representations. Let be the radius of the time neighborhood, which is generally taken as . The concept of temporal neighborhood is introduced, enabling the training process to model the spatiotemporal correlations between consecutive video frames.

[0065] (2)

[0066] (3)

[0067] In addition to reconstruction error, this invention also incorporates other self-supervised evaluation metrics (such as contrast loss between images). As shown in Formula 3, where, For image Three-dimensional feature representation; For image Three-dimensional feature representation; Indicates and Different times; The cosine similarity function; The temperature hyperparameter is used. Image contrast loss encourages feature representations from different moments in the same video to be closer in semantic space, while feature representations from different videos are further apart.

[0068] (4)

[0069] By combining the reconstruction error and contrast loss, the overall self-supervised learning objective is obtained. As shown in Formula 4, where, and These are the weighting coefficients. On a dataset of unlabeled multi-view images of plants, the 3D representation network... After pre-training, prior knowledge of the three-dimensional structure of tomato plants can be obtained, which can be used to guide subsequent multi-view fusion modeling.

[0070] In this embodiment, the motion control of the mobile monitoring vehicle and self-supervised learning are used to continuously collect multi-view images of tomato plants, and a three-dimensional representation network is pre-trained to obtain three-dimensional prior knowledge.

[0071] In one embodiment, the multi-view fusion method for three-dimensional reconstruction of plants provided in this application may further include:

[0072] Step 210: Extract features from the multi-view time-series images of the plant to be monitored to obtain the visual features of the images;

[0073] Step 220: Concatenate the feature maps from multiple perspectives to obtain a multi-view feature tensor; the feature maps are determined based on the visual features of the image;

[0074] Step 230: Fuse the multi-view feature tensors to obtain multi-view image fusion features.

[0075] Specifically, determine from Video sequences captured from various perspectives , where each video sequence Include Frame Image First, pre-trained convolutional neural networks (such as ResNet and VGG) are used to extract visual features from each frame of the image. Then, the feature maps from different viewpoints at the same time are concatenated along the viewpoint dimension to obtain the multi-view feature tensor. .

[0076] Then, the multi-view feature tensor The input is then fused into the attention aggregation module. This attention aggregation module consists of... The attention layer is obtained by concatenating several attention layers, each containing three parallel fully connected layers (FC), which are used to compute the three elements of the attention mechanism: query. 、key Sum of values ,in, This is a layer index.

[0077] In the Each attention layer first compresses the input features through spatial average pooling (AvgPool) to... Then, the data is fed in parallel into three fully connected layers to generate a query matrix, a key matrix, and a value matrix, as shown in Equation 5. , and These are the implicit space dimensions for queries, keys, and values, typically taken as... , For the number of attention heads; Indicates the first The layer is a fully connected layer used to generate the query matrix; Indicates the first The layer is a fully connected layer used to generate the key matrix; Indicates the first The layer is a fully connected layer used to generate the value matrix; express OK The space of real matrices in the column; express OK The space of real matrices in the column; express OK The space of real matrix columns.

[0078] (5)

[0079] Then, the attention weight matrix is ​​calculated using scaled dot-product attention (an attention mechanism). As shown in Formula 6, where, Indicates the first The first perspective on the Attention weights for each perspective; Key matrix transpose; express A 3D real matrix.

[0080] (6)

[0081] Finally, the output features are obtained by weighted summation of the value matrix using the attention matrix. , ,go through After iterative updates of the layer attention, the output is... Then, a fully connected layer is used to obtain the aggregated compact feature representation. This is used for subsequent 3D reconstruction networks.

[0082] To further improve the accuracy of 3D reconstruction of tomato plants, this invention, in addition to considering the commonly used Chamfer distance (a measure of the difference or similarity between two sets of points or shapes), also introduces normal vector consistency loss and semantic segmentation consistency loss in the design of the loss function. Among these, the normal vector consistency loss... The definition is shown in Formula 7, where, For point The predicted normal vector; This corresponds to the true normal vector; Represents the dot product of vectors; Indicates time The predicted set of point clouds; This indicates the size of the point cloud set.

[0083] (7)

[0084] (8)

[0085] Semantic segmentation consistency loss The definition is shown in Formula 8, where, For point Predict semantic labels (such as plant leaves, stems, and fruits). For the corresponding real semantic tags, This is a probability prediction.

[0086] (9)

[0087] Finally, this invention employs a weighted combination of multi-task loss functions, as shown in Formula 9, where, The multi-task loss is a weighted combination. The Chamfer distance loss is used. 、 、 These are the weighting coefficients, which can be tuned through cross-validation.

[0088] This embodiment uses adaptive weighted fusion of image features from different perspectives.

[0089] In one embodiment, the multi-view fusion method for three-dimensional reconstruction of plants provided in this application may further include:

[0090] Step 310: Divide the continuous three-dimensional space corresponding to the plant to be monitored to obtain a regular grid;

[0091] Step 320: Determine the occupying voxels in the regular grid;

[0092] Step 330: Extract the coordinates of the occupied voxels to obtain the voxelized three-dimensional information of the plant;

[0093] Step 340: Based on the plant's three-dimensional prior knowledge, the multi-view image fusion features, and the plant's voxelized three-dimensional information, perform three-dimensional reconstruction of the plant to be monitored to obtain a three-dimensional reconstruction model of the plant to be monitored.

[0094] Specifically, NeRF is an emerging implicit neural representation that learns the geometry and appearance of arbitrarily complex scenes through a continuous mapping of 5D coordinates between 3D position and 2D viewpoint. Given a set of calibrated multi-view images, NeRF reconstructs a 3D scene by optimizing the following continuum rendering equation (e.g., Equation 10), where, For rays The predicted color on; Cumulative transmittance; Position on the ray; The direction of view; Indicates the distal end of the ray; Indicates the near end point of the ray. and Representing positions respectively The density and color at a given location are parameterized by two multilayer perceptrons (MLPs).

[0095] (10)

[0096] This invention proposes a fast optimization algorithm for Markov Random Fields (MRF) based on NeRF priors. To reduce computational complexity, the continuous three-dimensional space of the plant to be monitored is first divided into... Individual elements form a regular grid. Each voxel The occupancy rate is represented as a binary variable. The problem is transformed into solving for the MRF distribution of voxel occupancy. A pre-trained NeRF network is used. It can calculate each voxel density and color As a priori guiding MRF energy function The construction is shown in Formula 11, where, and These are the univariate potential energy term and the bivariate potential energy term, respectively. Voxel representation The predicted occupancy rate; Indicates and Adjacent voxels; Indicates the occupancy rate of adjacent voxels; Voxel representation The color; Represents the set of neighborhood relations of a voxel.

[0097] (11)

[0098] (12)

[0099] To efficiently solve the problem of minimizing the energy function of the MRF, this invention employs a graph cut algorithm. First, the MRF is mapped to a flow network, where voxels and their neighbors correspond to nodes and edges, respectively, and the potential function value corresponds to the edge capacity. According to the maximum flow minimum cut theorem, finding the minimum cut of this network is equivalent to solving for the optimal configuration of the MRF. As shown in Formula 12.

[0100] This invention uses the augmented path method Boykov-Kolmogorov algorithm (an algorithm for solving the maximum flow problem) to solve the maximum flow problem, with a time complexity of O(n log n). ,in, The maximum flow value of the network is related to the scene complexity. After solving, the minimum cut divides the voxels into two sets: occupied and unoccupied, corresponding to the foreground and background regions of the 3D model, respectively.

[0101] Finally, the voxels occupied By extracting the vertex coordinates, the voxelized 3D structure of the plant can be obtained. Considering the discontinuities on the voxel surface, algorithms such as Marching Cubes can be used to further generate a triangular mesh model with denser vertices and a smoother topology.

[0102] This embodiment combines NeRF implicit representation with MRF graph optimization, achieving a dual guidance of prior knowledge and data constraints, and effectively balancing the accuracy and speed of 3D reconstruction.

[0103] In one embodiment, the multi-view fusion method for three-dimensional reconstruction of plants provided in this application may further include:

[0104] Step 400: Add a temporal feature field to the first neural radiation field to obtain a second neural radiation field; the first neural radiation field is used to perform three-dimensional reconstruction of the plant to be monitored; the second neural radiation field is used to perform N-dimensional reconstruction of the plant to be monitored, where N>3; the temporal feature field is determined based on the temporal morphology of the plant to be monitored.

[0105] Step 500: Determine the change information of the time sequence pattern in the time dimension;

[0106] Step 600: Based on the second neural radiation field and the change information, perform N-dimensional reconstruction of the plant to be monitored to obtain an N-dimensional reconstruction model of the plant to be monitored.

[0107] Specifically, this invention addresses the dynamic characteristics of three-dimensional morphological changes over time in tomato plant growth modeling by proposing a 4D spatiotemporal reconstruction scheme based on Temporal Feature Fields (TFF). The implicit temporal feature field constructed in this invention can map plant morphology at different growth stages into a compact low-dimensional manifold space, thereby achieving efficient encoding and decoding of the spatiotemporal evolution of plants.

[0108] This invention introduces an additional implicit temporal feature field based on the original NeRF. This is used to model the continuous changes in plant state over time. (The last two sentences appear to be fragments and cannot be translated accurately.) The features are represented as and its spatial characteristics The splicing forms a spatiotemporally consistent NeRF input. The extended spatiotemporal NeRF can be expressed as shown in Equation 13, where, For the spatial representation parameters of NeRF; These are the parameters of the temporal feature field; Represents the extended spacetime NeRF; Indicates the direction of view; Indicates time Color output; Indicates time Density output.

[0109] (13)

[0110] This means that NeRF's original 5D input is extended to 6D, giving NeRF the ability to depict time-varying scenes. Intuitively, Controlled at a given moment Lower NeRF spatial density field and radiation field The shape and distribution of the plant affect the rendered image of the plant at that moment.

[0111] To optimize the aforementioned spatiotemporal NeRF, it is necessary to consider different growth stages of the plant. Acquire images from multiple perspectives. (Note:) For the first Collected from various periods A collection of images, For the first Collected from various periods The set of camera parameters corresponding to each image.

[0112] The optimization objective of spatiotemporal consistency NeRF is to minimize the reconstruction error. As shown in Formula 14, where, To utilize Rendering the The first period The reconstruction results of the images are shown in Formula 15, where, Indicates the first Cumulative transmittance over a period of time; Indicates time Lower radiation In parameters Density at that location; Indicates time Lower radiation In parameters Location, perspective direction The color of the time.

[0113] (14)

[0114] (15)

[0115] To optimize implicit function parameters and implicit temporal feature fields This invention employs a layered, alternating strategy. Its pseudocode is as follows:

[0116]

[0117] Each episode (a complete data processing cycle or a training / testing experience) consists of two phases:

[0118] 1. For each period Fixed NeRF parameters Only optimize the feature representation of this period. Minimize the reconstruction error within the period This stage independently learns the optimal implicit encoding for each plant stage.

[0119] 2. Fixed characteristic representation for all periods Jointly optimize NeRF parameters across all periods. Minimize the accumulated reconstruction error This stage uses feature-guided optimization to generate a spatiotemporally consistent NeRF.

[0120] Through the above alternating learning, NeRF and implicit feature representation are adapted to each other in spatial and temporal dimensions, ultimately resulting in a continuous and smooth 4D spatiotemporal implicit scene representation of the plant that simultaneously fits multiple temporal observation data.

[0121] During the inference phase, given any query time... Its time series characteristics By inputting the optimized 4D NeRF, multi-view images of the plant at that moment can be reconstructed using volume rendering. Furthermore, by using an implicit field-based MRF optimization method to serialize the NeRF output in the spatial dimension, the query moment can be obtained. High-fidelity voxel model of plants at any time By using time series with uniform / adaptive sampling By repeating the single-time inference process, the spatiotemporal voxel manifold describing the complete plant growth process can be obtained. To achieve continuous interpolation of this manifold, 4D convolutional neural networks or graph neural networks can be used to learn the spatiotemporal correlation between voxels at different times, thereby generating plant growth animations at arbitrary time granularity.

[0122] The 4D plant reconstruction paradigm based on temporal feature fields proposed in this embodiment describes the intrinsic laws of continuous evolution of plants over time by introducing implicit dynamic features, and designs a hierarchical optimization strategy oriented towards spatiotemporal consistency, realizing end-to-end reconstruction from multiple temporal images to a high-fidelity 4D model.

[0123] In one embodiment, the multi-view fusion method for three-dimensional reconstruction of plants provided in this application may further include:

[0124] Step 700: Obtain the action information of the image acquisition device; the image acquisition device is used to acquire multi-view time-series images of the plant to be monitored;

[0125] Step 800: Obtain the posture information and real three-dimensional model of the plant to be monitored;

[0126] Step 900: Optimize the image acquisition device based on the motion information, the posture information, and the real 3D model.

[0127] Specifically, this invention uses a Markov Decision Process (MDP) to model the camera layout optimization task. Each acquisition time... The camera selection process is abstracted into a quadruple. .

[0128] state space : Includes the plant's posture parameters at the current moment (such as height) Projected area (etc.), a set of viewpoint indexes for acquired images. Information such as The above elements together depict the environmental conditions in which the decision is made.

[0129] Action space :Depend on Each action is composed of a combination of candidate camera viewpoints. For one A two-dimensional vector, where Indicates the selection of the first One perspective, This indicates that no selection is made.

[0130] State transition probability : Description in state Take action below Then, the environment transitions to the next state. The probability of the next state. In this problem, the next state depends only on the plant growth dynamics and the update of the collected view set, and is independent of the action selection. Therefore, it degenerates into a state transition function. .

[0131] reward function : Measurement in state Take action below The reward function designed in this invention comprehensively considers the balance between reconstruction accuracy and cost, and its specific form is shown in Formula 16, where, For the first A realistic 3D model of the plant at all times. To utilize perspective combinations The resulting reconstruction model. The intersection-union ratio (IUU) between the two models reflects the reconstructed similarity. express Norm, which is the number of selected viewpoints. This is a balancing factor that controls the relative importance of reconstruction accuracy and cost.

[0132] (16)

[0133] Within the MDP framework, the agent's goal is to find a policy function. This allows a series of actions taken under the guidance of this strategy to maximize the expected (objective) cumulative reward. As shown in Formula 17, where, Discount factor; Indicating in strategy The following expectations; Indicates the time step as Discount factor at time; Indicates the state Next strategy The given action.

[0134] (17)

[0135] Since both the state space and action space are high-dimensional discrete spaces, this invention uses the policy gradient method in deep reinforcement learning, namely the Proximal Policy Optimization (PPO) algorithm, to optimize and solve the above objectives.

[0136] This invention constructs a parameterized policy network. As an intelligent agent (an agent capable of sensing its environment, making decisions, and taking actions to achieve a specific goal, which can be software, hardware, or a system), where, These are the network weights. The network uses state features... As input, the output is the probability of choosing an action from each viewpoint in this state. Considering the high-dimensional discreteness of the action space, the output layer of the policy network uses the Sigmoid function (activation function) to model the action probabilities, as shown in Equation 18, where... For the first Network parameters corresponding to each viewpoint; For the Sigmoid function; Indicates the Bernoulli distribution; Indicates the first Choose an action from a different perspective; Indicates the first The inner product of network parameters and state features from each perspective.

[0137] (18)

[0138] To support gradient-based policy optimization, this invention utilizes the Gumbel-Softmax reparameterization technique (a method for maintaining gradient propagation when handling discrete random variable sampling in deep learning) to continuously relax discrete actions, transforming them into differentiable continuous distribution sampling.

[0139] The PPO algorithm achieves a good balance between training stability and sampling efficiency by limiting the policy update magnitude. Its core idea is to construct a surrogate objective (a type of objective function used to optimize a policy in reinforcement learning) that makes the current policy... Compared to the previous strategy The difference is constrained, as shown in Equation 19, where, For probability ratios, Let be the generalized dominance function, representing the action. The advantage over the average action value is shown in Formula 20, where, For action value functions, These are state-value functions, representing the states in the state... Take action below And the expected cumulative reward of random actions; Indicating in strategy Next state and actions The expectation. The function will give the probability ratio in Cut within the specified range to avoid excessively large policy updates.

[0140] (19)

[0141] (20)

[0142] The PPO algorithm alternates between policy evaluation and policy improvement in the following two steps:

[0143] 1. Strategy Evaluation: Fixed Strategy Parameters Estimating the state-value function using the Monte Carlo method and action value function As shown in Formulas 21 and 22, where, Indicates from time arrive The discounts are cumulative; Indicates at time The state; Indicates at time The actions taken; Indicates the discount factor; Indicates the parameter The gradient.

[0144] ;(twenty one)

[0145] ;(twenty two)

[0146] ;(twenty three)

[0147] 2. Strategy Improvement: Based on estimation The PPO objective function is obtained by applying the stochastic gradient ascent algorithm. conduct Further optimization is performed, updating the policy network parameters as shown in Equation 23, where, The learning rate is used. Constrained optimization reduces the risk of the policy prematurely converging to a local optimum.

[0148] By iterating through the above two steps, a near-optimal camera layout combination strategy is finally obtained. In practical applications, given the growth process of a plant to be reconstructed, the agent observes the state at each time step. The learned policy network is used to infer the optimal viewpoint selection in real time. This guides the complete process of multi-view data acquisition and 4D reconstruction.

[0149] The camera layout optimization algorithm based on reinforcement learning proposed in this embodiment can autonomously learn the strategy of viewpoint selection from a large amount of historical plant data, and adaptively adjust the strategy according to the current plant status, thereby maximizing the 4D reconstruction quality while saving hardware costs.

[0150] The following describes the plant three-dimensional reconstruction device with multi-view fusion provided by the present invention. The plant three-dimensional reconstruction device with multi-view fusion described below can be referred to in correspondence with the plant three-dimensional reconstruction method with multi-view fusion described above.

[0151] Please refer to Figure 3 The present invention also provides a multi-view fusion plant three-dimensional reconstruction device, comprising:

[0152] The plant 3D prior knowledge acquisition module 301 is used to acquire plant 3D prior knowledge based on a 3D representation network; the 3D representation network is trained based on unlabeled multi-view images of the plant.

[0153] The multi-view image fusion feature acquisition module 302 is used to extract features from the multi-view time-series images of the plant to be monitored, and obtain multi-view image fusion features.

[0154] The plant 3D reconstruction module 303 is used to perform 3D reconstruction of the plant to be monitored based on the plant's 3D prior knowledge and the multi-view image fusion features, so as to obtain a 3D reconstruction model of the plant to be monitored.

[0155] Optionally, the multi-view fusion plant 3D reconstruction device further includes:

[0156] The first determining module is used to determine the global image reconstruction error and image contrast loss based on the unlabeled multi-view images of the plants.

[0157] The self-supervised learning objective determination module is used to determine the self-supervised learning objective based on the global reconstruction error of the image and the contrast loss of the image.

[0158] A 3D representation network acquisition module is used to train a 3D representation network based on the self-supervised learning objective.

[0159] Optionally, the multi-view image fusion feature acquisition module includes:

[0160] The feature extraction unit is used to extract features from multi-view time-series images of the plant to be monitored, and obtain the visual features of the images.

[0161] The feature map stitching unit is used to stitch feature maps from multiple perspectives to obtain a multi-view feature tensor; the feature maps are determined based on the visual features of the image.

[0162] The feature fusion unit is used to fuse the multi-view feature tensors to obtain multi-view image fusion features.

[0163] Optionally, the plant three-dimensional reconstruction module includes:

[0164] The regular grid determination unit is used to divide the continuous three-dimensional space corresponding to the plant to be monitored into a regular grid;

[0165] Occupied voxel determination unit, used to determine occupied voxels in the regular grid;

[0166] The coordinate extraction unit is used to extract the coordinates of the occupied voxels to obtain the voxelized three-dimensional information of the plant.

[0167] The three-dimensional reconstruction model determination unit is used to perform three-dimensional reconstruction of the plant to be monitored based on the plant's three-dimensional prior knowledge, the multi-view image fusion features, and the plant's voxelized three-dimensional information, to obtain the three-dimensional reconstruction model of the plant to be monitored.

[0168] Optionally, the multi-view fusion plant 3D reconstruction device further includes:

[0169] A three-dimensional reconstruction module is used to add a temporal feature field to a first neural radiation field to obtain a second neural radiation field; the first neural radiation field is used to perform three-dimensional reconstruction of the plant to be monitored; the second neural radiation field is used to perform N-dimensional reconstruction of the plant to be monitored, where N>3; the temporal feature field is determined based on the temporal morphology of the plant to be monitored.

[0170] The change information determination module is used to determine the change information of the temporal pattern in the time dimension;

[0171] The N-dimensional reconstruction module is used to perform N-dimensional reconstruction of the plant to be monitored based on the second neural radiation field and the change information, so as to obtain the N-dimensional reconstruction model of the plant to be monitored.

[0172] Optionally, the multi-view fusion plant 3D reconstruction device further includes:

[0173] A multi-view time-series image acquisition module is used to acquire motion information of an image acquisition device; the image acquisition device is used to acquire multi-view time-series images of the plant to be monitored.

[0174] The first acquisition module is used to acquire the posture information and real three-dimensional model of the plant to be monitored.

[0175] An image acquisition device optimization module is used to optimize the image acquisition device based on the motion information, the posture information, and the real 3D model.

[0176] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a multi-view fusion method for 3D reconstruction of plants. This method includes: obtaining 3D prior knowledge of the plant based on a 3D representation network; the 3D representation network is trained based on unlabeled multi-view images of the plant; extracting features from the multi-view time-series images of the plant to be monitored to obtain multi-view image fusion features; and performing 3D reconstruction of the plant to be monitored based on the plant 3D prior knowledge and the multi-view image fusion features to obtain a 3D reconstruction model of the plant to be monitored.

[0177] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0178] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-view fusion plant 3D reconstruction method provided by the above methods. The method includes: obtaining plant 3D prior knowledge based on a 3D representation network; the 3D representation network is trained based on unlabeled multi-view images of the plant; extracting features from the multi-view time-series images of the plant to be monitored to obtain multi-view image fusion features; and performing 3D reconstruction of the plant to be monitored based on the plant 3D prior knowledge and the multi-view image fusion features to obtain a 3D reconstruction model of the plant to be monitored.

[0179] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for three-dimensional reconstruction of plants using multi-view fusion provided by the methods described above. This method includes: obtaining three-dimensional prior knowledge of the plant based on a three-dimensional representation network; the three-dimensional representation network being trained based on unlabeled multi-view images of the plant; extracting features from multi-view temporal images of the plant to be monitored to obtain multi-view image fusion features; and performing three-dimensional reconstruction of the plant to be monitored based on the plant's three-dimensional prior knowledge and the multi-view image fusion features to obtain a three-dimensional reconstruction model of the plant to be monitored.

[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0181] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-view fused plant three-dimensional reconstruction method, characterized in that, The method comprises the following steps: obtaining plant three-dimensional prior knowledge based on a three-dimensional representation network; the three-dimensional representation network is trained based on unannotated multi-view images of plants; extracting features from multi-view time sequence images of a plant to be monitored to obtain multi-view image fusion features; based on the plant three-dimensional prior knowledge and the multi-view image fusion features, performing three-dimensional reconstruction on the plant to be monitored to obtain a three-dimensional reconstruction model of the plant to be monitored; the loss function of three-dimensional reconstruction adopts a weighted combination multi-task loss function, and the weighted combination multi-task loss is determined according to a Chamfer distance loss, a normal vector consistency loss and a semantic segmentation consistency loss; wherein the normal consistency loss is defined as follows, wherein is a predicted normal vector of a point ; is a corresponding ground truth normal vector; denotes the vector inner product; denotes the time instant at which the point cloud set is predicted; denotes the size of the point cloud set; ; Semantic segmentation consistency loss The definition is shown in the following formula, where, For point Predicted semantic labels, For the corresponding real semantic tags, For probability prediction; ; ; wherein, is a weighted combination of multi-task losses; is a Chamfer distance loss; , , is a weight coefficient; before the plant three-dimensional prior knowledge is obtained based on the three-dimensional representation network, the following steps are included: based on the unannotated multi-view images of plants, determining an image global reconstruction error and an image contrast loss; based on the image global reconstruction error and the image contrast loss, determining a self-supervised learning target; ; ; wherein, is the image global reconstruction error, is the three-dimensional representation network; is the image reconstruction network; denotes finding the minimum of the reconstruction error in the temporal neighborhood; denotes the image of the video segment at time ; denotes the image of the video segment at time ; is the radius of the temporal neighborhood, is the image contrast loss, is the three-dimensional feature representation of the image ; is the three-dimensional feature representation of the image ; denotes a different time than ; is the cosine similarity function; is the temperature hyperparameter; training a three-dimensional representation network based on the self-supervised learning target.

2. The multi-view fused plant 3D reconstruction method of claim 1, wherein, the step of extracting features from multi-view time sequence images of a plant to be monitored to obtain multi-view image fusion features comprises the following steps: extracting features from multi-view time sequence images of a plant to be monitored to obtain image visual features; splicing feature maps of multi-views to obtain a multi-view feature tensor; the feature maps are determined based on the image visual features; fusing the multi-view feature tensor to obtain multi-view image fusion features.

3. The multi-view fused plant 3D reconstruction method of claim 1, wherein, the step of performing three-dimensional reconstruction on the plant to be monitored based on the plant three-dimensional prior knowledge and the multi-view image fusion features to obtain a three-dimensional reconstruction model of the plant to be monitored comprises the following steps: dividing a continuous three-dimensional space corresponding to the plant to be monitored to obtain a regular grid; determining an occupied voxel in the regular grid; extracting coordinates of the occupied voxel to obtain plant voxelized three-dimensional information; based on the plant three-dimensional prior knowledge, the multi-view image fusion features and the plant voxelized three-dimensional information, performing three-dimensional reconstruction on the plant to be monitored to obtain a three-dimensional reconstruction model of the plant to be monitored.

4. The multi-view fused plant 3D reconstruction method of claim 1, wherein, after the three-dimensional reconstruction on the plant to be monitored based on the plant three-dimensional prior knowledge and the multi-view image fusion features to obtain a three-dimensional reconstruction model of the plant to be monitored, the following steps are included: adding a time sequence feature field to a first neural radiation field to obtain a second neural radiation field; the first neural radiation field is used for three-dimensional reconstruction of the plant to be monitored; the second neural radiation field is used for N-dimensional reconstruction of the plant to be monitored, and N>3; the time sequence feature field is determined based on a time sequence morphology of the plant to be monitored; determining change information of the time sequence morphology in a time dimension; based on the second neural radiation field and the change information, performing N-dimensional reconstruction on the plant to be monitored to obtain an N-dimensional reconstruction model of the plant to be monitored.

5. The multi-view fused plant 3D reconstruction method of claim 1, wherein, after the three-dimensional reconstruction on the plant to be monitored based on the plant three-dimensional prior knowledge and the multi-view image fusion features to obtain a three-dimensional reconstruction model of the plant to be monitored, the following steps are included: Action information of an image acquisition device is acquired, the image acquisition device being configured to acquire multi-view time-series images of the plant to be monitored; Obtaining posture information and a real three-dimensional model of the plant to be monitored; Optimizing the image acquisition device based on the action information, the posture information and the real three-dimensional model.

6. A multi-view fused plant three-dimensional reconstruction device, characterized by comprising: Comprise: A plant three-dimensional prior knowledge obtaining module is configured to obtain plant three-dimensional prior knowledge based on a three-dimensional representation network; The three-dimensional representation network is trained based on unannotated multi-view images of plants; The plant three-dimensional prior knowledge is obtained based on the three-dimensional representation network, and the method comprises the following steps: Based on the unannotated multi-view images of plants, determine the image global reconstruction error and the image contrast loss; Based on the image global reconstruction error and the image contrast loss, determine the self-supervised learning target; ; ; wherein, is the image global reconstruction error, is the three-dimensional representation network; is the image reconstruction network; denotes finding the minimum of the reconstruction error in the temporal neighborhood; denotes the image of the video segment at time ; denotes the image of the video segment at time ; is the radius of the temporal neighborhood, is the image contrast loss, is the three-dimensional feature representation of the image ; is the three-dimensional feature representation of the image ; denotes a different time than ; is the cosine similarity function; is the temperature hyperparameter; Based on the self-supervised learning target, train the three-dimensional representation network; A multi-view image fusion feature obtaining module is configured to extract features from multi-view time-series images of the plant to be monitored to obtain multi-view image fusion features; A plant three-dimensional reconstruction module is configured to reconstruct the plant to be monitored in three dimensions based on the plant three-dimensional prior knowledge and the multi-view image fusion features to obtain a three-dimensional reconstruction model of the plant to be monitored; the loss function of three-dimensional reconstruction adopts a weighted combination multi-task loss function, and the weighted combination multi-task loss is determined according to a Chamfer distance loss, a normal vector consistency loss and a semantic segmentation consistency loss; wherein the normal consistency loss is defined as follows, wherein is the predicted normal vector of a point ; and is the corresponding ground truth normal vector; denotes the vector inner product; denotes the time instant at which the point cloud set is predicted; and denotes the size of the point cloud set. ; Semantic segmentation consistency loss is defined as follows, where, is the predicted semantic label for a point is the predicted semantic label for a point is the corresponding true semantic label, is the probability prediction; ; ; wherein, is a weighted combination of multi-task losses; is a Chamfer distance loss; , , is a weight coefficient.

7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor executes the computer program to realize the multi-view fusion plant three-dimensional reconstruction method according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the multi-view fusion plant three-dimensional reconstruction method according to any one of claims 1 to 5.

9. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the multi-view fusion plant three-dimensional reconstruction method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image labeling method, computer equipment and storage medium

    CN117152753A

  • Underwater multi-view three-dimensional reconstruction method

    CN119295645A