Method and system for multi-view projection contour based generative adversarial point cloud completion network

By employing a multi-view contour constraint method based on generative adversarial networks, the problem of missing internal structural information in plant 3D reconstruction was solved, achieving efficient and accurate plant 3D reconstruction and improving the accuracy of plant phenotypic parameter measurement.

CN119941972BActive Publication Date: 2026-05-08HANGZHOU DIANZI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2024-11-20
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies for 3D reconstruction of plants suffer from problems such as missing information on the internal structure of the plant and low reconstruction efficiency. In particular, due to the occlusion caused by the diversity of leaves and stems and the limitations of sensor accuracy, the information on the internal structure of the plant cannot be completely collected.

Method used

A multi-view contour constraint method based on generative adversarial networks is adopted. By collecting point clouds of plants without self-occlusion, performing non-rigid transformation and label addition, multi-view contour projection images are extracted. Multi-resolution feature encoding and generator are used to predict missing regions. The network parameters are optimized through multi-view projection contour constraint adversarial training to achieve point cloud completion.

Benefits of technology

It achieves accurate reconstruction of the three-dimensional structure of plants, improves reconstruction efficiency and accuracy, reduces equipment complexity and interference, enhances anti-interference ability, and improves the accuracy of plant phenotypic parameter measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941972B_ABST
    Figure CN119941972B_ABST
Patent Text Reader

Abstract

The application discloses a method and system for generating a generative adversarial point cloud completion network based on multi-view projection profiles, and the method comprises the following steps: S1, collecting plant point clouds without self-occlusion structure; S2, performing non-rigid transformation on the plant point clouds to obtain plant point cloud data with self-occlusion, and adding labels to the plant point clouds corresponding to the classified plant parts; S3, extracting multi-view profile projection images of each plant point cloud to obtain occluded areas and non-occluded areas corresponding to complete point clouds; S4, inputting incomplete point clouds into a multi-resolution feature encoder to fuse and establish feature codes with missing areas; S5, inputting the feature codes of the incomplete point clouds into a generator of a generative adversarial network for point cloud completion to predict missing areas corresponding to the point clouds; and S6, using multi-view projection profiles of real missing areas to constrain the prediction results, and then inputting the constrained point clouds into a point cloud and profile discriminator of the generative adversarial network for point cloud completion to perform adversarial training and optimize network parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of three-dimensional phenotypic reconstruction technology of agricultural plants, specifically involving a plant point cloud completion method based on multi-view contour constraints of generative adversarial networks. Background Technology

[0002] In recent years, 3D digitization technology has shown significant application prospects in fields such as medicine, military aerospace, and agriculture. In particular, digital agriculture has become an important trend in my country's agricultural development. 3D plant reconstruction can acquire various 3D phenotypic features under different conditions, including height, width, leaf angle, leaf area, and canopy volume. These features not only reflect the genetic characteristics of crops but also the influence of their growth environment and field management. This technology has significant application value in plant breeding, crop growth management, and virtual visualization. Traditional 3D plant phenotypic measurements mainly rely on manual measurement, which is limited by the accuracy of measuring tools and manpower, resulting in time-consuming, labor-intensive, and inaccurate measurement results. Constructing 3D plant models digitally allows for the rapid, accurate, and efficient acquisition of plant phenotypic information.

[0003] The motion-based structure reconstruction method is a general approach for reconstructing 3D plant models. After fixing the plant at a selected location, a camera captures color images of the stationary plant from different perspectives. By extracting image feature points and performing feature matching, and adjusting and optimizing the matching strategy, the structural data of the plant's 3D space is obtained. However, the technical implementation steps involved in feature extraction, feature matching, and optimization are the reasons for the high computational load and low reconstruction efficiency of this method. Obtaining the 3D structure of a plant depends on depth information. An RGB-D camera is a 3D sensor that combines a color image sensor and a depth image sensor. It extracts the depth information of the target point based on the Time-of-Flight (TOF) principle. It actively projects a laser onto the object being measured, and the camera's sensor receives the diffuse reflection of the laser light from the rear end of the object's surface. The distance from the laser emission to reception is calculated as the depth. This active laser projection method for acquiring depth information has advantages such as strong anti-interference capability, high stability, and fast reconstruction speed, and has been widely used in target grasping, 3D navigation, and 3D reconstruction. In actual data acquisition, a fixed RGB-D camera is used to obtain 3D point cloud data from different perspectives by rotating the object under test. The point cloud data from different perspectives are then stitched and fused to obtain a complete 3D point cloud model of the object under test. In the 3D reconstruction of plants, due to the diverse distribution, size, position, and orientation of leaves and stems, acquiring images at certain rotational angle intervals is necessary to address the reconstruction challenges posed by leaf diversity. This involves processing a large amount of point cloud data and is inherently somewhat indiscriminate. Furthermore, information acquisition from a single perspective can lead to information loss due to occlusion of the plant's internal structure caused by its own structure. Simultaneously, limited by sensor accuracy, unoccluded areas may lose some information, resulting in further information loss. To ensure the accuracy of plant 3D phenotypic measurement, it is necessary to predict, complete, and reconstruct the internal structure to obtain complete and sufficient plant point cloud data, thereby improving reconstruction efficiency and performance. Based on this, this invention proposes a point cloud completion technique based on multi-view contour constraints using generative adversarial networks, which can be used for the measurement and evaluation of plant 3D phenotypic characteristics. Summary of the Invention

[0004] To overcome the problems existing in the prior art, this invention proposes a method and system for predicting and completing the three-dimensional point cloud of plants based on a generative adversarial point cloud completion network with multi-view projection contours, in order to solve problems such as the lack of three-dimensional point cloud information of the internal structure of plants and the lack of data during the sensor acquisition process.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] A method for generative adversarial point cloud completion networks based on multi-view projected contours includes the following steps:

[0007] S1. Collect point clouds of plants whose internal structure does not have self-occlusion;

[0008] S2. Perform non-rigid transformation on the plant point cloud, obtain a large amount of plant point cloud data with self-occlusion based on the biological structure of the plant, and add labels to the plant parts corresponding to the classification of the point cloud.

[0009] S3. Extract the multi-view contour projection image of each plant point cloud, and obtain the occluded area (i.e. the real missing area point cloud) and the unoccluded area (i.e. the incomplete point cloud) corresponding to the complete point cloud based on its positional relationship with the virtual camera.

[0010] S4. Input the incomplete point cloud into the multi-resolution feature encoder, fuse and establish feature codes with missing regions;

[0011] S5. The feature encoding of the incomplete point cloud is fed into the generator of the generative adversarial network for point cloud completion to predict the missing regions corresponding to the point cloud.

[0012] S6. Constrain the prediction results using the multi-view projection contours of the real missing regions to ensure that the generation is within the effective range. Then, input the constrained point cloud into the point cloud discriminator and contour discriminator of the generative adversarial network for point cloud completion for adversarial training to optimize the network parameters.

[0013] Preferably, step S1 is as follows:

[0014] S1.1. Fix the rotating stage, image acquisition equipment, and plant without complex obstruction structures to ensure that the plant is within the visible field of view of the acquisition equipment;

[0015] S1.2. Obtain the transformation relationship between adjacent viewpoints through motion recovery structure, and extract the depth information of the plant part in the depth map using the mask of RGB image;

[0016] S1.3. Obtain complete point cloud data of the plant without complex occlusion structure by using the principle of multi-view stereo imaging.

[0017] Preferably, step S1.3 specifically includes:

[0018] Generate 3D point cloud data for the current viewpoint based on depth camera intrinsics and depth images:

[0019]

[0020] Where, x depth y depth z depth These are the x, y, and z coordinates of the object in the point cloud coordinate system, respectively, and f xdepth f ydepth c represents the axial and radial focal lengths of the depth image acquisition device, respectively.xdepth c ydepth These are the principal point coordinates of the depth map image, u depth v depth These are the horizontal and vertical pixel coordinates of the depth map, respectively.

[0021] Preferably, step S2 specifically includes:

[0022] S2.1. For a single plant point cloud without a self-occlusion structure, each time a new complete plant point cloud with a complex self-occlusion structure is obtained by manually copying, twisting, rotating, and pruning the point clouds of some leaves and stems.

[0023] S2.2. A complete plant consists of three main parts: stem, stalk, and leaves. A corresponding label value is assigned to the point cloud of each part.

[0024] S2.3. Repeat the above two steps to obtain the required number of complete plant point cloud datasets. The number of repetitions depends on the specific number of plant point clouds to be collected. For example, if three complete plant point clouds are required, the repetitions should be repeated three times. The process ends after one complete plant point cloud has been collected.

[0025] Preferably, in step S3, the distance between the camera and the plant is determined by referring to the camera position and intrinsic parameter information when the data is collected in step S1. The position of the virtual camera in the point cloud coordinate system is constructed with the distance as the radius. Based on the view of each virtual camera, the occluded area (i.e. the real missing area point cloud) and the unoccluded area (i.e. the incomplete point cloud) under different view of the same plant are determined.

[0026] Preferably, step S4 specifically includes:

[0027] S4.1. For the incomplete input point cloud, the IFPS algorithm is used to downsample the number of points in each point cloud to 2048, 1024, and 512, representing three resolutions from high to low.

[0028] S4.2. Use CMLP (Combined Multi-Layer Perception) with multiple fully connected layers to extract features from point clouds of different resolutions;

[0029] S4.3. Concatenate and stitch the features obtained at different resolutions, and input them into a multilayer perceptron (MLP) to extract the final 1920-dimensional feature vector.

[0030] Preferably, step S5 specifically includes:

[0031] S5.1. The 1920-dimensional feature vector obtained by multi-resolution feature encoding is input into the generator of the generative adversarial network for point cloud completion. The generator passes the input feature vector through four fully connected linear layers to obtain feature vectors with dimensions of 1024, 512, 256 and 256 respectively, which can extract global and local features at the same time.

[0032] S5.2. The four features of different dimensions are passed through different layers in the generator. The feature with a dimension of 1024 is regarded as the first layer, the feature with a dimension of 512 is regarded as the second layer, and so on. The feature with the first dimension of 256 is regarded as the third layer, and the feature with the second dimension of 256 is regarded as the fourth layer.

[0033] S5.3. Output point clouds of size M1×3 from layers 4 and 3, and stitch them together to obtain point cloud PC of size M2×3. primary , representing the low-resolution predicted value of the missing part, where:

[0034] 2M1=M2

[0035] Where M1 and M2 are the number of points used to generate point clouds at different resolutions, and their settings are determined by the final M.

[0036] S5.4. The second layer output is a point cloud of size 3M2×3. Compare it with the PC. primary After stitching, a point cloud PC of size M×3 is obtained. secondary , representing the medium-resolution predicted value of the missing portion, where:

[0037] 4M² = M

[0038] Where M is the number of points generated in the final point cloud, and the choice of M will affect the settings of M1 and M2;

[0039] S5.5. The first layer outputs a point cloud of size M×3. Combine this with the PC... secondary After stitching, a point cloud PC with a size of 2M×3 is obtained. detail , representing the high-resolution predicted value of the missing part;

[0040] Preferably, step S6 specifically includes:

[0041] S6.1. Based on the multi-view projected contour image of the missing region under the current viewpoint, points in the predicted point cloud of the missing region that exceed the range of the projected contour are regarded as erroneously generated points, and their values ​​are set to 0 for constraint generation.

[0042] S6.2. Feed the constrained predicted point cloud into the point cloud discriminator and contour discriminator in the generative adversarial network for point cloud completion to train the adversarial loss and optimize the network parameters.

[0043] Preferably, the CMLP, generator, and discriminator described in steps S4, S5, and S6 are jointly trained in supervised training using a multi-objective loss function, which is as follows:

[0044] L=λ com L com +λ points L points +λ contour L contour

[0045] Where, λ com , λ points , λ contour L respectively com L points L coontour The weight, L com L represents the Chamfer Distance (CD) loss (typically used when comparing the actual and predicted point clouds) between the three predicted point clouds and the actual point clouds containing missing regions. points L represents the binary cross-entropy loss between the predicted point cloud and the actual missing region point cloud. contour The binary cross-entropy loss represents the difference between the predicted multi-view projection contours of the point cloud and the actual multi-view projection contours of the missing region point cloud, as detailed below:

[0046]

[0047] Among them, PC detaol PC primary PC secondary These represent the generated high, low, and medium resolution point clouds, respectively. α and β are the weights of the CD loss at low and medium resolutions, respectively. PC gt Represents the ground truth value of the truly missing region. These represent the ground truth values ​​obtained by performing one and two iterations of IFPS (Iterative Farthest Point Sample) downsampling on the ground truth values ​​of the missing regions, respectively. The CD loss between two point clouds is represented by D(), where D() represents the point cloud discriminator, G() represents the generator, and x j This indicates an incomplete point cloud belonging to the input, y i The point cloud represents the actual missing region, and S represents the size of the dataset.

[0048] L front L side L topThese represent the binary cross-entropy losses of the predicted point cloud and the projected contours of the point cloud of the actual missing region, based on the three viewpoints of the virtual camera position. Their viewpoint direction vectors are parallel to the x, y, and z coordinate axes. Specific details are as follows:

[0049]

[0050] Among them, L front、side、top L represents front L side L top Both can be calculated according to this formula. Set1 represents the predicted point cloud of the missing region, and Set2 represents the actual point cloud of the missing region. x and y belong to the points in Set1 and Set2, respectively. j y i The definitions of G() are as follows, and in particular, D in the formula here... c () represents the projected contour discriminator, whose input I(y) i ), I(x j () is two-dimensional image data.

[0051] This invention also discloses a system for a generative adversarial point cloud completion network based on multi-view projected contours, used to perform the above method, which includes the following modules:

[0052] Unoccluded plant point cloud acquisition module: Acquires plant point clouds without self-occlusion.

[0053] The self-occluding plant point cloud data acquisition module performs non-rigid transformation on the plant point cloud, acquires plant point cloud data with self-occlusion based on the plant's biological structure, and adds labels to the plant parts corresponding to the point cloud classification.

[0054] Multi-view contour projection image extraction module: Extracts multi-view contour projection images of each plant point cloud, and obtains the occluded area (i.e., the real missing area point cloud) and the unoccluded area (i.e., the incomplete point cloud) corresponding to the complete point cloud based on its positional relationship with the virtual image acquisition device.

[0055] The feature encoding module for missing regions: inputs the incomplete point cloud into the multi-resolution feature encoder, fuses and establishes feature encodings for the missing regions;

[0056] Missing Region Prediction Module: Input the feature encoding of the incomplete point cloud into the generator of the generative adversarial network for point cloud completion, and predict the missing regions corresponding to the point cloud.

[0057] Adversarial training module: The prediction results are constrained by multi-view projection contours of real missing regions. The constrained point cloud is then fed into the point cloud discriminator and contour discriminator of the generative adversarial network for point cloud completion for adversarial training to optimize the network parameters.

[0058] This invention has the following characteristics and beneficial effects:

[0059] This invention employs a generative adversarial multi-view contour-constrained plant point cloud completion scheme to reconstruct the complete 3D structure of plants non-contactly and non-destructively. Addressing the problems of severe occlusion in current 3D plant reconstructions and the inability to capture complete 3D information about the plant's internal structure from multi-view image acquisition, leading to accumulated data processing errors and inaccurate reconstruction results, this invention uses point cloud completion to predict missing regions, eliminating the need for cumbersome camera calibration and achieving accurate plant reconstruction. Furthermore, this invention requires simple equipment, has strong anti-interference capabilities, and produces a high-precision plant reconstruction model, effectively improving the accuracy of plant phenotypic parameter measurements. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 This is a flowchart of a generative adversarial point cloud completion network based on multi-view projection contours, according to an embodiment of the present invention.

[0062] Figure 2 This is a schematic diagram illustrating the data collection and non-rigid transformation to generate different plants in an embodiment of the present invention.

[0063] Figure 3 This is a schematic diagram illustrating the placement of a virtual camera at a reference position to obtain the occluded area according to an embodiment of the present invention.

[0064] Figure 4 This is a schematic diagram of the Multi-view Contour Constraint Generative Point Cloud Completion Network (MCCGPCN) model structure in an embodiment of the present invention.

[0065] Figure 5 This is a schematic diagram of the generator of MCCGPCN in an embodiment of the present invention.

[0066] Figure 6 This is a schematic diagram of the overall structure of the point cloud discriminator and contour discriminator of the MCCGPCN in this embodiment of the invention.

[0067] Figure 7This is an example of the result of plant point cloud completion in an embodiment of the present invention.

[0068] Figure 8 This is a system block diagram of a generative adversarial point cloud completion network based on multi-view projection contours, according to an embodiment of the present invention. Detailed Implementation

[0069] The preferred embodiments of the present invention will now be described in detail. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0070] This embodiment provides a multi-view contour-constrained plant point cloud completion method based on generative adversarial networks, such as... Figure 1 As shown, it includes the following steps:

[0071] S1. Collect point clouds of plants without self-shading structures;

[0072] S2. Perform non-rigid transformation on the plant point cloud to obtain plant point cloud data with self-occlusion, and add labels to the plant parts corresponding to the point cloud classification.

[0073] S3. Extract the multi-view contour projection image of each plant point cloud, and obtain the occluded area (i.e. the real missing area point cloud) and the unoccluded area (i.e. the incomplete point cloud) corresponding to the complete point cloud based on its positional relationship with the virtual image acquisition device.

[0074] S4. Input the incomplete point cloud into the multi-resolution feature encoder, fuse and establish feature codes with missing regions;

[0075] S5. Input the feature encoding of the incomplete point cloud into the generator of the generative adversarial network for point cloud completion to predict the missing regions corresponding to the point cloud.

[0076] S6. Constrain the prediction results using the multi-view projection contours of the real missing regions, and then feed the constrained point cloud into the point cloud discriminator and contour discriminator of the generative adversarial network for point cloud completion for adversarial training to optimize the network parameters.

[0077] Specifically, by collecting plant data samples where leaves and stems do not obscure the internal structure, complete 3D point cloud data of the plant is obtained. The point clouds are then transformed using non-rigid transformation software. Through operations such as cropping, copying, and twisting, different complete 3D point clouds with mutual occlusion within the plant are generated to construct a dataset. Multi-view contour projection images are extracted from each plant point cloud to obtain the occluded and unoccluded regions of the complete point cloud. The point clouds of the unoccluded regions are processed by a multi-resolution encoder to obtain their feature codes. The feature codes are then fed into the generator of MCCGPCN (Multi-view Contour Constraint Generative Point Cloud Completion Network). The predicted missing region point clouds are constrained by the multi-view projection contour images. The constrained point clouds are then fed into the discriminator of MCCGPCN to train the adversarial loss and further optimize the generative adversarial network.

[0078] In this embodiment, it includes two parts: a generative point cloud completion network for dataset generation and multi-view contour constraints.

[0079] 1) Dataset generation for plant 3D reconstruction:

[0080] Since there are no readily available public datasets for the 3D point cloud prediction and completion task, and the cost of collecting large datasets is very high, this invention uses a multi-view 3D reconstruction method to collect a small number of complete 3D point clouds of plants. Simultaneously, it combines plant samples representing different structural configurations and uses the open-source software Blender to perform non-rigid transformations on the point clouds to expand the dataset from different plant point cloud samples. The specific implementation is as follows:

[0081] a) such as Figure 2 As shown, fix the rotating platform and camera, select a plant whose leaves and stems do not obscure the internal structure of the plant in the direction of the camera's optical axis, and fix the plant on the rotating platform.

[0082] b) Use a fixed RGB-D camera to acquire RGB and depth images of the plant rotating one revolution at fixed intervals. Use the mask of the RGB image to extract the depth information of the plant part. Obtain the complete point cloud data of the single plant without complex occlusion structure through the principle of multi-view stereo imaging.

[0083] c) For the complete point cloud data of a single plant, non-rigid transformations are performed on the point cloud using the open-source software Blender. Specific transformation methods include, but are not limited to: manually selecting a portion of the leaf point cloud and bending, twisting, and stretching this portion based on possible leaf bending under natural conditions to obtain a new leaf point cloud with a shape different from the original data; manually selecting a portion of the leaf point cloud and trimming it (removing that portion of the point cloud) to obtain a new leaf point cloud with some missing leaf shapes compared to the original; manually selecting any complete leaf and copying, rotating, and moving it to another part of the stem to obtain a new complete point cloud data with an increased number of leaves. Based on these non-rigid transformations of the collected complete point clouds of a single plant, a large number of complete point clouds of plants with different structural configurations are generated.

[0084] d) Using the open-source software CloudCompare, add label values ​​to the generated complete point cloud of the plant one by one, with each label value corresponding to a different part of the plant;

[0085] e) such as Figure 3 As shown, based on the camera position when collecting data from a single plant, and with the plant point cloud coordinate system as a reference, a virtual camera position is established under the same collection radius. The observation angle is determined according to the direction of the camera optical axis. Based on each virtual camera angle, the occluded area (i.e., the real missing area point cloud) and the unoccluded area (i.e., the incomplete point cloud) of the same plant are determined under different angles.

[0086] f) Divide the generated dataset into training and validation sets according to a certain ratio.

[0087] 2) Generative Point Cloud Completion Network MCCGPCN with Multi-view Contour Constraints:

[0088] The Generative Point Cloud Completion Network (MCCGPCN) for multi-view contour constraints of plants consists of two parts: a feature extraction network and a multi-view contour constraint point cloud generation network. The network structure is as follows: Figure 4 As shown in the figure. The feature extraction network is a multi-resolution perceptron network consisting of three CMLP (Combined Multi Layer Perception) networks connected in parallel, and the multi-view contour constraint point cloud generation network is a generative adversarial network that generates multi-view contour point cloud prediction results.

[0089] Given the true missing region point cloud and incomplete point cloud of the same plant, iterative farthest point sampling (IFPS) is used to downsample the point clouds, resulting in point clouds with [2048, 1024, 512] points, corresponding to three different resolutions. This reduces the amount of data processed while primarily preserving the geometric information of the point clouds without causing the loss of key information. The point clouds at different resolutions are then passed through parallel multilayer perceptrons (CMLPs), each CMLP consisting of a [64-128-256-512-1024]-dimensional multilayer perceptron (MLP), yielding features for the three point clouds at different resolutions. These three features are then concatenated and passed through an MLP to obtain a 1920-dimensional point cloud feature vector F. This feature vector F contains both global features and local detailed features of the complete plant point cloud.

[0090] The point cloud feature vector F is input into the generator of the Multi-View Contour Constrained Point Cloud Generation Network (MCCGPCN). The generator structure is as follows: Figure 5 As shown, the generator processes the input feature vector through four fully connected linear layers to obtain feature vectors with dimensions of 1024, 512, 256, and 256, representing coarse and fine features globally and locally, respectively. The four different dimensional features are passed through different layers in the generator, with the 1024-dimensional feature vector considered as layer 1, the 512-dimensional feature vector as layer 2, and so on, increasing sequentially. The fourth and third layers output point clouds of size M1×3, which are then concatenated to obtain a point cloud PC of size M2×3. primary , representing the low-resolution predicted value of the missing part, where:

[0091] 2M1=M2

[0092] The second layer outputs a point cloud of size 3M2×3. This is then compared with the PC... primary After stitching, a point cloud PC of size M×3 is obtained. secondary , representing the medium-resolution predicted value of the missing portion, where:

[0093] 4M² = M

[0094] The first layer outputs a point cloud of size M×3. This is then compared with the PC... secondary After stitching, a point cloud PC with a size of 2M×3 is obtained. detail , representing the high-resolution predicted value of the missing part;

[0095] Based on the multi-view projected contour image of the missing region from the current perspective, points in the predicted point cloud that exceed the projected contour range are considered erroneously generated points, and their values ​​are set to 0 to achieve constrained generation. The constrained predicted point cloud is then fed into the point cloud discriminator and contour discriminator in the generative adversarial network for point cloud completion for adversarial loss training to optimize network parameters. The structures of the point cloud discriminator and contour discriminator are as follows: Figure 6 As shown;

[0096] The multi-objective loss function for supervising the CMLP, generator, and discriminator is as follows:

[0097] L=λ com L com +λ points L points +λ contour L contour

[0098] Among them, L com L represents the CD loss of the predicted point clouds at three different resolutions compared to the actual missing region point clouds. points L represents the binary cross-entropy loss between the predicted point cloud and the actual missing region point cloud. contour The binary cross-entropy loss represents the difference between the predicted multi-view projection contours of the point cloud and the actual multi-view projection contours of the missing region point cloud, as detailed below:

[0099]

[0100] Among them, PC gt Represents the ground truth value of the truly missing region. These represent the ground truth values ​​obtained by performing one and two IFPS downsampling operations on the ground truth values ​​of the truly missing regions, respectively. The CD loss between two point clouds is represented by D(), where D() represents the point cloud discriminator, G() represents the generator, and x j This indicates an incomplete point cloud belonging to the input, y i The point cloud represents the true missing region, S represents the size of the dataset, and L represents the size of the dataset. front L side L top Let represent the binary cross-entropy loss of the projected contours of the predicted point cloud and the point cloud of the actual missing region based on the three viewpoints of the virtual camera position, respectively. Their viewpoint direction vectors are parallel to the x, y, and z coordinate axes, specifically:

[0101]

[0102]

[0103] Where Set1 represents the predicted point cloud of the missing region, Set2 represents the actual point cloud of the missing region, and x and y belong to points in Set1 and Set2 respectively. j y i and D c The definition of () is as follows, where D is... c () represents the projected contour discriminator, whose input I(y) i ), I(x j () is two-dimensional image data.

[0104] like Figure 8 As shown, this embodiment discloses a system for a generative adversarial point cloud completion network based on multi-view projected contours, used to execute the above method embodiment, which includes the following modules:

[0105] Unoccluded plant point cloud acquisition module: Acquires plant point clouds without self-occlusion.

[0106] The self-occluding plant point cloud data acquisition module performs non-rigid transformation on the plant point cloud, acquires plant point cloud data with self-occlusion based on the plant's biological structure, and adds labels to the plant parts corresponding to the point cloud classification.

[0107] Multi-view contour projection image extraction module: Extracts multi-view contour projection images of each plant point cloud, and obtains the occluded area (i.e., the real missing area point cloud) and the unoccluded area (i.e., the incomplete point cloud) corresponding to the complete point cloud based on its positional relationship with the virtual image acquisition device.

[0108] The feature encoding module for missing regions: inputs the incomplete point cloud into the multi-resolution feature encoder, fuses and establishes feature encodings for the missing regions;

[0109] Missing Region Prediction Module: Input the feature encoding of the incomplete point cloud into the generator of the generative adversarial network for point cloud completion, and predict the missing regions corresponding to the point cloud.

[0110] Adversarial training module: The prediction results are constrained by multi-view projection contours of real missing regions. The constrained point cloud is then fed into the point cloud discriminator and contour discriminator of the generative adversarial network for point cloud completion for adversarial training to optimize the network parameters.

[0111] Other aspects of this embodiment can be found in the above method embodiments.

[0112] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and these variations still fall within the protection scope of the present invention.

Claims

1. A method for generative adversarial point cloud completion networks based on multi-view projected contours, characterized in that, Includes the following steps: S1. Collect point clouds of plants without self-shading structures; S2. Perform non-rigid transformation on the plant point cloud to obtain plant point cloud data with self-occlusion behavior, and add labels to the plant parts corresponding to the point cloud classification. S3. Extract the multi-view contour projection image of each plant point cloud, and obtain the occluded area (i.e. the real missing area point cloud) and the unoccluded area (i.e. the incomplete point cloud) corresponding to the complete point cloud based on its positional relationship with the virtual image acquisition device. S4. Input the incomplete point cloud into the multi-resolution feature encoder, fuse and establish feature codes with missing regions; S5. Input the feature encoding of the incomplete point cloud into the generator of the generative adversarial network for point cloud completion to predict the missing regions corresponding to the point cloud. S6. Constrain the prediction results using the multi-view projection contours of the real missing regions, and then feed the constrained point cloud into the point cloud discriminator and contour discriminator of the generative adversarial network for point cloud completion for adversarial training to optimize the network parameters. Step S4 specifically includes: S4.

1. For the incomplete input point cloud, the IFPS algorithm is used to downsample the number of points in each point cloud to 2048, 1024, and 512, representing three resolutions from high to low. S4.

2. Use CMLP with multiple fully connected layers to extract features from point clouds at different resolutions; S4.

3. The features obtained at different resolutions are concatenated and stitched together, and then input into a multilayer perceptron for extraction, ultimately forming a 1920-dimensional feature vector; Step S5 specifically includes: S5.

1. The 1920-dimensional feature vector obtained by multi-resolution feature encoding is input into the generator of the generative adversarial network for point cloud completion. The generator passes the input feature vector through four fully connected linear layers to obtain feature vectors with dimensions of 1024, 512, 256 and 256 respectively. S5.

2. The four feature vectors of different dimensions are passed through different layers in the generator. The feature vector with a feature dimension of 1024 is regarded as the first layer, the feature vector with a feature dimension of 512 is regarded as the second layer, the feature vector with the first feature dimension of 256 is regarded as the third layer, and the feature vector with the second feature dimension of 256 is regarded as the fourth layer. S5.

3. The output sizes of the 4th and 3rd layers are The point cloud, after being stitched together, yields a size of Point clouds , representing the low-resolution predicted value of the missing part, where: in, The number of points used to generate point clouds at different resolutions; S5.

4. The output size of the second layer is The point cloud, and its relationship with After splicing, the resulting size is Point clouds , representing the medium-resolution predicted value of the missing portion, where: in, It represents the number of points generated in the final point cloud; S5.

5. The output size of the first layer is The point cloud, and its relationship with After splicing, the resulting size is Point clouds , representing the high-resolution predicted value of the missing part; The CMLP, generator, and discriminator are jointly trained in supervised training using a multi-objective loss function, which is as follows: in, They are respectively The weight it accounts for This represents the CD loss between the predicted point clouds at three different resolutions and the actual point cloud showing the missing regions. This represents the binary cross-entropy loss between the predicted point cloud and the actual missing region point cloud. The binary cross-entropy loss represents the difference between the predicted multi-view projection contours of the point cloud and the actual multi-view projection contours of the missing region point cloud, as detailed below: in, These represent the generated high-resolution, low-resolution, and medium-resolution point clouds, respectively. , These represent the weights for CD loss at low and medium resolutions, respectively. Represents the ground truth value of the truly missing region. , These represent the ground truth values ​​obtained by performing one and two IFPS downsampling operations on the ground truth values ​​of the truly missing regions, respectively. This represents the CD loss between two point clouds. , This represents a point cloud discriminator. Represents a generator. This indicates an incomplete point cloud belonging to the input. Point cloud representing the truly missing regions, Indicates the size of the dataset. The binary cross-entropy loss represents the projected contours of the predicted point cloud and the actual missing region point cloud from three perspectives based on the location of the virtual image acquisition device, respectively. Their perspective direction vectors and The coordinate axes are parallel to each other, specifically: in, This represents the point cloud of the predicted missing regions. Point cloud representing the truly missing regions, Belonging to The point in the middle, Indicates the projected contour discriminator, input It is two-dimensional image data.

2. The method for generative adversarial point cloud completion network based on multi-view projection contours according to claim 1, characterized in that, Step S1 is as follows: S1.

1. Fix the rotating stage, image acquisition equipment, and plants without complex obstruction structures to ensure that the plants are within the visible field of view of the acquisition equipment; S1.

2. Obtain the transformation relationship between adjacent viewpoints through motion recovery structures, and extract the depth information of plants in the depth map using a mask of RGB images; S1.

3. Obtain complete point cloud data of plants without self-occlusion through multi-view stereo imaging.

3. The method for generative adversarial point cloud completion network based on multi-view projection contours according to claim 2, characterized in that: In step S1.3, 3D point cloud data for the current viewpoint is generated based on the intrinsic parameters of the depth image acquisition device and the depth image: in, , , These are the x, y, and z coordinates of the object in the point cloud coordinate system, respectively. These are the axial and radial focal lengths of the depth image acquisition device, respectively. These are the coordinates of the principal point in the depth map image. These are the horizontal and vertical pixel coordinates of the depth map, respectively.

4. The method for generative adversarial point cloud completion network based on multi-view projection contours according to any one of claims 1-3, characterized in that: Step S2 specifically includes: S2.

1. For a single plant point cloud without self-occlusion structure, each time a part of the leaf and stem point cloud is copied, twisted, rotated and pruned to obtain a new complete plant point cloud with self-occlusion structure. S2.

2. A complete plant consists of three main parts: stem, stalk, and leaves. A corresponding label value is assigned to the point cloud of each part. S2.

3. Repeat steps S2.1 and S2.2 above several times to obtain the required number of complete plant point cloud datasets.

5. The method for generative adversarial point cloud completion network based on multi-view projection contours according to claim 2 or 3, characterized in that: In step S3, by referring to the location and intrinsic parameter information of the image acquisition device when collecting data in step S1, the distance between the image acquisition device and the plant is determined. The location of the virtual image acquisition device in the point cloud coordinate system is constructed with this distance as the radius. Based on the viewpoint of each virtual image acquisition device, the occluded area (i.e., the real missing area point cloud) and the unoccluded area (i.e., the incomplete point cloud) under different viewpoints of the same plant are determined.

6. The method for generative adversarial point cloud completion network based on multi-view projected contours according to any one of claims 1-3, characterized in that: Step S6 specifically includes: S6.

1. Based on the multi-view projected contour image of the missing region under the current viewpoint, points in the predicted point cloud of the missing region that exceed the range of the projected contour are regarded as erroneously generated points, and their values ​​are set to 0 for constraint generation. S6.

2. Feed the constrained predicted point cloud into the point cloud discriminator and contour discriminator of the generative adversarial network for point cloud completion to train the adversarial loss and optimize the network parameters.

7. A system based on a generative adversarial point cloud completion network with multi-view projected contours, used to perform the method as described in any one of claims 1-6, characterized in that, Includes the following modules: Unoccluded plant point cloud acquisition module: Acquires plant point clouds without self-occlusion. The self-occluding plant point cloud data acquisition module performs non-rigid transformation on the plant point cloud, acquires plant point cloud data with self-occlusion based on the plant's biological structure, and adds labels to the plant parts corresponding to the point cloud classification. Multi-view contour projection image extraction module: Extracts multi-view contour projection images of each plant point cloud, and obtains the occluded area (i.e., the real missing area point cloud) and the unoccluded area (i.e., the incomplete point cloud) corresponding to the complete point cloud based on its positional relationship with the virtual image acquisition device. The feature encoding module for missing regions: inputs the incomplete point cloud into the multi-resolution feature encoder, fuses and establishes feature encodings for the missing regions; Missing Region Prediction Module: Input the feature encoding of the incomplete point cloud into the generator of the generative adversarial network for point cloud completion, and predict the missing regions corresponding to the point cloud. Adversarial training module: The prediction results are constrained by multi-view projection contours of real missing regions. The constrained point cloud is then fed into the point cloud discriminator and contour discriminator of the generative adversarial network for point cloud completion for adversarial training to optimize the network parameters.

Citation Information

Patent Citations

  • Method and equipment for determining morphological structure of rape group

    CN118864564A