Method and system for generative adversarial point cloud completion network based on multi-view projection contour
Through the multi-view contour constraint method based on the generative adversarial network, the problems of large system computing volume and low reconstruction efficiency in plant three-dimensional reconstruction are solved, and accurate plant reconstruction and high-precision phenotypic parameter measurement are achieved.
Patent Information
- Application Number
- CN202411660426.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-20
AI Technical Summary
The prior art has problems such as large system computing and low reconstruction efficiency in plant three-dimensional reconstruction. Especially when processing image information acquisition from multiple perspectives, it is impossible to collect three-dimensional information of the complete internal structure of the plant, resulting in accumulated data processing errors and inaccurate reconstruction results.
The point cloud completion method based on the multi-view contour constraint of a generative adversarial network is adopted to reconstruct the plants through contactless and non-destructive three-dimensional structure, predict the point cloud data of the missing area, and conduct adversarial training through the multi-view projection contour constraint generator and discriminator to optimize network parameters.
Accurate reconstruction of plants is achieved, the accuracy of plant phenotypic parameters measurement is improved, equipment requirements are simplified, anti-interference ability is enhanced, and the accuracy of reconstruction model is improved.
Smart Images

Figure CN119941972A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of agricultural plant three-dimensional phenotype reconstruction, and specifically relates to a plant point cloud completion method based on multi-view contour constraints of a generative adversarial network. Background Art
[0002] In recent years, three-dimensional digital technology has important application prospects in the fields of medicine, military aerospace, agriculture, etc. In particular, digital agriculture has become an important trend in the development of agriculture in my country. Plant three-dimensional reconstruction can obtain three-dimensional phenotypic characteristics under various actual conditions, including height, width, leaf inclination, leaf area and canopy volume. These characteristics not only reflect the genetic characteristics of crops, but also reflect the influence of their growth environment and field management. This technology has important application value in the fields of plant breeding, crop growth management and virtual visualization. Traditional plant three-dimensional phenotypic measurement mainly relies on manual measurement, which is limited by the accuracy of measurement tools and manpower consumption. It is not only time-consuming and labor-intensive, but also accompanied by the limitation of inaccurate measurement results. By constructing a three-dimensional model of a plant in a digital way, plant phenotypic information can be obtained quickly, accurately and efficiently.
[0003] The method of restoring structure from motion is a general method for reconstructing the three-dimensional model of plants. After selecting a fixed plant position, the camera is used to collect color images of the stationary plant at different viewing angles. By extracting image feature points and feature matching, and adjusting and optimizing the matching strategy, the structural data of the plant in three-dimensional space is obtained. However, the technical implementation links involved, such as feature extraction, feature matching, and optimization, are the reasons for the large amount of system computation and low reconstruction efficiency of this method. The acquisition of the three-dimensional structure of plants depends on depth information. The RGB-D camera is a three-dimensional sensor that combines a color image sensor and a depth image sensor. It extracts the depth information of the target point based on the TOF principle, i.e. the time of flight principle. By actively projecting laser onto the object to be measured, the camera sensor receives the laser projected onto the diffuse reflection of the rear end of the surface of the object to be measured, and the corresponding distance is calculated as the depth according to the time difference from the laser emission to the reception. This technical solution of actively projecting laser to obtain depth information has the advantages of strong anti-interference ability, high stability, and fast reconstruction speed. It has been widely used in target capture, three-dimensional navigation, three-dimensional reconstruction and other fields. In the actual data collection process, the RGB-D camera is fixed, and the three-dimensional point cloud data under different viewing angles is obtained by rotating the object to be measured, and the point cloud data of different viewing angles are spliced and fused to obtain a complete three-dimensional point cloud model of the object to be measured. In the three-dimensional reconstruction of plants, since leaves and stems have diverse distributions, sizes, positions and orientations, images at certain rotation angle intervals are collected to solve the reconstruction challenges brought by the diversity of leaves, which involves the processing of a large amount of point cloud data and is blind. However, information collection under a single viewing angle will cause the internal structure of the plant to be blocked by the structure of the plant itself, resulting in the loss of information on the internal structure of the plant; at the same time, due to the limitations of sensor accuracy and other reasons, the unobstructed area will lose some information to a certain extent, resulting in information loss. In order to ensure the accuracy of the three-dimensional phenotypic measurement of the plant, its internal structure needs to be predicted, completed and reconstructed to obtain complete and sufficient plant point cloud data and improve reconstruction efficiency and performance. Based on this, the present invention proposes a point cloud completion technology with multi-view contour constraints based on a generative adversarial network, which can be used for the measurement and evaluation of the three-dimensional phenotype of the plant. Summary of the invention
[0004] In order to overcome the problems existing in the prior art, the present invention proposes a method and system for predicting and completing the three-dimensional point cloud of a plant based on a generative adversarial point cloud completion network based on multi-view projection contours, so as to solve the problems such as the missing three-dimensional point cloud information of the internal structure of the plant and the missing data during the sensor acquisition process.
[0005] In order to solve the above technical problems, the technical solution of the present invention is:
[0006] A method for generating adversarial point cloud completion network based on multi-view projection contours, comprising the following steps:
[0007] S1. Collect the plant point cloud without self-occlusion of its internal structure;
[0008] S2. Perform non-rigid transformation on the plant point cloud, obtain a large amount of plant point cloud data with self-occlusion according to the biological structure of the plant, and add labels to the corresponding plant parts according to the classification of the point cloud;
[0009] S3. extract the multi-view contour projection image of each plant point cloud, and obtain the blocked area corresponding to the complete point cloud, i.e., the real missing area point cloud, and the unblocked area, i.e., the incomplete point cloud, according to the positional relationship between the image and the virtual camera;
[0010] S4. Input the incomplete point cloud into the multi-resolution feature encoder, fuse and establish feature encoding with missing areas;
[0011] S5. Send the feature encoding of the incomplete point cloud to the generator of the generative adversarial network for point cloud completion to predict the missing area corresponding to the point cloud;
[0012] S6. Use the multi-view projection contour of the real missing area to constrain the prediction result to ensure that it is generated within the valid range, and then input the constrained point cloud into the point cloud discriminator and contour discriminator of the generative adversarial network of point cloud completion for adversarial training to optimize network parameters.
[0013] Preferably, step S1 is specifically as follows:
[0014] S1.1. Fix the rotating stage, image acquisition equipment and plants without complex shielding structures to ensure that the plants are within the visible field of view of the acquisition equipment;
[0015] S1.2. Obtain the transformation relationship of adjacent viewpoints by restoring the structure from motion, and use the mask of the RGB image to extract the depth information of the plant part in the depth map;
[0016] S1.3. Obtain the complete point cloud data of the plant without complex occlusion structure through the principle of multi-view stereo imaging.
[0017] Preferably, step S1.3 specifically includes:
[0018] Generate 3D point cloud data of the current view based on the depth camera internal parameters and depth image:
[0019]
[0020] Among them, x depth ,y depth 、z depth are the x, y, and z coordinate values of the object in the point cloud coordinate system, respectively. xdepth 、f ydepth are the axial and radial focal lengths of the depth image acquisition device, respectively, and cxdepth 、c ydepth is the principal point coordinate of the depth map image, u depth 、v depth They are the horizontal and vertical pixel coordinates of the depth map, respectively.
[0021] Preferably, step S2 specifically comprises:
[0022] S2.1. For a single plant point cloud without self-occlusion structure, manually copy, twist, rotate, trim, etc. the point cloud of some leaves and stems each time to obtain a new complete plant point cloud with complex self-occlusion structure;
[0023] S2.2. The complete plant consists of three main parts: stem, bar and leaf. The point cloud of each part is assigned a corresponding label value;
[0024] S2.3. Repeat the above two steps to obtain the required number of complete plant point cloud data sets. The number of repetitions depends on the number of plant point clouds that need to be collected. For example, if three complete plant point clouds need to be collected, it needs to be repeated three times. The end time is when one complete plant point cloud is collected.
[0025] Preferably, in step S3, by referring to the camera position and internal reference information when collecting data in step S1, the distance between the camera and the plant is determined, and the multi-view virtual camera position in the point cloud coordinate system is constructed with the distance as the radius, and the obscured area, i.e., the real missing area point cloud and the unobstructed area, i.e., the incomplete point cloud, under different viewing angles of the same plant are determined according to each virtual camera viewing angle;
[0026] Preferably, step S4 specifically comprises:
[0027] S4.1. The incomplete point cloud input is downsampled using the IFPS algorithm, and the number of points in each point cloud is downsampled to 2048, 1024, and 512, representing three resolutions from high to low;
[0028] S4.2. Use CMLP (Combined Multi-Layer Perception) with multiple fully connected layers to extract features for point clouds of different resolutions;
[0029] S4.3. The features obtained at different resolutions are concatenated and input into the multi-layer perceptron MLP to extract and finally form a 1920-dimensional feature vector;
[0030] Preferably, step S5 specifically includes:
[0031] S5.1. The 1920-dimensional feature vector obtained by multi-resolution feature encoding is input into the generator of the generative adversarial network for point cloud completion. The generator passes the input feature vector through 4 fully connected linear layers to obtain feature vectors with dimensions of 1024, 512, 256, and 256 respectively, which can extract global features and local features at the same time;
[0032] S5.2. Pass the four features of different dimensions through different layers in the generator, and regard the feature dimension of 1024 as the first layer, the feature dimension of 512 as the second layer, and so on. The feature dimension of the first 256 is regarded as the third layer, and the feature dimension of the second 256 is regarded as the fourth layer.
[0033] S5.3. The 4th and 3rd layers output point clouds of size M1×3, which are concatenated to obtain a point cloud PC of size M2×3 primary , represents the low-resolution prediction value of the missing part, where:
[0034] 2M1=M2
[0035] Among them, M1 and M2 are the number of points for generating point clouds of different resolutions, and their settings are determined by the final M;
[0036] S5.4. The second layer outputs a point cloud of size 3M2×3, which is compared with the PC primary After stitching, we get a point cloud PC of size M×3 secondary , represents the medium-resolution prediction value of the missing part, where:
[0037] 4M2=M
[0038] Among them, M is the number of points to generate the final point cloud. The choice of M will affect the settings of M1 and M2;
[0039] S5.5. The first layer outputs a point cloud of size M×3, which is compared with the PC secondary After stitching, we get a point cloud PC with a size of 2M×3 detail , represents the high-resolution prediction value of the missing part;
[0040] Preferably, step S6 specifically comprises:
[0041] S6.1. Based on the multi-view projection contour image of the true value of the missing area under the current view, the points in the point cloud predicted by the missing area that are beyond the projection contour range are regarded as incorrectly generated points, and their values are set to 0 for constraint generation;
[0042] S6.2. Send the constrained predicted point cloud into the point cloud discriminator and contour discriminator in the generative adversarial network of point cloud completion for adversarial loss training to optimize network parameters.
[0043] Preferably, the CMLP, the generator and the discriminator described in step S4, step S5 and step S6 are jointly supervised trained by a multi-objective loss function, and the multi-objective loss function is as follows:
[0044] L=λ com L com +λ points L points +λ contour L contour
[0045] Among them, λ com , points , contour L com , L points , L coontour The weight, L com It represents the CD loss of the three resolutions of the predicted point cloud and the point cloud of the actual missing area (CD loss is usually used when comparing the actual value and predicted value of the point cloud, specifically Chamfer Distance), L points represents the binary cross entropy loss between the predicted point cloud and the real missing area point cloud, L contour It represents the binary cross entropy loss between the predicted multi-view projection contour of the point cloud and the multi-view projection contour of the real missing area point cloud, as follows:
[0046]
[0047] Among them, PC detaol 、PC primary 、PC secondary They represent the generated high, low, and medium resolution point clouds, respectively. α and β are the weights of the CD loss of low and medium resolution, respectively. PC gt represents the ground truth value of the real missing area, They represent the ground truth obtained by downsampling the ground truth of the missing area once and twice by IFPS (this is a method of downsampling point cloud data, the Chinese name is Iterative Farthest Point Sample, specifically Iterative Farthest Point Sample), represents the CD loss between two point clouds, D() represents the point cloud discriminator, G() represents the generator, and x j represents the incomplete point cloud belonging to the input, y i represents the point cloud of the real missing area, S represents the size of the data set,
[0048] L front , L side , L topThey represent the binary cross entropy loss of the projection contours of the predicted point cloud and the real missing area point cloud according to the three perspectives of the virtual camera position. Their perspective direction vectors are parallel to the x, y, and z coordinate axes. The specific details are:
[0049]
[0050] Among them, L front、side、top Indicates L front , L side , L top can be calculated according to this formula. Set1 represents the point cloud of the missing area generated by the prediction, and Set2 represents the point cloud of the real missing area. x and y belong to the points in Set1 and Set2 respectively. j ,y i and G() are defined as follows. In particular, D c () represents the projection contour discriminator, whose input I(y i )、I(x j ) is two-dimensional image data.
[0051] The present invention also discloses a system of a generative adversarial point cloud completion network based on multi-view projection contours, which is used to execute the above method and includes the following modules:
[0052] Plant point cloud acquisition module without self-occlusion: collects plant point clouds without self-occlusion structures;
[0053] Self-occluded plant point cloud data acquisition module: Perform non-rigid transformation on the plant point cloud, obtain the plant point cloud data with self-occlusion according to the biological structure of the plant, and add labels to the corresponding plant parts according to the classification of the point cloud;
[0054] Multi-view contour projection image extraction module: extract the multi-view contour projection image of each plant point cloud, and obtain the obscured area corresponding to the complete point cloud, i.e. the real missing area point cloud, and the unobstructed area, i.e. the incomplete point cloud, according to its positional relationship with the virtual image acquisition device;
[0055] Feature coding establishment module for missing areas: Input the incomplete point cloud into the multi-resolution feature encoder, fuse and establish feature coding with missing areas;
[0056] Missing area prediction module: The feature encoding of the incomplete point cloud is input into the generator of the generative adversarial network for point cloud completion to predict the missing area corresponding to the point cloud;
[0057] Adversarial training module: Use the multi-view projection contour of the real missing area to constrain the prediction results, and then send the constrained point cloud to the point cloud discriminator and contour discriminator of the generative adversarial network of point cloud completion for adversarial training to optimize network parameters.
[0058] The present invention has the following characteristics and beneficial effects:
[0059] The present invention adopts a generative adversarial multi-view contour constraint plant point cloud completion scheme to reconstruct the complete three-dimensional structure of the plant in a non-contact and non-destructive manner. In view of the current problems of severe occlusion in plant three-dimensional reconstruction and multi-view acquisition of image information, the inability to acquire the complete three-dimensional information of the internal structure of the plant, resulting in accumulated data processing errors and inaccurate reconstruction results, the present invention adopts a point cloud completion method to predict the missing area, without the need for tedious camera calibration, to achieve accurate reconstruction of the plant. At the same time, the present invention requires simple equipment, strong anti-interference ability, and high precision of the plant reconstruction model, which effectively improves the accuracy of plant phenotypic parameter measurement. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0061] Figure 1 This is a flow chart of a method for a generative adversarial point cloud completion network based on multi-view projection contours according to an embodiment of the present invention.
[0062] Figure 2 Schematic diagram of data collection and non-rigid transformation to generate different plants in an embodiment of the present invention.
[0063] Figure 3 It is a schematic diagram of placing a virtual camera according to a reference position to obtain an occlusion area according to an embodiment of the present invention.
[0064] Figure 4 This is a schematic diagram of the model structure of the multi-view contour constrained generative point cloud completion network MCCGPCN (Multi-view Contour Constraint Generative Point Cloud Completion Network) according to an embodiment of the present invention.
[0065] Figure 5 It is a schematic diagram of the structure of the generator of MCCGPCN according to an embodiment of the present invention.
[0066] Figure 6 Schematic diagram of the overall structure of the point cloud identifier and contour identifier of MCCGPCN in an embodiment of the present invention.
[0067] Figure 7This is an example of the result of plant point cloud completion according to an embodiment of the present invention.
[0068] Figure 8 This is a system block diagram of a generative adversarial point cloud completion network based on multi-view projection contours according to an embodiment of the present invention. DETAILED DESCRIPTION
[0069] The preferred embodiments of the present invention are described in detail below. It should be noted that the embodiments and features in the embodiments of the present invention can be combined with each other without conflict.
[0070] This embodiment provides a multi-view contour constraint plant point cloud completion method based on a generative adversarial network. Figure 1 As shown, the following steps are included:
[0071] S1. Collecting plant point clouds without self-occluding structures;
[0072] S2. Perform non-rigid transformation on the plant point cloud to obtain plant point cloud data with self-occlusion, and add labels to the plant parts corresponding to the classification of the point cloud;
[0073] S3. extract the multi-view contour projection image of each plant point cloud, and obtain the obstructed area corresponding to the complete point cloud, i.e., the real missing area point cloud, and the unobstructed area, i.e., the incomplete point cloud, according to the positional relationship between the image acquisition device and the plant point cloud;
[0074] S4. Input the incomplete point cloud into the multi-resolution feature encoder, fuse and establish feature encoding with missing areas;
[0075] S5. Input the feature encoding of the incomplete point cloud into the generator of the generative adversarial network for point cloud completion to predict the missing area corresponding to the point cloud;
[0076] S6. Use the multi-view projection contour of the real missing area to constrain the prediction result, and then send the constrained point cloud to the point cloud discriminator and contour discriminator of the generative adversarial network of point cloud completion for adversarial training to optimize the network parameters.
[0077] Specifically, by collecting plants whose leaves and stems do not obstruct their internal structures, we obtain complete plant 3D point cloud data samples, and use software to perform non-rigid transformation on the point cloud. Through cropping, copying, twisting and other operations, we generate different complete 3D point clouds that occlude each other inside the plant to construct a data set. Multi-view contour projection images are extracted from each plant point cloud to obtain the occluded and unoccluded areas of the complete point cloud. The point cloud of the unoccluded area is passed through a multi-resolution encoder to obtain its feature code, and the feature code is sent to the generator of MCCGPCN (Multi-view Contour ConstraintGenerative Point Cloud Completion Network). The predicted point cloud of the missing area is constrained by the multi-view projection contour image, and the constrained point cloud is sent to the discriminator of MCCGPCN for adversarial loss training to further optimize the generative adversarial network.
[0078] In this embodiment, it includes two parts: data set generation and a generative point cloud completion network with multi-view contour constraints.
[0079] 1) Dataset generation for plant 3D reconstruction:
[0080] Since there is no ready-made public dataset for the task of 3D point cloud prediction and completion, and the acquisition cost of a dataset with a large sample size is very expensive, the present invention collects a small amount of complete 3D point clouds of plants based on the method of multi-view 3D reconstruction, and combines plants with different structural structures, and uses the open source software Blender to perform non-rigid transformation on the point cloud to expand different plant point cloud samples to realize dataset generation. The specific implementation is as follows:
[0081] a) If Figure 2 As shown, fix the rotating platform and the camera, select a plant whose leaves and stems do not block the internal structure of the plant itself in the direction of the camera optical axis, and fix the plant on the rotating platform;
[0082] b) Using a fixed RGB-D camera to collect RGB images and depth images of the plant rotating at a fixed interval angle, using the mask of the RGB image to extract the depth information of the plant part, and obtaining the complete point cloud data of the single plant without complex occlusion structure through the principle of multi-view stereo imaging;
[0083] c) For the complete point cloud data of a single plant, the point cloud is non-rigidly transformed through the open source software Blender. The transformation methods include but are not limited to: manually selecting a portion of the point cloud of the leaf, bending, twisting, and stretching the portion of the point cloud according to the possible bending of the leaf under natural conditions, to obtain a new leaf point cloud with a shape different from the original data; manually selecting a portion of the point cloud of the leaf, trimming or removing the portion of the point cloud, to obtain a new leaf point cloud with a shape that is missing from the original; manually selecting any complete leaf, copying, rotating, and moving it to another part of the stem, to obtain a new complete point cloud data with an increased number of leaves. Based on the non-rigid transformation of the complete point cloud of a single plant collected, a large number of complete point clouds of plants with different structural structures are generated;
[0084] d) Using the open source software CloudCompare, label values are added one by one to the generated complete point cloud of the plant, and each label value corresponds to a different part of the plant;
[0085] e) If Figure 3 As shown, according to the camera position when collecting a single plant, the plant point cloud coordinate system is used as a reference to establish a virtual camera position under the same collection radius, and the observation angle is determined according to the direction of the camera optical axis. According to each virtual camera angle, the blocked area, i.e., the real missing area point cloud, and the unblocked area, i.e., the incomplete point cloud, under different viewing angles of the same plant are determined;
[0086] f) Divide the generated data set into training set and validation set according to a certain ratio.
[0087] 2) Multi-view contour-constrained generative point cloud completion network MCCGPCN:
[0088] The generative point cloud completion network MCCGPCN for multi-view contour constraints of plants consists of two parts: feature extraction network and multi-view contour constraint point cloud generation network. The network structure is as follows: Figure 4 As shown in the figure, the feature extraction network is a multi-resolution perceptron network with three CMLP (Combined Multi Layer Perception) in parallel, and the multi-view contour constraint point cloud generation network is a generative adversarial network for multi-view contour point cloud prediction generation results;
[0089] Input the real missing area point cloud and incomplete point cloud of the same plant, use iterative farthest point sampling IFPS (Iterative Farthest Point Sampling) to downsample the point cloud to obtain point clouds with the number of points [2048, 1024, 512], which correspond to 3 different resolutions. While reducing the amount of data processing, it mainly retains the geometric information of the point cloud without causing the loss of key information. The point clouds of different resolutions are passed through parallel multi-layer perceptrons CMLP, and each CMLP is composed of a multi-layer perceptron MLP of [64-128-256-512-1024] dimensions to obtain the features of point clouds of 3 different resolutions. These 3 features are then cascaded and passed through MLP to obtain a 1920-dimensional point cloud feature vector F. This feature contains the global features and local detail features of the complete plant point cloud;
[0090] The point cloud feature vector F is input into the generator of the multi-view contour constraint point cloud generation network MCCGPCN. The generator structure is as follows: Figure 5 As shown. The generator passes the input feature vector through 4 fully connected linear layers to obtain feature vectors with dimensions of 1024, 512, 256, and 256, respectively, which are global and local rough and fine features. The features of 4 different dimensions are passed through different layers in the generator, and the feature dimension of 1024 is regarded as the first layer, and the feature dimension of 512 is regarded as the second layer, and so on. The 4th and 3rd layers output point clouds of size M1×3, which are spliced to obtain a point cloud PC of size M2×3 primary , represents the low-resolution prediction value of the missing part, where:
[0091] 2M1=M2
[0092] The second layer outputs a point cloud of size 3M2×3, which is compared with the PC primary After stitching, we get a point cloud PC of size M×3 secondary , represents the medium-resolution prediction value of the missing part, where:
[0093] 4M2=M
[0094] The first layer outputs a point cloud of size M×3, which is combined with the PC secondary After stitching, we get a point cloud PC with a size of 2M×3 detail , represents the high-resolution prediction value of the missing part;
[0095] According to the multi-view projection contour image of the true value of the missing area under the current view, the points in the point cloud predicted by the missing area that exceed the projection contour range are regarded as incorrectly generated points, and the values of these points are set to 0 to achieve constraint generation. The constrained predicted point cloud is sent to the point cloud discriminator and contour discriminator in the generative adversarial network of point cloud completion for adversarial loss training to optimize network parameters. The structures of the point cloud discriminator and contour discriminator are as follows: Figure 6 As shown;
[0096] The multi-objective loss function for supervising CMLP, generator and discriminator is as follows:
[0097] L=λ com L com +λ points L points +λ contour L contour
[0098] Among them, L com It represents the CD loss of the predicted point cloud with three resolutions and the point cloud in the real missing area. points represents the binary cross entropy loss between the predicted point cloud and the real missing area point cloud, L contour It represents the binary cross entropy loss between the predicted multi-view projection contour of the point cloud and the multi-view projection contour of the real missing area point cloud, as follows:
[0099]
[0100] Among them, PC gt represents the ground truth value of the real missing area, They represent the ground truth values obtained by downsampling the ground truth values of the missing areas once and twice by IFPS, respectively. represents the CD loss between two point clouds, D() represents the point cloud discriminator, G() represents the generator, and x j represents the incomplete point cloud belonging to the input, y i represents the point cloud of the real missing area, S represents the size of the dataset, and L front , L side , L top They represent the binary cross entropy loss of the projection contours of the predicted point cloud and the point cloud of the real missing area according to the three perspectives of the virtual camera position, and their perspective direction vectors are parallel to the x, y, and z coordinate axes. Specifically:
[0101]
[0102]
[0103] Among them, Set1 represents the predicted missing area point cloud, Set2 represents the real missing area point cloud, x and y belong to the points in Set1 and Set2 respectively, x j ,y i and D c The definition of () is as follows: D c () represents the projection contour discriminator, whose input I(y i )、I(x j ) is two-dimensional image data.
[0104] like Figure 8 As shown, this embodiment discloses a system of a generative adversarial point cloud completion network based on multi-view projection contours, which is used to execute the above method embodiment, and includes the following modules:
[0105] Plant point cloud acquisition module without self-occlusion: collects plant point clouds without self-occlusion structures;
[0106] Self-occluded plant point cloud data acquisition module: Perform non-rigid transformation on the plant point cloud, obtain the plant point cloud data with self-occlusion according to the biological structure of the plant, and add labels to the corresponding plant parts according to the classification of the point cloud;
[0107] Multi-view contour projection image extraction module: extract the multi-view contour projection image of each plant point cloud, and obtain the obscured area corresponding to the complete point cloud, i.e. the real missing area point cloud, and the unobstructed area, i.e. the incomplete point cloud, according to its positional relationship with the virtual image acquisition device;
[0108] Feature coding establishment module for missing areas: Input the incomplete point cloud into the multi-resolution feature encoder, fuse and establish feature coding with missing areas;
[0109] Missing area prediction module: The feature encoding of the incomplete point cloud is input into the generator of the generative adversarial network for point cloud completion to predict the missing area corresponding to the point cloud;
[0110] Adversarial training module: Use the multi-view projection contour of the real missing area to constrain the prediction results, and then send the constrained point cloud to the point cloud discriminator and contour discriminator of the generative adversarial network of point cloud completion for adversarial training to optimize network parameters.
[0111] For other contents of this embodiment, please refer to the above method embodiment.
[0112] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions and variations of these embodiments are made without departing from the principles and spirit of the present invention, and still fall within the scope of protection of the present invention.
Claims
1. A method for generative adversarial point cloud completion network based on multi-view projection contours, characterized in that: The following steps are involved: S1. Collecting plant point clouds without self-occluding structures; S2. Perform non-rigid transformation on the plant point cloud to obtain plant point cloud data with self-occlusion, and add labels to the plant parts corresponding to the classification of the point cloud; S3. extract the multi-view contour projection image of each plant point cloud, and obtain the obstructed area corresponding to the complete point cloud, i.e., the real missing area point cloud, and the unobstructed area, i.e., the incomplete point cloud, according to the positional relationship between the image acquisition device and the plant point cloud; S4. Input the incomplete point cloud into the multi-resolution feature encoder, fuse and establish feature encoding with missing areas; S5. Input the feature encoding of the incomplete point cloud into the generator of the generative adversarial network for point cloud completion to predict the missing area corresponding to the point cloud; S6. Use the multi-view projection contour of the real missing area to constrain the prediction result, and then send the constrained point cloud to the point cloud discriminator and contour discriminator of the generative adversarial network of point cloud completion for adversarial training to optimize the network parameters.
2. The method of generative adversarial point cloud completion network based on multi-view projection contour according to claim 1, characterized in that: Step S1 is specifically as follows: S1.
1. Fix the rotating stage, image acquisition equipment and plants without complex shielding structures to ensure that the plants are within the visible field of view of the acquisition equipment; S1.
2. Obtain the transformation relationship between adjacent viewpoints through motion recovery structure, and use the mask of RGB image to extract the depth information of plants in the depth map; S1.
3. Obtain complete point cloud data of plants without self-occlusion through multi-view stereo imaging.
3. The method of claim 2, wherein: In step S1.3, the three-dimensional point cloud data of the current viewing angle is generated according to the internal parameters of the depth image acquisition device and the depth image: Among them, x depth ,y depth 、z depth are the x, y, and z coordinate values of the object in the point cloud coordinate system, respectively. xdepth 、f ydepth are the axial and radial focal lengths of the depth image acquisition device, c xdepth 、c ydepth is the principal point coordinate of the depth map image, u depth 、v depth They are the horizontal and vertical pixel coordinates of the depth map, respectively.
4. The method for generative adversarial point cloud completion network based on multi-view projection contour according to any one of claims 1 to 3, characterized in that: Step S2 specifically includes: S2.
1. For a single plant point cloud without self-occlusion structure, a new complete plant point cloud with self-occlusion structure is obtained by copying, twisting, rotating and trimming the point cloud of part of the leaves and stems each time; S2.
2. The complete plant consists of three main parts: stem, bar and leaf. The point cloud of each part is assigned a corresponding label value; S2.
3. Repeat the above steps S2.1 and S2.2 several times to obtain the required number of complete plant point cloud datasets.
5. The method of generative adversarial point cloud completion network based on multi-view projection contour according to claim 2 or 3, characterized in that: In step S3, by referring to the position of the image acquisition device and the internal reference information when collecting data in step S1, the distance between the image acquisition device and the plant is determined, and the multi-perspective virtual image acquisition device position in the point cloud coordinate system is constructed with the distance as the radius. According to the perspective of each virtual image acquisition device, the obscured area, i.e., the real missing area point cloud and the unobstructed area, i.e., the incomplete point cloud, of the same plant at different perspectives are determined.
6. The method for generative adversarial point cloud completion network based on multi-view projection contour according to any one of claims 1 to 3, characterized in that: Step S4 specifically includes: S4.
1. The incomplete point cloud input is downsampled using the IFPS algorithm, and the number of points in each point cloud is downsampled to 2048, 1024, and 512, representing three resolutions from high to low; S4.
2. Extract features using CMLP with multiple fully connected layers for point clouds of different resolutions; S4.
3. The features obtained at different resolutions are concatenated and input into a multi-layer perceptron for extraction, ultimately forming a 1920-dimensional feature vector.
7. The method of claim 6, wherein: Step S5 specifically includes: S5.
1. The 1920-dimensional feature vector obtained by multi-resolution feature encoding is input into the generator of the generative adversarial network for point cloud completion, and the generator passes the input feature vector through 4 fully connected linear layers to obtain feature vectors with dimensions of 1024, 512, 256, and 256 respectively; S5.
2. The feature vectors of four different dimensions are passed through different layers in the generator, and the feature vector with a feature dimension of 1024 is regarded as the first layer, the feature vector with a feature dimension of 512 is regarded as the second layer, the feature vector with a feature dimension of the first 256 is regarded as the third layer, and the feature vector with a feature dimension of the second 256 is regarded as the fourth layer; S5.
3. The 4th and 3rd layers output point clouds of size M1×3, which are concatenated to obtain a point cloud PC of size M2×3 primary , represents the low-resolution prediction value of the missing part, where: 2M1=M2 Among them, M1 and M2 are the number of points used to generate point clouds of different resolutions; S5.
4. The second layer outputs a point cloud of size 3M2×3, which is compared with the PC primary After stitching, we get a point cloud PC of size M×3 secondary , represents the medium-resolution prediction value of the missing part, where: 4M2=M Where M is the number of points used to generate the final point cloud; S5.
5. The first layer outputs a point cloud of size M×3, which is compared with the PC secondary After stitching, we get a point cloud PC with a size of 2M×3 detail , represents the high-resolution prediction value of the missing part.
8. The method for generative adversarial point cloud completion network based on multi-view projection contour according to any one of claims 1 to 3, characterized in that: Step S6 specifically includes: S6.
1. Based on the multi-view projection contour image of the true value of the missing area at the current view, the points in the point cloud predicted from the missing area that are beyond the projection contour range are regarded as incorrectly generated points, and their values are set to 0 for constraint generation; S6.
2. Send the constrained predicted point cloud to the point cloud discriminator and contour discriminator of the generative adversarial network for point cloud completion for adversarial loss training to optimize the network parameters.
9. The method of claim 7, wherein: The CMLP, generator and discriminator are jointly supervised trained through a multi-objective loss function, which is as follows: L=λ com L com +λ points L points +λ contour L contour Among them, λ com , points , contour L com , L points , L coontour The weight, L com It represents the CD loss of the predicted point cloud with three resolutions and the real missing area point cloud, L points represents the binary cross entropy loss between the predicted point cloud and the real missing area point cloud, L contour It represents the binary cross entropy loss between the predicted multi-view projection contour of the point cloud and the multi-view projection contour of the real missing area point cloud, as follows: Among them, PC detail 、PC primary 、PC secondary They represent the generated high, low, and medium resolution point clouds, respectively. α and β are the weights of the CD loss of low and medium resolution, respectively. PC gt represents the ground truth value of the real missing area, They represent the ground truth values obtained by downsampling the ground truth values of the missing areas once and twice by IFPS, respectively. represents the CD loss between two point clouds, i = 1, 2, 3, D() represents the point cloud discriminator, G() represents the generator, x j represents the incomplete point cloud belonging to the input, y i represents the point cloud of the real missing area, S represents the size of the dataset, and L front , L side , L top They represent the binary cross entropy loss of the projection contours of the predicted point cloud and the point cloud of the real missing area according to the three perspectives of the virtual image acquisition device position, and their perspective direction vectors are parallel to the x, y, and z coordinate axes, specifically: Among them, Set1 represents the point cloud of the missing area generated by prediction, Set2 represents the point cloud of the real missing area, x and y belong to the points in Set1 and Set2 respectively, D c () represents the projection contour discriminator, input I(y i )、I(x j ) is two-dimensional image data.
10. A system for a generative adversarial point cloud completion network based on multi-view projection contours, for executing the method according to any one of claims 1 to 9, characterized in that: Includes the following modules: Plant point cloud acquisition module without self-occlusion: collects plant point clouds without self-occlusion structures; Self-occluded plant point cloud data acquisition module: Perform non-rigid transformation on the plant point cloud, obtain the plant point cloud data with self-occlusion according to the biological structure of the plant, and add labels to the corresponding plant parts according to the classification of the point cloud; Multi-view contour projection image extraction module: extract the multi-view contour projection image of each plant point cloud, and obtain the obscured area corresponding to the complete point cloud, i.e. the real missing area point cloud, and the unobstructed area, i.e. the incomplete point cloud, according to its positional relationship with the virtual image acquisition device; Feature coding establishment module for missing areas: Input the incomplete point cloud into the multi-resolution feature encoder, fuse and establish feature coding with missing areas; Missing area prediction module: The feature encoding of the incomplete point cloud is input into the generator of the generative adversarial network for point cloud completion to predict the missing area corresponding to the point cloud; Adversarial training module: Use the multi-view projection contour of the real missing area to constrain the prediction results, and then send the constrained point cloud to the point cloud discriminator and contour discriminator of the generative adversarial network of point cloud completion for adversarial training to optimize network parameters.
Citation Information
Patent Citations
Tree point cloud completion method based on deep learning
CN114491697A
Three-dimensional point cloud reconstruction method and system based on point cloud completion technology
CN114612619A
Multi-scale greenhouse plant point cloud completion method based on generative adversarial network inverse mapping
CN115439490A
Blade point cloud completion method based on multilayer attention fusion and edge structure guidance
CN118396895A
Method and equipment for determining morphological structure of rape group
CN118864564A
Cited By
Blade point cloud reconstruction method and system based on projection constraint and hybrid supervision
CN121582516A
A leaf point cloud reconstruction method and system based on projection constraint and mixed supervision
CN121582516B