Method for obtaining plant parameters and use thereof
Patent Information
- Application Number
- CN202411007697.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-07-25
AI Technical Summary
但是,通用的基于深度学习的多任务点云分割网络存在实时可处理点数相对较少,网络参数量大,训练时间较长等问题,不能直接用于田间玉米植株的语义与实例分割任务
[0004]本申请的目的在于解决上述问题中的至少之一。
Smart Images

Figure CN118968061B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural information technology, specifically to a method for obtaining plant parameters and its application. Background Technology
[0002] Stable and rapid detection and segmentation of plants are prerequisites for agricultural autonomous robots to perform agricultural tasks, assess crop status, and conduct crop phenotypic research. Semantic segmentation of point clouds refers to dividing each point in the point cloud into different semantic categories. Instance segmentation of point clouds, based on semantic segmentation, further divides each point in the point cloud into different individuals with different semantic meanings. For unmanned farms, the map processing requirements vary depending on the application scenario. For example, in intelligent agricultural machinery operation navigation tasks, semantic segmentation of the farmland point cloud map is sufficient to find passable areas and obstacles for the intelligent agricultural machinery. However, in farmland crop status assessment tasks, instance segmentation of the crops planted in the field is necessary to meet the parameter detection requirements of individual crops. Therefore, multi-task point cloud segmentation of farmland point cloud maps, including semantic segmentation and instance segmentation, is essential.
[0003] Scholars have researched semantic segmentation and instance segmentation of crops based on 3D point clouds, which can be divided into methods based on manually designed features and methods based on deep learning. Methods based on manually designed features represent global and local information using manually designed feature descriptors, which are then used for point cloud segmentation and classification tasks. While achieving a certain level of accuracy, these methods suffer from poor generalization and low intelligence. In recent years, with the powerful feature extraction capabilities demonstrated by deep learning technology in the image domain, more and more researchers have begun to apply deep learning to point cloud segmentation and classification tasks. Deep learning-based methods can learn high-level latent features in point cloud data structures through neural networks based on existing datasets, achieving high accuracy in point cloud segmentation tasks and have been widely applied in tasks such as autonomous driving and robot navigation. However, general-purpose deep learning-based multi-task point cloud segmentation networks suffer from problems such as a relatively small number of points that can be processed in real time, a large number of network parameters, and long training times, making them unsuitable for direct application to semantic and instance segmentation tasks of corn plants in the field. Therefore, it is necessary to develop new segmentation networks. Summary of the Invention
[0004] The purpose of this application is to solve at least one of the above-mentioned problems.
[0005] Therefore, this application constructs a multi-task point cloud segmentation network (MaizeNet) for field plants, including semantic segmentation and instance segmentation, and extracts planting parameters of plants based on the segmentation results, such as corn at different growth stages, to provide an information basis for agricultural autonomous robots to perform agricultural tasks, evaluate crop status and crop phenotype.
[0006] Therefore, in a first aspect, the present invention proposes a method for obtaining plant parameters. According to an embodiment of the present invention, the method includes: 1) acquiring point cloud data of the plant to obtain a point cloud map; 2) preprocessing the point cloud map of the plant to obtain a preprocessed point cloud map, the preprocessing including filtering and downsampling; 3) performing data annotation and data augmentation on the preprocessed point cloud map to obtain a point cloud dataset; 4) performing semantic segmentation and instance segmentation on the point cloud dataset to obtain a point cloud model; 5) extracting plant parameters from the point cloud model according to point cloud processing techniques, wherein the parameters include at least one of height, density, row spacing, and plant spacing. The method according to an embodiment of the present invention obtains a point cloud model by performing semantic segmentation and instance segmentation on the point cloud dataset obtained after data augmentation, and obtains plant parameters based on the obtained point cloud model, such as planting parameters like height, density, row spacing, and plant spacing.
[0007] In a second aspect, the present invention provides a device for acquiring plant parameters. According to an embodiment of the present invention, the device includes: a point cloud map acquisition module, used to collect point cloud data of the plant to obtain a point cloud map;
[0008] A point cloud map preprocessing module, connected to the point cloud map acquisition module, is used to preprocess the point cloud map of the plant to obtain a preprocessed point cloud map. The preprocessing includes filtering and downsampling. A point cloud dataset acquisition module, connected to the point cloud map preprocessing module, is used to perform data annotation and data augmentation on the preprocessed point cloud map to obtain a point cloud dataset. A second segmentation processing module, connected to the point cloud dataset acquisition module, is used to perform a second segmentation processing on the point cloud dataset to obtain a point cloud model. A plant parameter extraction module, connected to the second segmentation processing module, is used to extract plant parameters from the point cloud model according to point cloud processing technology. The parameters include at least one of height, density, row spacing, and plant spacing. The device according to the embodiment of the present invention can effectively obtain the planting parameters of the plant, providing support for intelligent machine operation.
[0009] In a third aspect, the present invention provides an electronic device. According to an embodiment of the invention, the electronic device includes a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program to implement the method for obtaining plant parameters as described in the first aspect. Alternatively, embodiments of this application also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the method of the above-described method embodiment.
[0010] In a fourth aspect, the present invention provides a computer-readable storage medium. According to an embodiment of the invention, the computer-readable storage medium is used to store a computer program; the computer program causes a computer to perform the method for obtaining plant parameters described in the first aspect above.
[0011] In other words, when implemented using software, it can be implemented wholly or partially as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in accordance with this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer instructions can be transferred from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0012] Those skilled in the art will recognize that the modules and algorithmic steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0013] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0014] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0015] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a diagram of the main structure of the MaizeNet network according to an embodiment of the present invention;
[0018] Figure 2 This describes the data processing flow inside the encoder in the main structure of the MaizeNet network according to an embodiment of the present invention.
[0019] Figure 3 This is a structural diagram of the local spatial coding unit within the encoder of the MaizeNet network main structure according to an embodiment of the present invention;
[0020] Figure 4 This is a structural diagram of the self-attention unit within the encoder of the MaizeNet network main structure according to an embodiment of the present invention;
[0021] Figure 5 This is a structural diagram of the feature fusion layer according to an embodiment of the present invention;
[0022] Figure 6 A schematic diagram of different types of cornfields according to an embodiment of the present invention;
[0023] Figure 7 This is a schematic diagram of the random midpoint displacement method according to an embodiment of the present invention;
[0024] Figure 8 This is a schematic diagram illustrating the extraction of corn plant height according to an embodiment of the present invention;
[0025] Figure 9 This is a schematic diagram illustrating the extraction of row spacing and plant spacing for maize plants according to an embodiment of the present invention.
[0026] Figure 10 This is a schematic diagram of the device for obtaining plant parameters according to an embodiment of the present invention;
[0027] Figure 11 This is a schematic diagram of the process for obtaining plant parameters according to an embodiment of the present invention. Detailed Implementation
[0028] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein. In embodiments of this application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0030] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0031] In this document, the term “optionally” generally means that an event or condition described below may, but may not, occur, and the description includes both cases in which the event or condition occurs and cases in which the event or condition does not occur.
[0032] The endpoints and any values of the ranges disclosed herein are not limited to the precise ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, the endpoint values of the various ranges, the endpoint values of the various ranges and individual point values, and individual point values can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed herein.
[0033] This application constructs a multi-task maize plant point cloud segmentation network—MaizeNet—based on a self-attention mechanism. Local features are extracted through local spatial coding and a self-attention mechanism, and semantic and instance features are fused through feature fusion. Secondly, a field maize point cloud map (a 32-line vehicle-mounted LiDAR is mounted on top of an agricultural robot, tilted downwards at a certain angle. An industrial control computer receives the point cloud data via UDP and performs point cloud registration using a laser odometry or RTK-GNSS to obtain the field maize point cloud map) undergoes data preprocessing and manual segmentation to obtain maize plant point clouds. A field maize point cloud dataset is then obtained through data augmentation, and MaizeNet is trained and tested using this augmented dataset. Finally, the maize plant point cloud is segmented using MaizeNet, and planting parameters such as plant height, density, row spacing, and plant spacing are extracted based on the results. The specific implementation is illustrated in the following examples, which can be combined with each other. Similar or identical concepts or processes may not be described again in some embodiments.
[0034] In some specific embodiments, this invention proposes a method for obtaining plant parameters, the method comprising: 1) collecting point cloud data of the plant to obtain a point cloud map; 2) preprocessing the point cloud map of the plant to obtain a preprocessed point cloud map, the preprocessing including filtering and downsampling; 3) performing data annotation and data augmentation on the preprocessed point cloud map to obtain a point cloud dataset; 4) performing semantic segmentation and instance segmentation on the point cloud dataset to obtain a point cloud model; 5) extracting plant parameters from the point cloud model according to point cloud processing techniques, wherein the parameters include at least one of height, density, row spacing, and plant spacing. According to the method of this invention, a point cloud model is obtained by performing semantic segmentation and instance segmentation on the point cloud dataset obtained after data augmentation, and plant parameters are obtained based on the obtained point cloud model, such as planting parameters like height, density, row spacing, and plant spacing.
[0035] According to embodiments of the present invention, the above method may further include at least one of the following additional technical features:
[0036] According to an embodiment of the present invention, the point cloud data of the plants is collected using a three-dimensional lidar. In some specific embodiments, a vehicle-mounted lidar (32 lines) is mounted on the top of an agricultural robot, tilted downwards at a certain angle during installation. After the industrial control computer receives the point cloud data via UDP protocol, it performs point cloud registration using a laser odometry or RTK-GNSS to obtain a field corn point cloud map.
[0037] According to an embodiment of the present invention, the point cloud data of the plant includes point cloud data of any growth stage of the plant. Those skilled in the art will understand that the type of plant is not particularly limited and can be any type of plant, such as corn, sorghum, etc.
[0038] According to an embodiment of the present invention, the filtering process is a Gaussian filtering process. Therefore, the Gaussian filtering method is used to filter out some noise points and outliers.
[0039] According to an embodiment of the present invention, the downsampling process is voxel downsampling, wherein the voxel side length is set to 0.02m. Thus, the voxel downsampling method is applied to process the point cloud data.
[0040] For example, the data annotation includes at least one of xyz coordinates, semantic labels, and instance labels. For instance, in a specific implementation, the inventors write a program to save the annotated data as a txt file, which sequentially contains the xyz coordinates, semantic labels, and instance labels for each point within a sub-interval. The semantic labels are represented by 0 and 1, respectively, representing soil and plant; the instance labels start from 0 and include soil instances and plant instances. Since the number of soil instances is 1, the total number of instances is the number of plants plus 1.
[0041] According to an embodiment of the present invention, the data augmentation is performed using the random midpoint shift method. When the dataset is small, the results of the deep learning network may overfit the training data, reducing the network's generalization ability. The data augmentation method can expand the amount of training data by processing existing data, thereby improving the network's accuracy while maintaining its generalization ability.
[0042] For example, refer to Figure 6 The purpose of data augmentation was to generate a 5m × 5m point cloud of corn in a field. First, based on common corn planting types in the data collection environment, the collected farmland maps were divided into three categories: standard, aisle, and random. The standard type represents farmland where corn is planted with specified row and plant spacing; the aisle type represents a scene where the data collection vehicle travels along a road between two farmlands, with corn plants on either side of the road planted with specified row and plant spacing; the random type represents farmland where corn plants are randomly planted throughout the entire field. Standard farmland was the most common, followed by aisle farmland, while random farmland was the least common. (Reference) Figure 7 The random midpoint displacement method, a fractal interpolation algorithm, was used to randomly generate soil point clouds. This method requires only a few parameters to generate complex surfaces and is characterized by ease of implementation, recursion, and high resolution. The random midpoint displacement method is further subdivided into the triangle midpoint displacement method and the square midpoint displacement method. Finally, planting points for each corn plant were generated, and corn plants were randomly selected and placed on the soil surface. The type of cornfield determined the selection method for planting points, with row spacing of 0.6m, 0.65m, 0.7m, and 0.75m, and plant spacing of 0.5m, 0.55m, and 0.6m. For each planting point, a small random error was added to simulate planting point deviations during plant growth. Additionally, to simulate missing plants, there was a 1% probability that no corn plant would be placed at each planting point. Since corn plants planted simultaneously in actual cornfields can vary in height, plants at similar growth stages were selected to avoid significant height differences. 15,000 cornfield data points, each 5m x 5m in size, were generated using data augmentation methods.
[0043] According to an embodiment of the present invention, the second segmentation process is performed using the MaizeNet segmentation network. In this application, the inventors construct a multi-task point cloud segmentation network (MaizeNet), which includes semantic segmentation and instance segmentation, and extracts plant planting parameters based on the segmentation results, providing an information basis for agricultural autonomous robots to perform agricultural tasks, evaluate crop status, and assess crop phenotypes.
[0044] According to an embodiment of the present invention, the MaizeNet segmentation network includes a feature fusion layer.
[0045] According to an embodiment of the present invention, the MaizeNet segmentation network further includes an encoder and a decoder.
[0046] According to an embodiment of the present invention, the number of encoders and decoders is 3 to 6, preferably 4.
[0047] According to an embodiment of the present invention, the encoder includes a local spatial encoding, a self-attention unit, and a random downsampling unit.
[0048] According to an embodiment of the present invention, the decoder includes semantic branches and instance branches.
[0049] According to an embodiment of the present invention, the feature fusion layer is used for information interaction between the instance segmentation feature matrix and the semantic segmentation feature matrix. Those skilled in the art will understand that local feature extraction is crucial for high-precision point cloud segmentation, and local feature extraction methods incorporating self-attention mechanisms significantly improve segmentation accuracy. Furthermore, methods that combine semantic and instance features based on feature fusion can improve the segmentation accuracy of both tasks.
[0050] For example, see Figure 1MaizeNet primarily consists of an encoder, a decoder, and a feature fusion layer. The entire network employs skip connections to fully utilize network feature information. MaizeNet has four encoders, whose main functions are to extract local features and perform downsampling. A Transformer-based self-attention unit is incorporated into the local feature extraction process, which assigns attention weights to points based on the feature distribution of their neighbors. After local feature extraction, the feature map is randomly downsampled, reducing the number of points to one-quarter of the input number each time. The decoder has two branches: semantic and instance, processing semantic and instance features respectively. It upsamples the local features of the point cloud output from the encoder and concatenates them with local features at different resolutions. The feature fusion layer improves the segmentation accuracy of both tasks by combining the semantic and instance features output from the decoder. After the feature fusion layer, the semantic branch outputs a semantic prediction for each point, and the class with the highest probability is the semantic label for that point. The instance prediction branch obtains instance labels through MeanShift clustering.
[0051] According to an embodiment of the present invention, the information interaction includes the following operations:
[0052] (1) Input P sem With P ins The nearest neighbor feature matrix is obtained by the KNN method. and
[0053] (2) Attention scores are calculated using the self-attention method. and
[0054] (3) Calculate attention score and The weighted characteristic matrix F of the characteristic matrix of the nearest neighbor points of the other branch si With F is ;
[0055] (4) Based on the F si With P sem F is With P ins Obtain O through a fully connected network ins With O sem .
[0056] According to an embodiment of the present invention, F si and F is It is obtained through the following formula: Where K is the number of nearest neighbors.
[0057] According to an embodiment of the present invention, the total loss function is calculated using the following formula: L = αL c +βL v , where α is the weight coefficient of the semantic segmentation loss function; β is the weight coefficient of the instance segmentation loss function.
[0058] According to an embodiment of the present invention, the weight coefficient of the semantic segmentation loss function is 0.3 to 0.7, preferably 0.5.
[0059] According to an embodiment of the present invention, the weight coefficient of the instance segmentation loss function is 0.3 to 0.7, preferably 0.5.
[0060] According to an embodiment of the present invention, the semantic segmentation loss function is the standard cross-entropy loss function.
[0061] According to an embodiment of the present invention, the instance segmentation loss function is an instance embedding loss function.
[0062] According to an embodiment of the present invention, the standard cross-entropy loss function is calculated using the following formula: Among them, L c This represents the semantic segmentation loss function. p i semantic tags, Indicates network prediction p i Let N represent the probability of a point on the ground, and N represent the number of points, with i representing the i-th point. According to a specific embodiment of the invention, the ground surface includes soil and debris on the soil other than the plants, such as weeds.
[0063] According to an embodiment of the present invention, the instance embedding loss function is obtained by the following formula:
[0064]
[0065] Among them, L v L represents the instance segmentation loss function. var L represents the loss function for embedding identical instances. diff L represents the loss function that excludes different instances. reg Let N represent the regularization term, I represent the number of real instances, and N represent the number of instances. t u represents the number of points contained in the t-th instance. t Let e represent the average embedding of the t-th instance. i Denotes the embedding of the i-th point; δ s δ represents the maximum distance allowed for point embedding within the same instance. d Represents the maximum average embedding range across different instances; ||·||1 indicates the calculation of L1 distance; [x] + This indicates taking the maximum value in (0, x).
[0066] According to an embodiment of the present invention, the height is the difference between the height of the highest point and the lowest point of the plant.
[0067] According to an embodiment of the present invention, the highest point is the highest point of the leaf or the highest point of the tassel.
[0068] According to an embodiment of the present invention, the height is determined by the formula The calculation yielded, where, This represents the height of the plant at coordinates (x, y). This represents the height of the highest point of the plant at coordinates (x, y). This represents the height of the lowest point of the plant at coordinates (x, y).
[0069] According to an embodiment of the present invention, the density is calculated based on the area of the plot where the plant is located and the number of plants.
[0070] According to an embodiment of the present invention, the density is calculated using the formula ρ = N / S, where ρ represents density, N represents the number of plants, and S represents the area of the plot.
[0071] According to an embodiment of the present invention, the row spacing is calculated from the first plant in each row to the last plant in each row.
[0072] According to an embodiment of the present invention, for a plant with coordinates (x, y), the row spacing of the plant is determined based on the three plants closest to it in the right row, and the coordinates of the plant are (x+1, y), (x+1, y+1) and (x+1, y-1).
[0073] According to an embodiment of the present invention, the row spacing is obtained by fitting a straight line to the planting points of the above three plants using the RANSAC method to obtain a fitted straight line. The distance from the plant with coordinates (x, y) to the straight line is calculated based on the fitted straight line and used as the row spacing between the plant with coordinates (x, y) and the plant in the row to its right. The average row spacing of each plant in the same row is used as the row spacing of the plant in the row.
[0074] According to an embodiment of the present invention, for a plant with coordinates (x, y), the plant spacing is determined based on the three plants closest to it in the row above, and the coordinates of the plant are (x, y+1), (x-1, y+1) and (x+1, y+1).
[0075] According to an embodiment of the present invention, the plant spacing is obtained by fitting a straight line to the planting points of the above three plants using the RANSAC method to obtain a fitted straight line. The distance from the plant with coordinates (x, y) to the straight line is calculated based on the fitted straight line and used as the plant spacing between the plant with coordinates (x, y) and the plants in the row above it. The average plant spacing of each plant in the same row is used as the plant spacing of the plants in the row.
[0076] In some specific implementations, the present invention proposes a device for acquiring plant parameters, such as... Figure 10 As shown, the device includes a point cloud map acquisition module 100, used to collect point cloud data of the plant to obtain a point cloud map; a point cloud map preprocessing module 200, connected to the point cloud map acquisition module 100, used to preprocess the point cloud map of the plant to obtain a preprocessed point cloud map, the preprocessing including filtering and downsampling; and a point cloud dataset acquisition module 300, connected to the point cloud preprocessing module 200, used to... The preprocessed point cloud map undergoes data annotation and augmentation to obtain a point cloud dataset. A second segmentation processing module 400, connected to the point cloud dataset acquisition module 300, performs a second segmentation process on the point cloud dataset to obtain a point cloud model. A plant parameter extraction module 500, connected to the second segmentation processing module 400, extracts plant parameters from the point cloud model using point cloud processing technology. These parameters include at least one of height, density, row spacing, and plant spacing. The device according to this embodiment can effectively obtain plant planting parameters, providing support for intelligent machine operation.
[0077] According to some specific embodiments of the present invention, the point cloud map acquisition module includes a three-dimensional lidar unit. In some specific embodiments, a vehicle-mounted lidar (32 lines) is mounted on the top of an agricultural robot, tilted downwards at a certain angle during installation. After the industrial control computer receives the point cloud data via UDP protocol, it performs point cloud registration using a laser odometry or RTK-GNSS to obtain a point cloud map of corn in the field.
[0078] According to some specific embodiments of the present invention, the filtering process is Gaussian filtering. Therefore, the Gaussian filtering method is used to filter out some noise points and outliers.
[0079] According to some specific embodiments of the present invention, the downsampling process is voxel downsampling, wherein the voxel side length is set to 0.02m. Thus, the voxel downsampling method is applied to process point cloud data.
[0080] According to some specific embodiments of the present invention, the data annotation includes at least one of xyz coordinates, semantic labels, and instance labels.
[0081] According to some specific embodiments of the present invention, the data augmentation is performed using the random midpoint displacement method.
[0082] According to some specific embodiments of the present invention, the second segmentation processing module includes a MaizeNet segmentation network unit.
[0083] According to some specific embodiments of the present invention, the MaizeNet segmentation network unit includes a feature fusion layer.
[0084] According to some specific embodiments of the present invention, the MaizeNet segmentation network unit further includes an encoder and a decoder.
[0085] According to some specific embodiments of the present invention, the number of encoders and decoders is 3 to 6, preferably 4.
[0086] According to some specific embodiments of the present invention, the encoder includes local spatial coding, a self-attention unit, and a random downsampling unit.
[0087] According to some specific embodiments of the present invention, the decoder includes semantic branches and instance branches.
[0088] According to some specific embodiments of the present invention, the feature fusion layer is used for information interaction between the instance segmentation feature matrix and the semantic segmentation feature matrix.
[0089] According to some specific embodiments of the present invention, the information interaction includes the following operations:
[0090] (1) Input P sem With P ins The nearest neighbor feature matrix is obtained by the KNN method. and
[0091] (2) Attention scores are calculated using the self-attention method. and
[0092] (3) Calculate attention score and The weighted characteristic matrix F of the characteristic matrix of the nearest neighbor points of the other branch si With F is ;
[0093] (4) Based on the F si With P sem F is With Pins Obtain O through a fully connected network ins With O sem .
[0094] According to some specific embodiments of the present invention, F si and F is It is obtained through the following formula: Where K is the number of nearest neighbors.
[0095] According to some specific embodiments of the present invention, the total loss function is calculated by the following formula: L = αL C +βL V , where α is the weight coefficient of the semantic segmentation loss function; β is the weight coefficient of the instance segmentation loss function.
[0096] According to some specific embodiments of the present invention, the weight coefficient of the semantic segmentation loss function is 0.3 to 0.7, preferably 0.5.
[0097] According to some specific embodiments of the present invention, the weight coefficient of the instance segmentation loss function is 0.3 to 0.7, preferably 0.5.
[0098] According to some specific embodiments of the present invention, the semantic segmentation loss function is the standard cross-entropy loss function.
[0099] According to some specific embodiments of the present invention, the instance segmentation loss function is the instance embedding loss function.
[0100] According to some specific embodiments of the present invention, the standard cross-entropy loss function is calculated using the following formula: Among them, L c This represents the semantic segmentation loss function. p i semantic tags, Indicates network prediction p i Let N be the probability of the point being on the ground, N be the number of points, and i be the i-th point.
[0101] According to some specific embodiments of the present invention, the instance embedding loss function is obtained by the following formula:
[0102]
[0103]
[0104] Among them, L v L represents the instance segmentation loss function. var L represents the loss function for embedding identical instances. diffL represents the loss function that excludes different instances. reg Let N represent the regularization term, I represent the number of real instances, and N represent the number of instances. t u represents the number of points contained in the t-th instance. t Let e represent the average embedding of the t-th instance. i Denotes the embedding of the i-th point; δ s δ represents the maximum distance allowed for point embedding within the same instance. d Represents the maximum average embedding range across different instances; ||·||1 indicates the calculation of L1 distance; [x] + This indicates taking the maximum value in (0, x).
[0105] According to some specific embodiments of the present invention, an electronic device is provided, comprising a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program to implement the method for obtaining plant parameters as described in the first aspect. Alternatively, embodiments of this application also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the method of the above-described method embodiments.
[0106] According to some specific embodiments of the present invention, a computer-readable storage medium is provided for storing a computer program that causes a computer to perform the method for obtaining plant parameters described in the first aspect above. When implemented using software, it can be implemented wholly or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in accordance with the embodiments of this application are generated.
[0107] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0108] Example 1: Multi-task point cloud segmentation network structure for maize plants in the field (MaizeNet)
[0109] 1. Main Structure
[0110] The main structure of the MaizeNet network is as follows: Figure 1As shown, MaizeNet mainly consists of an encoder, a decoder, and a feature fusion layer. The entire network uses skip connections to fully utilize network feature information. MaizeNet has four encoders, whose main functions are to extract local features and downsample. A Transformer-based self-attention unit is incorporated into the local feature extraction process, which assigns attention weights to points based on the feature distribution of neighboring points. After local feature extraction, the feature map is randomly downsampled, reducing the number of points to one-quarter of the input number each time. The decoder has two branches: semantic and instance, processing semantic and instance features respectively. It upsamples the local features of the point cloud output from the encoder and concatenates them with local features at different resolutions. The feature fusion layer improves the segmentation accuracy of both tasks by combining the semantic and instance features output from the decoder. After the feature fusion layer, the semantic branch outputs a semantic prediction for each point, and the class with the highest probability is the semantic label for that point. The instance prediction branch obtains instance labels through MeanShift clustering.
[0111] 2. Encoder Structure
[0112] The encoder's role is to extract local features and perform random downsampling. Local feature extraction is the core of the deep learning network and directly determines the network's segmentation accuracy. MaizeNet's encoder mainly consists of three neural units: a local spatial encoding unit, a self-attention unit, and a random downsampling unit. The data processing flow within each encoder is as follows: Figure 2 As shown, the point cloud input is P = {x1,...,x} n ,...,x N},in 3+d represents the number of channels in the input point cloud, d represents the number of high-dimensional feature channels, N represents the number of points, and K represents the number of nearest neighbors. out This represents the number of output channels.
[0113] (1) Local spatial coding unit
[0114] By extracting nearest neighbors from keypoints and explicitly embedding their coordinates into the keypoint features, keypoints can acquire the spatial locations of their neighbors, thereby learning complex local structures. The structure of a local spatial coding unit is as follows: Figure 3 As shown (using the local spatial coding module proposed by RandLA-Net, from the paper [RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds]). Here, the point cloud input is P = {x1,...,x...} n ,...,x N},in 3+d represents the number of channels in the input point cloud, and N represents the number of points. The original coordinates of the point cloud occupy 3 channels, denoted as p. n The point cloud coordinates are transformed into high-dimensional features of dimension d using a Multi-Layer Perceptron (MLP) network, denoted as f. n The output dimension of this invention is d=8. After point-by-point feature extraction, by analyzing p... n Nearest neighbor search is performed to further aggregate local features. For the nth point x n Using the K-Nearest Neighbors (KNN) method based on Euclidean distance, the nearest neighbors of a point and their features are collected, represented as... In this invention, k = 16.
[0115] For nearest neighbor set Perform adjacent point position encoding. For center point p n The set of nearest neighbors is represented as Encode the positions of adjacent points according to the formula:
[0116]
[0117] In the formula, MLP(·) — fully connected network, — Connect operation, ||·||2 — Calculate L2 distance.
[0118] The location codes of adjacent points are concatenated with the features of nearest neighbors to obtain a local feature vector. This step can be represented as:
[0119]
[0120] Finally, the output of the local spatial coding unit is The output feature size is [N, k, d] out ], where d out The number of output channels is given, and the feature dimensions of the four encoders are [32, 128, 256, 512].
[0121] (2) Self-attention unit
[0122] By capturing potential salient regions in low-level features and fusing information from the center point and its neighbors to achieve local feature extraction, network performance can be improved while saving computational resources. The self-attention unit structure is as follows: Figure 4 As shown, the plus and minus signs inside the circles represent addition and subtraction operations, corresponding to the calculation of formula (3).
[0123] Research shows that additive attention mechanisms outperform multiplicative attention mechanisms; therefore, subtraction is used to calculate the attention score between K and Q. Simultaneously, position embeddings are added to the calculation to obtain point p. n Attention score matrix
[0124]
[0125] In the formula, ReLU(·) is the calibration linear activation function. As mentioned earlier, in MLP, Q, K, and V represent... Figure 4 The three matrices Query, Key, and Value in the diagram have the meanings of k as described above, and δ represents the positional encoding.
[0126] Then, the attention score is calculated as a weighted sum of the attention score and the context vector:
[0127]
[0128] In this context, K represents the meaning as described above, and n is the point number.
[0129] Finally, the output features C = {c1,...,c...} of the self-attention mechanism are obtained through aggregation operations. n ,...,c N The output feature size is [N, d]. out ], where d out The number of output channels is [32, 128, 256, 512] for the four encoders.
[0130] (3) Random downsampling
[0131] Since MaizeNet employs a multi-scale local feature extraction method, a subset of the encoder input is extracted at the end of each encoder layer using a random downsampling method. This reduces the number of points to 1 / 4 of the input points, and this subset is used as the output of the next encoder. Random downsampling involves randomly selecting n points, each with an equal probability of being selected, thus achieving a fixed number of points. Therefore, the input points of the four encoders are [69632, 17408, 4352, 1088], and the output features are... The number of output points are [17408, 4352, 1088, 272], and the feature dimension of the output is [32, 64, 256, 512].
[0132] 3. Decoder Structure
[0133] The decoder's function is to upsample the input features and connect local features at different scales. The decoder has two branches: semantic and instance, corresponding to the semantic segmentation task and the instance segmentation task, respectively. Each layer of the decoder outputs the connected semantics or semantic features.
[0134] The four decoders correspond to the four encoders. Each decoder layer has two inputs: the feature output from the previous layer and the input from the current layer. With the corresponding encoder output for Upsampling is performed using the KNN method to obtain... Then, and Connect them. Finally, feature extraction is performed through a fully connected network, which is represented as:
[0135]
[0136] In formula (5), UpSample(·) represents the upsampling operation. This indicates a connection operation, as described above. MLP, as described above, is a fully connected operation.
[0137] The decoder for semantic and instance branches ultimately outputs P. sem With P ins The number of points output by each decoder layer is [1088,4352,17408,69632], and the feature dimension of the output is [256,64,32,8].
[0138] 4. Feature Fusion Layer
[0139] Semantic segmentation and instance segmentation of point clouds share some similarities but also present conflicts. Accurate semantic segmentation helps improve the accuracy of instance segmentation because points with different semantic labels should have different instance labels. However, instance segmentation requires further assignment of different instance labels to points with the same semantic label. Therefore, a feature fusion layer is designed to improve the segmentation accuracy of both tasks by enabling information interaction between the instance segmentation feature matrix and the semantic segmentation feature matrix. The structure of the feature fusion layer is as follows: Figure 5 As shown.
[0140] In this process, the feature maps of the two branches are combined (added, concatenated, or multiplied) using formulas (6) and (7) to perform feature fusion. Specifically, the input to the feature fusion layer is P. sem With P ins The nearest neighbor feature matrix is obtained through the KNN method. and Attention scores were calculated using the self-attention method. and And calculate the weighted feature matrix F of the attention score and the feature matrix of the nearest neighbors of the other branch. si With F is This step is represented as:
[0141]
[0142] Here, ins represents the instance, sem represents the semantics, and finally, F is connected to each of them. si With P sem F is With P ins Feature extraction is performed using a fully connected network to obtain the network's final output O. ins With O sem This step is described as follows:
[0143]
[0144] The final output of the feature fusion layer is O sem With O ins For semantic segmentation, since the training data includes two categories, land surface and plant, the semantic segmentation output size is [69632, 2]. For instance segmentation, the final output is the location embedding learned by the network, so the semantic segmentation output size is [69632, 8]. The final instance segmentation result is obtained through the MeanShift clustering method.
[0145] 5. Loss Function
[0146] The choice of loss function directly affects the accuracy of the network. For multi-task segmentation networks, different loss functions are selected for each task. For semantic segmentation tasks, the standard cross-entropy loss function is used. The specific formula is as follows:
[0147]
[0148] in,
[0149] For the instance segmentation task, the unknown category instance embedding learning strategy from the paper (Wang X, Liu S, Shen X, et al. Associatively Segmenting Instances and Semantics in Point Clouds[C]. 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2019.) is adopted, using the instance embedding loss function, the specific formula of which is as follows:
[0150]
[0151] In equations (10)-(14), L c L represents the semantic segmentation loss function; v α represents the instance segmentation loss function; β represents the weight coefficient of the semantic segmentation loss function; p i The semantic labels are 0 for the ground and 1 for plants;
[0152] —Network prediction p i The probability of L being the surface; var — Same instance embedding loss function; L diff —Different instance rejection loss function; L reg —Regularization term; I —Number of actual instances; N t —The number of points contained in the t-th instance;
[0153] u t —The average embedding of the t-th instance; e i —The embedding of the i-th point; δ s —The maximum distance allowed for point embedding within the same instance; δ d —The maximum average embedding range across different instances; ||·||1 —Calculating the L1 distance; [x] + —Take the maximum value in (0, x).
[0154] Therefore, the total loss function L is:
[0155] L=αL c +βL v Formula (15)
[0156] In equation (15), α is the weight coefficient of the semantic segmentation loss function, which is 0.5; β is the weight coefficient of the instance segmentation loss function, which is 0.5.
[0157] Example 2: Network Training and Testing
[0158] 1. Hardware and software parameters for network training
[0159] (1) The hardware is an Nvidia Quadro P4000 graphics processor;
[0160] (2) In terms of software, functional code was written based on Python3, Pytorch1.4, CUDA10.1, and CUDNN8.0.
[0161] (3) During the training phase, the input point cloud of MaizeNet contains only the coordinate information of the points. The input point cloud is downsampled to a fixed number of points, 69632. The training batch size is 2. The Adam method is used as the solver. The initial learning rate is 0.008, the momentum is set to 0.9, and the number of iterations is set to 7.
[0162] 2. Dataset Construction
[0163] (1) Data collection and preprocessing
[0164] Point cloud data was collected from farmland planted with corn using an RS-LiDAR-32 3D lidar system. During the data collection period, the corn plants were at different growth stages and had varying heights. Farmland point cloud maps were obtained by registering the corn point cloud data. Due to the influence of planting density, growth stage, accessibility, and actual terrain, the farmland area, number of points, and point cloud density included in the generated point cloud maps varied. For each point cloud map, Gaussian filtering was used to remove some noise and outliers, and voxel downsampling was applied to process the point cloud data, with the voxel side length set to 0.02m. The preprocessed point cloud data was then used as input to the segmentation network.
[0165] (2) Corn point cloud data annotation and augmentation
[0166] ① Data labeling
[0167] First, the point cloud map was divided into 5m×5m sub-regions. Then, the corn instances in each sub-region were segmented using a segmentation function. The segmented corn instances were saved separately as files with the .txt extension, named according to the segmentation order. After all corn instances in the sub-region were segmented, the surface point cloud was saved as a file with the .txt extension named "00000". After manual annotation, 2500 complete corn plant point clouds were obtained from 10 point cloud maps. Due to mutual occlusion between plants, some corn plant point clouds were partially missing. To ensure the authenticity of the dataset, the few missing corn plant point clouds were not filled or deleted; the number of partially missing corn plant point clouds was 250. Therefore, a total of 2750 corn plants were obtained from the 10 maps. A program was written to save the annotated data as a .txt file, which contained the xyz coordinates, semantic labels, and instance labels of each point in the sub-region. The semantic labels are represented by 0 and 1, which represent soil and plants respectively; the instance labels are counted starting from 0 and include soil instances and plant instances. Since the number of soil instances is 1, the total number of instances is the number of plants plus 1.
[0168] ② Data augmentation
[0169] When the dataset is small, deep learning networks tend to overfit the training data, reducing their generalization ability. Data augmentation methods can expand the training data by processing existing data, improving network accuracy while maintaining its generalization ability.
[0170] The purpose of data augmentation was to generate a 5m × 5m point cloud of corn in a field. First, based on common corn planting types in the data collection environment, the collected farmland maps were divided into three categories: standard, aisle, and random. The standard type represents farmland where corn is planted with specified row and plant spacing; the aisle type represents a scene where the data collection vehicle travels along a road between two farmlands, with corn plants on either side of the road planted with specified row and plant spacing; the random type represents farmland where corn plants are randomly planted throughout the entire field. The standard type is the most common, followed by the aisle type, and the random type is the least common. The three types of corn farmland are shown below. Figure 6 As shown.
[0171] See Figure 7 The random midpoint displacement method, a fractal technique, is used to randomly generate soil point clouds. The random midpoint displacement method is a type of random fractal interpolation algorithm that requires only a few parameters to generate complex surfaces. It is easy to implement, recursive, and has high resolution. The random midpoint displacement method is further subdivided into the triangle midpoint displacement method and the square midpoint displacement method.
[0172] Finally, planting points for each corn plant were generated, and corn plants were randomly selected and placed on the curved surface of the soil. The cornfield type determined the method for selecting planting points, with row spacing of 0.6m, 0.65m, 0.7m, and 0.75m, and plant spacing of 0.5m, 0.55m, and 0.6m respectively. For each planting point, a small random error was added to simulate planting point deviations that occur during plant growth. Additionally, to simulate missing plants, there was a 1% probability that no corn plant would be placed at each planting point. Since corn plants planted simultaneously in actual cornfields can vary in height, plants at similar growth stages were selected to avoid significant height differences. Data augmentation methods were used to generate 15,000 cornfield datasets of 5m × 5m size.
[0173] 3. Network performance testing
[0174] Maize plants were divided into three categories: Category I, Category II, and Category III. Category I plants were mostly in the early jointing stage, with a plant height between 0.3m and 0.6m; Category II plants were mostly in the jointing stage, with a plant height between 0.6m and 0.9m; and Category III plants were mostly in the tasseling and maturity stages, with a plant height between 0.9m and 2m.
[0175] (1) Comparison of semantic segmentation results
[0176] MaizeNet is compared to PointNet (which extracts global features using a learnable multilayer perceptron (MLP) and max pooling, where max pooling ensures the network's output is independent of the input order. Since global features cannot capture local details, many networks improve prediction accuracy by extracting finer local features; see: Qi CR, Yi L, Su H, et al. Pointnet++: Deephierarchical feature learning on point sets in a metric space[J]. Advances in neural information processing systems, 2017, 30.)++ and RandLANet (which extracts local features through a local feature aggregation module, which captures local point features using an attention pooling module; see: Hu Q, Yang B, Xie L, et al. RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds[C]. 2020 IEEE / CVF Conference on Computer Vision and Pattern). This paper compares the semantic segmentation results of two deep learning networks, Recognition (CVPR). IEEE, 2020., and evaluates their accuracy, recall, F1 score, and overall accuracy (OA). The test results are shown in Table 1.
[0177] Table 1: Test results of semantic segmentation accuracy of three networks
[0178]
[0179] (2) Comparison of instance segmentation results
[0180] MaizeNet's semantic segmentation results are compared with those of two conventional deep learning networks: ASIS and the improved RandLANet (which improves the RandLANet network by increasing the number of RandLANet decoders to two and adding the proposed feature fusion layer at the end of the decoders). The results are evaluated using four metrics: precision, recall, mean coverage (mCov), and mean weighted coverage (mWCov).
[0181] Table 2: Comparison of Segmentation Accuracy of Different Network Instances
[0182]
[0183]
[0184] Based on the test results above, MaizeNet can directly process cornfield point cloud maps to achieve fast and accurate semantic and instance segmentation, and the segmentation accuracy is better than conventional methods in this field.
[0185] (3) Ablation test
[0186] A self-attention unit is proposed in the feature extraction part, and a feature fusion layer is proposed after the decoder. The effectiveness of these two parts is verified by ablation experiments.
[0187] ① Self-attention unit:
[0188] To verify the effectiveness of the self-attention unit, MaizeNet and RandLANet were selected as the base networks. MaizeNet with the self-attention unit removed and RandLANet with the self-attention unit integrated were compared. The evaluation metrics for semantic segmentation and instance segmentation of the four networks were calculated, as shown in Tables 3 and 4. Method 1 is MaizeNet, Method 2 is MaizeNet with the self-attention module removed, Method 3 is RandLANet, and Method 4 is an improved RandLANet with the self-attention module added.
[0189] Table 3: Semantic segmentation results of the self-attention module experiment
[0190] 1 0.973 0.979 0.974 0.979 2 0.935 0.943 0.939 0.937 3 0.957 0.963 0.959 0.966 4 0.963 0.974 0.968 0.970
[0191] Table 4: Segmentation Results of Experimental Instances for the Self-Attention Module
[0192] 1 0.901 0.926 0.845 0.829 2 0.821 0.837 0.733 0.704 3 0.875 0.926 0.778 0.766 4 0.894 0.934 0.803 0.792
[0193] As shown in Table 3, the semantic segmentation results of networks with added self-attention units consistently demonstrate higher semantic segmentation accuracy than those without. This indicates that self-attention units enhance the ability to extract local features, thereby improving semantic segmentation accuracy. Table 4 shows that the instance segmentation results of networks with added self-attention modules all exhibit improved instance segmentation accuracy, with a more significant improvement observed in MaizeNet. This is because the RandLANet feature extraction method already includes an attention module; replacing this module with the self-attention unit proposed in this study also improves instance segmentation accuracy, suggesting that this model is more suitable for extracting fine local features from field corn point clouds.
[0194] ② Feature Fusion Layer: To verify the effectiveness of the feature fusion layer, similar to the previous methods, MaizeNet and the improved RandLANet were selected as the base networks. They were compared with MaizeNet (without the feature fusion layer) and the improved RandLANet (with the feature fusion layer added). For specific operation methods, please refer to [link / reference]. Figure 4 and Figure 5 The experimental results are shown in Tables 5 and 6. In the tables, Method 1 is MaizeNet, Method 2 is MaizeNet with the feature fusion layer removed, Method 3 is the improved RandLANet, and Method 4 is the improved RandLANet with the feature fusion layer added.
[0195] Table 5: Experimental semantic segmentation results of the feature fusion layer
[0196] 1 0.973 0.979 0.974 0.979 2 0.908 0.913 0.910 0.912 3 0.957 0.963 0.959 0.966 4 0.971 0.975 0.973 0.975
[0197] Table 6: Segmentation Results of Experimental Examples in Feature Fusion Layer
[0198] 1 0.901 0.926 0.845 0.829 2 0.781 0.801 0.681 0.674 3 0.763 0.793 0.675 0.671 4 0.875 0.926 0.778 0.766
[0199] As shown in Table 5, the semantic segmentation results indicate that the feature fusion layer improves the semantic segmentation accuracy of both MaizeNet and RandLANet, especially significantly improving the semantic segmentation accuracy of MaizeNet. Table 6 shows that the instance segmentation results also demonstrate that the feature fusion layer significantly improves the instance segmentation accuracy of MaizeNet and RandLANet, indicating that the proposed feature fusion method can fully combine the advantages of semantic and instance features to improve the segmentation accuracy of both tasks.
[0200] Example 3: Extraction of Field Maize Planting Parameters
[0201] Based on the segmentation results of Example 2, a method for extracting field maize parameters is proposed. The extracted parameters include row spacing, plant spacing, planting density, and plant height. At the same time, the maize planting line is fitted and missing plants are marked. This information can provide a reference for the autonomous operation of intelligent agricultural machinery.
[0202] 1. Extraction of plant height
[0203] Each corn instance is numbered by row and column and labeled M. x,y The height of a corn plant is defined as the difference between its highest and lowest points. Depending on the growth stage, the highest point may be the height of the highest leaf or the highest point of the tassel. Therefore, in corn example M... x,y height Determined by formula (16):
[0204]
[0205] 2. Planting density extraction
[0206] The area S of the plot is calculated based on the coordinates of its four corners. The planting density ρ is determined by the number of plants N and S in the plot according to formula (17):
[0207] ρ=N / S Formula (17)
[0208] 3. Extraction of row spacing and plant spacing
[0209] When calculating row spacing, the three closest instances M from the rightmost row of the plant are selected. 3,4 M 3,5 With M 3,6 The RANSAC method was used to fit straight lines to the planting points of these three instances, and the distances from the points to the lines were calculated as M. 2,5 The line spacing from the right-hand line is denoted as . Similarly, when calculating the plant spacing, calculate the plant spacing M. 2,5 The distance to the line fitted to the three nearest instances above is denoted as . The calculation methods for row spacing and plant spacing are as follows: Figure 9 As shown.
[0210] Row spacing is calculated starting with the first plant in each row and ending with the last plant. The average row spacing calculated for each plant is used as the row spacing for that row. Outliers are removed during the calculation to account for the possibility of missing plants. Similarly, plant spacing is also calculated based on the average plant spacing calculated for each plant.
[0211] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "embodiment," or "specific embodiment," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment. Furthermore, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments and features described in this specification.
[0212] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for obtaining plant parameters, characterized in that, The method includes: 1) Collect point cloud data of the plant to obtain a point cloud map; 2) Preprocess the point cloud map of the plant to obtain a preprocessed point cloud map. The preprocessing includes filtering and downsampling. 3) Perform data annotation and data augmentation on the preprocessed point cloud map to obtain a point cloud dataset; 4) Perform a second segmentation process on the point cloud dataset to obtain a point cloud model; 5) Extract the parameters of the plant from the point cloud model using point cloud processing technology, wherein the parameters include at least one of height, density, row spacing, and plant spacing, wherein: The point cloud data of the plant was acquired using a three-dimensional lidar, and the point cloud data of the plant includes point cloud data of any growth stage of the plant. The filtering process is a Gaussian filtering process; The downsampling process is voxel downsampling, wherein the voxel side length is set to 0.02 m; The data annotation includes at least one of xyz coordinates, semantic labels, and instance labels; The data augmentation was performed using the random midpoint displacement method; The second segmentation process includes semantic segmentation and instance segmentation. The second segmentation process is performed through the MaizeNet segmentation network. The MaizeNet segmentation network includes a feature fusion layer, an encoder, and a decoder. The number of encoders and decoders is 3 to 6. The encoder includes local spatial encoding, self-attention units, and random downsampling units. The decoder includes semantic branches and instance branches. The feature fusion layer is used for information interaction between the instance segmentation feature matrix and the semantic segmentation feature matrix. The information interaction includes the following operations: (1) Input and The nearest neighbor feature matrix is obtained by the KNN method. and ; (2) Attention scores were calculated using the self-attention method. and ; (3) Calculate attention score and The weighted feature matrix of the feature matrix of the nearest neighbor points of the other branch and ; (4) Based on the above and , and Obtained through a fully connected network and , The The output features of the semantic segmentation branch decoder, the The output features of the segmented branch decoder are used to segment the instance.
2. The method according to claim 1, characterized in that, and It is obtained through the following formula: , , where K is the number of nearest neighbors.
3. The method according to claim 1, characterized in that, The total loss function is calculated using the following formula: , where α is the weight coefficient of the semantic segmentation loss function; β is the weight coefficient of the instance segmentation loss function.
4. The method according to claim 3, characterized in that, The weight coefficients of the semantic segmentation loss function are 0.3 to 0.7; and / or The weight coefficients of the instance segmentation loss function are 0.3 to 0.
7.
5. The method according to claim 3, characterized in that, The semantic segmentation loss function is the standard cross-entropy loss function; and / or The instance segmentation loss function is the instance embedding loss function.
6. The method according to claim 5, characterized in that, The standard cross-entropy loss function is calculated using the following formula: ,in, This represents the semantic segmentation loss function. express semantic tags, Indicates network prediction The probability of the point being on the ground, where N represents the number of points, and i represents the i-th point; and / or The instance embedding loss function is obtained using the following formula: , , , , in, This represents the instance segmentation loss function. The loss function represents the embedding of identical instances. This represents the loss function that excludes different instances. Represents the regularization term. I Indicates the actual number of instances. Indicates the first t The number of points contained in each instance. Indicates the first t Average embedding of each instance Indicates the first i Embedding of points; This represents the maximum distance that points are allowed to embed within the same instance; Indicates the maximum range of average embeddings across different instances; This indicates the calculation of the L1 distance; This means taking (0, x The maximum value in ).
7. The method according to claim 1, characterized in that, The height is the difference between the height of the highest point and the lowest point of the plant.
8. The method according to claim 7, characterized in that, The highest point refers to either the highest point of the leaf or the highest point of the tassel.
9. The method according to claim 7 or 8, characterized in that, The height is determined by the formula. The calculation yielded, where, This represents the height of the plant at coordinates (x, y). This represents the height of the highest point of the plant at coordinates (x, y). This represents the height of the lowest point of the plant at coordinates (x, y).
10. The method according to claim 1, characterized in that, The density is calculated based on the area of the plot where the plant is located and the number of plants.
11. The method according to claim 1, characterized in that, The density is obtained through the formula The calculation yielded, where, Indicates density, N S represents the number of plants, and S represents the area of the plot.
12. The method according to claim 1, characterized in that, The row spacing is calculated from the first plant in each row to the last plant in each row.
13. The method according to claim 1, characterized in that, For a plant with coordinates (x, y), the row spacing of the plant is determined based on the three plants closest to it in the row to its right, and the coordinates of the plant are (x+1, y), (x+1, y+1) and (x+1, y-1).
14. The method according to claim 13, characterized in that, The row spacing is obtained by fitting a straight line to the planting points of the above three plants using the RANSAC method. The distance from the plant with coordinates (x, y) to the straight line is calculated based on the fitted line and used as the row spacing between the plant with coordinates (x, y) and the plants in the row to its right. The average row spacing of each plant in the same row is used as the row spacing of the plant in the row.
15. The method according to claim 1, characterized in that, For a plant with coordinates (x, y), the plant spacing is determined based on the three plants closest to it in the row above, and the coordinates of the plant are (x, y+1), (x-1, y+1), and (x+1, y+1).
16. The method according to claim 15, characterized in that, The plant spacing is obtained by fitting a straight line to the planting points of the above three plants using the RANSAC method. The distance from the plant with coordinates (x, y) to the straight line is calculated based on the fitted line, and is used as the plant spacing between the plant with coordinates (x, y) and the plants in the row above it. The average plant spacing of each plant in the same row is used as the plant spacing of the plant in the row.
17. A device for acquiring plant parameters, characterized in that, The device includes: The point cloud map acquisition module is used to collect point cloud data of the plant to obtain a point cloud map. A point cloud map preprocessing module, which is connected to the point cloud map acquisition module, is used to preprocess the point cloud map of the plant to obtain a preprocessed point cloud map. The preprocessing includes filtering and downsampling. A point cloud dataset acquisition module, which is connected to the point cloud map preprocessing module, is used to perform data annotation and data augmentation on the preprocessed point cloud map to obtain a point cloud dataset. The second segmentation processing module is connected to the point cloud dataset acquisition module and is used to perform a second segmentation processing on the point cloud dataset to obtain a point cloud model. A plant parameter extraction module, connected to the second segmentation processing module, is used to extract plant parameters from the point cloud model using point cloud processing technology. The parameters include at least one of height, density, row spacing, and plant spacing. The point cloud map acquisition module includes a 3D LiDAR; The filtering process is a Gaussian filtering process; The downsampling process is voxel downsampling, wherein the voxel side length is set to 0.02m; The data annotation includes at least one of xyz coordinates, semantic labels, and instance labels; The data augmentation was performed using the random midpoint displacement method; The second segmentation processing module includes a MaizeNet segmentation network, which includes a feature fusion layer, an encoder, and a decoder. The number of encoders and decoders is 3 to 6. The encoder includes a local spatial encoding, a self-attention unit, and a random downsampling unit. The decoder includes a semantic branch and an instance branch. The feature fusion layer is used for information interaction between the instance segmentation feature matrix and the semantic segmentation feature matrix; The information interaction includes the following operations: (1) Input and The nearest neighbor feature matrix is obtained by the KNN method. and ; (2) Attention scores were calculated using the self-attention method. and ; (3) Calculate attention score and The weighted feature matrix of the feature matrix of the nearest neighbor points of the other branch and ; (4) Based on the above and , and Obtained through a fully connected network and , The The output features of the semantic segmentation branch decoder, the The output features of the segmented branch decoder are used to segment the instance.
18. The device according to claim 17, characterized in that, and It is obtained through the following formula: , , where K is the number of nearest neighbors.
19. The device according to claim 17, characterized in that, The total loss function is calculated using the following formula: , where α is the weight coefficient of the semantic segmentation loss function; β is the weight coefficient of the instance segmentation loss function.
20. The device according to claim 19, characterized in that, The weight coefficients of the semantic segmentation loss function are 0.3 to 0.7; and / or The weight coefficients of the instance segmentation loss function are 0.3 to 0.
7.
21. The device according to claim 19, characterized in that, The semantic segmentation loss function is the standard cross-entropy loss function; and / or the instance segmentation loss function is the instance embedding loss function.
22. The device according to claim 21, characterized in that, The standard cross-entropy loss function is calculated using the following formula: ,in, This represents the semantic segmentation loss function. express semantic tags, Indicates network prediction Let N be the probability of the point being on the ground, N be the number of points, and i be the i-th point.
23. The device according to claim 21, characterized in that, The instance embedding loss function is obtained using the following formula: , , , , in, This represents the instance segmentation loss function. The loss function represents the embedding of identical instances. This represents the loss function that excludes different instances. Represents the regularization term. I Indicates the actual number of instances. Indicates the first t The number of points contained in each instance. Indicates the first t Average embedding of each instance Indicates the first i Embedding of points; This represents the maximum distance that points are allowed to embed within the same instance; Indicates the maximum range of average embeddings across different instances; This indicates the calculation of the L1 distance; This means taking (0, x The maximum value in ).
24. An electronic device, comprising a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the computer program to implement the method for obtaining plant parameters as described in any one of claims 1 to 16.
25. A computer-readable storage medium, characterized in that, Used to store computer programs; The computer program causes the computer to perform the method for obtaining plant parameters as described in any one of claims 1 to 16.
Citation Information
Patent Citations
Crop phenotypic parameter extraction method and device
CN111696122A