Soybean breeding plot instance segmentation method based on deep learning and prior knowledge

By applying deep learning spatiotemporal feature alignment network and traction line aggregation method in soybean breeding communities, the problem that traditional methods are difficult to achieve high-precision soybean breeding community instance segmentation in complex environments is solved, and high-precision, stable and robust segmentation effects are achieved, and good migration capabilities are provided.

CN120107587APending Publication Date: 2025-06-06NANJING AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510181730.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In large-scale breeding practice, traditional canopy semantic segmentation methods are difficult to achieve high-precision and stable instance segmentation of soybean breeding cells in complex environments, dynamic canopy, and confusion of row spacing and cell intervals.

Method used

The deep learning-based spatiotemporal feature alignment network (STANet) model is used to combine the traction line aggregation method (TLA), and coordinate attention, spatiotemporal feature alignment and planting design prior knowledge are combined to perform soybean canopy semantic segmentation and cell instance segmentation.

Benefits of technology

It realizes high-precision soybean canopy segmentation in full breeding period in complex environments, ensuring the accuracy and continuity of canopy segmentation in different periods and under multiple environmental conditions, and supports robust segmentation of different planting designs, with the ability to migrate across New Year, across locations and across data types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107587A_ABST
    Figure CN120107587A_ABST
Patent Text Reader

Abstract

The invention discloses a soybean breeding plot instance segmentation method based on deep learning and prior knowledge, and the method comprises the following steps: collecting data through employing an unmanned plane, constructing a time sequence data set, and dividing the time sequence data set into a training set, a verification set and a test set; through a spatial-temporal feature alignment network STANet model, performing canopy semantic segmentation on the time sequence image, capturing multi-scale spatial information and time sequence information, separating a soybean canopy from a background, performing model training, and improving segmentation precision; and based on the canopy semantic segmentation map, introducing prior knowledge of soybean plot planting design by using a pull line aggregation method TLA, generating a horizontal pull line and a vertical pull line, and performing plot instance segmentation. Through the stability, mobility and universality of STANet-TLA, excellent germplasm resources are screened, the breeding period is shortened, and the breeding efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a soybean breeding plot instance segmentation method, and in particular to a soybean breeding plot instance segmentation method based on deep learning and prior knowledge. Background Art

[0002] Soybean, a key food and oil crop, provides more than a quarter of the world's protein supply for human consumption and animal feed.

[0003] Canopy structure plays a central role in light interception and subsequent biochemical processes in soybeans, directly affecting the final yield. Its dynamic characteristics are closely linked to the soybean growth period, especially the soybean grain filling process during canopy closure, which has a decisive influence on material accumulation and yield formation. Studies have confirmed that various phenotypic characteristics of soybean canopies, including structure, spectrum, texture and temperature, are closely related to yield and can be used as effective indicators for predicting yield. In addition, the dynamic changes in canopy structural traits can reveal differences between varieties, providing breeders with a way to identify potential phenotypic traits associated with target traits. These findings have opened up new perspectives for crop breeding using quantitative growth curves.

[0004] In large-scale breeding practice, efficient monitoring of soybean canopy development relies on high-throughput, time-series phenotyping techniques. As an emerging remote sensing tool, unmanned aerial vehicles (UAVs) have revolutionized field crop development monitoring with their low cost, high efficiency, large-scale coverage, and spatiotemporal monitoring capabilities. UAVs have been successfully used to extract canopy structural traits, spectral indices, and texture features, and then estimate crop biomass and yield. However, despite the progress made in these studies in phenotypic trait extraction and predictive modeling, their application in variety screening in actual breeding trials is still insufficient.

[0005] Automatic segmentation of breeding plots is the basis of phenotypic analysis, but it faces many challenges. The instance segmentation problem in the field of computer vision requires that different varieties be divided into independent regions of interest (ROIs) and different rows of the same variety be integrated into a single ROI. Traditional canopy semantic segmentation methods rely on complex feature design and threshold settings, and perform poorly in environments with variable lighting, leaf color, and background. Although deep learning methods have made significant progress in image segmentation, their direct application to crop breeding plot boundary extraction still faces unique challenges such as complex environments, dynamic canopies, and confusion between row spacing and plot spacing. Summary of the invention

[0006] Purpose of the invention: The purpose of the present invention is to provide a soybean breeding plot instance segmentation method based on deep learning and prior knowledge, so as to improve the accuracy and stability of soybean breeding plot segmentation across years, locations, data types and different planting designs.

[0007] Technical solution: The soybean breeding plot instance segmentation method based on deep learning and prior knowledge described in the present invention comprises the following steps:

[0008] (1) Use drones to collect data and construct a time series dataset, which is divided into training set, validation set, and test set;

[0009] (2) Through the spatiotemporal feature alignment network STANet model, the time series images are semantically segmented to capture multi-scale spatial information and time series information, separate the soybean canopy from the background, perform model training, and improve segmentation accuracy;

[0010] (3) Based on the canopy semantic segmentation map, the traction line aggregation method TLA is used to introduce the prior knowledge of soybean plot planting design, generate horizontal and vertical traction lines, and perform plot instance segmentation.

[0011] Preferably, the STANet model in step (2) is an encoder-decoder architecture, the encoder includes a deformable convolution module DCN and a coordinate attention mechanism module CA, which are used to extract multi-scale and time series features of the canopy; the decoder includes 4 consecutive spatiotemporal alignment modules STAM, the first spatiotemporal alignment generates a multi-scale feature by horizontally connecting and fusing the feature map of the last layer of the encoder and its last moment feature, the multi-scale feature map is upsampled by a factor of 2 through the nearest neighbor interpolation strategy, and the final feature map uses a 1×1 convolutional layer to generate a canopy semantic segmentation mask.

[0012] Preferably, the encoder uses ResNet101 as the backbone network, and the input is a time series image. The features generated after processing by the deformable convolution module are then processed by the coordinate attention mechanism module. ResNet101 encodes the processed features to generate time series features.

[0013] Preferably, the coordinate attention mechanism module includes coordinate information embedding and coordinate attention generation, wherein the coordinate information embedding decomposes the input two-dimensional features into two one-dimensional features in the horizontal and vertical directions respectively through average pooling, and the coordinate attention is generated by the above two one-dimensional features.

[0014] Preferably, the spatiotemporal alignment module includes a feature alignment module and a convolutional long short-term memory module. With time series features and multi-scale features as input, the horizontally connected time series features are spatially aligned and fused with the multi-scale features through the feature alignment module, and the aligned time features are input into the convolutional long short-term memory to extract the time series features.

[0015] Preferably, the second spatiotemporal alignment in the spatiotemporal alignment module fuses the upsampled feature map and the horizontally connected time series feature map with the same spatial size to generate multi-scale features, the third spatiotemporal alignment fuses the upsampled feature map and the horizontally connected time series feature map with the same spatial size to generate multi-scale features, and the fourth spatiotemporal alignment fuses the upsampled feature map and the horizontally connected time series feature map with the same spatial size to generate multi-scale features.

[0016] Preferably, the feature alignment module takes the horizontally connected time series features and multi-scale feature maps as input, extracts the key weights of each channel of the time series features through average pooling, 1×1 convolution and sigmoid activation function, retains the key features and removes irrelevant features; multiplies the key weights with the input time series features, and adds the product to the input time series features to obtain the key features; adjusts the key feature channels through 1×1 convolution to generate feature selection results, and aligns the selected features with the multi-scale features.

[0017] Preferably, the convolutional long short-term memory module is used to process time series information and establish long-distance dependencies, the input is a three-dimensional tensor with a time dimension, and the output is a multi-scale feature map.

[0018] Preferably, the convolutional long short-term memory constructs time series data for time series images in the following steps: selecting an RGB image to fill the position at time t, and saving the segmentation result of the selected image as a matrix label, and then selecting an RGB image before or at the time t to fill the position at time t-1, and then filling the position at time t-2 in sequence.

[0019] Preferably, the traction line polymerization method in step (3) specifically comprises the following steps:

[0020] (31) Merge the canopy semantic segmentation results into a unified semantic segmentation map for the entire study area;

[0021] (32) Rotate the merged semantic segmentation map so that the planting row direction is aligned with the horizontal coordinate system of the image;

[0022] (33) Generate horizontal traction lines and vertical traction lines, and optimize the traction lines;

[0023] (34) The semantic segmentation map is divided into horizontal sub-maps using optimized horizontal traction lines, each horizontal sub-map is vertically divided into individual cells based on optimized vertical traction lines, and the segmented image is rotated back to its original orientation using the rotation angle determined in step (32) to complete the cell segmentation.

[0024] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: by integrating coordinate attention, spatiotemporal feature alignment and planting design prior knowledge, high-precision soybean canopy segmentation in the entire growth period under complex environments is achieved; the accuracy and continuity of soybean canopy segmentation under various environmental conditions during the entire growth period are strong for soybean canopy changes in different periods; at the same time, robust segmentation can be achieved for breeding plots with different planting designs; based on the stability, mobility and universality of STANet-TLA, the migration of data sets across years, locations and data types can be achieved, which helps to screen excellent germplasm resources, shorten the breeding cycle and improve breeding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is the STANet model architecture diagram described in the present invention.

[0026] Figure 2 Construct a graph for the time series data described in the present invention.

[0027] Figure 3 This is a process diagram of the traction line polymerization method TLA described in the present invention.

[0028] Figure 4 Quantitative analysis results of STANet-TLA accuracy at different days after seeding.

[0029] Figure 5 Qualitative analysis results of STANet-TLA accuracy at different days after seeding.

[0030] Figure 6 Canopy semantic segmentation results for the SoyUAV2022SYb dataset 37 days after sowing.

[0031] Figure 7 Plot instance segmentation results for a two-row planting design.

[0032] Figure 8 Plot instance segmentation results for a three-row planting design.

[0033] Fig. 9 Analyze and compare the results of STANet model with other models. DETAILED DESCRIPTION

[0034] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0035] The embodiment of the present invention provides a soybean breeding plot instance segmentation method based on deep learning and prior knowledge, comprising the following steps:

[0036] (1) Use drones to collect data and construct a time series dataset, which is divided into training set, validation set, and test set.

[0037] SoyUAV is a time series drone dataset for soybean canopy semantic segmentation. The dataset contains four subsets: SoyUAV2022SY, SoyUAV2023SY, SoyUAV2022NJ, and SoyUAVchm.

[0038] SoyUAV2022SY is a field trial conducted in Sanya in 2022. The plot planting design is 1.2m long, two rows, 977 varieties, two replicates, and a total of 1944 plots. It contains drone data at 17 time points, covering the entire growth period of soybeans. SoyUAV2022SY is divided into subsets A and B, of which SoyUAV2022Sya contains 8814 samples for model training and verification, and SoyUAV2022SYb contains 2934 samples for testing;

[0039] SoyUAV2023SY contains 2724 samples. It is a field experiment conducted in Sanya in 2023. The planting design is consistent with the Sanya experiment in 2022. It contains drone data at 22 time points, covering the entire growth period of soybean, which is used for the migration analysis of the model in cross-year datasets; SoyUAV2022NJ contains 849 samples. It is a field experiment conducted in Nanjing in 2022. The plot planting design is 0.7m long, three rows, 405 varieties, two replicates, and a total of 910 plots. It contains drone data at 8 time points, covering the entire growth period of soybean, which is used for the migration analysis of the model in cross-location datasets; SoyUAVchm contains 5998 samples. It is a canopy height model (CHM) dataset generated by the 2022 Sanya experiment data. The CHM data is a single-band floating point type, which is used for the migration analysis of the model in cross-data type datasets.

[0040] (2) Through the training of the spatiotemporal feature alignment network (STANet) model, canopy semantic segmentation is performed to capture multi-scale spatial information and time series information, separate the soybean canopy from the background, and perform model training to improve segmentation accuracy.

[0041] STANet is an encoder-decoder architecture for soybean canopy semantic segmentation. Figure 1 As shown, it includes deformable convolution (DCN), coordinate attention mechanism (CA) and spatiotemporal alignment module (STAM) to capture multi-scale spatial information and time series information.

[0042] The encoder uses ResNet101 as the backbone network, and introduces the deformable convolution (DCN) and coordinate attention (CA) modules. The time series image is input to the encoder and processed by a deformable convolution module with a convolution kernel of 3×3 to expand the image receptive field and improve the feature extraction capability. The features generated by the deformable convolution module are then processed by the coordinate attention module to capture the detailed canopy boundary features of the image. After being processed by the coordinate attention module, ResNet101 further encodes the features to generate time series features, recorded as {C2, C3, C4, C5}. Among them, the last layer of ResNet101 deletes the average pooling and fully connected layers to reduce the loss of spatial information.

[0043] The coordinate attention mechanism (CA) consists of two components, namely, coordinate information embedding and coordinate attention generation, such as Figure 1 b. Coordinate information embedding decomposes the input two-dimensional features (H×W) into two one-dimensional features (H and W) in the horizontal and vertical directions respectively through average pooling. The two one-dimensional feature vectors can capture long-distance spatial dependencies and retain the location information of the object. The formula is as follows:

[0044]

[0045] in, and Respectively represent the position information of the cth channel in the vertical (H) and horizontal (W) directions. c (h, i) and x c (w, j) represents the i-th and j-th pixel values ​​in the horizontal and vertical directions respectively.

[0046] The coordinate attention map is generated by the two one-dimensional features mentioned above. The two one-dimensional features are connected along the spatial dimension (C×(H+W)). The connected feature map is compressed in the channel dimension through 1×1 convolution. Batch normalization (BatchNorm) and non-linear activation (Non-linear) are used to divide the compressed feature map into two branches. The two branches are convolved and activated to generate two directional weights (H: and W: ). These two weights are multiplied by the input features to obtain the output features, as follows:

[0047]

[0048] Among them, y c (i, j) represents the result of the coordinate attention mechanism, x c (i, j) represents the input feature map, is the weight in the vertical direction, is the weight in the horizontal direction.

[0049] The output features are used as the input of the ResNet101 backbone network to generate a series of time series features, namely C2, C3, C4, and C5.

[0050] The decoder adopts four consecutive STAMs as a top-down path. The first STAM generates a multi-scale feature P5 by horizontally connecting the feature map C5 of the last layer of the encoder: B×T×C×H×W and the last moment feature of C5 (B×C×H×W). The multi-scale feature map is upsampled by a factor of 2 through the nearest neighbor interpolation strategy. The second STAM fuses the upsampled feature map with the horizontally connected time series feature map with the same spatial size to generate a multi-scale feature P4. The third STAM fuses the upsampled feature map with the horizontally connected time series feature map with the same spatial size to generate a multi-scale feature P3. The fourth STAM fuses the upsampled feature map with the horizontally connected time series feature map with the same spatial size to generate a multi-scale feature P2. The final feature map P2 uses a 1×1 convolutional layer to generate a canopy semantic segmentation mask.

[0051] STAM consists of a feature alignment module (FAM) and a convolutional long short-term memory (ConvLSTM), such as Figure 1 As shown in Figure 3, with time series features {C2, C3, C4, C5} and multi-scale features as input, the time series features of each horizontal connection are spatially aligned and fused with the multi-scale features through FAM, and then the aligned time series features are input into the ConvLSTM module to extract the time series features.

[0052] The feature alignment module (FAM) uses lightweight convolution operations to solve the changes in the spatial features of the canopy at different times. The input components of FAM are: the horizontally connected time series features C i-1 and multi-scale feature maps like Figure 1 c. Time series features C are extracted by average pooling, 1×1 convolution and sigmoid activation function i-1 The key weight u of each channel is used to selectively retain key features and remove irrelevant features. The formula is as follows:

[0053]

[0054] Among them, f m represents average pooling, 1×1 convolution, and sigmoid activation function.

[0055] Then the key weight u is combined with the input time series feature C i-1 The product is further added to the input time series features to obtain the key features. The key feature channels are adjusted by 1×1 convolution to generate feature selection results. The formula is as follows:

[0056]

[0057] Among them, f s Represents a 1×1 convolution.

[0058] The features to be selected Align with multi-scale features. and concatenate, and then perform a 1×1 convolution to generate a pixel shift map Δ i , the formula is as follows:

[0059]

[0060] Among them, f o Represents a 1×1 convolution.

[0061] Through iterative training, the offset between multi-scale features and time series features is detected and used to achieve alignment, and Δ is processed by deformable convolution (DCN) and linear rectification unit (ReLU) i and To obtain the alignment result The formula is as follows:

[0062]

[0063] Among them, f a Represents deformable convolution and rectified linear unit.

[0064] Align the results Add to To create the final aligned feature map.

[0065] Convolutional Long Short-Term Memory (ConvLSTM) is used to process time series information and establish long-distance dependencies. The input of ConvLSTM is a three-dimensional tensor with a time dimension, including data at the current moment and its adjacent past moments. The output of ConvLSTM is a multi-scale feature map.

[0066] This embodiment selects three time sequence feature graphs as a combination to construct a time series matrix [time t-2 , time t-1 , time t ], for time series images, such as Figure 2As shown in a, a random RGB image is selected to fill the position at time t, and the segmentation result of the selected image is saved as a matrix label. Then, an RGB image before or at the time t is selected to fill the position at time t-1, and then fill the position at time t-2. For common scenarios of single image segmentation without sequence information, such as Figure 2 As shown in b, a feature map is randomly selected and replicated three times, and the corresponding labels are saved as matrix labels.

[0067] (3) Based on the canopy semantic segmentation map, the traction line aggregation method TLA is used to introduce the prior knowledge of soybean plot planting design, generate horizontal and vertical traction lines, and perform plot instance segmentation.

[0068] Based on the canopy semantic segmentation results of STANet, the traction line aggregation method TLA is used to segment a single breeding plot in combination with prior knowledge constraints, such as Figure 3 As shown, the following steps are included:

[0069] (31) Image merging: Merge the semantic segmentation results into a unified semantic segmentation map of the entire study area;

[0070] (32) Image rotation: The merged semantic segmentation map is rotated so that the planting row direction is aligned with the image horizontal coordinate system. The angle between the canopy row and the horizontal axis is measured by a protractor as the rotation angle, which is a constant and can be used across time series images because they share the same geographic coordinate system in a single experiment;

[0071] (33) Cell segmentation: Generate horizontal traction lines (HTL) and vertical traction lines (VTL), and optimize the traction lines (TL).

[0072] (331) The horizontal traction line is generated based on the cell row detection and the number of rows in each cell. The cell row detection is based on the binary image of the canopy semantic segmentation, where the canopy is represented as 1 and the background is represented as 0. The specific steps are as follows:

[0073] (3311) Count the number of pixels in each row of the canopy in the merged image, and use the row with the canopy pixel number gradient closest to 0 as the preliminary horizontal traction line, numbered as {1, 2, 3, ..., i}, then the number of canopy rows is N = i / 2;

[0074] (3312) Assuming the number of rows in each cell is m, the positions of the retained horizontal traction lines are {2×m×j+1, 2×m×(j+1), j∈(0, N-1)};

[0075] (3313) There are two reserved traction lines between two adjacent cells. The center line of the cell interval is further determined by taking the average of the vertical coordinates of the two horizontal traction lines while keeping the horizontal coordinate unchanged.

[0076] (332) The vertical traction line is generated based on the cell column detection and cell row length. The specific steps are as follows:

[0077] (3321) Count the number of canopy pixels in each column of the merged image, and use the column whose canopy pixel gradient change is closest to 0 as the preliminary vertical traction line, which is numbered as {1, 2, 3, ..., x}; then retain the vertical traction line of each cell based on the prior knowledge of the cell row length;

[0078] (3322) Let the length of the cell row be y, the horizontal row spacing be z, and the distance from each vertical pull line to the first vertical pull line be D = {d1, d2, ..., dx}, then the number of cells is K = dx / (y+z);

[0079] (3323) Filter the original vertical traction line, the filtering condition is n×(y+z) <dx<n×(y+z)+y,n∈(0,K-1);

[0080] (3324) Two traction lines are retained between two adjacent cells. The center line of the cell interval is further determined by averaging the horizontal coordinates of the two perpendicular traction lines while keeping the vertical coordinates unchanged.

[0081] (333) Pull line optimization: Ensure that the canopy of a single variety will not be segmented into different small areas. Specifically, for each intersection pixel p of the pull line and the canopy semantic segmentation map, evaluate whether it belongs to the canopy p = 1 or the background p = 0 pixel by pixel; if p = 0, it means that the pull line is between the small area intervals; if p = 1, it means that the pull line intersects with the small area canopy, and it is necessary to determine the appropriate distance and direction based on its relative position to the two nearest pull lines, and move the pixel to adapt to the small area interval. Figure 3 As shown in e, when adjusting the horizontal traction line, the number of canopy rows R from the intersecting horizontal traction line to its upper and lower adjacent horizontal traction lines is calculated. up and R down To determine whether the moving direction is upward or downward. The method for determining the number of canopy rows is consistent with the above step (3311). up >R down , the horizontal traction line moves up, and the moving distance is the number of pixels between the intersection pixel point c = 1 and its adjacent canopy interval c = 0, and vice versa. When adjusting the vertical traction line, if Figure 3 f shows the number of canopy pixels (P left , P right ) calculates the moving direction as left or right. left >P right , then the vertical traction line moves to the right, and the moving distance is P right, and vice versa. This process optimizes the intersecting pixels, and other pixels on the line remain unchanged, and finally the intersecting line is optimized into a continuous line segment.

[0082] The semantic segmentation map is divided into horizontal sub-maps using optimized horizontal traction lines, and each horizontal sub-map is vertically divided into individual cells based on optimized vertical traction lines, such as Figure 3 g. At the same time, the segmented image is rotated back to its original direction using the rotation angle determined in step (32) to complete the cell segmentation.

[0083] In this embodiment, during the process of canopy semantic segmentation, STANet uses the SoyUAV2022SY dataset and PyTorch 1.11.0 framework for training, validation, and testing of its own accuracy. STANet was trained using the Adam optimizer with a cross entropy loss function. The training process spanned 200 iterations, with a batch size of 4 and an initial learning rate of 0.001. If the validation loss did not improve for 5 consecutive cycles, the learning rate was reduced by 50%. At the same time, the minimum learning rate was capped at 6.25×10-5, and the F1 score was used to evaluate the accuracy of model validation. The best test model was selected based on the highest F1 achieved on the validation dataset, and all experiments were conducted on a Windows server equipped with an NVIDIA GeForce GTX 3060 GPU (12GB memory), a 6-core 2.50GHz CPU, and 16GB RAM. The accuracy evaluation of STANet was evaluated using overall accuracy (OA), intersection over union (IoU), precision (Pre), recall (Rec), and F1. These evaluation indicators range from 0 to 100%, and higher values ​​indicate better performance.

[0084] like Figure 4 As shown in the figure, STANet converged after 48 iterations of training. The average OA, IoU, Pre, Rec and F1 of STANet in soybean canopy semantic segmentation during the whole growth period reached 94.47%, 85.43%, 90.75%, 93.34% and 91.89%, respectively. In all periods, OA was greater than 90%, and between DAS_46 and DAS_61, IoU and F1 were greater than 88% and 93%, respectively. According to the results of canopy semantic segmentation, the average OA, IoU, Pre, Rec and F1 of TLA in plot segmentation during the whole growth period reached 97.21%, 93.31%, 95.00%, 96.42% and 95.13%, respectively. In all periods, OA was greater than 92%, and at the same time, IoU and F1 were always higher than 92% and 94% in the early and middle stages, respectively. Qualitative analysis shows that STANet-TLA can effectively perform canopy semantic segmentation and cell instance segmentation under different periods and different cell planting designs, with a cell recognition rate of up to 100%, such as Figure 5As shown in Figure 2, the qualitative analysis results of STANet-TLA’s accuracy at different days after seeding; Figure 6 As shown in Figure 2, the canopy semantic segmentation results of the SoyUAV2022SYb dataset 37 days after sowing; Figure 7 As shown in , this is the segmentation result of a plot with two rows of soybeans planted; Figure 8 The following is an example of a plot with three rows of soybeans. STANet can accurately segment the soybean canopy even when there are weeds in the early stage and yellow leaves in the later stage. Similarly, TLA shows particularly good performance throughout the growth stage, especially in the early and middle stages.

[0085] In this example, the segmentation accuracy of STANet is compared with that of eight existing advanced networks, such as Fig. 9 As shown in Figure 2, the segmentation accuracy of representative areas with CC less than 30%, between 30-70%, and greater than 70% was qualitatively analyzed at different stages. These existing networks are all implemented in PyTorch, and all networks are trained and tested using the same dataset, evaluation criteria, and experimental environment as STANet.

[0086] The comparison results of STANet models are shown in Table 1.

[0087] Table 1: Segmentation accuracy comparison results

[0088]

[0089] This example analyzes the accuracy of STANet's migration in different years, locations, and data types through a transfer learning strategy, and directly evaluates the performance of STANet on three target datasets using a model trained on the SoyUAV2022SY dataset. For each target dataset, the pre-trained model trained on the SoyUAV2022SY dataset is fine-tuned, and the fine-tuned model is used to evaluate the migration of STANet on each target dataset.

[0090] The results of the transferability analysis of the STANet model in different years, locations, and data types are shown in Table 2, where a check mark (√) indicates that STANet was trained or further fine-tuned based on the dataset.

[0091] Table 2: Mobility analysis results

[0092]

Claims

1. A soybean breeding plot instance segmentation method based on deep learning and prior knowledge, characterized in that: The method comprises the following steps: (1) Use drones to collect data and construct a time series dataset, which is divided into training set, validation set, and test set; (2) Through the spatiotemporal feature alignment network STANet model, the time series images are semantically segmented to capture multi-scale spatial information and time series information, separate the soybean canopy from the background, perform model training, and improve segmentation accuracy; (3) Based on the canopy semantic segmentation map, the traction line aggregation method TLA is used to introduce the prior knowledge of soybean plot planting design, generate horizontal and vertical traction lines, and perform plot instance segmentation.

2. The soybean breeding plot instance segmentation method based on deep learning and prior knowledge according to claim 1, characterized in that: The STANet model in step (2) is an encoder-decoder architecture, where the encoder includes a deformable convolution module DCN and a coordinate attention mechanism module CA, which are used to extract multi-scale and time series features of the canopy; The decoder consists of 4 consecutive spatiotemporal alignment modules STAM. The first spatiotemporal alignment generates a multi-scale feature by horizontally connecting and fusing the feature map of the last layer of the encoder and its last moment feature. The multi-scale feature map is upsampled by a factor of 2 through the nearest neighbor interpolation strategy. The final feature map uses a 1×1 convolutional layer to generate a canopy semantic segmentation mask.

3. The soybean breeding plot instance segmentation method based on deep learning and prior knowledge according to claim 2, characterized in that: The encoder uses ResNet101 as the backbone network, and the input is a time series image. The features generated after processing by the deformable convolution module are then processed by the coordinate attention mechanism module. ResNet101 encodes the processed features to generate time series features.

4. The soybean breeding plot instance segmentation method based on deep learning and prior knowledge according to claim 2, characterized in that: The coordinate attention mechanism module includes coordinate information embedding and coordinate attention generation. The coordinate information embedding decomposes the input two-dimensional features into two one-dimensional features in the horizontal and vertical directions through average pooling, and the coordinate attention is generated by the above two one-dimensional features.

5. The soybean breeding plot instance segmentation method based on deep learning and prior knowledge according to claim 2, characterized in that: The spatiotemporal alignment module includes a feature alignment module and a convolutional long short-term memory module. It takes time series features and multi-scale features as inputs, spatially aligns and fuses the horizontally connected time series features with the multi-scale features through the feature alignment module, and inputs the aligned time features into the convolutional long short-term memory to extract the time series features.

6. The soybean breeding plot instance segmentation method based on deep learning and prior knowledge according to claim 2, characterized in that: The second spatiotemporal alignment in the spatiotemporal alignment module fuses the upsampled feature map and the horizontally connected time series feature map with the same spatial size to generate multi-scale features, the third spatiotemporal alignment fuses the upsampled feature map and the horizontally connected time series feature map with the same spatial size to generate multi-scale features, and the fourth spatiotemporal alignment fuses the upsampled feature map and the horizontally connected time series feature map with the same spatial size to generate multi-scale features.

7. The soybean breeding plot instance segmentation method based on deep learning and prior knowledge according to claim 5, characterized in that: The feature alignment module takes the horizontally connected time series features and multi-scale feature maps as input, extracts the key weights of each channel of the time series features through average pooling, 1×1 convolution and sigmoid activation function, retains the key features and removes irrelevant features; Multiply the key weights with the input time series features and add the product to the input time series features to obtain the key features; The key feature channels are adjusted through 1×1 convolution to generate feature selection results, aligning the selected features with the multi-scale features.

8. The soybean breeding plot instance segmentation method based on deep learning and prior knowledge according to claim 5, characterized in that: The convolutional long short-term memory module is used to process time series information and establish long-distance dependencies. The input is a three-dimensional tensor with a time dimension, and the output is a multi-scale feature map.

9. The soybean breeding plot instance segmentation method based on deep learning and prior knowledge according to claim 5, characterized in that: The convolutional long short-term memory constructs time series data for time series images in the following steps: selecting an RGB image to fill the position at time t, and saving the segmentation result of the selected image as a matrix label, and then selecting an RGB image before or at the time t to fill the position at time t-1, and then filling the position at time t-2 in sequence.

10. The soybean breeding plot instance segmentation method based on deep learning and prior knowledge according to claim 1, characterized in that: The traction line polymerization method in step (3) specifically comprises the following steps: (31) Merge the canopy semantic segmentation results into a unified semantic segmentation map for the entire study area; (32) Rotate the merged semantic segmentation map so that the planting row direction is aligned with the horizontal coordinate system of the image; (33) Generate horizontal traction lines and vertical traction lines, and optimize the traction lines; (34) The semantic segmentation map is divided into horizontal sub-maps using optimized horizontal traction lines, each horizontal sub-map is vertically divided into individual cells based on optimized vertical traction lines, and the segmented image is rotated back to its original orientation using the rotation angle determined in step (32) to complete the cell segmentation.