Crop breeding plot segmentation method based on rotating target detection

By building a rotary segmentation network model, combining the directional RPN module and the RotatedAlignment module, the problem of difficult to segment the rotation angle breeding cells in the existing technology is solved, and a higher segmentation accuracy and automation level is achieved.

CN120107599APending Publication Date: 2025-06-06NANJING AGRICULTURAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510300911.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to accurately segment crop breeding communities with rotation angles, resulting in low segmentation accuracy and unable to meet the needs of high-throughput phenotypic identification.

Method used

Using a method based on rotation object detection, a rotation segmentation network model is constructed, combined with the directional RPN module and the RotatedAlignment module, accurate detection and segmentation of breeding cells at any angle is achieved.

Benefits of technology

It improves the accuracy, robustness and automation level of crop breeding cells, and can extract breeding experimental cells more accurately to meet the needs of high-throughput phenotype identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107599A_ABST
    Figure CN120107599A_ABST
Patent Text Reader

Abstract

The invention discloses a crop breeding cell segmentation method based on rotation target detection, and the method comprises the steps: collecting data, constructing a data set, constructing a rotation segmentation network model based on a Mask RCNN network architecture, training the rotation segmentation network model through a training set, and obtaining a rotation target detection result. The rotation segmentation network model comprises an SW-GC Block module fusing a global context attention mechanism, a directional RPN module, a RotatedAlliance module and an output module; inputting the breeding plot image into the trained rotary segmentation network model to realize detection and segmentation of the breeding plot; an adaptive rotation target detection mechanism is introduced for a traditional instance segmentation algorithm by using a directional RPN module and a RotatedAlliance module, so that the algorithm can generate a candidate frame at any angle, the problem that the traditional algorithm is insufficient in rotation target segmentation performance is solved, an SW-GC Block module fused with a global context attention mechanism is designed, and the algorithm is more efficient and efficient. The extraction of global semantic information is enhanced, the loss of feature information is reduced, and the segmentation precision under a complex background is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a crop breeding plot segmentation method, and in particular to a crop breeding plot segmentation method based on rotating target detection. Background Art

[0002] Phenotyping is a key step in the selection of excellent varieties. It usually relies on manual measurement, which is time-consuming, labor-intensive and subjective, greatly limiting the efficiency of breeding. In recent years, the rapid development of crop high-throughput phenotyping technology uses different sensors and analysis platforms to collect and analyze multi-dimensional phenotypic data with high efficiency and precision, effectively accelerating the breeding process. The drone phenotyping platform equipped with RGB cameras is widely used due to its advantages such as large monitoring range, high efficiency and low cost.

[0003] In recent years, researchers have used different sensors and analysis platforms to efficiently and accurately collect and analyze multi-dimensional phenotypic data of crops, thereby significantly shortening the breeding process. Rapid and accurate extraction of crop breeding plots is a key step in conducting high-throughput phenotypic identification based on drone platforms. However, the automatic extraction of individual breeding plots still faces three major challenges: first, the color and texture of the exposed soil inside the plot before the canopy closes are similar to the characteristic information of the gap area between the plots; second, the color of crops is similar to the weed background color in the early growth stage, and similar to the soil background color in the later stage, making it difficult to identify; third, the irregular layout of field planting may cause the plot in the drone RGB to have an angle deviation from the due south and due north. These challenges make it difficult for traditional segmentation algorithms based on color and edges to achieve ideal results.

[0004] In the prior art, deep learning technology, especially semantic segmentation algorithms represented by convolutional neural networks (CNNs), have become ideal tools for plot segmentation because they can automatically extract multi-dimensional image features; among them, instance segmentation algorithms based on horizontal target detection are widely used in the field because they can accurately detect and segment individual objects in images. However, when faced with breeding plots with irregular planting directions in the field, this method still has obvious limitations and cannot accurately segment plots with rotation angles, resulting in low segmentation accuracy and unable to meet the needs of high-throughput phenotyping. Summary of the invention

[0005] Purpose of the invention: The purpose of the present invention is to provide a crop breeding plot segmentation method based on rotating target detection to accurately extract breeding experimental plots through UAV RGB images, thereby improving the accuracy, robustness and automation level of crop breeding plot segmentation.

[0006] Technical solution: The crop breeding plot segmentation method based on rotating target detection described in the present invention comprises the following steps:

[0007] (1) constructing a rotating target detection dataset for crop breeding plots, annotating the dataset, and dividing the dataset into a training set, a test set, and a validation set according to proportion and growth period;

[0008] (2) constructing a rotation segmentation network model based on the MaskRCNN network architecture, and training the rotation segmentation network model through the training set, wherein the rotation segmentation network model includes a multi-scale feature extraction backbone network, a SW-GC Block module integrating a global context attention mechanism, a directional RPN module, a RotatedAlignment module, and an output module;

[0009] (3) The breeding plot image is input into the trained rotation segmentation network model to realize the detection and segmentation of the breeding plot.

[0010] Preferably, the specific method of constructing the data set in step 1 is:

[0011] (11) Use a drone platform equipped with an RGB camera to collect single image data of the breeding plot, and stitch the single orthophotos to generate an orthophoto of the entire field based on GPS coordinate information and GCP correction;

[0012] (12) Crop the orthophoto of the entire field, add instance segmentation mask labels to each breeding plot in the image, and annotate them through manual visual interpretation;

[0013] (13) The images and the corresponding annotation masks are cropped by moving window cropping. After data enhancement, the images are divided into training set, validation set, and test set according to proportion and growth period.

[0014] Preferably, a multi-task loss function optimization algorithm is used to train the directional RPN module and the output module described in step 2.

[0015] Preferably, the formula for training the directional RPN module is:

[0016]

[0017] Among them, N represents the total number of samples in a mini batch, the default value is 256, i represents the index of the anchor, and L cls represents the cross entropy loss of anchor classification, p i Indicates the probability that the i-th anchor belongs to the foreground, represents the true label of the i-th anchor, i.e. background or foreground; t i represents the regression parameters of the i-th anchor point bounding box, represents the offset of the ground-truth bounding box of the i-th anchor point, L regRepresents the Smooth L1 loss for bounding box regression.

[0018] Preferably, the formula for training the output module is:

[0019] L head =L cls (p,p * )+L reg (t k ,t * )+L seg (m k ,m * );

[0020] Among them, L seg represents the cross entropy loss of segmentation, and the classification score vector p = (p 0 ,p 1 ,…,p k ), generated by the first branch of the output head, represents the probability distribution of K+1 categories, p * represents the true category label, t k represents the regression offset vector for each of the K object categories, t * represents the regression parameters of the ground-truth object bounding box, and for each of the K object categories, the segmentation mask m k Generated by the second branch of the output head, m k Represents the corresponding segmentation label of each category, and K is 1.

[0021] Preferably, the multi-scale feature extraction backbone network in step 2 includes 4 stages, and each stage includes a SW-GCBlock module that integrates the global context attention mechanism.

[0022] Preferably, the output module in step 2 also includes a detection branch for outputting candidate region category probabilities, detection box regression parameters and offset angles, and a segmentation branch for outputting a binary mask of the candidate region, representing an instance segmentation result of the current region.

[0023] Preferably, the process of constructing the SW-GC Block module integrating the global context attention mechanism in step 2 is as follows:

[0024] Learning global features: transform the input features through a 1x1 convolutional layer and use a global average pooling layer to capture global context information;

[0025] Learning different channel weights: transform the global feature descriptor through a 1x1 convolution layer, normalize and nonlinearly transform the features using layer normalization and ReLU activation function, and finally restore the number of channels through a 1x1 convolution layer;

[0026] The global context information is fused with the original input features by using the addition operation, and the enhanced feature map is subjected to the original Swin Transformer Block operation: the feature map enhanced with global context information is input into LN, and the SW-MSA operation is performed after layer normalization. The output of the SW-MSA is added to the input before layer normalization to form a residual connection. The result of the residual connection is normalized for the second time, and the normalized features are sent to the multi-layer perceptron to introduce nonlinear processing and feature transformation to enhance the learning and representation capabilities of the algorithm; the output of the MLP is added again to the features before the second layer normalization to form another residual connection to construct a complete SW-GC Block.

[0027] Preferably, the specific method for detecting and segmenting the breeding plot in step 3 is:

[0028] (31) extracting semantic feature maps of different levels through the multi-scale feature extraction backbone network, effectively fusing high-level semantic information into low-level feature maps, and obtaining a fused multi-scale feature map through a maximum pooling operation;

[0029] (32) inputting the fused multi-scale feature map into the directional RPN module, and then outputting directional target candidate boxes of different sizes and areas;

[0030] (33) extracting a feature area corresponding to the oriented target candidate box from the multi-scale feature map based on the RotatedAlignment module and mapping it to a fixed size;

[0031] (34) The fixed-size directional target candidate box region features are input into the output module to realize the detection and segmentation of the breeding plot.

[0032] Preferably, the output of the directional target candidate frame in step 32 specifically includes: inputting the fused multi-scale feature map into the directional RPN module, at each feature level, the directional RPN module adds a convolution layer for the regression branch and the classification branch, and allocates horizontal anchor points with different areas and sizes at each spatial position of the feature map, and then outputs the offset relative to the center position, width, height and angle of each anchor point through the regression branch, and generates the directional target candidate frame from the horizontal anchor point by decoding the output of the regression branch and using the additional offset based on the inheritance of the horizontal regression mechanism.

[0033] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: 1. By combining the algorithms of the directional RPN module and the RotatedAlignment module, an adaptive rotation target detection mechanism is introduced, so that the algorithm can generate candidate frames at any angle, solve the problem that the traditional algorithm has insufficient performance in segmenting rotated targets, and improve the accuracy, robustness and automation level of crop breeding plot segmentation; 2. By integrating the Swin Transformer module with the global context attention mechanism, the algorithm's extraction of global semantic information is enhanced, the loss of feature information is reduced, the segmentation accuracy under complex backgrounds is improved, and the problem of information loss in plot segmentation is effectively solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the process of the present invention;

[0035] Figure 2 This is a schematic diagram of the rotation segmentation network of the present invention;

[0036] Figure 3 It is a schematic diagram of a comparative experiment between the present invention and five instance segmentation algorithms based on horizontal target detection. DETAILED DESCRIPTION

[0037] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.

[0038] The study area is located at the Baima Experimental Base of Nanjing Agricultural University, Baima Town, Lishui District, Nanjing City, Jiangsu Province (119°18'E, 31°62'N). From 2020 to 2021, a total of 960 wheat breeding plots were planted, including 160 wheat varieties. The size of each plot is 1.5m×1.5m, the number of planting rows is 6, the row spacing is 0.25m, and the spacing between plots is 0.5m. DJIMavic 2pro (camera resolution is 5472 pixels×3648 pixels) was used as the flight platform, the flight altitude was 20m, the heading overlap rate and the lateral overlap rate were both 80%, and the camera angle was 80°. From February 22, 2021 to May 8, 2021, wheat at different growth stages was photographed by cross-flight, once a week, and all were taken under clear and windless conditions.

[0039] (1) Single image data of the breeding plot was collected in the flight mode, and then the images were spliced ​​and GCP corrected using Metashape software. The useless background outside the study area in the image was cropped out using ArcMap software. Finally, a total of 10 orthophoto RGB images of the plot at different growth stages were obtained.

[0040] ArcMap software is used to manually label images and construct datasets through visual interpretation, which mainly includes four steps:

[0041] In the first step, 3 images of different growth periods were selected from the 10 images to construct the test set, and the remaining 7 images were used to construct the training set and validation set.

[0042] In the second step, ArcMap software was used to crop the image using a moving window. Although the edge feature information of the community was important, the amount of data was large enough, so the window's X and Y moving steps were set to 0 (no image overlap).

[0043] In the third step, in order to take into account the size of the breeding plot and the size of the training video memory in the image, the image cropping size is selected as 1024 pixels × 1024 pixels. The cropping methods of the training set, validation set, and test set are all consistent. The training set and validation set are divided in a ratio of 7:3 to ensure that the trained network has good robustness. The remaining 3 images are used as test sets. After obtaining the initial training set, validation set, and test set, in order to explore the impact of the planting direction of the breeding plot on the network, the image needs to be rotated. First, each image in the data set is adjusted to an image with no tilt angle in the plot, which is defined as 0°, and then rotated clockwise from 0° to 7 angles (0°, 15°, 30°, 45°, 60°, 75°, 90°) at intervals of 15°. The rotation method is performed through the OpenCV library. The width and height of the rotated image are between 1000-1500 pixels.

[0044] In the fourth step, the mask of each breeding plot instance in each image was split and made into JSON format labels required for training. After data enhancement, it was divided into training set, validation set and test set according to proportion and growth period.

[0045] (2) A rotation segmentation network model is constructed based on the MaskRCNN network architecture. The rotation segmentation network includes a multi-scale feature extraction backbone network, a SW-GC Block module that integrates the global context attention mechanism, a directional RPN module, a RotatedAlignment module, and an output module.

[0046] In order to avoid the problem of unstable training process and slow convergence caused by random initialization of the network, the pre-trained weights on the ImageNet1K dataset were used to initialize the network parameters; then two NVIDIA GeForceRTX 2080Ti GPUs (11GB video memory, Intel(R)Xeon(R)Platinum 8255C CPU and 80G memory) were used to train the network; the batch size was 8, the epoch was set to 12, the learning rate was set to 0.01 and reduced by 10 times in the 8th and 11th epochs; the optimizer was the SGD algorithm, the momentum was 0.9, and the weight decay was 0.0001.

[0047] A multi-task loss function optimization algorithm is used to train the rotation segmentation network model, including L for training the directional RPN module. rpn and the L of the training output module head .

[0048] Among them, L rpn Defined as:

[0049]

[0050] Among them, N cls Indicates the total number of samples in a mini batch. The default value is 256. i indicates the index of the anchor. L cls represents the cross entropy loss of anchor classification, p i Indicates the probability that the i-th anchor belongs to the foreground; represents the true label of the i-th anchor, i.e. background or foreground; L reg represents the Smooth L1 loss for bounding box regression, t i represents the regression parameters of the i-th anchor point bounding box, Represents the offset of the ground-truth bounding box of the i-th anchor point.

[0051] L head Defined as:

[0052] L head =L cls (p,p * )+L reg (t k ,t * )+L seg (m k ,m * );

[0053] L cls represents the cross entropy loss of proposal classification, L reg represents the SmoothL1 loss for bounding box regression, L seg represents the cross entropy loss of segmentation, and the classification score vector p = (p 0 ,p 1 ,…,p k ), generated by the first branch of the output head, represents the probability distribution of K+1 categories; p * represents the true category label, t k represents the regression offset vector for each of the K object categories, t * represents the regression parameters of the ground-truth object bounding box. For each of the K object categories, the segmentation mask m kGenerated by the second branch of the output head, m k Represents the corresponding segmentation label of each category, and K is 1.

[0054] Scale feature extraction backbone network: In order to solve the problem of over-segmentation of wheat before ridge closure due to insufficient global semantic information, a multi-scale feature extraction backbone network is designed, which consists of the SW-GCBlock module integrating the global context attention mechanism.

[0055] The input image is first divided into 4×4 image blocks through Patch Partition, converting the image from traditional pixel representation to a series of patch representations that can be processed by the algorithm, in preparation for subsequent feature extraction and processing; then, each 4×4×3 (3 is the number of image channels) patch will be sent to the LinearEmbedding layer, mapped to a high-dimensional feature vector through linear transformation, and the feature map size becomes H / 4×W / 4×96, which converts the original pixel value of each patch into a high-dimensional representation that can be further processed by the algorithm; then the feature map is sent to the SW-GC Block in Stage 1 for feature extraction. After two repeated SW-GC Block extraction processes, the feature extraction of Stage 1 is completed; then the feature map is sent to PatchMerging in Stage 2. Patch Merging is between different stages of the algorithm, aiming to reduce the spatial dimension of the feature map and increase the feature dimension, thereby realizing hierarchical feature representation, enabling the algorithm to capture information of different scales from local to global. The feature map size is reshaped to H / 8×W / 8×192 and then passes through two SW-GC Blocks to complete the feature extraction of Stage 2; the feature processing flow of Stage 3 and Stage 4 is similar to that of Stage 2. The number of stacked SW-GC Blocks is 2 and 6 respectively, and the feature map size is reshaped to H / 16×W / 16×384 and H / 32×W / 32×768 respectively.

[0056] SW-GC Block module integrating global contextual attention mechanism: The context Block module aims to integrate global contextual attention into the original Swin TransformerBlock to enhance the global modeling capability.

[0057] First, the global features are learned, and the input features are transformed through a 1x1 convolution layer to reduce the feature dimension, and then the global average pooling (GAP) layer is used to capture the global context information; then the weights of different channels are learned, and the global feature descriptors are transformed through a 1x1 convolution layer to reduce the number of channels; then the dependencies between channels are learned, and the features are normalized and nonlinearly transformed using layer normalization and ReLU activation functions, and then the number of channels is restored through a 1x1 convolution layer; finally, the global context information is fused with the original input features through an addition operation.

[0058] The enhanced feature map is subjected to the original Swin TransformerBlock operation. First, the feature map enhanced with global context information is input into LN, and then the SW-MSA (Shifted Window Multi-head Self-Attention) operation is performed after layer normalization. This is a self-attention mechanism that aims to calculate attention in a local window of fixed size. In order to cover different areas of the image, adjacent SWBlocks will shift their windows in turn. Then the output of SW-MSA is added to the input before layer normalization to form a residual connection. This helps prevent the gradient vanishing problem in deep networks. Then, the result of the residual connection is normalized for the second time, and the normalized features are sent to the Multilayer Perceptron (MLP) to introduce nonlinear processing and feature transformation to enhance the learning and representation capabilities of the algorithm. Finally, the output of the MLP is added to the features before the second layer normalization to form another residual connection to achieve a complete SW-GCBlock.

[0059] Directional RPN module: To address the problem that traditional instance segmentation algorithms cannot accurately segment breeding plots in any direction, a directional RPN module that is adaptive to the cell angle is designed. The directional RPN module aims to generate high-quality directional region proposals.

[0060] The directional RPN module takes the feature map fused from the multi-scale feature pyramid as input. At each feature level, the directional RPN adds a 3×3 convolutional layer and two 1×1 convolutional layers (one for the regression branch and the other for the classification branch). Then, the horizontal anchor points are allocated, and three anchor points with different areas (32 2 , 64 2 , 128 2 , 256 2 , 512 2 ) and horizontal anchor points with aspect ratios (1:2, 1:1, 2:1); then after the regression branch, the output is the offset δ = (δ x ,δy ,δ w ,δ h ,δ α ,δ β ), including the center position (δ x ,δ y ), width (δ w ) and height (δ h ) and the angular offset (δ α ,δ β ), so there are 6A outputs in the end (A is the number of horizontal anchors, 1:2, 1:1, 2:1); finally, by decoding the output of the regression branch, on the basis of inheriting the horizontal regression mechanism, the additional offset is used to generate directional proposals from the horizontal anchors.

[0061] RotatedAlignment module: To address the problem that oriented targets cannot be aligned, which brings difficulties to subsequent segmentation and detection, the RotatedAlignment module is introduced to correct and align rotated targets.

[0062] The RotatedAlignment module first calculates the coordinates of the four corner points of the rotated box, then maps the four corner points of the rotated box to the coordinate space on the feature map, and scales it according to the size of the feature map. Finally, based on the coordinates of the four corner points of the rotated box, bilinear interpolation is used to calculate the value of each pixel, and finally the features that are precisely aligned with the rotated candidate area are obtained for detection (BoxHead) and segmentation (MaskHead).

[0063] Output modules: including Boxhead for detection and Maskhead for segmentation.

[0064] For BoxHead, the aligned feature map is fed into a 3x3 convolutional layer to reduce the channel dimension and extract richer semantic information. Then, the spatial feature map is further extracted through two 3x3 convolutional layers. Batch normalization and ReLU activation function are used after each convolutional layer. Finally, two 1x1 convolutional layers are used to generate the prediction score of each category and the regression of the rotated bounding box (offset value); for Maskhead, the aligned feature map is stacked with four 3x3 convolutional layers with 128 channels to extract high-level semantic features. Then, two 1x1 convolutional layers are used to predict the downsampled mask (14x14 size) of each instance. Finally, the predicted mask is bilinearly upsampled to the original image size to obtain the final segmentation mask.

[0065] (3) Breeding plot segmentation: input the image patches into the rotation segmentation network model in sequence to extract the single breeding plot in the whole field orthophoto, which specifically includes the following steps:

[0066] The first step is to extract multi-scale features

[0067] First, the small image is divided into 4×4 image patches through Patch Partition and Linearizending operations, the number of patch channels is adjusted to convert into embedding vectors, and position encoding is added to each embedding vector to perceive the relative position relationship between different patches.

[0068] Then, the feature map stacked by the embedding vectors is sent to the multi-scale feature extraction backbone network containing 4 stages to extract semantic feature maps of different levels. Each stage includes the SW-GCBlock module that integrates the global context attention mechanism. Patch Merging operations are added between different stages to reduce the spatial dimension of the feature map and increase the feature dimension. For the feature map of each semantic level, the number of channels of the feature map of the corresponding level of the backbone network is adjusted through 1x1 convolution to match the feature map of the top-down path. Then the two are added together to effectively integrate the high-level semantic information into the low-level feature map. The feature map of C4 is subjected to the maximum pooling operation to obtain P5, and finally 5 fused feature maps {P1, P2, P3, P4, P5} are obtained.

[0069] The second step is to generate directional target candidate boxes

[0070] To address the problem that traditional instance segmentation algorithms cannot achieve accurate segmentation of breeding plots in any direction, a directional RPN module with adaptive cell angle is introduced. It takes the multi-scale feature map {P1, P2, P3, P4, P5} output by the feature pyramid as input and outputs directional target candidate boxes of different sizes and areas.

[0071] Based on the RotatedAlignment module, the feature area corresponding to the oriented target candidate box is extracted from the multi-scale feature map, and then mapped to a fixed size for subsequent detection and segmentation.

[0072] Step 3: Detection and segmentation

[0073] The fixed-size directional target candidate box region features are input into the detection branch and the segmentation branch respectively. The detection branch outputs the category probability, detection box regression parameters and offset angle of each candidate region, and the segmentation branch outputs the binary mask of each candidate region, which represents the instance segmentation result of the current region.

[0074] Step 4: Compare test results

[0075] Compared with five instance segmentation algorithms based on horizontal target detection (MaskRCNN, MS RCNN, PointRend, Cascade MR and HTC), the present invention can more accurately extract wheat breeding plots.

[0076] The AP@0.5, F1-score, Accuracy and IoU indicators of the present invention are 0.917, 0.959, 0.966 and 0.912 respectively. In terms of AP@0.5 indicator, OSNet is 3.2%, 3.0%, 2.1%, 3.9% and 1.5% higher than MaskRCNN, MS RCNN, PointRend, Cascade MaskRCNN and HTC respectively; in terms of F1-score indicator, OSNet is 1.2%, 1.6%, 1.1%, 1.6% and 1.2% higher than Mask RCNN, MSRCNN, PointRend, Cascade MaskRCNN and HTC respectively; in terms of Accuracy indicator, OSNet is 1.2%, 1.6%, 1.1%, 1.6% and 1.2% higher than MaskRCNN, MS RCNN, PointRend,

[0077] Cascade Mask RCNN and HTC are 1.1%, 1.3%, 1.0%, 1.3%, and 1% higher respectively; in terms of IoU indicators,

[0078] OSNet is 1.3%, 2%, 1.1%, 1.9%, and 1.3% higher than Mask RCNN, MS RCNN, PointRend, Cascade Mask RCNN, and HTC, respectively.

Claims

1. A crop breeding plot segmentation method based on rotating target detection, characterized in that: The following steps are involved: (1) constructing a rotating target detection dataset for crop breeding plots, annotating the dataset, and dividing the dataset into a training set, a test set, and a validation set according to proportion and growth period; (2) constructing a rotation segmentation network model based on the Mask RCNN network architecture, and training the rotation segmentation network model through the training set, wherein the rotation segmentation network model includes a multi-scale feature extraction backbone network, a SW-GC Block module integrating a global context attention mechanism, a directional RPN module, a Rotated Alignment module, and an output module; (3) The breeding plot image is input into the trained rotation segmentation network model to realize the detection and segmentation of the breeding plot.

2. The crop breeding plot segmentation method according to claim 1, characterized in that: The specific method for constructing the data set in step 1 is: (11) Use a drone platform equipped with an RGB camera to collect single image data of the breeding plot, and stitch the single orthophotos to generate an orthophoto of the entire field based on GPS coordinate information and GCP correction; (12) Crop the orthophoto of the entire field, add instance segmentation mask labels to each breeding plot in the image, and annotate them through manual visual interpretation; (13) The images and the corresponding annotation masks are cropped by moving window cropping. After data enhancement, the images are divided into training set, validation set, and test set according to proportion and growth period.

3. The crop breeding plot segmentation method according to claim 1, characterized in that: The directional RPN module and output module described in step 2 are trained using a multi-task loss function optimization algorithm.

4. The crop breeding plot segmentation method according to claim 3, characterized in that: The loss function formula for training the directional RPN module is: Among them, N represents the total number of samples in a mini batch, the default value is 256, i represents the index of the anchor, and L cls represents the cross entropy loss of anchor classification, p i Indicates the probability that the i-th anchor belongs to the foreground, represents the true label of the i-th anchor, i.e. background or foreground; t i represents the regression parameters of the i-th anchor point bounding box, represents the offset of the ground-truth bounding box of the i-th anchor point, L reg Represents the Smooth L1 loss for bounding box regression.

5. The crop breeding plot segmentation method according to claim 3, characterized in that: The loss function formula for training the output module is: L head =L cls (p,p*)+L reg (t k ,t * )+L seg (m k ,m * ); Among them, L seg represents the cross entropy loss of segmentation, and the classification score vector p = (p0, p1, ..., p k ), generated by the first branch of the output head, represents the probability distribution of K+1 categories, p * represents the true category label, t k represents the regression offset vector for each of the K object categories, t * represents the regression parameters of the ground-truth object bounding box, and for each of the K object categories, the segmentation mask m k Generated by the second branch of the output head, m k Represents the corresponding segmentation label of each category, and K is 1.

6. The crop breeding plot segmentation method according to claim 1, characterized in that: The multi-scale feature extraction backbone network in step 2 includes 4 stages, each of which includes the SW-GC Block module that integrates the global context attention mechanism.

7. The crop breeding plot segmentation method according to claim 1, characterized in that: The output module in step 2 includes a detection branch for outputting the category probability of the candidate region, the regression parameters of the detection box and the offset angle, and a segmentation branch for outputting the binary mask of the candidate region, representing the instance segmentation result of the current region.

8. The crop breeding plot segmentation method according to claim 1, characterized in that: The construction process of the SW-GC Block module integrating the global context attention mechanism described in step 2 is as follows: Learning global features: transform the input features through a 1x1 convolutional layer and use a global average pooling layer to capture global context information; Learning different channel weights: transform the global feature descriptor through a 1x1 convolution layer, normalize and nonlinearly transform the features using layer normalization and ReLU activation function, and finally restore the number of channels through a 1x1 convolution layer; The global context information is fused with the original input features by using the addition operation, and the enhanced feature map is subjected to the original Swin Transformer Block operation: the feature map enhanced with global context information is input into LN, and the SW-MSA operation is performed after layer normalization. The output of the SW-MSA is added to the input before layer normalization to form a residual connection. The result of the residual connection is normalized for the second time, and the normalized features are sent to the multi-layer perceptron to introduce nonlinear processing and feature transformation to enhance the learning and representation capabilities of the algorithm; the output of the MLP is added again to the features before the second layer normalization to form another residual connection to construct a complete SW-GC Block.

9. The crop breeding plot segmentation method according to claim 1, characterized in that: The specific method for detecting and segmenting the breeding plot in step 3 is: (31) extracting semantic feature maps of different levels through the multi-scale feature extraction backbone network, effectively fusing high-level semantic information into low-level feature maps, and obtaining a fused multi-scale feature map through a maximum pooling operation; (32) inputting the fused multi-scale feature map into the directional RPN module, and then outputting directional target candidate boxes of different sizes and areas; (33) extracting a feature area corresponding to the oriented target candidate box from the multi-scale feature map based on the RotatedAlignment module and mapping it to a fixed size; (34) The fixed-size directional target candidate box region features are input into the output module to realize the detection and segmentation of the breeding plot.

10. The crop breeding plot segmentation method according to claim 8, characterized in that: The output of the directional target candidate frame described in step 32 specifically includes: inputting the fused multi-scale feature map into the directional RPN module, at each feature level, the directional RPN module adds a convolution layer for the regression branch and the classification branch, and allocates horizontal anchor points with different areas and sizes at each spatial position of the feature map. The regression branch then outputs the offset relative to the center position, width, height and angle of each anchor point, and generates a directional target candidate frame from the horizontal anchor point by decoding the output of the regression branch and inheriting the horizontal regression mechanism and using the additional offset.

Citation Information

Cited By

  • Deep learning-based intelligent detection method and system for color steel rooms along railway

    CN121214262A