Method for detecting phenotypic consistency of male sterile line of hybrid soybean
By improving the YOLOv5s model, adding the LSKNet attention mechanism and small object detection branch, and using the WIoU v3 loss function, the XLW-YOLO model was constructed, which solved the problem of low phenotype consistency detection efficiency of hybrid soybean male sterile lines, and achieved high-precision field phenotype detection, supporting large-scale hybrid breeding.
Patent Information
- Application Number
- CN202510455332.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the phenotypic consistency detection efficiency of hybrid soybean male sterile lines is low, and cannot meet the needs of large-scale hybrid breeding, hindering the utilization of hybrid advantages.
Based on the YOLOv5s model, combined with the LSKNet attention mechanism and small object detection branch, and using the WIoU v3 loss function, a phenotypic consistency detection model XLW-YOLO of the hybrid soybean male sterile line was constructed, and high-precision detection of combyl, leaves, hairs and flowers were achieved through image acquisition, data annotation and model training.
The accuracy and efficiency of phenotypic consistency detection of hybrid soybean male sterile lines is improved, and the phenotypic characteristics of hybrid soybeans can be detected in real time in the field to meet the needs of hybrid breeding.
Smart Images

Figure CN120472436A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting by using image data, and in particular to a method for detecting the phenotypic consistency of hybrid soybean male sterile lines. Background Art
[0002] Utilizing hybrid vigor in soybeans is a key technology for increasing soybean yields. Breeders utilize cytoplasmic-nuclear male sterile lines to develop soybean hybrids, but the challenge is maintaining homozygous male sterile lines. To improve the purity of soybean hybrids, manual removal of weeds is necessary during soybean growth. According to the "Procedure for Identification of Hybrid Soybean Cytoplasmic-nuclear Male Sterile Lines," phenotypic traits to be identified include leaf shape and the color of the hypocotyl, pubescence, and flower. While identification can be performed manually, its low efficiency hinders large-scale hybrid breeding and the utilization of heterosis. Therefore, utilizing deep learning technology to enable computerized identification of phenotypic consistency in field hybrid soybeans is crucial. Summary of the Invention
[0003] The present invention aims to overcome the deficiencies of the prior art and provides a method for detecting the phenotypic consistency of hybrid soybean male sterile lines.
[0004] The method for detecting phenotypic consistency of hybrid soybean male sterile lines of the present invention is achieved by the following steps:
[0005] (1) Original image acquisition
[0006] Visible light images of soybean plants growing naturally in the field were collected at different growth stages, including 1,000 hypocotyl images during the seedling stage, 1,000 leaf images during the vegetative growth stage, 1,000 fuzz images during the reproductive growth stage, and 1,000 flower images. The soybean plant image data was acquired using the Azure Kinect DK image acquisition sensor, with a resolution of 1280×720 pixels.
[0007] (2) Dataset annotation
[0008] The data samples obtained in step (1) were manually annotated using LabelImg software according to the annotation strategy to construct a dataset, and the annotated dataset was expanded using methods such as image flipping, translation, brightness adjustment, and noise addition using Python's OpenCV library to expand it to 8,000 image data. During the enhancement process, the bounding box coordinates of the XML annotation file were synchronously updated; the four datasets of hypocotyl, leaf, hair, and flower were divided into sub-training sets and sub-test sets at an 8:2 ratio, and then the four sub-training sets were merged into the total training set, and the four sub-test sets were merged into the total test set;
[0009] (3) Construction of a phenotypic consistency detection model for hybrid soybean male sterile lines
[0010] The total training set obtained in step (2) was used as the model input data, and the YOLOv5s model was used as the basic network structure. The LSKNet attention mechanism module was added to the layer before the SPPF of the Backbone network, and a small target detection branch XNet for the 160×160 resolution feature map was added to the Neck network. The CIoU loss function used in the YOLOv5s model was changed to the WIoU v3 loss function to establish the hybrid soybean male sterile line phenotypic consistency detection model XLW-YOLO.
[0011] (4) Detection of phenotypic consistency of hybrid soybean male sterile lines
[0012] By collecting images of soybean plants growing naturally in the field at different growth stages and inputting them into the detection model, the phenotypic consistency of hybrid soybean male sterile lines can be detected.
[0013] As a further improvement of the present invention, a small target detection branch XNet for 160×160 resolution feature maps is added to the Neck network of the YOLOv5s model in step (3). The specific process is as follows:
[0014] (1) For the 512-channel feature map with 80×80 resolution in the Neck network, a 256-channel feature map is generated through the C3 module;
[0015] (2) Perform 1×1 convolution compression on the optimized 256-channel feature map to generate a 128-channel feature map;
[0016] (3) Upsample the compressed feature map by a factor of 2 to generate a 128-channel feature map with a resolution of 160×160;
[0017] (4) The upsampled feature map is concatenated with the 160×160 resolution, 128-channel feature map output by the second layer of the Backbone network according to the channel dimension to generate a 256-channel fused feature map;
[0018] (5) The fused feature map is processed by the C3 module and then input into the Head network for prediction.
[0019] As a further improvement of the present invention, the LSKNet attention mechanism is added in step (3), which decomposes the convolution process of the input feature X into a series of deep convolutions through large kernel convolutions. In order to adjust the number of channels to match the same output dimension, F is used. 1×1 The convolutional layer obtains different kernel features The calculation is shown in formulas (1)-(3).
[0020] U0=X,U i+1 =F i dw (U i ) (1)
[0021]
[0022] Among them, F i dw (·) is a kernel with k i and extension d i Depthwise convolution;
[0023] The spatial feature descriptor SA is obtained by performing channel connection and pooling operations on different kernel features, using F 2→N Convolutional layer and Sigmoid activation function to obtain spatial selection mask The spatial selection mask is used to weight different kernel features, and the attention feature S is obtained through the convolution layer F. The input feature X is element-wise multiplied with the attention feature S to obtain the final output Y. The calculation is shown in formulas (4)-(8):
[0024]
[0025] Y=X·S (8)
[0026] Among them, P avg (·) and P max (·) is average pooling and maximum pooling, SA avg and SA max is the average and maximum pooling spatial feature descriptor, and σ(·) is the sigmoid function.
[0027] As a further improvement of the present invention, the WIoUv3 position loss function in step (3) has a WIoUv1 calculation formula as shown in equations (9)-(12):
[0028] L WIoUv1 =R WIoU L IoU (9)
[0029]
[0030] S u =wh+w gt h gt -W i H i (12)
[0031] Among them, (x, y, w, h) are the coordinates, width and height of the center point of the prediction box, (x gt ,ygt ,w gt ,h gt ) are the coordinates, width and height of the target frame center point, (W g ,H g ) is the width and length of the union of the prediction box and the target box, (W i ,H i ) is the width and length of the intersection of the prediction box and the target box;
[0032] A non-monotonic focusing coefficient γ is constructed using the outlier degree β of the anchor box to reduce the gradient penalty for low-quality samples. The WIoU v3 loss function is proposed by applying the non-monotonic focusing coefficient γ to WIoUv1. The calculation formulas are shown in Equations (13)-(15):
[0033]
[0034] L WIoUv3 =γL WIoUv1 (15)
[0035] Among them, α and δ are hyperparameters, β is the outlier degree, and γ is the non-monotonic focusing coefficient. The non-monotonic focusing coefficient γ reduces the competitiveness of high-quality samples and also reduces the harmful gradients generated by low-quality samples.
[0036] The present invention's hybrid soybean male sterile line phenotypic consistency detection method incorporates the LSKNet attention mechanism into the YOLOv5s model. This allows the model to effectively focus on small targets, such as flowers, in the mid-target detection layer, enhancing the model's ability to focus on target areas within each detection layer. The addition of a small target detection branch also allows the model to focus more on small targets. The resulting red highlight areas precisely encapsulate the hypocotyl, hairs, and flower targets, achieving optimal spatial overlap between the focus area and the actual target position. Leaf shape, hypocotyl color, hair color, and flower color can be accurately detected by the XLW-YOLO model with high confidence and precision. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 These are the original soybean images: (a) hypocotyl image, (b) leaf image, (c) hair image, and (d) flower image.
[0038] Figure 2 The leaf shapes of soybeans are: (a) round leaf, (b) ovate leaf, (c) elliptical leaf, (d) lanceolate leaf;
[0039] Figure 3 is the labeling strategy: (a) hypocotyl, (b) leaf, (c) pubescence, (d) flower;
[0040] Figure 4 It is the XLW-YOLO model network structure;
[0041] Figure 5 is the bounding box ratio distribution map;
[0042] Figure 6 It is the LSK Block network structure;
[0043] Figure 7 It is the LSK model network structure;
[0044] Figure 8 is the bounding box regression map;
[0045] Figure 9 is the flower detection heat map of different improved methods;
[0046] Figure 10 XLW-YOLO multi-target detection layer heatmaps of different sizes;
[0047] Figure 11 It is the result of hybrid soybean phenotypic consistency test. DETAILED DESCRIPTION
[0048] The following is a further description of the method for detecting phenotypic consistency of hybrid soybean male sterile lines according to the present invention with reference to the accompanying drawings:
[0049] 1. Original image acquisition
[0050] Visible light images of soybean plants growing naturally in the field were collected at different growth stages. To obtain rich image data and meet the requirements for phenotypic consistency testing of different soybean varieties, this example collected data from 250 soybean varieties. For each soybean variety, four hypocotyl images were collected during the seedling stage, four leaf images during the vegetative growth stage, and four pubescent images and four flower images during the reproductive growth stage, resulting in a total of 16 images. For each of the 250 varieties, 1,000 hypocotyl images, 1,000 leaf images, 1,000 pubescent images, and 1,000 flower images were obtained. The soybean plant image data was acquired using the Azure Kinect DK image acquisition sensor with a resolution of 1280×720 pixels.
[0051] To achieve information-based detection of hybrid soybean field phenotypic consistency, naturally grown soybean plants in the field were used as research objects. Azure Kinect DK image acquisition sensors were used to acquire visible light image data of soybean plants at a resolution of 1280 × 720 pixels. 1000 hypocotyl images were collected during the soybean seedling stage, 1000 leaf images during the vegetative growth stage, 1000 pubescent images during the reproductive growth stage, and 1000 flower images. Figure 1 shown.
[0052] According to the requirements of phenotypic consistency identification, the hypocotyl color is divided into purple and green; the leaf shape is divided into round and pointed; the hair color is divided into brown and white; the flower color is divided into purple and white. Round, oval and elliptical leaves are defined as round, and lanceolate leaves are defined as pointed. Figure 2 shown.
[0053] 2. Dataset Annotation
[0054] Use LabelImg software to manually label the data samples obtained in step 1 according to the labeling strategy to construct a data set, such as Figure 3 As shown in the figure, the labeling strategy is to mark purple and green hypocotyls, round and pointed leaves, brown and white hairs, purple and white flowers respectively. Due to the different curvature of the leaves, errors may occur when observing the leaf shape from different angles, so leaves parallel to the ground are selected for labeling. The image data set obtained during the vegetative growth period cannot distinguish the hair color, and only the leaf shape index is marked in the data set labeling process. Due to the limited visual angle of the image data and insufficient feature richness, the labeled data set is expanded by methods such as image flipping, translation, brightness adjustment and noise addition to expand to 8000 image data. The bounding box coordinates of the XML labeling file are updated synchronously during the enhancement process. The four data sets (hypocotyl, leaf, hairs and flowers) are divided into sub-training sets and sub-test sets according to the ratio of 8:2. The four sub-training sets are then merged into the total training set, and the four sub-test sets are merged into the total test set.
[0055] 3. Construction of phenotypic consistency detection model for hybrid soybean male sterile lines
[0056] YOLO (You Only Look Once) is a typical single-stage target detection algorithm. Due to its fast detection speed, it has been widely used in real-time detection and is used as the basic input model. The YOLOv5s model consists of four parts: Input layer, Backbone layer, Neck layer and Head layer. The image is preprocessed through the Input layer, the Backbone layer extracts target features of different dimensions from the processed data, the Neck layer mainly performs feature fusion on the multi-scale feature map obtained by the Backbone layer, and the Head layer combines the features of the Neck layer to generate a bounding box and classify the target. The total training set obtained in step (2) is input into the hybrid soybean phenotypic consistency detection model XLW-YOLO, as shown in Figure 2. Figure 4However, for small-sized target detection tasks such as hypocotyls, hairs, and flowers, the deeper feature maps in the YOLOv5s model may lose feature information of small targets due to multiple downsampling. Therefore, this method enhances positioning capability and detection accuracy while ensuring the network can detect targets, ensuring its future application in hybrid soybean shearing robots. The following improvements were made: a small target detection branch was added to the YOLOv5s network; an LSKNet attention mechanism module was added to the layer before the SPPF backbone network; and the CIou (Complete Intersection over Union) loss function was replaced with the WIoU v3 loss function.
[0057] 3.1 Small Target Detection Branch
[0058] By introducing more convolutional layers, pooling layers, and feature fusion layers into the YOLOv5s network, the small target detection branch has fewer downsampling times, which makes the resolution of small targets on the feature map higher, thereby retaining more detail information.
[0059] The bounding box ratio distribution diagram shows the aspect ratio of the target in the dataset relative to the entire image. By normalizing the size of the bounding box, it can be seen that most of the targets in the image are small targets, such as Figure 5 As shown in the figure. The YOLOv5s model has a relatively large downsampling rate, making it difficult for deep feature maps to learn the features of small objects. Therefore, the YOLOv5s model is prone to errors and omissions when detecting small objects in complex backgrounds. To address this issue, a small object detection branch is added to the YOLOv5s model. First, the 80×80 resolution 512-channel feature map in the Neck network is converted into a 256-channel feature map through the C3 module. The optimized 256-channel feature map is compressed by 1×1 convolution to generate a 128-channel feature map. The compressed feature map is then upsampled by a factor of 2 to generate a 160×160 resolution 128-channel feature map. The upsampled feature map is then concatenated with the 160×160 resolution, 128-channel feature map output by the second layer of the Backbone network according to the channel dimension to generate a 256-channel fused feature map. The fused feature map is processed by the C3 module and input into the Head network for prediction.
[0060] 3.2LSKNet Attention Mechanism
[0061] In the task of detecting hybrid soybean phenotypic consistency, the size and shape of the target are uncertain. Traditional convolutional neural networks use a fixed convolution kernel size. Although this simplifies the network structure, it lacks the ability to adapt to input features of varying scales. The size of the convolution kernel determines the size of its receptive field, and a fixed receptive field limits the network's comprehensive understanding and analysis of the input features. By adding the LSKNet attention mechanism module to the layer before the SPPF in the Backbone network, LSKNet uses a spatial selection mechanism to dynamically determine the size of the convolution kernel based on the input. This allows the model to adjust the receptive field of each target as needed, thereby better capturing the target's characteristics. Furthermore, LSKNet excels in processing targets with spatial variations. The model can better adapt to changes in different scenarios and targets, improving performance in practical applications.
[0062] Since different types of detection targets have different requirements for background information, the model needs to adapt to and select background areas of different sizes. LSKNet can dynamically adjust the receptive field and can effectively process the relevant background information required by different targets. LSK Block is the basic module of LSKNet, which mainly includes two sub-blocks: LK Selection and FFN. Figure 6 As shown in Figure 2. LKSelection can dynamically adjust the receptive field of the network, while FFN enhances the ability of channel mixing and feature refinement. The LSK model consists of Large kernel convolutions and Spatial kernel selection, which is integrated in LKSelection, as shown in Figure 2. Figure 7 Large kernel convolutions enable the model to extract features with different contextual information from different input fragments. Spatial kernel selection enhances the network's ability to capture spatial relationships.
[0063] The input feature X is decomposed into a series of depth convolutions through large kernel convolutions. In order to adjust the number of channels to match the same output dimension, F is used. 1×1 The convolutional layer obtains different kernel features The calculation is shown in formulas (1)-(3).
[0064] U0=X,U i+1 =F i dw (U i ) (1)
[0065]
[0066] Among them, F idw (·) is a kernel with k i and extension d i Depthwise convolution.
[0067] The spatial feature descriptor SA is obtained by performing channel connection and pooling operations on different kernel features, using F 2→N Convolutional layer and Sigmoid activation function to obtain spatial selection mask The spatial selection mask is used to weight different kernel features, and the attention feature S is obtained through the convolution layer F. The input feature X is element-wise multiplied with the attention feature S to obtain the final output Y. The calculation is shown in Equations (4)-(8).
[0068]
[0069]
[0070] Y=X·S (8)
[0071] Among them, P avg (·) and P max (·) is average pooling and maximum pooling, SA avg and SA max is the average and maximum pooling spatial feature descriptor, and σ(·) is the sigmoid function.
[0072] 3.3Wise-IoU Loss Function
[0073] A loss function is an evaluation metric used to assess the degree of discrepancy between a model's predicted value and the true value. The CIoU loss function is used in the YOLOv5s model. While the CIoU loss function performs well in object detection, it uses the same loss calculation method for both high-quality and low-quality samples. However, geometric metrics such as distance and aspect ratio penalize low-quality samples, resulting in reduced generalization performance. Since the experimental data was acquired in the field, images from brightly lit, backlit, overcast, and dusky environments contain a higher number of low-quality samples. The WIoU loss function reduces the penalty for low-quality samples, mitigating their negative impact on model performance and enabling the model to focus more on learning high-quality samples. Furthermore, to meet the future application requirements of hybrid soybean shearing robots in the field, improved model localization performance is also required. The WIoU loss function dynamically adjusts the gradient gain based on the quality of the anchor boxes, assigning smaller gradient gains to high-quality anchor boxes and larger gradient gains to average-quality anchor boxes. This refined gradient allocation strategy helps the model focus on those bounding boxes that require optimization, thereby improving model localization performance. The calculation formulas of WIoUv1 are shown in equations (9)-(12).
[0074] LWIoUv1 =R WIoU L IoU (9)
[0075]
[0076] S u =wh+w gt h gt -W i H i (12)
[0077] Among them, (x, y, w, h) are the coordinates, width and height of the center point of the prediction box, (x gt ,y gt ,w gt ,h gt ) are the coordinates, width and height of the target frame center point, (W g ,H g ) is the width and length of the union of the prediction box and the target box, (W i ,H i ) is the width and length of the intersection of the prediction box and the target box, such as Figure 8 shown.
[0078] A non-monotonic focusing coefficient γ is constructed using the outlier degree β of the anchor box to reduce the gradient penalty for low-quality samples. The WIoU v3 loss function is proposed by applying the non-monotonic focusing coefficient γ to WIoUv1. The calculation formulas are shown in Equations (13)-(15).
[0079]
[0080] L WIoUv3 =γL WIoUv1 (15)
[0081] Here, α and δ are hyperparameters, β is the outlier degree, and γ is the nonmonotonic focusing coefficient. The nonmonotonic focusing coefficient γ reduces the competitiveness of high-quality samples while also reducing the harmful gradients generated by low-quality samples. This allows WIoU v3 to dynamically and nonmonotonically focus on common samples, thereby improving the model's generalization ability and overall performance. To this end, the WIoU v3 loss function is used instead of the CIoU loss function.
[0082] 4. Test platform construction and training parameters
[0083] To meet the requirements of model training, an image testing platform was built with the following configuration: an NVIDIA GeForce RTX 4090 graphics card, an Intel Core i9-13900K processor, and 128GB of RAM. The model used PyTorch version 1.13.1, Python version 3.8.18, Cuda version 11.7, and Cudnn version 8500. The training parameters were epochs, batch size, workers, and image size of 300, 64, 16, and 640×640, respectively. The hyperparameters learning rate, momentum, and weight decay were set to 0.01, 0.973, and 0.0005, respectively. The hyperparameters for WIoU v3 were α = 1.9 and δ = 3.
[0084] By incorporating the LSKNet attention mechanism module, a small object detection branch, and the WIoU v3 position loss function into the YOLOv5s network, we developed the XLW-YOLO model for detecting phenotypic consistency in hybrid soybeans. This model can detect hypocotyl color during the seedling stage, leaf shape during the vegetative stage, and leaf shape, fuzz color, and flower color during the reproductive stage. This research aims to utilize information technology to identify phenotypic consistency in hybrid soybeans in the field, providing technical support for soybean breeders in hybrid selection.
[0085] 5. Evaluation indicators
[0086] To evaluate the performance of the hybrid soybean phenotypic consistency detection model, we used precision (P), recall (R), F1 value, mean average precision (mAP), detection speed, and model size as evaluation indicators. The calculation formulas for precision (P), recall (R), and F1 value are as follows:
[0087]
[0088] Among them, T P is the number of samples correctly predicted as positive, F P is the number of negative samples predicted as positive samples, F N is the number of positive samples predicted as negative samples.
[0089] AP (Average Precision) reflects the accuracy of each category prediction, and its value is the area enclosed by the PR curve and the horizontal and vertical axes. The calculation formula is as follows:
[0090]
[0091] Where r is the integral variable, which is the integral of the product of recall and precision.
[0092] mAP is the average AP of all detected target categories, reflecting the overall detection performance of the model. Among them, mAP 0.5 represents the mean average precision when the IoU threshold is 0.5. The calculation formula is as follows:
[0093]
[0094] Here, S is the number of all categories. There are eight categories in the study: green hypocotyl, purple hypocotyl, brown hair, white hair, purple flower, white flower, pointed leaf and round leaf, so S = 8.
[0095] The detection speed indicator is mainly used to evaluate the speed at which the model processes images. FPS represents the number of images the model can process per second. A larger FPS value indicates a faster image processing speed and better model performance. The calculation formula is as follows:
[0096]
[0097] Among them, Frames represents the total number of frames processed in one cycle, and Time is the time of one cycle.
[0098] 6. Phenotypic consistency calculation method
[0099] To evaluate the detection effect of the model, it is necessary to use the model to identify the phenotypic consistency of the acquired images. According to the "Procedure for Identification of Hybrid Soybean Cytoplasmic-Nuclear Interaction Male Sterility Lines", the phenotypic consistency calculation formula is as follows:
[0100]
[0101] Wherein, C is the phenotypic consistency, the unit is percentage (%); Z T is the total number of observed samples, in units of plants; Z is the number of mixed plants, in units of plants. The calculation results are accurate to one decimal place.
[0102] Below in conjunction with test, the effect of the present invention is further described:
[0103] 1. YOLO model comparison test
[0104] Comparative tests were conducted using the YOLO family of object detection algorithms. The results are shown in Table 1. The YOLOv5s model outperformed the other models in terms of P-value, F1-score, and mAP, reaching 89.4%, 86.2%, and 89.4%, respectively. Through comprehensive comparative analysis, the YOLOv5s model was selected as the base model, and improvements were made to enhance its detection performance. The improved XLW-YOLO model achieved significantly higher P-value, R-value, F1-score, and mAP, compared to the original YOLOv5s model, reaching 94.0%, 90.6%, 92.3%, and 94.8%, respectively. The model achieved a detection speed of 135 FPS, which is sufficient for real-time field detection tasks.
[0105] Table 1 Model detection results
[0106] Precision(%) Recall (%) F1score (%) mAP50 (%) YOLOv4 46.4 89.7 61.2 67.0 YOLOv5s 89.4 83.3 86.2 89.4 YOLOv7 69.7 66.3 68.0 70.3 YOLOv8n 74.5 68.4 71.3 73.8 YOLOv10s 87.1 79.6 83.2 87.2 XLW-YOLO 94.0 90.6 92.3 94.8
[0107] 2. Model Improvement Performance Verification
[0108] Heatmaps are an effective tool for visualizing the performance, attention regions, and attention distribution of deep learning network models. Figure 9 The heatmaps for flower detection using different improved methods are shown. The heatmaps show that the LSKNet attention mechanism allows the model to focus more on the flower region compared to the YOLOv5s model. YOLOv5s is unable to focus on the flower region at the 40×40 detection layer, but the addition of the LSKNet attention mechanism effectively changes this situation. It can also be seen that adding a small target detection branch allows the model to focus more on small targets, and allows the highlighted region to closely surround the flower target region. Figure 9 It can be seen that the LSKNet attention mechanism plays an important role in each detection layer, and the small target detection branch can make the attention mechanism pay more attention to small target detection.
[0109] Figure 10 This is a heat map of the XLW-YOLO model's multi-target detection layers of different sizes. The heat map shows that the red highlighted area in the 160×160 detection layer closely surrounds the small targets of the hypocotyl, hairs, and flowers, and the target regions of interest highly overlap with the actual locations. The red highlighted area in the 80×80 detection layer also includes the hypocotyl, hairs, and small flower targets, but its region of interest is larger than the actual size of the targets. The 40×40 target detection layer pays more attention to soybean leaves during the vegetative growth phase and less attention to soybean leaves during the peak flowering phase. The 20×20 detection layer fails to detect the hypocotyl and flower small targets, but performs well in detecting leaves during the peak flowering phase. Experiments demonstrate that the model's output layers of different sizes better focus on targets of different sizes, and adding a small target detection branch improves the model's detection performance for small hypocotyl, hairs, and flowers.
[0110] 3. Model Performance Verification
[0111] In order to verify the performance of the XLW-YOLO model, different index test set data were tested. The test results are as follows Figure 11 As shown in the figure, the XLW-YOLO model accurately detects soybean leaf shape, hypocotyl color, hair color, and flower color with high confidence. Compared with the XLW-YOLO model, the YOLOv5s model has some missed detections and lower confidence.
[0112] 4. Hybrid soybean phenotypic consistency detection
[0113] XLW-YOLO and YOLOv5s models were used to detect 100 purple hypocotyl index image data and compared with manually recorded data. The test results are shown in Table 2. The phenotypic consistency was calculated according to formula (22). The phenotypic consistency of manual counting, YOLOv5s, and XLW-YOLO model detection was 98.9%. The number of samples and the number of hybrid plants detected by YOLOv5s and XLW-YOLO models were lower than those of manual counting, indicating that there were omissions. However, this omission is acceptable because leaf shape, hair color, and flower color will be tested during the subsequent growth period. Therefore, the XLW-YOLO model can meet the needs of field hybrid soybean phenotypic consistency identification.
[0114] Table 2 Phenotypic consistency test results
[0115] Detection method Number of samples (plants) Number of hybrid plants (plants) Phenotypic consistency (%) Manual counting 567 7 98.9% YOLOv5s 554 6 98.9% XLW-YOLO 560 6 98.9%
Claims
1. The method for detecting phenotypic consistency of hybrid soybean male sterile lines is achieved through the following steps: (1) Original image acquisition For soybean plants growing naturally in the field, visible light images were collected at different growth stages, including: 1,000 hypocotyl images were collected during the soybean seedling stage, 1,000 leaf images were collected during the vegetative growth stage, 1,000 pubescent images were collected during the reproductive growth stage, and 1,000 flower images were collected. The Azure Kinect DK image acquisition sensor was used to acquire soybean plant image data with a resolution of 1280 × 720 pixels. (2) Dataset annotation The data samples obtained in step (1) were manually annotated using LabelImg software according to the annotation strategy to construct a dataset, and the annotated dataset was expanded using the Python OpenCV library by image flipping, translation, brightness adjustment, and noise addition methods to expand it to 8,000 image data. During the enhancement process, the bounding box coordinates of the XML annotation file were synchronously updated; the four datasets of hypocotyl, leaf, hair, and flower were divided into sub-training sets and sub-test sets at an 8:2 ratio, and then the four sub-training sets were merged into the total training set, and the four sub-test sets were merged into the total test set; (3) Construction of a phenotypic consistency detection model for hybrid soybean male sterile lines The total training set obtained in step (2) was used as the model input data, and the YOLOv5s model was used as the basic network structure. The LSKNet attention mechanism module was added to the layer before the SPPF of the Backbone network, and a small target detection branch XNet for the 160×160 resolution feature map was added to the Neck network. The CIoU loss function used in the YOLOv5s model was changed to the WIoU v3 loss function to establish the hybrid soybean male sterile line phenotypic consistency detection model XLW-YOLO. (4) Detection of phenotypic consistency of hybrid soybean male sterile lines By collecting images of soybean plants growing naturally in the field at different growth stages and inputting them into the detection model, the phenotypic consistency of hybrid soybean male sterile lines can be detected.
2. The method for detecting phenotypic consistency of hybrid soybean male sterile lines according to claim 1, characterized in that In step (3), a small target detection branch XNet for 160×160 resolution feature maps is added to the Neck network of the YOLOv5s model. The specific process is as follows: (1) For the 512-channel feature map with 80×80 resolution in the Neck network, a 256-channel feature map is generated through the C3 module; (2) Perform 1×1 convolution compression on the optimized 256-channel feature map to generate a 128-channel feature map; (3) Upsample the compressed feature map by a factor of 2 to generate a 128-channel feature map with a resolution of 160×160; (4) The upsampled feature map is concatenated with the 160×160 resolution, 128-channel feature map output by the second layer of the Backbone network according to the channel dimension to generate a 256-channel fused feature map; (5) The fused feature map is processed by the C3 module and then input into the Head network for prediction.
3. The method for detecting phenotypic consistency of hybrid soybean male sterile lines according to claim 1, characterized in that In step (3), the LSKNet attention mechanism is added to decompose the convolution process of the input feature X into a series of deep convolutions through large kernel convolutions. In order to adjust the number of channels to match the same output dimension, F is used. 1×1 The convolutional layer obtains different kernel features The calculation is shown in formulas (1)-(3). U0=X, U i+1 =F i dw (U i ) (1) Among them, F i dw (·) is a kernel with k i and extension d i Depthwise convolution; The spatial feature descriptor SA is obtained by performing channel connection and pooling operations on different kernel features, using F 2→N Convolutional layer and Sigmoid activation function to obtain spatial selection mask The spatial selection mask is used to weight different kernel features, and the attention feature S is obtained through the convolution layer F. The input feature X is element-wise multiplied with the attention feature S to obtain the final output Y. The calculation is shown in formulas (4)-(8): Y=X·S (8) Among them, P avg (·) and P max (·) is average pooling and maximum pooling, SA avg and SA max is the average and maximum pooling spatial feature descriptor, and σ(·) is the sigmoid function.
4. The method for detecting phenotypic consistency of hybrid soybean male sterile lines according to claim 1, characterized in that The WIoUv3 position loss function in step (3) is calculated using the WIoUv1 formula as shown in equations (9)-(12): L WIoUv1 =R WIoU L IoU (9) S u =wh+w gt h gt -W i H i (12) Among them, (x, y, w, h) are the coordinates, width and height of the center point of the prediction box, (x gt ,y gt ,w gt ,h gt ) are the coordinates, width and height of the target frame center point, (W g ,H g ) is the width and length of the union of the prediction box and the target box, (W i ,H i ) is the width and length of the intersection of the prediction box and the target box; A non-monotonic focusing coefficient γ is constructed using the outlier degree β of the anchor box to reduce the gradient penalty for low-quality samples. The WIoU v3 loss function is proposed by applying the non-monotonic focusing coefficient γ to WIoUv1. The calculation formulas are shown in Equations (13)-(15): L WIoUv3 =γL WIoUv1 (15) Among them, α and δ are hyperparameters, β is the outlier degree, and γ is the non-monotonic focusing coefficient. The non-monotonic focusing coefficient γ reduces the competitiveness of high-quality samples and also reduces the harmful gradients generated by low-quality samples.