Improved soybean seed testing system and method based on Transformer-faster-RCNN
By improving the soybean seed testing system based on Transformer-faster-RCNN, combined with the pod recognition network and image restoration network, the problems of large errors and long time consumption in counting the number of soybean grains and stems were solved, achieving more efficient and accurate phenotypic recognition and target detection.
Patent Information
- Application Number
- CN202510950125.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Existing technologies for counting soybean grains and stems have problems such as large errors, long time consumption, insufficient model generalization ability, low multi-scale target detection accuracy, and high data labeling costs, making it impossible to quickly and accurately count the number of soybeans.
A soybean seed testing system based on the improved Transformer-faster-RCNN is adopted, including a pod recognition network and a pod image restoration network. The Swin-Transform structure is used to extract features, enhance data diversity and refine the training strategy. The pod image restoration network solves the occlusion problem between pods and improves the efficiency and accuracy of phenotypic recognition.
It significantly improves the efficiency and accuracy of phenotypic recognition in complex scenarios, reduces missed detections or false detections, enhances the robustness of the model and the accuracy of target detection, and can more flexibly adapt to different scenarios and changes in soybean pod shapes.
Smart Images

Figure CN120431413B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of soybean seed testing, and in particular relates to a soybean seed testing system and method based on an improved Transformer-faster-RCNN. Background Art
[0002] Currently, there are two main types of high-throughput soybean bean and stem counting methods:
[0003] The first method involves threshing the soybean pods. The number of stem nodes is manually counted, and the beans are then spread out on a flat surface. Because the beans are all round, they avoid blocking each other and are distributed roughly evenly. A visible light camera is then used to capture the image, and digital image processing techniques are used to count the beans. This method takes a long time to thresh, resulting in smaller beans, and some beans may be damaged or lost, which can affect the accuracy of bean counting.
[0004] The second method involves removing pods from mature soybean plants, manually counting the number of stem nodes, and then neatly placing the pods on a flat plate. These pods are then imaged using a visible light camera and digital image processing techniques to identify pods that are not obstructed by each other. Pods containing different types of beans are then manually classified, and the total number of beans is finally calculated. While this method avoids missed detections, it requires extensive manual labor, is time-consuming, and is incompatible with the extraction of obstructed pods.
[0005] As can be seen above, none of the aforementioned methods can quickly and accurately count soybean pods and stems. Direct counting results in large errors due to the small size of the pods, while indirect counting relies on manual placement and classification, which is time-consuming. Existing methods also suffer from limitations such as insufficient model generalization, low multi-scale object detection accuracy, and high data annotation costs. Summary of the Invention
[0006] In light of this, the present invention aims to provide a soybean seed identification system and method based on an improved Transformer-Faster-RCNN. This system uses a pod recognition network to accurately detect and count the number of pods and stems in their naturally mature state, while a pod image restoration network addresses the issue of pod occlusion. By optimizing the network structure, enhancing data diversity, and refining training strategies, the efficiency and accuracy of phenotypic recognition in complex scenarios are significantly improved.
[0007] To achieve the above object, the technical solution created by the present invention is implemented as follows:
[0008] A soybean seed testing system based on an improved Transformer-faster-RCNN includes a pod recognition network and a pod image restoration network. The pod recognition network includes a backbone module with a Swin-Transform structure, a neck module with a pyramid structure, and a detection module. The pod image is input into the backbone module for feature extraction to obtain extracted features at multiple scales. The residual features at multiple scales are input into the neck module for feature transformation using a multi-scale pyramid to obtain transformed features at multiple scales. The transformed features at multiple scales are input into the detection module for target recognition to obtain a stem image in the pod image and an independent image of each pod in the pod image. The pod image restoration network restores the independent pod image. The pod image restoration network includes a discriminator and a generator. The generator is used to restore the input independent pod image, and the discriminator is used to determine the authenticity of the restored independent pod image.
[0009] Furthermore, the backbone module includes a feature splitting block, a first extraction block, a second extraction block, a third extraction block and a fourth extraction block; wherein the feature splitting block performs image splitting on the input image to obtain multiple non-overlapping image blocks of equal size; the multiple image blocks are sequentially subjected to feature extraction by the first extraction block, the second extraction block, the third extraction block and the fourth extraction block; the output features of the second extraction block, the third extraction block and the fourth extraction block are input into the neck module.
[0010] Furthermore, in the first extraction block, the input multiple image blocks are processed with multiple groups of linear embedding windows and SwinTransformer blocks to obtain output features; in the second extraction block, the third extraction block and the fourth extraction block, the input features are processed with multiple groups of merged adjacent windows and SwinTransformer blocks to obtain output features.
[0011] Furthermore, in the neck module: after the output features of the fourth extraction block are upsampled, they are channel-joined with the output features of the third extraction block to obtain a spliced feature; after the spliced feature is upsampled, it is channel-joined with the output features of the second extraction block to obtain a first transformation feature; after the first transformation feature is convolved to compress the feature size, it is channel-joined with the spliced feature to obtain a second transformation feature; after the second transformation feature is channel-joined with the output features of the fourth extraction block, a third transformation feature is obtained; and the three transformation features are input into the detection module for target recognition.
[0012] Furthermore, in the detection module, the transformation features of multiple scales are integrated, and the integrated features are subjected to RPN operation to obtain candidate pod targets and stem targets; the obtained pod targets and stem targets are combined with the integrated features to perform ROI alignment, and the aligned content is fully connected to obtain the stem image, pod independent image and their respective categories.
[0013] Furthermore, the generator adopts a U-Net structure. In the discriminator, the generator's restored pod image and the corresponding real pod image are subjected to multiple continuous convolution operations to complete feature extraction. The extracted features are subjected to subsequent image segmentation operations in PatchGAN to complete feature comparison, and an authenticity score is given to evaluate the authenticity of the pod image. The authenticity score is used to guide the generator to complete the restoration of the pod image.
[0014] A soybean seed testing method based on Transformer-faster-RCNN improvement, including:
[0015] S1: Obtain a pod image dataset and preprocess it to obtain a training set that includes pod occlusion conditions;
[0016] S2: Using the training set obtained in step S1, the soybean seed detection system improved based on Transformer-faster-RCNN provided by the present invention is trained to obtain a soybean seed detection model;
[0017] S3: According to the training result of step S2, adjusting the hyperparameters when training the soybean seed selection model until an optimal soybean seed selection model is obtained;
[0018] S4: Input the pod image to be tested into the optimal soybean seed testing model obtained in step S4 to obtain a repaired independent pod image of each pod in the pod image, as well as a stem image; calculate and count the traits of the stem image and all independent pod images, thus completing the soybean seed testing.
[0019] Furthermore, step S1 includes: collecting mature soybean plants in different scenes, scattering the pods on the soybean plants arbitrarily above the background board, placing the stems of the soybean plants below the background board, and shooting with a shooting device to obtain a pod image; removing redundant parts in the pod image, and cropping the pod image into an image of uniform pixel size, annotating the pods and stems in the cropped image to obtain an annotation file; performing image enhancement on the cropped image and the annotation file to obtain a training set, which includes the cropped pod image and the corresponding annotation file.
[0020] Furthermore, in step S2, the soybean recognition network is trained using the recognition loss function, and the pod image restoration network is trained using the restoration objective function; wherein the recognition loss function is:
[0021] ;
[0022] Among them, L({p i},{t i}) represents the recognition loss function, p i Indicates the probability that a target exists in the i-th bounding box during target recognition, Indicates the corresponding true label, t i represents the compensation between the i-th predicted bounding box and the corresponding true bounding box coordinates, Indicates the coordinates of the real bounding box where the target exists, N cls Indicates the batch size during training, N reg Represents the number of recognized targets, and λ represents the adjustment weight of the recognition loss function; represents the cross entropy loss function, namely:
[0023] ;
[0024] Represents the Focal_loss loss function, which is composed of smooth L1 The losses are accumulated, that is:
[0025] ;
[0026] ;
[0027] The repair objective function is:
[0028] ;
[0029] in, represents the repair objective function, D represents the discriminator, G represents the generator, LOSS1 represents the L1 loss function, Indicates the repair weight; LOSS cGAN (G,D) represents the repair loss function, which is:
[0030] ;
[0031] Among them, x represents the real image, y represents the corresponding constraint, z represents the noise, E x,y [logD(x,y)] represents the discriminator’s expectation of the real image, E x,z [log(1-D(x,G(x,y)))] represents the discriminator’s expectation of the image restored by the generator.
[0032] Furthermore, in step S4, the pod traits include pod length, pod width, pod circumference and pod projected area; the pod trait extraction process includes: grayscale the repaired pod independent image, extracting the R channel image with the maximum contrast between the pod grains and the background; binarizing the R channel image using a threshold segmentation algorithm to obtain a pod binary image; extracting the pod outer contour based on the pod binary image, including the coordinates of each point on the contour; and calculating the pod length, width, circumference and pod projected area based on the obtained contour coordinates.
[0033] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0034] (1) In the soybean seed testing system improved based on Transformer-faster-Rcnn created by the present invention, a backbone module for extracting features using the Swin-Transform structure is used. The Transformer can establish connections between all patches through the self-attention mechanism, thereby capturing the global information in the image; in addition, the Transformer's self-attention mechanism can dynamically adjust the feature extraction process, focusing on different image areas according to the different input content. This enables the model to adapt more flexibly to different scenes and changes in the appearance of different soybean pod types; the Transformer can effectively model long-distance dependencies, which is very useful for tasks that require understanding complex scenes. In target recognition tasks, the model needs to be able to capture the complex relationship between the target and its surrounding environment, and accurately distinguish the target information and background information of soybeans in order to more accurately identify the target;
[0035] (2) In the soybean seed testing system based on the improved Transformer-faster-RCNN described in the present invention, bidirectional fusion is used to make each layer of the feature pyramid have strong semantics and fine-grained details at the same time, such as significantly improving the target extraction effect in dense scenes; this structure can more accurately extract the feature information of targets of different pod scales, reducing missed detections or false detections caused by changes in pod scale. It improves the positioning accuracy of the feature map for soybean targets, provides better positioning points in the bounding box regression task, and generates candidate boxes in the target area;
[0036] (3) In the soybean seed classification method based on the improved Transformer-faster-RCNN described in the present invention, Focal_loss is used to replace BCE_loss in the category classification process. Focal_loss is improved based on the standard BCE loss function. Its core idea is to make the model pay more attention to difficult-to-classify samples by dynamically adjusting sample weights. Focal_loss introduces a modulation factor to reduce the loss contribution of easy-to-classify samples, making the model pay more attention to difficult-to-classify samples. This design helps to solve the problem of category imbalance and improve the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings, which constitute part of the present invention, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0038] Figure 1 This is a schematic diagram of the soybean seed testing system improved based on Transformer-faster-RCNN according to an embodiment of the present invention;
[0039] Figure 2 This is a schematic diagram of the structure of the peapod identification network described in an embodiment of the present invention;
[0040] Figure 3 A schematic structural diagram of the backbone module according to an embodiment of the present invention;
[0041] Figure 4 A schematic diagram of a peapod image restoration network according to an embodiment of the present invention;
[0042] Figure 5 A schematic diagram of a generator according to an embodiment of the present invention;
[0043] Figure 6 A schematic diagram of a discriminator according to an embodiment of the present invention;
[0044] Figure 7 A schematic diagram of the process of the soybean seed testing method improved based on Transformer-faster-RCNN according to an embodiment of the present invention;
[0045] Figure 8 Creating the pod image described in the embodiment of the present invention;
[0046] Figure 9 This is a diagram of the pod target recognition result described in the embodiment of the present invention;
[0047] Figure 10 This is a partial enlarged view of the pod target recognition described in the embodiment of the present invention;
[0048] Figure 11 A partially enlarged view of the stem target identification according to an embodiment of the present invention;
[0049] Figure 12 A comparison chart of the results of the pod image restoration described in the embodiment of the present invention;
[0050] Figure 13 This is a diagram of the process of extracting pod traits according to an embodiment of the present invention;
[0051] Figure 14 The data fitting effect diagram of the bad pod number statistics described in the embodiment of the present invention is created
[0052] Figure 15 This is a data fitting effect diagram of the statistics of the number of pods of a type described in an embodiment of the present invention;
[0053] Figure 16 This is a data fitting effect diagram of the second type of pod number statistics described in the embodiment of the present invention;
[0054] Figure 17 This is a data fitting effect diagram of the statistics of the number of three types of pods described in the embodiment of the present invention;
[0055] Figure 18 This is a data fitting effect diagram of the statistics of the number of four types of pods described in the embodiment of the present invention;
[0056] Figure 19 This is a data fitting effect diagram of the statistics of the number of all pods described in the embodiment of the present invention;
[0057] Figure 20 This is a data fitting effect diagram for the statistics of the number of all beans described in the embodiment of the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation of the present invention.
[0059] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0060] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second" and the like are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, features defined as "first", "second" and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0061] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art can understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0062] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments.
[0063] like Figures 1 to 6 As shown, the soybean seed testing system based on the improved Transformer-faster-RCNN described in the embodiment of the present invention includes a pod recognition network and a pod image restoration network; wherein, the pod recognition network includes a backbone module of a Swin-Transform structure, a neck module of a pyramid structure, and a detection module, the backbone module is used to extract soybean image features, the neck module is used to fuse and transform features extracted by the feature extraction network at different depths, and the detection module is used to classify the targets detected by the system and optimize the position of the detected target bounding box. Specifically: the pod image is input into the backbone module for feature extraction to obtain extracted features of multiple scales; the residual features of multiple scales are input into the neck module for feature transformation of a multi-scale pyramid to obtain transformed features of multiple scales; the transformed features of multiple scales are input into the detection module for target recognition to obtain a stem image in the pod image and an independent image of each pod in the pod image.
[0064] In some embodiments, the backbone module includes a feature segmentation block, a first extraction block, a second extraction block, a third extraction block, and a fourth extraction block. The feature segmentation block performs image segmentation on the input image to obtain multiple non-overlapping image blocks of equal size; the multiple image blocks are sequentially subjected to feature extraction by the first extraction block, the second extraction block, the third extraction block, and the fourth extraction block; and the output features of the second, third, and fourth extraction blocks are input into the neck module.
[0065] In some embodiments, in the first extraction block, the input multiple image blocks are processed with multiple groups of linear embedding windows and Swin Transformer blocks to obtain output features; in the second extraction block, the third extraction block and the fourth extraction block, the input features are processed with multiple groups of merged adjacent windows and Swin Transformer blocks to obtain output features.
[0066] In an embodiment of the present invention, in the first extraction block, the input multiple image blocks are processed by two groups of linear embedding windows and Swin Transformer blocks to obtain output features, wherein the linear embedding windows cut the image into small blocks and map them into high-dimensional vectors to form a serialized input. In the second extraction block, the input features are processed by two groups of merged adjacent windows and Swin Transformer blocks to obtain output features; in the third extraction block, the input features are processed by six groups of merged adjacent windows and Swin Transformer blocks to obtain output features; in the fourth extraction block, the input features are processed by two groups of merged adjacent windows and Swin Transformer blocks to obtain output features. The merging adjacent windows operation can adjust the resolution and number of channels to construct a multi-scale representation. The Swin Transformer block in the embodiment of the present invention adopts the Swin Transformer block in the paper "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows" published in the "2021 IEEE / CVF International Conference on Computer Vision". The window-based multi-head self-attention (W-MSA) in the Swin Transformer block can efficiently capture local features, and the shifted window self-attention (SW-MSA) can realize information fusion between windows and enhance global modeling.
[0067] In some embodiments, the neck module adopts a PAFPN (Panoptic Feature Pyramid Network) structure. Specifically, after the output feature of the fourth extraction block is upsampled, it is channel-joined with the output feature of the third extraction block to obtain a joint feature; after the joint feature is upsampled, it is channel-joined with the output feature of the second extraction block to obtain a first transformation feature; after the first transformation feature is convolved to compress the feature size, it is channel-joined with the joint feature to obtain a second transformation feature; after the second transformation feature is channel-joined with the output feature of the fourth extraction block, a third transformation feature is obtained; and the three transformation features are input into the detection module for target recognition.
[0068] In the detection module of some embodiments, transformation features of multiple scales are integrated, and the integrated features are subjected to RPN operation to obtain candidate pod targets and stem targets; the obtained pod targets and stem targets are combined with the integrated features to perform ROI alignment, and the aligned content is subjected to full connection operation to obtain stem images, pod independent images and their respective categories.
[0069] In the detection module of an embodiment of the present invention, the process of performing RPN operations on the integrated features includes: performing a convolution operation with a convolution kernel of 3×3 on the integrated features, and then performing two convolution operations with a convolution kernel of 1×1 on the convolved features in parallel to complete target classification and target bounding box regression, respectively. The process of performing a full connection operation on the aligned content includes: inputting the aligned content into the fully connected layer twice in succession, and the output features of the second fully connected layer enter the two fully connected layers in parallel to complete target category recognition and target bounding box determination.
[0070] The pod image restoration network restores the pod independent image; the pod image restoration network includes a discriminator and a generator. The generator is used to restore the input pod independent image, and the discriminator is used to judge the authenticity of the restored pod independent image.
[0071] In some embodiments, the peapod image restoration network uses a PIX2PIX network model. Specifically, the generator uses a U-Net structure. In the discriminator, multiple convolutions are used to replace the image blocking operation in PatchGAN. That is, the generator's restored peapod independent image and the corresponding real peapod independent image are subjected to multiple consecutive convolution operations to complete feature extraction. The extracted features are then subjected to subsequent image blocking operations in PatchGAN to complete feature comparison. An authenticity score is given to evaluate the authenticity of the peapod independent image, and the authenticity score is used to guide the generator to complete the restoration of the peapod independent image. In this embodiment of the present invention, after performing four consecutive convolution operations on the input image, feature extraction is completed and the extracted features are input into the PatchGAN module.
[0072] The present invention also provides a soybean seed testing method based on Transformer-faster-Rcnn improvement, such as Figure 7 As shown, including:
[0073] S1: Obtain a pod image dataset and preprocess the pod image dataset to obtain a training set that includes pod occlusion conditions. In some embodiments, step S1 includes:
[0074] S11: Capturing mature soybean plants in different scenarios, randomly spreading the soybean pods on the plants above a background plate, placing the soybean stems below the background plate, and photographing them using a camera to obtain images of the pods. In this embodiment of the present invention, the scenarios for capturing mature soybean plants include natural growth scenarios, adverse stress scenarios (such as drought, salinization, waterlogging, etc.), and biological stress scenarios (such as insect pests, bacteria, viruses, etc.), and photographing the pods and stems on the background plate using a camera.
[0075] like Figure 8 As shown, there are many different pod occlusion scenarios, with significant randomness, including but not limited to: no occlusion or only slight interference; occlusion of the top, middle, tail, or sides of a single pod; and occlusion of multiple pods side by side or staggered. This makes dataset creation challenging. Insufficient occlusion modeling in a dataset can lead to unknown network repair scenarios, resulting in repair failures and, indirectly, the inability to extract phenotypic traits from the pods. Therefore, embodiments of the present invention consider all possible occlusion scenarios and image each type of occlusion separately.
[0076] S12: Redundant portions of the pod image are removed, and the pod image is cropped to a uniform pixel size. The pods and stems in the cropped image are labeled to generate a labeling file. In this embodiment of the present invention, the collected pod image is labeled with the corresponding pod and stem type labels using the LabelImg tool. Bean pod types include bad pods, first-class pods, second-class pods, third-class pods, and fourth-class pods.
[0077] S13: Perform image enhancement on the cropped image and the annotation file to obtain a training set, which includes the cropped bean pod image and the corresponding annotation file. Among them, image enhancement includes HSV color gamut adjustment, random image scaling, random image cropping, and mosaic operation of partial image areas on the cropped image and the annotation file. In an embodiment of the present invention, image enhancement is used to simulate the same occlusion situation at different angles, HSV color gamut adjustment is used to simulate soybean pictures imaged in different environments, including adjusting the brightness, contrast, saturation, and hue parameters of the soybean picture for simulation, and mosaic operation of partial image areas is used to simulate the imaging blur caused by the failure of the imaging device to focus. The present invention can simulate the situation of bean pod occlusion in various situations by performing data enhancement on the original soybean image, and greatly expands the data set. Each enhancement method gives the bean pod image repair network new data, thereby improving the robustness of the repair network.
[0078] S2: Using the training set obtained in step S1, the soybean seed detection system based on the improved Transformer-faster-RCNN provided by the present invention is trained to obtain a soybean seed detection model. In some embodiments, in step S2, the soybean recognition network is trained using a recognition loss function, and the pod image restoration network is trained using a restoration objective function; wherein the recognition loss function is:
[0079] ;
[0080] Among them, L({p i},{t i}) represents the recognition loss function, p i Indicates the probability that a target exists in the i-th bounding box during target recognition, Indicates the corresponding true label, t i represents the compensation between the i-th predicted bounding box and the corresponding true bounding box coordinates, Indicates the coordinates of the real bounding box where the target exists, N cls Indicates the batch size during training, N reg Represents the number of recognized targets, and λ represents the adjustment weight of the recognition loss function; represents the cross entropy loss function, namely:
[0081] ;
[0082] Represents the Focal_loss loss function, which is composed of smooth L1 The losses are accumulated, that is:
[0083] ;
[0084] ;
[0085] The repair objective function is:
[0086] ;
[0087] in, represents the repair objective function, D represents the discriminator, G represents the generator, LOSS1 represents the L1 loss function, Indicates the repair weight; LOSS cGAN (G,D) represents the repair loss function, which is:
[0088] LOSS cGAN (G,D)=E x,y [logD(x,y)]+E x,z [log(1-D(x,G(x,y)))];
[0089] Among them, x represents the real image, y represents the corresponding constraint, z represents the noise, E x,y [logD(x,y)] represents the discriminator’s expectation of the real image, E x,z [log(1-D(x,G(x,y)))] represents the discriminator’s expectation of the image restored by the generator.
[0090] S3: According to the training results of step S2, the hyperparameters of the soybean seed testing model are adjusted during training until the optimal soybean seed testing model is obtained. In the embodiment of the present invention, the hyperparameters of the training include: using the adam optimizer, adopting the annealing cosine training strategy, the final learning rate is reduced to 1e-6, the training size is 640×640, the batch size is 32, the number of categories is 6, the initial learning rate is 1e-2, and the number of iterations is 400. According to the loss rate change curve, when the loss tends to be stable with the increase of the number of iterations, the network converges until 400 iterative trainings are completed to obtain the current soybean seed testing model.
[0091] S4: Input the pod image to be tested into the optimal soybean seed testing model obtained in step S4 to obtain a repaired independent pod image of each pod in the pod image, as well as a stem image; calculate and count the traits of the stem image and all independent pod images, thus completing the soybean seed testing.
[0092] The pod identification result in the embodiment of the present invention is as follows: Figures 9 to 11 As shown, the pod repair effect in the embodiment of the present invention is as follows Figure 12 shown. Figure 9 and Figure 10 The bounding boxes of different colors represent the different types of pods that have been identified. The types of pods identified and their class probabilities are marked above the bounding boxes. Figure 10 Among them, soybeon_1 represents the first type of pod, soybeon_2 represents the second type of pod, soybeon_3 represents the third type of pod, and soybeon_4 represents the fourth type of pod ( Figure 10 not shown), with Figure 10 For example, the label "soybeon_2 0.88" on one of the bounding boxes indicates that the pod in the bounding box is identified as a Class 2 pod, and the probability that it is a Class 1 pod is 0.88, or 88%. Figure 9 and Figure 11 The green box in the image represents the identified stem and its probability of being a stem. Figure 11 For example, if one of the bounding boxes is labeled "stem 0.72", it means that the content in the bounding box is a stem, and the probability that it is a stem is 0.72, or 72%. Figure 12 The middle left column shows different degrees of occlusion (no occlusion with slight interference, side occlusion, and side-by-side occlusion of multiple pods). Figure 12 The right column shows the corresponding repaired results. Figure 12 It can be seen from the figure that the present invention can effectively restore pods with different degrees of occlusion, laying a good foundation for subsequent pod trait statistics.
[0093] In some embodiments, the pod traits in step S4 include pod length, pod width, pod circumference and pod projected area; the extraction process of pod traits is as follows: Figure 13 As shown, it includes: graying the restored pod independent image, extracting the grayscale image of the R channel with the largest contrast between the pod grains and the background for subsequent shape extraction, as shown in FIG. Figure 13 (a) in the figure; use the threshold segmentation algorithm to binarize the R channel image and obtain the pod binary image, as shown in Figure 13 (b) in the figure; extract the outer contour of the pod based on the binary image of the pod, including the coordinates of each point on the contour; calculate the length, width, perimeter and projection area of the pod based on the obtained contour coordinates, as shown in Figure 13 (c) in the.
[0094] In an embodiment of the present invention, the number of pods contained in different tags can also be read from the number of beans in the entire plant, multiplied by the number of pods determined by the tags, and then added together to obtain the number of beans in the entire plant. For example, a bad pod is defined as P0, which contains no pods, that is, the number of pods corresponding to the bad pod is 0; a first-class pod is defined as P1, which contains 1 pod, that is, the number of pods corresponding to the first-class pod is 1; a second-class pod is defined as P2, which contains 2 pods, that is, the number of pods corresponding to the second-class pod is 2; a third-class pod is defined as P3, which contains 3 pods, that is, the number of pods corresponding to the third-class pod is 2; a fourth-class pod is defined as P4, which contains 3 pods, that is, the number of pods corresponding to the fourth-class pod is 2. In this case, the number of beans in the entire plant = 0×P0+1×P2+2×P2+3×P3+4×P4.
[0095] In an embodiment of the present invention, the process of calculating the length of a pod includes, in the process of obtaining the length of the pod, arbitrarily selecting two points on the outer contour of the pod, and calculating the distance between the two points, and the maximum value of all distances is the length of the pod. In the process of obtaining the width of the pod, the normal slope of the straight line between the two points of the maximum distance is calculated, and then the distance between the intersection of the normal with the same normal slope and the outer contour of the pod under different intercepts is calculated, and the maximum value of all distances is the width of the pod. In the process of obtaining the perimeter of the pod, the sum of the distances between all adjacent pixel points on the outer contour of the pod is the perimeter of the pod. In an embodiment of the present invention, based on the points on the contour of the pod, if the two points are adjacent in the upper and lower or left and right directions, the distance between the two points is defined as 1, and if the two points are adjacent in the upper left, lower left, upper right or lower right directions, the distance between the two points is defined as , the sum of the distances between all adjacent pixels on the contour is the pod perimeter. The pod projected area is the number of pixels within the pod's outer contour.
[0096] In an embodiment of the present invention, the stem traits include the number of main stem nodes, the maximum internode distance, the minimum internode distance, the average internode distance, the total length of the stem nodes, the farthest internode distance, and the curvature. Specifically, the target frame parameters corresponding to the label value are read according to the return value obtained by identification, including the coordinates (x1, y1, x2, y2) of the original image corresponding to the upper left point and the lower right point of the target frame, and the length and width of the stem are obtained as x2-x1 and y2-y1 respectively, and the coordinate position of the original image where the center point of the stem is located is ((x2-x1) / 2+x1, (y2-y1) / 2+y1). According to the obtained center coordinates of the first and last stems, the approximate length of the main stem can be obtained; according to the center coordinates of the adjacent stems, the lengths of several main trunks can be obtained. Count the distances between adjacent stems (X1, X2, ..., X n ), and obtain the maximum value X by comparison MAX , X MAX is the maximum internode distance. Similarly, the minimum value X is obtained by comparison.MIN , X MIN The minimum internode distance is obtained by averaging the distances between adjacent stems (X1, X2, ..., X n ) the average value X MEAN , X MEAN is the average internode distance, and the total length of the stem node X is obtained by accumulation. SUM =X1+X2+…+X n ; Get the center coordinates of the first stem node (X1, Y1) and the center coordinates of the last stem node (X n ,Y n ), through the formula The ratio of the farthest internode distance to the total length of the stem node is the curvature, which is used to calculate the degree of curvature of the stem.
[0097] The statistical data fitting effects of the number of bad pods, first-class pods, second-class pods, third-class pods and fourth-class pods in the embodiment of the present invention correspond to Figures 14 to 18 The statistical data fitting effects of the number of all pods and the number of all beans correspond to Figure 19 and Figure 20 . Figures 14 to 20 The vertical axis represents the number of corresponding pods counted using the method provided by the present invention, and the horizontal axis represents the number of corresponding pods counted manually. Figures 14 to 20 It can be seen that the number of corresponding pods counted by the method provided by the present invention is close to the corresponding number counted manually, indicating that the method provided by the present invention can effectively complete the statistics of the number of pods or stem nodes. 2 The R-squared indicator reflects the accuracy of quantitative statistics, where R 2 Also known as the coefficient of determination or determination coefficient, it is an indicator used to measure the goodness of fit of the regression model to the observed data. It indicates the degree to which the independent variable in the model explains the dependent variable, and its value range is between 0 and 1. 2 The closer it is to 1, the better the model fits the data; the closer it is to 0, the weaker the model's ability to explain the data. The R values of the number of bad pods, first-class pods, second-class pods, third-class pods, and fourth-class pods, as well as the number of all pods and all beans counted using the method provided by the present invention are as follows: 2 The indicators are 0.9824, 0.9172, 0.9749, 0.8194, 0.9009, 0.9737 and 0.9678 respectively. It can be seen that R 2 The indices are all close to 1, which further illustrates that the fitting effect of the method provided by the present invention is very good, that is, there is a highly significant linear relationship between the method provided by the present invention and the manual recognition method.
[0098] In addition, the embodiment of the present invention uses five indicators, namely precision P (Precision), recall R (Recall), mAP_50, mAP_75, and mAP50-95, to evaluate the target detection performance of the soybean recognition network. Among them, precision P refers to the proportion of samples predicted as positive that are actually positive, which measures the accuracy of the prediction result; recall R is the proportion of true positive samples that are correctly predicted as positive, reflecting the model's ability to capture positive samples; mAP_50 represents the average precision when the IoU threshold is 0.5, reflecting the detection performance of the model on different categories; mAP_75 represents the average precision when the IoU threshold is 0.75, which has higher detection accuracy requirements than mAP_50; mAP_s represents the average precision for small targets, which evaluates the model's ability to detect small targets; mAP50-95 represents the average precision in the IoU threshold range of 0.5 to 0.95, which comprehensively measures the performance of the model at different IoU thresholds. The indicator evaluation results are shown in Table 1:
[0099] Table 1:
[0100]
[0101] Table 1 shows that comprehensive performance metrics are presented for all categories, with relatively balanced precision and recall overall. The mAP series of metrics demonstrates performance at different IoU thresholds. Among the specific categories, pods from class 1 to class 4, as well as stems, performed well across most metrics. In particular, class 4 pods achieved a precision of 0.932, demonstrating high accuracy in predicting this category.
[0102] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved. This is not limited herein.
[0103] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A soybean seed testing system based on Transformer-faster-RCNN, characterized in that: It includes pod recognition network and pod image restoration network; among them, The pod recognition network includes a trunk module of a Swin-Transform structure, a neck module of a pyramid structure, and a detection module; a pod image is input into the trunk module for feature extraction to obtain extracted features at multiple scales; residual features at multiple scales are input into the neck module for feature transformation at multiple scales to obtain transformed features at multiple scales; the transformed features at multiple scales are input into the detection module for target recognition to obtain a stem image in the pod image and an independent image of each pod in the pod image; The backbone module includes a feature segmentation block, a first extraction block, a second extraction block, a third extraction block, and a fourth extraction block; wherein the feature segmentation block performs image segmentation on the input image to obtain a plurality of non-overlapping image blocks of equal size; the plurality of image blocks are sequentially subjected to feature extraction by the first extraction block, the second extraction block, the third extraction block, and the fourth extraction block; the output features of the second extraction block, the third extraction block, and the fourth extraction block are input into the neck module; In the neck module: after performing an upsampling operation on the output features of the fourth extraction block, the output features of the third extraction block are channel-joined to obtain a joint feature; after performing an upsampling operation on the joint feature, the output features of the second extraction block are channel-joined to obtain a first transformed feature; after performing a convolution operation on the first transformed feature to compress the feature size, the first transformed feature is channel-joined with the joint feature to obtain a second transformed feature; after performing channel-joining on the second transformed feature and the output features of the fourth extraction block, the third transformed feature is obtained; the three transformed features are input into the detection module for target recognition; In the detection module, the transformation features of multiple scales are integrated, and the integrated features are subjected to RPN operation to obtain candidate pod targets and stem targets; the obtained pod targets and stem targets are aligned with the integrated features to perform ROI alignment, and the aligned content is subjected to full connection operation to obtain the stem image, the pod independent image and their respective categories; The peapod image restoration network restores the independent peapod image; the peapod image restoration network includes a discriminator and a generator, the generator is used to restore the input independent peapod image, and the discriminator is used to judge the authenticity of the restored independent peapod image.
2. The soybean seed testing system based on Transformer-faster-RCNN improvement according to claim 1 is characterized in that: In the first extraction block, multiple input image blocks are processed by multiple sets of linear embedding windows and SwinTransformer blocks to obtain output features; In the second extraction block, the third extraction block, and the fourth extraction block, the input features are processed by combining multiple groups of adjacent windows and the Swin Transformer block to obtain output features.
3. The soybean seed testing system based on Transformer-faster-RCNN improvement according to claim 1, characterized in that: The generator has a U-Net structure. In the discriminator, multiple consecutive convolution operations are performed on the independent pod images restored by the generator and the corresponding real independent pod images to complete feature extraction. The extracted features are subjected to subsequent image segmentation operations in PatchGAN to complete feature comparison, and an authenticity score is given to evaluate the authenticity of the independent pod images. The authenticity score is used to guide the generator to complete the restoration of the independent pod images.
4. A soybean seed testing method based on Transformer-faster-RCNN improvement, characterized in that: include: S1: Obtain a pod image dataset and preprocess the pod image dataset to obtain a training set containing pod occlusion conditions; S2: Using the training set obtained in step S1, the soybean seed detection system improved based on Transformer-faster-RCNN according to any one of claims 1 to 3 is trained to obtain a soybean seed detection model; S3: According to the training result of step S2, adjusting the hyperparameters when training the soybean seed selection model until an optimal soybean seed selection model is obtained; S4: Input the pod image to be tested into the optimal soybean seed testing model obtained in step S4 to obtain a restored independent pod image of each pod in the pod image, as well as a stem image; calculate and count the traits of the stem image and all independent pod images, thereby completing the soybean seed testing.
5. The soybean seed testing method based on Transformer-faster-RCNN improvement according to claim 4, characterized in that: Step S1 includes: Mature soybean plants are collected in different scenes, the pods on the soybean plants are randomly spread above a background plate, the stems of the soybean plants are placed below the background plate, and photographed using a camera to obtain pod images; removing redundant parts from the bean pod image, cropping the bean pod image into an image of uniform pixel size, and annotating the bean pods and stems in the cropped image to obtain an annotation file; Image enhancement is performed on the cropped images and the annotation files to obtain a training set, which includes the cropped pod images and the corresponding annotation files.
6. The soybean seed testing method based on Transformer-faster-RCNN improvement according to claim 5, characterized in that: In step S2, the pod recognition network is trained using a recognition loss function, and the pod image restoration network is trained using a restoration objective function; wherein the recognition loss function is: ; Among them, L({p i },{t i }) represents the recognition loss function, p i Indicates the probability that a target exists in the i-th bounding box during target recognition, Indicates the corresponding true label, t i represents the compensation between the i-th predicted bounding box and the corresponding true bounding box coordinates, Indicates the coordinates of the real bounding box where the target exists, N cls Indicates the batch size during training, N reg represents the number of identified targets, and λ represents the adjustment weight for adjusting the recognition loss function; represents the cross entropy loss function, namely: ; Represents the Focal_loss loss function, which is composed of smooth L1 The losses are accumulated, that is: ; ; The repair objective function is: ; in, represents the restoration objective function, D represents the discriminator, G represents the generator, represents the L1 loss function, represents the repair weight; Represents the repair loss function, which is: ; Among them, x represents the real image, y represents the corresponding constraint condition, and z represents the noise. represents the discriminator’s expectation of the real image, Represents the discriminator's expectation of the image restored by the generator.
7. The soybean seed testing method based on Transformer-faster-RCNN improvement according to claim 5, characterized in that: In step S4, the pod traits include pod length, pod width, pod circumference and pod projected area; the pod trait extraction process includes: graying the repaired pod independent image, extracting the R channel image with the maximum contrast between the pod grains and the background; binarizing the R channel image using a threshold segmentation algorithm to obtain a pod binary image; extracting the pod outer contour based on the pod binary image, including the coordinates of each point on the contour; and calculating the pod length, width, circumference and pod projected area based on the obtained contour coordinates.
Citation Information
Patent Citations
Soybean pod quantity statistical method based on machine vision
CN114724141A
Automated phenotyping for seed POD shatter and seed POD drop
WO2024173245A2