Improved soybean seed testing system and method based on Faster-R-CNN
By improving the Faster-R-CNN system, combined with ResNet101 and the cascaded attention mechanism, the problems of low efficiency and insufficient precision in counting the number of pods and stems in the existing technology are solved, and efficient and accurate detection and counting of pods and stems are achieved.
Patent Information
- Application Number
- CN202510950127.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Traditional soybean seed testing methods are inefficient and highly subjective, making it difficult to meet the needs of high-throughput breeding. Existing methods have limitations such as insufficient model generalization capabilities, low multi-scale target detection accuracy, and high data labeling costs, making it impossible to quickly and accurately count the number of soybean grains and stems.
A soybean seed testing system based on the improved Faster-R-CNN is adopted, including a pod recognition network and a pod image restoration network. By optimizing the network structure, enhancing data diversity and refining the training strategy, the ResNet101 network and cascade attention mechanism are used for feature extraction, combined with Focal_loss and PatchGAN for image restoration, to achieve accurate detection and counting of pods and stems.
It significantly improves the efficiency and accuracy of phenotypic recognition in complex scenarios, enhances feature extraction capabilities and detection accuracy, solves occlusion problems, and achieves fast and accurate pod and stem counting.
Smart Images

Figure CN120451521B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of soybean seed testing, and in particular relates to a soybean seed testing system and method improved based on Faster-R-CNN. Background Art
[0002] In agricultural breeding, soybean plant phenotypic data (such as pod number, type, and stem structure) is a key basis for evaluating variety quality. Traditional manual seed testing methods are inefficient and highly subjective, making them difficult to meet the needs of high-throughput breeding. Currently, there are two main methods for high-throughput soybean bean and stem counting:
[0003] The first method involves threshing the soybean pods. The number of stem nodes is manually counted, and the beans are then spread out on a flat surface. Because the beans are all round, they avoid blocking each other and are distributed roughly evenly. A visible light camera is then used to capture the image, and digital image processing techniques are used to count the beans. This method takes a long time to thresh, resulting in smaller beans, and some beans may be damaged or lost, which can affect the accuracy of bean counting.
[0004] The second method involves removing pods from mature soybean plants, manually counting the number of stem nodes, and then neatly placing the pods on a flat plate. These pods are then imaged using a visible light camera and digital image processing techniques to identify pods that are not obstructed by each other. Pods containing different types of beans are then manually classified, and the total number of beans is finally calculated. While this method avoids missed detections, it requires extensive manual labor, is time-consuming, and is incompatible with the extraction of obstructed pods.
[0005] As can be seen above, none of the aforementioned methods can quickly and accurately count soybean pods and stems. Direct counting results in large errors due to the small size of the pods, while indirect counting relies on manual placement and classification, which is time-consuming. However, existing methods suffer from limitations such as insufficient model generalization, low multi-scale object detection accuracy, and high data annotation costs. Summary of the Invention
[0006] In light of this, the present invention aims to provide a soybean seed identification system and method based on an improved Faster-R-CNN. This system uses a pod recognition network to accurately detect and count the number of pods and stems in their naturally mature state, while a pod image restoration network addresses the issue of pod occlusion. By optimizing the network structure, enhancing data diversity, and refining training strategies, the efficiency and accuracy of phenotypic identification in complex scenarios are significantly improved.
[0007] To achieve the above object, the technical solution created by the present invention is implemented as follows:
[0008] A soybean seed testing system based on an improved Faster-R-CNN includes a pod recognition network and a pod image restoration network; wherein the pod recognition network includes a trunk module, a neck module, and a detection module; the pod image is input into the trunk module for multiple residual feature extractions to obtain residual features of multiple scales; the residual features of multiple scales are input into the neck module for multi-scale pyramid feature transformation to obtain transformed features of multiple scales; the transformed features of multiple scales are input into the detection module for target recognition to obtain the stem image in the pod image and the unique image of each pod in the pod image. In the detection module, the transformation features of multiple scales are integrated, and the integrated features are subjected to RPN operation to obtain candidate pod targets and stem targets; the obtained pod targets and stem targets are aligned with the integrated features for ROI alignment, and the aligned content is fully connected to obtain the stem image, pod independent image and their respective categories; the pod image restoration network is used to restore the pod independent image; the pod image restoration network includes a discriminator and a generator, the generator is used to restore the input pod independent image, and the discriminator is used to judge the authenticity of the restored pod independent image.
[0009] Furthermore, the backbone module adopts the ResNet101 structure, including a cascaded shallow feature extraction unit, a first residual unit, a second residual unit, a third residual unit and a fourth residual unit; in the shallow feature extraction unit, the input image is subjected to shallow feature extraction through a convolution operation, and the convolved features are pooled; the output features of the shallow feature extraction unit are sequentially subjected to feature extraction by the first residual unit, the second residual unit, the third residual unit and the fourth residual unit, and the output features of the second residual unit, the third residual unit and the fourth residual unit enter the neck module; in each residual unit, multiple residual operations are performed on the input features; in each residual operation, multiple consecutive convolution operations are performed on the input features, and the features obtained by the last convolution operation are subjected to a cascaded attention operation; the output features of the cascaded attention operation are added to the corresponding elements of the input features to obtain the output features of the current residual operation.
[0010] Furthermore, in the cascade attention operation, the channel attention operation is first performed on the input features, and then the spatial attention operation is performed on the output features of the channel attention operation to obtain the output features; in the channel attention operation, the maximum pooling and average pooling in the spatial direction are performed on the input features in parallel, and then the two pooled features are respectively subjected to convolution operations to complete feature extraction; after adding the corresponding elements of the two extracted features, the sigmoid activation operation is performed on the added features, and the activated features are multiplied by the corresponding elements of the input features to obtain channel features; in the spatial attention operation, the maximum pooling and average pooling in the channel direction are performed on the channel features in parallel, and after the two pooled features are channel-spliced, the spliced features are sequentially subjected to convolution and sigmoid activation operations, and the activated features are multiplied by the corresponding elements of the channel features to obtain the output features.
[0011] Furthermore, in the neck module, the output features of the fourth residual unit are upsampled and then channel-joined with the output features of the third residual unit to obtain a joint feature; the joint feature is upsampled and then channel-joined with the output features of the second residual unit to obtain a first transformation feature; the first transformation feature is convolved to compress the feature size and then channel-joined with the joint feature to obtain a second transformation feature; the second transformation feature is channel-joined with the output features of the fourth residual unit to obtain a third transformation feature; the three transformation features are input into the detection module for target recognition.
[0012] Furthermore, the generator adopts a U-Net structure. In the discriminator, the generator's restored pod image and the corresponding real pod image are subjected to multiple continuous convolution operations to complete feature extraction. The extracted features are subjected to subsequent image segmentation operations in PatchGAN to complete feature comparison, and an authenticity score is given to evaluate the authenticity of the pod image. The authenticity score is used to guide the generator to complete the restoration of the pod image.
[0013] A soybean seed testing method based on Faster-R-CNN improvement, comprising:
[0014] S1: Obtain a pod image dataset and preprocess it to obtain a training set that includes pod occlusion conditions;
[0015] S2: Using the training set obtained in step S1, the soybean seed detection system improved based on Faster-R-CNN provided by the present invention is trained to obtain a soybean seed detection model;
[0016] S3: According to the training results of step S2, the hyperparameters of the soybean seed selection model are adjusted until the optimal soybean seed selection model is obtained;
[0017] S4: Input the pod image to be tested into the optimal soybean seed testing model obtained in step S4 to obtain a repaired independent pod image of each pod in the pod image, as well as a stem image; calculate and count the traits of the stem image and all independent pod images, thus completing the soybean seed testing.
[0018] Furthermore, step S1 includes: collecting mature soybean plants in different scenes, scattering the pods on the soybean plants arbitrarily above the background board, placing the stems of the soybean plants below the background board, and shooting with a shooting device to obtain a pod image; removing redundant parts in the pod image, and cropping the pod image into an image of uniform pixel size, annotating the pods and stems in the cropped image to obtain an annotation file; performing image enhancement on the cropped image and the annotation file to obtain a training set, which includes the cropped pod image and the corresponding annotation file.
[0019] Furthermore, in step S2, the soybean recognition network is trained using the recognition loss function, and the pod image restoration network is trained using the restoration objective function; wherein the recognition loss function is:
[0020] ;
[0021] Among them, L({p i},{t i}) represents the recognition loss function, p i Indicates the probability that a target exists in the i-th bounding box during target recognition, p i 'Indicates the corresponding true label, t i represents the compensation between the i-th predicted bounding box and the corresponding true bounding box coordinates, t i 'Indicates the coordinates of the real bounding box where the target is located, N cls Indicates the batch size during training, N reg Indicates the number of identified targets, λ indicates the adjustment weight of the recognition loss function; L cls (p i ,p i ') represents the cross entropy loss function, that is:
[0022] L cls (p i ,p i ')=-p i 'logp i -(1-p i ')log(1-p i );
[0023] L reg (t i ,t i') represents the Focal_loss loss function, which is represented by smooth L1 The losses are accumulated, that is:
[0024] ;
[0025] ;
[0026] The repair objective function is:
[0027] G'=argminmaxLOSS cGAN (G,D)+λ'LOSS1(G);
[0028] Among them, G' represents the restoration objective function, D represents the discriminator, G represents the generator, LOSS1 represents the L1 loss function, and λ' represents the restoration weight; LOSS cGAN (G,D) represents the repair loss function, which is:
[0029] LOSS cGAN (G,D)=E x,y [logD(x,y)]+E x,z [log(1-D(x,G(x,y)))];
[0030] Among them, x represents the real image, y represents the corresponding constraint, z represents the noise, E x,y [logD(x,y)] represents the discriminator’s expectation of the real image, E x,z [log(1-D(x,G(x,y)))] represents the discriminator’s expectation of the image restored by the generator.
[0031] Furthermore, in step S4, the pod traits include pod length, pod width, pod circumference and pod projected area; the pod trait extraction process includes: grayscale the repaired pod independent image, extracting the R channel image with the maximum contrast between the pod grains and the background; binarizing the R channel image using a threshold segmentation algorithm to obtain a pod binary image; extracting the pod outer contour based on the pod binary image, including the coordinates of each point on the contour; and calculating the pod length, width, circumference and pod projected area based on the obtained contour coordinates.
[0032] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0033] (1) The present invention creates a soybean seed identification system and method based on the improved Faster-R-CNN. A pod recognition network is used to accurately detect and count the number of pods and stems in their natural mature state. A pod image restoration network is used to address the problem of pod occlusion. By optimizing the network structure, enhancing data diversity, and refining the training strategy, the efficiency and accuracy of phenotypic recognition in complex scenarios are significantly improved.
[0034] (2) In the soybean seed testing system improved based on Faster-R-CNN created by the present invention, the original Resnet50 network is replaced with the Resnet101 network. Deepening the network depth is achieved by increasing the number of nonlinear transformation layers, so that the model has stronger feature abstraction ability and generalization performance, thereby meeting the feature extraction requirements of complex environments; in addition, in order to improve the accuracy and robustness of the algorithm, the cascaded attention mechanism combining space and channel is placed behind each Resnet residual layer, which can further focus on the target features of different types of pods in the feature map, thereby quickly and accurately focusing on the target area, giving the target area feature map a higher weight, thereby highlighting the target and suppressing other irrelevant areas, allowing the target frame to be more accurately positioned in the target area, and improving detection accuracy; the present invention also uses bidirectional fusion to make each layer of the feature pyramid have strong semantics and fine-grained details at the same time, such as significantly improving the target extraction effect in dense scenes; this structure can more accurately extract the feature information of targets of different pod scales, reducing missed detection or false detection caused by changes in pod scale. Improve the positioning accuracy of feature maps for soybean targets, provide better positioning points in the bounding box regression task, and generate candidate boxes in the target area;
[0035] (3) In the soybean seed classification method based on the improved Faster-R-CNN described in the present invention, Focal_loss is used to replace BCE_loss in the category classification process. Focal_loss is improved based on the standard BCE loss function. Its core idea is to make the model pay more attention to difficult-to-classify samples by dynamically adjusting sample weights. Focal_loss introduces a modulation factor to reduce the loss contribution of easy-to-classify samples, making the model pay more attention to difficult-to-classify samples. This design helps to solve the problem of category imbalance and improve the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings, which constitute part of the present invention, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0037] Figure 1This is a schematic diagram of the structure of the soybean seed testing system improved based on Faster-R-CNN according to an embodiment of the present invention;
[0038] Figure 2 This is a schematic diagram of the structure of the peapod identification network described in an embodiment of the present invention;
[0039] Figure 3 A schematic diagram of a residual operation according to an embodiment of the present invention;
[0040] Figure 4 A schematic diagram of the cascaded attention operation according to an embodiment of the present invention;
[0041] Figure 5 A schematic diagram of a peapod image restoration network according to an embodiment of the present invention;
[0042] Figure 6 A schematic diagram of a generator according to an embodiment of the present invention;
[0043] Figure 7 A schematic diagram of a discriminator according to an embodiment of the present invention;
[0044] Figure 8 A schematic diagram of the process of the soybean seed testing method improved based on Faster-R-CNN according to an embodiment of the present invention;
[0045] Figure 9 Creating the pod image described in the embodiment of the present invention;
[0046] Figure 10 This is a diagram of the pod target recognition result described in the embodiment of the present invention;
[0047] Figure 11 This is a partial enlarged view of the pod target recognition described in the embodiment of the present invention;
[0048] Figure 12 A partially enlarged view of the stem target identification according to an embodiment of the present invention;
[0049] Figure 13 A comparison chart of the results of the pod image restoration described in the embodiment of the present invention;
[0050] Figure 14 This is a diagram of the process of extracting pod traits according to an embodiment of the present invention;
[0051] Figure 15 The data fitting effect diagram of the bad pod number statistics described in the embodiment of the present invention is created
[0052] Figure 16 This is a data fitting effect diagram of the statistics of the number of pods of a type described in an embodiment of the present invention;
[0053] Figure 17 This is a data fitting effect diagram of the second type of pod number statistics described in the embodiment of the present invention;
[0054] Figure 18 This is a data fitting effect diagram of the statistics of the number of three types of pods described in the embodiment of the present invention;
[0055] Figure 19 This is a data fitting effect diagram of the statistics of the number of four types of pods described in the embodiment of the present invention;
[0056] Figure 20 This is a data fitting effect diagram of the statistics of the number of all pods described in the embodiment of the present invention;
[0057] Figure 21 This is a data fitting effect diagram for the statistics of the number of all beans described in the embodiment of the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation of the present invention.
[0059] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0060] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second" and the like are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, features defined as "first", "second" and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0061] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art can understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0062] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments.
[0063] like Figures 1 to 7 As shown, the soybean seed detection system based on Faster-R-CNN improved in the embodiment of the present invention includes a pod recognition network and a pod image restoration network. Figure 2 As shown, it includes a trunk module, a neck module, and a detection module. The trunk module is used to extract soybean image features, the neck module is used to fuse and transform the features extracted by the feature extraction network at different depths, and the detection module is used to classify the targets detected by the system and optimize the position of the detected target bounding box. Specifically: the pod image is input into the trunk module for multiple residual feature extractions to obtain residual features of multiple scales; the residual features of multiple scales are input into the neck module for multi-scale pyramid feature transformation to obtain transformed features of multiple scales; the transformed features of multiple scales are input into the detection module for target recognition to obtain the stem image in the pod image and the independent image of each pod in the pod image.
[0064] In some embodiments, the backbone module adopts a ResNet101 structure, including a cascaded shallow feature extraction unit, a first residual unit, a second residual unit, a third residual unit, and a fourth residual unit; in the shallow feature extraction unit, shallow feature extraction is performed on the input image through a convolution operation, and the convolved features are pooled; the output features of the shallow feature extraction unit are sequentially subjected to feature extraction by the first residual unit, the second residual unit, the third residual unit, and the fourth residual unit, and the output features of the second residual unit, the third residual unit, and the fourth residual unit enter the neck module. In each residual unit, multiple residual operations are performed on the input features; in each residual operation, multiple consecutive convolution operations are performed on the input features, and the features obtained by the last convolution operation are subjected to a cascaded attention operation; the output features of the cascaded attention operation are added to the input features or the corresponding elements of the input image to obtain the output features of the current residual operation.
[0065] In the backbone module of the embodiment of the present invention, in the shallow feature extraction unit, the input image is subjected to shallow feature extraction through a convolution operation with a convolution kernel of 7×7, a channel number of 64, and a step size of 2, and the convolved features are subjected to a 3×3 pooling operation, and the step size of the pooling operation is 2. In the first residual unit, the input features are subjected to 3 residual operations; in the second residual unit, the input features are subjected to 4 residual operations; in the third residual unit, the input features are subjected to 23 residual operations; in the fourth residual unit, the input features are subjected to 3 residual operations. Specifically, each residual operation is as follows: Figure 3 As shown in the figure, after continuous convolution operations with kernels of 1×1, 3×3 and 1×1 are performed on the input features, the features obtained by the last convolution operation are subjected to a cascaded attention (CBAM) operation; the output features of the cascaded attention operation are added to the corresponding elements of the input features to obtain the output features of the current residual operation.
[0066] In the cascade attention operation of some embodiments, a channel attention operation is first performed on the input features, and then a spatial attention operation is performed on the output features of the channel attention operation to obtain output features; in the channel attention operation, maximum pooling and average pooling in the spatial direction are performed on the input features in parallel, and then convolution operations are performed on the two pooled features respectively to complete feature extraction; after adding the corresponding elements of the two extracted features, a sigmoid activation operation is performed on the added features, and the activated features are multiplied by the corresponding elements of the input features to obtain channel features; in the spatial attention operation, maximum pooling and average pooling in the channel direction are performed on the channel features in parallel, and after channel splicing of the two pooled features, convolution and sigmoid activation operations are performed on the spliced features in sequence, and the activated features are multiplied by the corresponding elements of the channel features to obtain output features.
[0067] The cascade attention operation of the embodiment of the present invention is as follows Figure 4 As shown in the figure, in the channel attention operation, spatial maximum pooling and average pooling are performed on the input features in parallel. Then, a 1×1 convolution operation and ReLU activation are performed on each of the two pooled features. Finally, a 1×1 convolution operation is performed on the two extracted features to complete feature extraction. After adding the corresponding elements of the two extracted features, a sigmoid activation operation is performed on the added features. The activated features are multiplied with the corresponding elements of the input features to obtain the channel features. In the spatial attention operation, channel maximum pooling and average pooling are performed on the channel features in parallel. After the two pooled features are concatenated, the concatenated features are sequentially convolved with a 1×1 convolution operation and a sigmoid activation operation to obtain the output features.
[0068] In some embodiments, the neck module adopts a PAFPN (Panoptic Feature Pyramid Network) structure. Specifically, after the output feature of the fourth residual unit is upsampled, it is channel-joined with the output feature of the third residual unit to obtain a stitched feature; after the stitched feature is upsampled, it is channel-joined with the output feature of the second residual unit to obtain a first transformed feature; after the first transformed feature is convolved to compress the feature size, it is channel-joined with the stitched feature to obtain a second transformed feature; after the second transformed feature is channel-joined with the output feature of the fourth residual unit, a third transformed feature is obtained; and the three transformed features are input into the detection module for target recognition.
[0069] In the detection module of some embodiments, transformation features of multiple scales are integrated, and the integrated features are subjected to RPN (Region Proposal Network) operation to obtain candidate pod targets and stem targets; the obtained pod targets and stem targets are aligned with the integrated features for ROI, and the aligned contents are fully connected to obtain stem images, pod independent images, and their respective categories. In the present invention, the integration of transformation features can be: corresponding elements of transformation features of multiple scales are added to complete the integration of transformation features; or the transformation features of multiple scales can be first channel-spliced, and then the spliced features are subjected to convolution operation to complete the integration of transformation features.
[0070] In the detection module of an embodiment of the present invention, the process of performing RPN operations on the integrated features includes: performing a convolution operation with a convolution kernel of 3×3 on the integrated features, and then performing two convolution operations with a convolution kernel of 1×1 on the convolved features in parallel to complete target classification and target bounding box regression, respectively. The process of performing a full connection operation on the aligned content includes: inputting the aligned content into the fully connected layer twice in succession, and the output features of the second fully connected layer enter the two fully connected layers in parallel to complete target category recognition and target bounding box determination.
[0071] The pod image restoration network repairs the pod independent image. The pod image restoration network is as follows Figure 5 As shown, it includes a discriminator and a generator. The generator is used to repair the input pod independent image, and the discriminator is used to judge the authenticity of the repaired pod independent image.
[0072] In some embodiments, the pod image restoration network adopts the PIX2PIX network model. Specifically, the generator is as follows: Figure 6 The U-Net structure shown in Figure 2. The discriminator is as follows: Figure 7As shown, multiple convolutions replace the image blocking operation in PatchGAN. This involves performing multiple consecutive convolutions on the generator's restored pod image and the corresponding real pod image to extract features. The extracted features are then compared using the subsequent image blocking operation in PatchGAN. A authenticity score is then assigned to assess the authenticity of the pod image, and the authenticity score is used to guide the generator in restoring the pod image. In this embodiment of the present invention, feature extraction is completed after four consecutive convolutions are performed on the input image, and the extracted features are input into the PatchGAN module.
[0073] A soybean seed testing method based on Faster-R-CNN improvement, such as Figure 8 Shown, including:
[0074] S1: Obtain a pod image dataset and preprocess the pod image dataset to obtain a training set that includes pod occlusion conditions. In some embodiments, step S1 includes:
[0075] S11: Capturing mature soybean plants in different scenarios, randomly spreading the soybean pods on the plants above a background plate, placing the soybean stems below the background plate, and photographing them using a camera to obtain images of the pods. In this embodiment of the present invention, the scenarios for capturing mature soybean plants include natural growth scenarios, adverse stress scenarios (such as drought, salinization, waterlogging, etc.), and biological stress scenarios (such as insect pests, bacteria, viruses, etc.), and photographing the pods and stems on the background plate using a camera.
[0076] like Figure 9 As shown, there are many different pod occlusion scenarios, with significant randomness, including but not limited to: no occlusion or only slight interference; occlusion of the top, middle, tail, or sides of a single pod; and occlusion of multiple pods side by side or staggered. This makes dataset creation challenging. Insufficient occlusion modeling in a dataset can lead to unknown network repair scenarios, resulting in repair failures and, indirectly, the inability to extract phenotypic traits from the pods. Therefore, embodiments of the present invention consider all possible occlusion scenarios and image each type of occlusion separately.
[0077] S12: Redundant portions of the pod image are removed, and the pod image is cropped to a uniform pixel size. The pods and stems in the cropped image are labeled to generate a labeling file. In this embodiment of the present invention, the collected pod image is labeled with the corresponding pod and stem type labels using the LabelImg tool. Bean pod types include bad pods, first-class pods, second-class pods, third-class pods, and fourth-class pods.
[0078] S13: Perform image enhancement on the cropped image and the annotation file to obtain a training set, which includes the cropped bean pod image and the corresponding annotation file. Among them, image enhancement includes HSV color gamut adjustment, random image scaling, random image cropping, and mosaic operation of partial image areas on the cropped image and the annotation file. In an embodiment of the present invention, image enhancement is used to simulate the same occlusion situation at different angles, HSV color gamut adjustment is used to simulate soybean pictures imaged in different environments, including adjusting the brightness, contrast, saturation, and hue parameters of the soybean picture for simulation, and mosaic operation of partial image areas is used to simulate the imaging blur caused by the failure of the imaging device to focus. The present invention can simulate the situation of bean pod occlusion in various situations by performing data enhancement on the original soybean image, and greatly expands the data set. Each enhancement method gives the bean pod image repair network new data, thereby improving the robustness of the repair network.
[0079] S2: Using the training set obtained in step S1, the improved soybean seed detection system based on Faster-R-CNN provided by the present invention is trained to obtain a soybean seed detection model. In some embodiments, step S2 uses a recognition loss function to train the soybean recognition network, and a restoration objective function to train the pod image restoration network. During the detection process, the number of bad pods, single pods, and quadruple pods is far less than the number of double and triple pods. As a result, during training, samples with double and triple pods dominate the classification. This causes the network to focus only on the large number of (easy) samples and ignore the small number of (difficult) samples. This results in insufficient training of the small number of samples, and subsequent recognition accuracy of the small number of samples is low, failing to meet the recognition task requirements. To address this issue, the present invention replaces BCE_loss with Focal_loss during the category classification process. Focal_loss is an improvement on the standard BCE loss function. Its core concept is to dynamically adjust sample weights to make the model pay more attention to difficult-to-classify samples. Focal_loss introduces a modulation factor to reduce the loss contribution of easy-to-classify samples, allowing the model to pay more attention to difficult-to-classify samples. This design helps solve the problem of class imbalance and improves the robustness of the model. Specifically:
[0080] The recognition loss function is:
[0081] ;
[0082] Among them, L({p i},{t i}) represents the recognition loss function, p i Indicates the probability that a target exists in the i-th bounding box during target recognition, p i 'Indicates the corresponding real label, t irepresents the compensation between the i-th predicted bounding box and the corresponding true bounding box coordinates, t i 'Indicates the coordinates of the real bounding box where the target is located, N cls Indicates the batch size during training, N reg Indicates the number of identified targets, λ indicates the adjustment weight of the recognition loss function; L cls (p i ,p i ') represents the cross entropy loss function, that is:
[0083] L cls (p i ,p i ')=-p i 'logp i -(1-p i ')log(1-p i );
[0084] L reg (t i ,t i ') represents the Focal_loss loss function, which is represented by smooth L1 The losses are accumulated, that is:
[0085] ;
[0086] ;
[0087] The repair objective function is:
[0088] G'=argminmaxLOSS cGAN (G,D)+λ'LOSS1(G);
[0089] Among them, G' represents the restoration objective function, D represents the discriminator, G represents the generator, LOSS1 represents the L1 loss function, and λ' represents the restoration weight; LOSS cGAN (G,D) represents the repair loss function, which is:
[0090] LOSS cGAN (G,D)=E x,y [logD(x,y)]+E x,z [log(1-D(x,G(x,y)))];
[0091] Among them, x represents the real image, y represents the corresponding constraint, z represents the noise, E x,y [logD(x,y)] represents the discriminator’s expectation of the real image, E x,z[log(1-D(x,G(x,y)))] represents the discriminator’s expectation of the image restored by the generator.
[0092] S3: According to the training results of step S2, the hyperparameters of the soybean seed testing model are adjusted until the optimal soybean seed testing model is obtained. In the embodiment of the present invention, the hyperparameters of the training include: using the adam optimizer, the training size is 640×640, the batch size is 32, the number of categories is 6, the initial learning rate is 1e-2, the annealing cosine training strategy is adopted, the final learning rate is reduced to 1e-6, and the number of iterations is 400. According to the loss rate change curve, when the loss tends to be stable with the increase of the number of iterations, the network converges until 400 iterative trainings are completed to obtain the current soybean seed testing model.
[0093] S4: Input the pod image to be tested into the optimal soybean seed testing model obtained in step S4 to obtain a repaired independent pod image of each pod in the pod image, as well as a stem image; calculate and count the traits of the stem image and all independent pod images, thus completing the soybean seed testing.
[0094] The pod identification result in the embodiment of the present invention is as follows: Figures 10 to 12 As shown, the pod repair effect in the embodiment of the present invention is as follows Figure 13 shown. Figure 10 and Figure 11 The bounding boxes of different colors represent the different types of pods that have been identified. The types of pods identified and their class probabilities are marked above the bounding boxes. Figure 11 Among them, soybeon_1 represents the first type of pod, soybeon_2 represents the second type of pod, soybeon_3 represents the third type of pod, and soybeon_4 represents the fourth type of pod ( Figure 11 not shown), with Figure 11 For example, the label "soybeon_2 0.86" on one of the bounding boxes indicates that the pod in the bounding box is identified as a Class 2 pod, and the probability that it is a Class 1 pod is 0.86, or 86%. Figure 10 and Figure 12 The green box in the image represents the identified stem and its probability of being a stem. Figure 12 For example, if one of the bounding boxes is labeled "stem 0.69", it means that the content in the bounding box is a stem, and the probability that it is a stem is 0.69, or 69%. Figure 13 The middle left column shows different degrees of occlusion (no occlusion with slight interference, side occlusion, and side-by-side occlusion of multiple pods). Figure 13 The right column shows the corresponding repaired results. Figure 13 It can be seen from the figure that the present invention can effectively restore pods with different degrees of occlusion, laying a good foundation for subsequent pod trait statistics.
[0095] In some embodiments, the pod traits in step S4 include pod length, pod width, pod circumference and pod projected area; the extraction process of pod traits is as follows: Figure 14 As shown, it includes: graying the restored pod independent image, extracting the grayscale image of the R channel with the largest contrast between the pod grains and the background for subsequent shape extraction, as shown in FIG. Figure 14 (a) in the figure; use the threshold segmentation algorithm to binarize the R channel image and obtain the pod binary image, as shown in Figure 14 (b) in the figure; extract the outer contour of the pod based on the binary image of the pod, including the coordinates of each point on the contour; calculate the length, width, perimeter and projection area of the pod based on the obtained contour coordinates, as shown in Figure 14 (c) in the.
[0096] In an embodiment of the present invention, the number of pods contained in different tags can also be read from the number of beans in the entire plant, multiplied by the number of pods determined by the tags, and then added together to obtain the number of beans in the entire plant. For example, a bad pod is defined as P0, which contains no pods, that is, the number of pods corresponding to the bad pod is 0; a first-class pod is defined as P1, which contains 1 pod, that is, the number of pods corresponding to the first-class pod is 1; a second-class pod is defined as P2, which contains 2 pods, that is, the number of pods corresponding to the second-class pod is 2; a third-class pod is defined as P3, which contains 3 pods, that is, the number of pods corresponding to the third-class pod is 2; a fourth-class pod is defined as P4, which contains 3 pods, that is, the number of pods corresponding to the fourth-class pod is 2. In this case, the number of beans in the entire plant = 0×P0+1×P2+2×P2+3×P3+4×P4.
[0097] In an embodiment of the present invention, the process of calculating the length of a pod includes, in the process of obtaining the length of the pod, arbitrarily selecting two points on the outer contour of the pod, and calculating the distance between the two points, and the maximum value of all distances is the length of the pod. In the process of obtaining the width of the pod, the normal slope of the straight line between the two points of the maximum distance is calculated, and then the distance between the intersection of the normal with the same normal slope and the outer contour of the pod under different intercepts is calculated, and the maximum value of all distances is the width of the pod. In the process of obtaining the perimeter of the pod, the sum of the distances between all adjacent pixel points on the outer contour of the pod is the perimeter of the pod. In an embodiment of the present invention, based on the points on the contour of the pod, if the two points are adjacent in the upper and lower or left and right directions, the distance between the two points is defined as 1, and if the two points are adjacent in the upper left, lower left, upper right or lower right directions, the distance between the two points is defined as , the sum of the distances between all adjacent pixels on the contour is the pod perimeter. The pod projected area is the number of pixels within the pod's outer contour.
[0098] In an embodiment of the present invention, the stem traits include the number of main stem nodes, the maximum internode distance, the minimum internode distance, the average internode distance, the total length of the stem nodes, the farthest internode distance, and the curvature. Specifically, the target frame parameters corresponding to the label value are read according to the return value obtained by identification, including the coordinates (x1, y1, x2, y2) of the original image corresponding to the upper left point and the lower right point of the target frame, and the length and width of the stem are obtained as x2-x1 and y2-y1 respectively, and the coordinate position of the original image where the center point of the stem is located is ((x2-x1) / 2+x1, (y2-y1) / 2+y1). According to the obtained center coordinates of the first and last stems, the approximate length of the main stem can be obtained; according to the center coordinates of the adjacent stems, the lengths of several main trunks can be obtained. Count the distances between adjacent stems (X1, X2, ..., X n ), and obtain the maximum value X by comparison MAX , X MAX is the maximum internode distance. Similarly, the minimum value X is obtained by comparison. MIN , X MIN The minimum internode distance is obtained by averaging the distances between adjacent stems (X1, X2, ..., X n ) the average value X MEAN , X MEAN is the average internode distance, and the total length of the stem node X is obtained by accumulation. SUM =X1+X2+…+X n ; Get the center coordinates of the first stem node (X1, Y1) and the center coordinates of the last stem node (X n ,Y n ), through the formula The ratio of the farthest internode distance to the total length of the stem node is the curvature, which is used to calculate the degree of curvature of the stem.
[0099] The statistical data fitting effects of the number of bad pods, first-class pods, second-class pods, third-class pods and fourth-class pods in the embodiment of the present invention correspond to Figures 15 to 19 The statistical data fitting effects of the number of all pods and the number of all beans correspond to Figure 20 and Figure 21 . Figures 15 to 21 The vertical axis represents the number of corresponding pods counted using the method provided by the present invention, and the horizontal axis represents the number of corresponding pods counted manually. Figures 15 to 21 It can be seen that the number of corresponding pods counted by the method provided by the present invention is close to the corresponding number counted manually, indicating that the method provided by the present invention can effectively complete the statistics of the number of pods or stem nodes. 2 The R-squared indicator reflects the accuracy of quantitative statistics, where R 2Also known as the coefficient of determination or determination coefficient, it is an indicator used to measure the goodness of fit of the regression model to the observed data. It indicates the degree to which the independent variable in the model explains the dependent variable, and its value range is between 0 and 1. 2 The closer it is to 1, the better the model fits the data; the closer it is to 0, the weaker the model's ability to explain the data. The R values of the number of bad pods, first-class pods, second-class pods, third-class pods, and fourth-class pods, as well as the number of all pods and all beans counted using the method provided by the present invention are as follows: 2 The indicators are 0.9705, 0.9651, 0.9868, 0.9275, 0.7703, 0.981 and 0.9806 respectively. It can be seen that R 2 The indices are all close to 1, which further illustrates that the fitting effect of the method provided by the present invention is very good, that is, there is a highly significant linear relationship between the method provided by the present invention and the manual recognition method.
[0100] In addition, the embodiment of the present invention uses five indicators, namely precision P (Precision), recall R (Recall), mAP_50, mAP_75, and mAP50-95, to evaluate the target detection performance of the soybean recognition network. Among them, precision P refers to the proportion of samples predicted as positive that are actually positive, which measures the accuracy of the prediction results; recall R is the proportion of true positive samples that are correctly predicted as positive, reflecting the model's ability to capture positive samples; mAP_50 represents the average precision when the IoU threshold is 0.5, reflecting the detection performance of the model on different categories; mAP_75 represents the average precision when the IoU threshold is 0.75, which has higher detection accuracy requirements than mAP_50; mAP50-95 represents the average precision in the IoU threshold range of 0.5 to 0.95, which comprehensively measures the performance of the model at different IoU thresholds. The indicator evaluation results are shown in Table 1:
[0101] Table 1:
[0102]
[0103] Table 1 shows that while comprehensive performance indicators are provided for all categories, specific indicators vary across categories. For example, precision and recall are higher for pods from Class 1 to Class 4, indicating good detection performance. Overall indicators for soybeans are relatively low, but some indicators, such as mAP_50, perform well for stems.
[0104] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved. This is not limited herein.
[0105] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A soybean seed detection system based on Faster-R-CNN, characterized in that: It includes pod recognition network and pod image restoration network; among them, The pod recognition network includes a trunk module, a neck module, and a detection module; the pod image is input into the trunk module for multiple residual feature extractions to obtain residual features of multiple scales; the residual features of multiple scales are input into the neck module for multi-scale pyramid feature transformation to obtain transformation features of multiple scales; the transformation features of multiple scales are input into the detection module for target recognition to obtain a stem image in the pod image and an independent image of each pod in the pod image; in the detection module, the transformation features of multiple scales are integrated, and the integrated features are subjected to RPN operation to obtain candidate pod targets and stem targets; the obtained pod targets and stem targets are combined with the integrated features for ROI alignment, and the aligned content is subjected to full connection operation to obtain the stem image, the independent pod image and their respective categories; The backbone module adopts the ResNet101 structure, including a cascaded shallow feature extraction unit, a first residual unit, a second residual unit, a third residual unit, and a fourth residual unit; in each residual unit, multiple residual operations are performed on the input features; in each residual operation, multiple consecutive convolution operations are performed on the input features, and the features obtained by the last convolution operation are subjected to a cascaded attention operation; the output features of the cascaded attention operation are added to the corresponding elements of the input features to obtain the output features of the current residual operation; In the cascade attention operation, a channel attention operation is first performed on the input feature, and then a spatial attention operation is performed on the output feature of the channel attention operation to obtain the output feature; In the channel attention operation, the maximum pooling and average pooling in the spatial direction are performed on the input features in parallel, and then the two pooled features are respectively subjected to convolution operations to complete feature extraction; after adding the corresponding elements of the two extracted features, the added features are subjected to a sigmoid activation operation, and the activated features are multiplied with the corresponding elements of the input features to obtain channel features; In the spatial attention operation, the maximum pooling and average pooling in the channel direction are performed in parallel on the channel features. After the two pooled features are spliced in the channel, the spliced features are sequentially subjected to convolution and sigmoid activation operations. The activated features are multiplied by the corresponding elements of the channel features to obtain the output features. The pod image restoration network restores the independent pod image; the pod image restoration network includes a discriminator and a generator, and the generator is a U-Net structure; in the discriminator, the independent pod image restored by the generator and the corresponding real independent pod image are subjected to multiple continuous convolution operations to complete feature extraction; the extracted features are subjected to subsequent image block operations in PatchGAN to complete feature comparison, and an authenticity score is given to evaluate the authenticity of the independent pod image, and the authenticity score is used to guide the generator to complete the restoration of the independent pod image.
2. The soybean seed detection system improved based on Faster-R-CNN according to claim 1, characterized in that: In the shallow feature extraction unit, shallow feature extraction is performed on the input image through a convolution operation, and the convolved features are pooled; the output features of the shallow feature extraction unit are sequentially subjected to feature extraction by the first residual unit, the second residual unit, the third residual unit and the fourth residual unit, and the output features of the second residual unit, the third residual unit and the fourth residual unit enter the neck module.
3. The soybean seed testing system based on Faster-R-CNN according to claim 1, characterized in that: In the neck module: After performing an upsampling operation on the output feature of the fourth residual unit, channel-wise splicing is performed with the output feature of the third residual unit to obtain a spliced feature; After performing an upsampling operation on the spliced features, the spliced features are channel-concatenated with the output features of the second residual unit to obtain a first transformed feature; After performing a convolution operation on the first transformed feature to compress the feature size, the first transformed feature is then channel-joined with the spliced feature to obtain a second transformed feature; Performing channel concatenation on the second transformed feature and the output feature of the fourth residual unit to obtain a third transformed feature; The three transformation features are input into the detection module for target recognition.
4. A soybean seed testing method based on Faster-R-CNN, characterized in that: include: S1: Obtain a pod image dataset and preprocess the pod image dataset to obtain a training set containing pod occlusion conditions; S2: Using the training set obtained in step S1, training the soybean seed detection system improved based on Faster-R-CNN as described in any one of claims 1 to 3 to obtain a soybean seed detection model; S3: According to the training result of step S2, adjusting the hyperparameters when training the soybean seed selection model until an optimal soybean seed selection model is obtained; S4: Input the pod image to be tested into the optimal soybean seed testing model obtained in step S4 to obtain a restored independent pod image of each pod in the pod image, as well as a stem image; calculate and count the traits of the stem image and all independent pod images, thereby completing the soybean seed testing.
5. The soybean seed testing method based on Faster-R-CNN improvement according to claim 4, characterized in that: Step S1 includes: Mature soybean plants are collected in different scenes, the pods on the soybean plants are randomly spread above a background plate, the stems of the soybean plants are placed below the background plate, and photographed using a camera to obtain pod images; removing redundant parts from the bean pod image, cropping the bean pod image into an image of uniform pixel size, and annotating the bean pods and stems in the cropped image to obtain an annotation file; Image enhancement is performed on the cropped images and the annotation files to obtain a training set, which includes the cropped pod images and the corresponding annotation files.
6. The soybean seed testing method based on Faster-R-CNN improvement according to claim 4, characterized in that: In step S2, the pod recognition network is trained using a recognition loss function, and the pod image restoration network is trained using a restoration objective function; wherein the recognition loss function is: ; Among them, L({p i },{t i }) represents the recognition loss function, p i Indicates the probability that a target exists in the i-th bounding box during target recognition, p i 'Indicates the corresponding real label, t i represents the compensation between the i-th predicted bounding box and the corresponding true bounding box coordinates, t i 'Indicates the coordinates of the real bounding box where the target is located, N cls Indicates the batch size during training, N reg represents the number of identified targets, λ represents the adjustment weight for adjusting the recognition loss function; L cls (p i ,p i ') represents the cross entropy loss function, that is: L cls (p i ,p i ’)=-p i ’logp i -(1-p i ’)log(1-p i ); L reg (t i ,t i ') represents the Focal_loss loss function, which is represented by smooth L1 The losses are accumulated, that is: ; ; The repair objective function is: <h2 style=";text-align:left;direction:ltr">G'=argminmaxLOSS<h2 style=";text-align:left;direction:ltr"> cGAN <h2 style=";text-align:left;direction:ltr"> (G,D)+λ'LOSS1(G); Wherein, G' represents the restoration objective function, D represents the discriminator, G represents the generator, LOSS1 represents the L1 loss function, and λ' represents the restoration weight; LOSS cGAN (G,D) represents the repair loss function, which is: LOSS cGAN (G,D)=E x,y [logD(x,y)]+E x,z [log(1-D(x,G(x,y)))]; Among them, x represents the real image, y represents the corresponding constraint, z represents the noise, E x,y [logD(x,y)] represents the discriminator’s expectation of the real image, E x,z [log(1-D(x,G(x,y)))] represents the discriminator’s expectation of the image restored by the generator.
7. The soybean seed testing method based on Faster-R-CNN improvement according to claim 4, characterized in that: In step S4, the pod traits include pod length, pod width, pod circumference and pod projected area; the pod trait extraction process includes: graying the repaired pod independent image, extracting the R channel image with the maximum contrast between the pod grains and the background; binarizing the R channel image using a threshold segmentation algorithm to obtain a pod binary image; extracting the pod outer contour based on the pod binary image, including the coordinates of each point on the contour; and calculating the pod length, width, circumference and pod projected area based on the obtained contour coordinates.
Citation Information
Patent Citations
Soybean pod quantity statistical method based on machine vision
CN114724141A
Soybean pod test method, system and device based on deep learning
CN116434066A