Tomato maturity detection method based on edge space feature fusion and feature focusing
By introducing a multi-scale focused feature pyramid network and feature extraction module FPSC in the YOLOv8 model, combined with C2f-SCM and MCBC modules, the problem of low detection accuracy of tomatoes in complex environments in the prior art is solved, and higher detection accuracy and lower error detection rate are achieved.
Patent Information
- Application Number
- CN202510050008.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-30
AI Technical Summary
The existing tomato ripening degree detection method has poor detection accuracy in complex environments, insufficient generalization ability, and insufficient extraction of tomato image features at different growth stages, resulting in missed detection and missed detection problems.
The multi-scale focused feature pyramid network (FCPN) based on YOLOv8 and the feature extraction module FPSC are adopted, combined with the C2f-SCM and MCBC modules, to enhance the extraction and fusion capabilities of image features, improve the feature fusion layer and feature extraction module to improve detection accuracy.
It effectively improves the detection accuracy of obscured tomatoes, enhances the extraction ability of image features, improves the detection accuracy of tomato ripening in complex backgrounds, and reduces the rate of false detection and missed detection.
Smart Images

Figure CN120071332A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision target detection, and particularly to a tomato maturity detection method based on edge space feature fusion and feature focusing. Background Art
[0002] Tomato is a common fruit and vegetable, which is deeply favored by people because it can promote digestion, beautify the skin and help detoxify. Due to the serious aging of the social population and the increasing shortage of agricultural labor, the traditional manual picking method is difficult to meet the production needs. Therefore, it has become an urgent task to realize the mechanization and intelligentization of the tomato industry. Target detection can provide target information for mechanized picking, guiding the picking robot to move precisely, so as to realize the mechanized picking of tomatoes. The image preprocessing, tomato recognition and maturity detection methods provide an effective solution for realizing the intelligent picking of tomatoes.
[0003] In recent years, artificial intelligence has developed faster and faster, the popularity of computer vision has continued to grow, and the fruit detection applications based on convolutional neural networks have become more and more widespread. Based on the convolutional neural network framework, the features of the target fruit are trained, and the trained network model can basically achieve the accurate detection of the fruit image. At the same time, people introduce computer vision technology into the design and development of tomato picking robots, enabling the machine module to have strong learning ability and improving the error-free rate of fruit detection. Compared with traditional picking equipment, its advantages are obvious. The common CNN-based target detection algorithms can be divided into two categories. The first category is the two-stage algorithm based on region proposal generation, such as R-CNN, Faster R-CNN, etc. The second category is the one-stage algorithm, such as YOLO, SSD, etc.
[0004] However, the accuracy of the existing detection methods for tomato maturity detection in complex environmental backgrounds is not ideal. For tomatoes in different planting environments, their appearance features may vary, and the growth environments are also diverse. Therefore, when the detection model faces a new environment, its generalization ability is insufficient, which will lead to a significant decrease in detection accuracy. In addition, the existing algorithms do not extract the image features of tomatoes at different growth stages sufficiently, and cannot distinguish the features of tomatoes with different maturities well, resulting in misdetection. At the same time, there is also the problem of fruit occlusion for tomatoes. Most algorithms tend to cause misdetection and missed detection when facing the situation of fruit occlusion. Therefore, it is essential and challenging to improve the detection accuracy of occluded tomatoes. Therefore, improving the detection accuracy of fruit occlusion, enhancing the ability to extract image features, and improving the detection accuracy of tomato maturity in complex backgrounds cannot be ignored. Summary of the Invention
[0005] Objective of the Invention: The objective of the present invention is to provide a tomato maturity detection method based on edge space feature fusion and feature focusing. Aiming at the situation of low tomato maturity detection caused by misdetection and missed detection, the present invention can effectively improve the detection accuracy of occluded tomatoes, enhance the ability to extract image features, and the detection accuracy of tomatoes under complex backgrounds.
[0006] Technical Solution: The present invention provides a tomato maturity detection method based on edge space feature fusion and feature focusing, which includes the following steps:
[0007] S1: Collect tomato images of greenhouse cultivation to form a tomato dataset, and use the Make Sense online annotation tool to annotate the collected images with different maturities. The specific categories include three types: mature, semi-mature, and immature.
[0008] S2: Divide the dataset obtained in S1 into a training set, a validation set, and a test set, and then perform data augmentation on the training set, validation set, and test set respectively.
[0009] S3: Adopt YOLOv8 as the basic model, construct an FCPN (multi-scale focused feature pyramid network) to replace the feature fusion layer of the original YOLOv8, design a C2f-SCM module to replace C2f in the backbone network, construct an MCBC module, and construct a feature extraction module FPSC to replace the original SPPF module to obtain an improved YOLOv8 model.
[0010] S4: Use the labeled dataset obtained in S2 to train the improved YOLOv8 model. Set the training parameters during each training and conduct evaluation. The improved YOLOv8 model identifies tomatoes with different maturities from the input images, and trains the model using the loss function. After convergence, obtain the optimal model weights.
[0011] S5: Use the optimal model weights trained in S4 and input the test set for detection to detect tomatoes with different maturities from the tomato images to be detected.
[0012] Furthermore, the specific steps of S1 are as follows:
[0013] S11: Collect tomato images of greenhouse cultivation to form a tomato dataset, divide it into a training set, a validation set, and a test set, and use Mosaic data augmentation to scale, rotate, translate, and flip the input images to increase data diversity.
[0014] S12: Use the online annotation tool Make Sense to annotate the tomato dataset. For tomatoes with a dark red skin color that meet the picking conditions, label them as "mature"; for those with a lighter orange - red skin color that do not yet meet the picking conditions, label them as "semi - mature"; and for those with a green skin color that are still in the growth stage, label them as "immature".
[0015] S13: Resize the input image to the standard size of 640 * 640 pixels, and then send it into the improved YOLOv8 model.
[0016] Furthermore, the specific steps of S3 are as follows:
[0017] S31: In the feature extraction stage, design the C2f - SCM module. Replace the second - layer and fourth - layer C2f modules in the original model with the C2f - SCM module to enhance the ability to capture global context information in the image and improve the detection accuracy. Then construct the MCBC module to replace the C2f modules in the sixth and eighth layers of the backbone network. At the same time, construct the feature extraction module FPSC to replace the original SPPF module to capture more fine - grained features in the image and better capture local features and context information in the image.
[0018] S32: In the feature fusion stage, construct the FCPN (Multi - Scale Focused Feature Pyramid Network) to replace the feature fusion layer of the original YOLOv8. This pyramid network can make each scale of features have detailed context information and can capture rich information across multiple scales, which is beneficial for subsequent object detection.
[0019] Furthermore, the backbone network described in S31 is mainly responsible for feature extraction, including the CBS module, C2f - SCM module, MCBC module, and the feature extraction module FPSC.
[0020] The CBS module is a module composed of a two - dimensional convolutional layer, a batch normalization Bn layer, and a SiLU activation function, which is used to obtain image features. One CBS module is regarded as a standard convolutional module in the YOLOv8 model.
[0021] The C2f - SCM module is a new module constructed by combining the SobelConv branch for extracting edge information and the convolutional branch for extracting spatial information. This module can extract richer image feature representations.
[0022] The MCBC module is a new module constructed by integrating the MHDA module and the Bottleneck module. The MHDA module is a new module that integrates the multi-head attention mechanism and double convolution. The MCBC module can effectively extract the global features of the image and enhance the non-linear feature expression ability of the image;
[0023] The FPSC module is a constructed feature extraction module. This module extracts features of different scales by using convolutional layers with different dilation rates. Low dilation rates capture local details, and high dilation rates capture global context information. Compared with the original SPPF module, this module can better capture the detailed features in the image.
[0024] Furthermore, the neck network described in S32 is mainly used for feature fusion, including the CBS module, C2f module, upsampling Upsample module, Concat module, and FEFS module;
[0025] The role of the C2f module is to achieve feature fusion by splicing the features of different branches in the channel dimension. The spliced features will contain information from different branches, enriching the feature expression ability and helping to improve the detection accuracy;
[0026] The role of the upsampling Upsample module is to enlarge the size of the feature map while keeping the number of channels of the feature map unchanged, so that feature maps with different scales but the same number of channels can be fused;
[0027] The role of the Concat module is to increase the number of channels of the feature map while keeping the size of the feature map unchanged, so as to fuse the semantic information of the deep feature map and the detail information of the shallow feature map;
[0028] The FEFS module can accept inputs of three scales and uses a set of parallel depth convolutions with different receptive field sizes to capture rich cross-multi-scale information, so as to improve the subsequent object detection accuracy.
[0029] Furthermore, the training parameters described in S4 are specifically: the input image size imgsz = 640, the initial learning rate lr = 0.01, the learning rate momentum = 0.937, the weight decay coefficient weight_decay = 0.0005, the number of training iterations epoch = 400, the number of samples in the batch training dataset batchsize = 16, and the official pre-trained weights are used for transfer learning and fine-tuning.
[0030] Furthermore, the evaluation metrics described in S4 are mainly: mean average precision (mAP), precision (P), and recall (R). Among them, mAP represents the comprehensive weighted average of the average precision (AP) of all category detections. P represents the ratio of the number of correctly predicted positive samples to the actual number of positive samples, and R represents the ratio of the number of correctly predicted positive samples to the total number of predicted samples. The specific formulas are as follows: Where APi represents the average precision of the i-th category, K represents K categories, TP represents true positives, that is, positive samples predicted as positive by the model, FP represents false positives, that is, negative samples predicted as positive by the model, and FN represents false negatives, that is, positive samples predicted as negative by the model.
[0031] Compared with the existing technology, the beneficial effects that the present invention can bring are as follows:
[0032] (1) Aiming at the problem of low detection accuracy of the original model due to fruit overlap and occlusion, the present invention constructs an FCPN (multi-scale focused feature pyramid network) to replace the feature fusion layer of the original YOLOv8. This pyramid network enables each scale of features to have detailed context information and can capture rich information across multiple scales, thereby effectively improving the detection ability of occluded fruits.
[0033] (2) Aiming at the situation where the original model extracts image features insufficiently during the detection process, the present invention designs a C2f-SCM module, replaces the ordinary C2f modules in the second and fourth layers of the original model with the C2f-SCM module, and enhances the capture of feature representations in the image. At the same time, an MCBC module is constructed to replace the C2f modules in the sixth and eighth layers. The MCBC module can effectively extract the global features of the image and effectively improve the sufficiency of image feature extraction.
[0034] (3) Aiming at the problems of false detection and missed detection caused by the complex background of the image during the detection process of the original model, the present invention constructs a feature extraction module FPSC. The feature extraction module FPSC can capture more fine-grained features, better capture local details and global context information in the image, thereby improving the detection accuracy of tomato maturity under complex backgrounds and reducing the false detection rate and missed detection rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flowchart of the present invention;
[0036] Figure 2 is a network structure diagram of RSF-YOLOv8;
[0037] Figure 3 is a schematic structural diagram of the SCM module;
[0038] Figure 4 It is a schematic structural diagram of the MCBC module;
[0039] Figure 5 It is a schematic structural diagram of the feature extraction module FPSC;
[0040] Figure 6 It is a schematic diagram of FCPN (multi-scale focused feature pyramid network);
[0041] Figure 7 It is a schematic structural diagram of the FEFS module. Specific implementation manner
[0042] To better understand the present invention, the following further describes a tomato maturity detection method based on edge space feature fusion and feature focusing in the present invention in combination with the drawings in the examples of the present invention.
[0043] From Figure 1 it can be seen that the specific steps of the present invention are as follows:
[0044] Step 1: Obtain the data set and preprocess the data set.
[0045] In the present invention, the data set selected is a publicly available data set on the network, which includes 4,500 training set images, 550 validation sets, and 100 test sets, including three categories of mature, semi-mature, and immature. Data augmentation is performed on the data set, and the input images are randomly scaled, rotated, translated, and flipped to increase data diversity;
[0046] In YOLOv8, a method called letterbox is used to preprocess input images of different sizes. This method can scale the image to a specified size (scaling the length and width proportionally), and then add black borders on both sides of the image to make its size consistent with the size to be adjusted. The specific steps are as follows: Calculate the ratio of the length and width of the original image to the specified size, and select the smaller ratio as the scaling ratio. Multiply the length and width of the original image by the scaling ratio respectively to obtain the scaled size. Calculate the difference between the specified size and the scaled size, and divide the difference by 2 to obtain the width of the black border to be filled on both sides. Use the resize function of the OpenCV library to scale the original image, and then use the copyMakeBorder function to add black borders on both sides of the scaled image. In this way, input images of different sizes can be preprocessed into images of the same size for input into the YOLOv8 model for training or inference.
[0047] Step 2: Construct a network model for the tomato maturity detection method based on the improved YOLOv8, as Figure 2 shown.
[0048] The YOLOv8 network model is mainly divided into the input end Input, the backbone network Backbone, the neck network Neck, and the head network Head.
[0049] (1) Improve feature extraction, design the C2f-SCM module, and replace the C2f modules in the second layer and the fourth layer of the original model with the C2f-SCM module, so that the extracted image feature representation contains both rich edge information and spatial information, improving the detection accuracy; then construct the MCBC module to replace the C2f modules in the sixth layer and the eighth layer of the backbone network; at the same time, construct the feature extraction module FPSC to replace the original SPPF module to capture more fine-grained features and better capture local features and context information in the image;
[0050] The backbone network of RSF-YOLOv8 is mainly responsible for feature extraction, including the CBS module, the C2f-SCM module, the MCBC module, and the feature extraction module FPSC.
[0051] The CBS module is a module composed of a two-dimensional convolutional layer, a batch normalization Bn layer, and a SiLU activation function, which is used to obtain image features. A CBS module is regarded as a standard convolutional module in the YOLOv8 model;
[0052] The C2f-SCM is a new module obtained by replacing the first convolutional module in the original C2f module with the SCM module. The structure diagram of the SCM module is as Figure 3 shown. This module is a new module constructed by combining the SobelConv branch for extracting edge information and the dilated convolution branch for extracting spatial information. First, the input image passes through a SobelConv convolutional layer, which uses two 3×3 convolutional kernels, one for detecting horizontal edges and the other for detecting vertical edges, and can highlight edge information; at the same time, the input image also passes through a dilated convolutional layer to extract the spatial context information of the image; then the output feature maps of the two convolutions are concatenated, and the concatenated feature map passes through another convolutional layer to adjust the number of channels of the feature map; then the output of the convolutional layer is connected with the input image through a residual connection, which can retain the information of the original image and alleviate the gradient disappearance problem in the deep network; finally, the output of the residual connection is adjusted by a convolution for the number of channels of the output feature map.
[0053] The MCBC module is as Figure 4As shown, this module is a new module constructed by integrating the MHDA module and the Bottleneck module. The input image passes through an MHDA layer and a Bottleneck layer, and then the two output feature maps are concatenated to merge the features of different processing paths. Then, the concatenated feature map passes through a Conv layer to adjust the number of channels. The structure of the MHDA module is as Figure 4 shown. This module is a new module that integrates the multi-head attention mechanism and double convolution. This module receives the input feature map. First, the feature map passes through a double convolution layer to extract the features of the feature map. The feature map after double convolution then passes through a multi-head self-attention layer to capture more complex feature relationships. Then, the output is connected to the original input feature map through a residual connection. The output of the residual connection passes through another double convolution layer to further extract features. Finally, the feature map passes through a Conv layer to adjust the number of channels. The MCBC module can effectively extract the global features of the image and enhance the non-linear feature expression ability of the image;
[0054] The FPSC module described above is as Figure 5 shown. This module is a constructed feature extraction module. This module extracts features of different scales by using convolution layers with different dilation rates. Low dilation rates capture local details, and high dilation rates capture global context information. The input image first passes through a 1×1 convolution layer, which is usually used to adjust the number of channels without changing the spatial dimensions of the feature map. Then, it passes through a 3×3 convolution layer with a dilation rate of 1, which extracts local features in the feature map. Then, it passes through another 3×3 convolution layer, but the dilation rate of this convolution layer is 3, which reduces the spatial dimensions of the feature map and further extracts features. Next, it passes through a third 3×3 convolution layer with a dilation rate of 5, which further reduces the spatial dimensions of the feature map and extracts deeper features of the image. The feature maps after convolution with different strides are concatenated with the original input image, and finally pass through a 1×1 convolution layer to integrate the concatenated features and adjust the number of channels to match the output requirements. Compared with the original SPPF module, this module can better capture the detailed information in the image.
[0055] (2) Improve feature fusion by constructing an FCPN (Multi-scale Focused Feature Pyramid Network) to replace the feature fusion layer of the original YOLOv8. This pyramid network can make the features of each scale have detailed context information and can capture rich information across multiple scales, which is beneficial to subsequent object detection. As Figure 6As shown in the figure, the network receives three feature maps of different scales as inputs, denoted as P3, P4, and P5 respectively. The P5 feature map first passes through the FEFS module to improve the quality of the feature map and perform feature enhancement, and then passes through a 3×3 convolutional layer for further feature extraction and channel reduction; then it is fused with the P5 feature map, and the fused feature map passes through a C2f module for further feature extraction and enhancement. The P4 and P5 feature maps are upsampled after passing through the FEFS module to match the resolution of the P3 feature map. The upsampled feature maps are fused with the P3 feature map again, and the fused feature map passes through the C2f module again. Finally, all processed feature maps are fed into the detection head Detect, which is responsible for generating the final detection results.
[0056] The neck network of RMC-YOLOv8 is mainly used for feature fusion, including the CBS module, C2f module, upsampling Upsample module, Concat module, and FEFS module;
[0057] The role of the C2f module is to achieve feature fusion by splicing the features of different branches in the channel dimension. The spliced features will contain information from different branches, enriching the feature expression ability and helping to improve the detection accuracy;
[0058] The role of the upsampling Upsample module is to enlarge the size of the feature map while keeping the number of channels of the feature map unchanged, so that feature maps of different scales but the same number of channels can be fused;
[0059] The role of the Concat module is to increase the number of channels of the feature map while keeping the size of the feature map unchanged, so as to fuse the semantic information of the deep feature map and the detail information of the shallow feature map;
[0060] The structure of the FEFS module is as Figure 7 shown. This module can accept inputs of three scales and uses a set of parallel depth convolutions with different receptive field sizes to capture rich cross-multi-scale information to improve the subsequent object detection accuracy. The module first receives feature maps of different scales, and each input feature map passes through a convolutional layer to adjust the number of channels. The feature maps after convolutional processing are spliced together, and the spliced feature map is fed into a depthwise separable convolutional layer with four different kernel sizes. The outputs of the four depthwise separable convolutional layers are fused, and the fused feature map passes through a standard convolution to adjust the number of channels of the feature map; finally, the output of the convolutional layer is connected to the initial spliced feature map with a residual connection.
[0061] Step 3: Set the training parameters, train the model, and evaluate and compare the training results. Set the input image size as imgsz = 640, the initial learning rate lr = 0.01, the learning rate momentum = 0.937, the weight decay coefficient weight_decay = 0.0005, the number of training iterations epoch = 400, and the number of samples in the batch training dataset batchsize = 16. In the model training of the present invention, the official pre-trained weight YOLOv8n.pt is used for transfer learning and fine-tuning because using pre-trained weights can shorten the training cycle, accelerate the network convergence speed, and improve the training effect of the model.
[0062] The evaluation metrics mainly include: mean average precision mAP, precision P (Precision), and recall R (Recall). Among them, mAP represents the comprehensive weighted average of the average precision (AP) of all category detections, P represents the proportion of the number of correctly predicted positive samples to the actual number of positive samples, and R represents the proportion of the number of correctly predicted positive samples to the total number of predicted samples. The specific formulas are as shown in formulas (4) - (6):
[0063] (4)
[0064] (5)
[0065] (6)
[0066] Among them, APi represents the average precision of the i-th category, K represents K categories, TP represents true positives, that is, positive samples predicted as positive classes by the model, FP represents false positives, that is, negative samples predicted as positive classes by the model, and FN represents false negatives, that is, positive samples predicted as false classes by the model.
[0067] The model training evaluation results and ablation experiment results are shown in Table 1.
[0068] Table 1
[0069]
[0070] Among them, YOLOv8 represents the original model; +FCPN represents replacing the feature fusion layer; +C2f-SCM represents adding C2f-SCM to replace the C2f module; +MCBC represents adding the MCBC module; +FPSC represents adding FPSC to replace the SPPF module.
[0071] As can be seen from Table 1, the RSF-YOLOv8 model has significantly improved values in terms of precision P (Precision), recall R (Recall), and mean average precision mAP compared to the original YOLOv8 model.
[0072] The above are only the preferred specific embodiments of the present invention and are not intended to limit the present invention. After reading the content of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent transformations and modifications also fall within the scope defined by the claims of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for detecting tomato maturity based on edge space feature fusion and feature focusing, characterized in that: The following steps are involved: S1: Collect greenhouse tomato images to form a tomato dataset. Use the Make Sense online annotation tool to annotate the collected images with different maturity levels. The specific categories include mature, semi-mature, and immature. S2: Divide the data set obtained in S1 into a training set, a validation set, and a test set, and then perform data augmentation on the training set, validation set, and test set respectively; S3: Use YOLOv8 as the basic model, use FCPN to replace the original YOLOv8 feature fusion layer, design the C2f-SCM module to replace the C2f in the backbone network, build the MCBC module, construct the feature extraction module FPSC, replace the original SPPF module, and get the improved YOLOv8 model; S4: Use the improved YOLOv8 model to train the labeled data set obtained in S2. Set the training parameters and perform evaluation each time. The improved YOLOv8 model can identify tomatoes of different maturity from the input images. The loss function is used to train the model. After convergence, the optimal model weights are obtained. S5: Using the optimal model weights obtained by training in S4, the test set input is used for detection, and tomatoes of different maturity levels are detected from the tomato images to be detected.
2. The method for detecting tomato maturity based on edge space feature fusion and feature focusing according to claim 1, characterized in that: The specific process of S1 is as follows: S11: Collect tomato images grown in greenhouses to form a tomato dataset, which is divided into training set, validation set and test set. Use Mosaic data enhancement to scale, rotate, translate and flip the input images to increase data diversity. S12: Use the online annotation tool Make Sense to annotate the tomato dataset. Tomatoes with dark red skin color that meet the picking conditions are annotated as "mature", those with lighter orange-red skin color that do not meet the picking conditions are annotated as "semi-mature", and those with green skin color that are still in the growth stage are annotated as "immature"; S13: Scale the input image to the standard size of 640*640 pixels and then feed it into the improved YOLOv8 model.
3. The method for detecting tomato maturity based on edge space feature fusion and feature focusing according to claim 2, characterized in that: The specific process of S3 is as follows: S31: In the feature extraction stage, the SCM module is designed, and the original C2f module is improved with the SCM module to obtain a new module C2f-SCM, which replaces the C2f modules of the 2nd and 4th layers in the original model, so that the extracted image feature representation contains both rich edge information and spatial information, thereby improving the accuracy of detection; then the MCBC module is constructed to replace the C2f modules of the 6th and 8th layers of the backbone network; at the same time, the feature extraction module FPSC is constructed to replace the original SPPF module to capture more fine-grained features and better capture local features and contextual information in the image; S32: In the feature fusion stage, FCPN is constructed to replace the feature fusion layer of the original YOLOv8. The pyramid network can make the features of each scale have detailed contextual information and can capture rich information across multiple scales, which is conducive to the detection of subsequent targets. FCPN receives feature maps of three different scales as input, which are recorded as P3, P4, and P5 respectively. The P5 feature map first passes through the FEFS module to improve the quality of the feature map and perform feature enhancement. Then, a 3×3 convolution layer is used for further feature extraction and channel reduction. Then, the feature map is fused with the P5 feature map. The fused feature map passes through a C2f module for further feature extraction and enhancement. After passing through the FEFS module, the P4 and P5 feature maps are upsampled to match the resolution of the P3 feature map. The upsampled feature map is fused with the P3 feature map again. The fused feature map passes through the C2f module again. Finally, all processed feature maps are sent to the detection head Detect.
4. The method for detecting tomato maturity based on edge space feature fusion and feature focusing according to claim 3, characterized in that: Step S31: The backbone network is responsible for feature extraction, including a CBS module, a C2f-SCM module, a MCBC module, and a feature extraction module FPSC; The CBS module is a module consisting of a two-dimensional convolution layer, a batch normalization Bn layer and a SiLU activation function, which is used to obtain image features. A CBS module is regarded as a standard convolution module in the YOLOv8 model; The C2f-SCM module is a new module obtained by improving the original C2f module with the SCM module, so that the extracted image feature representation contains both rich edge information and spatial information, thereby improving the accuracy of detection; wherein the SCM module is a new module constructed by combining the SobelConv branch for extracting edge information and the convolution branch for extracting spatial information, and the module can extract important edge information and retain rich spatial details; The MCBC module is a new module constructed by integrating the MHDA module and the BottleNeck module, wherein the MHDA module is a new module integrating the multi-head attention mechanism and the double convolution, and the MCBC module can effectively extract the global features of the image and enhance the nonlinear feature expression ability of the image. The input image passes through an MHDA layer and a BottleNeck layer, and then the two output feature maps are spliced to merge the features of different processing paths; then the spliced feature map passes through a Conv layer to adjust the number of channels, wherein the MHDA module structure is shown in FIG4 , and the module is a new module integrating the multi-head attention mechanism and the double convolution, and the module receives the input feature map, and first the feature map passes through a double convolution layer to extract the features of the feature map; the feature map after the double convolution then passes through a multi-head self-attention layer to capture more complex feature relationships; then the output is residually connected with the original input feature map; the output of the residual connection passes through a double convolution layer again to further extract features; finally, the feature map passes through a Conv layer to adjust the number of channels; The FPSC module is a constructed feature extraction module. The module extracts features of different scales by using convolution layers with different expansion rates. Low expansion rates capture local details, and high expansion rates capture global context information. Compared with the original SPPF module, this module can better capture the detailed features in the image. The input image first passes through a 1×1 convolution layer, which is usually used to adjust the number of channels without changing the spatial dimension of the feature map; then, it passes through a 3×3 convolution layer with a dilation rate of 1, which will extract local features in the feature map; then, it passes through another 3×3 convolution layer, but the dilation rate of this convolution layer is 3, which will reduce the spatial dimension of the feature map and further extract features; next, it passes through a third 3×3 convolution layer with a dilation rate of 5, which will further reduce the spatial dimension of the feature map and extract deeper features of the image. The feature map that has undergone different length convolutions is spliced with the original input image, and finally passes through a 1×1 convolution layer to integrate the spliced features and adjust the number of channels to match the output requirements.
5. The method for detecting tomato maturity based on edge space feature fusion and feature focusing according to claim 3, characterized in that: The neck network described in step S32 is used for feature fusion, including a CBS module, a C2f module, an upsample Upsample module, a Concat module and a FEFS module; The function of the C2f module is to realize feature fusion by splicing the features of different branches in the channel dimension. The spliced features will contain information from different branches, enrich the expression ability of the features, and help improve the accuracy of detection; The function of the upsample module is to enlarge the size of the feature map while keeping the number of channels of the feature map unchanged, so that feature maps of different scales but the same number of channels can be fused; The function of the Concat module is to increase the number of channels of the feature map while ensuring that the size of the feature map remains unchanged, so as to fuse the semantic information of the deep feature map with the detail information of the shallow feature map; The FEFS module can accept inputs of three scales and utilizes a set of parallel depthwise convolutions with different receptive field sizes to capture rich information across multiple scales in order to improve subsequent target detection accuracy.
6. The method for detecting tomato maturity based on edge space feature fusion and feature focusing according to claim 3, characterized in that: The training parameters in S4 are specifically as follows: input image size imgsz=640, initial learning rate lr=0.01, learning rate momentum=0.937, weight decay coefficient weight_decay=0.0005, training iteration number epoch=800, batch training data set sample number batchsize=16, and use official pre-trained weights for transfer learning and fine-tuning.
7. The method for detecting tomato maturity based on edge space feature fusion and feature focusing according to claim 3, characterized in that: The evaluation indicators in S4 are mainly: mean average precision (mAP), precision P (Precision), and recall R (Recall), where mAP represents the comprehensive weighted average of the average precision (AP) of all categories of detection, P represents the ratio of the number of correctly predicted positive samples to the number of actual positive samples, and R represents the ratio of the number of correctly predicted positive samples to the total number of predicted samples. The specific formula is: Where APi represents the average precision of the i-th category, K represents K categories, TP represents true positive examples, that is, positive samples predicted by the model as positive, FP represents false positive examples, that is, negative samples predicted by the model as positive, and FN represents false negative examples, that is, positive samples predicted by the model as false.
Citation Information
Cited By
Visual positioning method for UHPC highway guardrail center hole
CN120726126A
Fruit identifying and positioning method applied to fruit picking robot
CN120877278A
A fruit recognition and positioning method applied to a fruit picking robot
CN120877278B
Intelligent detection method and system for property certificate table based on edge enhancement and large model
CN121191183A