Asparagus grading method and system based on asparagus grading model, asparagus grading model and storage medium
By introducing the SE attention mechanism and Slim-neck structure in the YOLOv8 model, combining the weight loss function of multiple weight factors, the accuracy and efficiency of asparagus grading are improved, and the problems of insufficient accuracy and high computing power overhead in the asparagus grading are solved.
Patent Information
- Application Number
- CN202510437902.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
The existing mainstream target detection models such as YOLOv8 have poor accuracy in asparagus grading and high computing power overhead, resulting in low efficiency in asparagus grading.
The SE attention mechanism module was introduced into the backbone network of the YOLOv8 model, and the Slim-neck structure was introduced into the neck network, combining the weight loss functions of multiple weight factors to optimize the feature extraction and bounding box regression accuracy of the asparagus grading model.
It improves the accuracy and efficiency of asparagus grading, reduces the computational complexity of the model, and realizes the precise grading of post-harvest asparagus.
Smart Images

Figure CN120339840A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular, to an asparagus grading method, an asparagus grading model, a system, and a storage medium based on an asparagus grading model. Background Art
[0002] Asparagus is a perennial plant with a long harvesting cycle, fast growth rate, a relatively concentrated tender stem harvesting period, and a very short shelf life of tender stems. After harvesting, asparagus needs to be graded according to its specifications and then packaged and sold within a short time. However, the grading of traditional asparagus specifications basically relies on manual labor, and it is difficult to grade different specifications of asparagus by sensory evaluation, which requires a large amount of capital and labor.
[0003] At present, some studies have applied computer vision to the asparagus grading scenario to improve grading accuracy and sorting efficiency. However, the mainstream object detection models such as YOLOv8 currently have poor accuracy in asparagus grading and large computing power consumption. Therefore, a new deep learning model architecture is needed to process asparagus images to improve the accuracy and practicality of asparagus grading. Summary of the Invention
[0004] The present invention provides an asparagus grading method based on an asparagus grading model, aiming to solve the problem of how to improve the grading accuracy of asparagus.
[0005] To achieve the above object, an asparagus grading method based on an asparagus grading model provided by this application includes:
[0006] Inputting an asparagus image into a trained asparagus grading model to extract the asparagus stem thickness features of each target asparagus in the asparagus image through the asparagus grading model;
[0007] Determining the grading result of the target asparagus according to the asparagus stem thickness features;
[0008] Among them, the backbone network of the asparagus grading model includes an SE attention mechanism module, and the neck network includes a Slim-neck module;
[0009] The weight loss function of the asparagus grading model is:
[0010] WIoU = IoU - α × d center_dist - β × d aspect_ratio - γ × d area_difference ;
[0011] Among them, WIoU is the weight loss function of the asparagus grading model, IoU is the intersection over union between the predicted box and the ground truth box, d center_dist represents the center point weight, d aspect_ratio represents the aspect ratio weight, d area_differenceDenote the area difference weight, and α, β, γ denote the weight factors.
[0012] Optionally, the Slim-neck module includes a vector convolution-group shuffle spatial pyramid pooling module and two layers of group shuffle convolutional layers. The processing process of the vector convolution-group shuffle spatial pyramid pooling module includes:
[0013] Reduce the dimension of the initial asparagus image;
[0014] Input the dimension-reduced asparagus image into the two layers of the group shuffle convolutional layers for processing, and perform Concat splicing on the first group shuffle convolutional feature map and the second group shuffle convolutional feature map output by the two layers of group shuffle convolutional layers to obtain the asparagus stem diameter feature.
[0015] Optionally, the processing process of the group shuffle convolutional layer includes:
[0016] Process the asparagus image input by the vector convolution-group shuffle spatial pyramid pooling module through a standard convolutional layer to obtain a first output feature map;
[0017] Input the output feature map into the depthwise separable convolutional layer to generate a second output feature map by performing convolutional operations on each channel independently through the depthwise separable convolutional layer;
[0018] Splice the first output feature map and the second output feature map to form a third feature map, rearrange the channel order of the third feature map and then output it to obtain the group shuffle convolutional feature map.
[0019] Optionally, the expression of the SE attention mechanism module includes:
[0020]
[0021] y = ReLU(W1·Z c )
[0022] s = Sigmoid(W2·Z c )
[0023]
[0024] In the formula, Z C is the compressed feature vector; y is the activation function; represents the weighted feature map obtained by multiplying the weight vector s element-wise with each channel of the feature map X; ReLU and Sigmoid are both activation functions; W1 and W2 are both weight matrices of dimension (C / r)×C; H, W, C are the length, width, and number of channels of the feature map, i represents the row index in the spatial dimension of the feature map, and j represents the column index in the spatial dimension of the feature map.
[0025] Optionally, before the step of inputting the asparagus image into the trained asparagus grading model, the following steps are further included:
[0026] Annotate the asparagus image dataset containing different grades, where the size of each annotation box is the minimum bounding rectangle of the asparagus plant;
[0027] Divide the annotated asparagus image dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1;
[0028] Perform flipping, mirroring, and noise addition operations on the divided training set to complete the preprocessing of the training set;
[0029] Train the asparagus grading model based on the preprocessed training set, the validation set, and the test set.
[0030] Optionally, the grading results include thick asparagus, medium asparagus, and thin asparagus.
[0031] In addition, to achieve the above object, the present application further provides an asparagus grading model, and the asparagus grading model includes:
[0032] A backbone network, the backbone network is framed by YOLOv8 and includes a SE attention mechanism module for extracting the stem thickness features of asparagus;
[0033] A neck network, including a Slim-neck module for reducing the computational overhead of the model;
[0034] Among them, the weight loss function of the asparagus grading model is:
[0035] WIoU = IoU - α×d center_dist -β×d aspect_ratio -γ×d area_difference ;
[0036] Among them, WIoU is the weight loss function of the asparagus grading model, IoU is the intersection over union between the predicted box and the ground truth box, d center_dist represents the center point weight, d aspect_ratio represents the aspect ratio weight, d area_difference represents the area difference weight, and α, β, and γ represent weight factors.
[0037] Optionally, the Slim-neck module includes a vector convolution-group shuffle spatial pyramid pooling module and two layers of group shuffle convolutional layers;
[0038] The vector convolution-group shuffle spatial pyramid pooling module is used to reduce the dimension of the initial asparagus image; the reduced-dimension asparagus image is input into two layers of the group shuffle convolutional layers for processing, and the first group shuffle convolutional feature map and the second group shuffle convolutional feature map output by the two layers of group shuffle convolutional layers are concatenated by Concat to obtain the asparagus stem thickness feature;
[0039] The group shuffle convolutional layer is used to process the asparagus image input by the vector convolution-group shuffle spatial pyramid pooling module through a standard convolutional layer to obtain a first output feature map with the number of channels halved; the output feature map is input into a depthwise separable convolutional layer to independently perform a convolutional operation on each channel to generate a second output feature map; the first output feature map and the second output feature map are concatenated to form a third feature map, and the channel order of the third feature map is rearranged and then output to obtain a group shuffle convolutional feature map.
[0040] In addition, to achieve the above object, the present application also provides an asparagus grading system, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the asparagus grading method described in any one of the above are implemented.
[0041] In addition, to achieve the above object, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the asparagus grading method based on the asparagus grading model described in any one of the above are implemented.
[0042] The present application has at least the following beneficial effects:
[0043] 1. Based on the YOLOv8 model, the SE attention mechanism module is introduced into the backbone network and the Slim-neck structure is introduced into the neck network. The effective feature extraction ability is improved by using the SE attention mechanism, and the Slim-neck module reduces the model parameter quantity and complexity;
[0044] 2. A weight loss function containing multiple weight factors is proposed to improve the regression accuracy of the bounding box in the object detection task, realize the precise grading of post-harvest asparagus, and provide technical support for the deployment of the asparagus grading device. Description of the Drawings
[0045] Figure 1 It is the SE attention mechanism structure diagram involved in the embodiment of the present application;
[0046] Figure 2 It is the vector convolution-group shuffle spatial pyramid pooling module structure diagram involved in the embodiment of the present application;
[0047] Figure 3 This is the structural diagram of the grouped shuffle convolutional layer involved in the embodiments of this application;
[0048] Figure 4 This is the schematic flowchart of the asparagus grading method for the asparagus grading model involved in the embodiments of this application;
[0049] Figure 5 This is the schematic diagram of the detection effect involved in the embodiments of this application;
[0050] Figure 6 This is the schematic diagram of the asparagus grading model architecture involved in the embodiments of this application;
[0051] Figure 7 This is the schematic diagram of the architecture of the hardware operating environment of the asparagus grading system involved in the embodiments of this application.
[0052] The realization, functional characteristics and advantages of the purpose of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0053] To better understand the above technical solutions, the exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0054] Embodiment 1
[0055] In this embodiment, an asparagus grading method based on an asparagus grading model is provided.
[0056] Referring to Figure 1 the shown SE attention mechanism structure diagram, in this embodiment, on the basis of the traditional YOLOv8 as the backbone network architecture, an SE (Squeeze-and-Excitation) attention mechanism module is introduced as an improvement.
[0057] It should be noted that the SE attention mechanism module can focus on important feature channels while ignoring unimportant feature channels. After testing, the SE attention mechanism module can strengthen the extraction of the asparagus stem diameter feature in the post-harvest asparagus grading task. This attention mechanism adaptively learns the importance of each channel and assigns corresponding weights to each channel, which can effectively improve the accuracy of the model, reduce the extraction of other unimportant features, and improve the feature extraction ability of the model.
[0058] Further and optionally, the implementation process of the SE attention mechanism is mainly divided into three steps, namely: compression, excitation, and reweighting. The expression of the SE attention mechanism module includes:
[0059]
[0060] y = ReLU(W1·Z c )
[0061] s = Sigmoid(W2·Z c )
[0062]
[0063] In the formula, Z C is the feature vector after compression; y is the activation function; represents the weighted feature map obtained by multiplying the weight vector s and each channel of the feature map X element by element. ReLU and Sigmoid are both activation functions; W1 and W2 are both weight matrices with dimensions of (C / r)×C; H, W, and C are the length, width, and number of channels of the feature map, i represents the row index in the spatial dimension of the feature map, and j represents the column index in the spatial dimension of the feature map.
[0064] On the other hand, the neck network part of the model introduces the Slim-neck module. Slim-neck is an architecture for optimizing the neck network as an improvement.
[0065] Further and optionally, the Slim-neck module includes a Vector Convolution-Group Shuffle Spatial Pyramid Pooling (VoV-GSCSP) module and two layers of group shuffle convolutional layers (GSConv).
[0066] Referring to Figure 2 the structural diagram of the Vector Convolution-Group Shuffle Spatial Pyramid Pooling module shown, the VoV-GSCSP module can fully extract the feature information of the shallow network and the deep network and further fuse it, so that the output feature map has both rich position information and semantic information at the same time. Therefore, the VoV-GSCSP module is used to replace the C2f module in the original YOLOv8 model.
[0067] The processing process of the VoV-GSCSP module includes:
[0068] Dimensionality reduction is performed on the initial asparagus image;
[0069] The dimensionality-reduced asparagus image is input into the two layers of the group shuffle convolutional layers for processing, and the first group shuffle convolutional feature map and the second group shuffle convolutional feature map output by the two layers of group shuffle convolutional layers are Concat spliced to obtain the asparagus stem thickness feature.
[0070] Optionally, a 1×1 classical convolution operation can be used for dimensionality reduction.
[0071] Referring to Figure 3 the shown grouped shuffle convolution layer structure diagram, the GSConv layer reduces the computational overhead of the model by combining group convolution and channel shuffle. The processing process of the GSConv layer includes:
[0072] Processing the asparagus image input by the vector convolution-grouped shuffle spatial pyramid pooling module through a standard convolution layer to obtain a first output feature map with the number of channels being C2 / 2;
[0073] Inputting the output feature map into a depthwise separable convolution layer to perform convolution operations on each channel independently through the depthwise separable convolution layer to generate a second output feature map with the number of channels being C2 / 2;
[0074] Concatenating the first output feature map and the second output feature map to form a third feature map with the number of channels being C2, rearranging the channel order of the third feature map and then outputting to obtain the grouped shuffle convolution feature map.
[0075] After constructing the asparagus grading model that meets the above architecture, training is carried out. During the training process, the weighted intersection over union (WIoU) weight loss function in this embodiment is selected to optimize the regression accuracy of the bounding box in the object detection task.
[0076] Specifically, the weight loss function is set with center distance weighting, aspect ratio weighting, and area difference weighting.
[0077] The center distance weight is used to measure the distance between the center points of the predicted box and the ground truth box, which can reflect the positioning situation of the box; the aspect ratio weight is used to measure the aspect ratio of the predicted box and the ground truth box, which can reflect the predicted shape situation of the box; the area difference weight is used to measure the area difference between the predicted box and the ground truth box, reflecting the size matching situation of the box.
[0078] The weight loss function of the asparagus grading model is:
[0079] WIoU = IoU - α×d center_dist - β×d aspect_ratio - γ×d area_difference ;
[0080] Among them, WIoU is the weight loss function of the asparagus grading model, IoU is the intersection over union between the predicted bounding box and the ground truth bounding box, d center_dist represents the center point weight, d aspect_ratio represents the aspect ratio weight, d area_difference represents the area difference weight, and α, β, γ represent weight factors.
[0081] Furthermore, IoU, as the criterion for the overlap between the predicted bounding box and the ground truth bounding box, is calculated as follows:
[0082]
[0083] where A represents the predicted bounding box and B represents the ground truth bounding box.
[0084] After training is completed, the trained asparagus grading model is used to perform the following steps as Figure 4 shown:
[0085] Step S10: Input the asparagus image into the trained asparagus grading model to extract the asparagus stem diameter features of each target asparagus in the asparagus image through the asparagus grading model;
[0086] Step S20: Determine the grading result of the target asparagus according to the asparagus stem diameter features.
[0087] In the technical solution provided in this embodiment, based on the YOLOv8 model, an SE attention mechanism module is introduced into the backbone network and a Slim-neck structure is introduced into the neck network. The effective feature extraction ability is improved by using the SE attention mechanism, and the Slim-neck module reduces the number of model parameters and complexity; on the other hand, a weight loss function containing multiple weight factors is proposed to improve the regression accuracy of the bounding box in the object detection task, realizing the precise grading of post-harvest asparagus and providing technical support for the deployment of the asparagus grading device.
[0088] Embodiment 2
[0089] In this embodiment, a method for training an asparagus grading model is proposed. This method is used before step S10 and specifically includes:
[0090] S30: Label the asparagus image dataset containing different grades, where the size of each labeled bounding box is the minimum circumscribed rectangle of the asparagus plant;
[0091] S40: Divide the labeled asparagus image dataset into a training set, a validation set, and a test set according to a ratio of 7:2:1;
[0092] S50: Perform flipping, mirroring, and adding noise operations on the divided training set to complete the preprocessing of the training set;
[0093] S60, train the asparagus grading model based on the preprocessed training set, the validation set, and the test set.
[0094] Example 3
[0095] In this example, an ablation experiment is conducted on the asparagus grading model proposed in the first example to verify its processing accuracy. Referring to Figure 5 the schematic diagram of the detection effect shown, the results of the ablation experiment are shown in Table 1:
[0096] Table 1. Results of the ablation experiment
[0097]
[0098] It can be seen that the precision of the original YOLOv8 model is 95.6%, the recall rate is 97.1%, and the mean average precision is 85.8%. The average precisions of thin asparagus, medium asparagus, and thick asparagus are 99.0%, 98.0%, and 99.1% respectively. After adding the SE attention mechanism to the tenth layer of the backbone network, the precision, recall rate, and mean average precision are increased by 0.9, 1.0, and 0.7 percentage points respectively, and the average precisions of thin asparagus, medium asparagus, and thick asparagus are increased by 0.5, 1.1, and 0.4 percentage points respectively, which can effectively improve the ability of feature extraction.
[0099] After introducing Slim-neck in the neck, the precision and mean average precision are increased by 0.9 and 4.8 percentage points respectively, the recall rate decreases by 1.7 percentage points, the average precision of thin asparagus increases by 0.4 percentage points, the average precision of medium asparagus decreases by 0.2 percentage points, and the average precision of thick asparagus remains unchanged, which can effectively improve the precision of the model.
[0100] After replacing the loss function with WIoU Loss, the precision of the model decreases by 2.5 percentage points, the recall rate increases by 0.5 percentage points, the mean average precision decreases by 0.6 percentage points, and the average precisions of thin asparagus, medium asparagus, and thick asparagus increase by 0.5, 0.6, and 0.2 percentage points respectively.
[0101] Finally, after adding the SE attention mechanism, introducing the Slim-neck module, and replacing the loss function with WIoU Loss simultaneously, the precision, recall rate, and mean average precision are increased by 3.4, 1.4, and 5.5 percentage points respectively, and the average precisions of thin asparagus, medium asparagus, and thick asparagus are increased by 0.5, 1.3, and 0.4 percentage points respectively, which can improve the ability to extract the stem diameter features of asparagus and reduce the false detection at the boundaries of the three grades of asparagus.
[0102] Example 4
[0103] Referring toFigure 6 Schematic diagram of the asparagus grading model architecture. In this embodiment, the asparagus grading model includes:
[0104] A backbone network, which is based on YOLOv8 and includes an SE attention mechanism module for extracting the asparagus stem thickness features;
[0105] A neck network, including a Slim-neck module for reducing the model's computational overhead;
[0106] Among them, the weight loss function of the asparagus grading model is:
[0107] WIoU = IoU - α×d center_dist -β×d aspect_ratio -γ×d area_difference ;
[0108] Among them, WIoU is the weight loss function of the asparagus grading model, IoU is the intersection over union between the predicted box and the ground truth box, d center_dist represents the center point weight, d aspect_ratio represents the aspect ratio weight, d area_difference represents the area difference weight, and α, β, γ represent the weight factors.
[0109] Further and optionally, the Slim-neck module includes a vector convolution-group shuffle spatial pyramid pooling module and two layers of group shuffle convolutional layers;
[0110] The vector convolution-group shuffle spatial pyramid pooling module is used to reduce the dimension of the initial asparagus image; the dimension-reduced asparagus image is input into the two layers of group shuffle convolutional layers for processing, and the first group shuffle convolutional feature map and the second group shuffle convolutional feature map output by the two layers of group shuffle convolutional layers are Concat spliced to obtain the asparagus stem thickness features;
[0111] The group shuffle convolutional layer is used to process the asparagus image input by the vector convolution-group shuffle spatial pyramid pooling module through a standard convolutional layer to obtain a first output feature map with halved number of channels; the output feature map is input into a depthwise separable convolutional layer to perform a convolutional operation on each channel independently through the depthwise separable convolutional layer to generate a second output feature map; the first output feature map and the second output feature map are spliced to form a third feature map, and after rearranging the channel order of the third feature map, it is output to obtain a group shuffle convolutional feature map.
[0112] As an implementation scheme, Figure 7 It is a schematic diagram of the architecture of the hardware operating environment of the asparagus grading system involved in the embodiment of the present application.
[0113] As Figure 7As shown, the asparagus grading system may include: a processor 1001, such as a CPU, a memory 1005, a user interface 1003, a network interface 1004, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0114] Those skilled in the art can understand that Figure 7 the asparagus grading system architecture shown in does not constitute a limitation on the asparagus grading system, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0115] As Figure 7 shown, the memory 1005, as a storage medium, may include an operating system, a network communication module, a user interface module, and a computer program. Among them, the operating system is a program for managing and controlling the hardware and software resources of the asparagus grading system, and the operation of the computer program and other software or programs.
[0116] In Figure 7 the asparagus grading system shown, the user interface 1003 is mainly used to connect to the terminal and communicate with the terminal for data; the network interface 1004 is mainly used to connect to the background server and communicate with the background server for data; the processor 1001 can be used to call the computer program stored in the memory 1005.
[0117] In this embodiment, the asparagus grading system includes: a memory 1005, a processor 1001, and a computer program stored on the memory and executable on the processor, where:
[0118] When the processor 1001 calls the computer program stored in the memory 1005, it performs the following operations:
[0119] Input the asparagus image into the trained asparagus grading model to extract the asparagus stem thickness features of each target asparagus in the asparagus image through the asparagus grading model;
[0120] Determine the grading result of the target asparagus according to the asparagus stem thickness features;
[0121] Among them, the backbone network of the asparagus grading model includes an SE attention mechanism module, and the neck network includes a Slim-neck module;
[0122] The weight loss function of the asparagus grading model is:
[0123] WIoU = IoU - α × d center_dist -β × d aspect_ratio -γ × d area_difference ;
[0124] Among them, WIoU is the weight loss function of the asparagus grading model, IoU is the intersection over union between the predicted box and the ground truth box, and d center_dist represents the center point weight, and d aspect_ratio represents the aspect ratio weight, and d area_difference represents the area difference weight, and α, β, and γ represent weight factors.
[0125] When the processor 1001 calls the computer program stored in the memory 1005, it performs the following operations:
[0126] Reduce the dimension of the initial asparagus image;
[0127] Input the dimension-reduced asparagus image into the two-layer grouped shuffle convolutional layer for processing, and concatenate the first grouped shuffle convolutional feature map and the second grouped shuffle convolutional feature map output by the two-layer grouped shuffle convolutional layer to obtain the asparagus stem thickness feature.
[0128] When the processor 1001 calls the computer program stored in the memory 1005, it performs the following operations:
[0129] Process the asparagus image input to the vector convolutional-grouped shuffle spatial pyramid pooling module through a standard convolutional layer to obtain a first output feature map;
[0130] Input the output feature map into the depthwise separable convolutional layer to generate a second output feature map by performing convolutional operations on each channel independently through the depthwise separable convolutional layer;
[0131] Concatenate the first output feature map and the second output feature map to form a third feature map, rearrange the channel order of the third feature map, and then output it to obtain the grouped shuffle convolutional feature map.
[0132] When the processor 1001 calls the computer program stored in the memory 1005, it performs the following operations:
[0133] Label the asparagus image dataset containing different grades, where the size of each labeled box is the minimum bounding rectangle of the asparagus plant;
[0134] Divide the labeled asparagus image dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1;
[0135] Perform flipping, mirroring, and adding noise operations on the divided training set to complete the preprocessing of the training set;
[0136] Train the asparagus grading model based on the preprocessed training set, the validation set, and the test set.
[0137] In addition, those of ordinary skill in the art can understand that all or part of the processes in the methods of implementing the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the asparagus grading system to implement the process steps of the embodiments of the above methods.
[0138] Therefore, the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements each step of the asparagus grading method based on the asparagus grading model as described in the above embodiments.
[0139] Among them, the computer-readable storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.
[0140] It should be noted that since the storage medium provided in the embodiments of the present application is the storage medium used for implementing the methods of the embodiments of the present application, based on the methods introduced in the embodiments of the present application, those skilled in the art can understand the specific structure and deformation of the storage medium, so it will not be elaborated here. Any storage medium used in the methods of the embodiments of the present application belongs to the scope to be protected by the present application.
[0141] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0142] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a device for implementing the functions specified in multiple blocks.
[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a device for implementing the functions specified in multiple blocks.
[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a device for implementing the functions specified in multiple blocks.
[0145] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0146] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0147] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and variations.
Claims
1. An asparagus grading method based on an asparagus grading model, characterized in that, The method includes the following steps: Input the asparagus image into the trained asparagus grading model to extract the asparagus stem diameter features of each target asparagus in the asparagus image through the asparagus grading model; Determine the grading result of the target asparagus according to the asparagus stem diameter features; Among them, the backbone network of the asparagus grading model includes an SE attention mechanism module, and the neck network includes a Slim-neck module; The weight loss function of the asparagus grading model is: WIoU = IoU - α×d center_dist - β×d aspect_ratio - γ×d area_difference ; Among them, WIoU is the weight loss function of the asparagus grading model, IoU is the intersection over union between the predicted box and the ground truth box, d center_dist represents the center point weight, d aspect_ratio represents the aspect ratio weight, d area_difference represents the area difference weight, and α, β, and γ represent weight factors.
2. The method according to claim 1, wherein The Slim-neck module includes a vector convolution-group shuffle spatial pyramid pooling module and two layers of group shuffle convolutional layers. The processing process of the vector convolution-group shuffle spatial pyramid pooling module includes: Reduce the dimension of the initial asparagus image; Input the dimension-reduced asparagus image into the two layers of the group shuffle convolutional layers for processing, and perform Concat splicing on the first group shuffle convolutional feature map and the second group shuffle convolutional feature map output by the two layers of group shuffle convolutional layers to obtain the asparagus stem diameter features.
3. The method according to claim 2, wherein The processing process of the group shuffle convolutional layer includes: Process the asparagus image input by the vector convolution-group shuffle spatial pyramid pooling module through a standard convolutional layer to obtain a first output feature map; Input the output feature map into a depthwise separable convolutional layer to perform convolutional operations on each channel independently through the depthwise separable convolutional layer to generate a second output feature map; Splice the first output feature map and the second output feature map to form a third feature map, rearrange the channel order of the third feature map and then output it to obtain a group shuffle convolutional feature map.
4. The method according to claim 1, wherein The expression of the SE attention mechanism module includes: y = ReLU(W1·Z c ) s = Sigmoid(W2·Z c ) where Z C is the compressed feature vector; y is the activation function; represents the weighted feature map obtained by element-wise multiplying and weighting each channel of the weight vector s with the feature map X; ReLU and Sigmoid are both activation functions; W1 and W2 are both weight matrices of dimension (C / r)×C; H, W, and C are the length, width, and number of channels of the feature map, i represents the row index in the spatial dimension of the feature map, and j represents the column index in the spatial dimension of the feature map.
5. The method according to claim 1, wherein Before the step of inputting the asparagus image into the trained asparagus grading model, it further includes: Annotate the asparagus image dataset containing different grades, where the size of each annotation box is the minimum bounding rectangle of the asparagus plant; Divide the annotated asparagus image dataset into a training set, a validation set, and a test set according to a ratio of 7:2:1; Perform flipping, mirroring, and noise addition operations on the divided training set to complete the preprocessing of the training set; Based on the preprocessed training set, the validation set, and the test set, train the asparagus grading model.
6. The method according to claim 1, wherein The grading results include thick asparagus, medium asparagus, and thin asparagus.
7. An asparagus grading model, characterized in that, The asparagus grading model includes: A backbone network, which is framed by YOLOv8 and includes an SE attention mechanism module for extracting asparagus stem diameter features; A neck network, including a Slim-neck module for reducing the computational overhead of the model; Among them, the weight loss function of the asparagus grading model is: WIoU = IoU - α×d center_dist - β×d aspect_ratio - γ×d area_difference ; Among them, WIoU is the weight loss function of the asparagus grading model, IoU is the intersection over union between the predicted box and the ground truth box, d center_dist represents the center point weight, d aspect_ratio represents the aspect ratio weight, d area_difference represents the area difference weight, and α, β, and γ represent weight factors.
8. The asparagus grading model according to claim 7, wherein The Slim-neck module includes a vector convolution-group shuffle spatial pyramid pooling module and two layers of group shuffle convolutional layers; The vector convolution-group shuffle spatial pyramid pooling module is used to reduce the dimension of the initial asparagus image; input the dimension-reduced asparagus image into the two layers of the group shuffle convolutional layers for processing, and perform Concat splicing on the first group shuffle convolutional feature map and the second group shuffle convolutional feature map output by the two layers of group shuffle convolutional layers to obtain the asparagus stem diameter features; The grouped shuffle convolutional layer is used to process the asparagus image input by the vector convolution-grouped shuffle spatial pyramid pooling module through a standard convolutional layer to obtain a first output feature map with the number of channels halved; the output feature map is input into a depthwise separable convolutional layer to perform convolution operations on each channel independently through the depthwise separable convolutional layer to generate a second output feature map; the first output feature map and the second output feature map are concatenated to form a third feature map, and after rearranging the channel order of the third feature map, it is output to obtain a grouped shuffle convolutional feature map.
9. An asparagus grading system, characterized in that, The asparagus grading system includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the asparagus grading method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the steps of the asparagus grading method based on the asparagus grading model according to any one of claims 1 to 6.