Method and apparatus for identifying small fruit, terminal device, and storage medium
Through the improved YOLOv8 model, combined with the small object detection head and selective kernel attention module, the problem of poor recognition of small fruits is solved, and more efficient judgment and picking of fruit maturity is achieved, reducing labor intensity and cost.
Patent Information
- Application Number
- PCT/CN2024/105773
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-07-16
- Publication Date
- 2025-07-03
AI Technical Summary
The prior art has poor results in identifying small fruits, especially in lychee picking robots, and it is difficult to effectively identify and judge the fruit maturity.
The improved YOLOv8 model is adopted, combining the small object detection head, selective kernel attention module and RepVGG module to build a fruit recognition model. Through feature map fusion and training sample images, the detection and maturity judgment of small fruits are improved.
It improves the identification accuracy and maturity judgment ability of small fruits, improves the picking efficiency, and reduces the labor intensity and cost of manual operations.
Smart Images

Figure CN2024105773_03072025_PF_FP_ABST
Abstract
Description
Method, device, terminal equipment and storage medium for identifying small fruits Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a method, device, terminal equipment and storage medium for identifying small fruits. Background Art
[0002] [Corrected 24 July 2024, in accordance with Regulation 26] In 2022, the domestic lychee planting area reached 526,100 hectares, remaining essentially stable; lychee production was estimated at 2.2227 million tons. Traditional lychee harvesting requires manual labor, which is labor-intensive. With an aging population and rising labor costs, harvesting efficiency is expected to decline. To improve harvesting efficiency, reduce damage, and address labor shortages, deep learning-based picking robots are being deployed.
[0003] Deep learning-based object detection methods have demonstrated promising performance on public datasets, but existing conventional models rarely consider the case of handling small objects. When faced with small objects like lychees, insufficient feature extraction often occurs, resulting in poor recognition results.
[0004] Summary of the Invention
[0005] The present invention provides a method, device, terminal equipment and storage medium for identifying small fruits, so as to solve the technical problem that the existing technology has poor recognition effect on small fruits.
[0006] In order to solve the above technical problems, an embodiment of the present invention provides a method for identifying small fruits, comprising:
[0007] Acquire fruit images;
[0008] Inputting the fruit image into a fruit recognition model so that the fruit recognition model recognizes the position of each fruit in the fruit image and determines whether each fruit is ripe;
[0009] Among them, the fruit recognition model is constructed based on the YOLOv8 model; the fruit recognition model includes: a small target detection head; the small target detection head is connected to the StageLayer1 module and the second-level upsampling module of the YOLOv8 model, and when identifying the fruit image, it receives the output of the StageLayer1 module and the second-level upsampling module, and generates a feature map based on the output of the StageLayer1 module and the second-level upsampling module.
[0010] As a preferred solution, the fruit recognition model further includes: a selective kernel attention module;
[0011] The selective kernel attention module is connected to the SPPF module, the first-level sampling module and the first-level detection head of the YOLOv8 model;
[0012] The selective kernel attention module is used to receive the output of the SPPF module when recognizing the fruit image;
[0013] According to the output of the SPPF module, feature maps of different scales are obtained through multiple parallel convolution branches with different kernel sizes;
[0014] Determine a selection weight based on information of all convolution branches; fuse the feature maps of different scales based on the selection weight to obtain a fused feature map;
[0015] The fused feature map is transmitted to the first-stage upsampling module and the first-stage detection head.
[0016] As a preferred solution, the fruit recognition model further includes: a plurality of RepVGG modules;
[0017] The RepVGG module is used to replace the c2f module of the YOLOv8 model; and
[0018] During the training of the fruit recognition model, the RepVGG module includes a 3×3 convolution branch, a 1×1 convolution branch, an identity mapping branch, a BN layer, and a ReLU activation; wherein the 3×3 convolution branch, the 1×1 convolution branch, and the identity mapping branch are parallel to each other; and
[0019] During the inference process of the fruit recognition model, the RepVGG module includes: a 3×3 convolutional layer; wherein, the 3×3 convolutional layer is obtained by a structural reparameterization operation of the 3×3 convolutional branch, the 1×1 convolutional branch and the identity mapping branch.
[0020] The structure reparameterization operation includes:
[0021] Fusing the 3×3 convolution branch with the BN layer to generate a first 3×3 convolution BN branch; fusing the 1×1 convolution branch with the BN layer to generate a 1×1 convolution BN branch; fusing the identity mapping branch with the BN layer to generate an identity mapping BN branch;
[0022] Convert the convolution size of the 1×1 convolutional BN branch to 3×3 to obtain a second 3×3 convolutional BN branch;
[0023] The first 3×3 convolutional BN branch, the second 3×3 convolutional BN branch, and the identity mapping BN branch are fused to obtain the 3×3 convolutional layer.
[0024] As a preferred solution, the training process of the fruit recognition model includes:
[0025] Acquire several sample fruit images; wherein all target objects in the sample fruit images have been marked as corresponding rectangular frames and corresponding labels indicating whether the fruits are ripe are set;
[0026] The fruit recognition model is trained based on the sample fruit images.
[0027] On the basis of the above embodiment, another embodiment of the present invention provides a device for identifying small fruits, characterized in that it includes: an image acquisition module and a recognition module;
[0028] The image acquisition module is used to acquire fruit images;
[0029] The recognition module is used to input the fruit image into a fruit recognition model so that the fruit recognition model recognizes the fruit image;
[0030] Among them, the fruit recognition model is constructed based on the YOLOv8 model; the fruit recognition model includes: a small target detection head; the small target detection head is connected to the StageLayer1 module and the second-level upsampling module of the YOLOv8 model, and when identifying the fruit image, it receives the output of the StageLayer1 module and the second-level upsampling module, and generates a fused feature map based on the output of the StageLayer1 module and the second-level upsampling module.
[0031] As a preferred solution, the fruit recognition model further includes: a selective kernel attention module;
[0032] The selective kernel attention module is connected to the SPPF module, the first-level sampling module and the first-level detection head of the YOLOv8 model;
[0033] The selective kernel attention module is used to receive the output of the SPPF module when recognizing the fruit image;
[0034] According to the output of the SPPF module, feature maps of different scales are obtained through multiple parallel convolution branches with different kernel sizes;
[0035] Determine a selection weight based on information of all convolution branches; fuse the feature maps of different scales based on the selection weight to obtain a fused feature map;
[0036] The fused feature map is transmitted to the first-stage upsampling module and the first-stage detection head.
[0037] As a preferred solution, the fruit recognition model further includes: a plurality of RepVGG modules;
[0038] The RepVGG module is used to replace the c2f module of the YOLOv8 model; and
[0039] During the training of the fruit recognition model, the RepVGG module includes a 3×3 convolution branch, a 1×1 convolution branch, an identity mapping branch, a BN layer, and a ReLU activation; wherein the 3×3 convolution branch, the 1×1 convolution branch, and the identity mapping branch are parallel to each other; and
[0040] During the inference process of the fruit recognition model, the RepVGG module includes: a 3×3 convolutional layer; wherein, the 3×3 convolutional layer is obtained by a structural reparameterization operation of the 3×3 convolutional branch, the 1×1 convolutional branch and the identity mapping branch.
[0041] The structure reparameterization operation includes:
[0042] Fusing the 3×3 convolution branch with the BN layer to generate a first 3×3 convolution BN branch; fusing the 1×1 convolution branch with the BN layer to generate a 1×1 convolution BN branch; fusing the identity mapping branch with the BN layer to generate an identity mapping BN branch;
[0043] Convert the convolution size of the 1×1 convolutional BN branch to 3×3 to obtain a second 3×3 convolutional BN branch;
[0044] The first 3×3 convolutional BN branch, the second 3×3 convolutional BN branch, and the identity mapping BN branch are fused to obtain the 3×3 convolutional layer.
[0045] As a preferred solution, the training process of the fruit recognition model includes:
[0046] Acquire several sample fruit images; wherein all target objects in the sample fruit images have been marked as corresponding rectangular frames and corresponding labels indicating whether the fruits are ripe are set;
[0047] The fruit recognition model is trained based on the sample fruit images.
[0048] Based on the above embodiments, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method for identifying small fruits described in the above embodiment of the invention.
[0049] Based on the above embodiment, another embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method for identifying small fruits described in the above embodiment of the invention.
[0050] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0051] The present invention obtains a fruit image and inputs the fruit image into a fruit recognition model, so that the fruit recognition model identifies the location of each fruit in the fruit image and determines whether each fruit is ripe. The fruit recognition model is constructed based on the YOLOv8 model and includes a small target detection head. The small target detection head is connected to the StageLayer1 module and the second-level upsampling module of the YOLOv8 model. When recognizing the fruit image, the small target detection head receives the output of the StageLayer1 module and the second-level upsampling module and generates a feature map based on the output of the StageLayer1 module and the second-level upsampling module. The addition of the small target detection layer in the present invention makes the network focus more on the detection of small targets, thereby improving the detection effect of small fruits. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] FIG1 is a schematic flow chart of a method for identifying small fruits provided by one embodiment of the present invention;
[0053] Figure 2 is the original YOLOv8 network architecture diagram;
[0054] Fig. 3 is a network framework diagram of a fruit recognition model of the present invention;
[0055] FIG4 is a block diagram of the selective kernel attention module of the present invention;
[0056] FIG5 is a block diagram of the RepVGG module of the present invention;
[0057] FIG6 is a schematic diagram of the structure reparameterization process of the RepVGG module of the present invention.
[0058] FIG7 is a schematic structural diagram of a device for identifying small fruits provided by an embodiment of the present invention.
[0059] Among them, the figure numbers of the drawings in the specification are as follows: small target detection head 1, selective kernel attention module 2 and RepVGG module 3. DETAILED DESCRIPTION
[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0061] Example 1
[0062] Please refer to FIG1 , which is a flow chart of a method for identifying small fruits according to an embodiment of the present invention, including:
[0063] S1. Acquire fruit images.
[0064] It should be noted that the fruit image is an RGB image and / or a depth image, which is acquired by a RealSense D435 depth camera.
[0065] S2, inputting the fruit image into a fruit recognition model, so that the fruit recognition model recognizes the position of each fruit in the fruit image and determines whether each fruit is ripe;
[0066] Among them, the fruit recognition model is constructed based on the YOLOv8 model; the fruit recognition model includes: a small target detection head; the small target detection head is connected to the StageLayer1 module and the second-level upsampling module of the YOLOv8 model, and when identifying the fruit image, it receives the output of the StageLayer1 module and the second-level upsampling module, and generates a feature map based on the output of the StageLayer1 module and the second-level upsampling module.
[0067] Please refer to Figures 2 and 3. Figure 2 is the original YOLOv8 network architecture diagram, and Figure 3 is the network framework diagram of the fruit recognition model of the present invention, which is an improved YOLOv8 network structure, which includes a small target detection head. The small target detection head receives the output of the StageLayer1 module and the second-level upsampling module, and generates a feature map of size 160×160×45 based on the output of the StageLayer1 module and the second-level upsampling module.
[0068] It's important to note that the original YOLOv8 network downsamples significantly. While downsampling with a backbone stride of 2 allows the network to capture more semantic information, it also loses a significant amount of detailed feature information, a problem that arises from a lack of shallow network information. However, this detailed information includes quality features of small objects, which can be overlooked during downsampling. Deeper feature maps struggle to learn the features of small objects. Adding a small object detection head to the algorithm concatenates the shallower feature maps with the deeper ones for detection. This allows the network to focus more on detecting small objects, improving detection performance.
[0069] In a preferred embodiment, the fruit recognition model further comprises: a selective kernel attention module;
[0070] The selective kernel attention module is connected to the SPPF module, the first-level sampling module and the first-level detection head of the YOLOv8 model;
[0071] The selective kernel attention module is used to receive the output of the SPPF module when recognizing the fruit image;
[0072] According to the output of the SPPF module, feature maps of different scales are obtained through multiple parallel convolution branches with different kernel sizes;
[0073] Determine a selection weight based on information of all convolution branches; fuse the feature maps of different scales based on the selection weight to obtain a fused feature map;
[0074] The fused feature map is transmitted to the first-stage upsampling module and the first-stage detection head.
[0075] Please refer to Figures 2, 3 and 4, wherein Figure 4 is a structural diagram of the selective kernel attention module of the present invention. The network framework diagram of the fruit recognition model of the present invention includes a selective kernel attention module, which is connected to the SPPF module, the first-level sampling module and the first-level detection head of the YOLOv8 model; the selective kernel attention module is used to receive the output of the SPPF module when recognizing the fruit graphic; based on the output of the SPPF module, multiple parallel convolution branches with different kernel sizes are used to obtain feature maps of different scales; based on the information of all convolution branches, a selection weight is determined; based on the selection weight, the feature maps of different scales are fused to obtain a fused feature map; and the fused feature map is transmitted to the first-level upsampling module and the first-level detection head.
[0076] It's important to note that the Selective Kernel Attention module introduces the Selective Row Kernel Attention mechanism, an attention mechanism that uses different kernel sizes in convolutional neural networks to capture multi-scale contextual information. In traditional convolutional neural networks, the receptive field size is fixed, which limits their ability to effectively capture both local and global contextual information. The Selective Row Kernel Attention mechanism addresses this limitation by introducing multiple parallel convolution branches, each using a different kernel size. These branches can capture information at different spatial scales, giving the model a better understanding of the input features. The key idea of Selective Row Kernel Attention is to leverage channel-wise attention across different kernel sizes. The attention mechanism learns the importance of each channel for each kernel size, allowing the network to selectively focus on the most informative kernel size. This adaptability allows the model to dynamically adjust the receptive field and gather relevant information from different scales. By incorporating Selective Row Kernel Attention into the convolutional neural network architecture, the model is able to simultaneously capture fine-grained local details and a wider range of global context, thereby improving performance in various computer vision tasks and accelerating model inference time.
[0077] In a preferred embodiment, the fruit recognition model further comprises: a plurality of RepVGG modules;
[0078] The RepVGG module is used to replace the c2f module of the YOLOv8 model; and
[0079] During the training of the fruit recognition model, the RepVGG module includes a 3×3 convolution branch, a 1×1 convolution branch, an identity mapping branch, a BN layer, and a ReLU activation; wherein the 3×3 convolution branch, the 1×1 convolution branch, and the identity mapping branch are parallel to each other; and
[0080] During the inference process of the fruit recognition model, the RepVGG module includes: a 3×3 convolutional layer; wherein, the 3×3 convolutional layer is obtained by a structural reparameterization operation of the 3×3 convolutional branch, the 1×1 convolutional branch and the identity mapping branch.
[0081] The structure reparameterization operation includes:
[0082] Fusing the 3×3 convolution branch with the BN layer to generate a first 3×3 convolution BN branch; fusing the 1×1 convolution branch with the BN layer to generate a 1×1 convolution BN branch; fusing the identity mapping branch with the BN layer to generate an identity mapping BN branch;
[0083] Convert the convolution size of the 1×1 convolutional BN branch to 3×3 to obtain a second 3×3 convolutional BN branch;
[0084] The first 3×3 convolutional BN branch, the second 3×3 convolutional BN branch, and the identity mapping BN branch are fused to obtain the 3×3 convolutional layer.
[0085] Please refer to Figures 3 and 5, where Figure 5 is a structural diagram of the RepVGG module of the present invention. The network framework diagram of the fruit recognition model of the present invention includes several RepVGG modules, which are used to replace the c2f module of the YOLOv8 model.
[0086] It should be noted that during training, each layer of the RepVGG module has three branches: the identity mapping branch, the 1×1 convolution branch, and the 3×3 convolution layer. During model training, the output is y = x + g(x) + f(x), where y represents the output, and x, g(x), and f(x) represent the corresponding identity mapping, 1×1 convolution, and 3×3 convolution, respectively. Each layer requires three parameter blocks, and for an n-layer network, 3n parameter blocks are required. Therefore, reparameterization is necessary to reduce the number of model parameters during inference. Structural reparameterization means using different structures for training and inference, but the same set of parameters. RepVGG converts the three-branch network into an equivalent, simplified single-branch network.
[0087] Please refer to FIG6 , which is a schematic diagram of the structure reparameterization process of the RepVGG module of the present invention. The structure reparameterization is mainly divided into three steps;
[0088] (1) Fusing the 3×3 convolution branch with the BN layer to generate a first 3×3 convolution BN branch; fusing the 1×1 convolution branch with the BN layer to generate a 1×1 convolution BN branch; fusing the identity mapping branch with the BN layer to generate an identity mapping BN branch;
[0089] (2) converting the convolution size of the 1×1 convolutional BN branch to 3×3 to obtain a second 3×3 convolutional BN branch;
[0090] (3) The first 3×3 convolutional BN branch, the second 3×3 convolutional BN branch, and the identity mapping BN branch are fused to obtain the 3×3 convolutional layer.
[0091] In a preferred embodiment, the training process of the fruit recognition model includes:
[0092] Acquire several sample fruit images; wherein all target objects in the sample fruit images have been marked as corresponding rectangular frames and corresponding labels indicating whether the fruits are ripe are set;
[0093] The fruit recognition model is trained based on the sample fruit images.
[0094] It should be noted that the litchi recognition target detection model is trained using a back-propagation iterative method to obtain model parameters suitable for litchi recognition target detection.
[0095] Example 2
[0096] Please refer to FIG7 , which is a schematic structural diagram of a device for identifying small fruits according to an embodiment of the present invention. The device includes: an image acquisition module and a recognition module;
[0097] The image acquisition module is used to acquire fruit images;
[0098] The recognition module is used to input the fruit image into a fruit recognition model so that the fruit recognition model recognizes the fruit image;
[0099] Among them, the fruit recognition model is constructed based on the YOLOv8 model; the fruit recognition model includes: a small target detection head; the small target detection head is connected to the StageLayer1 module and the second-level upsampling module of the YOLOv8 model, and when identifying the fruit image, it receives the output of the StageLayer1 module and the second-level upsampling module, and generates a fused feature map based on the output of the StageLayer1 module and the second-level upsampling module.
[0100] In a preferred embodiment, the fruit recognition model further comprises: a selective kernel attention module;
[0101] The selective kernel attention module is connected to the SPPF module, the first-level sampling module and the first-level detection head of the YOLOv8 model;
[0102] The selective kernel attention module is used to receive the output of the SPPF module when recognizing the fruit image;
[0103] According to the output of the SPPF module, feature maps of different scales are obtained through multiple parallel convolution branches with different kernel sizes;
[0104] Determine a selection weight based on information of all convolution branches; fuse the feature maps of different scales based on the selection weight to obtain a fused feature map;
[0105] The fused feature map is transmitted to the first-stage upsampling module and the first-stage detection head.
[0106] In a preferred embodiment, the fruit recognition model further comprises: a plurality of RepVGG modules;
[0107] The RepVGG module is used to replace the c2f module of the YOLOv8 model; and
[0108] During the training of the fruit recognition model, the RepVGG module includes a 3×3 convolution branch, a 1×1 convolution branch, an identity mapping branch, a BN layer, and a ReLU activation; wherein the 3×3 convolution branch, the 1×1 convolution branch, and the identity mapping branch are parallel to each other; and
[0109] During the inference process of the fruit recognition model, the RepVGG module includes: a 3×3 convolutional layer; wherein, the 3×3 convolutional layer is obtained by a structural reparameterization operation of the 3×3 convolutional branch, the 1×1 convolutional branch and the identity mapping branch.
[0110] The structure reparameterization operation includes:
[0111] Fusing the 3×3 convolution branch with the BN layer to generate a first 3×3 convolution BN branch; fusing the 1×1 convolution branch with the BN layer to generate a 1×1 convolution BN branch; fusing the identity mapping branch with the BN layer to generate an identity mapping BN branch;
[0112] Convert the convolution size of the 1×1 convolutional BN branch to 3×3 to obtain a second 3×3 convolutional BN branch;
[0113] The first 3×3 convolutional BN branch, the second 3×3 convolutional BN branch, and the identity mapping BN branch are fused to obtain the 3×3 convolutional layer.
[0114] In a preferred embodiment, the training process of the fruit recognition model includes:
[0115] Acquire several sample fruit images; wherein all target objects in the sample fruit images have been marked as corresponding rectangular frames and corresponding labels indicating whether the fruits are ripe are set;
[0116] The fruit recognition model is trained based on the sample fruit images.
[0117] Example 3
[0118] Accordingly, an embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method for identifying small fruits described in the above-mentioned embodiment of the invention.
[0119] Example 4
[0120] Accordingly, an embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method for identifying small fruits described in the above-mentioned embodiment of the invention.
[0121] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.
[0122] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0123] The device can be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The device can include, but is not limited to, a processor and a memory.
[0124] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the device, connecting various parts of the entire device using various interfaces and lines.
[0125] The memory can be used to store the computer program, and the processor realizes various functions of the device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0126] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0127] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
[0128] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A recognition method for small fruits, characterized in that, Including: Obtain a fruit image; Input the fruit image into a fruit recognition model, so that the fruit recognition model recognizes the positions of each fruit in the fruit image and determines whether each fruit is ripe; Among them, the fruit recognition model is constructed based on the YOLOv8 model; the fruit recognition model includes: a small object detection head; the small object detection head is connected to the StageLayer1 module and the second-level upsampling module of the YOLOv8 model, and when recognizing the fruit image, it receives the outputs of the StageLayer1 module and the second-level upsampling module, and generates a feature map according to the outputs of the StageLayer1 module and the second-level upsampling module.
2. The identification method for small fruits according to claim 1, characterized in that, The fruit recognition model further includes: a selective kernel attention module; The selective kernel attention module is connected to the SPPF module, the first-level sampling module and the first-level detection head of the YOLOv8 model; The selective kernel attention module is used to receive the output of the SPPF module when recognizing the fruit graph; According to the output of the SPPF module, obtain feature maps of different scales through multiple parallel convolutional branches with different kernel sizes; Determine the selection weights according to the information of all convolutional branches; fuse the feature maps of different scales according to the selection weights to obtain a fused feature map; Transmit the fused feature map to the first-level upsampling module and the first-level detection head.
3. The identification method for small fruits according to claim 1, wherein The fruit recognition model further includes: a number of RepVGG modules; The RepVGG module is used to replace the c2f module of the YOLOv8 model; and During the training process of the fruit recognition model, the RepVGG module includes a 3×3 convolutional branch, a 1×1 convolutional branch, an identity mapping branch, a BN layer and a ReLU activation; among them, the 3×3 convolutional branch, the 1×1 convolutional branch and the identity mapping branch are parallel to each other; and During the inference process of the fruit recognition model, the RepVGG module includes: a 3×3 convolutional layer; where the 3×3 convolutional layer is obtained by a structural reparameterization operation of the 3×3 convolutional branch, the 1×1 convolutional branch and the identity mapping branch. Among them, the structural reparameterization operation includes: Fuse the 3×3 convolutional branch with the BN layer to generate a first 3×3 convolutional BN branch; fuse the 1×1 convolutional branch with the BN layer to generate a 1×1 convolutional BN branch; fuse the identity mapping branch with the BN layer to generate an identity mapping BN branch; Convert the convolution size of the 1×1 convolutional BN branch to 3×3 to obtain a second 3×3 convolutional BN branch; Fuse the first 3×3 convolutional BN branch, the second 3×3 convolutional BN branch and the identity mapping BN branch to obtain the 3×3 convolutional layer.
4. The identification method for small fruits according to claim 1, wherein The training process of the fruit recognition model includes: Obtain a number of sample fruit images; among them, all target objects in the sample fruit images have been labeled with corresponding rectangular boxes and corresponding labels indicating whether the fruits are ripe are set; Train the fruit recognition model according to the sample fruit images.
5. An identification device for small fruits, characterized in that, Including: An image acquisition module and a recognition module; The image acquisition module is used to obtain fruit images; The recognition module is used to input the fruit images into the fruit recognition model so that the fruit recognition model recognizes the fruit images; Among them, the fruit recognition model is constructed based on the YOLOv8 model; the fruit recognition model includes: a small target detection head; the small target detection head is connected to the StageLayer1 module and the second-level upsampling module of the YOLOv8 model, and when recognizing the fruit images, it receives the outputs of the StageLayer1 module and the second-level upsampling module, and generates a fused feature map according to the outputs of the StageLayer1 module and the second-level upsampling module.
6. The small fruit recognition device according to claim 5, characterized in that The fruit recognition model further includes: a selective kernel attention module; The selective kernel attention module is connected to the SPPF module, the first-level sampling module and the first-level detection head of the YOLOv8 model; The selective kernel attention module is used to receive the output of the SPPF module when recognizing the fruit graphics; According to the output of the SPPF module, obtain feature maps of different scales through multiple parallel convolutional branches with different kernel sizes; Determine the selection weights according to the information of all convolutional branches; fuse the feature maps of different scales according to the selection weights to obtain a fused feature map; Transmit the fused feature map to the first-level upsampling module and the first-level detection head.
7. The small fruit recognition device according to claim 5, characterized in that The fruit recognition model further includes: a number of RepVGG modules; The RepVGG module is used to replace the c2f module of the YOLOv8 model; and During the training process of the fruit recognition model, the RepVGG module includes a 3×3 convolutional branch, a 1×1 convolutional branch, an identity mapping branch, a BN layer and a ReLU activation; among them, the 3×3 convolutional branch, 1×1 convolutional branch and identity mapping branch are parallel to each other; and During the inference process of the fruit recognition model, the RepVGG module includes: a 3×3 convolutional layer; where the 3×3 convolutional layer is obtained by a structural reparameterization operation of the 3×3 convolutional branch, 1×1 convolutional branch and identity mapping branch. Among them, the structural reparameterization operation includes: Fuse the 3×3 convolutional branch with the BN layer to generate a first 3×3 convolutional BN branch; fuse the 1×1 convolutional branch with the BN layer to generate a 1×1 convolutional BN branch; fuse the identity mapping branch with the BN layer to generate an identity mapping BN branch; Convert the convolution size of the 1×1 convolution BN branch to 3×3 to obtain a second 3×3 convolution BN branch; Fuse the first 3×3 convolution BN branch, the second 3×3 convolution BN branch, and the identity mapping BN branch to obtain the 3×3 convolution layer.
8. The small fruit recognition device according to claim 5, characterized in that, The training process of the fruit recognition model includes: Obtain a number of sample fruit images; among them, all target objects in the sample fruit images have been labeled with corresponding rectangular boxes, and corresponding labels for indicating whether the fruits are ripe are set; Train the fruit recognition model according to the sample fruit images.
9. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the recognition method for small fruits described in any one of claims 1 to 4.
10. A storage medium, characterized in that, The storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the recognition method for small fruits described in any one of claims 1 to 4.
Citation Information
Patent Citations
Small target detection and recognition method, device and system and storage medium
CN112508924A
Target detection method and device, equipment and medium
CN115631433A
Recognition method and device for small fruits, terminal equipment and storage medium
CN117789199A
Image enhancement device and method for convolutional network apparatus
US20180114294A1
Cited By
Small target detection method of adaptive receptive field
CN121482521A