Improved Yolov8 model-based Or orange defect segmentation method
By introducing ContextGuidedBlock and ACMIX modules in the Yolov8 model, the Wogan image segmentation network model was constructed, which solved the problems of high cost and complex operation of the existing Wogan defect detection system, and achieved high accuracy and low cost Wogan defect detection.
Patent Information
- Application Number
- CN202510260274.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
The existing Wogan defect detection system is costly, complex in operation and poor in applicability, making it difficult to use effectively with limited computing resources.
Using the Wogan defect segmentation method based on the improved Yolov8 model, the Wogan image segmentation network model is constructed, data enhancement and preprocessing is carried out, and the intelligent identification and segmentation of Wogan surface defects is realized.
It has achieved high accuracy of Wogan defect detection, and the classification accuracy reaches more than 95%. It is suitable for scenarios with limited computing resources, with low cost and strong applicability.
Smart Images

Figure CN120220136A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and specifically relates to a method for segmenting defects of ponkan oranges based on an improved Yolov8 model. Background Art
[0002] With the continuous expansion of the planting area and the continuous growth of market demand, ponkan oranges face many challenges during the production process, especially ensuring the fruit quality has become a key issue.
[0003] During the production process of ponkan oranges, due to factors such as weather conditions, soil conditions, pest and disease attacks, and improper management measures, various defects may occur in the fruits, such as roughness, mottling, sunburn, disease spots, mechanical damage, etc. Therefore, efficient and accurate defect detection technology is of great significance for optimizing ponkan orange production, reducing economic losses, and improving the overall quality of products. Visual detection is one of the most widely used technologies in current ponkan orange defect detection. However, the current defect detection systems have deficiencies such as high cost, complex operation, and poor applicability, requiring high requirements for production equipment and being difficult to use when computing resources are limited. Summary of the Invention
[0004] In order to overcome the defects and deficiencies existing in the prior art, the present invention provides a method for segmenting defects of ponkan oranges based on an improved Yolov8 model. The present invention can intelligently identify and segment the defects on the surface of ponkan oranges, and then classify the ponkan oranges according to different quality grades, and has the advantages of low cost, good adaptability, high accuracy, etc., expanding the application scenarios of deep learning technology.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] The present invention provides a method for segmenting defects of ponkan oranges based on an improved Yolov8 model, including the following steps:
[0007] Obtain the surface image data of ponkan oranges and perform data augmentation on the surface image data of ponkan oranges;
[0008] Preprocess the surface image data of ponkan oranges;
[0009] Divide the preprocessed surface image data of ponkan oranges into a training set, a validation set, and a test set;
[0010] Based on the Yolov8 model, construct a ponkan orange image segmentation network model, replace the C2f module in the Backbone part of the Yolov8 model with a ContextGuidedBlock module, add an ACMIX module to the Neck part of the Yolov8 model, and obtain the prediction result of defect segmentation through the Head part of the Yolov8 model;
[0011] Train the ponkan image segmentation network model based on the training set to obtain the trained ponkan image segmentation network model;
[0012] Perform defect recognition tests on the ponkan image segmentation network model based on the test set, and output the accuracy of ponkan defect segmentation;
[0013] Output the predicted result of defect segmentation based on the trained ponkan image segmentation network model.
[0014] As a preferred technical solution, perform data augmentation on the ponkan surface image data, specifically including:
[0015] Perform data augmentation operations on the ponkan surface image data, including random scaling, inversion, cropping, rotation, and optical transformation.
[0016] As a preferred technical solution, preprocess the ponkan surface image data, specifically including:
[0017] Perform noise reduction processing on the ponkan surface image data based on mean filtering;
[0018] Perform image enhancement operations on the ponkan surface image data after mean filtering, and improve the contrast between the defect area and the normal area based on piecewise linear transformation.
[0019] As a preferred technical solution, the ContextGuidedBlock module includes a local feature extractor, a surrounding context extractor, a joint feature extractor, and a global context extractor. The local feature extractor is used to extract the local features of the input image, the surrounding context extractor is used to extract the surrounding context information of the image, the joint feature extractor is used to fuse the local features and the surrounding context information and output them to the global context extractor, and the global context extractor is used to extract the global context information.
[0020] As a preferred technical solution, the ACMIX module converts the input image features into queries, keys, and values, projects the input image feature map based on 1x1 convolution operations to generate a set of intermediate features, calculates the attention weights through similarity matching, and combines them to obtain the result.
[0021] The present invention also provides a ponkan defect segmentation system based on an improved Yolov8 model, including: an image data acquisition module, a data augmentation module, an image preprocessing module, a data partitioning module, an image segmentation network model construction module, a model training module, a model testing module, and a defect segmentation prediction module;
[0022] The image data acquisition module is used to acquire ponkan surface image data;
[0023] The data augmentation module is used to perform data augmentation on the surface image data of ponkan oranges;
[0024] The image preprocessing module is used to preprocess the surface image data of ponkan oranges;
[0025] The data division module is used to divide the preprocessed surface image data of ponkan oranges into a training set, a validation set, and a test set;
[0026] The image segmentation network model construction module is used to construct a ponkan orange image segmentation network model based on the Yolov8 model, replace the C2f module in the Backbone part of the Yolov8 model with the ContextGuidedBlock module, add the ACMIX module to the Neck part of the Yolov8 model, and obtain the prediction result of defect segmentation through the Head part of the Yolov8 model;
[0027] The model training module is used to train the ponkan orange image segmentation network model based on the training set to obtain the trained ponkan orange image segmentation network model;
[0028] The model testing module is used to perform defect recognition testing on the ponkan orange image segmentation network model based on the test set and output the accuracy rate of ponkan orange defect segmentation;
[0029] The defect segmentation prediction module is used to output the prediction result of defect segmentation based on the trained ponkan orange image segmentation network model.
[0030] As a preferred technical solution, the data augmentation module is used to perform data augmentation on the surface image data of ponkan oranges, specifically including:
[0031] Perform data augmentation operations on the surface image data of ponkan oranges, including random scaling, inversion, cropping, rotation, and optical transformation.
[0032] As a preferred technical solution, the image preprocessing module is used to preprocess the surface image data of ponkan oranges, specifically including:
[0033] Perform noise reduction processing on the surface image data of ponkan oranges based on mean filtering;
[0034] Perform image enhancement operations on the surface image data of ponkan oranges after mean filtering, and improve the contrast between the defect area and the normal area based on piecewise linear transformation.
[0035] As a preferred technical solution, the ContextGuidedBlock module includes a local feature extractor, a surrounding context extractor, a joint feature extractor, and a global context extractor. The local feature extractor is used to extract local features of the input image, the surrounding context extractor is used to extract surrounding context information of the image, the joint feature extractor is used to fuse the local features and the surrounding context information and output them to the global context extractor, and the global context extractor is used to extract global context information.
[0036] As a preferred technical solution, the ACMIX module converts the input image features into queries, keys, and values, projects the input image feature map based on 1x1 convolution operations to generate a set of intermediate features, calculates attention weights through similarity matching, and combines them to obtain the result.
[0037] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0038] (1) By introducing the ContextGuidedBlock module into the Yolov8 network model, the present invention improves the performance of the model in semantic segmentation.
[0039] (2) By introducing the ACMIX (Self-Attention and Convolution) self-attention mechanism and convolution into the Yolov8 network model, the present invention effectively integrates the self-attention mechanism and convolution, and enhances the segmentation ability of the model for different defects.
[0040] (3) The classification accuracy of the network model proposed by the present invention is relatively high, reaching more than 95%, which can meet the requirements in actual production and can be applied in scenarios with limited computing resources, with strong applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic flowchart of the method for segmenting citrus reticulata blanco cv. ponkan defects based on the improved Yolov8 model of the present invention;
[0042] Figure 2 It is a schematic network structure diagram of the ContextGuidedBlock module of the present invention;
[0043] Figure 3 It is a schematic network structure diagram of the ACMIX module of the present invention;
[0044] Figure 4 It is a schematic network structure diagram of the improved Yolov8 model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0046] Embodiment 1
[0047] As Figure 1 shown, a method for segmenting defects of ponkan oranges based on an improved Yolov8 model includes the following steps:
[0048] S1: Obtain ponkan orange surface image data and perform data augmentation on the ponkan orange surface image data;
[0049] In this embodiment, the ponkan orange surface image data is obtained through offline collection and an online database, and attention should be paid to the balance of the ponkan orange surface image samples to avoid a situation where the number of samples of different categories varies greatly;
[0050] Specifically, ponkan orange images can be obtained online, and relevant ponkan orange image samples can be collected through various search engines and databases;
[0051] Specifically, when obtaining from offline, place the ponkan oranges in a light box. When the ponkan oranges pass through the defect detection production line, use multiple high-precision cameras to take multi-angle photos of them to obtain the required data set, ensuring that the image acquisition conditions for each ponkan orange are the same;
[0052] After collecting the ponkan orange images, perform data augmentation operations on the obtained ponkan orange surface images. Random scaling, inversion, cropping, rotation, optical transformation, etc. can be performed on each ponkan orange image to increase the number of the sample data set, enable the network model to be fully trained, and make the distribution of the sample data set more balanced, avoiding the problem that the model has a tendency during the training process due to an excessive number of samples of certain categories.
[0053] S2: Preprocess the ponkan orange surface image by means of mean filtering to exclude the interference of noise, including methods such as mean filtering and image enhancement;
[0054] In this embodiment, use the method of mean filtering to perform noise reduction processing on the ponkan orange surface image to avoid the influence of noise on image recognition. Specifically, give a template to the target pixel on the image. The template includes its surrounding adjacent pixels (8 pixels surrounding the target pixel form a filtering template, that is, including the target pixel itself), and then use the average value of all pixels in the template to replace the original pixel value.
[0055] For the filtered image, through the method of image enhancement, the contrast between the defective area and the normal area on the surface of the Orah orange is increased, so that the network model can more accurately focus on the area of interest and improve the detection accuracy of the model;
[0056] In this embodiment, an image enhancement operation is performed on the filtered image. Since the piecewise linear transformation is conceptually simple and intuitive and easy to implement, and it can be effectively used to enhance the contrast and adjust the brightness to improve the image quality, the method of piecewise linear transformation is used to make the defective area more prominent, increase the contrast between the defective area and the normal area, so that the network model can more accurately focus on the area of interest.
[0057] S3: By means of random division, the image data set is divided into a training set, a validation set and a test set;
[0058] In this embodiment, the image data set on the surface of the Orah orange is divided into a training set, a validation set and a test set according to the ratio of 3:1:1. The classification labels of the images on the surface of the Orah orange include four labels: lesions, rot, mechanical damage, and fruit stalk. In each data set, the images are saved in different folders according to different categories;
[0059] S4: Establish a basic network model for Orah orange defect segmentation, and then make structural adjustments and optimizations to the basic network model, and add a ContextGuidedBlock module and an ACMIX module;
[0060] As Figure 4 shown, the Orah orange image segmentation network model in this embodiment is improved based on the Yolov8 model. Image represents the input image. The whole network is generally divided into three modules, namely Backbone, Neck, and Head;
[0061] Among them, the Backbone part of the original Yolov8 model consists of five conv (convolution) modules, four C2f modules and one SPPF module. In this embodiment, the C2f module in the backbone part Backbone is replaced with a ContextGuidedBlock module (Con_G module) to improve the accuracy of semantic segmentation and enhance the model's segmentation ability for non-obvious defects;
[0062] As Figure 2As shown in the figure, the Con_G module consists of four parts: a local feature extractor floc(*), a surrounding context extractor fsur(*), a joint feature extractor fjoi(*), and a global context extractor fglo(*). n×n represents the convolution kernel size, convolutions represents discrete convolution, GAP represents average pooling operation, and FC represents a fully connected layer. The design of the ContextGuidedBlock module aims to make full use of local features, surrounding context, and global context. Through this structural design, it is possible to establish a connection between local and global contexts, which is crucial for accurately classifying each pixel in an image. In addition, this module also adopts residual learning to help learn complex features and improve the backpropagation of gradients during training.
[0063] Among them, the conv module contains a convolution with a 3×3 convolution kernel, a regularization function, and a SiLu activation function. The main functions of the convolution module are as follows:
[0064] Downsampling: The convolution layer in each convolution module performs downsampling operations using a convolution kernel with a stride of 2 to reduce the size of the feature map and increase the number of channels;
[0065] Non-linear representation: A Batch Normalization layer and a ReLU activation function are added after each convolution layer to enhance the non-linear representation ability of the model.
[0066] The SPPF module consists of two 3×3 convolutions, three 3×3 max pooling MaxPool, and a fully connected layer, aiming to improve the performance of deep convolutional neural networks in image classification and object detection tasks;
[0067] The Neck part is a key part in the model and plays an important role in feature extraction. The conv module is the same as the conv module in the Backbone part. The internal processing logic of the C2f module is to first perform a 1×1 convolution, and then go through a split process. In this embodiment, after being processed by n DarknetBottleneck modules, the results of the residual module and the backbone module are concatenated by Concat, and then processed by a convolution module for output. Contact represents the concatenation operation, and Unsample represents the upsampling operation;
[0068] Such as Figure 3As shown in the figure, in this embodiment, an ACMIX module is added to the Neck part and connected to the C2f module in the Neck part. The cube in the figure represents the input and output images. C represents the channels in an image, H represents the number of pixels in the vertical dimension of the image, and W represents the number of pixels in the horizontal dimension of the image. The upper part performs a convolution operation, and the lower part performs a self-attention operation. The input features are first converted into queries, keys, and values, which is implemented using a 1x1 convolution, and the attention weights are calculated through similarity matching. Finally, the results are obtained through combination. ACMIX is a hybrid model that combines the advantages of the self-attention mechanism and convolution operations. First, a 1x1 convolution is used to project the input feature map to generate a set of intermediate features, and then these intermediate features are reused and aggregated according to different paradigms, namely the self-attention and convolution methods. In this way, ACMIX can utilize the global perception ability of self-attention and capture local features through convolution, thereby improving the performance of the model while maintaining a low computational cost.
[0069] The Head completes the segmentation of defects by three segmentation nodes segment. The predicted image is detected for defects through the segment segmentation node, and the defects are segmented and then output.
[0070] S5: Set the training parameters, use the training set to train the model, obtain the best weights for verification, and obtain the network model with the best defect segmentation effect.
[0071] In this embodiment, training parameters such as epoch, batch size, and learning rate are set, and the Adam optimizer is used. After a large number of trainings and debuggings, the network model with the best defect segmentation effect is obtained.
[0072] In this embodiment, epoch is set to 300, batch size is set to 16, and the learning rate is set to 0.0002. The cosine annealing algorithm is used to dynamically adjust the learning rate to avoid the oscillation phenomenon caused by too fast gradient descent during training, thereby improving the training stability and generalization ability of the model. Since the Adam optimizer incorporates the concept of momentum, it accumulates the exponentially decaying average of the previous gradients to help accelerate learning. At the same time, it also uses the exponentially decaying average of the squared gradients to adaptively adjust the learning rate of each parameter, which has strong robustness and is widely used in deep learning tasks. Therefore, the Adam optimizer is used for training. After a large number of trainings and debuggings, the network model with the best defect recognition effect is obtained.
[0073] S6: Call the network model to perform defect recognition tests on the test set. Use the classification accuracy rate as the model evaluation criterion to verify the model performance. By comparing the recognition results with the true categories of ponkan defects, it can be detected whether the recognition method has the ability to segment ponkan defects, and the accuracy rate of defect segmentation is output. Finally, the design of the improved ponkan defect segmentation method based on Yolov8 is completed.
[0074] After obtaining a network model with a classification accuracy rate meeting the requirements in this embodiment, it can be deployed to the computer vision system of the ponkan sorting machine. After using the camera to collect ponkan images, the images are used as the input of the network model. According to the output of the model, the categories of ponkan defects can be obtained in real time. Then, the ponkans are classified according to the set ponkan grading standards, and the sorting work of ponkans can be realized in cooperation with the motion control module (such as a robotic arm), effectively alleviating the difficulty of screening and classification in the case of a large number of ponkans, effectively expanding the actual application scenarios, and having the advantages of low cost, small implementation difficulty, strong applicability, and good detection effect.
[0075] Embodiment 2
[0076] This embodiment provides a ponkan defect segmentation system based on an improved Yolov8 model, which is used to implement the ponkan defect segmentation method based on the improved Yolov8 model in the above Embodiment 1. The system includes: an image data acquisition module, a data enhancement module, an image preprocessing module, a data partitioning module, an image segmentation network model construction module, a model training module, a model testing module, and a defect segmentation prediction module;
[0077] In this embodiment, the image data acquisition module is used to acquire ponkan surface image data;
[0078] In this embodiment, the data enhancement module is used to perform data enhancement on the ponkan surface image data;
[0079] In this embodiment, the image preprocessing module is used to perform preprocessing on the ponkan surface image data;
[0080] In this embodiment, the data partitioning module is used to partition the preprocessed ponkan surface image data into a training set, a validation set, and a test set;
[0081] In this embodiment, the image segmentation network model construction module is used to construct a ponkan image segmentation network model based on the Yolov8 model, replace the C2f module in the Backbone part of the Yolov8 model with a ContextGuidedBlock module, add an ACMIX module to the Neck part of the Yolov8 model, and obtain the prediction result of defect segmentation through the Head part of the Yolov8 model;
[0082] In this embodiment, the model training module is used to train the ponkan image segmentation network model based on the training set to obtain the trained ponkan image segmentation network model;
[0083] In this embodiment, the model testing module is used to perform defect recognition testing on the ponkan image segmentation network model based on the test set and output the accuracy of ponkan defect segmentation;
[0084] In this embodiment, the defect segmentation prediction module is used to output the prediction result of defect segmentation based on the trained ponkan image segmentation network model.
[0085] In this embodiment, the data augmentation module is used to perform data augmentation on the ponkan surface image data, specifically including:
[0086] Performing data augmentation operations on the ponkan surface image data, including random scaling, inversion, cropping, rotation, and optical transformation.
[0087] In this embodiment, the image preprocessing module is used to preprocess the ponkan surface image data, specifically including:
[0088] Performing noise reduction processing on the ponkan surface image data based on mean filtering;
[0089] Performing image enhancement operations on the ponkan surface image data after mean filtering, and improving the contrast between the defect area and the normal area based on piecewise linear transformation.
[0090] In this embodiment, the ContextGuidedBlock module includes a local feature extractor, a surrounding context extractor, a joint feature extractor, and a global context extractor. The local feature extractor is used to extract the local features of the input image, the surrounding context extractor is used to extract the surrounding context information of the image, the joint feature extractor is used to fuse the local features and the surrounding context information and output them to the global context extractor, and the global context extractor is used to extract the global context information.
[0091] In this embodiment, the ACMIX module converts the input image features into queries, keys, and values, projects the input image feature map based on 1x1 convolution operation to generate a set of intermediate features, calculates the attention weights through similarity matching, and combines them to obtain the result.
[0092] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and shall be included in the protection scope of the present invention.
Claims
1. A defect segmentation method for mandarin oranges based on an improved Yolov8 model, characterized in that: The steps include: Obtaining surface image data of the mandarin orange, and performing data enhancement on the surface image data of the mandarin orange; Preprocess the surface image data of Wogan; The preprocessed Wogan surface image data are divided into training set, validation set and test set; Based on the Yolov8 model, a network model for Wogan image segmentation is constructed. The C2f module in the Backbone of the Yolov8 model is replaced with the ContextGuidedBlock module. The ACMIX module is added to the Neck part of the Yolov8 model. The prediction results of defect segmentation are obtained through the Head part of the Yolov8 model. The Wogan image segmentation network model is trained based on the training set to obtain a trained Wogan image segmentation network model; Based on the test set, the defect recognition test of the mandarin orange image segmentation network model is carried out, and the accuracy of mandarin orange defect segmentation is output; The prediction results of defect segmentation are output based on the trained Wogan image segmentation network model.
2. The defect segmentation method of mandarin orange based on the improved Yolov8 model according to claim 1, characterized in that: Data enhancement is performed on the surface image data of Wogan, including: Data enhancement operations are performed on the surface image data of the mandarin orange, including random scaling, inversion, cropping, rotation, and optical transformation.
3. The defect segmentation method of mandarin orange based on the improved Yolov8 model according to claim 1, characterized in that: The surface image data of Wogan is preprocessed, including: The surface image data of Wogan was subjected to noise reduction based on mean filtering; Image enhancement operation is performed on the surface image data of Wogan after mean filtering, and the contrast between defective areas and normal areas is improved based on piecewise linear transformation.
4. The defect segmentation method of mandarin orange based on the improved Yolov8 model according to claim 1, characterized in that: The ContextGuidedBlock module includes a local feature extractor, a surrounding context extractor, a joint feature extractor, and a global context extractor. The local feature extractor is used to extract local features of the input image, the surrounding context extractor is used to extract the surrounding context information of the image, the joint feature extractor is used to fuse the local features and the surrounding context information and output them to the global context extractor, and the global context extractor is used to extract global context information.
5. The defect segmentation method of mandarin orange based on the improved Yolov8 model according to claim 1, characterized in that: The ACMIX module converts the input image features into queries, keys and values, projects the input image feature map based on a 1x1 convolution operation, generates a set of intermediate features, calculates the attention weights through similarity matching, and combines them to obtain the results.
6. A defect segmentation system for mandarin oranges based on an improved Yolov8 model, characterized in that: include: Image data acquisition module, data enhancement module, image preprocessing module, data partitioning module, image segmentation network model building module, model training module, model testing module and defect segmentation prediction module; The image data acquisition module is used to acquire the surface image data of the mandarin orange; The data enhancement module is used to perform data enhancement on the surface image data of the mandarin orange; The image preprocessing module is used to preprocess the surface image data of the mandarin orange; The data division module is used to divide the pre-processed surface image data of mandarin oranges into a training set, a verification set and a test set; The image segmentation network model construction module is used to construct the Wogan image segmentation network model based on the Yolov8 model, replace the C2f module in the Backbone of the Yolov8 model with the ContextGuidedBlock module, add the ACMIX module to the Neck part of the Yolov8 model, and obtain the prediction result of defect segmentation through the Head part of the Yolov8 model; The model training module is used to train the mandarin orange image segmentation network model based on the training set to obtain the trained mandarin orange image segmentation network model; The model testing module is used to perform defect recognition test on the mandarin orange image segmentation network model based on the test set, and output the accuracy of mandarin orange defect segmentation; The defect segmentation prediction module is used to output the defect segmentation prediction result based on the trained Wogan image segmentation network model.
7. The citrus defect segmentation system based on the improved Yolov8 model according to claim 6 is characterized in that: The data enhancement module is used to enhance the surface image data of the mandarin orange, and specifically includes: Data enhancement operations are performed on the surface image data of the mandarin orange, including random scaling, inversion, cropping, rotation, and optical transformation.
8. The defect segmentation system of mandarin oranges based on the improved Yolov8 model according to claim 6 is characterized in that: The image preprocessing module is used to preprocess the surface image data of the mandarin orange, specifically including: The surface image data of Wogan was subjected to noise reduction based on mean filtering; Image enhancement operation is performed on the surface image data of Wogan after mean filtering, and the contrast between defective areas and normal areas is improved based on piecewise linear transformation.
9. The defect segmentation system for mandarin oranges based on the improved Yolov8 model according to claim 6, characterized in that: The ContextGuidedBlock module includes a local feature extractor, a surrounding context extractor, a joint feature extractor, and a global context extractor. The local feature extractor is used to extract local features of the input image, the surrounding context extractor is used to extract the surrounding context information of the image, the joint feature extractor is used to fuse the local features and the surrounding context information and output them to the global context extractor, and the global context extractor is used to extract global context information.
10. The defect segmentation system of mandarin oranges based on the improved Yolov8 model according to claim 6, characterized in that: The ACMIX module converts the input image features into queries, keys and values, projects the input image feature map based on a 1x1 convolution operation, generates a set of intermediate features, calculates the attention weights through similarity matching, and combines them to obtain the results.