Anti-attention guava semantic segmentation detection device
By using an anti-attention semantic segmentation detection device, the edge features of the target region are enhanced through feature extraction, pre-segmentation, and feature fusion, solving the segmentation problem in near-color scenes and achieving higher segmentation accuracy and robustness.
Patent Information
- Application Number
- CN202310484606.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2043-04-28
AI Technical Summary
Existing semantic segmentation algorithms struggle to effectively handle near-color scenes where the background and target colors are similar, especially in unstructured environments where fruits like guava are difficult to distinguish from the background. Traditional methods lack robustness, and deep learning models suffer from reduced prediction accuracy under uneven lighting and occlusion conditions.
An anti-attention guava semantic segmentation detection device is adopted. Through feature extraction, pre-segmentation, anti-attention image generation and feature fusion, the edge features of the target region are emphasized. An improved receptive field module and feature fusion decoder are used to enhance the feature representation of the edge region.
It improves semantic segmentation performance in near-color scenes, enhances edge features of target regions, and improves segmentation accuracy and robustness in complex environments.
Smart Images

Figure CN116597439B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of image processing, in particular to an anti-attention Opuntia vulgaris semantic segmentation detection device. BACKGROUND
[0002] Semantic segmentation is one of the important contents in computer vision, and with the rapid development of the technology, related algorithms and technologies are widely used in various fields. For example, semantic segmentation algorithms have been applied to industrial manufacturing, autonomous driving, agricultural production and other fields. For scenes where the background color and the target color are similar, existing semantic segmentation is difficult to segment the image of the near-color scene. SUMMARY
[0003] The purpose of the present disclosure is to provide an anti-attention Opuntia vulgaris semantic segmentation detection device, which aims to solve the above technical problems.
[0004] In order to achieve the above purpose, the first aspect of the present disclosure provides an anti-attention Opuntia vulgaris semantic segmentation detection method, which comprises: performing feature extraction on a to-be-segmented image to obtain a feature image; performing pre-segmentation on the feature image to obtain an initial segmentation image, wherein the initial segmentation image includes an initial predicted target region; erasing a preset region inside the initial predicted target region in the initial segmentation image to obtain an anti-attention image, wherein the edge region of the initial predicted target region is retained in the anti-attention image; and performing feature fusion on the initial segmentation image and the anti-attention image to obtain a semantic segmentation image, wherein the feature of the edge region on the target region in the semantic segmentation image is stronger than that of the edge region on the initial predicted target region in the initial segmentation image.
[0005] Optionally, the feature extraction of the to-be-segmented image to obtain a feature image comprises: inputting the to-be-segmented image into a multi-layer convolution layer layer by layer to obtain image features extracted by each convolution layer in the multi-layer convolution layer, wherein at least two convolution layers in the multi-layer convolution layer are different; and splicing a plurality of image features to obtain a feature image corresponding to each convolution layer in the multi-layer convolution layer.
[0006] Optionally, the multi-layer convolutional layer comprises a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer and a fifth convolutional layer, the first convolutional layer comprises a 1*1 convolution, a 1*3 convolution, a 3*1 convolution and a 3*3 convolution connected in sequence; the second convolutional layer comprises a 1*1 convolution, a 1*1 convolution, a 1*1 convolution and a 3*3 convolution connected in sequence; the third convolutional layer comprises a 1*1 convolution, a 1*1 convolution, a 1*1 convolution and a 3*3 convolution connected in sequence; the fourth convolutional layer comprises a 1*1 convolution; the fifth convolutional layer comprises a 1*1 convolution, wherein the 3*3 convolution in the first convolutional layer has a hole rate of 3; the 3*3 convolution in the second convolutional layer has a hole rate of 5; and the 3*3 convolution in the third convolutional layer has a hole rate of 7.
[0007] Optionally, the pre-segmentation of the feature images to obtain initial segmentation images comprises: performing feature fusion on deep feature images in the plurality of feature images to obtain fused feature images; and performing pre-segmentation on the fused feature images to obtain the initial segmentation images.
[0008] Optionally, the plurality of feature images comprise a first feature image output by a fifth convolutional layer, a second feature image output by a fourth convolutional layer, and a third feature image output by a third convolutional layer, and the feature fusion on deep feature images in the plurality of feature images to obtain fused feature images comprises: performing up-sampling processing on the first feature image to obtain an up-sampled first feature image, and performing up-sampling processing on the second feature image to obtain an up-sampled second feature image; performing point multiplication processing on the up-sampled first feature image and the second feature image to obtain a first point multiplication result; splicing the first point multiplication result and the up-sampled first feature image to obtain a splicing result; performing point multiplication processing on the first point multiplication result, the up-sampled second feature image and the third feature image to obtain a second point multiplication result; and performing feature fusion on the splicing result and the second point multiplication result to obtain the fused feature images.
[0009] Optionally, the feature fusion on the initial segmentation images and the attention inverse image to obtain a semantic segmentation image comprises: obtaining a feature tensor corresponding to the initial segmentation image; segmenting the feature tensor to obtain at least two segmentation tensors; and fusing the attention inverse image and the at least two segmentation tensors to obtain the semantic segmentation image.
[0010] Optionally, the feature fusion comprises third-order feature fusion.
[0011] The second aspect of the present disclosure provides an anti-attention guavas semantic segmentation detection device, the device comprising: an extraction module for feature extraction on an image to be segmented to obtain a feature image; a first segmentation module for pre-segmentation on the feature image to obtain an initial segmentation image, wherein the initial segmentation image comprises an initial predicted target region; an obtaining module for erasing a preset region inside the initial predicted target region in the initial segmentation image to obtain an anti-attention image, wherein the edge region of the initial predicted target region is reserved in the anti-attention image; a second segmentation module for feature fusion on the initial segmentation image and the anti-attention image to obtain a semantic segmentation image, wherein the feature of the edge region on the target region in the semantic segmentation image is stronger than that of the edge region on the initial predicted target region in the initial segmentation image.
[0012] The third aspect of the present disclosure provides a non-transitory computer readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method of the first aspect.
[0013] The fourth aspect of the present disclosure provides an electronic device, comprising: a memory having a computer program stored thereon; and a processor configured to execute the computer program in the memory to implement the steps of the method of the first aspect.
[0014] The anti-attention guavas semantic segmentation detection device provided by the present disclosure performs feature extraction on an image to be segmented to obtain a feature image; performs pre-segmentation on the feature image to obtain an initial segmentation image, wherein the initial segmentation image comprises an initial predicted target region; erases a preset region inside the initial predicted target region in the initial segmentation image to obtain an anti-attention image that retains the edge region of the initial predicted target region; and performs feature fusion on the initial segmentation image and the anti-attention image, i.e., enhances the edge of the initial predicted target region on the anti-attention image through the edge region on the initial segmentation image, to obtain a semantic segmentation image, wherein the feature of the edge region on the target region in the semantic segmentation image is stronger than that of the edge region on the initial predicted target region in the initial segmentation image. Through the above fusion, the feature of the edge region of the obtained semantic segmentation image is more obvious, and the semantic segmentation effect in a near-color scene is improved.
[0015] Other features and advantages of the present disclosure will be described in detail in the following detailed description section. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of the specification, and are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation on the present disclosure. In the drawings:
[0017] Figure 1 is a flow chart of a method for semantic segmentation of guava images according to an example embodiment;
[0018] Figure 2 is a flow chart of sub-steps of step S110 in Figure 1
[0019] Figure 3 is a schematic diagram of a framework of a multi-layer convolutional layer;
[0020] Figure 4 is a flow chart of sub-steps of step S120 in Figure 1
[0021] Figure 5 is a schematic diagram of a framework of a MFD model;
[0022] Figure 6 is an image of guava to be segmented;
[0023] Figure 7 is a feature image obtained by a DeepLab model;
[0024] Figure 8 is a feature image obtained by a HRNet model;
[0025] Figure 9 is a flow chart of sub-steps of step S140 in Figure 1
[0026] Figure 10 is a schematic diagram of a reverse attention module RAM;
[0027] Figure 11 is a heat map of an initial segmentation image and a reverse attention image;
[0028] Figure 12 is a schematic diagram of a Bottleneck output of ResNet;
[0029] Figure 13 is a schematic diagram of a Bottleneck output of Res2Net;
[0030] Figure 14 is a schematic diagram of a semantic segmentation model;
[0031] Figure 15 is a comparison chart before and after adding a reverse attention mechanism;
[0032] Figure 16 is a guava image;
[0033] Figure 17 is a schematic diagram of a to-be-segmented image of guavas;
[0034] Figure 18 is a schematic diagram of a to-be-segmented image of guavas;
[0035] Figure 19 is a rendering map of prediction results of different algorithms;
[0036] Figure 20 a schematic diagram of a semantic segmentation detection device for guavas according to an example embodiment is shown;
[0037] Figure 21 is a block diagram of an electronic device according to an example embodiment. DETAILED DESCRIPTION
[0038] The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present disclosure, and are not used to limit the present disclosure.
[0039] Semantic segmentation is one of the important contents in computer vision. With the rapid development of this technology, related algorithms and technologies have been widely applied in various fields. For example, semantic segmentation algorithms have been applied in industrial manufacturing, autonomous driving, agricultural production and other fields. For scenes where the color of the background is similar to the color of the target, existing semantic segmentation is difficult to segment the image of the near-color scene.
[0040] For example, computer vision technology is applied to agricultural production. The development of computer vision technology has increased the possibility of agricultural automation, and high-quality agricultural scene images and accurate scene understanding are the premise of completing subsequent advanced applications such as yield estimation, quality monitoring, maturity detection, and agricultural product grading. Different from the analysis in a specific environment in the laboratory, most fruits are cultivated in a natural environment, which is a typical unstructured environment. The environmental information of the natural environment is easily disturbed by factors such as temperature, humidity, and weather, and becomes complex and changeable. Various agricultural robots based on visual information processing are expected to work in such an environment. Therefore, the research on agricultural images in unstructured environments is the key to making agricultural intelligent technology practical. The visual information processing task in the planting scene can be divided into two categories according to the fruit phenotype characteristics: non-near-color background fruits and near-color background fruits. Guavas are typical crops with near-color fruits and backgrounds. Non-near-color fruits can be easily distinguished by the color difference between the target and the background, such as apples, oranges, tomatoes, and other fruits in the mature stage. However, for near-color fruits such as guavas, there are the following difficulties when detecting and analyzing in a natural environment:
[0041] (1) The color between the fruit target and the background is basically the same, that is, the canopy composed of new leaves and old leaves, and the fruit is also green, so the leaves and fruits are easily confused, and it is difficult to distinguish them using traditional methods;
[0042] (2) Uneven lighting increases the difference between the same fruit targets;
[0043] (3) When the light intensity is insufficient, the difference between the fruit features and the surrounding objects is further reduced, and the color feature sharing phenomenon will make it more difficult to distinguish the fruit from the background such as new leaves and old leaves.
[0044] (4) Limited by factors such as image acquisition angle and light angle, the shading of branches and near-color leaves on fruits will affect the detection effect.
[0045] It can be seen that the traditional image algorithm relying on limited manual experience to design features and thresholds is difficult to cope with complex changes in unstructured environments and lacks robustness. Since Hinton et al. demonstrated the feature representation capability of deep learning models in computer vision applications, many research results have shown that models can obtain features through downsampling, and with the increase of depth, models can extract more complex and abstract semantic features and obtain prediction capabilities close to humans. However, the performance of deep learning models based on visual information depends on the quality of the collected images, which puts higher requirements on image quality and preprocessing techniques.
[0046] Some scholars have also tried to use deep learning methods to analyze near-color fruits in unstructured environments. Jia Weikuan et al. (2021) proposed a SOLO (Segmenting Objects by Locations) algorithm, which converts the segmentation problem into a location classification problem, and proved that the prediction accuracy is greatly affected under the influence of adverse environmental factors such as night, direct light, and shading. Huang Xiaoyu et al. (2018) proposed a method of assisting near-color fruit segmentation through saliency maps, but the data set has a small amount of data. In the task of visual detection of green citrus in natural environments, research has also shown that deep learning-based image processing algorithms have the potential to complete scene understanding tasks (Han Wen et al., 2020). However, the above-mentioned related tasks of near-color fruits often directly use classical algorithms to train on self-collected data sets, and there are almost no similar cases for improvement for specific tasks, and there is little targeted innovation in the loss function guiding the training of deep models.
[0047] For the semantic segmentation task of guava fruits, the phenotypic characteristics are extremely important factors for determination. The key to improving the inference ability of the model lies in sufficient extraction and utilization of image information, and more attention to the position, edge and overlapping part of the target. However, the current attention mechanism focuses on reassigning the focus to the key area, but not the edge area of the target, which is not conducive to the segmentation of guavas with similar color and shape to the background.
[0048] Therefore, the present disclosure proposes a reverse attention guava semantic segmentation detection device, which implicitly erases the predicted area through a reverse attention map and a channel compression method, thereby guiding the neural network to learn the edge area, and refining the prediction details layer by layer through the reverse attention mechanism in a deep-to-shallow manner. In addition, when using a convolutional neural network for feature extraction, the classification prediction is completely from the context and has nothing to do with other areas due to the limitation of the receptive field mechanism. Therefore, the present disclosure adds an improved receptive field module to the model. In order to better utilize the advantages of high-level semantics in target positioning and quality, the present disclosure uses a feature fusion decoder to reuse the output feature map of the deep convolution.
[0049] The present disclosure provides a reverse attention guava semantic segmentation detection method, which can be applied to Figure 20 the reverse attention guava semantic segmentation detection device 100 shown in the figure, Figure 21 the electronic device 700 shown in the figure and the computer readable storage medium, in this embodiment, the electronic device is taken as an example, which can be a mobile terminal, a computer, a server, etc., please refer to Figure 1 , the reverse attention guava semantic segmentation detection method can include the following steps:
[0050] Step S110, feature extraction is performed on the image to be segmented to obtain a feature image.
[0051] Obtain the image to be segmented. It can be understood that the image to be segmented is an image to be segmented.
[0052] In an embodiment, the local storage location of the electronic device pre-stores a plurality of images, for example, the local storage location is a local album, wherein the plurality of images can be images taken by the camera of the electronic device, or images saved from the application, webpage, browser on the electronic device. The electronic device obtains the image from the local storage location according to the preset path, and the obtained image is used as the image to be segmented.
[0053] In another implementation, the server stores the image to be segmented, and the electronic device is pre-connected to the server. Based on the connection, the electronic device downloads the image to be segmented from the server to obtain the image to be segmented, and saves the image to be segmented locally.
[0054] The image to be segmented is a high-dimensional image containing a large amount of redundant information or sparse original data. Directly segmenting the image to be segmented is difficult and inefficient. Therefore, the image to be segmented is first subjected to feature extraction to obtain a feature image. Feature extraction can be understood as a process of reducing the dimensionality of the image to be segmented, or as a process of mapping the high-dimensional image to be segmented into a low-dimensional feature image. Subsequent processing of the low-dimensional feature image can shorten the processing time and reduce the processing difficulty.
[0055] In an implementation, a preset feature extraction algorithm can be used to extract features from the image to be segmented to obtain a feature image. For example, the preset feature extraction algorithm can be principal component analysis (PCA), singular value decomposition (SVD), linear discriminant analysis (LDA), etc.
[0056] In another implementation, training samples are collected to construct a deep learning model. The deep learning model is trained by the training samples to obtain a trained model. The trained model includes a feature extractor containing convolution. The feature extractor extracts features from the image to be segmented to obtain a feature image output by the feature extractor.
[0057] In step S120, the feature image is pre-segmented to obtain an initial segmentation image, wherein the initial segmentation image includes an initial predicted target region.
[0058] The feature image is preliminarily pre-segmented to obtain an initial segmentation image. The size of the initial segmentation image is consistent with the size of the image to be segmented. The initial segmentation image includes an initial predicted target region and a background region. It can be understood that semantic segmentation is a classification of the initial predicted target region and the background region in the image. The target region and the background region can be painted in different colors to distinguish them. For example, the target region can be painted orange, white, etc., and the background region can be painted blue, black, etc.
[0059] In an embodiment, semantic segmentation can be performed using a semantic segmentation algorithm. For example, a feature image is first identified, and pixels in the entire image are identified. Then, according to the distribution of the pixels on the feature image, a semantic segmentation in an outline mode is used to preliminarily segment the feature image, and the feature image is segmented into an initial predicted target region and a background region to obtain an initial segmentation image.
[0060] In another embodiment, semantic segmentation can be performed using a pre-trained model. The feature image is input into the model, and an initial segmentation image output by the model is obtained.
[0061] In step S130, a preset region inside the initial predicted target region in the initial segmentation image is erased to obtain an anti-attention image, wherein an edge region of the initial predicted target region is retained in the anti-attention image.
[0062] Erasing the preset region inside the initial predicted target region retains more edge regions on the initial predicted target region to obtain the anti-attention image. This facilitates more attention to the edge regions of the target region in the later segmentation, so that the edges of the target region are highlighted in the segmentation task of the near-color scene.
[0063] The anti-attention image can be obtained in the following manner:
[0064] S ra (a,b)=1-Sigmoid(P ori (a,b)) (1)
[0065] In formula (1), (a,b) represents the pixel coordinates in the anti-attention image, S ra (a,b) is the anti-attention image, P ori (a,b) is the initial segmentation image, and Sigmoid is a function.
[0066] This embodiment combines the anti-attention mechanism with the original prediction segmentation method to further highlight the edge regions that are initially ignored in the original prediction, and further distinguishes the target region from the background region.
[0067] In step S140, the feature image and the anti-attention image are fused to obtain a semantic segmentation image, wherein the feature of the edge region on the target region in the semantic segmentation image is stronger than the feature of the edge region on the initial predicted target region in the initial segmentation image.
[0068] The initial segmentation region and the anti-attention image are fused to obtain a semantic segmentation image. It can be understood that the edge region of the target region is emphasized in the anti-attention image, and the edge region of the target region on the anti-attention image is enhanced through the feature fusion of the edge region of the target region on the anti-attention image and the initial segmentation image containing the initial prediction of the target region, so that the feature of the edge region of the obtained semantic segmentation image is more obvious, and the semantic segmentation effect in the near color scene is improved.
[0069] The anti-attention guavas semantic segmentation detection method provided in the embodiment performs feature extraction on a to-be-segmented image to obtain a feature image; performs pre-segmentation on the feature image to obtain an initial segmentation image, wherein the initial segmentation image includes an initial prediction of a target region; erases a preset region inside the initial prediction of the target region in the initial segmentation image to obtain an anti-attention image that retains the edge region of the target region of the initial prediction image; and fuses features of the initial segmentation image and the anti-attention image, that is, the edge of the initial prediction of the target region on the anti-attention image is enhanced through the edge region on the initial segmentation image, to obtain a semantic segmentation image, wherein the feature of the edge region on the target region in the semantic segmentation image is stronger than the feature of the edge region on the initial prediction of the target region in the initial segmentation image. Through the above fusion, the feature of the edge region of the obtained semantic segmentation image is more obvious, the semantic segmentation effect in the near color scene is improved, and after the edge region of the target region is enhanced, the target region can be further highlighted even under the interference of factors such as background region color, light angle, collection angle, and a better semantic segmentation effect is obtained.
[0070] When using a fully connected network for prediction, each feature value of a specific layer of a neural network depends on the overall output of the previous layer, which will produce a large amount of computation and occupy a large amount of storage resources, and a convolutional neural network can reduce such invalid calculation by using a receptive field mechanism. Hubel et al. (1962) found through experiments that after local features are detected by early visual layers of the visual cortex, they are gradually combined into a more complex form through layering and are finally processed by the visual information system. Based on this principle, the receptive field mechanism is proposed. In the first convolutional layer, the receptive field and the effective receptive field are the same, but as the convolutional layers deepen, the two begin to differ. The effective receptive field is the region in the original image that affects the activation of the neuron, and can reflect the degree of influence of the input signal on the activation of the neuron, and the receptive field can be simply equivalent to the size of the filter of the previous layer.
[0071] In the calculation of a convolutional neural network considering a receptive field, it is assumed that the size of a feature map input into an lth convolutional neural network is H l ×W l ×C l, the filter size is k x k, then each neuron of the lth convolutional layer will be connected to a k x k x C l region, and k x k x C l weight values and a bias term b need to be learned. The receptive field is a three-dimensional tensor whose depth is equal to the channel dimension of the previous convolutional layer, but this factor is usually ignored in order to focus more on the spatial information of the image. The feature value of each unit of the convolutional neural network depends on a certain region of the output feature map of the previous layer, and this region is called the receptive field of the unit.
[0072] Since any place outside the receptive field cannot affect the feature value, it is crucial for image understanding to ensure that the receptive field covers the relevant image region, especially in dense prediction tasks such as semantic segmentation, in order to make reasonable predictions for each pixel, sufficient context information should be provided as much as possible. There are various ways to enhance the receptive field: Spatial Pyramid Pooling (SPP) increases the receptive field without changing the convolution kernel by using pooling layers of different sizes, and PSPNet takes advantage of this to realize scene understanding based on multi-scale feature fusion (Zhao et al., 2017), and ASPP improves the way of calculating the convolution kernel, directly increases the receptive field by adjusting the hole rate to obtain multi-scale target information (Chen et al., 2017). Liu et al. (2018) defined a method to measure the effective receptive field and designed a Dense Global Context Module to make the effective receptive field cover a larger area with higher density while reducing the parameters. The rest of the research on expanding the size of the receptive field mostly focuses on adjusting the convolution parameters and changing the network structure, and from the application level, they show the improvement of the prediction ability after the improvement in the open source dataset. ResNet50 used by Backbone is also one of the methods, which realizes downsampling by convolutional layer stacking while linearly increasing the size of the receptive field. In the process of mapping from image space to high-level feature space, even the same size of the convolution kernel will increase the receptive field by several times because of the smaller size of the feature map.
[0073] In the near-field scene such as the guava fruit image scene understanding task, it is tried to improve the performance of the model by improving the receptive field. For example, Ilyas et al. (2021) proposed an adaptive receptive field module to dynamically change the receptive field size of neurons, which helps better multi-scale context aggregation. The addition of this module in the segmentation task of strawberry fruit increases the mIoU by 0.0384. For another example, Parico et al. (2021) found that the reason why YOLOv4 and YOLOv4-CSP have higher classification accuracy in pear target detection is that the receptive field is improved. Jia et al. (2022) realized the segmentation of large-size targets by stacking standard convolution layers and combining an extended convolution module for specific feature maps. And Roy et al. (2022) directly enhanced the detection effect of YOLOv4 model on plant diseases through spatial pyramid pooling (SPP). And these existing fruit image processing methods rarely consider multi-proportional receptive field, and most of them directly adjust the convolution parameters instead of making the receptive field become a separate module embedded in the neural network. When observing and picking guava fruits in the natural environment, as the distance between the fruit and the image acquisition device changes, the proportion of the fruit as a whole in the image and the length-width ratio of the fruit itself present diversification, therefore, the present application specifically proposes a separate receptive field module, i.e. the following multi-layer convolution layer.
[0074] In another embodiment, referring to Figure 2 , the above step S110 can include the following steps:
[0075] Step S111, input the image to be segmented into the multi-layer convolution layer layer by layer, and obtain the image features extracted by each convolution layer in the multi-layer convolution layer, wherein at least two convolution layers in the multi-layer convolution layer are different.
[0076] The multi-layer convolution layer is arranged in order according to the level, and each convolution layer includes at least one convolution. The at least two convolution layers in the multi-layer convolution layer are different, which can be understood as that the number of convolutions of the at least two convolution layers is different, or the size of the convolutions is different. Alternatively, the convolution can be a dilated convolution. The dilated convolution mainly increases the receptive field in semantic segmentation.
[0077] For example, taking a 5-layer convolution layer as an example, the multi-layer convolution layer designed by the present embodiment for the application scene of guava is as shown in Figure 3 , which includes a first convolution layer, a second convolution layer, a third convolution layer, a fourth convolution layer and a fifth convolution layer connected in order, and the convolution in each convolution layer can be represented by , wherein j represents the row in the multi-layer convolution layer, and i represents the column in the multi-layer convolution layer. The first convolution layer includes 1x1 convolutions connected in order 1x3 convolution 3x1 convolution and 3x3 convolution The second convolutional layer comprises 1x1 convolution 1x1 convolution 1x1 convolution and 3x3 convolution The third convolutional layer comprises 1x1 convolution 1x1 convolution 1x1 convolution and 3x3 convolution The fourth convolutional layer comprises 1x1 convolution The fifth convolutional layer comprises 1x1 convolution The 3x3 convolution included in the first convolutional layer has a dilation rate of 3; the 3x3 convolution included in the second convolutional layer has a dilation rate of 5; and the 3x3 convolution included in the third convolutional layer has a dilation rate of 7.
[0078] Optionally, the 3x3 convolution included in the first convolutional layer, the 3x3 convolution included in the second convolutional layer, and the 3x3 convolution included in the third convolutional layer are all dilated convolutions. The significance of introducing dilated convolution is to obtain a larger receptive field area with the same amount of convolution parameters, and to capture more context information while generating higher resolution feature images.
[0079] In the multi-layer convolutional layer in the embodiment, the use of different sizes of convolution to calculate the multi-size change of the receptive field is the basic idea of the receptive field enhancement in the embodiment. For guavas in the image horizontally, a 3x1 convolution with a length greater than a width is set. For guavas in the image vertically, a 1x3 convolution with a width greater than a length is set. In the prior art, a convolution with the same length and width is usually used, for example, a 3x3 convolution is used. However, the 3x3 convolution increases the amount of calculation, and in the embodiment, a convolution with different length and width, i.e., a 1x3 convolution and a 3x1 convolution, is set according to the position of the guavas in the image, which can not only realize the processing of the guavas in the image, but also save the amount of calculation, reduce the calculation time, and reduce the parameters and increase the non-linear relationship.
[0080] continue to combine Figure 3 The to-be-segmented image is copied into 5 parts, and the 5 parts of the to-be-segmented image are respectively input into the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer. From the first convolutional layer to the fifth convolutional layer, the level gradually deepens. Figure 3 The five convolutional layers in the figure are not the same, and the purpose is to increase the number of convolutional layers while giving different size ratios of the receptive field. Each layer in the five convolutional layers is processed separately, and in each convolutional layer, the convolution is processed in turn according to the arrangement order, and the image features extracted by each layer of the multi-layer convolutional layer are obtained. For example, the fifth feature image output by the first convolutional layer is Layer1, the size is 256x256x64, the fourth feature image output by the second convolutional layer is Layer2, the size is 128x128x256, the third feature image output by the third convolutional layer is Layer3, the size is 64x64x512, the second feature image output by the fourth convolutional layer is Layer4, the size is 32x32x1024, and the first feature image output by the fifth convolutional layer is Layer5, the size is 16x16x2048. Among them, the feature images output by the first convolutional layer to the fifth convolutional layer become deeper layer by layer.
[0081] Step S112, the plurality of image features are spliced to obtain the feature image corresponding to each layer of the multi-layer convolutional layer.
[0082] For example, please continue to refer to Figure 3 , the output feature image is Among them, the in the figure is the con calculation, is + in the above formula.
[0083] Feature fusion is a common method to improve semantic segmentation effect, and the purpose is to extract and integrate input features from different sources, and finally combine a more discriminative and stronger feature to improve the prediction performance of the model. In the convolutional neural network, the high-level features of low resolution represent semantic information, and the low-level features of high resolution represent spatial details. Many methods try to fuse feature maps of different levels, typical methods such as Chen et al. (2017) proposed ASPP module in DeepLab to fuse multi-scale features to process objects of different sizes, Zhao et al. (2017) proposed pyramid pooling in PSPNet, encoder and decoder feature splicing and channel compression used in U-Net by Ronneberger et al. (2015) and so on.
[0084] There are two classical feature fusion methods, which are splicing and matrix addition, if two input features are represented as and then the new feature after splicing fusion can be represented as If matrix addition is used, it can be represented as R c= R1+ R2. According to the operation process, it can be divided into early fusion (typical representatives are FCN, ParseNet, HyperNet, etc.) which fuses features first and then predicts, and late fusion (such as SSD and MS-CNN) which reverses the order and combines the prediction results of different layers to improve the prediction ability.
[0085] In the semantic understanding task based on deep convolutional neural network, low-level features consume more computational resources than high-level features but have less effect on performance improvement. Poudel et al. (2018) found that using only the output of the shallow neural network for segmentation prediction paid too much attention to details and was ineffective due to the limitation of the receptive field, which could not provide sufficient context information and had a significant false positive rate. At the same time, Wu et al. (2019) analyzed the DSS model proposed by Hou et al. (2017) on the PASCAL-S dataset and found that as the number of connections increased, the performance gradually increased, but the improvement effect gradually slowed down, and the inference time significantly increased, which was not conducive to real-time application. Wu et al. (2019) also found that when using a convolutional neural network to realize saliency detection, integrating deep feature maps can obtain more accurate saliency maps, effectively suppress interfering features, and improve the expression ability of features. Therefore, in order to realize the semantic segmentation task more quickly and accurately, this study proposes a feature fusion decoder for high-level semantics (Mutual Fusion Decoder, MFD), which aims to discard shallow features to ensure high computational efficiency and refine deep features to improve feature representation ability.
[0086] In another embodiment, referring to Figure 4 , step S120 includes the following steps:
[0087] Step S121, fusing the deep feature images in the plurality of feature images to obtain a fused feature image.
[0088] In the multi-layer convolutional layer, from the shallow convolutional layer to the deep convolutional layer, the output deep feature image becomes deeper layer by layer. For example, continuing with Figure 3 , feature image Layer1 and feature image Layer2 can be considered as shallow feature images. Shallow feature images have high resolution and many details, but if shallow feature images are used for prediction, it is difficult to locate and qualify the target area because the receptive field mechanism limits the context area available for prediction for a single pixel. Feature image Layer3, feature image Layer4, and feature image Layer5 can be considered as deep feature images. Deep feature images have a large receptive field and pay more attention to global information. Fusing multiple deep feature images improves the reuse of deep information.
[0089] One approach is to fuse multiple deep feature images, namely feature image Layer 3, feature image Layer 4, and feature image Layer 5, using a trained MFD model. Figure 5 As shown, the deepest layer's first feature image, Layer 5, is upsampled, for example, as... Figure 5 As shown, upsampling is performed in both branches using bilinear interpolation combined with a 3×3 convolution to obtain the upsampled first feature image. One branch serves as a short-circuit connection, and the other branch is used for subsequent dot product with the feature image Layer4. The second feature image is then upsampled, for example, by using bilinear interpolation combined with a 3×3 convolution to obtain the upsampled second feature image. A dot product is then performed between the upsampled first and second feature images to obtain the first dot product result. Figure 5 In This represents the dot product process. The first dot product result and the upsampled first feature image are concatenated to obtain a concatenated result. The first dot product result, the upsampled second feature image, and the third feature image are then multiplied to obtain a second dot product result. The concatenated result and the second dot product result are then fused to obtain the fused feature image. In each of the above dot product and concatenation processes, feature enhancement is performed on the processed image.
[0090] Optionally, the fused feature images can be channel compressed using 3×3 convolution to enhance the nonlinear relationship between feature values at the same location on different feature images.
[0091] Step S122: Pre-segment the fused feature image to obtain the initial segmented image.
[0092] In one implementation, the fused feature image is upsampled using bilinear interpolation combined with three sequentially connected 3×3 convolutions to obtain an initial segmentation image P. ori .
[0093] The attention mechanism in neural networks originates from the special way of human visual information processing, that is, selectively focusing on important information and ignoring other visible information. Correspondingly, this mechanism is expected to tilt the computing resources to more important areas and reduce the attention to other information to prevent information overload under limited computing power. The research of attention mechanism can be traced back to 1980, and Mnih et al. (2014) formally applied it to the natural language processing field of neural networks and made significant progress. Subsequently, more attention structures appeared and were applied in the field of computer vision. Among them, the classic attention mechanism model includes SE-Net, which shows the relationship between feature maps, and guides the update of feature maps through squeezing and excitation operations. The later ECA-Net improves SE-Net by adding a cross-channel information interaction mechanism through one-dimensional convolution. SK-Net uses Split, Fuse, and Select to synthesize attention maps in a multi-path way. CBAM not only uses the channel attention mechanism like the above attention models, but also claims that the feature maps inside the channel also hide attention information worth using, so CBAM constructs two modules connected in parallel to focus on spatial attention and channel attention. In the Dual Attention Network, the spatial and channel attention modules are connected in parallel, and the attention map reflecting global dependency is obtained by replacing the pooling operation with matrix reconstruction, transposition, and multiplication operations Guo et al. (2022). Semantic segmentation is a pixel-level dense prediction task in scene understanding, and studies such as DANet, Acnet, and CCNet have also shown that attention mechanisms can enhance the prediction ability of models.
[0094] However, most current semantic segmentation methods focus on feature extraction and prediction of the target, and even the attention mechanism does not focus on the distinction between targets. In the scene of guava fruit segmentation applied in unstructured environments, as shown in Figure 6 , the difficulty lies in the confusion between foreground and background due to similar color, texture and other features. In fact, the color of the leaf and the color of the guava are both green, and the shape is oval. The feature image obtained by the DeepLab model is shown in Figure 7 , and the feature image obtained by the HRNet model is shown in Figure 8 . After visualizing the feature images obtained by the DeepLab model and the HRNet model, it can be found that visual similarity may lead to high-level semantic features being shared by the target and its background, and the network model produces almost equally strong responses to the background area near the guava fruit, as shown in Figure 7 and Figure 8The white box is shown. Therefore, the anti-attention guava semantic segmentation detection method provided by the present disclosure refines the anti-attention in stages, and implicitly erases the predicted area by combining the feature map embedding method, thereby strengthening the edge area and guiding the network to explore the missing part, thereby obtaining a clearer edge area.
[0095] In an embodiment, in combination with the anti-attention module, the input tensor embeds the anti-attention map, and the attention information is fused by dimension restoration using channel compression. In this process, the predicted preset area is implicitly erased, and the model is guided to pay attention to the edge area. Please refer to Figure 9 , step S140 includes the following steps:
[0096] Step S141, obtaining the feature tensor corresponding to the initial segmentation image.
[0097] The size of the feature tensor can be HxWxC in .
[0098] Step S142, segmenting the feature tensor to obtain at least two segmentation tensors.
[0099] The feature tensor is segmented and truncated using the torch.chunk(tensor, 1, dim=0) function to obtain at least two segmentation tensors. For example, the feature tensor is segmented and truncated from the middle by the above function to obtain two independent tensors with the same size, which are and and C1+C2=C in .
[0100] Step S143, fusing the anti-attention image with the at least two segmentation tensors to obtain the semantic segmentation image.
[0101] As a way, the anti-attention image is inserted into the at least two segmentation tensors for fusion to predict the semantic segmentation image.
[0102] Optionally, the above feature fusion can include third-order feature fusion, which includes low-order fusion, middle-order fusion and high-order fusion. The anti-attention module is a recurrent attention model (RAM), as shown in Figure 10 The anti-attention module includes a low-order fusion module, a middle-order fusion module and a high-order fusion module.
[0103] Wherein, in the low-order fusion stage, the anti-attention image S ra (a,b) is obtained by formula (1), as shown in Figure 11As shown, the anti-attention image better represents edge regions compared to the initial segmentation image. The feature tensor corresponding to the initial segmentation image is obtained; the feature tensor has dimensions H×W×C. in . H×W×C in The feature tensor is truncated in the middle using the `torch.chunk(tensor, 1, dim=0)` function, resulting in two tensors of the same size. and And C1 + C2 = C in The truncated anti-attention image S ra After embedding the cut-off points (a,b), the three parts are spliced together to obtain a size of H×W×(C). in +1) tensor, then use a 3×3 convolution to combine H×W×(C) in +1) Compress the number of channels in the tensor to obtain a size of H×W×C. in The tensor, the image obtained in the low-order fusion stage is p l The final image p obtained in the low-order fusion stage l It can be represented as:
[0104]
[0105] Among them, P ori For the initial segmented image, This is the tensor after the segmentation.
[0106] The reverse attention module ultimately outputs a feature map. It can be represented as:
[0107]
[0108] It can be seen that p l P is implicitly controlled through backattention maps and channel compression. ori Enhancements have been made.
[0109] The principle of the intermediate-order fusion stage is similar to that of the low-order fusion truncation stage. In the intermediate-order fusion stage, the final H×W×C obtained in the low-order fusion stage is... in The tensor is truncated in three places to obtain four tensors. At the end of each of the four tensors, H×W×(C) obtained from the low-order fusion stage is inserted. in +1) Tensor, then concatenate the four tensors and the inserted tensor to obtain a size of H×W×(C). in +4) tensors, H×W×(C) are convolved by 3×3. in +4) Compress the number of channels in the tensor to obtain a size of H×W×C. in The tensor, the image obtained in the intermediate fusion stage is p m .
[0110] In the high-order fusion stage, the HxWxC in tensor obtained in the middle-order fusion stage is truncated in three places to obtain an eight-part tensor, and the HxWxC in +4) tensor obtained by splicing in the middle-order fusion stage is inserted at the end of the eight-part tensor. Then, the eight-part tensor and the inserted tensor are spliced to obtain an HxWxC in +32) tensor, the channel number of the HxWxC in +32) tensor is compressed through 3x3 convolution to obtain an HxWxC in tensor, and the image obtained in the high-order fusion stage is p h .
[0111] Optionally, the design aspect of the present disclosure uses a convolutional encoder-decoder structure, and the original Backbone mode is as shown in Figure 12 . The Backbone used in the down-sampling of the encoder is Res2Net, a multi-scale neural network structure improved based on ResNet. It divides the output of the 1x1 convolution layer in the Bottleneck of ResNet into four groups, x1, x2, x3 and x4, replaces the original 3x3 convolution layer in Figure 12 with a multi-segment 3x3 convolution operation, and the specific operation is as shown in Figure 13 . x1 is directly used as the first segment y1 of the output, while x2, x3 and x4 are calculated through 3x3 convolution to obtain y2, y3 and y4 respectively, and then spliced with y1 in order to obtain the Bottleneck output. Among them, the input of convolution layers K3 and K4 is the addition of the output of the previous convolution and the input of the current segment. Finally, the increase of the convolution kernel will bring the increase of the receptive field, and Res2Net will obtain an output y with expanded feature scale due to the combination effect.
[0112] Optionally, to perform semantic segmentation on the image to be segmented, a multi-layer convolution layer as shown in Figure 3 , an MFD model framework as shown in Figure 5 , and a reverse attention module RAM as shown in Figure 10 may be pre-trained. A semantic segmentation model including a multi-layer convolution layer as shown in Figure 3 , a receptive field module framework, and a reverse attention module RAM as shown in Figure 10 may also be pre-trained. The receptive field module can include a feature extraction module (Receptive Field Block, RFB for short) and an MFD module, wherein the RFB module can perform feature fusion, and the MFD module is used for preliminary prediction. The semantic segmentation model can be as shown in Figure 14As shown, Res2Net as an encoder uses 5 Bottleneck to finally downsample the image to be segmented with size 512x512x3 to feature image with size 16x16x2048. Figure 14 The working principle of the three RAM modules in Figure 10 is shown.
[0113] Wherein the feature maps output by Layer 3, Layer 4 and Layer 5 will all be input to the corresponding RFB of the receptive field module, so as to enhance the size of the receptive field and the multi-scale feature extraction capability, and the three layers of feature maps after enhancement will be input to the feature fusion decoder for multiplexing high-level semantics and thus obtain a rough initial segmentation image P ori .
[0114] The semantic segmentation task based on deep learning often focuses on predicting the target pixel by pixel, even if the attention mechanism is used to reassign features, the focus is still concentrated on predicting the high-frequency response area. This makes it difficult for the model to actively focus on the distinction between targets and backgrounds without additional supervision. The guava fruit and the canopy have visual similarities. In unstructured environments, due to factors such as angle and light, the boundary area shares high-level features, making it more difficult to distinguish. Therefore, the model focuses on the distinction between target areas and background areas.
[0115] Deep features can better locate and qualify multi-scale guava targets, but as the number of layers in the Backbone increases, it is difficult to avoid losing some details originally preserved in shallow feature maps, which affects prediction accuracy. Therefore, this study selects the method of generating a rough prediction segmentation image using high-level feature maps, then implicitly erasing the predicted area using the reverse attention map, and finally using shallow feature maps to refine the details to realize the prediction area guidance from deep to shallow. In specific implementation, the rough initial segmentation image P ori generated directly by MFD and the Layer 5 feature map enhanced by RFB will be input to the reverse attention module RAM_L5, the initial segmentation image P ori After three-order fusion operation in RAM_L5, the current layer prediction O l5 is obtained, which is input to RAM_L4 as the prediction guidance map of RAM_L4 to obtain P l4 Through formula (1). Similarly, after implementing the same operation in the Layer 3 feature map, the model can refine P ori to the final prediction segmentation image P l3 after sufficient detail exploration, P l3 is the final guava fruit semantic segmentation image, that is, the image.
[0116] To train Figure 13 The semantic segmentation module shown in the figure, in the semantic segmentation data set designed in the present disclosure, there are a total of 1000 images, collected in a certain farm planting guavas, the guavas photographed are of the red guava variety. The collection device is a Nikon D5300 camera, the shooting resolution is 2992x2000 pixels, and before inputting the neural network, it is cropped and adjusted to 512x512 pixels. In the collection and production of the data set, the present disclosure fully considers the particularity of guava fruits in unstructured environments. When shooting, the distance from the fruit is 0.5-3 meters, and the lighting environment, canopy density, fruit angle and shielding degree are all different, so detailed manual annotation is required, and the annotation results are as shown in the figure. Figure 15
[0117] The model in the disclosure is built using PyTorch v1.7.1, and the CUDA version is 11.1. The training device is a GeForce RTX 3090ti graphics card with 24G of video memory. The optimizer used is Adam, the learning rate is set to 0.001, and the momentum parameters are set to 0.9 and 0.999. In this study, each comparative experimental model was trained for 200 epochs. It is worth noting that in order to prevent the segmentation model from being disturbed by meaningless discrimination results produced by the discriminator after random initialization in the initial stage, this study chooses to first freeze the discriminator and let the segmentation model train alone in the training set with a batch size of 1 for 500 iterations. Subsequently, the segmentation model is frozen, and the discriminator is also trained alone for 200 iterations using the combined input method introduced in Section 5.3.3, and then the subsequent training is completed using the alternating iteration training strategy of the generative adversarial network.
[0118] As follows, the semantic segmentation model provided by the present disclosure is trained in the following manner: In terms of learning rate planning, this study uses the Cosine Annealing Warm Restarts strategy Loshchilov et al. (2016), which can be expressed by the formula:
[0119]
[0120] Where η t is the current learning rate, η min and η max are the learning rate range, T cur represents the current number of epochs, T i represents how many epochs to start a restart, and T cur represents how many epochs have passed since the last restart, so when T cur = T i ηt = η min , when T cur = 0, η t = η max In a specific application, T_0 and T_mult are set to 50 and 2, respectively.
[0121] In addition, the present disclosure also conducts ablation experiments. In the ablation experiments in this part, we use the method of sequentially adding receptive field block (RFB), multi-feature decoder (MFD) and reverse attention module (RAM) to verify the promotion of each module to the inference ability of the model. In terms of evaluation criteria, in order to more fairly evaluate the performance of the model, five evaluation indexes including accuracy (Acc), precision (PC), recall (Recall), F1 and intersection over union (IoU) are selected for comprehensive evaluation. The statistical results are shown in Table 1:
[0122] Table 1
[0123] RFB MFD RAM A cc ]]> PC Recall F1 IoU × × × 0.9725 0.8956 0.8268 0.8652 0.7668 √ × × 0.9856 0.9332 0.9322 0.9252 0.8805 × √ × 0.9875 0.9346 0.9435 0.9321 0.8965 × × √ 0.9823 0.9321 0.9286 0.9212 0.8755 √ √ × 0.9924 0.9609 0.9622 0.9614 0.9261 √ √ √ 0.9939 0.9702 0.9678 0.9690 0.9401
[0124] As can be seen from the above table, after adding RFB, MFD and RAM modules respectively, the inference ability of the model is significantly improved. Among them, the maximum improvement is brought by adding MFD alone. The Recall is improved by 0.1167, the F1 is improved by 0.0669, and the IoU is improved by 0.1297. It can be seen that the three modules promote the fitting of the predicted region and the label region. However, the improvement in Acc is not obvious, only 0.0098 to 0.0150. In order to better reflect the improvement brought by the reverse attention mechanism, this experiment statistics the changes of evaluation indexes before and after the application of RAM (under the condition that RFB and MFD are added). After the application, the model achieves the best performance in the five evaluation indexes. Among them, the Acc reaches 0.9939, which is improved by 0.0214 compared with the original model and by 0.0015 compared with before the application. In PC, it is improved by 0.0093. It can be seen that the proportion of accurate positive samples is improved. The IoU is also significantly improved, reaching 1.4. It can be seen that the RAM module is beneficial to the refinement of the predicted region of the model.
[0125] Table 2 is the ablation experiment result of the reverse attention mechanism. As shown in Table 2:
[0126] Table 2
[0127] A cc ]]> PC Recall F1 IoU P ori ]]> 0.9924 0.9609 0.9622 0.9614 0.9261 P l5 (RAM_L5)]]> 0.9905 0.9587 0.9450 0.9518 0.9086 P l4 (RAM_L4)]]> 0.9933 0.9672 0.9647 0.9659 0.9344 P l3 (RAM_L3)]]> 0.9939 0.9702 0.9678 0.9690 0.9401
[0128] In order to further illustrate the role of the reverse attention module, this study conducts an ablation experiment on the RAM addition strategy. Table 2 corresponds Figure 7 The prediction results output by the RAM_L5 to RAM_L3 modules are recorded respectively. Overall, even if the Pori i has already achieved relatively reasonable inference performance, even reaching 0.9261 in IoU, and RAM can still further improve prediction performance. Although the P output of RAM_L5 is affected by the implicit erasure of some prediction information by the anti-attention module, l5 Performance compared to P ori There was a slight regression, but the prediction guidance strategy designed in this study, which proceeded from deep to shallow, showed improvement in P. l4 This began to show in performance, exceeding P in all five evaluation indicators. ori Furthermore, it improved by 0.0083 in terms of IoU. (This is attributed to P...) l4 Guided by the predicted segmentation map P generated by RAM_L3 l3 (i.e., semantic segmentation images) further demonstrate the effectiveness of the prediction guidance strategy, achieving the best performance across all evaluation metrics. It reaches 0.9939, 0.9702, 0.9678, 0.9690, and 0.9401 in Acc, PC, Recall, F1, and IoU, respectively, significantly outperforming the second-best performing P. l4 There was an improvement of 0.0006 to 0.0057.
[0129] To more intuitively demonstrate the effectiveness of the anti-attention mechanism in distinguishing guava fruits from the background and fruits from each other, this disclosure uses two sets of feature maps before and after the RAM_L4 module to generate a heatmap. Figure 15 In (a), the image before the anti-attention mechanism was added is not clear at the fruit edge, and there are even some interfering responses at some boundaries. In terms of distribution, the response distribution in the fruit area is uneven, and there is no trend of more concentrated high-frequency responses inside the fruit. Instead, in some feature images, high-frequency responses appear in the background and boundary areas. Figure 15 (b) is the graph after adding the anti-attention mechanism. The high-frequency responses are all distributed within the guava fruit and have a concentrated trend. The boundary area near the fruit is clear and there are no interfering high-frequency responses. There are no high-frequency responses in the background areas such as leaves and trunk at the boundary of a single fruit.
[0130] To intuitively demonstrate the improvement in model prediction ability brought about by the anti-attention mechanism, such as Figure 16 As shown, where Figure 16 In the image (a), the image to be segmented is... Figure 16 (b) is the initial segmented image. Figure 16 (c) is the semantic segmentation image. It can be seen that after layer-by-layer refinement, the fruit prediction effect is improved, the boundary between the fruit and the green background is clearer and neater, the missed detection phenomenon in the middle of the fruit is reduced, and the interference of light pollution on fruit prediction is also alleviated.
[0131] To more intuitively demonstrate the reasoning ability of the anti-attention-based semantic segmentation algorithm for the guava semantic segmentation task, this section will... l3 P l5 P ori It was compared with state-of-the-art algorithms including SINet, BiSeNet, and SETR, and also with AttSegGAN proposed in Chapter 5. The prediction statistics are shown in Table 3. It can be seen that P... l3 The prediction results achieved the best performance across all metrics, with an IoU of 0.0069, surpassing the second-place AttSegGAN. This indicates a better fit between the predicted and labeled regions, and more accurate prediction of boundary details. The comparison of model prediction capabilities is shown in Table 3.
[0132] Table 3
[0133] Acc PC Recall F1 IoU ours_P l3 ]]> 0.9939 0.9702 0.9678 0.9690 0.9401 ours_P l5 ]]> 0.9905 0.9587 0.9450 0.9518 0.9086 ours_P ori ]]> 0.9924 0.9609 0.9622 0.9614 0.9261 AttSegGAN 0.9938 0.9620 0.9670 0.9635 0.9332 U-Net 0.9926 0.9620 0.9631 0.9623 0.9279 SINet 0.9919 0.9612 0.9565 0.9586 0.9211 DeepLabv3+ 0.9912 0.9617 0.9495 0.9553 0.9150 HRNet+OCR 0.9885 0.9458 0.9381 0.9414 0.8903 BiSeNet 0.9910 0.9581 0.9506 0.9542 0.9129 SETR 0.9718 0.8826 0.8157 0.8452 0.7398
[0134] Regarding the intuitive demonstration of prediction results, this experiment selected guava images from the test set that were significantly affected by unstructured environments for prediction and display. The displayed images could be... Figure 17 The image shows overlapping fruits, obscured by green leaves, and instances where the edges of the fruits share color with the leaves. It could also be... Figure 18 The images shown are affected by factors such as uneven lighting and low brightness. In these images, the interference of color sharing and visual similarity is more obvious.
[0135] For the comparison algorithms, the state-of-the-art algorithm, U-Net algorithm, SINet algorithm, and BiSeNet algorithm, which performed well in Table 3, were selected and compared with the final prediction result P of this model. l3 The comparison was performed and a rendering was generated. In the rendering, green represents true positive (TP) pixels; transparent represents true negative (TN) pixels; red represents false positive (FP) pixels; and blue represents false negative (FN) pixels. Figure 19 (a) in the image is a rendered image generated by the U-Net algorithm. Figure 19 (b) is the rendered image corresponding to the SINet algorithm. Figure 19 (c) is the rendered image corresponding to the BiSeNet algorithm. Figure 19 (d) is the prediction result P of this disclosure. l3 .from Figure 19 As can be seen, the model's predictions are more accurate in the easily confused areas marked with yellow boxes. At the boundary where green leaves overlap with fruit, the model's predicted regions have clearer edges and fewer false positives and false negatives. In low-light images, the model identifies more true positive fruit areas and accurately distinguishes the fruit from the background even in significantly degraded brightness.
[0136] In the case of target interference such as branches, the model can more clearly avoid small backgrounds, and there are fewer false positive and false negative predictions in the relevant area, which benefits from the refinement of the attention mechanism. In terms of uneven illumination interference, the four algorithms accurately identify the fruits, which proves the robustness of the deep learning-based model of the present disclosure in unstructured environments, but the prediction of the present model is more clear at the boundary between the light and dark intersection and the fruit boundary, and there are more true positive areas.
[0137] The semantic segmentation algorithm provided by the present disclosure first, in the receptive field module design, in order to better obtain fruits of different scales in the image, the present research designs a receptive field module according to the specific situation of guavas in the image, and the proportion of the convolution kernel and the hollow rate (the meaning of the hollow convolution is to obtain a larger receptive field area with the same parameter amount of the convolution kernel) are selected according to the actual situation of the fruit. Compared with other receptive field enhancement mechanisms, the receptive module proposed in the present research is more direct and can be embedded in the network in a modular way. Secondly, in the feature fusion decoder, the main purpose is to enhance the reuse of high-level semantic information. Because the feature map output by the shallow neural network has high resolution and many details, it is difficult to predict the target using the feature map output by the deep neural network, because the receptive field mechanism limits the context area available for prediction of a single pixel. The present model improves the reuse of deep information by fusing the deep feature map of Layer 5 with Layer 4 and Layer 3, and refines the features using the Layer 4 and Layer 3 feature maps, while not using the output of Layer 2 and Layer 1 convolution layers. Finally, in the anti-attention module, for the guava fruit semantic segmentation task, the phenotype feature is the key factor for determination, and the key to improve the reasoning ability of the model lies in the sufficient extraction and utilization of image information, and more attention to the position, edge and overlapping part of the target. However, the current attention mechanism focuses on reassigning the focus to the key area, rather than the edge area of the target, which is not conducive to the segmentation of guavas with similar fruits and backgrounds. However, the present research proposes an anti-attention guava semantic segmentation detection method based on the anti-attention mechanism, which implicitly erases the predicted area through the anti-attention map and channel compression method, so as to guide the neural network to learn the edge area, and to refine the prediction details through the anti-attention mechanism layer by layer in a deep-to-shallow manner, and to realize more accurate fruit semantic segmentation.
[0138] To implement the above method embodiment, the present disclosure provides an anti-attention guava semantic segmentation detection device, such as Figure 20As shown, the attention-free guava semantic segmentation detection device 100 comprises an extraction module 110, a first segmentation module 120, an obtaining module 130, and a second segmentation module 140.
[0139] The extraction module 110 is configured to perform feature extraction on the image to be segmented to obtain a feature image.
[0140] The first segmentation module 120 is configured to perform pre-segmentation on the feature image to obtain an initial segmentation image, wherein the initial segmentation image comprises an initial predicted target region.
[0141] The obtaining module 130 is configured to erase a preset region inside the initial predicted target region in the initial segmentation image to obtain an attention-free image, wherein an edge region of the initial predicted target region is retained in the attention-free image.
[0142] The second segmentation module 140 is configured to perform feature fusion on the initial segmentation image and the attention-free image to obtain a semantic segmentation image, wherein the feature of the edge region on the target region in the semantic segmentation image is stronger than the feature of the edge region on the initial predicted target region in the initial segmentation image.
[0143] Optionally, the extraction module 110 comprises a feature extraction module and a first splicing module.
[0144] The feature extraction module is configured to input the image to be segmented into a plurality of convolution layers layer by layer to obtain image features extracted by each convolution layer in the plurality of convolution layers, wherein at least two convolution layers in the plurality of convolution layers are different.
[0145] The first splicing module is configured to splice a plurality of image features to obtain a feature image corresponding to each convolution layer in the plurality of convolution layers.
[0146] Optionally, the plurality of convolution layers comprise a first convolution layer, a second convolution layer, a third convolution layer, a fourth convolution layer, and a fifth convolution layer, the first convolution layer comprises a 1x1 convolution, a 1x3 convolution, a 3x1 convolution, and a 3x3 convolution connected in sequence; the second convolution layer comprises a 1x1 convolution, a 1x1 convolution, a 1x1 convolution, and a 3x3 convolution connected in sequence; the third convolution layer comprises a 1x1 convolution, a 1x1 convolution, a 1x1 convolution, and a 3x3 convolution connected in sequence; the fourth convolution layer comprises a 1x1 convolution; and the fifth convolution layer comprises a 1x1 convolution, wherein the 3x3 convolution in the first convolution layer has a hole rate of 3; the 3x3 convolution in the second convolution layer has a hole rate of 5; and the 3x3 convolution in the third convolution layer has a hole rate of 7.
[0147] Optionally, the first segmentation module 120 comprises a first fusion module and an initial segmentation module.
[0148] The first fusion module is configured to perform feature fusion on deep feature images in the plurality of feature images to obtain a fused feature image.
[0149] The initial segmentation module is configured to perform pre-segmentation on the fused feature image to obtain the initial segmentation image.
[0150] Optionally, the plurality of feature images comprises a first feature image output by a fifth convolutional layer, a second feature image output by a fourth convolutional layer, and a third feature image output by a third convolutional layer, and the fusion module comprises an up-sampling module, a first point multiplication module, a second splicing module, a second point multiplication module, and a second fusion module.
[0151] The up-sampling module is configured to perform up-sampling processing on the first feature image to obtain an up-sampled first feature image, and perform up-sampling processing on the second feature image to obtain an up-sampled second feature image.
[0152] The first point multiplication module is configured to perform point multiplication processing on the up-sampled first feature image and the second feature image to obtain a first point multiplication result.
[0153] The second splicing module is configured to splice the first point multiplication result and the up-sampled first feature image to obtain a splicing result.
[0154] The second point multiplication module is configured to perform point multiplication processing on the first point multiplication result, the up-sampled second feature image, and the third feature image to obtain a second point multiplication result.
[0155] The second fusion module is configured to perform feature fusion on the splicing result and the second point multiplication result to obtain the fused feature image.
[0156] Optionally, the second segmentation module 140 comprises a tensor obtaining module, a tensor segmentation module, and a tensor fusion module.
[0157] The tensor obtaining module is configured to obtain a feature tensor corresponding to the initial segmentation image.
[0158] The tensor segmentation module is configured to segment the feature tensor to obtain at least two segmentation tensors.
[0159] The tensor fusion module is configured to fuse the inverse attention image and the at least two segmentation tensors to obtain the semantic segmentation image.
[0160] Optionally, the feature fusion comprises third-order feature fusion.
[0161] With regard to the opacus attention guavus semantic segmentation detection device 100 in the above-described embodiments, the specific manner in which the various modules perform operations has been described in detail in the embodiments relating to the method, and will not be elaborated here.
[0162] Figure 21 is a block diagram of an electronic device according to an exemplary embodiment. As shown, the electronic device 700 can include a processor 701, a memory 702. The electronic device 700 can also include one or more of a multimedia component 703, an input / output (I / O) interface 704, and a communication component 705. Figure 21
[0163] The processor 701 is configured to control overall operations of the electronic device 700 to complete all or part of the steps of the anti-attention Opacus semantic segmentation detection method described above. The memory 702 is configured to store various types of data to support operations of the electronic device 700, which can include, for example, instructions for operating any application or method on the electronic device 700, and application-related data, such as contact data, sent and received messages, pictures, audio, video, and the like. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The multimedia component 703 can include a screen and an audio component. The screen can be, for example, a touch screen, and the audio component is configured to output and / or input audio signals. For example, the audio component can include a microphone configured to receive external audio signals. The received audio signals can be further stored in the memory 702 or transmitted through the communication component 705. The audio component also includes at least one speaker configured to output audio signals. The I / O interface 704 provides an interface between the processor 701 and other interface modules, which can be a keyboard, a mouse, a button, and the like. The buttons can be virtual buttons or physical buttons. The communication component 705 is configured to perform wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, and the like, or a combination of one or more of them, is not limited herein. Therefore, the corresponding communication component 705 can include a Wi-Fi module, a Bluetooth module, an NFC module, and the like.
[0164] In an exemplary embodiment, the electronic device 700 can be implemented by one or more Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), controller, microcontroller, microprocessor or other electronic elements for executing the above-mentioned anti-attention Opuntia vulgaris semantic segmentation detection method.
[0165] In another exemplary embodiment, a computer readable storage medium including program instructions is also provided, which, when executed by a processor, implements the steps of the above-mentioned anti-attention Opuntia vulgaris semantic segmentation detection method. For example, the computer readable storage medium can be the above-mentioned memory 702 including program instructions, which can be executed by the processor 701 of the electronic device 700 to complete the above-mentioned anti-attention Opuntia vulgaris semantic segmentation detection method.
[0166] In another exemplary embodiment, a computer readable storage medium including program instructions is also provided, which, when executed by a processor, implements the steps of the above-mentioned anti-attention Opuntia vulgaris semantic segmentation detection method. For example, the computer readable storage medium can be the above-mentioned memory 702 including program instructions, which can be executed by the processor 701 of the electronic device 700 to complete the above-mentioned anti-attention Opuntia vulgaris semantic segmentation detection method.
[0167] In summary, the anti-attention Opuntia vulgaris semantic segmentation detection device provided by the present disclosure extracts features from the image to be segmented to obtain a feature image; pre-segments the feature image to obtain an initial segmentation image, wherein the initial segmentation image includes an initial predicted target region; erases a preset region inside the initial predicted target region in the initial segmentation image to obtain an anti-attention image that retains the edge region of the initial predicted target region; and fuses the features of the initial segmentation image and the anti-attention image, i.e. enhances the edge of the initial predicted target region on the anti-attention image through the edge region on the initial segmentation image, to obtain a semantic segmentation image, wherein the feature of the edge region on the target region in the semantic segmentation image is stronger than that of the edge region on the initial predicted target region in the initial segmentation image. Through the above fusion, the feature of the edge region of the obtained semantic segmentation image is more obvious, and the semantic segmentation effect in the near-color scene is improved.
[0168] The preferred embodiments of the present disclosure are described in detail above with reference to the drawings, but the present disclosure is not limited to the specific details of the above-described embodiments. Various simple modifications can be made to the technical solutions of the present disclosure within the technical concept of the present disclosure, and these simple modifications all belong to the protection scope of the present disclosure.
[0169] In addition, it should be noted that each specific technical feature described in the above specific embodiments can be combined in any appropriate manner without contradiction. In order to avoid unnecessary repetition, various possible combinations are not described again by the present disclosure.
[0170] In addition, any combination of various different embodiments of the present disclosure can also be made, as long as it does not deviate from the idea of the present disclosure, and it should also be considered as disclosed by the present disclosure.
Claims
1. A method for anti-attention guava semantic segmentation detection, characterized in that, The method comprises: performing feature extraction on the image to be segmented to obtain a feature image; performing pre-segmentation on the feature image to obtain an initial segmentation image, wherein the initial segmentation image includes an initial predicted target region; erasing a preset region inside the initial predicted target region in the initial segmentation image to obtain an anti-attention image, wherein the edge region of the initial predicted target region is retained in the anti-attention image; performing feature fusion on the initial segmentation image and the anti-attention image to obtain a semantic segmentation image, wherein the feature of the edge region on the target region in the semantic segmentation image is stronger than the feature of the edge region on the initial predicted target region in the initial segmentation image; the performing feature fusion on the initial segmentation image and the anti-attention image to obtain a semantic segmentation image comprises: obtaining a feature tensor corresponding to the initial segmentation image; segmenting the feature tensor to obtain at least two segmentation tensors; fusing the anti-attention image and the at least two segmentation tensors to obtain the semantic segmentation image.
2. The method of claim 1, wherein, The performing feature extraction on the image to be segmented to obtain a feature image comprises: inputting the image to be segmented into a plurality of convolution layers layer by layer to obtain image features extracted by each convolution layer in the plurality of convolution layers, wherein at least two convolution layers in the plurality of convolution layers are different; splicing a plurality of image features to obtain a feature image corresponding to each convolution layer in the plurality of convolution layers.
3. The method of claim 2, wherein, The multi-layer convolutional layer comprises a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer and a fifth convolutional layer, the first convolutional layer comprises sequentially connected convolution, convolution, convolution, and convolution; the second convolutional layer comprises sequentially connected convolution, convolution, convolution, and convolution; the third convolutional layer comprises sequentially connected convolution, convolution, convolution, and convolution; the fourth convolutional layer comprises convolution; the fifth convolutional layer comprises convolution, wherein the first convolutional layer comprises convolution with a hole rate of 3; the second convolutional layer comprises convolution with a hole rate of 5; and the third convolutional layer comprises convolution with a hole rate of 7.
4. The method of claim 3, wherein, The performing pre-segmentation on the feature image to obtain an initial segmentation image comprises: performing feature fusion on deep feature images in a plurality of feature images to obtain a fused feature image; performing pre-segmentation on the fused feature image to obtain the initial segmentation image.
5. The method of claim 4, wherein, The plurality of feature images comprise a first feature image output by a fifth convolution layer, a second feature image output by a fourth convolution layer, and a third feature image output by a third convolution layer; the performing feature fusion on deep feature images in a plurality of feature images to obtain a fused feature image comprises: performing up-sampling processing on the first feature image to obtain an up-sampled first feature image, and performing up-sampling processing on the second feature image to obtain an up-sampled second feature image; performing point multiplication processing on the up-sampled first feature image and the second feature image to obtain a first point multiplication result; splicing the first point multiplication result and the up-sampled first feature image to obtain a splicing result; performing point multiplication processing on the first point multiplication result, the up-sampled second feature image, and the third feature image to obtain a second point multiplication result; performing feature fusion on the splicing result and the second point multiplication result to obtain the fused feature image.
6. The method of claim 1, wherein, The feature fusion comprises third-order feature fusion.
7. An anti-attention guava semantic segmentation detection device, characterized in that, The device comprises: an extraction module configured to perform feature extraction on an image to be segmented to obtain a feature image; a first segmentation module configured to perform pre-segmentation on the feature image to obtain an initial segmentation image, wherein the initial segmentation image includes an initial predicted target region; The obtaining module is configured to erase a preset region inside the initial predicted target region in the initial segmentation image to obtain an anti-attention image, wherein an edge region of the initial predicted target region is reserved in the anti-attention image. The second segmentation module is configured to perform feature fusion on the initial segmentation image and the anti-attention image to obtain a semantic segmentation image, wherein the feature of the edge region on the target region in the semantic segmentation image is stronger than the feature of the edge region on the initial predicted target region in the initial segmentation image; and the performing feature fusion on the initial segmentation image and the anti-attention image to obtain the semantic segmentation image comprises: acquiring a feature tensor corresponding to the initial segmentation image; segmenting the feature tensor to obtain at least two segmentation tensors; and fusing the anti-attention image with the at least two segmentation tensors to obtain the semantic segmentation image.
8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The program, when executed by a processor, implements the steps of the method of any one of claims 1-6.
9. An electronic device, comprising: The program, when executed by a processor, implements the steps of the method of any one of claims 1-6. The memory has a computer program stored thereon; The processor is configured to execute the computer program in the memory to implement the steps of the method of any one of claims 1-6.