An Edge Detection Method Based on Gradient and Fuzzy Supervision
By introducing gradient information and fuzzy supervision into the edge detection model, the model training process is optimized, and the problem of insufficient utilization of image supervision signals by the edge detection model is solved, and the accuracy and recall of edge detection are improved.
Patent Information
- Application Number
- CN202310384359.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-04-12
AI Technical Summary
The existing edge detection model fails to fully utilize image supervision signals during training, especially the gradient information in the edge label image and the blurred features of the deep output, resulting in poor detection results.
By introducing gradient information and fuzzy supervision of edge label images, the gradient loss function and fuzzy edge label images are constructed, and the training process of edge detection model is optimized, and the accuracy and recall rate of the model in edge detection tasks are improved.
Without increasing the number of model parameters and inference time, the accuracy and recall of the edge detection model are significantly improved, and are suitable for edge detection models with HED-like structures.
Smart Images

Figure CN116452624B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and particularly relates to an edge detection method based on gradient and fuzzy supervision. Background Art
[0002] Edge detection is a fundamental problem in computer vision technology. Extracting edge features in an image through edge detection helps improve the performance of high-level computer vision tasks such as image segmentation and object detection. Since the birth of the edge detection problem, the research methods for it have evolved from low-level feature-based methods that utilize derivative information in images, to traditional machine learning methods that artificially design various features, and then to data-driven deep learning methods. With the continuous improvement of methods, the performance of edge detection models has been continuously enhanced and has reached a level close to human perception of edges. The Holistically-Nested Edge Detection (HED) model is a deep learning model constructed based on a convolutional neural network, which achieves end-to-end edge detection and the highest edge detection accuracy at the time of publication. Most subsequent mainstream edge detection model structures are improved based on it. By analyzing the use of supervision signals by HED-like structure models, the present invention discovers that there is room for optimization in the training process of such structure models.
[0003] During the training process, an edge detection model uses an image as supervision information and treats the edge detection problem as a classification problem. It performs binary classification on each pixel in the edge prediction image to determine the probability that it belongs to the edge class or the non-edge class, and then uses the class of each pixel point in the edge label image to supervise the output. However, different from a pure classification problem, the edge label image contains not only the class information of pixel points but also the relative relationships between pixel points. For example, the gradient information of image pixels can reflect the association between local pixels. Previous HED-like models, such as the Pixel Difference Network (PiDiNet) model, ignored this additional information during training. Therefore, in the present invention, the gradient image in the X–Y direction in the edge label image is introduced as a supervision signal during the training process to achieve a more comprehensive utilization of supervision information, making the gray value distribution and gradient value distribution of the edge prediction image of the model simultaneously close to the edge label image.
[0004] The model of the HED-like structure supervises the edge prediction images of all stages of the model through the same edge label image, but there are differences in the edge prediction images of different stages. For example, the edge prediction image output by the shallow layer in the HED model contains more local details, and the edge prediction image output by the deep layer pays more attention to the semantic information of the objects in the image. It can be seen that different-depth stages learn features of different stages, and it is not the best choice to train these stages with the same supervision. In the paper "A computational approach to edge detection" by Liu et al., it is proposed to use the Canny algorithm to process the original image and artificially set parameters to output specific edge label images for different stages. However, reasonable parameter configuration in the method must be obtained through multiple experiments. In the Bi-Direction Cascade Network (BDCN) model, the supervision signals of each stage come from the combination of its previous and subsequent stages, and the comprehensive performance of the accuracy and recall rate of the edge detection model in the edge detection task is improved. However, this change will lead to the complication of the edge detection model structure, thereby increasing the number of model parameters and the inference time. In addition, the edge prediction image of the multi-stage edge detection model in the high-level stage is obtained through upsampling operations such as deconvolution or bilinear interpolation. Therefore, the edge part in the image is relatively blurred, but the clear edge label image is still used for supervision during training. To address the above problems, the present invention analyzes the edge detection problem as a regression problem, blurs the edge label image for supervising the deep output to match the output characteristics of the deep layer of the model, reduces the learning difficulty of the edge detection model, and makes the edge detection model easier to learn reasonable feature representations. Summary of the Invention
[0005] The purpose of the present invention is to provide an edge detection method based on gradient and fuzzy supervision, aiming to comprehensively improve the accuracy and recall rate indicators of the deep learning model in the edge detection task and solve the problem that the edge detection model in the field of computer vision lacks reasonable and comprehensive utilization of image supervision signals.
[0006] The technical solution for realizing the present invention is as follows: An edge detection method based on gradient and fuzzy supervision, the steps are as follows:
[0007] Step 1: Select the BSDS500 dataset. The BSDS500 dataset includes a training set, a validation set, and a test set. Augment the original images of the training set and the validation set respectively to obtain augmented images, normalize the augmented images to construct a model training set, and use the augmented images in the model training set as the input images of the edge detection model.
[0008] Meanwhile, since each original image in the BSDS500 dataset corresponds to multiple edge annotation images, during the construction of the model training set, multiple edge annotation images are fused by summing and taking the average to obtain the annotation image corresponding to the original image.
[0009] Proceed to Step 2.
[0010] Step 2: Construct an edge detection model of the HED-like structure. The edge detection model of the HED-like structure includes a backbone network and four branch networks; proceed to Step 3.
[0011] Step 3: Construct the loss function of the edge detection model of the HED-like structure. The loss function includes a basic loss function, a gradient loss function, a fuzzy supervision label, and an overall loss function.
[0012] Step 4: Based on the edge detection model of the HED-like structure and the loss function of the above-mentioned edge detection model, combined with the model training set, train the edge detection model to obtain a trained edge detection model, and proceed to Step 5.
[0013] Step 5: Select the test set in the publicly available BSDS500 dataset, and form a test sample set through normalization processing to evaluate the accuracy Pre and recall rate Rec of the edge detection model prediction, and obtain the edge prediction image corresponding to the test set.
[0014] Compared with the prior art, the significant advantages of the present invention are as follows:
[0015] (1) By using the gradient information in the edge label image and constructing a fuzzy edge label image for the edge prediction image in the deeper stage, the present invention realizes more sufficient and reasonable use of the edge label image, achieving the purpose of optimizing the training process of the edge detection model and improving the prediction effect of the edge detection model.
[0016] (2) On the premise of using the gradient information in the edge label image to improve the visual performance of the edge prediction image, the present invention does not increase the number of parameters and inference time of the edge detection model.
[0017] (3) The present invention has portability. Without changing the structure of the edge detection model, the method of the present invention can be applied to the edge detection model of the HED-like structure to comprehensively improve the accuracy and recall rate of the edge detection model. Description of the Drawings
[0018] Figure 1 is the overall flowchart of the present invention.
[0019] Figure 2 is the overall structure diagram of the present invention (PiDiNet model).
[0020] Figure 3 It is the attention module in the model of the present invention.
[0021] Figure 4 It is a schematic diagram of the gradient supervision of the present invention.
[0022] Figure 5 It is a schematic diagram of the generation of fuzzy labels of the present invention. Specific embodiments
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.
[0024] Combined with Figures 1 - 5 , the edge detection method based on gradient and fuzzy supervision of the present invention includes the following steps:
[0025] Step 1: Select the BSDS500 dataset. The BSDS500 dataset includes a training set, a validation set, and a test set. The original images in the training set and the validation set are respectively subjected to data augmentation to obtain augmented images. The augmented images are normalized to construct a model training set, and the augmented images in the model training set are used as the input images of the edge detection model.
[0026] Since each original image in the BSDS500 dataset corresponds to multiple edge annotation images, during the construction of the model training set, multiple edge annotation images are fused by summing and averaging to obtain the annotation image corresponding to the original image.
[0027] The present invention has no special requirements for the data augmentation method. It should be noted that local color changes to the original image are avoided. In the present invention, the original image is sequentially subjected to three data augmentation methods: rotation, scaling, and flipping. These methods are widely used in the construction of the training set of the edge detection model. In addition, the original image and the corresponding annotation image need to adopt exactly the same data augmentation operations.
[0028] Proceed to Step 2.
[0029] Step 2. Construct an edge detection model with a Holistically-Nested Edge Detection (HED) structure (hereinafter referred to as the edge detection model). The HED-like structure refers to an edge detection model structure with multi-stage feature fusion and allowing the introduction of deep supervision. In the present invention, the HED-like structure adopts a Pixel Difference Network (PiDiNet) model, and the input of the edge detection model is called the input image, the corresponding labeled image of the input image is called the edge label image, the output of the edge detection model is called the edge prediction image, and the input image is transmitted in the state of features (graphs) in the edge detection model. The edge detection model includes a backbone network and four branch networks. The specific construction steps are as follows:
[0030] Step 2.1. Construct the backbone network. Combining Figure 2 , the backbone network of the edge detection model is divided into four stages (the first stage, the second stage, the third stage, and the fourth stage in sequence), and each stage consists of four convolutional modules, that is, the backbone network consists of sixteen convolutional modules in total. The first convolutional module in the entire edge detection model consists of a single convolutional layer, and the structures of the remaining fifteen convolutional modules are similar to the residual convolutional module, and are called residual-like convolutional modules. The residual-like convolutional module includes a 3×3 convolution, a ReLU activation layer, and a 1×1 convolution, and the output of the residual-like convolutional module is the sum of the input feature and the output feature. In the backbone network, the maximum pooling layer is used to connect between each stage. Each maximum pooling layer compresses the length and width of the feature to half of the input image, and the maximum pooling layers before the second stage and the third stage respectively double the number of channels of the feature. Let C represent the number of channels of the first feature in the backbone network, then the feature channels of each stage in the backbone network are C, 2C, 4C, and 4C respectively. In addition, the backbone network performs shortcuts on the first residual-like convolutional module in the second stage, the third stage, and the fourth stage to optimize the learning of the features by the edge detection model.
[0031] Step 2.2. Construct the branch network. The branch network of the edge detection model is led out from the last residual-like convolutional module in each stage of the backbone network. Therefore, the edge detection model has the same number of branch networks as the number of stages, and the weights of each branch network are not shared. To learn the multi-stage edge feature representation, the edge detection model with the HED-like structure outputs the stage-based edge prediction images at each stage, and uses the edge label image to supervise the edge prediction images output at each stage. Specifically, each branch network in the PiDiNet model consists of an attention module and a Sigmoid activation layer.
[0032] Combining Figure 3, in the attention module, to obtain a richer edge feature representation, the features extracted at each stage are passed through a Compact Dilation Convolution Module (CDCM). In the CDCM, the features first pass through a ReLU activation layer, then through a 1×1 convolution, and finally through four dilated convolutions with different dilation coefficients in parallel to enrich the features. The features from different dilated convolutions are summed to obtain the feature output of the CDCM. Then, the feature output is passed through a Compact Spatial Attention Module (CSAM). Specifically, in the CSAM, the features pass through a ReLU activation layer, a 1×1 convolution, a 3×3 convolution, and a Sigmoid activation layer to generate a spatial attention mask. The spatial attention mask is multiplied by the features to suppress the feature representation of the non-edge regions in the edge prediction image, and the feature output of the CSAM is obtained.
[0033] At the end of the attention module, a 1×1 convolution is used to reduce the dimension of the edge features obtained through the CDCM and CSAM, obtaining single-channel features. Through bilinear interpolation, the single-channel features are upsampled to align with the input image. The single-channel features at each stage are concatenated and reduced in dimension at the channel level, and the edge prediction image of the multi-stage feature fusion is obtained by the Sigmoid activation layer. This edge prediction image is used as the final edge prediction image of the model, and the stage-wise edge prediction image is obtained by passing the upsampled single-channel features through the Sigmoid activation layer.
[0034] Go to step 3.
[0035] Step 3: Construct the loss function of the edge detection model. The loss function includes a basic loss function, a gradient loss function, a fuzzy supervision label, and an overall loss function. The specific steps are as follows:
[0036] Step 3.1: Construct the basic loss function. Different from the previous edge detection methods based on deep learning models that regard the edge detection problem as a binary classification problem, the present invention treats the edge detection problem as a regression problem. Therefore, the weighted binary cross-entropy loss function commonly used in binary classification problems is modified to a weighted L2 loss function as the regression loss function L fuse (P, Q) is as follows:
[0037]
[0038] where P represents the edge prediction image, Q represents the edge label image corresponding to P, and the edge prediction image and the edge label image are of the same size. The length of both the edge prediction image and the edge label image is represented by H, and the width is represented by W, p ijDenotes the grayscale value at (i, j) in the edge prediction image P, q ij Denotes the grayscale value at (i, j) in the edge label image Q, ||…||2 represents the L2 norm. Since in the edge detection problem, edge pixels account for a very small part of the edge label image, to address the imbalance in the number of edge pixels and non-edge pixels, it is necessary to weight the pixels with different grayscale values when calculating the L2 loss. w ij Denotes the weight at (i, j) in the edge label image Q, and its specific calculation formula is:
[0039]
[0040] Where Q+ represents the number of edges in the edge label image, Q- represents the number of non-edges in the edge label image, η is a hyperparameter specified by humans, and for the BSDS500 dataset, the present invention takes η = 0.3. Pixels with grayscale values lower than this hyperparameter and non-zero in the normalized edge label image are regarded as confused pixels and are not included in the calculation of the loss function. During the training stage of the edge detection model, while supervising the final edge prediction image of the model, deep supervision is introduced, that is, the edge prediction images at different stages of the edge detection model are supervised. Adding the regression loss functions of each stage of the edge detection model to the regression loss function of the final edge detection model, the basic loss function L bright (P, Q) expression is:
[0041]
[0042] Where M represents the total number of stages of the edge detection model. In the present invention, M = 4, w s ide m Denotes the weight of the regression loss function of the m-th stage of the edge detection model included in the basic loss function, wf use Denotes the weight of the regression loss function of the final edge detection model included in the basic loss function, L sidem Denotes the regression loss function of the edge detection model at the m-th stage, P m Denotes the edge prediction image at the m-th stage. In the present invention, it is set Therefore, the mathematical expressions of the aforementioned two weights are not explicitly given in the subsequent formulas.
[0043] Step 3.2, construct the gradient loss function. The present invention notices that in addition to directly providing grayscale value information, the edge label image also contains information such as divergence and gradient. However, existing edge detection models only use grayscale value information to supervise the training process, lacking the comprehensive utilization of the edge label image. The present invention calculates the gradient information in the edge label image and constrains the gradient of the generated edge prediction image during the training process, so that when the edge detection model focuses on the grayscale value at the edge, it also focuses on the significant gradient value at the junction of edge pixels and non-edge pixels, which is beneficial for the edge detection model to learn a better edge feature representation. Combining Figure 4 , the gradient of the edge prediction image consists of gradient values in the X–Y two directions. In implementation, the present invention adopts the first-order difference method to obtain the gradient values of the edge label image and the edge prediction image.
[0044]
[0045] Among them, represents the first-order difference of the edge label image Q in the X direction, which is called the gradient image of the edge label image in the X direction, represents the first-order difference of the edge label image Q in the Y direction, q ij represents the grayscale value at (i,j) in the edge label image Q. Similarly, the gradient images of the edge prediction image P in the X and Y directions are obtained by the first-order difference method and The gradient loss function L grad (P,Q) adopts a weighted L2 loss function, and its mathematical expression is as follows:
[0046]
[0047] Among them, represents the gradient value at (i,j) in the gradient map of the edge prediction image in the X direction, that is, the gradient value at (i,j), represents the gradient value at (i,j) in the gradient map of the edge label image in the Y direction, that is, the gradient value at (i,j).
[0048] Step 3.3, construct the fuzzy edge label image. To reduce the training difficulty of the edge detection model, for the edge prediction image in the deep stage, the present invention uses the fuzzy edge label image to supervise it. At this time, the edge label image is no longer only composed of pixels with grayscale values {0,1}, but is composed of pixels with grayscale values in the interval [0,1], and the edge detection problem can only be processed as a regression problem. Combining Figure 5 , the present invention uses the fuzzy edge label to construct an alternative label in bold. The construction process of the fuzzy edge label image is as follows:
[0049] Blur(Q, m) = Q * W m (6)
[0050] W m = Conv(2m + 1, 2m + 1) (7)
[0051] where W m represents an unbiased convolution with a convolution kernel size of (2m + 1) × (2m + 1), and the values of the elements in the convolution kernel are taken as 1 / (2m + 1). 2 , the number of stages m is determined by the stage depth supervised by the blurred supervision label. For example, for the output of the third stage of the edge detection model, m = 3 is taken. Simply blurring the edge label image will destroy the true edge positions in the edge label image. After convolution, it is necessary to ensure that the gray value at the original edge does not change. Therefore, the gray value at the original edge in the blurred edge label image is reset to 1 again.
[0052] Step 3.4, construct the overall loss function. On the premise that the edge detection model takes the basic loss function as the main body, the edge label image and the gradient image of the edge label image are used as supervision for the shallow stage and the final edge prediction image of the model, and the blurred edge label image is used as supervision for the edge prediction image of the deep stage. The overall loss function L final (P, Q) is as follows:
[0053]
[0054] where N represents the number of shallow stages of the edge detection model, represents the gradient loss function of the model in the nth stage. In the implementation of the present invention, N = 2 is taken. λ grad represents the weight of the gradient loss function in the overall loss function. In the implementation of the present invention, λ grad = 1.0.
[0055] Proceed to Step 4.
[0056] Step 4, based on the edge detection model of the class HED structure and the loss function of the above-mentioned edge detection model, combined with the model training set, train the edge detection model to obtain a trained edge detection model.
[0057] The relevant software versions used in the method of the present invention are PyTorch (1.9.0), Python (3.7.5), and the latest versions are used for other environment-dependent packages. The graphics card on the hardware device is Nvidia-RTX-TiTAN 24GB, and the CPU is Intel Xeon Gold6230R. Since the image sizes are not unified in the composition of the model training set, the batchsize in the training process of the edge detection model is set to 1, but backpropagation is performed once every 24 iterations to update the parameters of the edge detection model. The number of training epochs of the edge detection model is set to 14, the initial learning rate is set to 5e-3, and the learning rate is decayed to 0.1 times the previous learning rate at the 8th and 12th epochs of training. The optimizer used in the training process is Adam, the weight decay of the convolutional layer in the edge detection model is set to 1e-4, and there is no weight decay in the ReLU activation layer. In addition, in deep learning, the training of the model is affected by random factors. To reduce its influence, the random seed of the present invention is set to 2022 without a specific purpose. Through the training set constructed based on the BSDS500 dataset and combined with the loss function proposed by the present invention, the edge detection model is trained according to the aforementioned hyperparameter configuration to obtain a trained edge detection model. Proceed to step 5.
[0058] Step 5: Select the test set in the publicly available BSDS500 dataset, and form a test sample set through normalization processing to evaluate the accuracy (Precision, Pre) and recall rate (Recall, Rec) of the edge detection model prediction, and obtain the edge prediction image corresponding to the test set.
[0059] In the edge detection problem, the F value is used to comprehensively represent the accuracy Pre and the recall rate Rec. Among them, the accuracy Pre describes the proportion of correctly predicted pixels among the pixels predicted as edges. The recall rate Rec describes the proportion of pixels predicted as edges among all edge pixels. The mathematical expression of the F value is as follows:
[0060]
[0061] Although the definition of the F value is simple, the actual calculation process is relatively cumbersome. It does not directly calculate using the edge prediction image and the edge label image. Therefore, it is necessary to explain the calculation process of the edge detection evaluation index. For the edge prediction image output by the edge detection model, first perform non-maximum suppression (NMS) to obtain a relatively clear edge prediction image, and then perform morphological thinning on the processed edge prediction image to further clarify the edges. More intuitively, if there is a line segment with a width of 3 pixels in the edge prediction image after non-maximum suppression, the width of this line segment is only 1 pixel after morphological thinning. After non-maximum suppression and morphological thinning operations, the edge prediction image corresponding to the test set is obtained, that is, the edge prediction image post-processed by the edge detection model. To obtain the F value index, it is also necessary to match the edge label image and the edge prediction image. There is a concept of tolerance in the matching process, that is, the distance error between the edge points in the edge prediction image and the edge points in the edge label image can be considered matched within a certain range. Define the matching tolerance as d. When the Euclidean distance error between a certain set of edge pixel coordinates in the two images is within pixels, then this set of pixels is considered matched, and the matched pixels will not be repeatedly matched with other pixels.
[0062] For an original image, there may be edge label images made by multiple annotators. In the present invention, the set composed of multiple edge label images is simply referred to as the edge label set. Before calculating the F value, it is necessary to process the edge label set to generate the E graph and the G graph from the edge label set. Among them, the E graph represents the intersection of all edge label images in the edge label set, and the G graph represents the sum of all edge label images in the edge label set. When there is only one edge label image in the edge label set, the E graph and the G graph are equivalent. After obtaining the E graph and the G graph, the post-processed edge prediction image is respectively intersected with the E graph and the G graph to obtain the matching images E match and G match . The recall rate Rec and the accuracy Pre indexes of the edge detection model are statistically obtained from the above four images:
[0063]
[0064] Among them, nnz(·) represents the number of non-zero elements in the graph, and sum(·) represents the sum of all elements in the graph. Therefore, in formula (10), cntP represents the number of non-zero elements in the E match graph, sumP represents the sum of all elements in the E graph, and cntR represents the number of elements in the G matchThe number of non-zero elements in the figure, and sumR represents the sum of all elements in the G figure. After obtaining the recall rate Rec and accuracy rate Pre metrics of the model, the F value of the model's edge prediction image and the corresponding edge label image is calculated through Equation (9). In the actual scenario, the post-processed edge prediction image is usually used as the edge prediction image corresponding to the original image.
Claims
1. An edge detection method based on gradient and fuzzy supervision, characterized in that, The steps are as follows: Step 1: Select the BSDS500 dataset. The BSDS500 dataset includes a training set, a validation set, and a test set. Augment the original images in the training set and the validation set respectively to obtain augmented images. Normalize the augmented images to construct a model training set, and use the augmented images in the model training set as the input images for the edge detection model; Meanwhile, since each original image in the BSDS500 dataset corresponds to multiple edge annotation images, during the construction of the model training set, multiple edge annotation images are fused in a way of summing and taking the mean to obtain the annotation image corresponding to the original image; Proceed to Step 2; Step 2: Construct an edge detection model with a HED-like structure. The edge detection model with a HED-like structure includes a backbone network and four branch networks; Proceed to Step 3; Step 3: Construct the loss function of the edge detection model with a HED-like structure. The loss function includes a basic loss function, a gradient loss function, a fuzzy supervision label, and an overall loss function; Proceed to Step 4; Step 4: Based on the edge detection model with a HED-like structure and the loss function of the above edge detection model, combined with the model training set, train the edge detection model to obtain a trained edge detection model; Proceed to Step 5; Step 5: Select the test set in the publicly available BSDS500 dataset, and form a test sample set through normalization to evaluate the accuracy Pre and recall rate Rec of the edge detection model prediction, and obtain the edge prediction image corresponding to the test set.
2. The edge detection method based on gradient and fuzzy supervision according to claim 1, characterized in that: In Step 1, three data augmentation methods of rotation, scaling, and flipping are performed on the original images in sequence.
3. The edge detection method based on gradient and fuzzy supervision according to claim 2, characterized in that: In Step 2, construct an edge detection model with a HED-like structure. The HED-like structure uses the PiDiNet model, which includes a backbone network and four branch networks, specifically as follows: Step 2.1: Construct the backbone network: The backbone network is divided into four stages, namely the first stage, the second stage, the third stage, and the fourth stage. Each stage consists of four convolutional modules, that is, the backbone network consists of a total of sixteen convolutional modules; The first convolutional module in the entire edge detection model consists of a single convolutional layer, and the structures of the remaining fifteen convolutional modules are similar to the residual convolutional module, called the residual-like convolutional module; The residual-like convolutional module includes a 3×3 convolution, a ReLU activation layer, and a 1×1 convolution, and the output of the residual-like convolutional module is the sum of the input feature and the output feature; Step 2.2: Construct the branch network: The branch network is led out from the last residual-like convolutional module in each stage of the backbone network, and the weights of each branch network are not shared; Each branch network consists of an attention module and a Sigmoid activation layer.
4. The edge detection method based on gradient and fuzzy supervision according to claim 3, wherein: In the backbone network, max-pooling layers are used to connect between stages. Each max-pooling layer compresses the length and width of the features to half of the input image, and the max-pooling layers before the second stage and the third stage double the number of channels of the features respectively. Let C represent the number of channels of the first feature in the backbone network, then the number of feature channels in each stage of the backbone network is C, 2C, 4C, and 4C respectively. The backbone network shorts the first class residual convolution module in the second stage, the third stage, and the fourth stage to optimize the feature learning of the edge detection model.
5. The edge detection method based on gradient and fuzzy supervision according to claim 3, characterized in that: In the attention module of the branch network, the features extracted in each stage are passed through CDCM. In CDCM, the features first pass through a ReLU activation layer, then through a 1×1 convolution, and finally through four dilated convolutions with different dilation coefficients in parallel to enrich the features. The features from different dilated convolutions are summed to obtain the feature output of CDCM. Then the feature output is passed through CSAM. In CSAM, the features pass through a ReLU activation layer, a 1×1 convolution, a 3×3 convolution, and a Sigmoid activation layer to generate a spatial attention mask. The spatial attention mask is multiplied by the features to suppress the feature representation of non-edge regions in the edge prediction image, and the feature output of CSAM is obtained. At the end of the attention module, a 1×1 convolution is used to reduce the dimension of the edge features obtained after CDCM and CSAM to obtain single-channel features. Through bilinear interpolation, the single-channel features are upsampled to align with the input image. The single-channel features in each stage are concatenated and reduced in dimension at the channel level, and the edge prediction image of multi-stage feature fusion is obtained by a Sigmoid activation layer, and this edge prediction image is used as the final edge prediction image of the model. The stage edge prediction image is obtained by passing the upsampled single-channel features through a Sigmoid activation layer.
6. The edge detection method based on gradient and fuzzy supervision according to claim 3, wherein Construct the loss function of the edge detection model. The loss function includes a basic loss function, a gradient loss function, a blurred supervision label, and an overall loss function. The specific steps are as follows: Step 3.1: Construct the basic loss function: Treat the edge detection problem as a regression problem, that is, modify the weighted binary cross-entropy loss function commonly used in binary classification problems to a weighted L2 loss function as the regression loss function L of the edge detection model fuse (P, Q) is as follows: Where P represents the edge prediction image, Q represents the edge label image corresponding to P, and the edge prediction image and the edge label image have the same size; the length of both the edge prediction image and the edge label image is represented by H, and the width of both is represented by W, and p ij represents the gray value at (i, j) in the edge prediction image P, and q ij represents the gray value at (i, j) in the edge label image Q, and ||…||2 represents the 2-norm; w ij represents the weight at (i, j) in the edge label image Q, and its specific calculation formula is: Where Q + represents the number of edges in the edge label image, Q - represents the number of non - edges in the edge label image, η is a hyper - parameter specified by humans. During the training stage of the edge detection model, while supervising the final edge prediction image of the model, deep supervision is introduced, that is, supervising the edge prediction images at different stages of the edge detection model; adding the regression loss functions at different stages of the edge detection model to the final regression loss function of the edge detection model, the basic loss function L bright (P, Q) is expressed as: Among them, M represents the number of stages of the overall edge detection model. In the present invention, M = 4. represents the weight of the regression loss function of the m-th stage of the edge detection model incorporated into the basic loss function, wf use represents the weight of the final regression loss function of the edge detection model incorporated into the basic loss function. represents the regression loss function of the edge detection model at the m-th stage, P m represents the edge prediction image of the m-th stage, set Step 3.2: Construct the gradient loss function: The gradient of the edge prediction image consists of gradient values in the X–Y two directions. The gradient values of the edge label image and the edge prediction image are obtained by using the first-order difference method: Among them, represents the first-order difference of the edge label image Q in the X direction, which is called the gradient image of the edge label image in the X direction. represents the first-order difference of the edge label image Q in the Y direction, q ij represents the gray value at (i, j) in the edge label image Q. Similarly, the gradient images of the edge prediction image P in the X and Y directions are obtained in the way of first-order difference. and the gradient loss function L grad (P, Q) adopts a weighted L2 loss function, and its mathematical expression is as follows: Among them, represents the gradient value of the gradient map of the edge prediction image in the X direction at (i, j), that is, the gradient value at (i, j), represents the gradient value of the gradient map of the edge label image in the Y direction at (i, j), that is, the gradient value at (i, j); Step 3.3: Construct the blurred edge label image: For the edge prediction image in the deep stage, the blurred edge label image is used to supervise it. At this time, the edge label image is composed of pixels with gray values in the range of [0,1], and the edge detection problem can only be treated as a regression problem. Use the blurred edge label to construct an expression to replace the bold label. The construction process of the blurred edge label image is as follows: Blur(Q,m) = Q * W m (6) W m = Conv(2m + 1, 2m + 1) (7) Among which W m represents a bias-free convolution with a convolution kernel size of (2m + 1) × (2m + 1), and the values of the elements in the convolution kernel are taken as 1 / (2m + 1) 2 , and the number of stages m is determined by the stage depth supervised by the fuzzy supervision label; Step 3.4: Construct the overall loss function: On the premise that the edge detection model takes the basic loss function as the main body, the edge prediction images in the shallow stage and the final edge prediction image of the model are supervised by using the edge label image and the gradient image of the edge label image, and the edge prediction image is supervised by using the blurred edge label image; the overall loss function \(L\) of the edge detection model final (P, Q) is as follows: where N represents the number of shallow stages of the edge detection model, represents the gradient loss function of the model in the nth stage, and λ grad represents the weight of the gradient loss function in the overall loss function.
7. The edge detection method based on gradient and fuzzy supervision according to claim 6, characterized in that, In step 4, based on the edge detection model of the class HED structure and the loss function of the above edge detection model, combined with the model training set, the edge detection model is trained to obtain a trained edge detection model, specifically as follows: Since the image sizes are not unified in the composition of the model training set, the batch size during the training of the edge detection model is set to 1, but backpropagation is performed once every 24 iterations to update the parameters of the edge detection model; the number of training epochs of the edge detection model is set to 14, the initial learning rate is set to 5e-3, and the learning rate is decayed to 0.1 times the previous learning rate at the 8th and 12th epochs of training; The optimizer used in the training process is Adam, the weight decay of the convolutional layer in the edge detection model is set to 1e-4, and there is no weight decay in the ReLU activation layer.
8. The edge detection method based on gradient and fuzzy supervision according to claim 7, characterized in that In step 5, the test set in the publicly available BSDS500 dataset is selected, and the test sample set is formed through normalization processing to evaluate the accuracy Pre and recall rate Rec predicted by the edge detection model, and the edge prediction image corresponding to the test set is obtained, specifically as follows: In the edge detection problem, the F value is used to comprehensively represent the accuracy Pre and recall rate Rec, and the mathematical expression of the F value is as follows: For the edge prediction image output by the edge detection model, first perform non-maximum suppression to obtain a relatively clear edge prediction image, and then perform morphological thinning operation on the processed edge prediction image to further clarify the edges; after non-maximum suppression and morphological thinning operations, the edge prediction image corresponding to the test set is obtained, that is, the edge prediction image post-processed by the edge detection model; Define the tolerance of matching as d. When the Euclidean distance error between the coordinates of a set of edge pixels in two images is within pixels, it is considered that this set of pixels is matched, and the matched pixels will not be repeatedly matched with other pixels; For an original image, there may be multiple edge label images created by multiple annotators. The set composed of multiple edge label images is simply referred to as the edge label set. Before calculating the F value, the edge label set needs to be processed to generate an E map and a G map from the edge label set. Among them, the E map represents the intersection of all edge label images in the edge label set, and the G map represents the sum of all edge label images in the edge label set. When there is only one edge label image in the edge label set, the E map and the G map are equivalent. After obtaining the E map and the G map, the post-processed edge prediction image is intersected with the E map and the G map respectively to obtain the matching images E match and G match ; The recall rate Rec and accuracy rate Pre indicators of the edge detection model are obtained by statistics from the aforementioned four images: Among them, nnz(·) represents the number of non-zero elements in the graph, and sum(·) represents the sum of all elements in the graph. Therefore, cntP in Equation (10) represents the number of non-zero elements in graph E match The number of non-zero elements in the graph, sumP represents the sum of all elements in graph E, cntR represents the number of non-zero elements in graph G match The number of non-zero elements in the graph, sumR represents the sum of all elements in graph G; after obtaining the recall rate Rec and accuracy Pre indicators of the model, the F value of the model's edge prediction image and the corresponding edge label image is calculated through Equation (9).
Citation Information
Patent Citations
Multi-scale edge detection method under deep supervision
CN112580661A
Implicit edge prior-based scale progressive image completion method
CN113298733A