Edge Detection Method and Device Based on Bidirectional Feature Supervision
By adopting a bidirectional feature supervision method in edge detection, multi-level edge feature extraction and bidirectional confidence mapping are performed, and the problems of error detection and low accuracy in the prior art are solved, achieving more efficient and accurate edge detection.
Patent Information
- Application Number
- CN202211649783.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-12-21
AI Technical Summary
The prior art has problems with error detection in edge detection, especially in scenarios where low contrast and uneven grayscale distribution are distributed, and the annotation-based supervision mode has a great impact on the position accuracy of different scales, resulting in slow network convergence speed and low detection accuracy.
The edge detection method based on bidirectional feature supervision is adopted, and the supervision of different scales is achieved through multi-level edge feature extraction processing and bidirectional edge confidence mapping processing, thereby improving the network convergence speed and detection accuracy.
It improves the accuracy of edge detection, shortens the network training time, and improves the convergence speed and detection accuracy of edge detection networks.
Smart Images

Figure CN116205938B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and particularly to an edge detection method and device based on bidirectional feature supervision in image processing. Background Art
[0002] Edge detection is an important application in image processing. An edge is the boundary line between different regions and is a set of pixel points with significant changes in surrounding (local) gray values, having two attributes: amplitude and direction.
[0003]
[0004] Where G x is the gradient in the X direction, and G y is the gradient in the Y direction. This is not an absolute definition, mainly used to illustrate that an edge is a local feature and that significant changes in surrounding gray values generate an edge. A contour is a description of the complete boundary of an object, and edge points are connected one by one to form a contour. According to the human visual characteristic, when looking at an object, generally the contour information of the object is obtained first, and then the detailed information in the object is obtained.
[0005] Traditional image processing for edge detection mainly includes several steps: (1) calculating the gradient of the image; (2) performing threshold judgment or non-maximum suppression according to the gradient assignment; (3) obtaining the edge information of the object and performing contour tracking, morphological operations, etc. Judging based on traditional thresholds is likely to cause false detection of edges in scenarios with low contrast, uneven gray distribution, etc.
[0006] For edge detection, currently a supervised mode based on annotation is adopted, but the supervised mode based on annotation has a great influence on the position accuracy of different scales, resulting in a slow network convergence speed during the training of the edge detection network and low detection accuracy. Summary of the Invention
[0007] The present application provides an edge detection method and device based on bidirectional feature supervision, which has the characteristics of helping to improve the edge detection accuracy and having a fast network convergence speed during the training of the edge detection network.
[0008] According to a first aspect, an edge detection method based on bidirectional feature supervision provided in an embodiment includes:
[0009] Obtain an image to be detected;
[0010] Perform multi-level edge feature extraction processing on the image to be detected in sequence. Each level of edge feature extraction processing includes: for the input image of the current level, obtain a first edge feature map after performing the first edge feature extraction processing, and obtain a second edge feature map after performing the second edge feature extraction processing; for the input image of each non-last level, obtain a third edge feature map after performing the third edge feature extraction processing; wherein, in the edge feature extraction processing of adjacent two levels, the third edge feature map obtained by the previous-level edge feature extraction processing is used as the input image of the next-level edge feature extraction processing.
[0011] For the first edge feature map output by the edge feature extraction processing of each non-first level: fuse the first edge feature map of the current level with the first edge feature map of the previous level to obtain the first fusion feature map of the current level, and perform the first edge confidence mapping processing on the first fusion feature map of the current level to obtain the first mapping feature map of the current level.
[0012] For the first edge feature map output by the edge feature extraction processing of the first level, perform the first edge confidence mapping processing to obtain the first mapping feature map of the first level.
[0013] For the second edge feature map output by the edge feature extraction processing of each non-last level: fuse the second edge feature map of the current level with the second edge feature map of the next level to obtain the second fusion feature map of the current level, and perform the second edge confidence mapping processing on the second fusion feature map of the current level to obtain the second mapping feature map of the current level.
[0014] For the second edge feature map output by the edge feature extraction processing of the last level, perform the second edge confidence mapping processing to obtain the second mapping feature map of the last level.
[0015] Perform channel splicing and fusion processing on the first mapping feature maps and the second mapping feature maps of each level, and then perform 1×1 convolution operation to obtain the edge detection result.
[0016] In one embodiment, in the edge feature extraction processing of adjacent two levels, using the third edge feature map obtained by the previous-level edge feature extraction processing as the input image of the next-level edge feature extraction processing includes: performing pooling processing on the third edge feature map obtained by the previous-level edge feature extraction processing in the edge feature extraction processing of adjacent two levels, and using the result as the input image of the next-level edge feature extraction processing.
[0017] In one embodiment, for the input image of the current level, obtaining a first edge feature map after performing the first edge feature extraction processing and obtaining a second edge feature map after performing the second edge feature extraction processing includes:
[0018] For the first edge feature extraction process: The input image of this level is sequentially subjected to gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a first feature map; the input image of this level is sequentially subjected to gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a second feature map; the first feature map and the second feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a first edge feature map;
[0019] For the second edge feature extraction process: The input image of this level is sequentially subjected to the above-mentioned gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a second feature map; the input image of this level is sequentially subjected to the above-mentioned gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a first feature map; the second feature map and the first feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a second edge feature map;
[0020] For each input image of each non-last level, after performing the third edge feature extraction process to obtain a third edge feature map, it includes: For each input image of each non-last level, taking the first feature map as the third edge feature map, or taking the second feature map as the third edge feature map, or performing feature fusion processing on the first feature map and the second feature map and then taking it as the third edge feature map.
[0021] In one embodiment, for the input image of this level, after performing the first edge feature extraction process to obtain a first edge feature map and performing the second edge feature extraction process to obtain a second edge feature map, it includes:
[0022] For the first edge feature extraction process: The input image of this level is sequentially subjected to gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a first feature map; the input image of this level is sequentially subjected to gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a second feature map; the input image of this level is sequentially subjected to at least two 3×3 convolution processes and at least one multi-scale feature fusion process to obtain a third feature map; the first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a first edge feature map;
[0023] For the second edge feature extraction process: After the input image at this level is successively subjected to the gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing, a second feature map is obtained; After the input image at this level is successively subjected to at least two 3×3 convolution processes and at least one multi-scale feature fusion process, a third feature map is obtained; After the input image at this level is successively subjected to the gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing, a first feature map is obtained; The first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a second edge feature map;
[0024] For the input image of each non-last level, after performing the third edge feature extraction process to obtain a third edge feature map, it includes: For the input image of each non-last level, any one of the first feature map, the second feature map, or the third feature map is used as the third edge feature map, or, any two of the first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then used as the third edge feature map, or, the first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then used as the third edge feature map.
[0025] In one embodiment, for the input image at this level, after performing the first edge feature extraction process to obtain a first edge feature map and performing the second edge feature extraction process to obtain a second edge feature map, it includes:
[0026] For the edge feature extraction process of non-last levels,
[0027] For the first edge feature extraction process: The first feature map obtained after the input image at this level is successively subjected to the gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; The second feature map obtained after the input image at this level is successively subjected to the gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; The first feature map and the second feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a first edge feature map;
[0028] For the second edge feature extraction process: The second feature map obtained after the input image at this level is successively subjected to the gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; The first feature map obtained after the input image at this level is successively subjected to the gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; The second feature map and the first feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a second edge feature map;
[0029] For the input images of each level other than the last level, after performing the third edge feature extraction process, a third edge feature map is obtained, including: using the first feature map as the third edge feature map, or using the second feature map as the third edge feature map, or performing feature fusion on the first feature map and the second feature map and then using the result as the third edge feature map;
[0030] For the edge feature processing of the last level,
[0031] For the first edge feature extraction process: the input image of this level is sequentially subjected to gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a first feature map; the input image of this level is sequentially subjected to gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a second feature map; the input image of this level is sequentially subjected to at least two 3×3 convolution processes and at least one multi-scale feature fusion process to obtain a third feature map; the first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a first edge feature map;
[0032] For the second edge feature extraction process: the input image of this level is sequentially subjected to gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a second feature map; the input image of this level is sequentially subjected to at least two 3×3 convolution processes and at least one multi-scale feature fusion process to obtain a third feature map; the input image of this level is sequentially subjected to gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a first feature map; the first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a second edge feature map.
[0033] In one embodiment, the gradient difference convolution processing in the first multi-angle direction includes: dividing the corresponding pixel points within the range of a 5×5 image filtering kernel into 9 parts, namely the upper left 2×2, the upper right 2×2, the lower left 2×2, the lower right 2×2, the upper middle 1×2, the lower middle 1×2, the left middle 2×1, the right middle 2×1, and the middle point. Taking the difference between the gray values of two pixel points in the 135-degree direction of the upper left 2×2 to form the gray value of the upper left 2×2 as an upper left pixel point, taking the difference between the gray values of two pixel points in the 45-degree direction of the upper right 2×2 to form the gray value of the upper right 2×2 as an upper right pixel point, taking the difference between the gray values of two pixel points in the 45-degree direction of the lower left 2×2 to form the gray value of the lower left 2×2 as a lower left pixel point, taking the difference between the gray values of two pixel points in the 135-degree direction of the lower right 2×2 to form the gray value of the lower right 2×2 as a lower right pixel point, taking the difference between the gray values of two pixel points in the upper middle 1×2 to form the gray value of the upper middle 1×2 as an upper middle pixel point, taking the difference between the gray values of two pixel points in the lower middle 1×2 to form the gray value of the lower middle 1×2 as a lower middle pixel point, taking the difference between the gray values of two pixel points in the left middle 2×1 to form the gray value of the left middle 2×1 as a left middle pixel point, taking the difference between the gray values of two pixel points in the right middle 2×1 to form the gray value of the right middle 2×1 as a right middle pixel point, and taking the difference between the gray value of the middle point pixel and itself to form the gray value of the middle point as a middle point pixel, thereby obtaining a 3×3 feature map;
[0034] The gradient difference convolution processing in the second multi-angle direction includes: dividing the corresponding pixel points within the range of a 5×5 image filtering kernel into 9 parts, namely the upper left 2×2, upper right 2×2, lower left 2×2, lower right 2×2, upper middle 1×2, lower middle 1×2, left middle 2×1, right middle 2×1, and the middle point. Subtracting the gray values of two pixel points in the 45-degree direction of the upper left 2×2 to form the gray value of the upper left 2×2 as an upper left pixel point. Subtracting the gray values of two pixel points in the 135-degree direction of the upper right 2×2 to form the gray value of the upper right 2×2 as an upper right pixel point. Subtracting the gray values of two pixel points in the 135-degree direction of the lower left 2×2 to form the gray value of the lower left 2×2 as a lower left pixel point. Subtracting the gray values of two pixel points in the 45-degree direction of the lower right 2×2 to form the gray value of the lower right 2×2 as a lower right pixel point. Subtracting the gray values of two pixel points in the upper middle 1×2 to form the gray value of the upper middle 1×2 as an upper middle pixel point. Subtracting the gray values of two pixel points in the lower middle 1×2 to form the gray value of the lower middle 1×2 as a lower middle pixel point. Subtracting the gray values of two pixel points in the left middle 2×1 to form the gray value of the left middle 2×1 as a left middle pixel point. Subtracting the gray values of two pixel points in the right middle 2×1 to form the gray value of the right middle 2×1 as a right middle pixel point. Subtracting the gray value of the middle point pixel from itself to form the gray value of the middle point as a middle point pixel point, thereby obtaining a 3×3 feature map.
[0035] In one embodiment, the multi-scale feature fusion processing includes: obtaining a feature map of a preset channel after convolution operation on the feature map, then respectively performing atrous convolution operations with two or more different dilation rates on the obtained feature map of the preset channel, and then summing the results of the corresponding two or more atrous convolution operations; the dilation rate is set according to the requirements of the edge detection model network structure.
[0036] In one embodiment, the edge detection network on which the edge detection method is based is obtained by training in combination with a loss function based on a network architecture including edge feature extraction processing and edge confidence mapping processing; the loss function can be expressed as:
[0037]
[0038] where N represents the number of levels of edge feature extraction processing, represents the self-supervised weight of the first mapped feature map corresponding to the s-th level of edge feature extraction processing, 1≤s≤N; represents the self-supervised weight of the second mapped feature map corresponding to the s-th level of edge feature extraction processing; Y kis the edge annotation information of the k-th level, which is 1 or 0, 1 ≤ k ≤ N; P 1 i is the confidence of the first mapped feature map corresponding to the predicted edge feature extraction process of the i-th layer, 1 ≤ i ≤ N; P 2 i is the confidence of the second mapped feature map corresponding to the predicted edge feature extraction process of the i-th layer.
[0039] In one embodiment, the edge detection method further includes, for the obtained detection result, connecting a plurality of contours in the detection result through one or more morphological operations among opening operation, closing operation and thinning morphological operation.
[0040] According to the second aspect, in one embodiment, an edge detection method is provided, including sub-pixel edge extraction based on the edge detection result obtained in the above method, including:
[0041] Obtain the gray-scale gradient value n along the X-axis of the image at the center point of the pixel x and the gray-scale gradient value n along the Y-axis of the image y ;
[0042] Calculate the offset parameter t:
[0043]
[0044] where g x and g y are the first-order derivatives at the center point of the pixel, g xx , g xy and g yy are the second-order derivatives at the center point of the pixel, and
[0045]
[0046]
[0047] where represents the Kronecker product, (x I , y I ) is the image coordinate of the center point of the pixel, g(x I , y I ) is the gray-scale value of the center point of the pixel, k x , k y , k xx , k xy and k yy are preset convolution kernels;
[0048] Then the sub-pixel coordinates of the center point of the pixel are (x I ′, y I′) = (x I + tn x , y I + tn y );
[0049] Obtain the sub - pixel coordinates of the center points of each pixel to achieve sub - pixel edge extraction of the image.
[0050] In one embodiment, the preset convolution kernel can be set as follows:
[0051]
[0052]
[0053] According to the third aspect, in one embodiment, an edge detection device based on bidirectional feature supervision is provided, including
[0054] An image acquisition module, configured to acquire the image to be detected;
[0055] A feature extraction module, including multiple levels, configured to perform multi - level edge feature extraction processing on the image to be detected in sequence; Each level of edge feature extraction processing includes: for the input image of this level, after performing the first edge feature extraction processing, a first edge feature map is obtained, and after performing the second edge feature extraction processing, a second edge feature map is obtained; for the edge feature processing of other levels except the last level, after performing the third edge feature extraction processing, a third edge feature map is obtained; wherein, in the edge feature extraction processing of adjacent two levels, the third edge feature map obtained by the previous - level edge feature extraction processing is used as the input image of the next - level edge feature extraction processing.
[0056] A first edge mapping module, including multiple levels, corresponding one - to - one with the multi - level feature extraction module, configured to perform a first edge confidence mapping process on the input image;
[0057] A second edge mapping module, including multiple levels, corresponding one - to - one with the multi - level feature extraction module, configured to perform a second edge confidence mapping process on the input image;
[0058] Wherein
[0059] For the first edge feature map output by the feature extraction module of each level except the first level: the first edge feature map of this level is fused with the first edge feature map of the previous level to obtain the corresponding first fusion feature map of this level, and the first fusion feature map of this level is input into the corresponding first edge mapping module for the first edge confidence mapping process to obtain the first mapping feature map of this level;
[0060] For the feature extraction module of the first level, the first edge feature map output by performing edge feature extraction on the input image is input into the corresponding first edge mapping module for first edge confidence mapping processing to obtain the first mapping feature map of the first level;
[0061] For the second edge feature map output by the feature extraction module of each other level except the last level: The second edge feature map of this level is fused with the second edge feature map of the next level to obtain the corresponding second fusion feature map of this level, and the second fusion feature map of this level is input into the corresponding second edge mapping module for second edge confidence mapping processing to obtain the second mapping feature map of this level;
[0062] For the feature extraction processing module of the last level, the second edge feature map output by performing edge feature extraction on the input image is input into the corresponding second edge mapping module for second edge confidence mapping processing to obtain the second mapping feature map of the last level;
[0063] The channel splicing module is used to perform channel splicing and fusion processing on the first mapping feature maps and the second mapping feature maps of each level;
[0064] The 1×1 convolution module is used to perform 1×1 convolution operation on the feature map after channel splicing and fusion processing to obtain the edge detection result.
[0065] According to the fourth aspect, in an embodiment, a computer storage medium is provided. A program is stored on the medium, and the program can be executed by a processor to implement any of the above-mentioned edge detection methods based on bidirectional feature supervision and / or any of the above-mentioned edge detection methods.
[0066] The beneficial effects of this application are as follows: The image to be detected is subjected to multi-level edge feature extraction processing, and both the first edge feature extraction processing and the second edge feature extraction processing are performed at each level. Hierarchical supervision is used to achieve supervision of different scales. For the first edge feature map obtained by the first edge feature extraction processing at each level, through low-level to high-level supervision, the low level can predict details and the high level can predict the global. For the second edge feature map obtained by the second edge feature extraction processing at each level, through high-level to low-level supervision, the high level can predict details and the low level can predict the global. Through the above two-way supervision modes from low to high and from high to low, the network convergence speed in the training process of the edge detection network is improved, and the edge detection accuracy of the trained edge detection network is also improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is a schematic flowchart of an edge detection method in an embodiment of this application;
[0068] Figure 2 Schematic diagram of the principle of an edge detection method in an embodiment of the present application;
[0069] Figure 3 Schematic diagram of the principle of the first edge feature extraction process and the second edge feature extraction process in an embodiment of the present application;
[0070] Figure 4 Schematic diagram of the principle of the first edge feature extraction process and the second edge feature extraction process in an embodiment of the present application;
[0071] Figure 5 Schematic diagram of the 5×5 convolution gray range in an embodiment of the present application;
[0072] Figure 6 Schematic diagram of the 5×5 convolution area division in an embodiment of the present application;
[0073] Figure 7 Schematic diagram of the processing process of obtaining a 3×3 feature map by performing difference in one direction in the Roberts gradient difference convolution process in an embodiment of the present application;
[0074] Figure 8 Schematic diagram of the processing process of obtaining another 3×3 feature map by performing difference in another direction in the Roberts gradient difference convolution process in an embodiment of the present application;
[0075] Figure 9 Schematic diagram of the multi-scale feature fusion processing flow in an embodiment of the present application;
[0076] Figure 10 Schematic diagram of the principle of the multi-scale feature fusion processing in an embodiment of the present application;
[0077] Figure 11 Schematic diagram of the principle of constructing the loss function in an embodiment of the present application.
[0078] Figure 12 Block diagram of the edge detection device structure in an embodiment of the present application. Detailed implementation manners
[0079] The present application will be further described in detail below in conjunction with the specific embodiments and the accompanying drawings. Similar elements in different embodiments are labeled with related similar element numbers. In the following embodiments, many detailed descriptions are provided to enable a better understanding of the present application. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, or methods. In some cases, some operations related to the present application are not shown or described in the specification, which is to avoid overwhelming the core part of the present application with excessive descriptions. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the descriptions in the specification and the general technical knowledge in the art.
[0080] In addition, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can also be reordered or adjusted in a manner obvious to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for clearly describing a certain embodiment and do not mean a necessary sequence, unless it is stated that a certain sequence must be followed.
[0081] The serial numbers assigned to the components herein, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meaning.
[0082] For edge detection, currently, a supervised mode based on annotation is adopted. That is, when training an edge detection network to construct a loss function, the loss function is constructed based on the annotation information of the input images in the training set. This supervised mode has a great impact on the position accuracy of different scales, resulting in, on the one hand, a slow network convergence speed during the training of the edge detection network, and on the other hand, a low edge detection accuracy achieved by the trained edge detection network.
[0083] To reduce the impact of the traditional supervised mode on the position accuracy of different scales, the present application provides an edge detection method based on bidirectional feature supervision. Please refer to Figure 1 and Figure 2 , and this edge detection method includes:
[0084] Step 101: Obtain the image to be detected.
[0085] It can be understood that for the image to be detected, before performing edge detection, those skilled in the art can decide whether to perform preprocessing and how to perform preprocessing according to the requirements.
[0086] Step 102: Perform multi-level edge feature extraction processing and two-way edge confidence mapping processing on the image to be detected to obtain respective mapped feature maps. Specifically, it includes:
[0087] Perform multi-level edge feature extraction processing on the image to be detected in sequence. Each level of edge feature extraction processing includes: for the input image of the current level, obtain the first edge feature map after performing the first edge feature extraction processing, and obtain the second edge feature map after performing the second edge feature extraction processing; for the input images of every other level except the last level, obtain the third edge feature map after performing the third edge feature extraction processing; wherein, in the edge feature extraction processing of adjacent two levels, the third edge feature map obtained by the previous level of edge feature extraction processing is used as the input image of the next level of edge feature extraction processing.
[0088] For the first edge feature map output by the edge feature extraction processing of every other level except the first level: fuse the first edge feature map of the current level with the first edge feature map of the previous level to obtain the first fused feature map of the current level, and perform the first edge confidence mapping processing on the first fused feature map of the current level to obtain the first mapped feature map of the current level; for the first edge feature map output by the edge feature extraction processing of the first level, perform the first edge confidence mapping processing to obtain the first mapped feature map of the first level.
[0089] For the second edge feature map output by the edge feature extraction processing of every other level except the last level: fuse the second edge feature map of the current level with the second edge feature map of the next level to obtain the second fused feature map of the current level, and perform the second edge confidence mapping processing on the second fused feature map of the current level to obtain the second mapped feature map of the current level; for the second edge feature map output by the edge feature extraction processing of the last level, perform the second edge confidence mapping processing to obtain the second mapped feature map of the last level.
[0090] In the confidence mapping processing, in order to obtain the confidence considered to be an edge, it is desired to map to a score within the range of 0 to 1, which can be achieved by using fully connected and sigmoid operations, or other mapping methods, such as fully connected and tanh operations.
[0091] Using multi-scale representations is crucial for improving edge detection of objects at different scales. In this application, multi-level edge feature extraction processing is performed on the image to be detected, and for each level, first-edge feature extraction processing and second-edge feature extraction processing are carried out. Hierarchical supervision is used to achieve supervision of different scales. For the first-edge feature map obtained from the first-edge feature extraction processing at each level, low-level downsampling is used to perform high-level supervision. The low level can predict details, and the high level can predict the global situation. For the second-edge feature map obtained from the second-edge feature extraction processing at each level, high-level upsampling is used to perform low-level supervision. The high level can predict details, and the low level can predict the global situation. Through the above two-way supervision modes from low to high and from high to low, the network convergence speed during the training of the edge detection network is improved, and the accuracy of edge detection of the trained edge detection network is also improved.
[0092] Step 103: Perform channel splicing processing on each mapping feature map to obtain a channel-fused feature map; that is, perform channel splicing and fusion processing on the first mapping feature map and the second mapping feature map at each level to obtain a channel-fused feature map.
[0093] Step 104: Perform 1×1 convolution operation (denoted by Conv in the attached figure) on the channel-fused feature map to obtain the edge detection result.
[0094] Therefore, the multi-scale edge detection method adopted in this application, combined with the above-mentioned two-way full-directional cascade network supervision mode from low to high and from high to low, not only helps to improve the network convergence speed during the training of the edge detection network, but also helps to improve the accuracy that can be achieved in detecting and analyzing the edges and contours of objects. It can be understood that although Figure 2 in the illustrated embodiment, four-level edge feature extraction processing and the corresponding full-directional cascade network are adopted, those skilled in the art can make equivalent transformations based on this solution, and all equivalent transformations based on it are within the protection scope of this application.
[0095] To increase the receptive field, in one embodiment, in the edge feature extraction processing of adjacent two levels, the third-edge feature map obtained from the previous-level edge feature extraction processing is used as the input image for the next-level edge feature extraction processing, including performing pooling processing on the third-edge feature map obtained from the previous-level edge feature extraction processing in the edge feature extraction processing of adjacent two levels and using it as the input image for the next-level edge feature extraction processing.
[0096] To enrich the supervision mode of the omnidirectional cascaded network, existing methods can be adopted for the multi-scale feature generation method. This application also provides the following three new methods, which are described below. It can be understood that those skilled in the art can make equivalent transformations based on the following three new methods, and all equivalent transformations based thereon are within the protection scope of this application.
[0097] For the input image of the current level, after the first edge feature extraction process to obtain the first edge feature map and after the second edge feature extraction process to obtain the second edge feature map, it may include:
[0098] Method 1, please refer to Figure 3 .
[0099] For the first edge feature extraction process: The first feature map obtained by sequentially subjecting the input image of the current level to gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; the second feature map obtained by sequentially subjecting the input image of the current level to gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; the first feature map and the second feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain the first edge feature map.
[0100] For the second edge feature extraction process: The second feature map and the first feature map obtained in the first edge feature extraction process are subjected to feature fusion processing and then 1×1 convolution processing to obtain the second edge feature map.
[0101] Method 2, please refer to Figure 4 .
[0102] For the first edge feature extraction process: The first feature map obtained by sequentially subjecting the input image of the current level to gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; the second feature map obtained by sequentially subjecting the input image of the current level to gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; the third feature map obtained by sequentially subjecting the input image of the current level to at least two 3×3 convolution processes and at least one multi-scale feature fusion process; the first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain the first edge feature map.
[0103] For the second edge feature extraction process: The first feature map, the second feature map, and the third feature map obtained in the above first edge feature extraction process are subjected to feature fusion processing and then 1×1 convolution processing to obtain the second edge feature map.
[0104] Method 3, in combination with Figure 3 andFigure 4 。
[0105] The first edge feature extraction process and the second edge feature extraction process of non - the last level follow the solution of Method 1, and the first edge feature extraction process and the second edge feature extraction process of the last level follow the solution of Method 2.
[0106] Among them, in Method 3, the first edge feature extraction process and the second edge feature extraction process of non - the last level follow the solution of Method 1, which can improve the convergence speed of the network during training and the accuracy of edge detection on the basis of ensuring sufficient feature extraction ability; the first edge feature extraction process and the second edge feature extraction process of the last level follow the solution of Method 2 to enhance the feature extraction ability of the last level.
[0107] Based on the above three methods, for the third edge feature map, in Method 1, for the input images of each level other than the last level, after performing the third edge feature extraction process, the third edge feature map is obtained, including: for the input images of each level other than the last level, taking the first feature map as the third edge feature map, or taking the second feature map as the third edge feature map, or taking the first feature map and the second feature map after feature fusion processing as the third edge feature map.
[0108] Among them, the third edge feature map after fusing the first feature map and the second feature map has stronger feature expression ability, higher detection stability and accuracy, and the fusion also increases the feature extraction and expression ability of edges in different directions.
[0109] In Method 2, for the input images of each level other than the last level, after performing the third edge feature extraction process, the third edge feature map is obtained, including: for the input images of each level other than the last level, taking any one of the first feature map, the second feature map or the third feature map as the third edge feature map, or taking any two of the first feature map, the second feature map and the third feature map after feature fusion processing as the third edge feature map, or taking the first feature map, the second feature map and the third feature map after feature fusion processing as the third edge feature map.
[0110] Similarly, the third edge feature map after fusion processing has stronger feature expression ability, higher detection stability and accuracy, and the fusion also increases the feature extraction and expression ability of edges in different directions.
[0111] In Method 3, for the third edge feature map, for the input images of each level other than the last level, the third edge feature map is obtained after performing the third edge feature extraction process in the same processing method as in Method 1.
[0112] Use any of the above three methods to generate multi-scale features and enrich the omnidirectional cascade network, rather than using a deeper convolutional neural network. This method encourages learning multi-scale representations at different levels, detecting their edges, and depicting the edges according to multiple scales. Additionally, a compact network with fractional parameters can be generated.
[0113] Next, for the first multi-angle direction gradient difference convolution processing and the second multi-angle direction gradient difference convolution processing among the above three methods; it should be noted that those skilled in the art can adopt existing methods, and this application also provides the following new methods, which will be described below. It can be understood that those skilled in the art can make equivalent transformations based on the following methods, and all equivalent transformations based on them are within the protection scope of this application.
[0114] For the first multi-angle direction gradient difference convolution processing, please refer to Figure 5 , it can be that the corresponding pixel points within the 5×5 image filter kernel are divided into the upper left 2×2 (i.e., the area where x Figure 5 、x 1 、x 2 、x 6 and x 7 are located), the upper right 2×2 (i.e., the area where x Figure 5 、x 4 、x 5 、x 9 and x 10 are located), the lower left 2×2 (i.e., the area where x Figure 5 、x 16 、x 17 、x 21 and x 22 are located), the lower right 2×2 (i.e., the area where x Figure 5 、x 19 、x 20 、x 24 and x 25 are located), the upper middle 1×2 (i.e., the area where x Figure 5 、x 3 and x 8 are located), the lower middle 1×2 (i.e., the area where x Figure 5 、x 18 and x 23 are located), the left middle 2×1 (i.e., the area where x Figure 5 、x 11 and x 12 are located), the right middle 2×1 (i.e., the area where x Figure 5 、x 14 and x 15 are located) and the middle point (i.e., x Figure 5 in 13A total of 9 parts in the area where it is located), and these 9 parts correspond one by one Figure 6 to the 9 corresponding pixel area ranges of A, C, G, I, B, H, D, F, and E shown
[0115] Please refer to Figure 7 , subtract the grayscale values of the two pixels in the 135-degree direction of the upper-left 2×2 to form the grayscale value of the upper-left 2×2 as a single upper-left pixel, that is, take the area where x 1 , x 2 , x 6 and x 7 are located as a single pixel, and use the difference between x Figure 5 and x 1 in 7 as the grayscale value of this pixel; for the upper-right 2×2, subtract the grayscale values of the two pixels in the 45-degree direction to form the grayscale value of the upper-right 2×2 as a single upper-right pixel, that is, take the area where x 4 , x 5 , x 9 and x 10 are located as a single pixel, and use the difference between x Figure 5 and x 5 in 9 as the grayscale value of this pixel; for the lower-left 2×2, subtract the grayscale values of the two pixels in the 45-degree direction to form the grayscale value of the lower-left 2×2 as a single lower-left pixel, that is, take the area where x 16 , x 17 , x 21 and x 22 are located as a single pixel, and use the difference between x Figure 5 and x 17 in 21 as the grayscale value of this pixel; for the lower-right 2×2, subtract the grayscale values of the two pixels in the 135-degree direction to form the grayscale value of the lower-right 2×2 as a single lower-right pixel, that is, take the area where x 19 , x 20 , x 24 and x 25 are located as a single pixel, and use the difference between x Figure 5 and x 19 in 25 as the grayscale value of this pixel; for the upper-middle 1×2, subtract the grayscale values of the two pixels to form the grayscale value of the upper-middle 1×2 as a single upper-middle pixel, that is, take the area where x 3 and x 8 are located as a single pixel, and use the difference between x Figure 5 and x 3 in 8The difference is used as the grayscale value of this pixel; the grayscale values of the two pixels in the lower middle 1×2 are subtracted to form the grayscale value of the lower middle 1×2 as a single lower middle pixel, that is, the area where x 18 and x 23 are located is regarded as a pixel, and Figure 5 the difference between x 23 -x 18 in is used as the grayscale value of this pixel; the grayscale values of the two pixels in the left middle 2×1 are subtracted to form the grayscale value of the left middle 2×1 as a single left middle pixel, that is, the area where x 11 and x 12 are located is regarded as a pixel, and Figure 5 the difference between x 11 -x 12 in is used as the grayscale value of this pixel; the grayscale values of the two pixels in the right middle 2×1 are subtracted to form the grayscale value of the right middle 2×1 as a single right middle pixel, that is, the area where x 14 and x 15 are located is regarded as a pixel, and Figure 5 the difference between x 15 -x 14 in is used as the grayscale value of this pixel; the grayscale value of the middle point pixel is subtracted from itself to form the grayscale value of the middle point as a single middle point pixel, that is, the area where x 13 is located is regarded as a pixel, and the difference between x 13 -x 13 is 0, which is used as the grayscale value of this pixel; according to the above gradient difference, a 3×3 feature map is obtained; it should be noted that the minuend and the subtrahend in the subtraction here can be interchanged, and the grayscale value is taken as the absolute value of the difference after subtraction.
[0116] For the second multi-angle direction gradient difference convolution processing, similar to the above one-direction multi-angle direction gradient difference convolution, referring to Figure 5 , the corresponding pixels within the 5×5 image filter kernel are divided into 9 parts, and these 9 parts correspond one by one to Figure 6 the 9 corresponding pixel area ranges of A, C, G, I, B, H, D, F, and E shown. Please refer to Figure 8 , the grayscale values of the two pixels in the 45-degree direction of the upper left 2×2 are subtracted to form the grayscale value of the upper left 2×2 as a single upper left pixel, that is, the area where x 1 , x 2 , x 6 and x 7 are located is regarded as a pixel, and Figure 5 the difference between x 2 -x 6The difference is used as the grayscale value of this pixel; the difference between the grayscale values of the two pixels in the 135-degree direction of the upper-right 2×2 is used to form the grayscale value of the upper-right 2×2 as a single upper-right pixel, that is, taking x 4 、x 5 、x 9 and x 10 's area as a single pixel, and taking the difference between x Figure 5 in x 4 -x 10 as the grayscale value of this pixel; the difference between the grayscale values of the two pixels in the 135-degree direction of the lower-left 2×2 is used to form the grayscale value of the lower-left 2×2 as a single lower-left pixel, that is, taking x 16 、x 17 、x 21 and x 22 's area as a single pixel, and taking the difference between x Figure 5 in x 16 -x 22 as the grayscale value of this pixel; the difference between the grayscale values of the two pixels in the 45-degree direction of the lower-right 2×2 is used to form the grayscale value of the lower-right 2×2 as a single lower-right pixel, that is, taking x 19 、x 20 、x 24 and x 25 's area as a single pixel, and taking the difference between x Figure 5 in x 20 -x 24 as the grayscale value of this pixel; the difference between the grayscale values of the two pixels in the upper-middle 1×2 is used to form the grayscale value of the upper-middle 1×2 as a single upper-middle pixel, that is, taking x 3 and x 8 's area as a single pixel, and taking the difference between x Figure 5 in x 3 -x 8 as the grayscale value of this pixel; the difference between the grayscale values of the two pixels in the lower-middle 1×2 is used to form the grayscale value of the lower-middle 1×2 as a single lower-middle pixel, that is, taking x 18 and x 23 's area as a single pixel, and taking the difference between x Figure 5 in x 23 -x 18 as the grayscale value of this pixel; the difference between the grayscale values of the two pixels in the left-middle 2×1 is used to form the grayscale value of the left-middle 2×1 as a single left-middle pixel, that is, taking x 11 and x 12 's area as a single pixel, and taking the difference between x Figure 5 in x 11 -x 12The difference is used as the grayscale value of this pixel; the difference between the grayscale values of the two pixels in the 2×1 area in the middle right is formed as the grayscale value of this 2×1 area in the middle right as a middle right pixel, that is, taking the area where x 14 and x 15 are located as a pixel, and taking the difference between Figure 5 x 15 -x 14 in Figure 5 as the grayscale value of this pixel; the difference between the grayscale value of the middle point pixel and itself is formed as the grayscale value of this middle point as a middle point pixel, that is, taking the area where x 13 is located as a pixel, and taking the difference between x 13 -x 13 which is 0 as the grayscale value of this pixel; according to the above gradient difference, the other 3×3 feature map is obtained; it should be noted that in the difference here, the minuend and the subtrahend of the two numbers can be interchanged, and the grayscale value is taken as the absolute value of the difference.
[0117] It should be noted that for the above two multi-angle direction gradient difference convolutions, any one of the first multi-angle direction gradient difference convolution processes can be selected, and the second multi-angle direction gradient difference convolution process can select another one, and the same is true for the edge feature extraction process at each level, that is, it is not necessary that the first multi-angle direction gradient difference convolution process for the edge feature extraction at each level all adopt the same multi-angle direction gradient difference convolution process, nor is it necessary that the second angle direction gradient difference convolution process for the edge feature extraction at each level all adopt the same multi-angle direction gradient difference convolution process.
[0118] In traditional image processing, the edge detection uses a filter kernel to extract edges by finding the boundary pixels of brighter and darker regions. The filter kernel finds the parts with obvious gradient changes in the image. Currently, the relatively standard filters include the Sobel filter kernel considering horizontal and vertical directions:
[0119]
[0120] The Prewitt filter kernel considering horizontal and vertical directions:
[0121]
[0122] The Laplace filter kernel considering diagonal directions:
[0123]
[0124] The Roberts filter kernel for detecting edge information in two diagonal directions:
[0125]
[0126] The above filtering kernels only consider the horizontal and vertical directions or the diagonal direction. The edge detection directions considered are relatively single, only using the gradient features of the edges, lacking the ability to express global features, and the ability to detect edges is relatively weak. In this application, by using the above two gradient difference convolutions in multiple angular directions, differential operations in multiple angular directions can be incorporated, improving the edge detection ability for edges in each angle. In addition, considering the gray range of 5×5 and then converting the 5×5 gray range into eigenvalue of 3×3 can consider a larger gray range, reduce the influence of noise, and increase the receptive field at the same time.
[0127] In one embodiment, the multi-scale feature fusion processing can use existing methods. This application also provides a new multi-scale feature fusion processing method. Please refer to Figure 9 and Figure 10 , including;
[0128] Step 201: Obtain a feature map of a preset channel after performing a convolution operation on the input image.
[0129] Step 202: Respectively perform atrous convolution operations with more than two different dilation rates on the obtained feature map of the preset channel. The dilation rate is set according to the requirements of the edge detection network.
[0130] In Figure 10 the embodiment, atrous convolution operations with 4 different dilation rates are respectively performed on the obtained feature map of the preset channel. The 4 different dilation rates are 3, 5, 7, and 9 respectively. It should be noted that these 4 dilation rates are only an example. In other applications and / or instances, the dilation rate can be specifically set according to the requirements of the edge detection model network.
[0131] Step 203: Sum the results of the above atrous convolution operations with more than two different dilation rates to obtain a first feature map or a second feature map.
[0132] In one embodiment, the edge detection network on which the edge detection method is based is a network architecture composed of edge feature extraction processing and edge confidence mapping processing, and is obtained by training in combination with a loss function; the construction method of the loss function can use the existing construction method of the loss function. This application also provides a new construction method of the loss function. Please refer to Figure 11 , which can be expressed as:
[0133]
[0134] where N represents the number of levels of edge feature extraction processing, represents the self-supervised weight of the first mapped feature map corresponding to the s-th level of edge feature extraction processing, representing top-down supervision, 1≤s≤N; Represents the self-supervised weight of the second mapped feature map corresponding to the s-th level of edge feature extraction processing, indicating bottom-up supervision; Y k Is the edge information of the k-th level annotation, which is 1 or 0, 1 ≤ k ≤ N; P 1 i Is the confidence of the first mapped feature map corresponding to the predicted i-th level of edge feature extraction processing, 1 ≤ i ≤ N; Is the confidence of the second mapped feature map corresponding to the predicted i-th level of edge feature extraction processing.
[0135] In one embodiment, to improve the accuracy and continuity of edge detection, the edge detection method further includes, for the obtained detection result, performing a morphological operation on one or more of opening operation, closing operation, and thinning morphological operation to connect multiple contours in the detection result.
[0136] In scenarios such as human key point extraction, vehicle tracking, and instance segmentation, it is also necessary to obtain the contour morphology information of the object for subsequent applications in scenarios such as position estimation and VR games. The edges obtained by current edge detection are basically pixel-level edges, and the accuracy of edge points cannot meet the applications in scenarios such as high-precision measurement and pose estimation.
[0137] To improve the edge detection accuracy, in some embodiments of the present application, second-order surface fitting is used to obtain the sub-pixel level accuracy of the edge. Using the sub-pixel method can further improve the edge detection accuracy, which has special significance in scenarios such as some high-precision applications in semiconductors and panel displays.
[0138] The sub-pixel coordinates of the pixel center point can be obtained using existing sub-pixel coordinate calculation methods. The present application also provides a new sub-pixel coordinate calculation method, which will be described below.
[0139] For the gray value of any point in the image, performing a second-order approximation estimate on it can be expressed as:
[0140]
[0141] Where (x I , y I ) represents the image coordinates, g(x I , y I ) is the gray value of the point (x I , y I ), g 0 Is a constant, g x And g y Are the first-order derivatives at the point (x I , y I ), g xx , gxy and g yy is the second derivative at the point (x I , y I ).
[0142] Therefore, the second-order approximate estimate of the gray value of the sub-pixel can be expressed as:
[0143]
[0144] where t is the offset parameter, n x and n y are the gray gradient values along the X-axis and Y-axis of the image at the point (x I , y I ), respectively.
[0145] Let its partial derivative be equal to 0, and we can get:
[0146]
[0147] We can get: Then the sub-pixel coordinates of the point (x I , y I ) are (x′ I , y′ I ) = (x I + tn x , y I + tn y ).
[0148] Therefore, for the pixel center point with image coordinates (x I , y I ), we can first obtain the gray gradient value n I along the X-axis of the image at the pixel center point (x I , y x and the gray gradient value n y along the Y-axis of the image; the gray gradient value n x along the X-axis of the image and the gray gradient value n y along the Y-axis of the image can be obtained by solving the eigenvector corresponding to the minimum eigenvalue of the second derivative matrix ; then the offset parameter t is calculated according to the following formula:
[0149]
[0150] where g x and g y are the first derivatives at the pixel center point, g xx , g xy and g yy are the second derivatives at the pixel center point, and
[0151]
[0152]
[0153] wherein represents the Kronecker product, and g(x I , y I ) is the grayscale value of the center point of the pixel, and k x , k y , k xx , k xy , and k yy are preset convolution kernels, and the convolution kernels can be set according to actual needs.
[0154] Then, the sub-pixel coordinates of the center point of the pixel (x I , y I ) are (x′ I , y′ I ) = (x I + tn x , y I + tn y ).
[0155] Obtain the sub-pixel coordinates of the center points of each pixel to achieve sub-pixel edge extraction of the image.
[0156] In one embodiment, the preset convolution kernel can be set as follows:
[0157]
[0158]
[0159] In an embodiment of the present application, an edge detection device based on bidirectional feature supervision is provided. Please refer to Figure 2 . Only Figure 12 is taken as an example for illustration in the embodiments of the present application, which does not mean that the present application is limited thereto. As shown in Figure 12 , the edge detection device of this embodiment includes an image acquisition module 01, a feature extraction module 02, a first edge mapping module 03, a second edge mapping module 04, a channel splicing module 05, and a 1×1 convolution module 06.
[0160] Among them, the image acquisition module 01 is used to acquire the image to be detected; the feature extraction module 02 has multiple levels. Except for the last-level feature extraction module, each level of the feature extraction module is used to perform edge feature extraction processing on the input image. After the first edge feature extraction processing, the first edge feature map is obtained. After the second edge feature processing, the second edge feature map is obtained. After the third edge feature processing, the third edge feature map is obtained, and the third edge feature map is used as the input of the next-level feature extraction module. For the last-level feature extraction module, after the first edge feature extraction processing on the input image, the first edge feature map is obtained, and after the second edge feature processing, the second edge feature map is obtained.
[0161] The first edge mapping module 03 includes multiple levels, corresponding one-to-one with the multiple levels of the feature extraction module, and is used to perform the first edge mapping degree mapping processing on the input image. The second edge mapping module also includes multiple levels and is used to perform the second edge mapping execution degree processing on the input image.
[0162] Among them, for the first edge feature map output by each level of the feature extraction module other than the first level: the first edge feature map of this level is fused with the first edge feature map of the previous level to obtain the first fusion feature map corresponding to this level, and the first fusion feature map of this level is input into the corresponding first edge mapping module for the first edge confidence degree mapping processing to obtain the first mapping feature map of this level; for the first-level feature extraction module, the first edge feature map output after the edge feature extraction processing of the input image is input into the corresponding first edge mapping module for the first edge confidence degree mapping processing to obtain the first mapping feature map of the first level.
[0163] For the second edge feature map output by each level of the feature extraction module other than the last level: the second edge feature map of this level is fused with the second edge feature map of the next level to obtain the second fusion feature map corresponding to this level, and the second fusion feature map of this level is input into the corresponding second edge mapping module for the second edge confidence degree mapping processing to obtain the second mapping feature map of this level; for the last-level feature extraction processing module, the second edge feature map output after the edge feature extraction processing of the input image is input into the corresponding second edge mapping module for the second edge confidence degree mapping processing to obtain the second mapping feature map of the last level.
[0164] The channel splicing module 05 is used to perform channel splicing and fusion processing on the first mapping feature maps and the second mapping feature maps at all levels; the 1×1 convolution module 06 is used to perform 1×1 convolution operation on the feature map after the channel splicing and fusion processing to obtain the edge detection result.
[0165] Based on the above edge detection device, the image to be detected is subjected to multi-level edge feature extraction processing, and the first edge feature extraction processing and the second edge feature extraction processing are performed at each level. The hierarchical supervision method is used to realize the supervision of different scales. For the first edge feature map obtained by the first edge feature extraction processing at each level, through the supervision from the low level to the high level, the low level can predict details and the high level can predict the global situation. For the second edge feature map obtained by the second edge feature extraction processing at each level, through the supervision from the high level to the low level, the high level can predict details and the low level can predict the global situation. Through the above two-way supervision modes from low to high and from high to low, the network convergence speed in the training process of the edge detection network is improved, and the edge detection accuracy of the trained edge detection network is also improved.
[0166] In order to increase the receptive field, in one embodiment, a pooling module is further included, which is used to process the third edge feature map obtained by the above feature extraction processing module of the non-last level and use it as the input of the next-level feature extraction module.
[0167] In one embodiment of the present application, a computer storage medium is provided, on which a program is stored, and the program can be executed by a processor to implement any of the above edge detection methods based on bidirectional feature supervision, and / or any of the above edge detection methods.
[0168] Those skilled in the art can understand that all or part of the functions of the above methods can be implemented in a hardware manner or in a computer program manner. When all or part of the functions in the above embodiments are implemented in a computer program manner, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, magnetic disk, optical disk, hard disk, etc. The above functions are realized by the computer executing the program. For example, the program is stored in the memory of the device, and when the processor executes the program in the memory, the above all or part of the functions can be realized. In addition, when all or part of the functions in the above embodiments are implemented in a computer program manner, the program can also be stored in a storage medium such as a server, another computer, magnetic disk, optical disk, flash drive or mobile hard disk, downloaded or copied and saved to the memory of the local device, or the system of the local device is updated. When the processor executes the program in the memory, the above all or part of the functions in the above embodiments can be realized.
[0169] The above uses specific examples to elaborate on the present application, which is only used to help understand the technical solution of the present application and does not limit the present application. For those skilled in the art in the technical field, according to the idea of the present application, several simple deductions, deformations or substitutions can also be made.
Claims
1. An edge detection method based on bidirectional feature supervision, characterized in that, it includes, obtaining an image to be detected; performing multi-level edge feature extraction processing on the image to be detected in sequence. Each level of edge feature extraction processing includes: for the input image of this level, obtaining a first edge feature map after performing the first edge feature extraction processing, and obtaining a second edge feature map after performing the second edge feature extraction processing; for the input image of each non-last level, obtaining a third edge feature map after performing the third edge feature extraction processing; wherein, in the edge feature extraction processing of adjacent two levels, the third edge feature map obtained by the previous-level edge feature extraction processing is used as the input image of the next-level edge feature extraction processing; For the first edge feature map output by the edge feature extraction processing of each non-first level: fusing the first edge feature map of this level with the first edge feature map of the previous level to obtain the first fusion feature map of this level, and performing the first edge confidence mapping processing on the first fusion feature map of this level to obtain the first mapping feature map of this level; For the first edge feature map output by the edge feature extraction processing of the first level, performing the first edge confidence mapping processing to obtain the first mapping feature map of the first level; For the second edge feature map output by the edge feature extraction processing of each non-last level: fusing the second edge feature map of this level with the second edge feature map of the next level to obtain the second fusion feature map of this level, and performing the second edge confidence mapping processing on the second fusion feature map of this level to obtain the second mapping feature map of this level; For the second edge feature map output by the edge feature extraction processing of the last level, performing the second edge confidence mapping processing to obtain the second mapping feature map of the last level; Performing channel splicing and fusion processing on the first mapping feature maps and the second mapping feature maps of each level, and then performing 1×1 convolution operation to obtain the edge detection result.
2. The method according to claim 1, characterized in that, in the edge feature extraction processing of adjacent two levels, using the third edge feature map obtained by the previous-level edge feature extraction processing as the input image of the next-level edge feature extraction processing includes: performing pooling processing on the third edge feature map obtained by the previous-level edge feature extraction processing in the edge feature extraction processing of adjacent two levels, and using the result as the input image of the next-level edge feature extraction processing.
3. The method according to claim 1, characterized in that, for the input image of this level, obtaining a first edge feature map after performing the first edge feature extraction processing, and obtaining a second edge feature map after performing the second edge feature extraction processing, includes: For the first edge feature extraction process: The input image of this level is sequentially subjected to gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a first feature map; The input image of this level is sequentially subjected to gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a second feature map; The first feature map and the second feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a first edge feature map; For the second edge feature extraction process: The input image of this level is sequentially subjected to the above-mentioned gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a second feature map; The input image of this level is sequentially subjected to the above-mentioned gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a first feature map; The second feature map and the first feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a second edge feature map; For the input images of each non-last level, after performing the third edge feature extraction process to obtain a third edge feature map, it includes: For the input images of each non-last level, using the first feature map as the third edge feature map, or using the second feature map as the third edge feature map, or using the first feature map and the second feature map after feature fusion processing as the third edge feature map.
4. The method according to claim 1, wherein, For the input image of this level, after performing the first edge feature extraction process to obtain a first edge feature map and performing the second edge feature extraction process to obtain a second edge feature map, it includes: For the first edge feature extraction process: The input image of this level is sequentially subjected to gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a first feature map; The input image of this level is sequentially subjected to gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a second feature map; The input image of this level is sequentially subjected to at least two 3×3 convolution processes and at least one multi-scale feature fusion process to obtain a third feature map; The first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a first edge feature map; For the second edge feature extraction process: After the input image of this level is successively subjected to the gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing, a second feature map is obtained; After the input image of this level is successively subjected to at least two 3×3 convolution processings and at least one multi-scale feature fusion processing, a third feature map is obtained; After the input image of this level is successively subjected to the gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing, a first feature map is obtained; After the first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then 1×1 convolution processing, a second edge feature map is obtained; For the input images of every other level except the last level, after the third edge feature extraction process, a third edge feature map is obtained, including: For the input images of every other level except the last level, any one of the first feature map, the second feature map, or the third feature map is used as the third edge feature map, or, any two of the first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then used as the third edge feature map, or, the first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then used as the third edge feature map.
5. The method according to claim 1, characterized in that, For the input image of this level, after the first edge feature extraction process, a first edge feature map is obtained, and after the second edge feature extraction process, a second edge feature map is obtained, including: For the edge feature extraction process of every other level except the last level, For the first edge feature extraction process: The first feature map obtained by successively subjecting the input image of this level to the gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; The second feature map obtained by successively subjecting the input image of this level to the gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; After the first feature map and the second feature map are subjected to feature fusion processing and then 1×1 convolution processing, a first edge feature map is obtained; For the second edge feature extraction process: The second feature map obtained by successively subjecting the input image of this level to the gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; The first feature map obtained by successively subjecting the input image of this level to the gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing; After the second feature map and the first feature map are subjected to feature fusion processing and then 1×1 convolution processing, a second edge feature map is obtained; For the input images of every other level except the last level, after the third edge feature extraction process, a third edge feature map is obtained, including: Using the first feature map as the third edge feature map, or, using the second feature map as the third edge feature map, or, subjecting the first feature map and the second feature map to feature fusion processing and then using it as the third edge feature map; For the edge feature processing of the last level, For the first edge feature extraction process: The input image at this level is sequentially subjected to gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a first feature map; The input image at this level is sequentially subjected to gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a second feature map; The input image at this level is sequentially subjected to at least two 3×3 convolution processes and at least one multi-scale feature fusion process to obtain a third feature map; The first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a first edge feature map; For the second edge feature extraction process: The input image at this level is sequentially subjected to gradient difference convolution processing in the second multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a second feature map; The input image at this level is sequentially subjected to at least two 3×3 convolution processes and at least one multi-scale feature fusion process to obtain a third feature map; The input image at this level is sequentially subjected to gradient difference convolution processing in the first multi-angle directions, 3×3 convolution processing, and multi-scale feature fusion processing to obtain a first feature map; The first feature map, the second feature map, and the third feature map are subjected to feature fusion processing and then 1×1 convolution processing to obtain a second edge feature map.
6. The method according to any one of claims 3 to 5, characterized in that, The gradient difference convolution processing in the first multi-angle direction includes: dividing the corresponding pixel points within the range of a 5×5 image filtering kernel into 9 parts, namely the upper left 2×2, the upper right 2×2, the lower left 2×2, the lower right 2×2, the upper middle 1×2, the lower middle 1×2, the left middle 2×1, the right middle 2×1, and the middle point. Subtracting the gray values of two pixel points in the 135-degree direction of the upper left 2×2 to form the gray value of the upper left 2×2 as an upper left pixel point. Subtracting the gray values of two pixel points in the 45-degree direction of the upper right 2×2 to form the gray value of the upper right 2×2 as an upper right pixel point. Subtracting the gray values of two pixel points in the 45-degree direction of the lower left 2×2 to form the gray value of the lower left 2×2 as a lower left pixel point. Subtracting the gray values of two pixel points in the 135-degree direction of the lower right 2×2 to form the gray value of the lower right 2×2 as a lower right pixel point. Subtracting the gray values of two pixel points in the upper middle 1×2 to form the gray value of the upper middle 1×2 as an upper middle pixel point. Subtracting the gray values of two pixel points in the lower middle 1×2 to form the gray value of the lower middle 1×2 as a lower middle pixel point. Subtracting the gray values of two pixel points in the left middle 2×1 to form the gray value of the left middle 2×1 as a left middle pixel point. Subtracting the gray values of two pixel points in the right middle 2×1 to form the gray value of the right middle 2×1 as a right middle pixel point. Subtracting the gray value of the middle point pixel from itself to form the gray value of the middle point as a middle point pixel, thereby obtaining a 3×3 feature map; The gradient difference convolution processing in the second multi-angle direction includes: dividing the corresponding pixel points within the range of a 5×5 image filtering kernel into 9 parts, namely, the upper-left 2×2, the upper-right 2×2, the lower-left 2×2, the lower-right 2×2, the upper-middle 1×2, the lower-middle 1×2, the left-middle 2×1, the right-middle 2×1, and the middle point. Subtracting the gray values of two pixel points in the 45-degree direction of the upper-left 2×2 to form the gray value of the upper-left 2×2 as an upper-left pixel point; subtracting the gray values of two pixel points in the 135-degree direction of the upper-right 2×2 to form the gray value of the upper-right 2×2 as an upper-right pixel point; subtracting the gray values of two pixel points in the 135-degree direction of the lower-left 2×2 to form the gray value of the lower-left 2×2 as a lower-left pixel point; subtracting the gray values of two pixel points in the 45-degree direction of the lower-right 2×2 to form the gray value of the lower-right 2×2 as a lower-right pixel point; subtracting the gray values of two pixel points in the upper-middle 1×2 to form the gray value of the upper-middle 1×2 as an upper-middle pixel point; subtracting the gray values of two pixel points in the lower-middle 1×2 to form the gray value of the lower-middle 1×2 as a lower-middle pixel point; subtracting the gray values of two pixel points in the left-middle 2×1 to form the gray value of the left-middle 2×1 as a left-middle pixel point; subtracting the gray values of two pixel points in the right-middle 2×1 to form the gray value of the right-middle 2×1 as a right-middle pixel point; subtracting the gray value of the middle point pixel from itself to form the gray value of the middle point as a middle point pixel, thereby obtaining a 3×3 feature map.
7. The method according to any one of claims 3 to 5, wherein, the multi-scale feature fusion processing includes: obtaining a feature map of a preset channel after convolution operation on the feature map, then respectively performing atrous convolution operations with two or more different dilation rates on the obtained feature map of the preset channel, and then summing the results of the corresponding two or more atrous convolution operations; the dilation rate is set according to the requirements of the edge detection model network structure.
8. The method according to claim 1, wherein, the edge detection network on which the edge detection method is based is obtained by training in combination with a loss function based on a network architecture composed of edge feature extraction processing and edge confidence mapping processing; the loss function can be expressed as: Among them, N represents the number of levels of edge feature extraction processing. represents the self-supervised weight of the first mapped feature map corresponding to the s-th level of edge feature extraction processing, where 1 ≤ s ≤ N; represents the self-supervised weight of the second mapped feature map corresponding to the s-th level of edge feature extraction processing; Y k is the edge annotation information of the k-th level, which is 1 or 0, where 1 ≤ k ≤ N; P 1 i is the confidence of the first mapped feature map corresponding to the predicted i-th layer of edge feature extraction processing, where 1 ≤ i ≤ N; is the confidence of the second mapped feature map corresponding to the predicted i-th layer of edge feature extraction processing.
9. The method according to any one of claims 3 to 5, wherein, the edge detection method further includes, for the obtained detection result, connecting multiple contours in the detection result through one or more morphological operations such as opening operation, closing operation, and thinning morphological operation.
10. An edge detection method, wherein, it includes sub-pixel edge extraction based on the edge detection result obtained in any one of claims 1 to 9, including: Obtain the gray gradient value n along the X-axis of the image at the center point of the pixel x and the gray gradient value n along the Y-axis of the image y ; calculating the offset parameter t: where g x and g y are the first derivatives at the center point of the pixel, and g xx , g xy and g yy are the second derivatives at the center point of the pixel, and Among them represents the Kronecker product, (x I , y I ) is the image coordinate of the center point of the pixel, g(x I , y I ) is the grayscale value of the center point of the pixel, k x , k y , k xx , k xy , and k yy are preset convolution kernels; Then the sub-pixel coordinates of the center point of the pixel are (x I ′, y I ′) = (x I + tn x , y I + tn y ); obtaining the sub-pixel coordinates of the center points of each pixel point to achieve sub-pixel edge extraction of the image.
11. The method according to claim 10, wherein, The preset convolution kernel can be set as follows:
12. An edge detection device based on bidirectional feature supervision, characterized in that it includes an image acquisition module for acquiring an image to be detected; a feature extraction module, including multiple levels, for successively performing edge feature extraction processing at multiple levels on the image to be detected; the edge feature extraction processing at each level includes: for the input image of this level, obtaining a first edge feature map after performing first edge feature extraction processing, and obtaining a second edge feature map after performing second edge feature extraction processing; for the edge feature processing at other levels except the last level, obtaining a third edge feature map after performing third edge feature extraction processing; wherein, in the edge feature extraction processing of adjacent two levels, the third edge feature map obtained by the previous level of edge feature extraction processing is used as the input image of the next level of edge feature extraction processing; a first edge mapping module, including multiple levels, corresponding one-to-one with the multiple-level feature extraction module, for performing first edge confidence mapping processing on the input image; a second edge mapping module, including multiple levels, corresponding one-to-one with the multiple-level feature extraction module, for performing second edge confidence mapping processing on the input image; wherein for the first edge feature map output by the feature extraction module of each level except the first level: fusing the first edge feature map of this level with the first edge feature map of the previous level to obtain the corresponding first fusion feature map of this level, and inputting the first fusion feature map of this level into the corresponding first edge mapping module for first edge confidence mapping processing to obtain the first mapping feature map of this level; for the feature extraction module of the first level, inputting the first edge feature map output by the input image after edge feature extraction processing into the corresponding first edge mapping module for first edge confidence mapping processing to obtain the first mapping feature map of the first level; for the second edge feature map output by the feature extraction module of each level except the last level: fusing the second edge feature map of this level with the second edge feature map of the next level to obtain the corresponding second fusion feature map of this level, and inputting the second fusion feature map of this level into the corresponding second edge mapping module for second edge confidence mapping processing to obtain the second mapping feature map of this level; for the feature extraction processing module of the last level, inputting the second edge feature map output by the input image after edge feature extraction processing into the corresponding second edge mapping module for second edge confidence mapping processing to obtain the second mapping feature map of the last level; a channel splicing module for performing channel splicing and fusion processing on the first mapping feature maps and the second mapping feature maps of each level; a 1×1 convolution module for performing 1×1 convolution operation on the feature map after channel splicing and fusion processing to obtain an edge detection result.
13. A computer storage medium, characterized in that a program is stored on the medium, and the program can be executed by a processor to implement the edge detection method based on bidirectional feature supervision as described in any one of claims 1-9, and / or the edge detection method as described in any one of claims 10-11.
Citation Information
Patent Citations
Object-level edge detection method based on deep residual network
CN110706242A
Edge detection method, device and equipment based on bidirectional cascade network
CN112581486A
Cited By
Edge detection method, system and equipment based on improved loss function and medium
CN120852459A
Edge detection method, system, device and medium based on improved loss function
CN120852459B