Image edge detection method and system, electronic equipment and storage medium
Through multi-stage and multi-level edge detection methods, using feature information of different sizes and abstract levels, the problems of low accuracy of image edge detection and high computing resource consumption in the prior art are solved, and a more accurate, reliable and efficient detection effect is achieved.
Patent Information
- Application Number
- CN202411860667.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has problems such as low accuracy, high computing resources consumption, and inability to detect real-time in image edge detection, especially on mobile terminals and edge sides.
Through multi-stage and multi-level edge detection methods, feature information at different sizes and abstract levels is used to extract, fusion, size adjustment and multi-stage detection to reduce false detection and missed detection and improve the accuracy and reliability of edge detection.
It realizes more accurate and reliable image edge detection, which is suitable for different scenarios, improves the robustness and adaptability of detection, and reduces the consumption of computing resources.
Smart Images

Figure CN120014286A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and in particular to a method, system, electronic device and storage medium for image edge detection. Background Art
[0002] With the development of intelligent technology, image edge detection has become one of the key technologies in image processing and is crucial to improving the accuracy and efficiency of computer vision systems.
[0003] Traditional edge detection algorithms use different edge detection operators such as canny to manually set parameter values for extraction effects. They perform differently on different scene images. Some scenes have too many noisy edges, and some scenes cannot extract the required edges. They are universal, but not accurate enough, and the effects are uncontrollable. Current edge detection algorithms based on deep learning usually rely on more complex models, which can improve detection accuracy, but consume a lot of computing resources and have slow computing speeds. GPUs are needed to increase computing resources and speed, and real-time detection is not possible on mobile terminals and edge devices.
[0004] Therefore, a method is needed to improve the accuracy of image edge detection. Summary of the invention
[0005] The present application provides a method, system, electronic device and storage medium for image edge detection. The multi-stage and multi-level edge detection method fully utilizes feature information of different sizes and different abstract levels to improve the accuracy and completeness of edge detection. By fusing the detection results of multiple stages, false detections and missed detections can be reduced, making the edge detection results more accurate and reliable.
[0006] In a first aspect of the present application, a method for image edge detection is provided, which is applied to an image edge detection platform. The method comprises: Acquire an original image, and perform feature extraction on the original image to obtain a plurality of first feature maps; Perform channel addition fusion on the first feature maps of the same size to obtain a plurality of second feature maps of different sizes; Adjusting the second feature map to the same size as the original image through deconvolution to obtain a third feature map, and performing weighted addition fusion on channels of multiple third feature maps to obtain a fourth feature map; Edge detection is performed on the second feature map, the third feature map, and the fourth feature map respectively to obtain corresponding edge detection results, and all edge detection results are combined to obtain an edge detection image.
[0007] By adopting the above technical solution, the original image is feature extracted to obtain multiple first feature maps, which may capture the information of the image at different scales and different abstract levels. Subsequently, the first feature maps of the same size are added and fused on the channel to form second feature maps of different sizes. This process enhances the diversity and richness of the features and helps to understand the image content more comprehensively. The second feature map is adjusted to the same size as the original image by deconvolution to obtain the third feature map. This step ensures the consistency of the feature map and the original image in the spatial dimension during the subsequent processing, which facilitates pixel-level analysis and comparison. Multiple third feature maps are weighted added and fused on the channel to obtain the fourth feature map. Weighted fusion can perform differentiated processing according to the contribution of different feature maps to the edge detection task, thereby improving the pertinence and effectiveness of the fusion results. At the same time, this process also helps to reduce redundant information and improve the compactness and robustness of feature representation. Edge detection is performed on the second feature map, the third feature map and the fourth feature map respectively to obtain their respective edge detection results. Subsequently, all edge detection results are combined to obtain the final edge detection image. The embodiment of the present invention effectively utilizes the multi-level information in the image through multiple steps such as feature extraction, fusion, resizing, and multi-stage detection, and improves the adaptability and robustness of the edge detection algorithm to changes in image content. Whether it is a simple image or a complex scene, the embodiment of the present invention can maintain a good detection effect.
[0008] Optionally, extracting features from the original image to obtain a plurality of first feature maps includes: The original image is scaled at different ratios to obtain a plurality of intermediate images, preliminary feature extraction is performed on the first intermediate image, and deep feature extraction is performed on the second intermediate image to obtain a plurality of first feature maps, wherein the first intermediate image is an intermediate image obtained by scaling the original image at a first scaling ratio, and the second intermediate image is other intermediate images among the plurality of intermediate images except the first intermediate image, the preliminary features include edges, corners and textures of the image, and the deep features include semantic information of the image.
[0009] By adopting the above technical solution, the original image is scaled at different ratios to obtain multiple intermediate images. This multi-scale processing helps to capture the feature information of the image at different scales. Preliminary feature extraction is performed on the first intermediate image. The preliminary features mainly include low-level features such as edges, corners and textures of the image. These features are crucial for the basic structure and shape recognition of the image. Deep feature extraction is performed on the second intermediate image to obtain multiple first feature maps. Deep features mainly include semantic information of the image, such as high-level features such as object categories and scene layouts. These features help to understand the overall content and contextual relationship of the image. Through multi-scale scaling and feature extraction, various feature information of the image at different scales can be captured. This multi-scale and multi-level feature representation helps to more comprehensively describe the image content and improve the accuracy and robustness of subsequent processing tasks. The preliminary features (edges, corners, textures) provide basic structure and shape information for the image, which is the basis for subsequent processing tasks. The deep features (semantic information) further enhance the feature expression ability of the image, so that the image can be understood and analyzed more accurately. In the field of computer vision and image processing, this multi-scale, multi-level feature extraction method is widely used in tasks such as image classification, target detection, and scene understanding. By extracting rich feature information, the performance of these tasks can be significantly improved. Since different application scenarios have different requirements for image features, the multi-scale, multi-level feature extraction method can flexibly adjust the scaling and feature extraction strategy according to specific needs to adapt to different application scenarios and needs.
[0010] Optionally, the performing channel addition fusion on the first feature maps of the same size to obtain a plurality of second feature maps of different sizes includes: The number of convolution kernels is increased or decreased to increase or reduce the dimension of the channels of the first feature map to the same value, and the elements at corresponding positions of two or more first feature maps with the same size and the same number of channels are added to obtain the second feature map.
[0011] By adopting the above technical solution, by adding the elements of the corresponding positions of two or more first feature maps with the same size and the same number of channels, the complementarity of feature information can be achieved. This fusion method can combine the advantages of different feature maps so that the second feature map contains richer and more comprehensive feature information. In the process of addition and fusion, the same feature information will be strengthened, while inconsistent feature information may offset or weaken each other. This mechanism helps to highlight the important features in the image and improve the saliency of the features. Before fusion, the number of channels of the first feature map is adjusted by increasing or decreasing the number of convolution kernels to reach the same value. This standardization process enables feature maps from different sources to be fused in the same dimension, avoiding the fusion difficulties caused by the mismatch of the number of channels. By adjusting the number of channels, the size and complexity of the fused second feature map can be flexibly controlled. This helps to generate feature maps with different characteristics according to different application scenarios and requirements. In the process of fusing multiple first feature maps into the second feature map, since only a simple addition operation is performed, the amount of calculation is relatively small. This helps to reduce the computational burden of subsequent processing steps and improve the overall processing efficiency. The second feature map obtained by fusion has a more compact and optimized feature representation. This representation helps subsequent processing steps extract key information faster, improving processing speed and accuracy. Although the fusion of the first feature maps of the same size is directly described, this method can be extended to multi-scale feature processing. By extracting and fusing features at different scales, feature maps with different resolutions and abstraction levels can be obtained, which can better meet the needs of multi-scale feature processing.
[0012] Optionally, adjusting the second feature map to the same size as the original image by deconvolution to obtain the third feature map includes: Determining parameters required for a deconvolution operation according to a size difference between the second feature map and the original image, the parameters including the number of channels of the input feature map, the number of channels of the output feature map, the size of the convolution kernel, the step size, and the amount of padding; A deconvolution layer is configured using the parameters, and the second feature map is passed as input to the deconvolution layer to output a third feature map.
[0013] By adopting the above technical solution, the size of the second feature map can be accurately adjusted to the same size as the original image through the deconvolution operation. This size consistency restoration is the basis for subsequent pixel-level analysis, comparison or fusion operations, ensuring the accuracy and reliability of the processing results. The deconvolution operation provides flexible size adjustment capabilities, and the size of the output feature map can be adjusted as needed to adapt to different application scenarios and requirements. In the deconvolution process, by reasonably configuring parameters such as the convolution kernel size, step size and padding number, the loss of feature information can be minimized to ensure that the third feature map can retain the key information in the second feature map. The deconvolution operation itself has a certain feature enhancement effect. Through appropriate parameter configuration, the third feature map can be made more significant and prominent in certain feature dimensions, which helps to better extract and utilize these features in subsequent processing steps. By adjusting the second feature map to the same size as the original image, subsequent processing steps (such as edge detection, target recognition, etc.) can be performed on the same size, thereby reducing the additional calculation amount caused by size mismatch. As one of the preprocessing steps, the deconvolution operation can optimize the overall processing flow and make the subsequent processing steps smoother and more efficient. The deconvolution operation enables the model to process input images of different sizes, enhancing the adaptability and generalization ability of the model. When the third feature map is fused with other feature maps (such as feature fusion that may be performed in subsequent steps), due to the consistency of size, the fusion operation can be performed more conveniently, improving the fusion effect.
[0014] Optionally, performing weighted addition fusion on channels of the plurality of the third feature maps to obtain a fourth feature map includes: Assigning a weight value to each channel of each of the third feature maps; Multiplying the value of each of the third feature maps on the target channel by the weight value of the third feature map on the target channel to obtain a first value, and adding all the first values to obtain a second value of the fourth feature map on the target channel, where the target channel is any channel corresponding to the third feature map; The fourth characteristic map is obtained according to the plurality of the second values.
[0015] By adopting the above technical solution, a weight value is assigned to each channel of each third feature map, and weighted addition fusion is performed accordingly. This weighted fusion method can fully consider the importance difference of different feature maps on different channels, so that the fused fourth feature map is more optimized and efficient in feature expression. Through weighted addition fusion, the feature information between different third feature maps can complement each other to form a more comprehensive and rich feature representation. This feature complementarity helps to improve the accuracy and robustness of subsequent processing tasks. The allocation of weight values can be flexibly adjusted according to the requirements of specific tasks and the characteristics of feature maps. This customizability enables the weighted addition fusion method to adapt to different application scenarios and requirements, and improve the adaptability and generalization ability of the algorithm. Weighted addition fusion at the channel level can retain the feature information on each channel and optimize the feature expression by weight adjustment. This fine-grained fusion method helps to better utilize the information in the feature map. In the process of weighted addition fusion, some unimportant feature information can be suppressed by weight adjustment, thereby reducing the interference of redundant information on subsequent processing steps. This helps to improve processing efficiency and accuracy. In the fusion process, only a simple weighted addition operation is required, and the amount of calculation is relatively small. This helps optimize the use of computing resources and improve the overall processing speed. Weighted addition fusion can integrate the information of multiple third feature maps, thereby suppressing the impact of noise and interference on feature expression to a certain extent. This noise resistance helps improve the robustness and stability of the model. In complex scenarios, different feature maps may contain different useful information. Through weighted addition fusion, this information can be fully utilized to improve the processing capability and effect of the model in complex scenarios.
[0016] Optionally, the method further includes: Obtaining an edge label for each pixel in the original image; Calculate a scaling ratio between the second feature map and the edge label, and perform nearest neighbor interpolation scaling on the edge label according to the calculated scaling ratio to obtain a scaled edge label; Use a preset loss function to calculate the difference between the second feature map and the scaled edge label to obtain a total loss value; update the parameters of the image edge detection platform according to the total loss value and perform the above steps until the total loss value is less than a threshold.
[0017] By adopting the above technical solution, the edge label of each pixel in the original image is obtained, providing accurate supervision information for image edge detection. This helps the model learn the characteristics of the edge in the image more accurately during the training process. The scaling ratio between the second feature map and the edge label is calculated, and the edge label is scaled by nearest neighbor interpolation to ensure the consistency of the two in spatial size. This helps to evaluate the performance of the model more fairly in the subsequent difference calculation. The difference between the second feature map and the scaled edge label is calculated using a preset loss function to obtain the total loss value. This step is a key link in the model training process, and the parameters of the model are optimized by minimizing the total loss value. The parameters of the image edge detection platform are updated according to the total loss value, and the above steps are performed until the total loss value is less than the set threshold. This process gradually improves the performance of the model in the edge detection task through iterative optimization. By gradually reducing the total loss value, the model gradually converges to the optimal solution during the training process, which improves the stability and reliability of the training. Because accurate edge labels are used as supervision information and the model performance is evaluated by an effective loss function, the trained model has high accuracy in the edge detection task.
[0018] Optionally, performing edge detection on the second feature map, the third feature map, and the fourth feature map respectively to obtain corresponding edge detection results, and combining all edge detection results to obtain an edge detection image includes: Performing edge detection on the second feature map to obtain a first result, performing edge detection on the third feature map to obtain a second result, and performing edge detection on the fourth feature map to obtain a third result; Assigning a first weight to the first result, assigning a second weight to the second result, and assigning a third weight to the third result, wherein the third weight is greater than the second weight, and the second weight is greater than the first weight; The first result, the second result, and the third result are respectively multiplied by corresponding weights to obtain target detection results, and an edge detection image is obtained according to the target detection results.
[0019] By adopting the above technical solution, the feature maps at different stages contain image information at different levels. The second feature map may retain more original image details, while the third and fourth feature maps may contain higher-level semantic information and more abstract features. By performing edge detection on these feature maps and fusing the results, the feature advantages of each stage can be fully utilized, feature complementarity can be achieved, and the accuracy and robustness of edge detection can be improved. By assigning different weights to the edge detection results at different stages (the third weight is greater than the second weight, and the second weight is greater than the first weight), the importance difference of different feature maps in edge detection can be reflected. This weight distribution method helps to optimize the fusion results, making the final edge detection image more accurate and reliable. Since the second feature map is closer to the original image, it may retain more image details, but it may also contain more noise. Through weighted fusion, while retaining certain details, the semantic information of subsequent feature maps can be used to suppress noise and improve the accuracy of edge detection. The third and fourth feature maps contain richer semantic information, which helps to identify complex structures and edge patterns in images. Through weighted fusion, these semantic information can be integrated into the final edge detection image to improve the accuracy and robustness of detection.
[0020] In a second aspect of the present application, a system for image edge detection is provided, comprising a feature extraction module, a first fusion module, a second fusion module and an edge detection module, wherein: A feature extraction module is configured to obtain an original image and perform feature extraction on the original image to obtain a plurality of first feature maps; A first fusion module is configured to perform channel addition fusion on first feature maps of the same size to obtain a plurality of second feature maps of different sizes; A second fusion module is configured to adjust the second feature map to the same size as the original image through deconvolution to obtain a third feature map, and perform weighted addition fusion on channels of multiple third feature maps to obtain a fourth feature map; The edge detection module is configured to perform edge detection on the second feature map, the third feature map and the fourth feature map respectively, obtain corresponding edge detection results, and combine all edge detection results to obtain an edge detection image.
[0021] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes any one of the methods described above.
[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions, and when the instructions are executed, any of the methods described above is executed.
[0023] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By extracting features from the original image, multiple first feature maps are obtained, which contain information at different levels of the image. The first feature maps of the same size are added and fused on the channels to obtain multiple second feature maps of different sizes. This process achieves the initial integration of features and provides a rich source of information for subsequent edge detection; the second feature map is adjusted to the same size as the original image through deconvolution to obtain the third feature map, and multiple third feature maps are further weighted added and fused on the channels to obtain the fourth feature map. This process achieves cross-scale feature fusion, so that the final feature map contains both detailed information of the image and higher-level semantic information; 2. Perform edge detection on the second feature map, the third feature map, and the fourth feature map respectively to obtain their respective edge detection results. This multi-stage detection method helps to capture edge information at different levels and improve the comprehensiveness of detection; perform weighted fusion on the edge detection results at different stages to obtain the final edge detection image. Through reasonable weight distribution, important features can be emphasized, noise and unnecessary edges can be suppressed, thereby improving the accuracy and robustness of detection; 3. The parameters and algorithms of the steps of feature extraction, feature fusion and edge detection can be flexibly adjusted according to the requirements of the specific task and the characteristics of the image to adapt to different application scenarios and data sets; with the development of deep learning technology and the emergence of new feature extraction and edge detection methods, the embodiments of the present invention can easily integrate new technologies and algorithms to further improve the performance of edge detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flowchart of the method for image edge detection disclosed in the embodiment of the present application; Figure 2 It is a schematic diagram of the structure of a lightweight backbone network in the image edge detection platform disclosed in the embodiment of the present application; Figure 3 is a schematic diagram of a depth-separable convolution unit of a lightweight backbone network disclosed in an embodiment of the present application; Figure 4 It is a module schematic diagram of the system for image edge detection disclosed in the embodiment of the present application; Figure 5 It is a structural schematic diagram of an electronic device disclosed in an embodiment of the present application.
[0025] Explanation of the reference numerals: 401, feature extraction module; 402, first fusion module; 403, second fusion module; 404, edge detection module; 501, processor; 502, communication bus; 503, user interface; 504, network interface; 505, memory. DETAILED DESCRIPTION
[0026] In order to enable technicians in this field to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0027] In the description of the embodiments of the present application, words such as "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "for example" or "for example" is intended to present related concepts in a specific way.
[0028] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0029] This embodiment discloses a method for image edge detection, which is applied to an image edge detection platform. Figure 1 is a flow chart of the method for image edge detection disclosed in the embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps: S110, acquiring an original image, and performing feature extraction on the original image to obtain a plurality of first feature maps; In the embodiment of the present invention, a feature fusion unit and an edge detection unit are added to the lightweight backbone network, and edge detection is performed on the image based on the fusion of multi-level features of the lightweight backbone network. In the embodiment of the present invention, the lightweight backbone network is MobileNetV2, but this is not a limitation to the lightweight backbone network. In other embodiments, other lightweight backbone networks may also be used, such as MobileNetV1, MobileNetV3, ShuffleNet, etc. Figure 2 is a schematic diagram of the structure of a lightweight backbone network in the image edge detection platform disclosed in the embodiment of the present application, such as Figure 2 As shown in the figure, the lightweight backbone network consists of one ordinary convolution layer and seven bottlenecks (depthwise separable convolution units). The parameters t, c, n, and s represent the expansion ratio of the convolution kernel in the first 1x1 convolution layer, the number of feature map output channels, the number of bottleneck module repetitions, and the channel-by-channel convolution step size in the first bottleneck module. Figure 3 is a schematic diagram of a depth-separable convolution unit of a lightweight backbone network disclosed in an embodiment of the present application, such as Figure 3 As shown in the figure, the depth-wise separable convolution unit combines 3x3 channel-by-channel convolution and 1x1 normal convolution, which greatly reduces the number of parameters and computation while improving the number of channels.
[0030] Feature extraction is a key step in image processing, which aims to extract useful information or patterns from the original image for subsequent processing or analysis. These features can be the color, texture, shape, edge, etc. of the image, which can describe the important attributes or content of the image. In the process of feature extraction, various image processing techniques and algorithms are usually used, such as filters, convolution, pooling, activation functions, etc. These techniques and algorithms can analyze the local or global characteristics of the image and convert them into a higher-level representation, namely the feature map. The original image is usually multi-channel, for example, RGB images contain three color channels: red, green, and blue. During feature extraction, these channels are processed separately or in combination. In deep learning, convolutional layers are commonly used feature extraction tools. Multiple different feature maps can be generated by convolving the original image with multiple convolution kernels (also called filters). Each convolution kernel learns and captures a specific pattern or feature in the image.
[0031] Optionally, extracting features from the original image to obtain a plurality of first feature maps includes: The original image is scaled at different ratios to obtain a plurality of intermediate images, preliminary feature extraction is performed on the first intermediate image, and deep feature extraction is performed on the second intermediate image to obtain a plurality of first feature maps, wherein the first intermediate image is an intermediate image obtained by scaling the original image at a first scaling ratio, and the second intermediate image is other intermediate images among the plurality of intermediate images except the first intermediate image, the preliminary features include edges, corners and textures of the image, and the deep features include semantic information of the image.
[0032] The original image is scaled according to different ratios to generate multiple intermediate images. Figure 2As shown, these scaling ratios are 1 / 2, 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 16, 1 / 32, and 1 / 32 of the original image, respectively. In this way, eight intermediate images of different scales are obtained. The first two feature layers of the lightweight backbone network (i.e., ordinary convolutional layer and bottleneck0) are used to extract preliminary image features, and the following six bottlenecks (i.e., bottleneck1, bottleneck2, bottleneck3, bottleneck4, bottleneck5, and bottleneck6) further extract deep features for subsequent feature fusion and edge detection. Preliminary feature extraction is performed on the first intermediate image (i.e., the image with a scaling ratio of 1 / 2 of the original image). The purpose of this step is to capture preliminary features of the image, such as edges, corners, and textures. These features are crucial for subsequent image analysis and processing because they provide basic structure and shape information of the image. Perform deep feature extraction on the remaining intermediate images (i.e., images with a scale ratio of 1 / 4, 1 / 8, 1 / 16, 1 / 16, 1 / 32, and 1 / 32 of the original image). This step is usually achieved through a series of complex network structures (such as multiple convolutional layers, pooling layers, activation functions, etc. in convolutional neural networks), aiming to extract deep features of the image, such as semantic information. These deep features are crucial for understanding image content and performing advanced image analysis (such as image classification, object detection, etc.). Through the above-mentioned preliminary feature extraction and deep feature extraction processes, multiple first feature maps are obtained. The first feature map will be used as input for subsequent processing steps (such as feature fusion, edge detection, image classification, etc.).
[0033] By scaling the original image at different ratios, multiple intermediate images are obtained. This method enables image feature extraction to be performed at multiple scales, increasing the flexibility of image processing. Images of different scales can capture feature information at different levels, which helps to improve the accuracy and robustness of subsequent processing. Preliminary feature extraction is performed on the first intermediate image, focusing on low-level features such as edges, corners, and textures of the image. These features are of great significance for the basic shape and structure analysis of the image, and provide important clues for subsequent processing. Deep feature extraction is performed on the second intermediate image, focusing on the semantic information of the image. Semantic information refers to high-level features such as objects, scenes, and the relationship between them contained in the image, which is crucial for the understanding and recognition of the image. By fusing preliminary features and deep features, a richer and more comprehensive feature expression can be obtained. This multi-level feature fusion method helps to capture complex information in the image and improve the accuracy and robustness of feature expression. Image feature extraction at different scales helps to enhance the recognition of features. For example, edge features may be clearer in small-scale images, while semantic features may be more obvious in large-scale images. By combining feature information at multiple scales, feature expression can be made more comprehensive and accurate.
[0034] S120, performing channel addition fusion on the first feature maps of the same size to obtain a plurality of second feature maps of different sizes; The first feature map usually has multiple channels, each of which contains a specific feature information. Channel addition fusion usually refers to adding multiple feature maps (which may have the same size and number of channels) in the channel dimension or performing some form of fusion (such as weighted sum, concatenation, etc.) to generate a new feature map. This operation does not change the size of the feature map, but increases or modifies its number of channels or the information within the channel.
[0035] Optionally, the performing channel addition fusion on the first feature maps of the same size to obtain a plurality of second feature maps of different sizes includes: The number of convolution kernels is increased or decreased to increase or reduce the dimension of the channels of the first feature map to the same value, and the elements at corresponding positions of two or more first feature maps with the same size and the same number of channels are added to obtain the second feature map.
[0036] For feature maps generated by different bottlenecks that need to be fused (for example, the first feature map generated by bottleneck3 and the first feature map generated by bottleneck4), if their number of channels (that is, the number of convolution kernels) is different, it is necessary to increase or decrease the number of convolution kernels through additional convolution layers (possibly 1x1 convolutions) so that all feature maps involved in the fusion have the same number of channels. This step is to ensure that when the channels are added, the elements of each corresponding channel are comparable. After ensuring that all feature maps of the same size have the same number of channels, the next step is to add these feature maps in the channel dimension, for example, scaling them to the original Figure 1 / 16 The first feature maps of bottleneck3 and bottleneck4 are added and scaled to the original Figure 1 / 32, add the first feature maps of bottleneck5 and bottleneck6, and add channel 1 of the two first feature maps to obtain channel 1 of the second feature map. Specifically, two or more feature maps with the same size and the same number of channels are added together at the same position on each corresponding channel. For example, if two feature maps have 256 values on each channel (assuming they are both 256x256), then add all 256 values of the two feature maps on the first channel to obtain the value of the fused feature map on the first channel, and so on, until all channels are added together. Figure 2 As shown, bottleneck3 and bottleneck4 produce a scaled Figure 1 / 16, bottleneck5 and bottleneck6 produce the second feature map scaled to the original Figure 1 / 32 second feature map, that is, the final zoomed to the original Figure 1 The second feature maps of four different sizes: / 4, 1 / 8, 1 / 16, and 1 / 32.
[0037] By adding and fusing two or more first feature maps of the same size and the same number of channels on the channel, feature information from different sources or different network layers can be integrated. This integration helps to enhance the expressive power of feature maps, so that the fused feature maps can contain richer information, which is beneficial to subsequent classification, detection or recognition tasks. In some cases, different feature maps may contain some redundant information. By adding and fusing, this redundancy can be reduced to a certain extent, making the fused feature maps more compact and efficient. This helps to reduce the computational complexity of the model and improve the inference speed. Feature fusion can be regarded as a way of data enhancement, introducing more feature combinations during model training. This diversity helps the model learn more robust feature representations, thereby improving the generalization ability of the model and enabling it to perform well when facing unseen data. Although the above content directly describes the channel addition fusion of feature maps of the same size, this technique is usually combined with multi-scale feature processing. In practical applications, the network may extract multiple feature maps of different sizes and adjust them to the same size through operations such as upsampling, downsampling or convolution for channel addition fusion. This multi-scale feature fusion helps the model capture feature information at different levels in the image and improves the model's ability to understand complex scenes.
[0038] S130, adjusting the second feature map to the same size as the original image through deconvolution to obtain a third feature map, and performing weighted addition fusion on channels of multiple third feature maps to obtain a fourth feature map; Deconvolution is a special convolution operation that is usually used to enlarge the size of the feature map to the size of the original input image or larger. In convolutional neural networks, convolutional layers and pooling layers usually cause the size of the feature map to gradually decrease, while deconvolution layers can achieve an increase in size. Deconvolution is not the inverse operation of convolution (because convolution operations are usually irreversible), but it can approximately restore the size of the feature map through learned parameters. In this step, the second feature map obtained previously is used as the input of the deconvolution layer. Through the deconvolution operation, the size of the second feature map is enlarged to the same size as the original image, thereby obtaining a third feature map. These third feature maps now contain feature information with the same spatial resolution as the original image, but their number of channels may still be the same as the original second feature map. After obtaining multiple third feature maps of the same size as the original image, the third feature maps are weighted and added together on the channels. This means that for each spatial position (i.e., each pixel position of the image), the values of the corresponding channels from different third feature maps will be weightedly added to generate a fourth feature map. This weight can be fixed (e.g., equally weighted) or learnable (i.e., optimized as part of the network parameters).
[0039] Optionally, adjusting the second feature map to the same size as the original image by deconvolution to obtain the third feature map includes: Determining parameters required for a deconvolution operation according to a size difference between the second feature map and the original image, the parameters including the number of channels of the input feature map, the number of channels of the output feature map, the size of the convolution kernel, the step size, and the amount of padding; A deconvolution layer is configured using the parameters, and the second feature map is passed as input to the deconvolution layer to output a third feature map.
[0040] Number of channels of the input feature map: This refers to the number of channels of the second feature map, which directly determines the dimension of the input data of the deconvolution layer. Number of channels of the output feature map: This is usually determined according to the specific task or network design. In some cases, you may want to keep the same number of channels as the input feature map, or adjust it according to the needs of subsequent layers. However, in the feature map that is ultimately the same size as the original image, if you plan to fuse multiple feature maps by weighted addition on the channels, then the number of output channels may need to match the fusion strategy. Size of the convolution kernel: The size of the convolution kernel determines the scope of local connections during deconvolution. Larger convolution kernels can capture a wider range of spatial information, but they also increase the amount of computation and the number of parameters. The size of the convolution kernel needs to be selected according to the size of the feature map and the target size. Step size: The step size determines the multiple by which the feature map size is enlarged during deconvolution. In order to adjust the feature map size to the same as the original image, the step size needs to match the size difference between the original image and the second feature map. Typically, the step size is set to a value greater than 1 to achieve an increase in size. Amount of padding: Padding is used to add extra zero values to the edges of the input feature map before the deconvolution operation to control the size of the output feature map. By adjusting the amount of padding, the size of the output feature map can be further fine-tuned to ensure that it is exactly the same size as the original image.
[0041] Once the above parameters are determined, they can be used to configure the deconvolution layer. In deep learning frameworks (such as TensorFlow, PyTorch, etc.), there are usually dedicated functions or layers to implement deconvolution operations. After the deconvolution layer is configured, the second feature map is passed as input to the deconvolution layer. The deconvolution layer will deconvolve the input feature map according to the configured parameters and output a third feature map of the same size as the original image. This process is a key step in multi-scale and multi-level feature fusion. It allows the feature information extracted by the network at different scales to be unified to the same size in the final layer for further fusion and training. By weighted addition and fusion of feature maps of the same size obtained after multiple deconvolutions on the channel, feature information from different scales and levels can be further integrated to improve the performance of the model.
[0042] like Figure 2As shown, the above obtained scaling is the original Figure 1 The second feature of four different sizes: 1 / 4, 1 / 8, 1 / 16, 1 / 32 Figure 1 On the one hand, it is used for direct training. On the other hand, it is deconvolved to a feature layer with the same size as the original image, and the third feature maps obtained after the four deconvolutions are weightedly added and fused on the channel. The fourth feature map obtained after the fusion and the above four third feature maps are also added to the training.
[0043] Through the deconvolution operation, the second feature map with a smaller size can be enlarged to the same size as the original image, so that subsequent processing (such as feature fusion, classification, etc.) can be performed at the same spatial resolution, thus avoiding the performance degradation that may be caused by size mismatch. The deconvolution operation not only enlarges the size of the feature map, but also retains the key information in the original feature map to a certain extent through the weight learning of the convolution kernel. Through the nonlinear transformation of the deconvolution layer, the feature map may also be enhanced while its size is enlarged, which helps the model better capture the details and contextual information in the image. According to the size difference between the second feature map and the original image, the parameters of the deconvolution operation (such as convolution kernel size, step size, padding, etc.) can be flexibly adjusted to adapt to different input sizes and output requirements. In the multi-scale feature fusion framework, the deconvolution operation enables feature maps of different scales to be fused at the same size, thereby making full use of the feature information at different scales and improving the robustness and generalization ability of the model.
[0044] Optionally, performing weighted addition fusion on channels of the plurality of the third feature maps to obtain a fourth feature map includes: Assigning a weight value to each channel of each of the third feature maps; Multiplying the value of each of the third feature maps on the target channel by the weight value of the third feature map on the target channel to obtain a first value, and adding all the first values to obtain a second value of the fourth feature map on the target channel, where the target channel is any channel corresponding to the third feature map; The fourth characteristic map is obtained according to the plurality of the second values.
[0045] Assign a weight value to each channel of each third feature map. These weight values are used to control the contribution of different feature maps in the fusion process. The weight value can be a preset fixed value or a learnable parameter that is automatically optimized through the network training process. In the case of learnability, these weight values will be trained as part of the network to find the optimal fusion strategy. Select a target channel, which can be any channel corresponding to the third feature map. Since all third feature maps have been adjusted to the same size, they have the same number of channels and can be fused channel by channel. For each third feature map, multiply its value on the target channel by the weight value of the feature map on the target channel to obtain a first value. This operation is performed independently at each spatial position of the feature map (that is, each pixel position of the image). Add the first values calculated on the target channel of all third feature maps to obtain the second value of the fourth feature map on the target channel. This operation is performed within the range of the entire feature map, that is, the values at all spatial positions are accumulated. For each channel of the third feature map, the above operation is repeated, that is, assigning weights, calculating first values, and adding to obtain the second value. In this way, a corresponding second value can be generated for each channel of the fourth feature map. The second values on all channels are combined to obtain the final fourth feature map. This feature map contains the comprehensive information from multiple third feature maps and has the same size as the original image.
[0046] Through weighted addition fusion, feature information from different network layers or different processing paths can be effectively integrated to form a more comprehensive and rich feature representation. The learnability of weight values makes the fusion process highly flexible, and the fusion strategy can be automatically adjusted according to the characteristics of specific tasks and data sets. The fusion of multiple feature maps helps the model capture more contextual information and detailed features, thereby improving the performance of the model, such as classification accuracy, detection accuracy, etc. Although the fusion process increases the amount of calculation, since all feature maps have been adjusted to the same size, they can be efficiently processed in parallel, thereby maintaining a high computational efficiency.
[0047] Optionally, the method further includes: Obtaining an edge label for each pixel in the original image; Calculate a scaling ratio between the second feature map and the edge label, and perform nearest neighbor interpolation scaling on the edge label according to the calculated scaling ratio to obtain a scaled edge label; Use a preset loss function to calculate the difference between the second feature map and the scaled edge label to obtain a total loss value; update the parameters of the image edge detection platform according to the total loss value and perform the above steps until the total loss value is less than a threshold.
[0048] In order to train an image edge detection model, we first need to know the pixels in the image that belong to the edge. This is usually done by some form of annotation process (such as manual annotation or automatic annotation algorithm) to generate edge labels. The edge labels corresponding to the original image are obtained from the annotated data. These labels are usually binary images of the same size as the original image, where edge pixels are labeled as 1 (or other non-zero values) and non-edge pixels are labeled as 0. Since different feature layers in the model may have different spatial resolutions (i.e. sizes), the edge labels need to be scaled to the same size as each feature layer in order to calculate the loss on these layers. The scaling factor is calculated based on the ratio of the size of the second feature map (or any other intermediate feature layer) to the size of the original image. Using the calculated scaling factor, the edge labels are scaled by nearest neighbor interpolation to generate scaled edge labels of the same size as the second feature map (or other feature layer). Nearest neighbor interpolation is a simple and computationally efficient image scaling method suitable for pixel-level tasks such as edge detection.
[0049] The performance of the model is evaluated by comparing the edges predicted by the model (i.e., the second feature map) with the true edges (i.e., the scaled edge labels), and the loss value is calculated accordingly. Weighted binary cross entropy is selected as the loss function because it can handle binary classification problems (edge / non-edge) and allows the influence of positive and negative samples to be balanced by weights. The difference between the second feature map and the scaled edge labels is calculated using the selected loss function. The difference here represents the inconsistency between the model prediction and the true edge. If the model contains multiple intermediate feature layers and each layer participates in the loss calculation (such as through feature fusion strategies of different scales), the loss values of all layers need to be weighted and summed according to the preset weights to obtain the total loss value. The model's ability to detect image edges is improved by adjusting the model's parameters to minimize the total loss value. The backpropagation algorithm and optimizer (such as SGD, Adam, etc.) are used to update the parameters of the image edge detection platform (i.e., the weights and biases of the model) according to the calculated total loss value. The above steps (including obtaining edge labels, calculating scaling, loss calculation, and parameter updating) are repeated until the total loss value is less than a preset threshold or other stopping conditions (such as an upper limit on the number of iterations) are reached.
[0050] The public datasets BSDS500 and PASCAL VOC are used, and the training set, validation set, and test set are divided into 8:1:1. The pixel values of the input image are normalized according to the original image size without scaling. The label is a grayscale image. The pixel value equal to 0 is set as a negative sample, the pixel value greater than or equal to 127.5 is set as a positive sample, and other pixel values are set as ignored samples. Due to the different ratios of positive and negative samples, the weights of positive and negative samples are set to balance the proportion of positive and negative samples in the total loss value. The weight of the positive sample is the ratio of the number of negative sample pixels to the sum of the number of positive and negative sample pixels, the weight of the negative sample is the ratio of the number of positive sample pixels to the sum of the number of positive and negative sample pixels, and the weight of the ignored sample is 0. The weighted binary cross entropy is used to calculate the model loss. For the feature layers of different scales in the middle of the model, not only the deconvolution feature layer is used as the output feature layer to calculate the loss with the original input image, but also the label is scaled to the scale of the feature map by the nearest neighbor to calculate the loss with the original feature layer. Since some output feature layers correspond to the original input image without scaling, the loss weight of these feature layers is greater than that of other layers. The loss function formula is as follows: Among them, L is the loss function, y is the actual value, Prediction value, N is the number of input images, M is i is the number of output feature layers of the i-th input image, α j is the j-th output feature layer loss weight, F i,j is the total number of pixels in the jth output feature layer of the ith input image, i.e., the product of the length and width of the feature map, w i,j,k is the kth sample weight of the jth output feature layer of the ith input image, y i,j,k is the true label of the kth sample of the jth output feature layer of the ith input image (positive samples are 1 and negative samples are 0), The model prediction value for the kth sample of the jth output feature layer of the ith input image.
[0051] The stochastic gradient descent algorithm SGD is used to back-propagate the loss value to update the model parameters.
[0052] The F1-score evaluation indicator is used, and the formula is as follows: Among them, F1 is the comprehensive value, P is the precision, and R is the recall rate.
[0053] S140, performing edge detection on the second feature map, the third feature map, and the fourth feature map respectively to obtain corresponding edge detection results, and combining all edge detection results to obtain an edge detection image.
[0054] Apply edge detection algorithms to the second feature map, the third feature map, and the fourth feature map, respectively, to obtain multiple independent edge detection results. Combine multiple edge detection results to form a unified edge detection image. The fused edge detection image can be post-processed, such as morphological operations (dilation, erosion, opening operation, closing operation, etc.) to further improve the continuity and accuracy of the edge. That is, in the end, one feature map scaled to 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image and five feature maps with the same size as the original image are obtained, for a total of nine feature maps. All nine feature maps can be used for training to improve the accuracy of edge detection, and the fused feature map with the same size as the original image can be used for inference prediction.
[0055] Optionally, performing edge detection on the second feature map, the third feature map, and the fourth feature map respectively to obtain corresponding edge detection results, and combining all edge detection results to obtain an edge detection image includes: Performing edge detection on the second feature map to obtain a first result, performing edge detection on the third feature map to obtain a second result, and performing edge detection on the fourth feature map to obtain a third result; Assigning a first weight to the first result, assigning a second weight to the second result, and assigning a third weight to the third result, wherein the third weight is greater than the second weight, and the second weight is greater than the first weight; The first result, the second result, and the third result are respectively multiplied by corresponding weights to obtain target detection results, and an edge detection image is obtained according to the target detection results.
[0056] In order to effectively combine these edge detection results from different scales, different weights need to be assigned to them. The weight allocation is based on the importance of the feature map, the confidence of the edge detection results, and their contribution to the final edge detection image. In an embodiment of the present invention, the third weight is greater than the second weight, and the second weight is greater than the first weight. This weight allocation strategy reflects that in the fusion process, higher-level feature maps (such as the fourth feature map) usually contain richer information and higher confidence, and therefore should be given a larger weight. The edge detection results (first result, second result, and third result) of each feature map are multiplied by the corresponding weights to obtain a weighted edge detection result. This process adjusts the edge information of different scales according to its importance. The weighted edge detection results are added (or other forms of fusion are performed) to obtain a target detection result. This target detection result combines edge information from different scales to form a more comprehensive and accurate edge representation. Based on the target detection result, the final edge detection image can be directly generated. This image clearly shows the edge information in the input image, and because it combines multi-scale feature information, it is usually more accurate than a single-scale edge detection result.
[0057] By performing edge detection on feature maps of different levels (second feature map, third feature map, fourth feature map) respectively, edge information of different scales and abstraction levels in the image can be captured. Shallow feature maps (such as the second feature map) usually contain more details and texture information, while deep feature maps (such as the fourth feature map) contain higher-level semantic information. Combining this information can locate edges more comprehensively and accurately. Feature maps of different levels have different sensitivities to interference factors such as noise and illumination changes in the image. By weighted merging of edge detection results of different levels, these interferences can be smoothed to a certain extent and the robustness of edge detection can be improved. By assigning different weights to edge detection results of different levels (the third weight is greater than the second weight, and the second weight is greater than the first weight), the contribution of feature maps of each level can be adjusted according to specific application scenarios and requirements. This flexibility enables this method to better adapt to different image analysis tasks.
[0058] The embodiment of the present invention can automatically extract edge information more accurately by training a large amount of data without manually setting parameters. For special scenes, more accurate edge information can be obtained by fine-tuning the training pictures of special scenes. The embodiment of the present invention can greatly improve the edge detection speed on the CPU through a model structure based on a lightweight backbone network such as MobileNetV2, and the speed on the GPU is also better than the existing technology. The embodiment of the present invention improves the expressiveness of features by fusing multiple levels of features. The embodiment of the present invention proposes multi-scale feature training. This multi-scale does not scale the input image to expand the training set data, but scales the label map corresponding to the input image to the feature map size for different scale feature layers, so as to train the expressiveness of features of different levels and scales.
[0059] This embodiment also discloses a system for image edge detection. Figure 4 is a module schematic diagram of the image edge detection system disclosed in the embodiment of the present application, such as Figure 4 As shown, the system includes a feature extraction module 401, a first fusion module 402, a second fusion module 403 and an edge detection module 404, wherein: A feature extraction module 401 is configured to obtain an original image and perform feature extraction on the original image to obtain a plurality of first feature maps; A first fusion module 402 is configured to perform channel addition fusion on first feature maps of the same size to obtain a plurality of second feature maps of different sizes; The second fusion module 403 is configured to adjust the second feature map to the same size as the original image through deconvolution to obtain a third feature map, and perform weighted addition and fusion of multiple third feature maps on channels to obtain a fourth feature map; the edge detection module 404 is configured to perform edge detection on the second feature map, the third feature map and the fourth feature map respectively to obtain corresponding edge detection results, and combine all edge detection results to obtain an edge detection image.
[0060] Optionally, the feature extraction module 401 is configured to: The original image is scaled at different ratios to obtain a plurality of intermediate images, preliminary feature extraction is performed on the first intermediate image, and deep feature extraction is performed on the second intermediate image to obtain a plurality of first feature maps, wherein the first intermediate image is an intermediate image obtained by scaling the original image at a first scaling ratio, and the second intermediate image is other intermediate images among the plurality of intermediate images except the first intermediate image, the preliminary features include edges, corners and textures of the image, and the deep features include semantic information of the image.
[0061] Optionally, the first fusion module 402 is configured to: The number of convolution kernels is increased or decreased to increase or reduce the dimension of the channels of the first feature map to the same value, and the elements at corresponding positions of two or more first feature maps with the same size and the same number of channels are added to obtain the second feature map.
[0062] Optionally, the second fusion module 403 is configured to: Determining parameters required for a deconvolution operation according to a size difference between the second feature map and the original image, the parameters including the number of channels of the input feature map, the number of channels of the output feature map, the size of the convolution kernel, the step size, and the amount of padding; A deconvolution layer is configured using the parameters, and the second feature map is passed as input to the deconvolution layer to output a third feature map.
[0063] Optionally, the second fusion module 403 is configured to: Assigning a weight value to each channel of each of the third feature maps; Multiplying the value of each of the third feature maps on the target channel by the weight value of the third feature map on the target channel to obtain a first value, and adding all the first values to obtain a second value of the fourth feature map on the target channel, where the target channel is any channel corresponding to the third feature map; The fourth characteristic map is obtained according to the plurality of the second values.
[0064] Optionally, the system further includes an adjustment module, wherein the adjustment module is configured to: Obtaining an edge label for each pixel in the original image; Calculate a scaling ratio between the second feature map and the edge label, and perform nearest neighbor interpolation scaling on the edge label according to the calculated scaling ratio to obtain a scaled edge label; Use a preset loss function to calculate the difference between the second feature map and the scaled edge label to obtain a total loss value; update the parameters of the image edge detection platform according to the total loss value and perform the above steps until the total loss value is less than a threshold.
[0065] Optionally, the edge detection module 404 is configured to: Performing edge detection on the second feature map to obtain a first result, performing edge detection on the third feature map to obtain a second result, and performing edge detection on the fourth feature map to obtain a third result; Assigning a first weight to the first result, assigning a second weight to the second result, and assigning a third weight to the third result, wherein the third weight is greater than the second weight, and the second weight is greater than the first weight; The first result, the second result, and the third result are respectively multiplied by corresponding weights to obtain target detection results, and an edge detection image is obtained according to the target detection results.
[0066] It should be noted that: when the device provided in the above embodiment realizes its function, only the division of the above functional modules is used as an example. In actual application, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0067] This embodiment also discloses an electronic device, referring to Figure 5 The electronic device may include: at least one processor 501 , at least one communication bus 502 , a user interface 503 , a network interface 504 , and at least one memory 505 .
[0068] The communication bus 502 is used to realize the connection and communication between these components.
[0069] The user interface 503 may include a display screen (Display) and a camera (Camera), and the optional user interface 503 may also include a standard wired interface and a wireless interface.
[0070] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0071] Among them, the processor 501 may include one or more processing cores. The processor 501 uses various interfaces and lines to connect various parts in the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 505, and calling data stored in the memory 505. Optionally, the processor 501 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 501 can integrate one or a combination of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 501, and it can be implemented separately through a chip.
[0072] Among them, the memory 505 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read-Only Memory). Optionally, the memory 505 includes a non-transitory computer-readable storage medium. The memory 505 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 505 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 505 may also be optionally at least one storage device located away from the aforementioned processor 501. As Figure 5 As shown, the memory 505 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program of the image edge detection method.
[0073] exist Figure 5In the electronic device shown, the user interface 503 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 501 can be used to call the application program of the image edge detection method stored in the memory 505. When executed by one or more processors 501, the electronic device executes one or more methods in the above-mentioned embodiments.
[0074] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the present application.
[0075] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0076] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0077] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0078] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0079] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory 505 and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned memory 505 includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk.
[0080] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field that are not recorded in the present disclosure. The description and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A method for image edge detection, characterized in that: Applied to an image edge detection platform, the method comprises: Acquire an original image, and perform feature extraction on the original image to obtain a plurality of first feature maps; Perform channel addition fusion on the first feature maps of the same size to obtain a plurality of second feature maps of different sizes; Adjusting the second feature map to the same size as the original image through deconvolution to obtain a third feature map, and performing weighted addition fusion on channels of multiple third feature maps to obtain a fourth feature map; Edge detection is performed on the second feature map, the third feature map, and the fourth feature map respectively to obtain corresponding edge detection results, and all edge detection results are combined to obtain an edge detection image.
2. The method for image edge detection according to claim 1, characterized in that: The extracting features from the original image to obtain a plurality of first feature maps comprises: The original image is scaled at different ratios to obtain a plurality of intermediate images, preliminary feature extraction is performed on the first intermediate image, and deep feature extraction is performed on the second intermediate image to obtain a plurality of first feature maps, wherein the first intermediate image is an intermediate image obtained by scaling the original image at a first scaling ratio, and the second intermediate image is other intermediate images among the plurality of intermediate images except the first intermediate image, the preliminary features include edges, corners and textures of the image, and the deep features include semantic information of the image.
3. The method for image edge detection according to claim 1, characterized in that: The adding and fusing the first feature maps of the same size on channels to obtain a plurality of second feature maps of different sizes comprises: The number of convolution kernels is increased or decreased to increase or reduce the dimension of the channels of the first feature map to the same value, and the elements at corresponding positions of two or more first feature maps with the same size and the same number of channels are added to obtain the second feature map.
4. The method for image edge detection according to claim 1, characterized in that: The step of adjusting the second feature map to the same size as the original image by deconvolution to obtain a third feature map comprises: Determining parameters required for a deconvolution operation according to a size difference between the second feature map and the original image, the parameters including the number of channels of the input feature map, the number of channels of the output feature map, the size of the convolution kernel, the step size, and the amount of padding; A deconvolution layer is configured using the parameters, and the second feature map is passed as input to the deconvolution layer to output a third feature map.
5. The method for image edge detection according to claim 1, characterized in that: The step of performing weighted addition fusion on channels of the plurality of the third feature maps to obtain a fourth feature map comprises: Assigning a weight value to each channel of each of the third feature maps; Multiplying the value of each of the third feature maps on the target channel by the weight value of the third feature map on the target channel to obtain a first value, and adding all the first values to obtain a second value of the fourth feature map on the target channel, where the target channel is any channel corresponding to the third feature map; The fourth characteristic map is obtained according to the plurality of the second values.
6. The method for image edge detection according to claim 1, characterized in that: The method further comprises: Obtaining an edge label for each pixel in the original image; Calculate a scaling ratio between the second feature map and the edge label, and perform nearest neighbor interpolation scaling on the edge label according to the calculated scaling ratio to obtain a scaled edge label; Calculate the difference between the second feature map and the scaled edge label using a preset loss function to obtain a total loss value; The parameters of the image edge detection platform are updated by the total loss value and the above steps are performed until the total loss value is less than a threshold.
7. The method for image edge detection according to claim 1, characterized in that: The performing edge detection on the second feature map, the third feature map, and the fourth feature map respectively to obtain corresponding edge detection results, and combining all edge detection results to obtain an edge detection image comprises: Performing edge detection on the second feature map to obtain a first result, performing edge detection on the third feature map to obtain a second result, and performing edge detection on the fourth feature map to obtain a third result; Assigning a first weight to the first result, assigning a second weight to the second result, and assigning a third weight to the third result, wherein the third weight is greater than the second weight, and the second weight is greater than the first weight; The first result, the second result, and the third result are respectively multiplied by corresponding weights to obtain target detection results, and an edge detection image is obtained according to the target detection results.
8. A system for image edge detection, characterized in that: It includes a feature extraction module, a first fusion module, a second fusion module and an edge detection module, wherein: A feature extraction module is configured to obtain an original image and perform feature extraction on the original image to obtain a plurality of first feature maps; A first fusion module is configured to perform channel addition fusion on first feature maps of the same size to obtain a plurality of second feature maps of different sizes; A second fusion module is configured to adjust the second feature map to the same size as the original image through deconvolution to obtain a third feature map, and perform weighted addition fusion on channels of multiple third feature maps to obtain a fourth feature map; The edge detection module is configured to perform edge detection on the second feature map, the third feature map and the fourth feature map respectively, obtain corresponding edge detection results, and combine all edge detection results to obtain an edge detection image.
9. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is executed.