Image edge detection method and device based on deep learning model, equipment and storage medium

By adopting a bidirectional cascaded structural network based on deep learning models in image edge detection, the problem of low detection accuracy in the prior art is solved, and high-accuracy edge detection of images of different quality is achieved.

CN119992120APending Publication Date: 2025-05-13GUANGZHOU SHIKUN ELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311494107.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has low detection accuracy in image edge detection and is susceptible to factors such as light and imaging quality, resulting in inaccurate detection.

Method used

Using an image edge detection method based on a deep learning model, a bidirectional cascaded structural network is formed through preprocessing modules and multi-layer feature extraction blocks, the input images are acquired and shallow and deep edge prediction images are output, and these feature images are fused to generate the final edge feature image.

Benefits of technology

It improves the accuracy and robustness of image edge detection, can show strong adaptability in images in different qualities and environments, and improves the accuracy of edge detection by fusing different levels of features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992120A_ABST
    Figure CN119992120A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of artificial intelligence, in particular to an image edge detection method and device based on a deep learning model, equipment and a storage medium, and the method comprises the steps: obtaining an input image, the deep learning model comprises a preprocessing module and multiple layers of feature extraction blocks, and the multiple layers of feature extraction blocks form a bidirectional cascade structure network; inputting the input image into a preprocessing module and each layer of feature extraction block in sequence, so that each layer of feature extraction block outputs a shallow layer edge prediction image and a deep layer edge prediction image; generating a shallow edge feature image according to the shallow edge prediction image output by the multi-layer feature extraction block; generating a deep edge feature image according to the deep edge prediction image output by the multi-layer feature extraction block; and fusing the shallow edge feature image and the deep edge feature image to obtain a final edge feature image. According to the embodiment of the invention, the edge detection and feature fusion capabilities are enhanced, and excellent performance and robustness can be provided when an image edge detection task is processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to an image edge detection method, device, equipment and storage medium based on a deep learning model. Background Art

[0002] Image edge detection methods are the pre-operation of many computer image tasks, including but not limited to image processing, target detection, and image segmentation. The existing image edge detection methods are traditional image processing edge detection, which finds the edge area of ​​objects in the image through image grayscale gradient and other shallow image information. The disadvantage of this method is that it is easily disturbed. Uncontrollable factors such as light and imaging quality will greatly affect the image edge detection task, resulting in inaccurate detection when performing image edge detection. Summary of the invention

[0003] One purpose of the embodiments of the present application is to provide an image edge detection method, apparatus, device and storage medium based on a deep learning model to solve the technical problem of low detection accuracy in image edge detection.

[0004] In a first aspect, a method for image edge detection based on a deep learning model is provided, comprising:

[0005] Acquire an input image, wherein the deep learning model includes a preprocessing module and a multi-layer feature extraction block, wherein the multi-layer feature extraction blocks constitute a bidirectional cascade structure network;

[0006] The input image is sequentially input into the preprocessing module and the feature extraction blocks of each layer, so that the feature extraction blocks of each layer output a shallow edge prediction image and a deep edge prediction image;

[0007] Generate a shallow edge feature image according to the shallow edge prediction images output by the multiple layers of feature extraction blocks;

[0008] Generate a deep edge feature image according to the deep edge prediction images output by the multiple layers of feature extraction blocks;

[0009] The shallow edge feature image and the deep edge feature image are fused to obtain a final edge feature image.

[0010] In combination with the first aspect, in a possible implementation method, the input image is sequentially input into the preprocessing module and the feature extraction blocks of each layer so that the feature extraction blocks of each layer output a shallow edge prediction image and a deep edge prediction image, including: inputting the input image sequentially into the preprocessing module to obtain an edge output image; inputting the edge output image sequentially into the feature extraction blocks of each layer so that the feature extraction blocks of each layer output a shallow edge prediction image, a deep edge prediction image and an edge output image of the current layer, and the edge output image of the current layer is the input image of the feature extraction blocks of the next layer.

[0011] In combination with the first aspect, in a possible implementation method, each layer of the feature extraction block includes multiple layers of feature extraction sublayers in parallel, and the edge output image is input into the feature extraction block of each layer in sequence so that each layer of the feature extraction block outputs a shallow edge prediction image and a deep edge prediction image, including: inputting the edge output image into each layer of the feature extraction sublayer in sequence to obtain an edge prediction subgraph corresponding to the feature extraction sublayer; merging the edge prediction subgraphs of the feature extraction sublayers of each layer to obtain a merged prediction subgraph; and generating a shallow edge prediction image and a deep edge prediction image according to the merged prediction subgraph.

[0012] In combination with the first aspect, in a possible implementation method, each layer of the feature extraction sublayer includes a depth-separable convolution block and a scale enhancement block connected in series, and the edge output image is input into each layer of the feature extraction sublayer in sequence to obtain an edge prediction subgraph corresponding to the feature extraction sublayer, including: inputting the edge output image into the depth-separable convolution block and the scale enhancement block in sequence to obtain an edge prediction subgraph corresponding to the feature extraction sublayer, wherein the output of the depth-separable convolution of the previous layer is used as the input of the depth-separable convolution of the next layer.

[0013] In combination with the first aspect, in a possible implementation, the depthwise separable convolution block supports an attention mechanism.

[0014] In combination with the first aspect, in a possible implementation method, the depthwise separable convolution block includes an ascending pointwise convolution sub-block, a depthwise separable convolution sub-block, and a descending pointwise convolution sub-block connected in sequence, and the inputting the edge output image into the depthwise separable convolution block and the scale enhancement block in sequence to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer includes: inputting the edge output image into the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block, the descending pointwise convolution sub-block, and the scale enhancement block in sequence to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer.

[0015] In combination with the first aspect, in a possible implementation manner, the depth-separable convolution sub-block adopts a nonlinear activation function.

[0016] In combination with the first aspect, in a possible implementation, the depthwise separable convolution block also includes an attention mechanism module, which is connected in series between the depthwise separable convolution sub-block and the descending pointwise convolution sub-block, and the edge output image is sequentially input into the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block, the descending pointwise convolution sub-block and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer, including: inputting the edge output image into the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block, the attention mechanism module, the descending pointwise convolution sub-block and the scale enhancement block in sequence to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer.

[0017] In a second aspect, an image edge detection device based on a deep learning model is provided, comprising:

[0018] An acquisition unit acquires an input image, wherein the deep learning model includes a preprocessing module and a multi-layer feature extraction block, wherein the multi-layer feature extraction blocks constitute a bidirectional cascade structure network;

[0019] A processing unit, used for sequentially inputting the input image into the preprocessing module and the feature extraction blocks of each layer, so that the feature extraction blocks of each layer output a shallow edge prediction image and a deep edge prediction image;

[0020] The processing unit is further used to generate a shallow edge feature image according to the shallow edge prediction images output by the multiple layers of feature extraction blocks;

[0021] The processing unit is further used to generate a deep edge feature image according to the deep edge prediction images output by the multiple layers of feature extraction blocks;

[0022] The fusion unit is used to fuse the shallow edge feature image and the deep edge feature image to obtain a final edge feature image.

[0023] In a third aspect, an embodiment of the present invention provides a computer device, including:

[0024] at least one processor; and,

[0025] a memory communicatively connected to the at least one processor; wherein,

[0026] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to the first aspect.

[0027] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method as described in the first aspect.

[0028] In the scheme implemented by the image edge detection method, device, equipment and storage medium based on the deep learning model, the input image is first obtained, and the deep learning model includes a preprocessing module and a multi-layer feature extraction block, wherein the multi-layer feature extraction blocks constitute a bidirectional cascade structure network, and then the input image is sequentially input into the preprocessing module and the feature extraction blocks of each layer, so that each layer of the feature extraction block outputs a shallow edge prediction image and a deep edge prediction image, and then a shallow edge feature image is generated according to the shallow edge prediction image output by the multi-layer feature extraction block, and a deep edge feature image is generated according to the deep edge prediction image output by the multi-layer feature extraction block, and finally the shallow edge feature image and the deep edge feature image are fused to obtain the final edge feature image. This method optimizes the deep learning model by bidirectional cascade connection, and thus has strong adaptability and robustness for images of different qualities and taken in different environments, and by fusing the shallow edge feature image and the deep edge feature image, the model can integrate features at different levels, thereby improving the accuracy of edge detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0030] Figure 1A is a schematic diagram of a VGG16 network in one embodiment of the present invention;

[0031] Figure 1B is a model network structure diagram in one embodiment of the present invention;

[0032] Figure 2A It is a flowchart of an image edge detection method based on a deep learning model in one embodiment of the present invention;

[0033] Figure 2B is an edge feature image in one embodiment of the present invention;

[0034] Figure 2C is a network diagram of a depthwise separable convolution in one embodiment of the present invention;

[0035] Figure 2D A network diagram of a scale enhancement module in one embodiment of the present invention;

[0036] Figure 2E is a schematic diagram of a deep separable network with an attention mechanism added in one embodiment of the present invention;

[0037] Figure 3 is a structural schematic diagram of an image edge detection device based on a deep learning model in one embodiment of the present invention;

[0038] Figure 4 It is a schematic diagram of the structure of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0040] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other, all within the scope of protection of the present application. In addition, although the functional module division is performed in the device schematic diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a sequence different from the module division in the device or the flow chart. Furthermore, the words "first", "second", "third", etc. used in this application do not limit the data and execution order, but only distinguish the same items or similar items with basically the same functions and effects.

[0041] The present invention is described in detail below through specific embodiments.

[0042] The technical solution of this application can be applied to various edge computing scenarios, etc.

[0043] See also Figure 1A , Figure 1A A VGG16 network diagram is provided for an embodiment of the present application. In public information, VGG (Visual Geometry Group) is a deep convolutional neural network architecture proposed by a visual geometry group in 2014. The VGG network is widely used in computer vision tasks such as image classification, target detection, and semantic segmentation. The simplicity and ease of implementation of its network structure make VGG one of the classic models in the field of deep learning.

[0044] The commonly used structure is VGG16, such as Figure 1A As shown, Figure 1A D in the figure corresponds to the network structure of VGG16. As shown in the figure, VGG16 has 13 convolutional layers, 5 maximum pooling layers, and 3 fully connected layers. The size of the convolutional layer is 3×3, and the step size is 1. The number of channels can be increased through the convolutional layer. The size of the pooling layer is 2×2, and the step size is 2. The function is to reduce the size of the feature map and improve the anti-interference ability of the network. The number of channels in the first two layers of the three fully connected layers is 4096, and the number of channels in the third layer is 1000, representing 1000 label categories, showing strong classification capabilities. All hidden layers are followed by ReLU nonlinear activation functions, which helps to increase the expressive power of the model and alleviate the gradient disappearance problem. For the VGG16 convolutional neural network, its 13 convolutional layers and 5 pooling layers are responsible for feature extraction, showing strong feature extraction capabilities. Finally, the 3 fully connected layers are responsible for completing the classification task.

[0045] Based on the above VGG backbone network, the present invention proposes a deep learning model, which removes the last three fully connected layers and one pooling layer of the VGG16 network compared to VGG16. Based on the fully connected layer in the above structure, it usually corresponds to a fixed output structure (such as a classification label). For pixel-level tasks (image segmentation) or tasks involving multiple different types of outputs, the fully connected layer can be removed to allow the model to have a more flexible output structure and reduce the amount of calculation, so that the model pays more attention to pixel-level features. In addition, removing a pooling layer can reduce information loss and retain more details.

[0046] For an example of the above deep learning model, see Figure 1B ,exist Figure 1BA deep learning model proposed in the present invention includes: a preprocessing module, which can be N convolutional layers; feature extraction blocks of Block1-Block5, that is, all convolutional layers used to extract features have a total of 13 layers. The multiple feature extraction blocks in the deep learning model constitute a bidirectional cascade structure network, and include a depthwise separable convolution module DSCM (Depthwise Separable Convolution module) and a scale enhancement module SEM (Scale Enhancement Module) with an attention mechanism, and the preset convolution layer is connected to the SEM; Block1 contains two DSCM modules, and Block2-Block5 contains three DSCM modules; except for the last DSCN module, the other DSCM modules have two outputs, one output is connected to the next layer of DSCM module, and the second output is to the scale enhancement module SEM. The shallow edge prediction image and the deep edge prediction image output by the feature extraction block of each layer are fused in fusion to obtain an edge feature image. Among them, the number of convolutional layers in the N convolutional layers of the above preprocessing is not uniquely limited. For example, it can contain the following: Figure 1B The three convolutional layers in are connected in sequence: the first convolutional layer is 1×1 with a stride of 32, the second convolutional layer is 3×3 with a stride of 32, and the third convolutional layer is 1×1 with a stride of 16. The above cases are only examples and are not limited to the only ones.

[0047] By using the deep learning model proposed in the present invention, a network with a bidirectional cascade structure is adopted to process the image edges, and a deep separable convolution module is introduced to replace the standard convolution block to reduce the computational workload and simplify the network complexity, so that the deep learning model is more lightweight and can be deployed on mobile devices. At the same time, an attention mechanism is added to the deep separable convolution so as not to reduce the edge detection performance.

[0048] In view of this, the present application proposes an image edge detection method based on a deep learning model to solve the above problems, which is described in detail below.

[0049] See also Figure 2A , Figure 2A A flowchart of an image edge detection method based on a deep learning model provided by an embodiment of the present invention is characterized in that the method comprises the following steps:

[0050] S10, obtaining an input image, wherein the deep learning model includes a preprocessing module and a multi-layer feature extraction block, wherein the multi-layer feature extraction block constitutes a bidirectional cascade structure network. The input image may be a still image or frame from a camera, a video stream, a data set or other image source, and the image will be fed into the deep learning model for further processing.

[0051] The preprocessing module includes a preprocessing layer connected in series, that is, each convolution layer in the multiple convolution layers is connected in series, and the multiple convolution layers can be a 1*1*32 convolution layer, a 3*3*32 convolution layer, or a 1*1*16 convolution layer. This is not a unique limitation, and can be as follows: Figure 1B as shown in .

[0052] Among them, the preprocessing module is used to preprocess the input image, that is, to convert the input image into the format and size required by the deep learning model, and to perform some basic image enhancement operations, such as contrast enhancement, denoising, normalization, etc., which helps to improve the deep learning model's recognition and processing capabilities for the input image.

[0053] Among them, the multi-layer feature extraction block may include five layers of feature extraction blocks, which can be as follows Figure 1B As shown in Block1-Block5 in the figure, each Block includes a depthwise separable convolution DSCM and a scale enhancement module SEM, and the SEM is also connected to a preset convolution layer, which can be a 1*1 convolution layer with a step size of 1. The number of preset convolution layers is not limited and can be 2, 3, 4, etc., and is not uniquely limited here.

[0054] Specifically, in the feature extraction block of each layer, the data output by each SEM is evenly divided and input into the preset convolution layer respectively to sum the data and obtain the output data.

[0055] Among them, the above-mentioned feature extraction block is used to extract useful information or features from the input image to provide support for subsequent tasks (such as classification, detection, segmentation, etc.). In the present invention, the feature extraction block can be used to simultaneously extract deep and shallow features of the input image.

[0056] Among them, in the bidirectional cascade structure network, features are not only transmitted during forward propagation, but also can exchange information during back propagation. This means that the features of the lower layer can directly affect the features of the higher layer, and vice versa. The bidirectional cascade structure network helps the model capture rich features from shallow to deep layers and optimize the performance of tasks such as edge detection.

[0057] It can be seen that the deep learning model can effectively process and analyze the input image and extract the features and information that are crucial to completing specific tasks. At the same time, the use of a bidirectional cascade structure network further enhances the model's feature extraction capability, ensuring that the features extracted from different levels can interact and optimize with each other, thereby improving the performance and accuracy of the model.

[0058] S20, sequentially inputting the input image into the preprocessing module and the feature extraction blocks of each layer, so that the feature extraction blocks of each layer output a shallow edge prediction image and a deep edge prediction image.

[0059] Among them, the preprocessing operations based on the preprocessing module in S10 may include: resizing: adjusting the size of the image to ensure that it meets the input requirements of the model; normalization: converting the pixel value range of the image to the range expected by the model (usually between 0 and 1); denoising: reducing noise in the image and improving image quality; other image enhancement operations: may include contrast enhancement, brightness adjustment, etc., which are not limited here.

[0060] Among them, in this deep learning model, edge prediction is carried out through two different directions of propagation, namely Ps2d (from shallow to deep) from shallow to deep layers and Pd2s (from deep to shallow) from deep to shallow layers. These two propagations reflect different focuses and details of edge detection. In the shallow layers, the network may pay more attention to the basic features of the image, such as edges, textures, and colors. These shallow features change rapidly in space, so they are usually related to the local structure of the image; in the deep layers, the network is more likely to focus on small details or specific parts of the image. For example, the primary layer may focus mainly on large structures, while deeper layers may begin to focus on smaller, more distinguishing features or details, such as the shape, structure, and relationship of objects.

[0061] Specifically, the shallow edge prediction image represents the image Pd2s (from deep to shallow) from the deep layer to the shallow layer, and the deep edge prediction image represents the image Ps2d (from shallow to deep) from the shallow layer to the deep layer.

[0062] It can be seen that the present invention combines shallow and deep predictions through feature extraction blocks, and the deep learning model can capture various features and details in the image more comprehensively and accurately, thereby making more accurate edge detection predictions.

[0063] S30, generating a shallow edge feature image according to the shallow edge prediction images output by the multiple layers of feature extraction blocks.

[0064] Among them, based on the five layers of feature extraction blocks in S10, and each layer of feature extraction blocks in S20 can output a shallow edge prediction image, it can be obtained that the multiple layers of the feature extraction blocks can output five shallow edge prediction images.

[0065] Optionally, the above operation of generating a shallow edge feature image is to sequentially superimpose a plurality of shallow edge prediction images to generate a shallow edge feature image.

[0066] The shallow edge prediction images of different layers are superimposed, which can be pixel-level summation or weighted superposition, which is not limited here. However, before superposition, the feature images output by different layers need to be scale-aligned to ensure that their spatial resolutions are consistent.

[0067] Among them, the above-mentioned shallow edge feature image integrates shallow edge feature images of edge information at multiple levels and scales.

[0068] It can be seen that by superimposing multiple layers of shallow edge prediction images output by the feature extraction blocks, it is possible to more accurately capture the fine and complex edge information in the image, more comprehensively express the edge characteristics of the image, and improve the output edge accuracy.

[0069] S40, generating a deep edge feature image according to the deep edge prediction images output by the multiple layers of feature extraction blocks.

[0070] Among them, based on the five layers of feature extraction blocks in S10, and each layer of feature extraction blocks in S20 can output a deep edge prediction image, it can be obtained that the multiple layers of the feature extraction blocks can output five deep edge prediction images.

[0071] Optionally, the operation of generating the deep edge feature image is to sequentially superimpose a plurality of deep edge prediction images to generate the deep edge feature image.

[0072] The deep edge prediction images of different layers are superimposed, which can be pixel-level summation or weighted superposition, which is not limited here. However, before superposition, the feature images output by different layers need to be scale-aligned to ensure that their spatial resolutions are consistent.

[0073] Among them, the above-mentioned deep edge feature image integrates deep edge feature images of edge information at multiple levels and scales.

[0074] It can be seen that by superimposing multiple layers of deep edge prediction images output by the feature extraction blocks, a deeper and more abstract deep edge feature image is generated, thereby improving the output edge accuracy.

[0075] S50, fusing the shallow edge feature image and the deep edge feature image to obtain a final edge feature image.

[0076] Among them, Figure 1B As shown, the fusion of S50 is performed in the fusion in the figure.

[0077] The above fusion operation is to integrate edge information of different levels and scales to obtain a comprehensive and accurate final edge feature image.

[0078] Specifically, the fusion process can be regarded as a weighted sum operation, and edge feature images at different levels are assigned different weights according to their importance and reliability. The specific weights can be obtained through training (the loss corresponding to each feature extraction block in the figure), or manually set according to prior knowledge; the weighted feature images are superimposed together to form the final edge feature image. In order to enhance the edge response and suppress noise, an activation function or nonlinear transformation can be applied to the result after fusion. Common choices include Sigmoid or ReLU functions, which can highlight strong edges and suppress weak edges. Finally, in order to keep the value range of the final edge feature image consistent and facilitate subsequent processing, normalization can be performed to map the pixel values ​​to a standard range, such as [0,1] or [-1,1].

[0079] Through this fusion process, the final edge feature image combines shallow detail information and deep global information, and can more accurately and robustly represent the edge structure in the image.

[0080] like Figure 2B As shown, Figure 2B It is a schematic diagram of an edge feature image, column A is the original image of the data set, column B is the reference edge image in the data set used for comparison, column C is the output result of the comparison algorithm HED, column D is the output result of the comparison algorithm RCF, column E is the output result of the comparison algorithm BDCN, and column F is the final output edge feature image of the present invention.

[0081] Optionally, after S30 and S40, the method further includes: performing loss calculation on the edge prediction images after the shallow edge prediction images output by the multiple layers of the feature extraction blocks and the deep edge prediction images output by the multiple layers of the feature extraction blocks, that is, performing loss calculation on the five shallow edge prediction images and the five deep edge prediction images that can be output by the multiple layers of the feature extraction blocks to obtain the loss value corresponding to each edge prediction image; after S50, outputting the final edge feature image and the loss value corresponding to each edge prediction image, if it is detected that any loss value does not reach the preset loss threshold, the input image in S10 is re-circulated through steps S10-S50 until the loss values ​​corresponding to all edge prediction images reach the preset loss threshold, that is, the final edge feature image is judged to be the most accurate edge feature image.

[0082] Among them, the above-mentioned preset loss threshold can be set manually or the result obtained after model training, and can be 0.3, which is the only limitation here.

[0083] It can be seen that the final edge feature image will highlight the edge information in the input image while reducing other unimportant details and noise, thereby providing a clear and robust feature representation for subsequent image processing and analysis tasks and improving the accuracy and effect of edge detection.

[0084] This method optimizes the deep learning model by performing bidirectional cascading, and thus has strong adaptability and robustness for images of different qualities and taken in different environments. By fusing shallow edge feature images with deep edge feature images, the model can integrate features at different levels and improve the accuracy of edge detection.

[0085] In a possible example, the step of sequentially inputting the input image into the preprocessing module and the feature extraction blocks of each layer so that the feature extraction blocks of each layer output a shallow edge prediction image and a deep edge prediction image includes: sequentially inputting the input image into the preprocessing module to obtain an edge output image; sequentially inputting the edge output image into the feature extraction blocks of each layer so that the feature extraction blocks of each layer output a shallow edge prediction image, a deep edge prediction image and an edge output image of the current layer, wherein the edge output image of the current layer is the input image of the feature extraction blocks of the next layer.

[0086] Among them, as described in the convolution in the preset processing module in S10, when the input image is sequentially input into the preprocessing module to obtain the edge output image, the input image can be sequentially input into 1*1*32 convolution, 3*3*32 convolution, and 1*1*16 convolution to obtain the edge output image, and the edge output image is equivalent to a feature matrix with edge features.

[0087] Optionally, each layer of the feature extraction block includes multiple layers of parallel feature extraction sublayers, and each layer of the feature extraction sublayer includes depth-separable convolution blocks and scale enhancement blocks connected in series in sequence, so that the edge output image of this layer represents the edge output image of each layer obtained only through multiple layers of depth-separable convolution blocks in each layer of the feature extraction block.

[0088] Optionally, each layer of the feature extraction block includes multiple layers of feature extraction sublayers in parallel, and each layer of the feature extraction sublayer includes a depth-separable convolution block and a scale enhancement block connected in series in sequence. The depth-separable convolution block also includes an attention mechanism module, so that the edge output image of this layer represents the edge output image of each layer obtained only through multiple layers of depth-separable convolution blocks with attention mechanism in each layer of feature extraction block.

[0089] Among them, the above-mentioned shallow edge prediction image, deep edge prediction image and edge output image of this layer are equivalent to a feature matrix with edge features.

[0090] For example, the edge output image is sequentially input into the feature extraction block of each layer, the first layer feature extraction block includes multiple layers of parallel feature extraction sublayers, each layer of the feature extraction sublayer includes a depth-separable convolution block and a scale enhancement block sequentially connected in series, and the depth-separable convolution block also includes an attention mechanism module. The specific process is to input the edge output image on one side into the first layer feature extraction block, and obtain a shallow edge prediction image and a deep edge prediction image through the depth-separable convolution block and the scale enhancement block sequentially connected in series, and at the same time input the edge output image into the depth-separable convolution block in the first layer feature extraction sublayer on the other side to obtain a first edge output image. If there are n layers of feature extraction sublayers in the first layer feature extraction block, and then there are n layers of parallel depth-separable convolutions in the n layers of feature extraction sublayers, then after the first edge output image goes through n layers of parallel depth-separable convolutions in sequence, the edge output image of the first layer can be obtained, that is, the edge output image of this layer. After obtaining the edge output image of the first layer, the edge output image of the first layer is input into the second layer feature extraction block on one side to obtain the shallow edge prediction image and the deep edge prediction image, and the edge output image of the first layer is simultaneously input into the first layer feature extraction sublayer on the other side, and the second edge output image is obtained through the depthwise separable convolution block and the scale enhancement block connected in series in sequence. If there are n layers of feature extraction sublayers in the second layer feature extraction block, and then there are n layers of parallel depthwise separable convolution in the n layers of feature extraction sublayers, then after the second edge output image has passed through n layers of parallel depthwise separable convolution in sequence, the edge output image of the second layer can be obtained, that is, the edge output image of this layer.

[0091] It can be seen that the feature extraction blocks of each layer extract shallow edge prediction images and deep edge prediction images of different dimensions, thereby optimizing the details of edge detection, making the edge detection effect better, and improving the accuracy and efficiency of edge detection.

[0092] In a possible example, each layer of the feature extraction block includes multiple layers of feature extraction sublayers connected in parallel, and the edge output image is sequentially input into the feature extraction block of each layer so that each layer of the feature extraction block outputs a shallow edge prediction image and a deep edge prediction image, including: sequentially inputting the edge output image into each layer of the feature extraction sublayer to obtain an edge prediction subgraph corresponding to the feature extraction sublayer, merging the edge prediction subgraphs of the feature extraction sublayers of each layer to obtain a merged prediction subgraph; and generating a shallow edge prediction image and a deep edge prediction image according to the merged prediction subgraph.

[0093] The multi-layer feature extraction sublayer is responsible for extracting edge information of the image from different perspectives and scales. The feature extraction sublayer can be a depth-wise separable convolution or a depth-wise separable convolution with an attention mechanism and a scale enhancement block. For example, Figure 1BAs shown in the DSCM in the figure, DSCM is a depth-separable convolution with an attention mechanism, and the SEM in the figure represents a scale enhancement block.

[0094] Optionally, the edge prediction subgraphs of the feature extraction sublayers of each layer are merged to obtain a merged prediction subgraph, and the edge prediction subgraphs of each layer of the feature extraction sublayer are evenly divided to obtain a first edge prediction subgraph and a second edge prediction subgraph of each layer of the feature extraction sublayer. In each layer of the feature extraction sublayer, multiple first edge prediction subgraphs and multiple second edge prediction subgraphs can be obtained.

[0095] Optionally, multiple first edge prediction sub-images are merged to obtain a first merged prediction sub-image; the first merged prediction sub-image is passed through a preset convolution layer to generate a shallow edge prediction image.

[0096] Optionally, multiple second edge prediction sub-images are merged to obtain a second merged prediction sub-image; and the second merged prediction sub-image is passed through a preset convolution layer to generate a deep edge prediction image.

[0097] Among them, the above generation process is data processing through a preset convolution layer, and each layer of the feature extraction block also includes a preset convolution layer, and the preset convolution layer can be a 1*1, step-size 1 convolution layer, which is not uniquely limited here. If the feature extraction sublayer can be a depth-separable convolution or a depth-separable convolution with an attention mechanism and a scale enhancement block, the above preset convolution layer is placed after the scale enhancement block, and data is transmitted in the scale enhancement block. The above 1*1, step-size 1 convolution layer is used to perform feature fusion and dimensionality reduction in the channel dimension while retaining spatial information. Through two 1*1, step-size 1 convolution layers, two different edge prediction images are generated, such as Ps2d: an edge prediction image from shallow to deep layers, capturing edge information from fine-grained to coarse-grained; Pd2s: an edge prediction image from deep to shallow layers, capturing edge information from coarse-grained to fine-grained.

[0098] For example, in the process of inputting the edge output image into each layer of the feature extraction sublayer in sequence to obtain the edge prediction subgraph corresponding to the feature extraction sublayer, assuming that each layer of the feature extraction block in the n layers has M layers of feature extraction sublayers, in the same layer of feature extraction blocks, the detailed steps are: the first edge prediction subgraph to be output by the first layer of feature extraction sublayer is divided into two to obtain two first edge prediction subgraphs and second edge prediction subgraphs to be output; and so on, at the Mth layer, the Mth edge prediction subgraph to be output by the Mth layer of feature extraction sublayer is divided into two to obtain two first edge prediction subgraphs and second edge prediction subgraphs to be output; the two first edge prediction subgraphs to be output by the first layer and the two first edge prediction subgraphs to be output by the Mth layer are merged to obtain a first merged prediction subgraph; the two second edge prediction subgraphs to be output by the first layer and the two second edge prediction subgraphs to be output by the Mth layer are merged to obtain a second merged prediction subgraph; finally, the first merged prediction subgraph is passed through a preset convolution layer to generate a shallow edge prediction image; and the second merged prediction subgraph is passed through a preset convolution layer to generate a deep edge prediction image.

[0099] It can be seen that the present invention focuses on different parts and features of the image through feature extraction sublayers at different layers, providing rich and diverse feature representations for subsequent edge detection. That is, by flexibly merging edge prediction subgraphs at different layers, the model can balance computational efficiency and detection accuracy to varying degrees.

[0100] In a possible example, each of the feature extraction sublayers includes a depthwise separable convolution block and a scale enhancement block connected in series, and inputting the edge output image into each of the feature extraction sublayers in sequence to obtain an edge prediction subgraph corresponding to the feature extraction sublayer includes: inputting the edge output image into the depthwise separable convolution block and the scale enhancement block in sequence to obtain an edge prediction subgraph corresponding to the feature extraction sublayer, wherein the output of the depthwise separable convolution of the previous layer is used as the input of the depthwise separable convolution of the next layer.

[0101] The schematic diagram of the depth separation convolution block can be referenced Figure 2C , the structure of the scale enhancement block can be as follows Figure 2D shown.

[0102] Among them, the deep separable convolution block is an efficient convolutional neural network structure, which decomposes the standard convolution into two parts: deep convolution and point-by-point convolution. Deep convolution is responsible for extracting spatial features, and point-by-point convolution is responsible for combining features of different channels. This decomposition method significantly reduces the number of parameters and computational complexity, and improves the efficiency of the network. In the present invention, k=3 (3x3 deep separable convolution) is used, which is not the only limitation here.

[0103] The scale enhancement block SEM is a module used to enhance the spatial feature representation in deep learning models. Its main purpose is to enhance the network's ability to capture spatial information. SEM can help different layers in the model highlight features of specific scales, making the detection of specific scale information more focused and further improving the understanding of local details and spatial structures.

[0104] For example, when there are three layers of feature extraction sublayers in the feature extraction block, based on the edge output image, the first feature matrix is ​​obtained through the first DSCM module in the first layer of feature extraction sublayer in the feature extraction block; the first feature matrix is ​​transmitted to the scale enhancement module corresponding to the first layer of DSCM module to obtain the second feature matrix; the first feature matrix is ​​transmitted to the second layer of DSCM module connected to the first layer of DSCM module as input to obtain the third feature matrix; the third feature matrix is ​​transmitted to the scale enhancement module corresponding to the second layer of DSCM module to obtain the fourth feature matrix; the fourth feature matrix is ​​transmitted to the third layer of DSCM module connected to the second layer of DSCM module as input to obtain the fifth feature matrix; the fifth feature matrix is ​​transmitted to the scale enhancement module corresponding to the third layer of DSCM module to obtain the sixth feature matrix; the above-mentioned second feature matrix, the fourth feature matrix and the sixth feature matrix are the output feature matrices corresponding to the three layers of feature extraction sublayers in the feature extraction block. That is, the second feature matrix is ​​the output feature matrix corresponding to the first layer of feature extraction sublayer, the fourth feature matrix is ​​the output feature matrix corresponding to the second layer of feature extraction sublayer, and the sixth feature matrix is ​​the output feature matrix corresponding to the third layer of feature extraction sublayer. The above feature matrices are all edge prediction subgraphs described in this solution.

[0105] For example, when there are two layers of feature extraction sublayers in the feature extraction block, based on the edge output image, the first feature matrix is ​​obtained through the first DSCM module in the first layer of feature extraction sublayer in the feature extraction block; the first feature matrix is ​​transmitted to the scale enhancement module corresponding to the first layer of DSCM module to obtain the second feature matrix; the first feature matrix is ​​transmitted to the second layer of DSCM module connected to the first layer of DSCM module as input to obtain the third feature matrix; the third feature matrix is ​​transmitted to the scale enhancement module corresponding to the second layer of DSCM module to obtain the fourth feature matrix; the above-mentioned second feature matrix and the fourth feature matrix are the output feature matrices corresponding to the two layers of feature extraction sublayers in the feature extraction block. That is, the second feature matrix is ​​the output feature matrix corresponding to the first layer of feature extraction sublayer, and the fourth feature matrix is ​​the output feature matrix corresponding to the second layer of feature extraction sublayer. The above-mentioned feature matrices are all edge prediction subgraphs described in this scheme.

[0106] It can be seen that the present invention adopts a depth-separable convolution block to separate the depth and cross-channel of the convolution operation, reduces the number of parameters, improves the parameter efficiency of the model, and makes the network more lightweight. It also adopts a scale enhancement block to extract and enhance features of different scales, so that the network can better process objects and details of various sizes in the image, and improve the performance of the model in complex scenes.

[0107] In one possible example, the depthwise separable convolutional block supports an attention mechanism.

[0108] Among them, the purpose of the depthwise separable convolutional block supporting the attention mechanism is to enhance the model's attention to important features and suppress the influence of unimportant features, thereby improving the performance of edge detection.

[0109] The attention mechanism evaluates the features at each position in the feature map to determine their importance. This is usually achieved through a small neural network that outputs an attention score map indicating the importance of the features at each position in the input feature map. The feature map is then weighted using the calculated attention scores so that important features are enhanced and unimportant features are suppressed.

[0110] Among them, you can refer to Figure 2E , Figure 2E Describes depthwise separable convolutions with attention.

[0111] It can be seen that by introducing the attention mechanism in the depthwise separable convolution block, the model can not only process feature information more efficiently, but also capture the edge information in the image more accurately and clearly, thereby improving the overall edge detection performance.

[0112] In a possible example, the depthwise separable convolution block includes an ascending pointwise convolution sub-block, a depthwise separable convolution sub-block, and a descending pointwise convolution sub-block connected in sequence, and inputting the edge output image into the depthwise separable convolution block and the scale enhancement block in sequence to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer includes: inputting the edge output image into the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block, the descending pointwise convolution sub-block, and the scale enhancement block in sequence to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer.

[0113] The depthwise separable convolution block can connect the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block and the descending pointwise convolution sub-block according to the inverted residual structure. The ascending pointwise convolution sub-block is used to increase the number of channels and transform and refine the input features.

[0114] A 1x1 convolution kernel can be used for convolution to increase the depth of the output feature map, thereby obtaining a representation of a high-dimensional feature space and providing richer information for deep convolution.

[0115] The depthwise separable convolution sub-block is used to extract more complex and abstract features in the increased-dimensional feature space. First, a depthwise convolution is performed to perform convolution operations on each channel independently; then these features are integrated through point-by-point convolution to form a new feature representation. That is, more complex image features are captured while maintaining computational efficiency.

[0116] Among them, the descending point-by-point convolution sub-block is used to reduce the depth of the feature map and convert the high-dimensional feature space back to the low-dimensional space. A 1x1 convolution kernel can be used for convolution to reduce the depth of the output feature map, thereby generating a more compact feature representation to prepare for subsequent processing.

[0117] For example, the depthwise separable convolution block consists of three parts, including two pointwise convolutions and a middle depthwise separable convolution. After the image passes through the 1x1 pointwise convolution, the number of channels will increase so that the middle 3x3 depthwise separable convolution can extract features. The second pointwise convolution converts the number of channels back to the number of channels at the input, and passes the corresponding activation function to obtain weights (to avoid gradient disappearance), and then outputs the result to the next module. Figure 2C As shown, N of each convolution represents the number of ascending multiples, C represents the number of channels, and the second PW is followed by a linear activation function.

[0118] It can be seen that the present invention can process images more finely and extract useful edge information.

[0119] In one possible example, the depthwise separable convolution sub-block adopts a nonlinear activation function.

[0120] Among them, the depth-separable convolution sub-block adopts a nonlinear activation function to increase the network's expressive power, enabling it to capture complex features and patterns, reduce the number of memory accesses, and thus significantly reduce latency costs. The nonlinear activation function introduces nonlinear transformations to help the network learn complex mapping relationships from input to output.

[0121] Optionally, the non-linear activation function is as follows:

[0122] It can be seen that in the depthwise separable convolution sub-block, the nonlinear activation function is usually applied after the depthwise convolution and point-by-point convolution, which helps the model learn the complex mapping relationship from the underlying features to the high-level features, thereby improving the performance of the edge detection task.

[0123] In a possible example, the depthwise separable convolution block also includes an attention mechanism module, which is connected in series between the depthwise separable convolution sub-block and the descending pointwise convolution sub-block, and the step of sequentially inputting the edge output image into the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block, the descending pointwise convolution sub-block and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer includes: sequentially inputting the edge output image into the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block, the attention mechanism module, the descending pointwise convolution sub-block and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer.

[0124] Among them, the above-mentioned depth-wise separable convolution block can connect the ascending point-by-point convolution sub-block, the depth-wise separable convolution sub-block, the descending point-by-point convolution sub-block and the attention mechanism module according to the inverted residual structure.

[0125] For example, first, the edge output image is input into the ascending point-by-point convolution sub-block, which expands the feature space by increasing the number of channels and provides richer information for the deep convolution; then, the processed feature map is transmitted to the depth-separable convolution sub-block, which includes depth-wise convolution and point-by-point convolution. The depth-wise convolution performs independent convolution on each input channel to capture spatial features; the point-by-point convolution integrates the information of different channels through 1x1 convolution to capture the dependencies between channels; then, the feature map is transmitted to the attention mechanism module. This module strengthens useful features and suppresses irrelevant information by focusing on important areas in the image (usually areas containing target edges), which helps the model locate and detect edges more accurately; the feature map processed by the attention mechanism module is then transmitted to the descending point-by-point convolution sub-block, which compresses the feature space by reducing the number of channels, thereby reducing the computational complexity; finally, the processed feature map is input to the scale enhancement block, which further processes the feature map, which may include up / down sampling, feature fusion, etc., to ensure that the feature map can maintain consistent performance at different scales.

[0126] It can be seen that this solution adds the deep separable convolution block of the attention mechanism module to enhance the model's ability to recognize important features in the image, thereby improving the performance of edge detection. Further, by focusing on the key areas in the image and performing feature processing through multi-layer convolution sub-blocks, the model is ensured to be efficient and accurate in performing edge detection tasks.

[0127] It should be noted that, in each of the above-mentioned embodiments, there is not necessarily a certain order between the above-mentioned steps. A person skilled in the art can understand, based on the description of the embodiments of the present application, that in different embodiments, the above-mentioned steps may have different execution orders, that is, they may be executed in parallel, may be executed interchangeably, and so on.

[0128] As another aspect of the embodiment of the present application, the embodiment of the present application provides an image edge detection device based on a deep learning model. The image edge detection device based on a deep learning model can be a software module, which includes several instructions stored in a memory, and the processor can access the memory and call the instructions for execution to complete the image edge detection method based on the deep learning model described in the above-mentioned various embodiments.

[0129] See also Figure 3 , Figure 3 Schematic diagram of the structure of an image edge detection device based on a deep learning model provided in an embodiment of the present application. Figure 3 As shown, the image edge detection device based on the deep learning model includes:

[0130] An acquisition unit 301 acquires an input image, wherein the deep learning model includes a preprocessing module and a multi-layer feature extraction block, wherein the multi-layer feature extraction block constitutes a bidirectional cascade structure network;

[0131] A processing unit 302 is used to sequentially input the input image into the preprocessing module and the feature extraction blocks of each layer, so that each feature extraction block outputs a shallow edge prediction image and a deep edge prediction image;

[0132] The processing unit 302 is further used to generate a shallow edge feature image according to the shallow edge prediction images output by the multiple layers of feature extraction blocks;

[0133] The processing unit 302 is further configured to generate a deep edge feature image according to the deep edge prediction images output by the multiple layers of feature extraction blocks;

[0134] The fusion unit 303 is used to fuse the shallow edge feature image and the deep edge feature image to obtain a final edge feature image.

[0135] This method optimizes the deep learning model by performing bidirectional cascading, and thus has strong adaptability and robustness for images of different qualities and taken in different environments. By fusing shallow edge feature images with deep edge feature images, the model can integrate features at different levels and improve the accuracy of edge detection.

[0136] In one embodiment, in the step of sequentially inputting the input image into the preprocessing module and the feature extraction blocks of each layer so that each feature extraction block outputs a shallow edge prediction image and a deep edge prediction image, the processing unit 302 is further used to: sequentially input the input image into the preprocessing module to obtain an edge output image; sequentially input the edge output image into the feature extraction blocks of each layer so that each feature extraction block of each layer outputs a shallow edge prediction image, a deep edge prediction image and an edge output image of the current layer, and the edge output image of the current layer is the input image of the feature extraction blocks of the next layer.

[0137] In one embodiment, the feature extraction block in each layer includes multiple layers of feature extraction sublayers in parallel, and the edge output image is sequentially input into the feature extraction block in each layer so that each feature extraction block outputs a shallow edge prediction image and a deep edge prediction image. The processing unit 302 is also used to: sequentially input the edge output image into each feature extraction sublayer to obtain an edge prediction subgraph corresponding to the feature extraction sublayer; merge the edge prediction subgraphs of the feature extraction sublayers in each layer to obtain a merged prediction subgraph; and generate a shallow edge prediction image and a deep edge prediction image according to the merged prediction subgraph.

[0138] In one embodiment, each of the feature extraction sublayers includes a depth-separable convolution block and a scale enhancement block connected in series in sequence, and the edge output image is sequentially input into each of the feature extraction sublayers to obtain an edge prediction subgraph corresponding to the feature extraction sublayer. The processing unit 302 is also used to: sequentially input the edge output image into the depth-separable convolution block and the scale enhancement block to obtain an edge prediction subgraph corresponding to the feature extraction sublayer, wherein the output of the depth-separable convolution block of the previous layer is used as the input of the depth-separable convolution block of the next layer.

[0139] In one embodiment, the depthwise separable convolutional block supports an attention mechanism.

[0140] In one embodiment, the depthwise separable convolution block includes an ascending pointwise convolution sub-block, a depthwise separable convolution sub-block, and a descending pointwise convolution sub-block connected in sequence, and the edge output image is sequentially input into the depthwise separable convolution block and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer. The processing unit 302 is also used to: sequentially input the edge output image into the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block, the descending pointwise convolution sub-block, and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer.

[0141] In one embodiment, the depthwise separable convolution sub-block adopts a non-linear activation function.

[0142] In one embodiment, the depthwise separable convolution block also includes an attention mechanism module, which is connected in series between the depthwise separable convolution sub-block and the descending pointwise convolution sub-block, and the edge output image is sequentially input into the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block, the descending pointwise convolution sub-block and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer. The processing unit 302 is also used to: sequentially input the edge output image into the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block, the attention mechanism module, the descending pointwise convolution sub-block and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer.

[0143] In some embodiments, the image edge detection device based on the deep learning model can also be constructed by hardware devices. For example, the image edge detection device based on the deep learning model can be constructed by one or more chips, and each chip can work in coordination with each other to complete the image edge detection method based on the deep learning model described in the above embodiments. For another example, the image edge detection device based on the deep learning model can also be constructed by various logic devices, such as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.

[0144] It should be noted that the above-mentioned image edge detection device based on a deep learning model can execute the image edge detection method based on a deep learning model provided in the embodiment of the present application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in the embodiment of the image edge detection device based on a deep learning model, please refer to the image edge detection method based on a deep learning model provided in the embodiment of the present application.

[0145] See also Figure 4 , Figure 4 1 is a schematic diagram of a computer device provided in an embodiment of the present application. The computer device includes one or more processors and a memory. The memory is connected to the one or more processors, for example, connected to the processor via a bus.

[0146] The processor is configured to support the computer device to perform the corresponding functions in the method in the above method embodiment. The processor can be a central processing unit (CPU), a network processor (NP), a hardware chip or any combination thereof. The above hardware chip can be an application specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0147] The memory is used to store program codes, etc. The memory may include volatile memory (VM), such as random access memory (RAM); the memory may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory may also include a combination of the above types of memory.

[0148] The memory can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the image edge detection method based on a deep learning model in the embodiment of the present application. The processor executes various functional applications and data processing of the image edge detection method based on a deep learning model and the image edge detection device based on a deep learning model by running the non-volatile software programs, instructions and modules stored in the memory, that is, realizes the functions of each module or unit of the image edge detection method based on a deep learning model and the image edge detection device based on a deep learning model provided in the above method embodiment.

[0149] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function. The data storage area may store data created based on the use of the image edge detection device based on the deep learning model, etc. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the image edge detection device based on the deep learning model via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0150] The one or more modules are stored in the memory, and when executed by the one or more processors, the image edge detection method based on the deep learning model in any of the above-mentioned method embodiments is executed, for example, the method steps described in the above-mentioned method embodiments are executed to realize the functions of the modules described in the above-mentioned device embodiments.

[0151] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method described in the above embodiment.

[0152] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0153] The above disclosure is only the preferred embodiment of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. An image edge detection method based on a deep learning model, characterized in that: include: Acquire an input image, wherein the deep learning model includes a preprocessing module and a multi-layer feature extraction block, wherein the multi-layer feature extraction blocks constitute a bidirectional cascade structure network; The input image is sequentially input into the preprocessing module and the feature extraction blocks of each layer, so that the feature extraction blocks of each layer output a shallow edge prediction image and a deep edge prediction image; Generate a shallow edge feature image according to the shallow edge prediction images output by the multiple layers of feature extraction blocks; Generate a deep edge feature image according to the deep edge prediction images output by the multiple layers of feature extraction blocks; The shallow edge feature image and the deep edge feature image are fused to obtain a final edge feature image.

2. The method according to claim 1, characterized in that The step of sequentially inputting the input image into the preprocessing module and the feature extraction blocks of each layer so that the feature extraction blocks of each layer output a shallow edge prediction image and a deep edge prediction image comprises: Inputting the input images into the preprocessing module in sequence to obtain edge output images; The edge output image is sequentially input into the feature extraction block of each layer, so that each layer of the feature extraction block outputs a shallow edge prediction image, a deep edge prediction image and an edge output image of the current layer, and the edge output image of the current layer is the input image of the feature extraction block of the next layer.

3. The method according to claim 2, characterized in that Each layer of the feature extraction block includes multiple layers of feature extraction sublayers connected in parallel, and the edge output image is sequentially input into each layer of the feature extraction block so that each layer of the feature extraction block outputs a shallow edge prediction image and a deep edge prediction image, comprising: Inputting the edge output image into each feature extraction sublayer in sequence to obtain an edge prediction subgraph corresponding to the feature extraction sublayer; Merging the edge prediction subgraphs of the feature extraction sub-layers of each layer to obtain a merged prediction subgraph; A shallow edge prediction image and a deep edge prediction image are generated according to the merged prediction sub-image.

4. The method according to claim 3, characterized in that Each of the feature extraction sublayers comprises a depth-separable convolution block and a scale enhancement block connected in series in sequence, and the edge output image is sequentially input into each of the feature extraction sublayers to obtain an edge prediction subgraph corresponding to the feature extraction sublayer, comprising: The edge output image is sequentially input into the depthwise separable convolution block and the scale enhancement block to obtain an edge prediction sub-image corresponding to the feature extraction sub-layer, wherein the output of the depthwise separable convolution block of the previous layer is used as the input of the depthwise separable convolution block of the next layer.

5. The method according to claim 4, characterized in that The depthwise separable convolutional block supports the attention mechanism.

6. The method according to claim 4, characterized in that The depth-separable convolution block includes ascending point-by-point convolution sub-blocks, depth-separable convolution sub-blocks, and descending point-by-point convolution sub-blocks connected in series in sequence, and the edge output image is sequentially input into the depth-separable convolution block and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer, including: The edge output image is sequentially input into the ascending point-by-point convolution sub-block, the depth-separable convolution sub-block, the descending point-by-point convolution sub-block and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer.

7. The method according to claim 6, characterized in that The depth-wise separable convolution sub-block adopts a non-linear activation function.

8. The method according to claim 6, characterized in that The depthwise separable convolution block further includes an attention mechanism module, which is connected in series between the depthwise separable convolution sub-block and the descending pointwise convolution sub-block, and the edge output image is sequentially input into the ascending pointwise convolution sub-block, the depthwise separable convolution sub-block, the descending pointwise convolution sub-block and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer, including: The edge output image is sequentially input into the ascending point-by-point convolution sub-block, the depth-separable convolution sub-block, the attention mechanism module, the descending point-by-point convolution sub-block and the scale enhancement block to obtain an edge prediction sub-graph corresponding to the feature extraction sub-layer.

9. An image edge detection device based on a deep learning model, characterized in that: include: An acquisition unit acquires an input image, wherein the deep learning model includes a preprocessing module and a multi-layer feature extraction block, wherein the multi-layer feature extraction block constitutes a bidirectional cascade structure network; A processing unit, used for sequentially inputting the input image into the preprocessing module and the feature extraction blocks of each layer, so that the feature extraction blocks of each layer output a shallow edge prediction image and a deep edge prediction image; The processing unit is further used to generate a shallow edge feature image according to the shallow edge prediction images output by the multiple layers of feature extraction blocks; The processing unit is further used to generate a deep edge feature image according to the deep edge prediction images output by the multiple layers of feature extraction blocks; The fusion unit is used to fuse the shallow edge feature image and the deep edge feature image to obtain a final edge feature image.

10. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory is connected to the processor, and the processor is used to execute one or more computer programs stored in the memory. When the processor executes the one or more computer programs, the computer device implements the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 8.