Roller cage shoe edge image extraction method and system based on multi-layer attention
By adopting a multi-layer attention mechanism method in the image extraction of roller tank ear edges, the problem of multi-scale feature fusion and insufficient detail edge extraction is solved, which significantly improves the accuracy and robustness of roller tank ear edge detection, providing effective technical support for roller tank ear wear status monitoring.
Patent Information
- Application Number
- CN202510220516.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The prior art has problems in the extraction of roller tank ear edge images that are difficult to fully integrate multi-scale features, insufficient extraction of details, and limited attention to key areas in the extraction of roller tank ear edges, making it difficult to provide effective support for monitoring roller tank ear wear status.
The multi-layered attention-based roller tank ear edge image extraction method is adopted. Through the operations of multi-layered feature extraction, foreground fusion, edge feature fusion and attention mechanism, the multi-scale features of the image are fully captured, the key edge areas are adaptively focused, background noise is suppressed, and edge detection results are optimized through post-processing.
It significantly improves the accuracy, robustness and adaptability of the edge detection of roller tank ears, and generates more complete, clear and noise-resistant edge information, providing important technical support for real-time monitoring and safety assessment of roller tank ear wear status.
Smart Images

Figure CN120107295A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer image processing, and in particular relates to a method and system for extracting edge images of a roller can ear based on multi-layer attention. Background Art
[0002] In the process of coal mining, the shaft hoisting system is a key equipment for transporting coal, equipment and personnel. The roller lugs, as an important part of the hoisting system, ensure the stability and safety of the system by running up and down along the rigid tankway. However, due to complex working conditions and long-term operation, the roller lugs are often worn, which may lead to a decrease in the efficiency of the hoisting system and even safety hazards. Especially in the wear state monitoring, the accurate detection of the edge of the roller lugs is very important, and its results directly affect the accuracy and reliability of the subsequent wear analysis.
[0003] Traditional edge detection methods, such as the Sobel operator, Prewitt operator, and Canny algorithm, achieve edge extraction by calculating the grayscale gradient of the image, and are widely used due to their high computational efficiency. However, these methods have many problems in actual industrial scenarios, such as being sensitive to illumination and noise, and the extracted edge lines are rough and incomplete, which cannot meet the accuracy requirements in complex scenarios. In addition, methods based on texture features are not able to adapt to the irregularity of the edges of industrial equipment, and edge information is easily lost.
[0004] In recent years, the rise of deep learning technology has provided new solutions for edge detection. With the feature learning ability of neural networks, large-scale data-driven methods have shown significant advantages in extracting edge features. However, such methods still face challenges when applied to the scene of roller can lug edge image extraction, such as the difficulty in fully integrating multi-scale features, insufficient extraction of detailed edges, and limited ability to focus on key areas, making it difficult to provide effective support for roller can lug wear status monitoring. Summary of the invention
[0005] The purpose of the present invention is to propose a roller can ear edge image extraction method based on multi-layer attention, which is conducive to the efficient extraction and detection of roller can ear edge features through multi-level feature extraction, foreground fusion, edge feature fusion and attention operations.
[0006] In order to achieve the above object, the present invention adopts the following technical scheme: A method for extracting the edge image of a roller can ear based on multi-layer attention comprises the following steps: Step 1. Preprocess the edge detection data set and the original data of the roller can ear to construct a training data set; Step 2. Use the multi-level feature extraction module to perform multi-level feature extraction and fusion on the images in the training data set to obtain fusion features; Step 3. Use the foreground feature fusion module to enhance the fusion features and obtain a foreground feature map; use the edge feature fusion module to enhance the fusion features and obtain a shallow edge feature map, a middle edge feature map, and a deep edge feature map; Step 4. Use the attention mechanism module to enhance the foreground feature map, shallow edge feature map, middle edge feature map, and deep edge feature map respectively; Step 5. Post-process and fuse the enhanced shallow edge feature map, the enhanced middle edge feature map, and the enhanced deep edge feature map to obtain a binary edge map; Step 6. Use the binary edge map and enhanced foreground feature map, combined with the labels in the training data set, to back-propagate and train the multi-level feature extraction module, foreground feature fusion module, edge feature fusion module, and attention mechanism module; Step 7. Extract the edge image of the roller can ear from the input image to obtain a visualization image of the edge of the roller can ear.
[0007] In addition, based on the roller can ear edge image extraction method based on multi-layer attention, the present invention also proposes a roller can ear edge image extraction system based on multi-layer attention, and its technical solution is as follows: A roller can ear edge image extraction system based on multi-layer attention includes an image sensor, a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it is used to implement the steps of the roller can ear edge image extraction method based on multi-layer attention mentioned above.
[0008] The present invention has the following advantages: As described above, the present invention relates to a method for extracting edge images of roller can ears based on multi-layer attention. Through a multi-level feature extraction module, the multi-scale features of the image are fully captured, so that shallow details to deep semantic information can be fully utilized; through the introduction of foreground feature fusion module, edge fusion module and attention mechanism, the foreground features and edge features are effectively separated and enhanced, which greatly improves the adaptability of the model to complex industrial scenes and the ability to focus on key areas; through the dynamic threshold and multi-scale fusion strategy of post-processing, the edge detection results are further optimized, making the edge information more complete, clear and noise-resistant. Compared with the traditional edge image extraction method, the method of the present invention does not require manual adjustment of parameters, and can realize automatic edge image extraction through data-driven, which significantly improves the accuracy, robustness and adaptability of edge detection, especially in the detection of irregular edges of industrial equipment such as roller can ears. It shows higher accuracy and efficiency, and provides important technical support for real-time monitoring and safety assessment of equipment wear status. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 It is a flowchart of the method for extracting the edge image of the ear of a roller can based on multi-layer attention in an embodiment of the present invention.
[0010] Figure 2 It is a system block diagram of the roller can ear edge image extraction method based on multi-layer attention in an embodiment of the present invention.
[0011] Figure 3 2 is a network structure diagram of the attention mechanism module in an embodiment of the present invention.
[0012] Figure 4 This is an image of a roller can ear used in the test in the embodiment of the present invention.
[0013] Figure 5 It is a visualized image of the edge of the roller can ear obtained by using the edge image extraction method of the present invention. DETAILED DESCRIPTION
[0014] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments: Example 1 In this embodiment, a roller can ear edge image extraction method based on multi-layer attention is proposed, which is used for edge feature extraction of roller can ears and is suitable for scenarios such as roller can ear wear detection during coal mining. The method of the present invention targets the specific needs of roller can ear edge detection, combines a multi-level feature extraction strategy, supports capturing rich edge features from different scales; uses an attention mechanism to adaptively focus on key edge areas and suppress background noise; generates high-quality edge detection results by fusing shallow, middle and deep edge information; and is adaptable to a variety of input data sources, has strong versatility and extensibility, and is suitable for a variety of industrial scenarios.
[0015] The method of the present invention aims to improve the accuracy and robustness of edge detection of roller can ears. It combines multi-level feature extraction, foreground and edge feature fusion and attention mechanism modules, and realizes efficient extraction of edge features of roller can ears through adaptive feature enhancement strategies. The method generally includes the following steps: importing and preprocessing edge detection data sets and original images of roller can ears to construct enhanced data sets; designing multi-level feature extraction modules to gradually extract and fuse features from shallow to deep layers; using foreground feature fusion modules and edge feature fusion modules to generate feature maps of different levels; enhancing features of each level through channel and spatial attention mechanisms, and generating high-precision binary edge maps through post-processing. This method significantly improves the detection capabilities of detail edges and global contours, improves detection accuracy and robustness, and provides effective support for the wear status monitoring of roller can ears.
[0016] like Figure 1 As shown, the method for extracting the edge image of the ear of a roller can based on multi-layer attention includes the following steps: Step 1. Data import and preprocessing.
[0017] Import and preprocess the edge detection dataset and the customized roller can ear raw data to construct the training dataset D.
[0018] The customized roller can ear raw data refers to the roller can ear raw image obtained from the acquisition device.
[0019] The imported edge detection dataset is the HED-BSDS edge detection dataset. The HED-BSDS edge detection dataset is a dataset used for edge detection tasks. The dataset includes 500 original images with a resolution of 481*321 pixels.
[0020] Each original image has corresponding edge annotations, and the HED-BSDS edge detection dataset is widely used in edge detection algorithms.
[0021] The preprocessing process is that the input module uniformly adjusts the image size and pixel value range of all images imported from the HED-BSDS edge detection dataset and the original images of the roller can ear, and generates the enhanced training dataset D by random cropping, rotation, horizontal flipping and adding noise.
[0022] Specifically, for all images imported from the edge detection dataset and the custom roller can ear raw data, the image size and pixel value range are uniformly adjusted using the input module. The image size is uniformly adjusted to ,in is the image height, is the image width, and the image pixel value range is scaled to [0, 1]. All images are copied once and placed in the original position. The copied images are randomly cropped, rotated, horizontally flipped, and Gaussian noise is added to obtain the enhanced dataset D.
[0023] Step 2. Multi-level feature extraction.
[0024] Use the multi-level feature extraction module Net1 to extract the images in the training dataset D Perform multi-level feature extraction and fusion to obtain fusion features .
[0025] The multi-level feature extraction module Net1 includes a shallow feature extraction module Net11, a sub-shallow feature extraction module Net12, a middle feature extraction module Net13, a sub-deep feature extraction module Net14 and a deep feature extraction module Net15.
[0026] The input module takes images from the training dataset D The input image is passed into the multi-level feature extraction module Net1 to extract features from shallow to deep layers. Specifically, the input image passes through five sub-networks, Net11, Net12, Net13, Net14, and Net15. Each of Net11 to Net15 consists of a convolutional layer to capture edge details and global semantic information layer by layer. Then, features at different levels are fused through jump connections to generate fused features. .
[0027] Among them, the input of the shallow feature extraction module Net11 is the image in the training dataset D , used to extract shallow features through convolution operations and ReLU activation functions, shallow features The specific extraction process is as follows: .
[0028] in is the convolution kernel, The size is 3×3, the stride is 1, and the padding is 1; represents the convolution operation, is the image in the training dataset D , is the bias, is the activation function ReLU.
[0029] The input of the sub-shallow feature extraction module Net12 is the shallow feature , used to extract sub-shallow features through convolution operations, maximum pooling operations and ReLU activation functions, sub-shallow features The specific extraction process is as follows: .
[0030] in is the convolution kernel, The size is 3×3, the stride is 1, and the padding is 1; It is a maximum pooling operation, which is used to reduce the spatial dimension of the feature map; For bias.
[0031] The input of the middle-level feature extraction module Net13 is the sub-shallow feature , used to extract mid-level features through convolution operations, maximum pooling operations, and ReLU activation functions. The specific extraction process is as follows: .
[0032] in is the convolution kernel, The size is 3×3, the stride is 1, and the padding is 1. For bias.
[0033] The input of the sub-deep feature extraction module Net14 is the mid-level features , used to extract sub-deep features through convolution operations, maximum pooling operations and ReLU activation functions, sub-deep features The specific extraction process is as follows: .
[0034] in is the convolution kernel, The size is 3×3, the stride is 1, and the padding is 1. For bias.
[0035] The input of the deep feature extraction module Net15 is the sub-deep feature , used to extract deep features through convolution operations, maximum pooling operations and ReLU activation functions, deep features The specific extraction process is as follows: .
[0036] in is the convolution kernel, The size is 3×3, the stride is 1, and the padding is 1. For bias.
[0037] The fusion in step 2 specifically refers to integrating features at different levels through jump connections to obtain fused features. , the specific process is: .
[0038] in, is the feature concatenation function, , , , They are , , , Feature map obtained by upsampling with bilinear interpolation.
[0039] Step 3. Feature fusion.
[0040] Use foreground feature fusion module Net2 to enhance fusion features , get the foreground feature map .
[0041] Use edge feature fusion module Net3 to enhance fusion features , get the shallow edge feature map , middle layer edge feature map and deep edge feature maps .
[0042] Step 3 uses the foreground feature fusion module Net2 and the edge feature fusion module Net3 to further divide the fused features into a foreground feature map and a multi-layer edge feature map. The purpose of dividing the foreground and edge features is to improve the accuracy and efficiency of image processing tasks. By clearly distinguishing the foreground and the edge, the algorithm can more accurately identify the shape and position of the object, reduce background interference, and thus enhance the edge detection effect. In addition, after the division, the foreground area can be processed intensively, thereby reducing the computational complexity and improving the robustness of the system in different environments.
[0043] In this embodiment, the foreground feature fusion module Net2 is a 1×1 convolutional layer.
[0044] Through the foreground feature fusion module Net2, the fusion feature After a convolution layer with a convolution kernel size of 1×1, the dimension is reduced and then processed by the ReLU activation function to generate a preliminary foreground feature map. , the specific process is: .
[0045] in, is the convolution kernel, The size is 1×1, is the bias, is the activation function ReLU.
[0046] In this embodiment, the edge feature fusion module Net3 is three 1×1 convolutional layers.
[0047] Through the edge feature fusion module Net3, the fusion feature According to different depth levels, the features are divided into shallow, middle and deep layers, and the shallow edge feature maps are generated respectively through a 1×1 convolution layer with a convolution kernel size of 1×1 and a ReLU activation function. , middle layer edge feature map , deep edge feature map , the specific process is: .
[0048] .
[0049] .
[0050] in, , , is the convolution kernel, , , The size is 1×1, , , For bias.
[0051] Step 4. Feature enhancement.
[0052] Use the attention mechanism module Net4 to enhance the foreground feature map , shallow edge feature map , middle layer edge feature map , deep edge feature map , and get the enhanced foreground feature map , Enhance shallow edge feature map , Enhance the middle layer edge feature map , Enhanced deep edge feature map .
[0053] The attention mechanism module Net4 is used to enhance features, highlight key areas from the channel and spatial dimensions, and generate enhanced foreground feature maps and enhanced shallow, middle, and deep edge feature maps.
[0054] Step 4.1. Calculate the channel attention mechanism.
[0055] Foreground feature map , shallow edge feature map , middle layer edge feature map , deep edge feature map Calculate channel weights separately: .
[0056] .
[0057] .
[0058] .
[0059] in, represents the global average pooling operation, is a multi-layer perceptron, is the activation function; , , , Foreground feature maps , shallow edge feature map , middle layer edge feature map , deep edge feature map Channel attention.
[0060] Figure 3 Channel attention shown include , , , .
[0061] Step 4.2. Compute the spatial attention mechanism.
[0062] Foreground feature map , shallow edge feature map , middle layer edge feature map , deep edge feature map Calculate spatial weights separately: .
[0063] .
[0064] .
[0065] .
[0066] in, Represents a 2D convolution operation; , , , Foreground feature maps , shallow edge feature map , middle layer edge feature map , deep edge feature map spatial attention. Figure 3 The spatial attention shown include , , , .
[0067] Step 4.3. Feature map , , , Enhanced to obtain enhanced foreground feature map , Enhance shallow edge feature map , Enhance the middle layer edge feature map , Enhanced deep edge feature map .
[0068] .
[0069] .
[0070] .
[0071] .
[0072] in, Represents an element-wise product operation.
[0073] like Figure 3 As shown in Figure 2, the input of the attention mechanism module Net4 is the foreground feature map output by the foreground feature fusion module Net2. , and the shallow edge feature map output by the edge feature fusion module Net3 , middle layer edge feature map , deep edge feature map The foreground information output of the attention mechanism module Net4 is the enhanced foreground feature map , shallow edge information output, shallow edge information output, shallow edge information output are respectively enhanced shallow edge feature maps , Enhance the middle layer edge feature map , Enhance the deep edge feature map .
[0074] Step 5. Post-processing and edge extraction.
[0075] To enhance the shallow edge feature map , Enhance the middle layer edge feature map , Enhance the deep edge feature map Post-process and fuse to obtain a binary edge map .
[0076] The post-processing module performs binarization processing on the enhanced edge feature map and fuses it to generate the final binary edge map. The specific process is as follows: Step 5.1. Enhance the shallow edge feature map , middle layer edge feature map , deep edge feature map , a dynamic threshold method is used for binarization processing. In this embodiment, the dynamic threshold method preferably adopts the Otsu method, that is, the maximum inter-class variance method.
[0077] Step 5.2. The binarized enhanced shallow feature map, the enhanced middle feature map and the enhanced deep feature map are fused element by element, specifically, the enhanced shallow feature map, the enhanced middle feature map and the enhanced deep feature map are ORed.
[0078] Step 6. Network training and optimization.
[0079] Using binary edge maps and enhanced foreground feature map , combined with the labels in the training data set D, back-propagation training is performed on the multi-level feature extraction module Net1, the foreground feature fusion module Net2, the edge feature fusion module Net3 and the attention mechanism module Net4.
[0080] Step 6 combines the binary edge map and the enhanced foreground feature map, and performs back-propagation training on the multi-level feature extraction module, feature fusion module, and attention mechanism module through the dataset label.
[0081] Step 7. Edge feature detection.
[0082] Use the trained multi-level feature extraction module Net1, edge feature fusion module Net3 and attention mechanism module Net4 to extract the input image The edge image of the roller can ear is extracted to obtain a visualized image of the edge of the roller can ear.
[0083] The real-time collected roller tank ear images are input into the trained network model, and the foreground segmentation results and high-precision edge feature maps are output, providing reliable edge feature extraction support for the wear status detection of roller tank ears.
[0084] Specifically, first, the input image Perform preprocessing and uniformly adjust the image size and pixel value range.
[0085] Then, the multi-level feature extraction module Net1 is used to extract and fuse multi-level features to obtain fused features.
[0086] Then the edge feature fusion module Net3 is used to enhance the fusion features to obtain shallow edge feature maps, middle edge feature maps and deep edge feature maps.
[0087] Then the attention mechanism module Net4 is used to enhance the shallow edge feature map, the middle edge feature map, and the deep edge feature map respectively to obtain the enhanced shallow edge feature map, the enhanced middle edge feature map, and the enhanced deep edge feature map.
[0088] Finally, the enhanced shallow edge feature map, the enhanced middle edge feature map, and the enhanced deep edge feature map are post-processed and fused to obtain a binary edge map, i.e., a visualization image of the edge of the roller can ear.
[0089] In addition, in order to verify the effectiveness of the roller can ear edge image extraction method based on multi-layer attention proposed in the present invention, the Canny edge detection algorithm, the global nested edge detection algorithm HED, the sparse coding gradient edge detection algorithm SCG and the method of the present invention are used to extract the roller can ear edge image.
[0090] The indicators for measuring the edge detection effect include the global optimal threshold (ODS, Optimal Dataset Scale), the single image optimal threshold (OIS, Optimal Image Scale) and the average precision (AP, Average Precision).
[0091] ODS evaluates edge detection results at a series of different scales and finds the corresponding evaluation results at the scale that achieves the best detection effect. OIS finds the optimal scale for each image separately and calculates the average of the evaluation results at these optimal scales. Unlike ODS, OIS focuses on the optimal scale of each image itself. AP calculates the area under the precision-recall curve, or PR curve. Precision refers to the ratio of detected true edge pixels to all detected edge pixels, and recall refers to the ratio of detected true edge pixels to all real edge pixels.
[0092] Table 1 Comparison results between the method of the present invention and the comparative method
[0093] It can be seen from Table 1 above that the roller can ear edge image extraction method based on multi-layer attention of the present invention is superior to other traditional edge detection algorithms in various indicators for measuring edge detection effects.
[0094] Example 2 This embodiment 2 describes a roller can ear edge image extraction system based on multi-layer attention, which is based on the same inventive concept as the roller can ear edge image extraction method based on multi-layer attention in the above-mentioned embodiment 1.
[0095] Specifically, the roller can ear edge image extraction system based on multi-layer attention includes an image sensor, a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, the steps of the roller can ear edge image extraction method based on multi-layer attention as described in Example 1 are implemented.
[0096] In this embodiment, the image sensor preferably uses a CMOS image sensor, and the processor preferably uses a data processor equipped with an RTX4060 graphics card. The roller can ear edge image extraction system based on multi-layer attention uses high-performance computing equipment and CMOS image sensors, supports real-time edge detection, can provide technical support for roller can ear wear status monitoring, and has strong practicality and scalability.
[0097] The memory can be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or an external storage device of any device or apparatus with data processing capabilities, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card, etc. equipped on the device; the processor can also be any device or apparatus with data processing capabilities, which will not be described in detail here.
[0098] Of course, the above description is only a preferred embodiment of the present invention, and the present invention is not limited to the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any technician familiar with the field under the guidance of this specification fall within the essential scope of this specification and should be protected by the present invention.
Claims
1. A method for extracting the edge image of a roller can ear based on multi-layer attention, characterized in that: The steps include: Step 1. Preprocess the edge detection data set and the original data of the roller can ear to construct a training data set; Step 2. Use the multi-level feature extraction module to extract and fuse multi-level features of the images in the training data set to obtain fused features; Step 3. Use the foreground feature fusion module to enhance the fusion features and obtain the foreground feature map; The edge feature fusion module is used to enhance the fusion features to obtain shallow edge feature maps, middle edge feature maps and deep edge feature maps; Step 4. Use the attention mechanism module to enhance the foreground feature map, shallow edge feature map, middle edge feature map, and deep edge feature map respectively; Step 5. Post-process and fuse the enhanced shallow edge feature map, the enhanced middle edge feature map, and the enhanced deep edge feature map to obtain a binary edge map; Step 6. Use the binary edge map and enhanced foreground feature map, combined with the labels in the training data set, to back-propagate and train the multi-level feature extraction module, foreground feature fusion module, edge feature fusion module, and attention mechanism module; Step 7. Extract the edge image of the roller can ear from the input image to obtain a visualization image of the edge of the roller can ear.
2. The method for extracting the edge image of the roller can ear based on multi-layer attention according to claim 1 is characterized in that: In step 1, the imported edge detection data set is the HED-BSDS edge detection data set, and the imported roller can ear original data is the roller can ear original image obtained from the acquisition device; the preprocessing process is specifically as follows: For all images imported from the edge detection dataset and the roller can ear original data, the image size and pixel value range are uniformly adjusted, and the enhanced training dataset is generated by random cropping, rotation, horizontal flipping and adding noise.
3. The method for extracting the edge image of the ear of a roller can based on multi-layer attention according to claim 1 is characterized in that: In step 2, the multi-level feature extraction module includes a shallow feature extraction module, a sub-shallow feature extraction module, a middle feature extraction module, a sub-deep feature extraction module and a deep feature extraction module; The input of the shallow feature extraction module is the image in the training dataset. The shallow feature extraction module is used to extract shallow features through convolution operations and ReLU activation functions. ; The input of the sub-shallow feature extraction module is the shallow feature. The sub-shallow feature extraction module is used to extract the sub-shallow features through convolution operation, maximum pooling operation and ReLU activation function. ; The input of the middle-level feature extraction module is the secondary shallow-level features. The middle-level feature extraction module is used to extract middle-level features through convolution operations, maximum pooling operations and ReLU activation functions. ; The input of the sub-deep feature extraction module is the mid-level features. The sub-deep feature extraction module is used to extract sub-deep features through convolution operations, maximum pooling operations and ReLU activation functions. ; The input of the deep feature extraction module is the sub-deep level features. The deep feature extraction module is used to extract deep features through convolution operations, maximum pooling operations and ReLU activation functions. .
4. The method for extracting the edge image of the ear of a roller can based on multi-layer attention according to claim 3 is characterized in that: In step 2, the specific process of fusing the multi-level features extracted from the image is as follows: , , , After bilinear interpolation upsampling, the feature maps are obtained respectively , , , ; Integrate features from different levels through skip connections , , , , , and obtain the fusion features.
5. The method for extracting the edge image of the ear of a roller can based on multi-layer attention according to claim 1 is characterized in that: The step 3 is specifically as follows: The foreground feature fusion module reduces the dimension of the fused features through the convolution layer, and then processes them through the ReLU activation function to generate a preliminary foreground feature map; The edge feature fusion module divides the fused features into shallow, middle and deep features according to different depth levels, and generates shallow edge feature maps, middle edge feature maps and deep edge feature maps respectively after being processed by convolutional layers and ReLU activation functions.
6. The method for extracting the edge image of the ear of a roller can based on multi-layer attention according to claim 1, characterized in that: The step 4 is specifically as follows: The foreground feature map, shallow edge feature map, middle edge feature map, and deep edge feature map are processed by global average pooling, multi-layer perceptron and The activation function calculates the channel weights and obtains the channel attention of the foreground feature map, the shallow edge feature map, the middle edge feature map, and the deep edge feature map; The foreground feature map, shallow edge feature map, middle edge feature map, and deep edge feature map are processed by global average pooling, two-dimensional convolution, and The activation function calculates the spatial weights to obtain the spatial attention of the foreground feature map, the shallow edge feature map, the middle edge feature map, and the deep edge feature map; The foreground feature map, shallow edge feature map, middle edge feature map, and deep edge feature map are respectively element-by-element multiplied with the corresponding channel attention and spatial attention to obtain an enhanced foreground feature map, an enhanced shallow edge feature map, an enhanced middle edge feature map, and an enhanced deep edge feature map.
7. The method for extracting the edge image of the ear of a roller can based on multi-layer attention according to claim 1, characterized in that: The step 5 is specifically as follows: The enhanced shallow edge feature map, the enhanced middle edge feature map, and the enhanced deep edge feature map are binarized using a dynamic threshold method; the enhanced shallow feature map, the enhanced middle feature map, and the enhanced deep feature map after the binarization process are fused element by element to obtain a binary edge map.
8. The method for extracting the edge image of the ear of a roller can based on multi-layer attention according to claim 1, characterized in that: In step 7, the trained multi-level feature extraction module, edge feature fusion module and attention mechanism module are used to extract the edge image of the roller can ear from the input image to obtain a visualized image of the edge of the roller can ear.
9. A roller can ear edge image extraction system based on multi-layer attention, comprising an image sensor, a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, the steps of the roller can ear edge image extraction method based on multi-layer attention as described in any one of claims 1 to 8 are implemented.
10. The roller can ear edge image extraction system based on multi-layer attention according to claim 9, characterized in that: The image sensor adopts a CMOS image sensor; The processor uses a data processor equipped with an RTX4060 graphics card.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer readable storage medium
CN111260666A
Automatic quantitative analysis method and system for lung digital pathological image
CN113222012A
MRI image hippocampus region segmentation method based on various losses and multi-scale features
CN113496496A
Remote sensing image road segmentation method fusing multi-scale features and double attention mechanism
CN117078943A
Camouflage target detection algorithm based on multi-scale cross-layer feature fusion network
CN117475134A