Image segmentation method suitable for underground coal mine railway track image
By combining the U-Net architecture and SAM encoding segmentation network in the coal mine underground railway track image segmentation, the problem of insufficient segmentation accuracy in the existing technology is solved, and a more efficient image segmentation effect is achieved.
Patent Information
- Application Number
- CN202510203465.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-02-24
AI Technical Summary
In the prior art, the U-Net architecture is difficult to effectively capture long-distance dependencies when segmenting railway tracks under coal mines, resulting in insufficient segmentation accuracy.
The main unit of the segmented network based on the U-Net architecture and the main unit of the SAM encoding are adopted to perform the first encoding process through the main encoding unit, the secondary unit of the segmented network is subjected to the second encoding process, and the results are transmitted to the main decoding unit for decoding process to generate a high-precision downhole track segmented image.
The segmentation accuracy of railway track images under coal mines has been improved, and the processing capability of complex images has been enhanced, especially under the requirements of real-time and continuous analysis.
Smart Images

Figure CN119919435A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an image segmentation method, in particular to an image segmentation method suitable for railway track images in underground coal mines. Background Art
[0002] Semantic segmentation aims to accurately classify each pixel in an image into different categories, thereby improving the accuracy and reliability of image analysis. In recent years, semantic segmentation has made significant progress through deep learning technology.
[0003] In order to meet the needs of underground coal mine work, railway tracks should generally be laid in coal mines. The railway tracks in coal mines are relatively complex, including many curved sections and intersections. Due to the particularity of underground coal mine work, the monitoring video data from the railway tracks in coal mines generally needs to be analyzed comprehensively, in real time, and continuously. When analyzing the monitoring video data, semantic segmentation is generally required for image processing. In order to meet the needs of monitoring video data analysis of underground coal mine railway tracks, the semantic segmentation method needs to have excellent real-time processing capabilities.
[0004] At present, semantic segmentation mostly uses the U-type architecture (U-Net architecture) based on convolutional neural network (CNN), that is, it mainly revolves around the U-Net architecture and the fully convolutional network (FCN), accompanied by a variety of subsequent adaptive improvements; specifically, the U-Net architecture has significant similarities with the feature pyramid network (FPN) framework. In the encoder of the U-Net architecture, the U-Net architecture adopts a divide-and-conquer strategy by decomposing the input image into feature maps of different scales, which are then passed to the decoder. At the same time, the decoder in the U-Net architecture exhibits an effective feature fusion mechanism, which integrates the multi-scale feature maps from the encoder into a unified single-scale feature map by performing step-by-step feature fusion at the corresponding resolution.
[0005] In the prior art, the U-Net architecture mainly relies on convolution and pooling operations when constructing the encoder and decoder architectures. This dependence limits its ability to effectively capture long-distance dependencies, thereby limiting the image segmentation accuracy of railway track images in underground coal mines. Summary of the invention
[0006] The purpose of the present invention is to overcome the deficiencies in the prior art and to provide an image segmentation method suitable for underground railway track images in coal mines, which can effectively segment underground railway track images in coal mines and improve the segmentation accuracy of underground railway track images in coal mines.
[0007] According to the technical solution provided by the present invention, an image segmentation method suitable for railway track images in coal mines, the image segmentation method comprises:
[0008] The downhole track target image to be segmented is obtained, and the acquired downhole track target image is loaded into the constructed image segmentation network, so as to use the image segmentation network to perform image segmentation on the downhole track target image, and generate a downhole track segmentation image after image segmentation, wherein the downhole track segmentation image includes a railway track segmentation marked area and a non-railway track segmentation marked area, wherein:
[0009] The image segmentation network includes a segmentation network main unit based on a U-Net architecture and a segmentation network auxiliary unit based on SAM coding, wherein the segmentation network main unit includes a main coding unit and a main decoding unit adaptively connected to the main coding unit, and the segmentation network auxiliary unit is adaptively connected to the main decoding unit;
[0010] When performing image segmentation on the downhole track target image, the main coding unit in the segmentation network main unit is used to perform a first coding process to generate a main coding process feature map after the first coding process, and at the same time, the segmentation network auxiliary unit is used to perform a second coding process on the downhole track target image to generate an auxiliary coding embedding feature map after the second coding;
[0011] The main coding processing feature map and the auxiliary coding embedding feature map are transmitted to the main decoding unit for decoding processing by the main decoding unit, and a downhole track segmentation image is generated after decoding processing, wherein:
[0012] When the main decoding unit performs decoding processing, it includes at least a first decoding processing and a plurality of second decoding processings performed in sequence;
[0013] When performing the first decoding process, the auxiliary coding embedding feature map is subjected to feature splicing processing with the main coding deconvolution processing feature map and the main coding attention processing feature map generated based on the main coding processing feature map, so as to generate a decoding reference feature map after feature splicing processing, wherein,
[0014] Performing deconvolution processing on the main coding processing feature map, and generating a main coding deconvolution processing feature map after the deconvolution processing;
[0015] Performing attention mechanism processing on the main coding processing feature map, and generating the main coding attention processing feature map after the attention mechanism processing;
[0016] The generated decoded reference feature map is subjected to a second decoding process, and a downhole track segmentation image is generated after the second decoding process.
[0017] The main coding unit includes a plurality of main coding network layers connected in sequence, wherein, in the main coding unit, the main coding network layer at the bottom of the U-Net architecture is configured as a main coding transition connection layer, and the remaining main coding network layers are respectively configured as main coding processing layers;
[0018] The main decoding unit comprises a plurality of main decoding processing layers connected in sequence, wherein:
[0019] In the main unit of the segmentation network, the main encoding processing layer in the main encoding unit corresponds one-to-one to the main decoding processing layer in the main decoding unit, and the main encoding processing layer is jump-connected to the corresponding main decoding processing layer through the channel-spatial attention module;
[0020] The main coding transition connection layer is adaptively connected to the main decoding processing layer at the bottom of the U-Net architecture in the main decoding unit, and transmits the generated main coding processing feature map to the corresponding main decoding processing layer through the main coding transition connection layer, and the main decoding processing layer at the bottom of the U-Net architecture also receives the auxiliary coding embedded feature map;
[0021] When the main decoding unit performs decoding processing, the main decoding processing layer at the bottom of the U-Net architecture is used to perform the first decoding processing, and the remaining main decoding processing layers are configured to perform the second decoding processing.
[0022] The main coding network layer includes a main coding convolution block and a main coding Ghost module connected in sequence, wherein:
[0023] When the main coding unit performs the first coding process, the coding feature extraction is performed in sequence through the main coding network layer, so as to use each main coding network layer to perform the first coding sub-processing, wherein, when extracting the coding feature, the main coding convolution block is firstly used to perform the coding convolution process, and after the coding convolution process, the main coding Ghost module is used to perform feature extraction, and the corresponding network layer feature map is generated after the feature extraction of the main coding Ghost module;
[0024] In the main coding unit, for two adjacent main coding network layers, along the direction of the opening of the U-Net architecture pointing to the bottom, the network layer feature map generated by the upper main coding network layer is subjected to maximum pooling processing to generate a maximum pooled feature map after maximum pooling processing, and the maximum pooled feature map is loaded into the main network coding layer below.
[0025] When the first decoding process is performed using the main decoding processing layer at the bottom of the U-Net architecture, there are:
[0026] The main decoding processing layer receives the main coding deconvolution processing feature map, and loads the received main coding sampling feature map into the correspondingly connected channel-spatial attention module, and the channel-spatial attention module also receives the network layer feature map output by the correspondingly connected main coding network layer;
[0027] Based on the received main encoding deconvolution processing feature map and the network layer feature map, the channel-spatial attention module performs attention mechanism processing to generate a main encoding attention processing feature map after the attention mechanism processing;
[0028] The auxiliary coding embedding feature map, the main coding deconvolution processing feature map, and the main coding attention processing feature map are feature spliced, and after the feature splicing processing, they are processed by Ghost to generate a decoding benchmark feature map.
[0029] For the main decoding processing layer that performs the second decoding process, the main decoding processing layer includes a decoding deconvolution block and a main decoding Ghost module connected in sequence, wherein:
[0030] When performing the second decoding process, a basic decoding feature map to be decoded is received, and thereafter, a decoding deconvolution block is used to perform a deconvolution process on the basic decoding feature map to generate a deconvolution-post decoding feature map after the deconvolution process;
[0031] Loading the deconvolution decoded feature map into the channel-spatial attention module of the current main decoding processing layer, and the channel-spatial attention module also receives the network layer feature map corresponding to the output of the main encoding network layer;
[0032] Based on the received post-deconvolution decoding feature map and the network layer feature map, the channel-spatial attention module performs attention mechanism processing to generate a decoding attention processing feature map after the attention mechanism processing;
[0033] The main decoding Ghost module is used to extract features from the decoding attention processing feature map to generate a basic decoding feature map after feature extraction;
[0034] In the main decoding unit, for two adjacent main decoding processing layers, along the bottom of the U-Net architecture pointing to the direction of the opening, the basic decoding feature map is generated by the main decoding processing layer below and transmitted to the main decoding processing layer above.
[0035] The channel-spatial attention module includes an attention first stitcher, a channel attention module, a spatial attention module, and an attention linear layer, wherein:
[0036] The attention first splicer is connected to the corresponding main encoding processing layer and the main decoding processing layer;
[0037] The channel attention module is connected to the output end of the first attention splicer, and the channel attention module adopts a residual connection;
[0038] The spatial attention module is connected to the output end of the channel attention module, and the spatial attention module adopts residual connection;
[0039] The output end of the spatial attention module is connected to the attention linear layer, and the attention linear layer is used as the output layer of the channel-spatial attention module.
[0040] The channel attention module includes a channel attention first maximum pooling module and a channel attention first average pooling module, wherein:
[0041] The channel attention first maximum pooling module and the channel attention first average pooling module are both connected to the output end of the attention first splicer;
[0042] The output end of the first maximum pooling module of the channel attention and the output end of the first average pooling module of the channel attention are connected to the first linear module of the channel attention, and the first linear module of the channel attention is connected to the second linear module of the channel attention through the LeakyReLu activation function of the channel attention;
[0043] The output end of the second linear module of the channel attention is connected to the second maximum pooling module of the channel attention and the second average pooling module of the channel attention, and the second maximum pooling module of the channel attention and the second average pooling module of the channel attention are both connected to the channel attention splicer;
[0044] The output end of the channel attention splicer is connected to the channel attention adder used to form the residual connection of the channel attention module through the channel attention sigmoid function, and is adaptively connected to the spatial attention module through the channel attention adder.
[0045] The spatial attention module includes a spatial attention first convolution block, a normalized activation function module, a spatial attention second convolution block, a normalization module and a spatial attention Sigmoid function connected in sequence, wherein:
[0046] The first spatial attention convolution block is connected to the output of the channel attention adder;
[0047] The convolution kernel sizes used in the first spatial attention convolution block and the second spatial attention convolution block are the same;
[0048] When the spatial attention module adopts residual connection, the spatial attention sigmoid function is connected to the input of the spatial attention adder, and the input of the first convolutional block of the spatial attention is also connected to the input of the spatial attention adder, and the output of the spatial attention adder is connected to the attention linear layer.
[0049] The segmentation network auxiliary unit includes a SAM encoding unit and a size adjustment unit adaptively connected to the SAM encoding unit, wherein:
[0050] The downhole track target image is coded by using the SAM coding unit, and a basic coding embedding feature map is generated after the coding process;
[0051] The generated basic coding embedding feature map is resized by using a resizing unit to generate an auxiliary coding embedding feature map after the feature map resizing, wherein the feature map size of the auxiliary coding embedding feature map is consistent with the corresponding feature map sizes of the main coding deconvolution processing feature map and the main coding attention processing feature map.
[0052] When acquiring downhole track target images, it includes:
[0053] A downhole track source image is acquired, and the acquired downhole track source image is preprocessed to generate a downhole track target image after the preprocessing, wherein:
[0054] The preprocessing of the downhole track source image includes bilateral filtering and / or contrast limiting adaptive histogram equalization enhancement processing;
[0055] When performing contrast-limited adaptive histogram equalization enhancement processing, including
[0056] Obtain a source image to be processed by equalization and enhancement, and divide the source image to be processed by equalization and enhancement into a plurality of source image sub-blocks. For any source image sub-block, there is:
[0057]
[0058] Wherein, (b1, b2) is the size of each source image sub-block, (H, W) is the size of the source image to be processed by equalization and enhancement, K is the number of sub-blocks of the divided source image, and the source image for equalization and enhancement processing is the downhole track source image or the downhole track filtered image generated by bilateral filtering;
[0059] Contrast limiting processing is performed on each source image sub-block, wherein the contrast limiting processing includes:
[0060] For the source image sub-block T, a grayscale histogram of the source image sub-block T is generated, and the frequency of the i-th grayscale level is limited based on the generated grayscale histogram, then:
[0061]
[0062] Where DB is the frequency limit threshold, H T (i) is the frequency of the source image sub-block T at the i-th grayscale level, and L is the number of grayscale levels of the source image to be processed by equalization enhancement;
[0063] After frequency limitation, the mirror grayscale mapping process is performed, and then:
[0064]
[0065] in, is the grayscale mapping of the i-th grayscale level; is the cumulative distribution function of the i-th gray level;
[0066] After the mirror grayscale mapping process, a mapped grayscale histogram is generated based on the grayscale mapping information, and a corresponding contrast-limited sub-block is generated based on the mapped grayscale histogram;
[0067] Any two contrast-limited sub-blocks are merged by linear interpolation to generate a downhole track target image after merging.
[0068] Advantages of the present invention: According to the characteristics of the downhole track target image, the constructed image segmentation network includes a segmentation network main unit and a segmentation network auxiliary unit. The first encoding process is performed by the main encoding unit in the segmentation network main unit, and the second encoding process is performed by the segmentation network auxiliary unit. Thereafter, the main encoding process feature map generated by the first encoding process and the auxiliary encoding embedding feature map generated by the second encoding process are sent to the main decoding unit, so that the downhole track segmentation image can be generated after the decoding process, so as to improve the accuracy of image segmentation of the downhole track target.
[0069] Using the Ghost module to perform feature extraction in the main unit of the segmentation network can reduce the overhead of the feature extraction operation performed by the main unit of the segmentation network during image segmentation, and can reduce the impact of the second encoding processing performed by the auxiliary unit of the segmentation network, thereby achieving the purpose of improving the efficiency of image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 The present invention is a flow chart of an embodiment of performing image segmentation on railway track images in underground coal mines.
[0071] Figure 2 This is an embodiment of the present invention for obtaining downhole trajectory source images.
[0072] Figure 3 This is a structural block diagram of an embodiment of the image segmentation network of the present invention.
[0073] Figure 4 This is a structural block diagram of an embodiment of the channel-spatial attention module of the present invention. DETAILED DESCRIPTION
[0074] The present invention will be further described below in conjunction with specific drawings and embodiments.
[0075] In order to effectively segment the railway track image in the coal mine and improve the segmentation accuracy of the railway track image in the coal mine, the present invention provides an image segmentation method suitable for the railway track image in the coal mine. Specifically, the image segmentation method includes:
[0076] The downhole track target image to be segmented is obtained, and the acquired downhole track target image is loaded into the constructed image segmentation network, so as to use the image segmentation network to perform image segmentation on the downhole track target image, and generate a downhole track segmentation image after image segmentation, wherein the downhole track segmentation image includes a railway track segmentation marked area and a non-railway track segmentation marked area, wherein:
[0077] The image segmentation network includes a segmentation network main unit based on a U-Net architecture and a segmentation network auxiliary unit based on SAM coding, wherein the segmentation network main unit includes a main coding unit and a main decoding unit adaptively connected to the main coding unit, and the segmentation network auxiliary unit is adaptively connected to the main decoding unit;
[0078] When performing image segmentation on the downhole track target image, the main coding unit in the segmentation network main unit is used to perform a first coding process to generate a main coding process feature map after the first coding process, and at the same time, the segmentation network auxiliary unit is used to perform a second coding process on the downhole track target image to generate an auxiliary coding embedding feature map after the second coding;
[0079] The main coding processing feature map and the auxiliary coding embedding feature map are transmitted to the main decoding unit for decoding processing by the main decoding unit, and a downhole track segmentation image is generated after the decoding processing, wherein:
[0080] When the main decoding unit performs decoding processing, it includes at least a first decoding processing and a plurality of second decoding processings performed in sequence;
[0081] When performing the first decoding process, the auxiliary coding embedding feature map is subjected to feature splicing processing with the main coding deconvolution processing feature map and the main coding attention processing feature map generated based on the main coding processing feature map, so as to generate a decoding reference feature map after feature splicing processing, wherein,
[0082] Performing deconvolution processing on the main coding processing feature map, and generating a main coding deconvolution processing feature map after the deconvolution processing;
[0083] Performing attention mechanism processing on the main coding processing feature map, and generating the main coding attention processing feature map after the attention mechanism processing;
[0084] The generated decoded reference feature map is subjected to a second decoding process, and a downhole track segmentation image is generated after the second decoding process.
[0085] Figure 1 , a flowchart of an embodiment of image segmentation of the present invention is shown in FIG. As can be seen from the figure, when performing image segmentation, the downhole track target image to be segmented should be obtained, and then the acquired downhole track target image should be loaded into the image segmentation network to perform image segmentation using the image segmentation network, and generate a downhole track segmentation image after image segmentation. It should be understood that the image segmentation network should be constructed in advance, and the constructed image segmentation network, the method and process of generating the downhole track segmentation image using the image segmentation network will be described in detail below.
[0086] Figure 2 A schematic diagram of a working scenario for acquiring underground track target images of the present invention is shown in the figure. It can be seen from the figure that an underground railway will be laid in the coal mine, and a signal receiver and an underground monitoring camera can be set on the side of the underground railway. The signal receiver and the underground monitoring camera can be set in an equal spacing manner. In the figure, the underground monitoring camera can be located above the underground railway of the coal mine through a camera mounting frame.
[0087] In specific implementation, a vehicle-mounted camera can also be installed on a transport cart running on the underground railway of a coal mine. After the vehicle-mounted camera is installed on the transport cart, the vehicle-mounted camera can be used to obtain images in the direction of the transport cart's advance. It can be seen that the acquired images should include the underground railway of the coal mine and the area outside the underground railway of the coal mine. The underground railway of the coal mine specifically refers to the area between two tracks and the two tracks, and the area outside the underground railway of the coal mine specifically refers to the area outside the two tracks of the underground railway of the coal mine.
[0088] In one embodiment of the present invention, an underground track target image can be generated based on an image taken by a vehicle-mounted camera, and therefore, the underground track target image should include the coal mine underground railway and the area outside the coal mine underground railway. It should be noted that when using an image segmentation network for image segmentation, the coal mine underground railway and the area outside the coal mine underground railway contained in the underground track target image are mainly segmented and labeled, and therefore, the underground track segmentation image includes a railway track segmentation and labeling area and a non-railway track segmentation and labeling area, wherein the railway track segmentation and labeling area specifically refers to the area labeled after the coal mine underground railway is segmented, and the non-railway track segmentation standard area specifically refers to the labeled area outside the two tracks.
[0089] It is known to those skilled in the art that, due to the low illumination and high dust in the coal mine environment, it is difficult to achieve an ideal image segmentation effect if the image captured by the vehicle-mounted camera is directly segmented. Therefore, the acquired underground track target image should generally be generated after processing. In one embodiment of the present invention, an underground track source image is acquired, and the acquired underground track source image is preprocessed to generate an underground track target image after preprocessing. It can be seen from the above description that the underground track source image is an image directly captured and generated by the vehicle-mounted camera, and the underground track source image is generally an RGB image.
[0090] In specific implementation, the preprocessing of the downhole track source image includes bilateral filtering and / or contrast-limited adaptive histogram equalization enhancement processing. Therefore, when preprocessing the downhole track source image, bilateral filtering and contrast-limited adaptive histogram equalization enhancement processing can be performed, or bilateral filtering and contrast-limited adaptive histogram equalization enhancement processing can be performed simultaneously. It should be noted that when bilateral filtering and contrast-limited adaptive histogram equalization enhancement processing are performed simultaneously, bilateral filtering can be performed first, and then contrast-limited adaptive histogram equalization enhancement processing can be performed; of course, contrast-limited adaptive histogram equalization enhancement processing can also be performed first, and then bilateral filtering can be performed, and the specific selection can be made according to needs.
[0091] In one embodiment of the present invention, when bilateral filtering is performed, there are:
[0092]
[0093] Among them, I filered (x, y) is the pixel value after filtering, I(x, y) is the pixel value in the original image, Ω is the neighborhood window during bilateral filtering, and W p (x, y) is the normalized weight of the pixel value in the neighborhood window Ω, F spatial (i, j) is the Gaussian kernel function in the spatial domain, F rang (I(x,y),I(x+i,y+j)) is the Gaussian kernel function in pixel domain;
[0094] When configuring bilateral filtering parameters based on mine environment parameters, it at least includes configuring the size of the neighborhood window Ω, the spatial domain Gaussian kernel function F based on the mine environment parameters. spatial (i, j), pixel domain Gaussian kernel function F rang (I(x,y),I(x+i,y+j)) is the width of the corresponding kernel function.
[0095] It is understandable that the original image can be an underground track source image or an image enhanced by contrast-limited adaptive histogram equalization. An example is given for configuring bilateral filtering parameters based on mine environment parameters. For example, if the mine environment brightness is greater than >100 (cd / m 2 ) and the mine dust concentration is less than 1000 (mg / m 3 ), the size of the neighborhood window Ω can be 5, and the spatial domain Gaussian kernel function F spatial The width of the (i, j) kernel function can be 3, and the pixel domain Gaussian kernel function F rang The width of the kernel function corresponding to (I(x,y),I(x+i,y+j)) can be 3; in other cases, the size of the neighborhood window Ω can be 5, and the spatial domain Gaussian kernel function F spatial The width of the (i, j) kernel function can be 9, and the pixel domain Gaussian kernel function F rang The width of the kernel function corresponding to (I(x,y),I(x+i,y+j)) can be 9. It can be understood that for the spatial domain Gaussian kernel function, once the width of the kernel function is determined, those skilled in the art can determine to use the spatial domain Gaussian kernel function for corresponding filtering processing.
[0096] From the above description, it can be seen that the size of the neighborhood window Ω can be configured based on the mine environment parameters. After that, during filtering, the normalized weight W can be calculated according to the pixel value of the current neighborhood window Ω. p (x, y). After that, the low-light source image can be subjected to bilateral filtering based on the above bilateral filtering processing formula. Of course, other methods can also be used to perform bilateral filtering on the low-light source image. The specific bilateral filtering processing method can be selected according to actual needs, and will not be listed here one by one.
[0097] In one embodiment of the present invention, when performing contrast limiting adaptive histogram equalization enhancement processing, it includes:
[0098] Obtain a source image to be processed by equalization and enhancement, and divide the source image to be processed by equalization and enhancement into a plurality of source image sub-blocks. For any source image sub-block, there is:
[0099]
[0100] Wherein, (b1, b2) is the size of each source image sub-block, (H, W) is the size of the source image to be processed by equalization and enhancement, K is the number of sub-blocks of the divided source image, and the source image for equalization and enhancement processing is the downhole track source image or the downhole track filtered image generated by bilateral filtering;
[0101] Contrast limiting processing is performed on each source image sub-block, wherein the contrast limiting processing includes:
[0102] For the source image sub-block T, a grayscale histogram of the source image sub-block T is generated, and the frequency of the i-th grayscale level is limited based on the generated grayscale histogram, then:
[0103]
[0104] Where DB is the frequency limit threshold, H T (i) is the frequency of the source image sub-block T at the i-th grayscale level, and L is the number of grayscale levels of the source image to be processed by equalization enhancement;
[0105] After frequency limitation, the mirror grayscale mapping process is performed, and then:
[0106]
[0107] in, is the grayscale mapping of the i-th grayscale level; is the cumulative distribution function of the i-th gray level;
[0108] After the mirror grayscale mapping process, a mapped grayscale histogram is generated based on the grayscale mapping information, and a corresponding contrast-limited sub-block is generated based on the mapped grayscale histogram;
[0109] Any two contrast-limited sub-blocks are merged by linear interpolation to generate a downhole track target image after merging.
[0110] In specific implementation, after the source image to be processed for equalization enhancement is divided into corresponding source image sub-blocks, contrast limitation processing should be performed on each source image sub-block. When performing contrast limitation processing, the grayscale histogram of each source image sub-block is first generated by using existing commonly used technical means. In the generated grayscale histogram, the horizontal axis is the grayscale level, and the vertical axis is the ratio of the number of pixels contained in any grayscale level to the number of pixels in the current source image sub-block. Therefore, for any source image sub-block, according to the generated grayscale histogram, the frequency H of the i-th grayscale level can be obtained. T (i) If the inverse of the ratio of the i-th gray level in the gray level histogram is taken, the corresponding frequency can be obtained. The value of the gray level number L of the source image to be processed by equalization enhancement can be 256.
[0111] After determining the frequency of each gray level, the frequency limit threshold DB can be used to limit the contrast. It can be seen that when performing contrast limitation, the frequency limit threshold DB should be configured. It should be noted that the frequency limit threshold DB should be related to the acquisition conditions of the source image to be processed by equalization and enhancement in the coal mine, so as to improve the accuracy of generating the underground track target image. In specific implementation, if the mine environment brightness is greater than> 100 (cd / m 2 ) and the mine dust concentration is less than 1000 (mg / m 3), the number of divisions K can be 32, and the frequency limit threshold DB can be 10; in other cases, the number of divisions K can be 64, and the frequency limit threshold DB can be 30.
[0112] After limiting the frequency of each gray level in the gray level histogram, a mirror gray level mapping process is required to determine the gray level mapping of each gray level in the mirror gray level mapping process, so that the gray level mapping information can be obtained based on all gray level mappings, and then a mapping gray level histogram can be generated according to the gray level mapping information. According to the corresponding relationship between the gray level histogram and the image, a contrast-limited sub-block can be generated, that is, for each source image sub-block, after the above-mentioned contrast limiting process, a corresponding contrast-limited sub-block can be generated.
[0113] For any source image sub-block, after the above-mentioned contrast limitation is used to generate the corresponding contrast-limited sub-block, any two contrast-limited sub-blocks can be merged by bilinear interpolation to generate a downhole track target image after merging. Specifically, when bilinear interpolation is used for merging, the contrast-limited sub-blocks in the same arrangement direction can be merged. The use of bilinear interpolation is consistent with the prior art, and the specific merging that can achieve linear interpolation shall prevail.
[0114] In order to reduce the impact of the underground environment of the coal mine, in the specific implementation, during preprocessing, it is preferred to perform bilateral filtering and contrast-limited adaptive histogram equalization enhancement processing at the same time. Of course, other preprocessing methods can also be used. The form of preprocessing can be selected according to needs, so as to effectively improve the segmentation accuracy of the underground track target image. They will not be listed here one by one.
[0115] In order to improve the segmentation accuracy, in one embodiment of the present invention, the image segmentation network includes a segmentation network main unit based on the U-Net architecture and a segmentation network auxiliary unit based on SAM coding. Specifically, the segmentation network main unit based on the U-Net architecture can be used to implement image segmentation, and the segmentation network auxiliary unit based on SAM coding can effectively improve the segmentation accuracy. For the segmentation network main unit based on the U-Net architecture, the segmentation network main unit may include a main encoding unit and a main decoding unit. The main encoding unit and the main decoding unit form a U-shaped network, that is, the main encoding unit can form an encoder of the U-shaped network, and the main decoding unit forms a decoder of the U-shaped network. Therefore, the role of the main encoding unit and the main decoding unit in image segmentation can be consistent with the prior art.
[0116] In a specific implementation, when performing image segmentation on a downhole track target image, a main encoding unit in a segmentation network main unit is used to perform a first encoding process to generate a main encoding process feature map after the first encoding process. At the same time, a segmentation network auxiliary unit is used to perform a second encoding process on the downhole track target image to generate an auxiliary encoding embedding feature map after the second encoding. Thereafter, the main encoding process feature map and the auxiliary encoding embedding feature map are transmitted to a main decoding unit to perform decoding process by the main decoding unit, and a downhole track segmentation image is generated after the decoding process.
[0117] It should be understood that when the main decoding unit is used for decoding processing, since the main coding processing feature map generated by the main coding unit and the auxiliary coding embedding feature map generated by the segmentation network auxiliary unit are used at the same time, it can be seen that the downhole track segmentation image generated by the decoding processing has a higher segmentation accuracy.
[0118] In one embodiment of the present invention, the main coding unit includes a plurality of main coding network layers connected in sequence, wherein, in the main coding unit, the main coding network layer at the bottom layer of the U-Net architecture is configured as a main coding transition connection layer, and the remaining main coding network layers are respectively configured as main coding processing layers;
[0119] The main decoding unit comprises a plurality of main decoding processing layers connected in sequence, wherein:
[0120] In the main unit of the segmentation network, the main encoding processing layer in the main encoding unit corresponds one-to-one to the main decoding processing layer in the main decoding unit, and the main encoding processing layer is jump-connected to the corresponding main decoding processing layer through the channel-spatial attention module;
[0121] The main coding transition connection layer is adaptively connected to the main decoding processing layer at the bottom of the U-Net architecture in the main decoding unit, and transmits the generated main coding processing feature map to the corresponding main decoding processing layer through the main coding transition connection layer, and the main decoding processing layer at the bottom of the U-Net architecture also receives the auxiliary coding embedded feature map;
[0122] When the main decoding unit performs decoding processing, the main decoding processing layer at the bottom of the U-Net architecture is used to perform the first decoding processing, and the remaining main decoding processing layers are configured to perform the second decoding processing.
[0123] In order to realize the first encoding process, the main encoding unit may include multiple main encoding network layers, and the multiple main encoding network layers are connected in sequence to form a U-Net architecture. Then, there is: there is a main encoding network layer located at the opening of the U-Net architecture, and there is a main encoding network layer located at the bottom of the U-Net architecture, and the remaining main encoding network layers are located between the opening and the bottom of the U-Net architecture. In specific implementation, the main encoding network layer at the bottom of the U-Net architecture is used as the main encoding transition connection layer, and the remaining main encoding network layers are used as main encoding processing layers, that is, except for the main encoding network layer used as the main encoding transition connection layer, among the remaining main encoding network layers, one main encoding network layer forms a main encoding processing layer.
[0124] In specific implementation, the main decoding unit should include multiple main decoding processing layers connected in sequence, wherein one main decoding processing layer should correspond to one encoding processing layer, and each main encoding processing layer is jump-connected to the corresponding main decoding processing layer through a channel-spatial attention module. It can be seen that the number of main decoding processing layers in the main decoding unit is less than the number of main encoding network layers in the main encoding unit. Therefore, for the constructed segmentation network main unit, the main decoding processing layer at the bottom of the U-Net architecture is connected to a corresponding main encoding processing layer, and is also adaptively connected to the main encoding transition connection layer, wherein the main encoding processing feature map generated by the first encoding process can be output through the main encoding transition connection layer.
[0125] In addition, the auxiliary coding embedded feature map generated by the segmentation network auxiliary unit should also be transmitted to the main decoding processing layer at the bottom of the U-Net architecture, so as to use the main decoding processing layer at the bottom of the U-Net architecture to perform the first decoding process. Therefore, the segmentation network auxiliary unit is adaptively connected to the main decoding unit, which specifically refers to the segmentation network auxiliary unit being connected to the main decoding processing layer at the bottom of the U-Net architecture in the main decoding unit. It should be noted that for the main decoding processing layer in the main decoding unit, except for the main decoding processing layer at the bottom of the U-Net architecture that performs the first decoding process, the remaining main decoding processing layers should be configured to perform the second decoding process, wherein when performing the second decoding process, each main decoding processing layer performs the same second decoding sub-process.
[0126] Depend on Figure 3As can be seen from the U-Net architecture in the prior art, the main encoding transition connection layer should be located at the bottom of the U-Net architecture. At this time, the main decoding processing layer connected to the main encoding transition connection layer is also located at the bottom of the U-Net architecture. When performing the second decoding processing, along the bottom of the U-Net architecture pointing to the opening direction of the U-Net architecture, multiple main decoding processing layers sequentially execute the corresponding second decoding sub-processing. For example, the main decoding processing layer closely connected to the main decoding processing layer at the bottom of the U-Net architecture executes the second decoding sub-processing first, and the main decoding processing layer at the opening of the U-Net architecture executes the second decoding sub-processing last, and the downhole track segmentation image can be generated after executing the corresponding second decoding sub-processing.
[0127] In one embodiment of the present invention, the main coding network layer includes a main coding convolution block and a main coding Ghost module connected in sequence, wherein:
[0128] When the main coding unit performs the first coding process, the coding feature extraction is performed in sequence through the main coding network layer, so as to use each main coding network layer to perform the first coding sub-processing, wherein, when extracting the coding feature, the main coding convolution block is firstly used to perform the coding convolution process, and after the coding convolution process, the main coding Ghost module is used to perform feature extraction, and the corresponding network layer feature map is generated after the feature extraction of the main coding Ghost module;
[0129] In the main coding unit, for two adjacent main coding network layers, along the direction of the opening of the U-Net architecture pointing to the bottom, the network layer feature map generated by the upper main coding network layer is subjected to maximum pooling processing to generate a maximum pooled feature map after maximum pooling processing, and the maximum pooled feature map is loaded into the main network coding layer below.
[0130] In specific implementation, the main coding network layers in the main coding unit can adopt the same structural form. For example, each main coding network layer adopts the form of a main coding convolution block and a main coding Ghost module. When the main coding unit performs the first coding process, along the direction of the opening of the U-Net architecture pointing to the bottom, each main coding network layer performs the first coding sub-processing in turn, that is, when performing the first coding process, all the main coding network layers perform the first coding sub-processing once, so that when all the main coding network layers perform the corresponding first coding sub-processing, it is deemed that the first coding process is completed.
[0131] For any main coding network layer, when the first coding sub-processing is performed once, the main coding convolution block is first used to perform coding convolution processing, and then the main coding Ghost module is used to extract features, and the corresponding network layer feature map can be generated after the main coding Ghost module performs feature extraction. If the current main coding network layer is not used as the main coding transition connection layer, the generated network layer feature map should be transmitted to the adjacent main coding network layer, and at the same time, the network layer feature map should be transmitted to the corresponding connected channel-spatial attention module. It should be noted that when the network layer feature map is transmitted to the adjacent main coding network layer, the transmission direction is the sequential direction of the main coding network layer performing the first coding sub-processing, wherein the transmission direction is the direction along the opening of the U-Net architecture pointing to the bottom, that is, the direction pointing to the main coding transition connection layer, and finally the main coding processing feature map is output by the main coding transition connection layer.
[0132] Figure 3 An embodiment of the main coding unit of the present invention is shown in the figure, which shows an embodiment in which the main coding unit includes five main coding network layers. It can be seen from the above description that one main coding network layer is configured as a main coding transition connection layer, and the remaining four main coding network layers are configured as main coding processing layers. At this time, the main decoding unit should include four main decoding processing layers. Specifically, the four main decoding processing layers correspond to the four main coding processing layers one by one, and each main coding processing layer is jump-connected with the corresponding main decoding processing layer through a channel-spatial attention module.
[0133] Figure 3 In , GXT is the downhole track target image. After obtaining the downhole track target image, the main coding network layer that performs the first coding sub-processing on the downhole track target image is configured as the fifth main coding network layer. At this time, the fifth main coding network layer should be used as the main coding processing layer, and the fifth main coding network layer is located at the opening of the U-Net architecture, such as Figure 3 shown. Figure 3 In the figure, CR4 is the main coding convolution block in the fifth main coding network layer, and BGh4 is the main coding Ghost module in the fifth main coding network layer. In the specific implementation, the main coding convolution block may include a coding convolution unit and a coding activation function. Among them, in the main coding convolution block CR4, the number of convolution kernels of the coding convolution unit may be 64, the size of each convolution kernel is 3*3, the coding activation function unit may adopt a ReLU activation function, and the main coding Ghost module may adopt a Ghost module commonly used in the prior art. The Ghost module adopted may refer to the corresponding description in GhostNet: More Features From Cheap Operations.
[0134] When the fifth main coding network layer adopts the above-mentioned parameter configuration, when the downhole track target image is subjected to the first coding sub-processing, the main coding convolution block CR4 is first used for coding convolution processing to generate a feature map E4-0. At this time, the number of channels of the feature map E4-0 is 64 dimensions, and the feature size of the feature map E4-0 is 1 / 2 of the corresponding feature size of the downhole track target image; the feature map E4-0 is subjected to feature extraction by the main coding Ghost module, and after the feature extraction, a network layer feature map E4 can be generated, wherein the feature size of the network layer feature map E4 is consistent with the feature size of the feature map E4-0, and the number of channels of the network layer feature map E4 is also 64 dimensions. It can be understood that when the main coding Ghost module is used to extract features from the feature map E4-0, the efficiency of the main coding network layer in feature extraction can be effectively improved.
[0135] Further, Figure 3 An embodiment in which the main coding network layer adopts the same structure is shown in FIG. Figure 3 In the figure, CR3 is the main coding convolution block in the fourth main coding network layer, and BGh3 is the main coding Ghost module in the fourth main coding network layer. As can be seen from the figure, the fourth main coding network layer is adjacent to the fifth main coding network layer. After the fifth main coding network layer performs the first coding sub-processing on the downhole track target image, the fourth main coding network layer then performs the first coding sub-processing. As can be seen from the above description, after the fifth main coding network layer performs the first coding sub-processing, the generated network layer feature map E4 should be transmitted to the fourth main coding network layer, and after the maximum pooling processing, the maximum pooling processed feature map is generated. Figure 3 In FIG. 1 , MP4 is a maximum pooling processing module (Maxpool) for performing maximum pooling processing on the network layer feature map E4, and E3-0 is a maximum pooling processing feature map generated after performing maximum pooling processing on the network layer feature map E4.
[0136] In specific implementation, the maximum pooling processing module MP4 can perform 2*2 maximum pooling processing, that is, the pooling window size of the maximum pooling processing module MP4 when performing maximum pooling processing is 2*2. Therefore, the number of channels for generating the maximum pooling processing feature map E3-0 is 64 dimensions, and the feature size of the maximum pooling processing feature map E3-0 should be 1 / 4 of the corresponding feature size of the downhole track target image.
[0137] The corresponding parameter configurations of the main coding convolution block CR3 and the main coding Ghost module BGh3 can refer to the corresponding descriptions of the main coding convolution block CR4 and the main coding Ghost module BGh4 respectively. Of course, when configuring the parameters, the corresponding number of channels and feature sizes in the following description should be obtained. Specifically, when the fourth main coding network layer performs the first coding sub-processing, the main coding convolution block CR3 is used to perform coding convolution processing and generate a feature map E3-1. At this time, the number of channels of the feature map E3-1 is 128 dimensions, and the feature size of the feature map E3-1 is consistent with the feature size of the maximum pooling processing feature map E3-0. Feature map E3-1 is feature extracted by the main coding Ghost module BGh3 to generate a network layer feature map E3. The number of channels of the network layer feature map E3 is 128 dimensions, and the feature size of the network layer feature map E3 is consistent with the feature size of the feature map E3-1.
[0138] Referring to the above description, we can get Figure 3 In the figure, MP3 is a maximum pooling processing module that performs maximum pooling processing on the network layer feature map E3. After the maximum pooling processing module MP3 performs maximum pooling processing on the network layer feature map E3, a maximum pooling processing feature map E2-0 can be generated, wherein the number of channels of the maximum pooling processing feature map E2-0 is 128 dimensions, and the feature size of the maximum pooling processing feature map E2-0 should be 1 / 8 of the corresponding feature size of the downhole track target image.
[0139] Figure 3 In the figure, CR2 is the main coding convolution block in the third main coding network layer, and BGh2 is the main coding Ghost module in the third main coding network layer. When the third main coding network layer performs the first coding sub-processing, the maximum pooling processing feature map E2-0 is first processed by the main coding convolution block CR2, and the feature map E2-1 is generated. At this time, the number of channels of the feature map E2-1 is 256 dimensions, and the feature size of the feature map E2-1 is consistent with the feature size of the maximum pooling processing feature map E2-0. The feature map E2-1 is extracted by the main coding Ghost module BGh2 to generate a network layer feature map E2. The number of channels of the network layer feature map E2 is 256 dimensions, and the feature size of the network layer feature map E2 is consistent with the feature size of the feature map E2-1.
[0140] Figure 3 In the figure, MP2 is a maximum pooling processing module that performs maximum pooling processing on the network layer feature map E2. After the maximum pooling processing module MP2 performs maximum pooling processing on the network layer feature map E2, a maximum pooling processing feature map E1-0 can be generated, wherein the number of channels of the maximum pooling processing feature map E1-0 is 256 dimensions, and the feature size of the maximum pooling processing feature map E1-0 should be 1 / 16 of the corresponding feature size of the downhole track target image.
[0141] Figure 3 In the figure, CR1 is the main coding convolution block in the second main coding network layer, and BGh1 is the main coding Ghost module in the second main coding network layer. When the second main coding network layer performs the first coding sub-processing, the main coding convolution block CR1 is first used to perform coding convolution processing on the maximum pooling processing feature map E1-0, and generate the feature map E1-1. At this time, the number of channels of the feature map E1-1 is 512 dimensions, and the feature size of the feature map E1-1 is consistent with the feature size of the maximum pooling processing feature map E1-0. The feature map E1-1 is extracted by the main coding Ghost module BGh1 to generate a network layer feature map E1. The number of channels of the network layer feature map E1 is 512 dimensions, and the feature size of the network layer feature map E1 is consistent with the feature size of the feature map E1-1.
[0142] Figure 3 In the figure, MP1 is a maximum pooling processing module that performs maximum pooling processing on the network layer feature map E1. After the maximum pooling processing module MP1 performs maximum pooling processing on the network layer feature map E1, a maximum pooling processing feature map E0-0 can be generated, wherein the number of channels of the maximum pooling processing feature map E0-0 is 512 dimensions, and the feature size of the maximum pooling processing feature map E0-0 should be 1 / 32 of the corresponding feature size of the downhole track target image.
[0143] CR0 is the main coding convolution block in the first main coding network layer, and BGh0 is the main coding Ghost module in the first main coding network layer. When the first main coding network layer performs the first coding sub-processing, the main coding convolution block CR0 is first used to perform coding convolution processing on the maximum pooling processing feature map E0-0, and generate the feature map E0-1. At this time, the number of channels of the feature map E0-1 is 1024 dimensions, and the feature size of the feature map E0-1 is consistent with the feature size of the maximum pooling processing feature map E0-0. The feature map E0-1 is extracted by the main coding Ghost module BGh0 to generate the network layer feature map E0. The number of channels of the network layer feature map E0 is 1024 dimensions, and the feature size of the network layer feature map E0 is consistent with the feature size of the feature map E0-1.
[0144] It should be noted that the first main encoding network layer is located at the bottom layer of the U-Net architecture. Therefore, there is no need to perform the above-mentioned maximum pooling processing on the network layer feature map E0, that is, the network layer feature map E0 is the main encoding processing feature map.
[0145] In one embodiment of the present invention, the segmentation network auxiliary unit includes a SAM encoding unit and a size adjustment unit adaptively connected to the SAM encoding unit, wherein:
[0146] The downhole track target image is coded by using the SAM coding unit, and a basic coding embedding feature map is generated after the coding process;
[0147] The generated basic coding embedding feature map is resized by using a resizing unit to generate an auxiliary coding embedding feature map after the feature map resizing, wherein the feature map size of the auxiliary coding embedding feature map is consistent with the corresponding feature map sizes of the main coding deconvolution processing feature map and the main coding attention processing feature map.
[0148] Figure 3 In the above, SAM is the SAM coding unit, and the SAM coding unit can adopt the existing commonly used form, such as the corresponding description of SAM coding in Faster segment anything: Towards lightweight sam for mobile applications. When the segmentation network auxiliary unit is used to perform the second coding process on the underground track target image, the SAM coding unit is first used for coding process, and the basic coding embedding feature map is generated after the coding process. Figure 3 The S in the figure is the basic coding embedding feature map generated by the coding process performed by the SAM coding unit, wherein the number of channels of the basic coding embedding feature map may be 512 dimensions.
[0149] For the generated basic coding embedded feature map, the feature size is adjusted by the size adjustment unit so that the feature size of the auxiliary coding embedded feature map is consistent with the feature map size of the main coding deconvolution processing feature map and the main coding attention processing feature map. For example, the feature size of the auxiliary coding embedded feature map can be 1 / 16 of the feature size of the downhole track target image. Figure 3 In the figure, ZC3 is the auxiliary coding embedding feature map.
[0150] It should be noted that the size adjustment unit only adjusts the characteristic size of the size adjustment unit. The size adjustment unit can adopt an existing commonly used form. The specific method and process of implementing the characteristic size adjustment can be consistent with the existing technology and will not be repeated here.
[0151] In one embodiment of the present invention, when the first decoding process is performed using the main decoding process layer at the bottom of the U-Net architecture, there are:
[0152] The main decoding processing layer receives the main coding deconvolution processing feature map, and loads the received main coding sampling feature map into the correspondingly connected channel-spatial attention module, and the channel-spatial attention module also receives the network layer feature map output by the correspondingly connected main coding network layer;
[0153] Based on the received main encoding deconvolution processing feature map and the network layer feature map, the channel-spatial attention module performs attention mechanism processing to generate a main encoding attention processing feature map after the attention mechanism processing;
[0154] The auxiliary coding embedding feature map, the main coding deconvolution processing feature map, and the main coding attention processing feature map are feature spliced, and after the feature splicing processing, they are processed by Ghost to generate a decoding benchmark feature map.
[0155] From the above description, it can be seen that the number of channel-spatial attention modules should be consistent with the number of main encoding processing layers in the main encoding unit. When the main encoding unit includes four main encoding processing layers, when the main decoding processing layer at the bottom layer of the U-Net architecture is used as the first main decoding processing layer, then along the bottom layer of the U-Net architecture pointing to the direction of the opening of the U-Net architecture, the remaining main decoding processing layers are the second main decoding processing layer, the third main decoding processing layer and the fourth main decoding processing layer.
[0156] From the above description, we can see that the main unit of the segmentation network should include four channel-spatial attention modules. Figure 3 An embodiment of the segmentation network main unit including four channel-spatial attention modules is shown in the figure. The four channel-spatial attention modules are channel-spatial attention module SCCBAM1 to channel-spatial attention module SCCBAM4. Specifically, the channel-spatial attention module SCCBAM1 is adaptively connected to the second main encoding network layer and the first main decoding processing layer, the channel-spatial attention module SCCBMA2 is adaptively connected to the third main encoding network layer and the second main decoding processing layer, the channel-spatial attention module SCCBAM3 is adaptively connected to the fourth main encoding network layer and the third main decoding processing layer, and the channel-spatial attention module SCCBMA4 is adaptively connected to the reading main encoding network layer and the fourth main decoding processing layer.
[0157] In specific implementation, the first main decoding processing layer should be configured to perform the first decoding processing. When the first main decoding processing layer performs the first decoding processing, the main coding processing feature map is first deconvolved. For example, a deconvolution module with a convolution kernel size of 2*2 can be used to perform deconvolution processing on the main coding processing feature map, so that the main coding deconvolution processing feature map can be obtained after the deconvolution processing. After the deconvolution processing, the number of channels of the main coding deconvolution processing feature map can be 512 dimensions, and the feature size of the main coding deconvolution processing feature map is 1 / 16 of the corresponding feature size of the downhole track target image. Figure 3 In the figure, ZC1 is the main coding deconvolution processing feature map, UC0 is the decoding deconvolution block for deconvolution processing the main coding processing feature map, and the decoding deconvolution block can adopt the existing commonly used deconvolution module, such as the decoding deconvolution block can adopt the deconvolution module with a convolution kernel size of 2*2.
[0158] Specifically, after obtaining the main coding deconvolution processing feature map, the main coding deconvolution processing feature map should be loaded into the channel-spatial attention module SCCBAM1, and the channel-spatial attention module SCCBAM1 also receives the network layer feature map E1 generated by the second main coding network layer. The channel-spatial attention module SCCBAM1 performs attention mechanism processing on the main coding deconvolution processing feature map and the network layer feature map E1 to generate the main coding attention processing feature map after the attention mechanism processing. It should be noted that the number of channels of the main coding attention processing feature map is 512 dimensions, and the feature size of the main coding attention main feature map is 1 / 16 of the corresponding feature size of the downhole track target image. Figure 3 In the figure, ZC2 is the main encoding attention processing feature map.
[0159] In the first main decoding processing layer, the auxiliary coding embedding feature map, the main coding deconvolution processing feature map, and the main coding attention processing feature map are feature spliced in the channel number dimension, and after the feature splicing process, the decoding reference feature map is generated after Ghost processing. From the above description, it can be seen that the generated decoding reference feature map should be transmitted to the second main decoding processing layer. Figure 3 In the figure, D0 is the decoding benchmark feature map, the number of channels of the decoding benchmark feature map should be 512 dimensions, and the feature size of the decoding benchmark feature map is 1 / 16 of the corresponding feature size of the downhole track target image.
[0160] It should be understood that in order to perform the above-mentioned first decoding process, the first main decoding processing layer should include a decoding deconvolution block, a feature map splicer and a decoding Ghsot module. Figure 3 In the example, JGh0 is the decoding Ghost module in the first main decoding processing layer, but the feature map splicer is not in Figure 3 As shown in the figure, the parameter setting of the decoding deconvolution block can refer to the above description, the feature map splicer can adopt the form in the prior art, and the decoding Ghost module can refer to the description of the above main encoding Ghost module.
[0161] In one embodiment of the present invention, for a main decoding processing layer that performs the second decoding process, the main decoding processing layer includes a decoding deconvolution block and a main decoding Ghost module connected in sequence, wherein:
[0162] When performing the second decoding process, a basic decoding feature map to be decoded is received, and thereafter, a decoding deconvolution block is used to perform a deconvolution process on the basic decoding feature map to generate a deconvolution-post decoding feature map after the deconvolution process;
[0163] Loading the deconvolution decoded feature map into the channel-spatial attention module of the current main decoding processing layer, and the channel-spatial attention module also receives the network layer feature map corresponding to the output of the main encoding network layer;
[0164] Based on the received post-deconvolution decoding feature map and the network layer feature map, the channel-spatial attention module performs attention mechanism processing to generate a decoding attention processing feature map after the attention mechanism processing;
[0165] The main decoding Ghost module is used to extract features from the decoding attention processing feature map to generate a basic decoding feature map after feature extraction;
[0166] In the main decoding unit, for two adjacent main decoding processing layers, along the bottom of the U-Net architecture pointing to the direction of the opening, the basic decoding feature map is generated by the main decoding processing layer below and transmitted to the main decoding processing layer above.
[0167] In a specific implementation, the second decoding process performed by all main decoding process layers is the same, and the difference from the first decoding process performed by the first main decoding process layer is that the feature map splicing using the feature map splicer can be omitted. From the above description, it can be known that for the second main decoding process layer, the received basic decoding feature layer should be the decoding reference feature map generated by the first main decoding process performed by the first main decoding process layer, for the third main decoding process layer, the basic decoding feature map should be generated by the second main decoding process layer, and for the fourth main decoding process layer, the basic decoding feature map should be generated by the third main decoding process layer.
[0168] Figure 3 In the figure, UC1 is the decoding deconvolution block in the second main decoding processing layer, JGh1 is the main decoding Ghost module in the second main decoding processing layer, D1-0 is the deconvolution decoding feature map generated by upsampling the decoding benchmark feature map using the decoding deconvolution block UC1, and D1-1 is the decoding attention processing feature map generated by the channel-spatial attention module SCCBMA2, wherein the number of channels of the deconvolution decoding feature map D1-0 is 256 dimensions, and the feature size of the deconvolution decoding feature map D1-0 is 1 / 16 of the corresponding feature size of the downhole track target image; the channel-spatial attention module SCCBMA2 performs attention mechanism processing on the deconvolution decoding feature map D1-0 and the network layer feature map E2, so that the decoding attention processing feature map D1-1 can be generated after the attention mechanism processing, and the number of channels of the decoding attention processing feature map D1-1 is 256 dimensions,
[0169] The main decoding Ghost module JGh1 performs feature extraction on the decoding attention processing feature map D1-1 to generate a basic decoding feature map D1. The number of channels of the basic decoding feature map D1 is 512 dimensions, and the feature size of the basic decoding feature map D1 is 1 / 16 of the corresponding feature size of the downhole track target image.
[0170] Figure 3 , UC2 is a decoding deconvolution block in the third main decoding processing layer, JGh2 is a main decoding Ghost module in the third main decoding processing layer, D2-0 is a deconvolution decoding feature map generated by upsampling the basic decoding feature map D1 using the decoding deconvolution block UC2, and D2-1 is a decoding attention processing feature map generated by the channel-spatial attention module SCCBMA3, wherein,
[0171] The number of channels of the decoded feature map D2-0 after deconvolution is 256 dimensions, and the feature size of the decoded feature map D2-0 after deconvolution is 1 / 8 of the corresponding feature size of the downhole track target image. The channel-spatial attention module SCCBMA3 performs attention mechanism processing on the decoded feature map D2-0 after deconvolution and the network layer feature map E3, so that after the attention mechanism processing, the decoding attention processing feature map D2-1 can be generated. The main decoding Ghost module JGh2 performs feature extraction on the decoding attention processing feature map D2-1 to generate the basic decoding feature map D2. The number of channels of the basic decoding feature map D2 is 256 dimensions, and the feature size of the basic decoding feature map D2 is 1 / 8 of the corresponding feature size of the downhole track target image.
[0172] Figure 3 , UC3 is a decoding deconvolution block in the fourth main decoding processing layer, JGh3 is a main decoding Ghost module in the fourth main decoding processing layer, D3-0 is a deconvolution decoding feature map generated by upsampling the basic decoding feature map D2 using the decoding deconvolution block UC3, and D3-1 is a decoding attention processing feature map generated by the channel-spatial attention module SCCBMA4, wherein,
[0173] The channel-spatial attention module SCCBMA4 performs attention mechanism processing on the post-deconvolution decoding feature map D3-0 and the network layer feature map E4, so that the decoding attention processing feature map D3-1 can be generated after the attention mechanism processing. The main decoding Ghost module JGh4 performs feature extraction on the decoding attention processing feature map D3-1 to generate a basic decoding feature map D3. The number of channels of the basic decoding feature map D3 can be 64 dimensions, and the feature size of the basic decoding feature map D3 can be 1 / 2 of the corresponding feature size of the downhole track target image.
[0174] It should be noted that after the basic decoding feature map D3 is generated, it can also be processed by the decoding output processing module to generate an underground railway track segmentation image after the output processing, wherein the decoding output processing module can be a decoding output processing convolution block and a decoding processing Sigmoid activation function. Specifically, the decoding output processing convolution module can use a 1*1 convolution kernel. During the output processing, the basic decoding feature map D3 should first be convolved by the decoding output convolution block. The decoding output convolution block can be used to increase the nonlinear characteristics while keeping the feature size of the basic decoding feature map D3 unchanged. Thereafter, it is processed by the decoding processing Sigmoid activation function to map the output to between 0 and 1 by using the over-decoding processing Sigmoid activation function, so as to realize the image segmentation of the present invention and generate an underground railway track segmentation image. Figure 3 UCS is the decoding output processing module.
[0175] In one embodiment of the present invention, the channel-spatial attention module includes an attention first stitcher, a channel attention module, a spatial attention module and an attention linear layer, wherein:
[0176] The attention first splicer is connected to the corresponding main encoding processing layer and the main decoding processing layer;
[0177] The channel attention module is connected to the output end of the first attention splicer, and the channel attention module adopts a residual connection;
[0178] The spatial attention module is connected to the output end of the channel attention module, and the spatial attention module adopts residual connection;
[0179] The output end of the spatial attention module is connected to the attention linear layer, and the attention linear layer is used as the output layer of the channel-spatial attention module.
[0180] In specific implementation, the channel-spatial attention modules SCCBAM1 to SCCBAM4 can adopt the same form. Figure 4 An embodiment of a channel-spatial attention module is shown in Figure 4 In the figure, Cat1 is the first attention splicer, Channel Attention is the channel attention module, Spatial Attention is the spatial attention module, XN3 is the attention linear layer, and the attention linear layer XN3 can use a multi-layer perceptron. Through the attention linear layer, the number of channels of the output feature map can be reduced to half the number of channels of the input feature map. Figure 4In the figure, From Encoder Layer specifically refers to the network layer feature map output from the main encoding processing layer, and From Lower Decoder Layer specifically refers to the feature map from the main decoding processing layer. For example, when the channel-spatial attention module SCCBAM1 is connected to the first main decoding processing layer, the feature map of From Lower Decoder Layer should be the main encoding deconvolution processing feature map, and when the channel-spatial attention module is connected to the second to fourth main decoding processing layers, the feature map of From Lower Decoder Layer should be the deconvolution post-decoding feature map. The main encoding deconvolution processing feature map and the deconvolution post-decoding feature map can be referred to the corresponding descriptions above, which will not be repeated here.
[0181] For any channel-spatial attention module, the channel-spatial attention module is adapted and connected to the corresponding main encoding processing layer and the main decoding processing layer through the first attention splicer. For example, for the channel-spatial attention module SCCBAM1, it is connected to the first main decoding processing layer and the main encoding processing layer formed by the fourth main encoding network layer through the first attention splicer. Other connection conditions can be referred to here and the corresponding descriptions above.
[0182] It can be understood that feature stitching can be performed through the attention-first stitcher, channel attention processing can be achieved through the channel attention module, and spatial attention processing can be achieved through the spatial attention module. That is, the above-mentioned attention mechanism processing should include channel attention processing and spatial attention processing.
[0183] It should be noted that when the corresponding main encoding processing layer and the main decoding processing layer are jump-connected through the channel-spatial attention module, since the channel attention processing and the spatial attention processing are realized at the same time, and when the attention mechanism processing is executed, the channel attention processing and the spatial attention processing are performed in sequence, compared with relying solely on a single attention mechanism, it can significantly improve the multi-dimensional interaction in the channel and spatial domains in image segmentation, enhance the ability to identify and extract key information features, effectively reduce the common spatial information loss in the coding convolution processing operation, and further improve the accuracy of image segmentation.
[0184] In one embodiment of the present invention, the channel attention module includes a channel attention first maximum pooling module and a channel attention first average pooling module, wherein:
[0185] The channel attention first maximum pooling module and the channel attention first average pooling module are both connected to the output end of the attention first splicer;
[0186] The output end of the first maximum pooling module of the channel attention and the output end of the first average pooling module of the channel attention are connected to the first linear module of the channel attention, and the first linear module of the channel attention is connected to the second linear module of the channel attention through the LeakyReLu activation function of the channel attention;
[0187] The output end of the second linear module of the channel attention is connected to the second maximum pooling module of the channel attention and the second average pooling module of the channel attention, and the second maximum pooling module of the channel attention and the second average pooling module of the channel attention are both connected to the channel attention splicer;
[0188] The output end of the channel attention splicer is connected to the channel attention adder used to form the residual connection of the channel attention module through the channel attention sigmoid function, and is adaptively connected to the spatial attention module through the channel attention adder.
[0189] Figure 4 In the figure, Maxpool1 is the first maximum pooling module of channel attention, and Maxpool2 is the second maximum pooling module of channel attention, wherein the first maximum pooling module of channel attention and the second maximum pooling module of channel attention can perform maximum pooling processing respectively. Avgpool1 is the first average pooling module of channel attention, and Avgpool2 is the second average pooling module of channel attention, wherein the first average pooling module of channel attention and the second average pooling module of channel attention can perform average pooling processing respectively. Figure 4 In the figure, XN1 is the first linear module of channel attention, LeakyReLu is the LeakyReLu activation function of channel attention, XN2 is the second linear module of channel attention, Cat2 is the channel attention splicer, S1 is the channel attention Sigmoid function, and Ad1 is the channel attention adder. Both the first linear module of channel attention and the second linear module of channel attention can use multi-layer perceptron to realize linear transformation processing.
[0190] Figure 4 In the figure, F1 is the feature map output by the attention first stitcher, and F2 is the feature map output by the channel attention adder. In the specific implementation, after the feature map F1 is generated by the attention first stitcher, two different spatial context representations can be generated by the channel attention first maximum pooling module and the channel attention first average pooling module: and in, It is the spatial context representation generated by the feature map F1 after the channel attention first maximum pooling module performs maximum pooling processing. The spatial context representation generated after the feature map F1 is processed by the channel attention first average pooling module. After that, the channel attention first linear module is used for dimensional compression to reduce the amount of parameter calculation. The channel attention first linear module can reduce the number of channels of the feature map by half, and the channel attention LeakyReLu activation function can enhance the expression of nonlinear relationships between channels and improve feature interaction capabilities. The channel attention second linear module can restore the feature size of the feature map to the feature dimension before being processed by the channel attention first linear module.
[0191] In one embodiment of the present invention, the spatial attention module includes a spatial attention first convolution block, a normalized activation function module, a spatial attention second convolution block, a normalization module and a spatial attention Sigmoid function connected in sequence, wherein:
[0192] The first spatial attention convolution block is connected to the output of the channel attention adder;
[0193] The convolution kernel sizes used in the first spatial attention convolution block and the second spatial attention convolution block are the same;
[0194] When the spatial attention module adopts residual connection, the spatial attention sigmoid function is connected to the input of the spatial attention adder, and the input of the first convolutional block of the spatial attention is also connected to the input of the spatial attention adder, and the output of the spatial attention adder is connected to the attention linear layer.
[0195] Figure 4 In the figure, C1 is the first convolution block of spatial attention, C2 is the second convolution block of spatial attention, BNLR is the normalization activation function module, BN is the normalization module, S2 is the spatial attention Sigmoid function, and Ad2 is the spatial attention adder. F3 is the feature map output by the spatial attention module, that is, the feature map output by the spatial attention adder Ad2. In the specific implementation, the size of the convolution kernel used by the first convolution block of spatial attention and the second convolution block of spatial attention can be 7*7, and the normalization activation function module BNLR can include a normalization layer and a LeakyReLu activation function. The feature map F3 can obtain the feature map x3 through the attention linear layer. It should be noted that the number of convolution kernels used by the first convolution block of spatial attention and the second convolution block of spatial attention is 1. By using a convolution kernel of size 7*7, the receptive field can be increased to sense features in a wider range.
[0196] When the spatial attention module adopts the above form, the first spatial attention convolution block is first used to perform a convolution operation on the feature map, and a single-channel feature map is output after the convolution operation to effectively capture a large range of spatial dependencies. At the same time, the spatial weight can be generated after the number of channels is reduced to a single channel; the normalized activation function module is used to eliminate the input offset problem in the deep layer of the network and improve the effectiveness of back propagation. Thereafter, the convolution operation is performed through the second spatial attention convolution block; finally, the convolution result is normalized to a spatial attention weight map of 0-1 through the spatial attention Sigmoid function to identify the importance of different spatial positions, so as to realize dynamic feature screening in the spatial dimension, so that the image segmentation network of the present invention pays more attention to semantic areas related to the task.
[0197] From the above description, it can be seen that when the spatial attention first convolution block and the spatial attention second convolution block are set in the spatial attention module, the ability to extract spatial features can be effectively enhanced. This is because a larger receptive field helps to capture information more comprehensively, thereby improving the ability to extract spatial features.
[0198] The above-mentioned image segmentation network can be constructed by the following method, which includes:
[0199] Constructing a basic image segmentation model and a segmentation model training data set for training the basic image segmentation model;
[0200] The model training condition parameters are configured, and the model training is performed based on the configured model training condition parameters and the segmentation model training data set. When the image segmentation basic model training reaches the target state, the corresponding image segmentation basic model configuration is used as the image segmentation network.
[0201] It can be understood that the basic model of image segmentation is the model to be built, and the image segmentation network is a model that trains the basic model of image segmentation to reach the target state. Therefore, the constructed basic model of image segmentation should be consistent with the image segmentation network. The situation of the image segmentation substrate model can refer to the description of the above-mentioned image segmentation network, which will not be repeated here.
[0202] In order to achieve image segmentation, a segmentation model training data set should also be constructed. The segmentation model training data set should include several training samples, where each training sample should be generated based on images containing railway tracks taken underground in a coal mine. Specifically, when generating training samples, an image containing railway tracks should be obtained first. Thereafter, a closed area containing railway tracks should be outlined on the image using tools such as VIA, and a label of the railway track area should be added to the closed area to use the label to indicate that the current closed area is a railway track area.
[0203] From the above description, it can be seen that each training sample should be an image containing a closed area and a label. The resolution of each training sample is 1280×800, and the corresponding label is stored in JSON format. The number of training samples in the segmentation model training data set can be selected as needed. Generally, the training samples can be produced in the above manner. In order to improve the diversity and robustness of the model training data set, after the training samples are produced in the above manner, new training samples can also be constructed by image transformation methods. The image training method may include random scaling (ranging from 0.75 to 1.25), rotation (within a certain angle) and horizontal flipping, etc. The image transformation method can be selected as needed and will not be repeated here.
[0204] It should be noted that the configured model training condition parameters should generally include the loss function and the training parameters necessary for model training, such as the number of training samples in each batch, the initial learning rate is set to 5e-3, and the planned decay rate is set to 0.004. The necessary training parameters can be selected as needed to meet the training requirements, and they are not listed here one by one. In order to cope with the challenges brought by imbalance, the loss function of the present invention adopts a hybrid loss, then:
[0205]
[0206] Among them, L z is the loss value of the loss function, H is the height of the training output segmentation image, W is the width of the training output segmentation image, and C is the number of categories represented by the segmentation; (a, b) represents the coordinate position of a pixel point in the training output segmentation image; p j (a, b) represents the predicted probability of the pixel at coordinate position (a, b) in the training output segmentation image belonging to category j; y j (a, b) is the actual situation where the pixel at coordinate position (a, b) in the training output segmentation image is of category j, and λ is the weighting parameter.
[0207] For any training sample, a corresponding training output segmentation image can be output through the image segmentation basic model. Based on the output training output segmentation image, the loss value of the above loss function can be calculated. In specific implementation, the prediction probability p j (a, b) can be directly obtained through the training output segmentation image output by the basic image segmentation model. In the training output segmentation image, when the pixel point at the coordinate position (a, b) belongs to category j, then y j The value of (a,b) should be 1, otherwise, y j The value of (a, b) should be 0. In addition, the value of the weighting parameter λ can be 0.6. From the above description of the image segmentation purpose of the present invention, it can be seen that the value of the above segmentation category number C should be 2.
[0208] It should be understood that the image size of the training output segmented image is consistent with the size of the corresponding image in the training sample. When the image of the training sample is 1280×800, the height H of the training output segmented image is 1280, and the width W of the training output segmented image is 800. In addition, the size of the downhole track target image should also be 1280×800.
[0209] It should be noted that when the image segmentation basic model is trained using the segmentation model training data set, when the value of the loss function tends to be stable, it can be considered that the basic training of the image segmentation has reached the target state. At this time, the image segmentation basic model that has reached the target state can be configured as an image segmentation network.
Claims
1. An image segmentation method suitable for railway track images in coal mines, characterized in that: The image segmentation method comprises: The downhole track target image to be segmented is obtained, and the acquired downhole track target image is loaded into the constructed image segmentation network, so as to use the image segmentation network to perform image segmentation on the downhole track target image, and generate a downhole track segmentation image after image segmentation, wherein the downhole track segmentation image includes a railway track segmentation marked area and a non-railway track segmentation marked area, wherein: The image segmentation network includes a segmentation network main unit based on a U-Net architecture and a segmentation network auxiliary unit based on SAM coding, wherein the segmentation network main unit includes a main coding unit and a main decoding unit adaptively connected to the main coding unit, and the segmentation network auxiliary unit is adaptively connected to the main decoding unit; When performing image segmentation on the downhole track target image, the main coding unit in the segmentation network main unit is used to perform a first coding process to generate a main coding process feature map after the first coding process, and at the same time, the segmentation network auxiliary unit is used to perform a second coding process on the downhole track target image to generate an auxiliary coding embedding feature map after the second coding; The main coding processing feature map and the auxiliary coding embedding feature map are transmitted to the main decoding unit for decoding processing by the main decoding unit, and a downhole track segmentation image is generated after decoding processing, wherein: When the main decoding unit performs decoding processing, it includes at least a first decoding processing and a plurality of second decoding processings performed in sequence; When performing the first decoding process, the auxiliary coding embedding feature map is subjected to feature splicing processing with the main coding deconvolution processing feature map and the main coding attention processing feature map generated based on the main coding processing feature map, so as to generate a decoding reference feature map after feature splicing processing, wherein, Performing deconvolution processing on the main coding processing feature map, and generating a main coding deconvolution processing feature map after the deconvolution processing; Performing attention mechanism processing on the main coding processing feature map, and generating the main coding attention processing feature map after the attention mechanism processing; The generated decoded reference feature map is subjected to a second decoding process, and a downhole track segmentation image is generated after the second decoding process.
2. The image segmentation method suitable for railway track images in coal mines according to claim 1 is characterized in that: The main coding unit includes a plurality of main coding network layers connected in sequence, wherein, in the main coding unit, the main coding network layer at the bottom of the U-Net architecture is configured as a main coding transition connection layer, and the remaining main coding network layers are respectively configured as main coding processing layers; The main decoding unit comprises a plurality of main decoding processing layers connected in sequence, wherein: In the main unit of the segmentation network, the main encoding processing layer in the main encoding unit corresponds one-to-one to the main decoding processing layer in the main decoding unit, and the main encoding processing layer is jump-connected to the corresponding main decoding processing layer through the channel-spatial attention module; The main coding transition connection layer is adaptively connected to the main decoding processing layer at the bottom of the U-Net architecture in the main decoding unit, and transmits the generated main coding processing feature map to the corresponding main decoding processing layer through the main coding transition connection layer, and the main decoding processing layer at the bottom of the U-Net architecture also receives the auxiliary coding embedded feature map; When the main decoding unit performs decoding processing, the main decoding processing layer at the bottom of the U-Net architecture is used to perform the first decoding processing, and the remaining main decoding processing layers are configured to perform the second decoding processing.
3. The image segmentation method suitable for railway track images in underground coal mines according to claim 2 is characterized in that: The main coding network layer includes a main coding convolution block and a main coding Ghost module connected in sequence, wherein: When the main coding unit performs the first coding process, the coding feature extraction is performed in sequence through the main coding network layer, so as to use each main coding network layer to perform the first coding sub-processing, wherein, when extracting the coding feature, the main coding convolution block is firstly used to perform the coding convolution process, and after the coding convolution process, the main coding Ghost module is used to perform feature extraction, and the corresponding network layer feature map is generated after the feature extraction of the main coding Ghost module; In the main coding unit, for two adjacent main coding network layers, along the direction of the opening of the U-Net architecture pointing to the bottom, the network layer feature map generated by the upper main coding network layer is subjected to maximum pooling processing to generate a maximum pooled feature map after maximum pooling processing, and the maximum pooled feature map is loaded into the main network coding layer below.
4. The image segmentation method suitable for railway track images in coal mines according to claim 3, characterized in that: When the first decoding process is performed using the main decoding processing layer at the bottom of the U-Net architecture, there are: The main decoding processing layer receives the main coding deconvolution processing feature map, and loads the received main coding sampling feature map into the correspondingly connected channel-spatial attention module, and the channel-spatial attention module also receives the network layer feature map output by the correspondingly connected main coding network layer; Based on the received main encoding deconvolution processing feature map and the network layer feature map, the channel-spatial attention module performs attention mechanism processing to generate a main encoding attention processing feature map after the attention mechanism processing; The auxiliary coding embedding feature map, the main coding deconvolution processing feature map, and the main coding attention processing feature map are feature spliced, and after the feature splicing processing, they are processed by Ghost to generate a decoding benchmark feature map.
5. The image segmentation method suitable for railway track images in coal mines according to claim 3, characterized in that: For the main decoding processing layer that performs the second decoding process, the main decoding processing layer includes a decoding deconvolution block and a main decoding Ghost module connected in sequence, wherein: When performing the second decoding process, a basic decoding feature map to be decoded is received, and thereafter, a decoding deconvolution block is used to perform a deconvolution process on the basic decoding feature map to generate a deconvolution-post decoding feature map after the deconvolution process; Loading the deconvolution decoded feature map into the channel-spatial attention module of the current main decoding processing layer, and the channel-spatial attention module also receives the network layer feature map corresponding to the output of the main encoding network layer; Based on the received post-deconvolution decoding feature map and the network layer feature map, the channel-spatial attention module performs attention mechanism processing to generate a decoding attention processing feature map after the attention mechanism processing; The main decoding Ghost module is used to extract features from the decoding attention processing feature map to generate a basic decoding feature map after feature extraction; In the main decoding unit, for two adjacent main decoding processing layers, along the bottom of the U-Net architecture pointing to the direction of the opening, the basic decoding feature map is generated by the main decoding processing layer below and transmitted to the main decoding processing layer above.
6. The image segmentation method suitable for railway track images in coal mines according to claim 2, characterized in that: The channel-spatial attention module includes an attention first stitcher, a channel attention module, a spatial attention module, and an attention linear layer, wherein: The attention first splicer is connected to the corresponding main encoding processing layer and the main decoding processing layer; The channel attention module is connected to the output end of the first attention splicer, and the channel attention module adopts a residual connection; The spatial attention module is connected to the output end of the channel attention module, and the spatial attention module adopts residual connection; The output end of the spatial attention module is connected to the attention linear layer, and the attention linear layer is used as the output layer of the channel-spatial attention module.
7. The image segmentation method suitable for railway track images in coal mines according to claim 6, characterized in that: The channel attention module includes a channel attention first maximum pooling module and a channel attention first average pooling module, wherein: The channel attention first maximum pooling module and the channel attention first average pooling module are both connected to the output end of the attention first splicer; The output end of the first maximum pooling module of the channel attention and the output end of the first average pooling module of the channel attention are connected to the first linear module of the channel attention, and the first linear module of the channel attention is connected to the second linear module of the channel attention through the LeakyReLu activation function of the channel attention; The output end of the second linear module of the channel attention is connected to the second maximum pooling module of the channel attention and the second average pooling module of the channel attention, and the second maximum pooling module of the channel attention and the second average pooling module of the channel attention are both connected to the channel attention splicer; The output end of the channel attention splicer is connected to the channel attention adder used to form the residual connection of the channel attention module through the channel attention sigmoid function, and is adaptively connected to the spatial attention module through the channel attention adder.
8. The image segmentation method for railway track images in underground coal mines according to claim 6, characterized in that: The spatial attention module includes a spatial attention first convolution block, a normalized activation function module, a spatial attention second convolution block, a normalization module and a spatial attention Sigmoid function connected in sequence, wherein: The first spatial attention convolution block is connected to the output of the channel attention adder; The convolution kernel sizes used in the first spatial attention convolution block and the second spatial attention convolution block are the same; When the spatial attention module adopts residual connection, the spatial attention sigmoid function is connected to the input of the spatial attention adder, and the input of the first convolutional block of the spatial attention is also connected to the input of the spatial attention adder, and the output of the spatial attention adder is connected to the attention linear layer.
9. The image segmentation method for railway track images in underground coal mines according to any one of claims 1 to 8, characterized in that: The segmentation network auxiliary unit includes a SAM encoding unit and a size adjustment unit adaptively connected to the SAM encoding unit, wherein: The downhole track target image is coded by using the SAM coding unit, and a basic coding embedding feature map is generated after the coding process; The generated basic coding embedding feature map is resized by using a resizing unit to generate an auxiliary coding embedding feature map after the feature map resizing, wherein the feature map size of the auxiliary coding embedding feature map is consistent with the corresponding feature map sizes of the main coding deconvolution processing feature map and the main coding attention processing feature map.
10. The image segmentation method for railway track images in underground coal mines according to any one of claims 1 to 8, characterized in that: When acquiring underground track target images, it includes: A downhole track source image is acquired, and the acquired downhole track source image is preprocessed to generate a downhole track target image after the preprocessing, wherein: The preprocessing of the downhole track source image includes bilateral filtering and / or contrast limiting adaptive histogram equalization enhancement processing; When performing contrast-limited adaptive histogram equalization enhancement processing, including Obtain a source image to be processed by equalization and enhancement, and divide the source image to be processed by equalization and enhancement into a plurality of source image sub-blocks. For any source image sub-block, there is: Wherein, (b1, b2) is the size of each source image sub-block, (H, W) is the size of the source image to be processed by equalization and enhancement, K is the number of sub-blocks of the divided source image, and the source image for equalization and enhancement processing is the downhole track source image or the downhole track filtered image generated by bilateral filtering; Contrast limiting processing is performed on each source image sub-block, wherein the contrast limiting processing includes: For the source image sub-block T, a grayscale histogram of the source image sub-block T is generated, and the frequency of the i-th grayscale level is limited based on the generated grayscale histogram, then: Where DB is the frequency limit threshold, H T (i) is the frequency of the source image sub-block T at the i-th grayscale level, and L is the number of grayscale levels of the source image to be processed by equalization enhancement; After frequency limitation, the mirror grayscale mapping process is performed, and then: in, is the grayscale mapping of the i-th grayscale level; is the cumulative distribution function of the i-th gray level; After the mirror grayscale mapping process, a mapped grayscale histogram is generated based on the grayscale mapping information, and a corresponding contrast-limited sub-block is generated based on the mapped grayscale histogram; Any two contrast-limited sub-blocks are merged by linear interpolation to generate a downhole track target image after merging.
Citation Information
Patent Citations
Conditional generative adversarial remote sensing image target segmentation method containing multi-level channel attention
CN111259906A
Medical image segmentation method and electronic equipment
CN114782440A
Lung lobe segmentation method and system combined with attention under multi-task learning framework
CN116030078A
Image processing method and system based on double-branch multi-scale semantic segmentation network
CN116580241A
Single image defogging method based on random mask convolution and attention mechanism
CN116721033A
Cited By
Unstructured environment significance semantic segmentation method and system
CN120495669A
A method and system for unstructured environment saliency semantic segmentation
CN120495669B