Power equipment semantic segmentation method based on visible light and infrared image feature fusion
Patent Information
- Application Number
- CN202410144812.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-01
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-02-01
AI Technical Summary
然而可见光图像在光照恶劣等情况下成像质量较差,往往并不能为模型训练提供更多更具有区分度的信息,常常会导致预测结果不准确
[0059] (1) The semantic segmentation method model for power equipment based on the feature fusion of visible light images and infrared images proposed in this invention makes full use of the complementary advantages of visible light images and infrared images, making the semantic fusion between visible light images and infrared images more complete and obtaining more comprehensive and accurate semantic information.
Smart Images

Figure CN118196405B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image segmentation, and in particular relates to a semantic segmentation method for power equipment based on the fusion of features from visible light images and infrared images. Background Technology
[0002] Image segmentation is the process of dividing a digital image into multiple regions with similar properties. Its purpose is to classify and label objects, regions, or features in an image for subsequent analysis and processing. Semantic segmentation is an important task in the field of image segmentation. It divides an input image into multiple regions and assigns a semantic category label to each pixel to indicate which object or region in the image that pixel belongs to.
[0003] Currently, most semantic segmentation models obtain semantic segmentation results solely through processing visible light images. However, visible light images often suffer from poor image quality under adverse lighting conditions, failing to provide sufficient discriminative information for model training and frequently leading to inaccurate predictions. During power equipment inspections, complex scenes with similar textures, occlusions, darkness, and smoke are frequently encountered. In recent years, with the widespread adoption of infrared cameras, infrared information has proven highly effective in addressing recognition ambiguity caused by poor lighting conditions. Infrared images contain ample semantic information and are less affected by illumination, eliminating or reducing optical noise interference in visible light images. Furthermore, visible light images can effectively supplement the detailed information of objects in thermal infrared images, serving as a crucial information complement to visible light images. Therefore, there is an urgent need for an image semantic segmentation method that can more fully integrate features from both visible light and infrared images, which is of great significance to the field of image segmentation. Summary of the Invention
[0004] Purpose of the invention: In view of the prior art, the present invention provides a semantic segmentation method for power equipment based on the fusion of visible light image and infrared image features, which can effectively combine visible light image features and infrared image features to make semantic segmentation more accurate.
[0005] Technical solution: This invention provides a semantic segmentation method for power equipment based on the fusion of visible light and infrared image features, comprising the following steps:
[0006] Step S1: Obtain a dataset containing pairs of visible light and infrared images of substation power equipment;
[0007] Step S2: Use the encoding module to extract features from the visible light image and the infrared image;
[0008] Step S3: Perform feature fusion and feature enhancement on the extracted visible light features and infrared features. After fusing and enhancing the high-level features using the feature fusion and enhancement module combined with superpixel segmentation, combine the semantic information in the high-level features and use the multi-dimensional cross-layer feature fusion and enhancement module to fuse and enhance the mid- and low-level features.
[0009] Step S4: Use the decoding module to perform decoding, and combine it with a multi-stage loss function to output accurate semantic segmentation results for power equipment.
[0010] Further, the step of obtaining the dataset in step S1 is as follows:
[0011] Step S11: Using visible light cameras and infrared cameras, acquire visible light and infrared registered images of substation power equipment against a complex background from different angles and at different times.
[0012] Step S12: Manually annotate the visible light and infrared image pairs obtained in step S11 to obtain the corresponding power equipment true value map;
[0013] Step S13: Combine the image information obtained in steps S11 and S12 to obtain the required dataset.
[0014] Further, in step S2, the encoding module consists of two independent ResNet152 residual networks forming a dual-channel encoding module to receive visible light and infrared images; the visible light and infrared images are input into the two residual networks, and feature extraction is performed through convolution and pooling operations to obtain the visible light image features. Features of infrared images Where i represents the feature level; as the feature extraction layer deepens, the extraction progresses from low-level features to high-level features.
[0015] The output features of the encoding module are divided into 5 layers for downsampling, reducing the spatial resolution of each layer and doubling the number of channels. The extracted visible light image features and infrared image features are respectively... and
[0016] Furthermore, in step S3, the feature fusion enhancement module combined with superpixel segmentation includes:
[0017] Based on the complementary advantages of visible light and infrared images, their high-level features are fused, key information in the feature map is located using a spatial attention mechanism, and the feature map is enhanced using the spatial attention weights of each modality, as shown in formula (1):
[0018]
[0019] In the formula: Indicates the characteristics of visible light; Indicates infrared features; i indicates which layer of features; This indicates element-wise multiplication. Indicates element-wise addition; F fuse Indicates the features after fusion; f conv (·) denotes a one-dimensional convolutional layer; SA(·) denotes a spatial attention mechanism; f cat (·) indicates a channel-level connection;
[0020] Secondly, an iterative edge thinning method is used to perform superpixel segmentation on the infrared image, and the superpixel weights of each segmented region are calculated to obtain a superpixel segmentation map (spsmap) of the infrared image, which is then used to form a superpixel mask. Where N S The superpixel mask for a single channel is SSM (Short Scale Array). i ∈R 1×H×W i = 1, 2, ..., N S ;
[0021] A prototype O is built based on each superpixel mask. i And concatenate each prototype to generate a corresponding prototype block. Generate prototype sampler vectors from prototype blocks Then multiply p by SSM to generate the superpixel-assisted prediction map SSM. pred This process can be represented by formula (2):
[0022]
[0023] In the formula: APO(·) represents the average pooling operator of the superpixel mask; MLP(·) represents the multilayer sensing module; σ is the sigmoid activation function;
[0024] The weight of each superpixel is calculated by using the position and feature information of the superpixel to fuse the features of different channels. At the same time, the superpixel features are used to assist the fusion features of visible light images and infrared images, so that the network pays more attention to the foreground information of the target and suppresses the background features, thereby improving the image clarity and recognition effect, as shown in formula (3):
[0025]
[0026] In the formula: This indicates that upsampling is performed using bilinear interpolation; concat(·) is the channel concatenation operator; F′ fuse This indicates the fusion features of the enhanced visible light image and infrared image.
[0027] Furthermore, the multi-dimensional cross-layer feature fusion enhancement module includes:
[0028] The multi-dimensional cross-layer module (MDCM) is used to enhance visible light features and infrared image features, and then combines the features of each modality to perform cross-modal fusion.
[0029] When dealing with intermediate features When performing feature fusion, the bilinear interpolation method is first used for upsampling, and then cross-modal enhancement between visible light features and infrared features is performed. The visible light features, infrared features and the enhanced features of the previous layer are aggregated and output, as shown in formula (4):
[0030]
[0031] In the formula: out i This represents the enhanced features of the i-th layer;
[0032] In the multi-dimensional cross-layer feature fusion enhancement module, the visible light image features of the previous layer are respectively... and infrared image features Bilinear interpolation is used for upsampling to unify the image sizes of the two layers. Then, the visible light and infrared image features from the two layers are linearly added together. The two-dimensional convolution operation is repeated twice, followed by another upsampling. Finally, the features obtained after processing the visible light and infrared branches are linearly added together to obtain the final image.
[0033] The feature output obtained after enhancing the previous layer i+1 ∈R C×2H×2W Perform an upsampling operation to make it consistent with Given the same dimensions, linearly add the two features together, then apply a sigmoid activation function with channel attention to obtain the self-trained feature map F. s The specific calculation is shown in equation (5):
[0034]
[0035] In the formula: l b This represents a two-dimensional convolutional layer; CA(·) is the spatial attention module, and σ is the sigmoid activation function.
[0036] Combined with F s 'and out i+1 The feature map of this layer can be calculated, as shown in formula (6):
[0037]
[0038] In the formula: out i ∈RC×4H×4W This indicates the output feature result of this layer;
[0039] When dealing with low-level features When performing feature fusion, the two are combined with the feature output of the previous layer for fusion, and then the attention-assisted module (AAM) is used for feature enhancement. This process can be represented by formula (7):
[0040]
[0041] In the formula: AAM(·) is the attention assist module.
[0042] Further, in step S4, the decoding module includes:
[0043] The decoding module is divided into 5 layers for progressive decoding. In each layer, regularization is first performed with parameter p = 0.1, followed by two convolution operations. The first convolution uses dilated convolution with a kernel size of 3×3 and a dilation parameter of 3. This operation halves the number of feature channels in the output feature while keeping the length and width unchanged. The second convolution is a normal convolution operation with a kernel size of 3×3 and a dilation parameter of 1. This operation keeps the number of feature channels and size of the output feature unchanged. Finally, bilinear interpolation is used to upsample the feature map, doubling the length and width of the feature.
[0044] Further, in step S4, the multi-stage loss function includes:
[0045] The multi-stage loss function consists of three parts, the first part being semantic supervision. The second part is superpixel segmentation supervision. The third part consists of two supervisions added during the training of the decoder: edge supervision. and semantic supervision Therefore, the total loss function is calculated as shown in formula (8):
[0046]
[0047] In the formula: α, β, δ are the proportion coefficients of each part in the total error, and their magnitudes depend on the importance of each part to the overall model;
[0048] For semantic supervision The weighted cross-entropy loss function is adopted, as shown in formula (9):
[0049]
[0050] In the formula: h is the height of the image; w is the width of the image; x ijw(x) represents the coordinates of a point in the image. ij p(x) represents the weight of the category represented by that coordinate point; ij q(x) is a binary number representing whether the predicted value is true or false; ij ) represents the probability of predicting it as the target category;
[0051] Superpixel segmentation supervision As shown in formula (10):
[0052]
[0053] Where: SSM bsmapi Represents a binary segmentation map, SSM gti This represents the ground truth map for superpixel segmentation, and the SSM is selected based on the spsmap. gti The value of SSM is taken when the pixels segmented in the superpixel segmentation are foreground pixels. gti The value is 1, and when the pixels segmented in the superpixel segmentation are the background, the SSM is taken. gti The value is 0;
[0054] For edge supervision Using the binary cross-entropy loss method for semantic supervision A hybrid loss function is adopted, including the weighted cross-entropy loss function and the Lovasz-Softmax loss function, as shown in equations (11) and (12):
[0055]
[0056]
[0057] In the formula: G edge For edge truth graphs; E pre E pre This represents the edge prediction result output during the decoding process; wbce (·) represents the weighted cross-entropy loss function; Lovasz (·) represents the Lovasz-Softmax loss function; This is the final semantic segmentation map obtained after decoding; G sem This is the truth graph for semantic segmentation.
[0058] The beneficial effects of the technical solution provided by this invention are:
[0059] (1) The semantic segmentation method model for power equipment based on the feature fusion of visible light images and infrared images proposed in this invention makes full use of the complementary advantages of visible light images and infrared images, making the semantic fusion between visible light images and infrared images more complete and obtaining more comprehensive and accurate semantic information.
[0060] (2) The feature fusion enhancement module proposed in this invention combines superpixel segmentation results to guide the fusion of infrared and visible light image features, enhances foreground information and suppresses background information during the fusion process, which enables the network to better fuse high-level features and more accurately capture the structural information in the image, thereby obtaining more accurate foreground and background segmentation results.
[0061] (3) The multi-dimensional cross-layer feature fusion enhancement module proposed in this invention combines high-level features to fuse and enhance low- and medium-level features, enabling the network to capture more detailed information.
[0062] (4) The present invention uses a multi-stage loss function to supervise the network, thereby improving the accuracy of network model segmentation. Attached Figure Description
[0063] Figure 1 This is a flowchart of a semantic segmentation method for power equipment based on the fusion of features from visible light and infrared images;
[0064] Figure 2 A schematic diagram of a multi-dimensional cross-layer feature fusion enhancement module;
[0065] Figure 3 This is a schematic diagram of the decoding module;
[0066] Figure 4 This is a flowchart of the overall process of a semantic segmentation network for power equipment based on the fusion of features from visible light and infrared images. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0068] Please refer to Figure 1 This invention provides a semantic segmentation method for power equipment based on the fusion of visible light and infrared image features. The specific process is as follows:
[0069] Step S1: Obtain a dataset containing image pairs of visible light and infrared substation power equipment. The specific steps are as follows:
[0070] Step S11: Using visible light cameras and infrared cameras, acquire visible light and infrared registered images of substation power equipment against a complex background from different angles and at different times.
[0071] Step S12: Manually annotate the visible light and infrared image pairs obtained in step S11 to obtain the corresponding power equipment true value map.
[0072] Step S13: Combine the image information obtained in steps S11 and S12 to obtain the required dataset.
[0073] Step S2: Use the encoding module to extract features from the visible light image and the infrared image. The specific steps are as follows:
[0074] Due to the modal differences between visible light and infrared images, a dual-encoding module composed of two independent ResNet152 residual networks is used to receive both images. The visible light and infrared images are input into the two residual networks, and feature extraction is performed through convolution and pooling operations to obtain the visible light image features. Features of infrared images Here, 'i' represents the feature layer number. As the feature extraction layers increase, the extraction progresses from low-level to high-level features. Low-level features mainly include image color, texture, and shape, while high-level features contain richer semantic information. This paper divides the output features of the encoding module into 5 layers for downsampling, reducing the spatial resolution of each layer while doubling the number of channels. The extracted visible light image features and infrared image features are respectively... and
[0075] Step S3: Perform feature fusion and feature enhancement on the extracted visible light and infrared features. The specific steps are as follows:
[0076] Step S31: Based on the complementary advantages of the visible light image and the infrared image, their high-level features are fused, the key information in the feature map is located using the spatial attention mechanism, and the feature map is enhanced using the spatial attention weights of each modality image, as shown in formula (1):
[0077]
[0078] In the formula: Indicates the characteristics of visible light; Indicates infrared features; i indicates which layer of features; This indicates element-wise multiplication. Indicates element-wise addition; F fuse Indicates the features after fusion; f conv (·) denotes a one-dimensional convolutional layer; SA(·) denotes a spatial attention mechanism; f cat (·) indicates a channel-level connection.
[0079] Secondly, an iterative edge refinement method is used to perform superpixel segmentation on the infrared image, and the superpixel weight of each segmented region is calculated to enhance the correlation between adjacent pixel features in the fusion of visible light and infrared features, thereby obtaining a superpixel segmentation map (spsmap) of the infrared image and forming a superpixel mask. Where N S The superpixel mask for a single channel is SSM (Short Scale Array). i ∈R 1×H×W i = 1, 2, ..., N S .
[0080] A prototype O is built based on each superpixel mask. i And concatenate each prototype to generate a corresponding prototype block. Generate prototype sampler vectors from prototype blocks Then multiply p by SSM to generate the superpixel-assisted prediction map SSM. pred This process can be represented by formula (2):
[0081]
[0082] In the formula: APO(·) represents the average pooling operator of the superpixel mask; MLP(·) represents the multilayer sensing module; σ is the sigmoid activation function.
[0083] To better utilize the auxiliary effects of each channel in the superpixel mask, the weight of each superpixel is calculated based on its position and feature information to fuse the features of different channels. At the same time, the superpixel features are used to assist in the fusion of visible light and infrared images, making the network pay more attention to the target foreground information and suppressing background features, thereby improving the image clarity and recognition effect, as shown in formula (3):
[0084]
[0085] In the formula: This indicates that bilinear interpolation is used for upsampling; concat(·) is the channel concatenation operator; F f ' use This indicates the fusion features of the enhanced visible light image and infrared image.
[0086] Step S32: Use the Multi-dimensional cross-layer module (MDCM) to enhance visible light features and infrared image features, and combine the features of each modality to perform cross-modal fusion.
[0087] When dealing with intermediate features When performing feature fusion, the bilinear interpolation method is first used for upsampling, and then cross-modal enhancement between visible light features and infrared features is performed. The visible light features, infrared features and the enhanced features of the previous layer are aggregated and output, as shown in formula (4):
[0088]
[0089] In the formula: out i This represents the enhanced features of the i-th layer.
[0090] In the multi-dimensional cross-layer feature fusion enhancement module, the visible light image features of the previous layer are respectively... and infrared image features Bilinear interpolation is used for upsampling to unify the image sizes of the two layers. Then, the visible light and infrared image features from the two layers are linearly added together, and the two-dimensional convolution operation is repeated twice, followed by another upsampling. The features obtained after processing the visible light and infrared branches are then linearly added together to obtain the final image size.
[0091] The feature output obtained after enhancing the previous layer i+1 ∈R C×2H×2W Perform an upsampling operation to make it consistent with Given the same dimensions, linearly add the two features together, then apply a sigmoid activation function with channel attention to obtain the self-trained feature map F. s The specific calculation is shown in equation (5):
[0092]
[0093] In the formula: l b This represents a two-dimensional convolutional layer; CA(·) is the spatial attention module, and σ is the sigmoid activation function.
[0094] Combined with F s 'and out i+1 The feature map of this layer can be calculated, as shown in formula (6):
[0095]
[0096] In the formula: out i ∈R C×4H×4W This indicates the output feature result of this layer.
[0097] When dealing with low-level features When performing feature fusion, the two are combined with the feature output of the previous layer for fusion, and then combined with the Attention Assistance Module (AAM) for feature enhancement. This process can be represented by formula (7):
[0098]
[0099] In the formula: AAM(·) is the attention assist module.
[0100] Step S4: Use the decoding module to perform decoding, and combine it with a multi-stage loss function to output accurate semantic segmentation results for power equipment. The specific steps are as follows:
[0101] Step S41: Perform step-by-step decoding using the decoding module. The decoding module consists of 5 layers. In each layer, regularization is first performed, with parameter p = 0.1. Then, two convolution operations are performed. The first convolution uses dilated convolution, employing a 3×3 kernel layer with a dilation parameter of 3. After this operation, the feature channels of the output feature are halved, while the length and width remain unchanged. The second convolution is a normal convolution operation, using a 3×3 kernel layer with a dilation parameter of 1. After this operation, the feature channels and size of the output feature remain unchanged. Finally, bilinear interpolation is used to upsample the feature map, doubling its length and width.
[0102] Step S42: Calculate the error between the predicted result and the true label in semantic segmentation using a multi-stage loss function. The multi-stage loss function consists of three parts, the first part being semantic supervision. The second part is superpixel segmentation supervision. The third part consists of two supervisions added during the training of the decoder: edge supervision. and semantic supervision Therefore, the total loss function is calculated as shown in formula (8):
[0103]
[0104] In the formula: α, β, δ are the proportion coefficients of each part in the total error, and their magnitudes depend on the importance of each part to the overall model.
[0105] For semantic supervision The weighted cross-entropy loss function is adopted, as shown in formula (9):
[0106]
[0107] In the formula: h is the height of the image; w is the width of the image; x ij w(x) represents the coordinates of a point in the image. ij p(x) represents the weight of the category represented by that coordinate point; ij q(x) is a binary number representing whether the predicted value is true or false; ij ) represents the probability of predicting the target category.
[0108] Superpixel segmentation supervision As shown in formula (10):
[0109]
[0110] Where: SSM bsmapi Represents a binary segmentation map, SSM gti This represents the ground truth map for superpixel segmentation, and the SSM is selected based on the spsmap. gti The value of SSM is taken when the pixels segmented in the superpixel segmentation are foreground pixels. gti The value is 1, and when the pixels segmented in the superpixel segmentation are the background, the SSM is taken. gti The value is 0.
[0111] For edge supervision Using the binary cross-entropy loss method for semantic supervision A hybrid loss function is adopted, including the weighted cross-entropy loss function and the Lovasz-Softmax loss function, as shown in equations (11) and (12):
[0112]
[0113]
[0114] In the formula: G edge For edge truth graphs; E pre E pre This represents the edge prediction result output during the decoding process; wbce (·) represents the weighted cross-entropy loss function; Lovasz (·) represents the Lovasz-Softmax loss function; This is the final semantic segmentation map obtained after decoding; G sem This is the truth graph for semantic segmentation.
[0115] By utilizing the obtained error and combining it with a multi-stage loss function, the error between the model's prediction information and the actual data is minimized, so as to output a highly accurate semantic segmentation result for power equipment.
[0116] The above description is merely a specific embodiment of the present invention, and the scope of protection of the present invention is not limited thereto. Anyone skilled in the art can exercise their rights within the scope of the technology disclosed in the present invention.
Claims
1. A semantic segmentation method for power equipment based on the fusion of visible light and infrared image features, characterized in that, Includes the following steps: Step S1: Obtain a dataset containing pairs of visible light and infrared images of substation power equipment; Step S2: Use the encoding module to extract features from the visible light image and the infrared image; The encoding module consists of two independent ResNet152 residual networks forming a dual-channel encoding module, which receives visible light and infrared images. The visible light and infrared images are input into the two residual networks, and feature extraction is performed through convolution and pooling operations to obtain the visible light image features. Features of infrared images Where i represents the feature layer number; as the feature extraction layers deepen, the extraction progresses from low-level features to high-level features; the output features of the encoding module are divided into 5 layers for downsampling, reducing the spatial resolution of each layer and doubling the number of channels. The extracted visible light image features and infrared image features are respectively... and ; Step S3: Perform feature fusion and feature enhancement on the extracted visible light features and infrared features. After fusing and enhancing the high-level features using the feature fusion and enhancement module combined with superpixel segmentation, combine the semantic information in the high-level features and use the multi-dimensional cross-layer feature fusion and enhancement module to fuse and enhance the mid- and low-level features. The feature fusion enhancement module combined with superpixel segmentation includes: Based on the complementary advantages of visible light and infrared images, their high-level features are fused, key information in the feature map is located using a spatial attention mechanism, and the feature map is enhanced using the spatial attention weights of each modality, as shown in formula (1): (1); In the formula: Indicates the characteristics of visible light; Indicates infrared features; i indicates which layer of features; This indicates element-wise multiplication. This indicates element addition; Indicates the characteristics after fusion; This represents a one-dimensional convolutional layer; This represents the spatial attention mechanism; Indicates channel-level connection; Secondly, an iterative edge thinning method is used to perform superpixel segmentation on the infrared image, and the superpixel weights of each segmented region are calculated to obtain the superpixel segmentation map of the infrared image. And form a superpixel mask. ,in Where is the number of channels, and the superpixel mask for a single channel is . ; A prototype is built based on each superpixel mask. And concatenate each prototype to generate a corresponding prototype block. Prototype sampler vectors are generated from prototype blocks. Then and Multiply to generate a superpixel-assisted prediction map. This process can be represented by formula (2): (2); In the formula: This represents the average pooling operator for the superpixel mask; This represents a multi-layer sensing module; It is the sigmoid activation function; The weight of each superpixel is calculated by using the position and feature information of the superpixel to fuse the features of different channels. At the same time, the superpixel features are used to assist the fusion features of visible light images and infrared images, so that the network pays more attention to the foreground information of the target and suppresses the background features, thereby improving the image clarity and recognition effect, as shown in formula (3): (3); In the formula: This indicates that the upsampling operation is performed using the bilinear interpolation method; Channel connection operator; This indicates the fusion features of the enhanced visible light image and infrared image; The multi-dimensional cross-layer feature fusion enhancement module includes: The Multi-dimensional Cross-Layer Module (MDCM) is used to enhance visible light and infrared image features and combine features from each modality for cross-modal fusion. When dealing with intermediate features When performing feature fusion for i=2,3,4, firstly, the bilinear interpolation method is used for upsampling, and then cross-modal enhancement between visible light features and infrared features is performed. The visible light features, infrared features, and the enhanced features of the previous layer are aggregated and output, as shown in formula (4): (4); In the formula: This represents the enhanced features of the i-th layer; In the multi-dimensional cross-layer feature fusion enhancement module, the visible light image features of the previous layer are respectively... and infrared image features Bilinear interpolation is used for upsampling to unify the image sizes of the two layers. Then, the visible light and infrared image features from the two layers are linearly added together. The two-dimensional convolution operation is repeated twice, followed by another upsampling. Finally, the features obtained after processing the visible light and infrared branches are linearly added together to obtain the final image. ; Features obtained after enhancing the previous layer Perform an upsampling operation to make it consistent with Given the same dimensions, the two are linearly added together, and then a sigmoid activation function with channel attention is used to obtain a self-trained feature map. For specific calculations, see equation (5): (5); In the formula: This represents a two-dimensional convolutional layer; For spatial attention modules, It is the sigmoid activation function; Combination and The feature result map of this layer can be calculated, as shown in formula (6): (6); In the formula: This indicates the output feature result of this layer; When dealing with low-level features When performing feature fusion, the two are combined with the feature output of the previous layer for fusion, and then the attention-assisted module (AAM) is used for feature enhancement. This process can be represented by formula (7): (7); In the formula: For attention assist module; Step S4: Use the decoding module to perform decoding, and combine it with a multi-stage loss function to output accurate semantic segmentation results for power equipment.
2. The semantic segmentation method for power equipment based on visible light and infrared image feature fusion according to claim 1, characterized in that, The steps for obtaining the dataset in step S1 are as follows: Step S11: Using visible light cameras and infrared cameras, acquire visible light and infrared registered images of substation power equipment against a complex background from different angles and at different times. Step S12: Manually annotate the visible light and infrared image pairs obtained in step S11 to obtain the corresponding power equipment true value map; Step S13: Combine the image information obtained in steps S11 and S12 to obtain the required dataset.
3. The semantic segmentation method for power equipment based on visible light and infrared image feature fusion according to claim 1, characterized in that, In step S4, the decoding module includes: The decoding module is divided into 5 layers for progressive decoding. At each layer, regularization is performed first, and parameters are set. Then, two more convolution operations are performed. The first convolution uses dilated convolution, with a 3×3 kernel and a dilation parameter of 3. After this operation, the number of feature channels in the output feature is halved, while the length and width remain unchanged. The second convolution is a normal convolution, using a 3×3 kernel and a dilation parameter of 1. After this operation, the number of feature channels in the output feature remains unchanged, and the size remains unchanged. Finally, bilinear interpolation is used to upsample the feature map, making the feature length and width twice as large as before.
4. The semantic segmentation method for power equipment based on visible light and infrared image feature fusion according to claim 1, characterized in that, In step S4, the multi-stage loss function includes: The multi-stage loss function consists of three parts, the first part being semantic supervision. The second part is superpixel segmentation supervision. The third part consists of two supervisions added during the training of the decoder: edge supervision. and semantic supervision Therefore, the total loss function is calculated as shown in formula (8): (8); In the formula: This is the proportion of each part in the total error, and its magnitude depends on the importance of each part to the overall model; For semantic supervision The weighted cross-entropy loss function is used, as shown in formula (9): (9); In the formula: h is the height of the image; w is the width of the image; These are the coordinates of a point in the image. The weight of the category represented by that coordinate point; It is a binary number representing whether the predicted value is true or false; This indicates the probability of predicting the target category; Superpixel segmentation supervision As shown in formula (10): (10); In the formula: Represents a binary segmentation image. Represents the ground truth map of superpixel segmentation, based on choose The value is taken when the pixels segmented in the superpixel segmentation are foreground pixels. The value is 1, which is taken when the pixels segmented in the superpixel segmentation are the background. The value is 0; For edge supervision The binary cross-entropy loss method is used for semantic supervision. A hybrid loss function is used, including weighted cross-entropy loss function and... The loss function is expressed as shown in equations (11) and (12): (11); (12); In the formula: For edge truth maps; This represents the edge prediction result output during the decoding process; The weighted cross-entropy loss function; express Loss function; This is the final semantic segmentation map obtained after decoding; This is the truth graph for semantic segmentation.