Intelligent substation inspection segmentation method based on balanced memory matching
By adopting the balanced memory matching mechanism and pixel-level balanced matching method in the intelligent inspection of substations, the problem of low target segmentation accuracy caused by complex image background is solved, and the precise segmentation of power equipment and high accuracy of substation inspection is achieved.
Patent Information
- Application Number
- CN202510269019.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-13
AI Technical Summary
During the intelligent inspection of substations, factors such as complex image background, uneven lighting, and occlusion have caused low accuracy in target segmentation.
The segmentation method based on balanced memory matching is adopted, and the memory database is used to collect the historical information of the target through the equalized memory matching mechanism. The pixel-level balanced matching method is used to ensure reliable transmission of reference frame information, obtain the best match between the reference frame and the current frame, enhance the target representation and suppress the influence of background clutter.
It realizes accurate segmentation of power equipment and improves the accuracy of intelligent inspection of substations.
Smart Images

Figure CN120147931A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of power grid operation and maintenance, and particularly to a substation intelligent inspection segmentation method based on balanced memory matching. Background Art
[0002] During the process of substation intelligent inspection, it is possible to determine whether the equipment needs maintenance by viewing the images of substation equipment. However, the background of the substation is relatively complex, the light distribution is uneven, the change range is large, and the images are prone to occlusion, resulting in low accuracy of target segmentation when processing substation images subsequently.
[0003] Therefore, a method capable of accurately segmenting substation intelligent inspection images is needed. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, this application provides a substation intelligent inspection segmentation method based on balanced memory matching. In the embodiments of this application, through a balanced memory matching mechanism, the memory bank is used to collect historical information of the target, and a pixel-level balanced matching method is adopted to ensure the reliable transmission of reference frame information to the current frame, obtain the best match between the reference frame and the current frame to enhance the target representation, and also suppress the adverse effects of background clutter, realizing the precise segmentation of power equipment, and further improving the accuracy of substation intelligent inspection.
[0005] The above application purpose of this application is achieved through the following technical solutions: A substation intelligent inspection segmentation method based on balanced memory matching includes the following steps: In response to the acquired substation inspection video, a data set is made according to the substation inspection video; Determine the power equipment to be segmented and the corresponding mask annotation; Extract the depth features of the current query frame image including the power equipment to be segmented and its background area, and extract the depth features of the stored historical reference frame image; Perform downsampling operation on the stored historical reference frame mask; Use the pixel-level balanced matching method to compare the query frame features and the reference frame features to obtain an enhanced composite feature map; Decode the composite feature map to predict the query frame segmentation result map.
[0006] Optionally, the depth features of the current query frame image including the power equipment to be segmented and its background area, and the depth features of the stored historical reference frame image are both extracted by a pre-trained ResNet50 convolutional neural network model.
[0007] Optionally, according to the predicted current query frame image and the segmentation result, calculate the current query frame mask quality score; Calculate the quality scores of each reference frame in the current memory bank; Compare the quality score of the current query frame mask with the quality scores of each reference frame mask in the memory bank; If the quality score of the current query frame mask is higher than the reference frame with the lowest mask quality score in the memory bank, remove the reference frame with the lowest mask quality score in the memory bank and add the current query frame to the memory bank.
[0008] Optionally, the calculation of the quality score of the current query frame mask includes: Input the predicted current frame mask and the image into the mask quality evaluation module, which consists of a score encoder, four convolutional layers and two fully connected layers; Obtain the quality score of the query frame mask according to the following process : where, denotes the concatenation operation, i represents the index of the current query frame, denotes the query frame image, denotes the predicted mask of the current frame, denotes the score feature map, , and respectively denote the score encoder, convolution and fully connected layer with Sigmoid non-linear function.
[0009] Optionally, the mask quality evaluation module defines the value of the quality score as the mask overlap degree between the segmentation mask and the ground truth annotation; The loss function of the training process of the mask quality evaluation module is: where, denotes the quality score of the segmentation result of the -th frame target, IOU denotes the mask overlap degree, denotes the segmentation result of the current frame, denotes the ground truth annotation, denotes the total number of targets.
[0010] Optionally, it further includes the following steps: Calculate the temporal consistency score of the reference frame: where k and i respectively denote the indices of the reference frame and the current frame; Calculate the total score of each reference frame in the memory bank: Compare the total scores of each reference frame in the memory bank, and remove the reference frame with the lowest total score in the memory bank.
[0011] Optionally, obtaining the enhanced composite feature map includes the following steps: Calculate the feature similarity between the reference frame and the query frame to obtain a similarity matrix; Multiply the similarity matrix by the reference frame mask after downsampling operation to obtain enhanced target background channel and foreground channel information; Perform query maximization operation on the enhanced background and foreground information to obtain the matching scores of the background and foreground; Concatenate the matching scores of the background and foreground along the channel to form a matching score map; Connect the matching score map with the depth feature of the query frame to obtain a composite feature map.
[0012] Optionally, the calculating the feature similarity between the reference frame and the query frame to obtain a similarity matrix includes the following steps: Calculate the feature similarity between the reference frame and the query frame: where p and q are the positions of each spatial pixel in the reference frame and query frame features respectively, is the reference frame feature, is the query frame feature, is the query frame index, is the reference frame index; Perform a normalization operation along the query dimension to obtain a similarity matrix .
[0013] In summary, the present application has the following beneficial technical effects: In the embodiment of the present application, through the balanced memory matching mechanism, the historical information of the target is collected by using the memory bank, and the pixel-level balanced matching method is adopted to ensure the reliable transmission of the reference frame information to the current frame, obtain the best match between the reference frame and the current frame to enhance the target representation, and also suppress the adverse effects of background clutter, realizing the accurate segmentation of power equipment, thereby improving the accuracy of intelligent substation patrol inspection. Brief Description of the Drawings
[0014] Figure 1 is a schematic flowchart of an embodiment of the present application; Figure 2 is a schematic flowchart of the pixel-level balanced matching method in an embodiment of the present application; Figure 3 is a schematic diagram of the mask quality evaluation model in an embodiment of the present application. Detailed Embodiment
[0015] The following is a further detailed description of this application in conjunction with the attached Figure 1 - attached Figure 3 drawings.
[0016] To better understand the technical solutions presented in the embodiments of this application, a brief introduction to existing substation inspections is provided first.
[0017] During the intelligent inspection process in a substation scenario, inspection technologies based on drones and robots can use computer vision technology to identify appearance defects of substation equipment (such as transformers, insulators, and meters), thus greatly reducing the risks for workers in high-risk environments such as substations. At the same time, the inspection quality and efficiency in dangerous environments have also been improved. It requires the identification and segmentation of substation power equipment targets to achieve real-time monitoring and analysis of equipment status. Therefore, researching accurate segmentation algorithms is of great significance for substation intelligent inspection.
[0018] However, for power equipment targets in substation inspection videos, due to interference from factors such as complex backgrounds, uneven light distribution, large variation ranges, and occlusion, it will have a negative impact on the accuracy, stability, and efficiency of power target segmentation. However, traditional matching methods do not consider the best match between the query frame and the reference frame, resulting in some reference frame pixels not being referenced and some pixels being referenced multiple times, making them vulnerable to background interference. Another typical matching method, namely the non-local matching method, can capture the subtle differences between adjacent pixels, but it does not consider the global structure and context information of the overall image, which to a certain extent limits its ability to process the overall similarity of images.
[0019] To solve the above technical problems, the embodiments of this application provide a substation intelligent inspection segmentation method based on balanced memory matching, including the following steps: S101. In response to the acquired substation inspection video, a dataset is made according to the substation inspection video; S102. Determine the power equipment to be segmented and the corresponding mask annotations; S103. Extract the depth features of the current query frame image containing the power equipment to be segmented and its background area, and extract the depth features of the stored historical reference frame images; S104. Perform a downsampling operation on the stored historical reference frame masks; S105. Use the pixel-level balanced matching method to compare the query frame features and the reference frame features to obtain an enhanced composite feature map; S106. Decode the composite feature map to predict the query frame segmentation result map.
[0020] The following is a further introduction in combination with specific usage scenarios.
[0021] First, step S101 is executed to obtain the substation inspection video, and a dataset is made according to the substation inspection video. The dataset includes electrical equipment scenarios such as instruments, insulators, and liquid level gauges captured by inspection robots or drones. Next, step S102 is executed to determine the power equipment to be segmented and the corresponding mask annotations, preparing for subsequent extraction. In step S103, the current query frame image is extracted to include the depth features of the power equipment to be segmented and its background area , and the depth features of the stored historical reference frame images are extracted. , where i refers to the query frame, refers to the reference frame. Exemplarily, the feature extraction of the query frame image and the historical reference frame images can be performed through the pre-trained convolutional neural network model ResNet50. Next, step S104 is executed to perform downsampling on the stored historical reference frame mask to obtain a mask that includes a background channel and a foreground channel. Continue to execute step S105, and use the pixel-level equalization matching method to compare the query frame features and the reference frame features to obtain an enhanced composite feature map. Finally, step S106 is executed to decode the composite feature map to predict the query frame segmentation result map. During the process of generating the segmentation mask, the low-resolution features in the query frame depth features are continuously merged to generate a more accurate segmentation mask.
[0022] Generally speaking, in the embodiment of this application, through the equalization memory matching mechanism, the historical information of the target is collected by using the memory bank, and the pixel-level equalization matching method is adopted to ensure the reliable transmission of the reference frame information to the current frame, obtain the best match between the reference frame and the current frame to enhance the target representation, and also suppress the adverse effects of background clutter, realizing the accurate segmentation of power equipment, and thus improving the accuracy of substation intelligent inspection.
[0023] In some possible implementation manners of the embodiment of this application, obtaining the enhanced composite feature map includes the following steps: Step S401: To explore the relationship between the reference frame and the query frame, that is, to obtain the best matching degree between them, the reference frame mask is passed to the query frame, and the query frame mask is predicted. First, the feature similarity between the reference frame and the query frame is calculated to obtain a similarity matrix: where p and q are the positions of each spatial pixel in the reference frame and query frame features respectively, is a reference frame feature, is a query frame feature, is a query frame index, is a reference frame index; Perform a normalization operation along the query dimension to obtain a similarity matrix , in the embodiments of the present application, the Softmax operation can be used for the normalization operation: According to step S401, for each pixel in the reference frame, the scores of all pixels in the query frame are normalized so that their sum becomes 1, which ensures that all reference frame pixels contribute equally to the query frame prediction. Therefore, if a reference frame pixel is referenced multiple times, its score will be reduced proportionally because they must share the same total score; Step S402: Multiply the similarity matrix with the reference frame mask after downsampling operation : To obtain enhanced target background channel and foreground channel information, where represents the Hadamard product, and represent the background channel and foreground channel of the downsampled historical frame mask respectively, refers to the reference frame; Step S403: Perform a query maximization operation on the enhanced background and foreground information to obtain the matching scores of the background and foreground; Step S404: Concatenate the matching scores of the background and foreground along the channel to form a matching score map. The final matching score of the i-th frame can be expressed as: ; Step S405: Connect the matching score map with the depth feature of the query frame to obtain a composite feature map .
[0024] After the above steps are completed, the composite feature map can be decoded to predict the query frame segmentation result map. During the process of generating the segmentation mask, the low-resolution features in the query frame depth feature are continuously merged to generate a more accurate segmentation mask.
[0025] As a feasible specific implementation manner of the embodiments of the present application, after generating the predicted segmentation result map, the following steps are further included: Step S107: According to the predicted current query frame image And the segmentation result , calculate the quality score of the current query frame mask. Specifically, the predicted current frame mask and the image are input into the mask quality evaluation module. The mask quality evaluation module consists of a score encoder, four convolutional layers, and two fully connected layers, and obtains the query frame mask quality score according to the following process : Among them, represents the concatenation operation, i represents the index of the current query frame, represents the query frame image, represents the predicted mask of the current frame, represents the score feature map, , and represent the score encoder, convolution, and fully connected layer with the Sigmoid non-linear function respectively; The mask quality evaluation module defines the value of the quality score as the mask overlap (IOU) between the segmentation mask and the ground truth annotation. The loss function in the training process of the mask quality evaluation module is: Among them, represents the quality score of the segmentation result of the th frame target, IOU represents the mask overlap, represents the segmentation result of the current frame, represents the ground truth annotation, represents the total number of targets; After obtaining the query frame mask quality score , take the average of the scores of all targets in a frame as the quality score of this frame: Among them, represents the total number of targets in the th frame, represents the segmentation result quality score of the th frame; Of course, in order to better measure the relativity of the segmentation results of all frames in the video, normalization processing is performed, which helps to remember more useful information in challenging scenarios. Specifically, the final quality score of each frame is its initial predicted score divided by the score of the first frame, which is expressed by the formula as follows: Among them, represents the segmentation result quality score of the th frame, Indicates the quality score of the segmentation result of the first frame; Step S108: Take a threshold and compare it with the final mask quality score to determine whether the final mask quality score is greater than the threshold. If the final mask quality score is greater than the threshold, then add this mask to the repository as a reference frame to implement dynamic update of the memory bank. In the embodiments of the present application, the threshold can take a value of , ensuring that the memory bank selectively adds those frames with mask quality scores greater than the threshold.
[0026] By introducing a mask quality evaluation module, the confidence of the current segmentation result is obtained, so as to selectively store frames with good segmentation quality, ensure the reliability of memory update, prevent the accumulation of segmentation template errors, achieve accurate segmentation of power equipment, and further improve the accuracy of intelligent substation inspection.
[0027] By storing historical frame information and evaluating mask quality, the best match between the reference frame and the current frame can also be obtained to enhance the target representation.
[0028] Step S108 can also be understood as storing frames with higher mask quality in the repository to implement dynamic update of the memory bank, so as to selectively store frames with good segmentation quality. On this basis, the following memory bank dynamic update method can also be adopted: Step S801: Calculate the quality score of each reference frame in the current memory bank , and the quality score of the reference frame can be the total score of the temporal consistency score and the accuracy score of the reference frame: The calculation method of the temporal consistency score of the reference frame is: where k and i represent the indices of the reference frame and the current frame respectively; Step S802: Compare the mask quality score of the current query frame with the mask quality scores of each reference frame in the memory bank; Step S803: If the mask quality score of the current query frame is higher than the reference frame with the lowest mask quality score in the memory bank, then remove the reference frame with the lowest mask quality score in the memory bank and add the current query frame to the memory bank.
[0029] As the number of frames increases, the increase of memory frames will limit the practicality of this algorithm in real scenarios. Therefore, limiting the size of the memory bank and performing dynamic update can effectively prevent memory explosion. Taking the total score of the reference frame Based on this, the stored frame with the lowest score is removed to dynamically update the memory bank and prevent memory explosion.
[0030] The embodiments of this specific implementation manner are all preferred embodiments of this application, and do not limit the protection scope of this application accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of this application shall be covered within the protection scope of this application.
Claims
1. A substation intelligent inspection segmentation method based on balanced memory matching, characterized in that: include: In response to the acquired substation inspection video, a data set is generated according to the substation inspection video; Determine the power equipment that needs to be segmented and the corresponding mask annotations; Extracting the depth features of the current query frame image containing the power equipment to be segmented and its background area, and extracting the depth features of the stored historical reference frame image; Performing a downsampling operation on the stored historical reference frame mask; The query frame features and the reference frame features are compared using a pixel-level balanced matching method to obtain an enhanced composite feature map. The composite feature map is decoded and the query frame segmentation result map is predicted.
2. According to claim 1, a substation intelligent inspection segmentation method based on balanced memory matching is characterized in that: The current query frame image contains the depth features of the power equipment to be segmented and its background area, as well as the depth features of the stored historical reference frame images, which are extracted by the pre-trained ResNet50 convolutional neural network model.
3. According to claim 1, a substation intelligent inspection segmentation method based on balanced memory matching is characterized in that: Also includes: Calculate the mask quality score of the current query frame according to the predicted current query frame image and the segmentation result; Calculate the quality score of each reference frame in the current memory bank; Compare the current query frame mask quality score with each reference frame mask quality score in the memory bank; If the mask quality score of the current query frame is higher than the reference frame with the lowest mask quality score in the memory library, the reference frame with the lowest mask quality score in the memory library is removed and the current query frame is added to the memory library.
4. According to claim 3, a substation intelligent inspection segmentation method based on balanced memory matching is characterized in that: Calculating the mask quality score of the current query frame includes: The predicted current frame mask and images Input mask quality assessment module,The mask quality assessment module consists of a score encoder, four convolutional layers, and two fully connected layers; The query frame mask quality score is obtained according to the following process : in, represents the connection operation, i represents the index of the current query frame, represents the query frame image, represents the current frame prediction mask, represents the score feature map, , and They represent the fractional encoder, convolution, and fully connected layers with Sigmoid nonlinear function, respectively.
5. The method for intelligent inspection and segmentation of substations based on balanced memory matching according to claim 4 is characterized in that: The mask quality assessment module defines the value of the quality score as the mask overlap between the segmentation mask and the true annotation; The loss function of the mask quality assessment module training process is: in, Indicates The quality score of the segmentation result of the frame target, IOU represents the mask overlap, Indicates the segmentation result of the current frame. Indicates the true mark. Indicates the total number of targets.
6. The method for intelligent inspection and segmentation of substations based on balanced memory matching according to claim 5 is characterized in that: Also includes: Compute the temporal consistency score for the reference frame: Where k and i represent the indexes of the reference frame and the current frame respectively; Compute the total score for each reference frame in the memory bank: The total score of each reference frame in the memory bank is compared, and the reference frame with the lowest total score in the memory bank is removed.
7. The method for intelligent inspection and segmentation of substations based on balanced memory matching according to claim 1 is characterized in that: The enhanced composite feature map includes: Calculate the feature similarity between the reference frame and the query frame to obtain a similarity matrix; Multiply the similarity matrix with the reference frame mask after downsampling to obtain enhanced target background channel and foreground channel information; Perform query maximization operation on the enhanced background and foreground information to obtain the matching scores of the background and foreground; The matching scores of the background and foreground are concatenated along the channels to form a matching score map; The matching score map is concatenated with the deep features of the query frame to obtain a composite feature map.
8. The method for intelligent inspection and segmentation of substations based on balanced memory matching according to claim 7 is characterized in that: Calculating the feature similarity between the reference frame and the query frame to obtain a similarity matrix includes: Calculate the feature similarity between the reference frame and the query frame: Among them, p and q are the positions of each spatial pixel in the reference frame and query frame features, respectively. is the reference frame feature, is the query frame feature, is the query frame index, is the reference frame index; Perform normalization along the query dimension to obtain a similarity matrix .