Multi-scale salient target detection method for electromechanical equipment of hydroelectric generating set

Through the improved SE-ResNet-50 model and multi-scale feature fusion module, combined with the boundary refinement module, the problem of insufficient multi-scale feature processing and boundary identification of hydropower unit equipment in complex environments is solved, and a significant target detection effect with higher accuracy and robustness is achieved.

CN120107732APending Publication Date: 2025-06-06THREE GORGES JINSHAJIANG CHUANYUN HYDROPOWER DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510170447.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has shortcomings in the multi-scale feature processing and boundary identification of hydroelectric unit equipment, resulting in insufficient accuracy and robustness of significant object detection in complex environments.

Method used

The improved SE-ResNet-50 model is adopted, combined with the multi-scale feature fusion module and boundary refinement module, and a significant object detection model is constructed through steps such as feature extraction, multi-scale fusion and boundary refinement, and a joint loss function is designed for training.

Benefits of technology

It improves the ability to capture multi-scale features, enhances the recognition accuracy of significant target boundaries, overcomes the boundary blur problem, and improves the overall accuracy and robustness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107732A_ABST
    Figure CN120107732A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale salient target detection method for electromechanical equipment of a hydroelectric generating set, and relates to the technical field of computer vision and image processing. According to the invention, the capturing capability of the multi-scale features is enhanced through SE-ResNet, and the problem of improper processing of the multi-scale features in the prior art is effectively solved; and a boundary refining module is introduced, so that the boundary identification precision of the salient target is improved, the edge processing of complex equipment is more accurate, and the defect of boundary blur is overcome. Moreover, through a multi-scale fusion module, the problem of feature pollution caused by cross-layer feature fusion in the prior art is avoided, and the features of different levels can be effectively fused, so that the overall precision and robustness of salient target detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision technology and image processing technology, and in particular to a multi-scale salient target detection method for electromechanical equipment of a hydropower unit. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and may not constitute prior art.

[0003] Hydropower units are the core components of modern power generation systems. Their key electromechanical equipment includes generators, rotors, stators, bearings, etc. The operating conditions of these equipment directly affect the safety and efficiency of the units. Therefore, real-time monitoring and identification of significant components of these equipment are essential for preventing failures and optimizing operations. Existing monitoring methods mainly rely on traditional means such as vibration analysis, temperature monitoring, and infrared imaging. These methods show limitations in complex working environments and are difficult to meet the requirements of automation and high precision.

[0004] Salient target detection technology has been widely used in the field of computer vision in recent years. By automatically extracting important target areas in images, it can be applied to the identification and monitoring of hydropower equipment. However, existing technologies are mostly designed for natural scenes, and they show obvious deficiencies when directly applied to the detection of electromechanical equipment of hydropower units. Specifically, the multi-scale features of hydropower equipment are difficult to handle uniformly, and the size, shape and complex background environment of the equipment make it difficult for existing algorithms to deal with them effectively; at the same time, due to complex on-site conditions such as lighting, the recognition accuracy of the edge of the equipment is poor, resulting in blurred boundaries of the generated salient targets, and the inability to accurately locate the key parts of the equipment.

[0005] Therefore, when dealing with the detection of significant targets of equipment in complex environments, the existing technology shows problems such as insensitivity to multi-scale targets and inaccurate boundary processing. These shortcomings limit the application of technology in the field of hydropower unit monitoring. The existing significant target detection performance needs to be further improved to better adapt to the efficient and accurate identification of key components of electromechanical equipment in complex environments. Summary of the invention

[0006] The purpose of the present invention is to provide a multi-scale salient target detection method for electromechanical equipment of hydropower units, in view of the problems existing in the prior art of salient target detection of important electromechanical equipment of hydropower units, such as improper multi-scale feature processing and fuzzy boundary identification, so as to solve the above problems.

[0007] The technical solution of the present invention is as follows:

[0008] A multi-scale salient target detection method for electromechanical equipment of a hydropower unit comprises the following steps:

[0009] Step 1: Collect data through historical images of the hydropower plant site and web crawling, and randomly divide the collected data into training set, validation set, and test set;

[0010] Step 2: Annotate the pixel-level salient regions of the training set and test set samples obtained in step 1 to construct a salient target detection dataset for electromechanical equipment of hydropower units;

[0011] Step 3: Construct a salient object detection model consisting of a feature extraction stage, a multi-scale feature fusion stage, and a boundary refinement module, and design a joint loss function;

[0012] Step 4: Input the training set in step 2 into the salient object detection model for training, optimize the network weights through back propagation, obtain the optimal detection model, and use the test set for performance evaluation.

[0013] Furthermore, the step 1 comprises:

[0014] Step 1.1: Collect images covering multi-scale equipment components, complex boundaries and different lighting conditions, and generate an image sample library through data expansion;

[0015] Step 1.2: Preprocess the image, including denoising, enhancement and multi-scale processing.

[0016] Furthermore, the step 2 comprises:

[0017] Step 2.1: Using the LabelMe tool, start the software and load the dataset; select the image to be labeled and switch to the polygon tool to draw the salient area of ​​the target pixel by pixel;

[0018] Step 2.2: Draw a polygonal mask around the key parts of the device in the image, ensure that the annotated area covers the salient parts of the target, and mark them with corresponding labels;

[0019] Step 2.3: After the annotation is completed, save the annotation information as a mask file in JSON or PNG format; generate a salient object detection dataset containing the original image and the corresponding mask.

[0020] Furthermore, the feature extraction stage of step 3 adopts the improved SE-ResNet-50 as the backbone network, which specifically includes:

[0021] Introduce SE branches in each residual block, generate channel-weighted feature maps through global average pooling, full connection layer compression and channel dimension recovery;

[0022] The encoder outputs four levels of feature maps, E1-E4, where the low-level feature maps contain edge texture information and the high-level feature maps contain global semantic information.

[0023] Furthermore, the multi-scale feature fusion stage of step 3 includes 6 multi-scale fusion modules, each of which performs the following operations:

[0024] Perform 3×3 convolution and normalization on high-level and low-level feature maps respectively;

[0025] Perform bilinear interpolation upsampling on the low-level feature map to the same resolution as the high-level feature map;

[0026] The processed feature maps are concatenated in the channel dimension and a fused saliency map is generated by a 3×3 convolution.

[0027] Furthermore, the boundary refinement module of step 3 includes:

[0028] The E1-E4 feature maps are reduced in dimension by 1×1 convolution and then upsampled to a uniform resolution;

[0029] Generate a boundary map by adding and fusing adjacent layer feature maps layer by layer;

[0030] The boundary map is fused with the salient map through a 1×1 convolution to output a salient map containing the precise outline of the salient target.

[0031] Furthermore, the training of the boundary refinement module uses the boundary labels generated by Canny edge detection as the supervision truth value.

[0032] Furthermore, the joint loss function design of step 3 includes:

[0033] Design a joint loss function that combines cross entropy loss, intersection-over-union loss, and boundary loss to supervise the overall generation of saliency maps;

[0034] Cross entropy loss formula:

[0035]

[0036] Among them, G ij and P ij are the values ​​of the true image G and the predicted image P at position (i, j), H and W are the height and width of the input image respectively;

[0037] The intersection-over-union loss formula is:

[0038]

[0039] L BCE and L IoU The joint loss formula is:

[0040] L(P,G)=ω*L BCE (P,G)+(1-ω)*L Iou (P,G)

[0041] Among them, ω weight coefficient adjusts the proportion of the two losses;

[0042] Boundary loss formula:

[0043] L(S 1 ,G)=L BCE (S 1 ,G)+L_IoU(S 1 ,G)

[0044]

[0045] Among them, E is the target boundary label;

[0046] The joint loss formula is:

[0047]

[0048] Among them, i is a hyperparameter used to adjust the weight of multi-level supervision.

[0049] Furthermore, the step 4 comprises:

[0050] Step 4.1: Use the training set in step 2 to train the constructed salient object detection model, then use the validation set in step 2 to verify the training results, and continuously adjust the parameters of the improved network training through back propagation to select the optimal weight file. When the overall loss value reaches the minimum and converges, the network model reaches the optimal value.

[0051] Step 4.2: Input the test set in step 2 into the trained salient object detection model, and use the mean absolute error, F β The value and S value indicators are used to test the model performance.

[0052] Furthermore, the implementation of the SE branch includes:

[0053] Perform global average pooling on the feature map to generate channel descriptors;

[0054] Modeling inter-channel dependencies through two fully connected layers;

[0055] The channel weights are generated by the Sigmoid function and multiplied channel by channel with the original feature map.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] The present invention enhances the ability to capture multi-scale features through SE-ResNet, effectively solving the problem of improper multi-scale feature processing in the prior art; introduces a boundary refinement module to improve the accuracy of boundary recognition of significant targets, especially in the edge processing of complex devices, overcoming the deficiency of blurred boundaries. In addition, through the multi-scale fusion module, the feature contamination problem caused by cross-layer feature fusion in the prior art is avoided, ensuring that features at different levels can be effectively fused, thereby improving the overall accuracy and robustness of significant target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 Schematic diagram of a multi-scale salient object detection network of the present invention;

[0059] Figure 2 It is a schematic diagram of the multi-scale fusion module;

[0060] Figure 3 Schematic diagram of the boundary refinement module. DETAILED DESCRIPTION

[0061] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0062] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.

[0063] Embodiment 1

[0064] See also Figure 1-3 , a multi-scale salient target detection method for electromechanical equipment of a hydropower unit, comprising the following steps:

[0065] Step 1: Collect data through historical images of the hydropower plant site and web crawling, and randomly divide the collected data into training set, verification set, and test set; that is, collect data through historical images of the hydropower plant site and web crawling, and randomly divide the collected data into training set, verification set, and test set according to a certain ratio;

[0066] Step 2: Annotate the training set and test set samples obtained in step 1 at the pixel level to construct a salient target detection dataset for electromechanical equipment of hydropower units; that is, annotate the salient regions of the samples obtained in step 1 at the pixel level, annotate the key components of the equipment, such as generators, rotors, stators, bearings, etc., and construct a salient target detection dataset for electromechanical equipment of hydropower units;

[0067] Step 3: Construct a salient object detection model consisting of a feature extraction stage, a multi-scale feature fusion stage, and a boundary refinement module, and design a joint loss function;

[0068] Step 4: Input the training set in step 2 into the salient object detection model for training, optimize the network weights through back propagation, obtain the optimal detection model, and use the test set for performance evaluation.

[0069] In this embodiment, specifically, step 1 includes:

[0070] Step 1.1: Collect images covering multi-scale equipment components, complex boundaries and different lighting conditions, and generate an image sample library through data expansion; that is, collect on-site historical images of the hydropower plant and images related to the electromechanical equipment of the hydropower units on the Internet, focusing on multi-scale equipment components, complex boundaries, and image types under different lighting conditions and background noise interference, and form a real image sample library after data expansion;

[0071] Step 1.2: Preprocess the image, including denoising, enhancement and multi-scale processing; that is, preprocess the collected data, including image denoising, image enhancement and multi-scale processing, to improve the detection ability of the model in complex environments.

[0072] In this embodiment, specifically, step 2 includes:

[0073] Step 2.1: Use the LabelMe tool (saliency labeling tool), start the software and load the dataset; select the image to be labeled and switch to the Polygon Tool to draw the salient area of ​​the target pixel by pixel;

[0074] Step 2.2: Draw a polygonal mask around the key components of the device in the image, ensure that the annotated area covers the salient parts of the target, and mark it with the corresponding label; that is, draw a polygonal mask around the key components of the device (such as generator, rotor, stator, etc.) in the image, ensure that the annotated area covers the salient parts of the target, and mark it with the corresponding label, such as "generator", "rotor", "stator", etc.;

[0075] Step 2.3: After the annotation is completed, save the annotation information as a mask file in JSON or PNG format; generate a salient object detection dataset containing the original image and the corresponding mask; after the annotation is completed, save the annotation information as a mask file in JSON or PNG format, which can represent the salient object area. Make sure that all the annotated mask files correspond to the original image one by one, and generate a salient object detection dataset containing the original image and the corresponding mask.

[0076] In this embodiment, specifically, the feature extraction stage of step 3 uses the improved SE-ResNet-50 as the backbone network, which specifically includes:

[0077] Introduce SE branches in each residual block, generate channel-weighted feature maps through global average pooling, full connection layer compression and channel dimension recovery;

[0078] The encoder outputs four levels of feature maps, E1-E4, where the low-level feature maps contain edge texture information and the high-level feature maps contain global semantic information;

[0079] It should be noted that ResNet-50 consists of multiple residual blocks, each of which includes a combination of convolution, normalization (Batch Normalization) and activation function (ReLU), and solves the problem of gradient disappearance in deep networks through cross-layer skip connections;

[0080] Specifically, each residual block consists of a 1x1 convolution, a 3x3 convolution, and another 1x1 convolution, which are used for dimensionality reduction, processing spatial features, and restoring the number of channels, respectively. Among them, the first convolution layer of the encoder is a 7x7 convolution layer, which is used to extract features from the input image, followed by a maximum pooling layer (Max Pooling) to reduce the size of the feature map. Each subsequent residual block gradually reduces the spatial resolution of the feature map, but increases the number of channels. The encoder has a total of five stages of feature extraction, and each stage outputs feature maps of different levels E1, E2, E3, E4;

[0081] To further enhance the ability to capture key features, an SE branch is introduced in each residual block. The SE branch enables the network to automatically focus on more important features by adjusting the weight of each channel. First, the feature map of each channel is subjected to global average pooling (Global Average Pooling) to reduce each channel to a single value, representing the global information of the channel; the pooled result is compressed through a fully connected layer (FC) and then activated by ReLU; then the channel dimension is restored through another fully connected layer, and the weight value is generated through the sigmoid function; finally, these weights are multiplied by the feature map of each channel to achieve channel weighting;

[0082] During the feature extraction process, each feature map is downsampled through a convolution with a step size of 2 to reduce the resolution of the feature map. Feature maps of different resolutions represent different levels of information: lower-level feature maps focus on detail information (such as edges and textures), and higher-level feature maps capture global information and semantic features.

[0083] In this embodiment, specifically, the multi-scale feature fusion stage of step 3 includes 6 multi-scale fusion modules, and each multi-scale fusion module performs the following operations:

[0084] Perform 3×3 convolution and normalization on high-level and low-level feature maps respectively;

[0085] Perform bilinear interpolation upsampling on the low-level feature map to the same resolution as the high-level feature map;

[0086] The processed feature maps are concatenated in the channel dimension and fused into a saliency map through 3×3 convolution.

[0087] That is, the main task of the multi-scale feature fusion stage is to fuse the adjacent layer features extracted above layer by layer to avoid feature contamination caused by direct cross-layer fusion of features between different layers. The multi-scale feature fusion stage consists of multiple multi-scale fusion modules (MFM), which realize feature fusion through upsampling and convolution operations;

[0088] Specifically, the MFM module consists of the following key operations:

[0089] First, perform 3x3 convolution on the input high-level and low-level features respectively, and apply normalization and activation function (ReLU) to extract features; perform bilinear interpolation up-sampling (Up-Sample) on the low-level feature map to keep its resolution consistent with the high-level feature map; concatenate the processed high-level and low-level feature maps in the channel dimension; finally, further fuse these cascaded features through 3x3 convolution to obtain the fused feature map;

[0090] The feature fusion module is gradually fused according to the hierarchical structure, and the fused saliency map S1 is generated through 6 multi-scale fusion modules. Each module fuses the features of adjacent levels to generate gradually fused feature maps D1_i and D2_i. This ensures the model's ability to perceive objects at different scales. Through layer-by-layer fusion, the model can gradually restore the features extracted from the encoder and generate high-quality saliency maps.

[0091] In this embodiment, specifically, the boundary refinement module of step 3 includes:

[0092] The E1-E4 feature maps are reduced in dimension by 1×1 convolution and then upsampled to a uniform resolution;

[0093] Generate a boundary map by adding and fusing adjacent layer feature maps layer by layer;

[0094] The boundary map and the salient map are fused by 1×1 convolution to output a salient map containing the precise outline of the salient target;

[0095] The task of the Boundary Refinement Module (BRM) is to learn the contour information of salient targets, improve the model's perception and prediction capabilities of complex boundaries, generate clear target boundaries through boundary learning, and finally generate a salient map containing more accurate edge information;

[0096] It should be noted that the input of the boundary refinement module is the feature maps E1, E2, E3, and E4 of different levels mentioned above. These feature maps are first reduced in dimension through 1x1 convolution, and then upsampled to a uniform image resolution through bilinear interpolation; adjacent hierarchical feature maps are fused layer by layer, using pixel addition to ensure that the details of the boundary features are not lost;

[0097] The fused boundary features generate a boundary map, which is consistent with the salient features generated above. Figure 1 The convolution operation is further performed through 1x1 convolution, and the resulting saliency map contains the precise outline of the salient target. When training the model, the boundary labels generated by the Canny edge detection algorithm are used as the true value to supervise the training. The model gradually optimizes its learning of the boundary by comparing it with the boundary labels.

[0098] In this embodiment, specifically, the joint loss function design of step 3 includes:

[0099] Design a joint loss function that combines cross entropy loss, intersection-over-union loss, and boundary loss to supervise the overall generation of saliency maps;

[0100] Cross entropy loss formula:

[0101]

[0102] Among them, G ij and P ij are the values ​​of the true image G and the predicted image P at position (i, j), H and W are the height and width of the input image respectively;

[0103] The intersection-over-union loss formula is:

[0104]

[0105] L BCE and L IoU The joint loss formula is:

[0106] L(P,G)=ω*L BCE (P,G)+(1-ω)*L Iou (P,G)

[0107] Among them, ω weight coefficient adjusts the proportion of the two losses;

[0108] Boundary loss formula:

[0109] L(S 1 ,G)=L BCE (S 1 ,G)+L_IoU(S 1 ,G)

[0110]

[0111] Among them, E is the target boundary label;

[0112] Combined loss formula:

[0113]

[0114] Among them, i is a hyperparameter used to adjust the weight of multi-level supervision.

[0115] In this embodiment, specifically, step 4 includes:

[0116] Step 4.1: Use the training set in step 2 to train the constructed salient object detection model, then use the validation set in step 2 to verify the training results, and continuously adjust the parameters of the improved network training through back propagation to select the optimal weight file. When the overall loss value reaches the minimum and converges, the network model reaches the optimal value.

[0117] Step 4.2: Input the test set in step 2 into the trained salient object detection model, and use the mean absolute error, F β The value and S value indicators are used to test the model performance.

[0118] In this embodiment, specifically, the evaluation index evaluates and tests the model;

[0119] The mean absolute error (MAE) is used to calculate the average pixel error between the predicted image and the true image, as shown in the following formula. The smaller the MAE, the better the prediction effect:

[0120]

[0121] Where P(i,j) and G(i,j) represent the pixel values ​​of the predicted image P and the true image G at position (i,j) respectively;

[0122] F β The value is the weighted harmonic mean of precision and recall, and the formula is as follows:

[0123]

[0124] where β 2 Set to 0.3, Precision and Recall represent the precision and recall respectively, and the formula is as follows:

[0125]

[0126] Among them, TP, FP, and FN represent the number of pixels where the salient area is predicted as the salient area, the non-salient area is predicted as the salient area, and the salient area is predicted as the non-salient area, respectively;

[0127] The S index is used to evaluate the similarity between the predicted image and the true image, including the region-based structural similarity S r With target-based structural similarity S o The calculation method is shown in the following formula, where α is set to 0.5:

[0128] S=(1-α)×S r +α×S o

[0129] Embodiment 2

[0130] This example compares and verifies the detection effect of the model of the present invention on multiple public data sets. The specific implementation steps are as follows:

[0131] Step S1: Use the publicly available salient object detection datasets: ECSSD, PASCAL-S, and DUTS for experiments. ECSSD contains 1,000 natural scene images with complex backgrounds, PASCAL-S contains 850 images with multiple objects and occlusions, and DUTS comes from SUN and ImageNetDET, including DUTS-TR with 10,553 images and DUTS-TE with 5,019 images;

[0132] Step S2: This embodiment selects four methods for comparative experiments, including EANet, ICON-R, DCENe and the method of the present invention. The four models are trained on the data sets used in step 1 respectively.

[0133] Step S3: Evaluate the performance of the three models using the test set. The main indicators include mean absolute error (MAE), F β Value, S value. The test results are shown in Table 1:

[0134] Table 1 Performance comparison of different models

[0135]

[0136] From the experimental results, we can see that the proposed method performs well on the three datasets (ECSSD, PASCAL-S, DURS-TE), especially in MAE, F β The proposed method outperforms other comparison methods in key indicators such as the α value and the S value. On the ECSSD dataset, the MAE of the proposed method is 0.031, the MAE on the PASCAL-S dataset is 0.055, and the MAE on the DURS-TE dataset is 0.034, which are lower than other methods, indicating that the proposed method can capture salient targets more accurately. In addition, the proposed method has a better performance in F β The performance in terms of value and S value also shows strong overall detection ability, and the boundary processing effect is significantly better than the comparison model.

[0137] This embodiment enhances the capture of multi-scale features through SE-ResNet, so that when processing complex electromechanical equipment of hydropower units, it can better identify salient targets of different scales. The boundary refinement module further improves the detection accuracy of the boundary of salient targets, making the boundary of salient targets clearer and overcoming the boundary fuzziness problem in the prior art. In addition, the multi-scale fusion module effectively avoids the problem of feature contamination and ensures the effective fusion of features at different levels, thereby improving the overall detection accuracy and the robustness of the model.

[0138] Embodiment 3

[0139] This embodiment also proposes a multi-scale salient target detection device for electromechanical equipment of a hydropower unit, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned multi-scale salient target detection method for electromechanical equipment of a hydropower unit are implemented; preferably, the computer program can be executed on a terminal device, such as a personal computer.

[0140] This embodiment also proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the above-mentioned multi-scale salient target detection method for electromechanical equipment of a hydropower unit; however, the device of the present invention is not limited to this. In this document, the readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, apparatus or device.

[0141] The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0142] The computer readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by an instruction execution system, an apparatus, or a device or used in combination with it. The program code contained on the readable storage medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0143] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0144] The above-mentioned embodiments only express the specific implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the protection scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the technical solution concept of the present application, and these all belong to the protection scope of the present application.

[0145] This background section is provided to generally present the context of the invention, and the work of the presently named inventors, the work to the extent described in this background section, and aspects of this section that did not constitute prior art at the time of application are neither explicitly nor implicitly admitted to be prior art to the present invention.

Claims

1. A multi-scale salient target detection method for electromechanical equipment of hydropower units, characterized in that: The following steps are involved: Step 1: Collect data through historical images of the hydropower plant site and web crawling, and randomly divide the collected data into training set, validation set, and test set; Step 2: Annotate the pixel-level salient regions of the training set and test set samples obtained in step 1 to construct a salient target detection dataset for electromechanical equipment of hydropower units; Step 3: Construct a salient object detection model consisting of a feature extraction stage, a multi-scale feature fusion stage, and a boundary refinement module, and design a joint loss function; Step 4: Input the training set in step 2 into the salient object detection model for training, optimize the network weights through back propagation, obtain the optimal detection model, and use the test set for performance evaluation.

2. A multi-scale salient target detection method for electromechanical equipment of a hydropower unit according to claim 1, characterized in that: The step 1 comprises: Step 1.1: Collect images covering multi-scale equipment components, complex boundaries and different lighting conditions, and generate an image sample library through data expansion; Step 1.2: Preprocess the image, including denoising, enhancement and multi-scale processing.

3. A multi-scale salient target detection method for electromechanical equipment of a hydropower unit according to claim 2, characterized in that: The step 2 comprises: Step 2.1: Using the LabelMe tool, start the software and load the dataset; select the image to be labeled and switch to the polygon tool to draw the salient area of ​​the target pixel by pixel; Step 2.2: Draw a polygonal mask around the key parts of the device in the image, ensure that the annotated area covers the salient parts of the target, and mark them with corresponding labels; Step 2.3: After the annotation is completed, save the annotation information as a mask file in JSON or PNG format; generate a salient object detection dataset containing the original image and the corresponding mask.

4. A multi-scale salient target detection method for electromechanical equipment of a hydropower unit according to claim 3, characterized in that: The feature extraction stage of step 3 uses the improved SE-ResNet-50 as the backbone network, which specifically includes: Introduce SE branches in each residual block, generate channel-weighted feature maps through global average pooling, full connection layer compression and channel dimension recovery; The encoder outputs four levels of feature maps, E1-E4, where the low-level feature maps contain edge texture information and the high-level feature maps contain global semantic information.

5. A multi-scale salient target detection method for electromechanical equipment of a hydropower unit according to claim 4, characterized in that: The multi-scale feature fusion stage of step 3 includes 6 multi-scale fusion modules, each of which performs the following operations: Perform 3×3 convolution and normalization on high-level and low-level feature maps respectively; Perform bilinear interpolation upsampling on the low-level feature map to the same resolution as the high-level feature map; The processed feature maps are concatenated in the channel dimension and a fused saliency map is generated by a 3×3 convolution.

6. A multi-scale salient target detection method for electromechanical equipment of a hydropower unit according to claim 5, characterized in that: The boundary refinement module of step 3 includes: The E1-E4 feature maps are reduced in dimension by 1×1 convolution and then upsampled to a uniform resolution; Generate a boundary map by adding and fusing adjacent layer feature maps layer by layer; The boundary map is fused with the salient map through a 1×1 convolution to output a salient map containing the precise outline of the salient target.

7. A multi-scale salient target detection method for electromechanical equipment of a hydroelectric unit according to claim 6, characterized in that: The training of the boundary refinement module uses the boundary labels generated by Canny edge detection as the supervision truth value.

8. A multi-scale salient target detection method for electromechanical equipment of a hydroelectric unit according to claim 7, characterized in that: The joint loss function design of step 3 includes: Design a joint loss function that combines cross entropy loss, intersection-over-union loss, and boundary loss to supervise the overall generation of saliency maps; Cross entropy loss formula: Among them, G ij and P ij are the values ​​of the true image G and the predicted image P at position (i, j), H and W are the height and width of the input image respectively; The intersection-over-union loss formula is: L BCE and L IoU The joint loss formula is: L(P,G)=ω*L BCE (P,G)+(1-ω)*L Iou (P,G) Among them, ω weight coefficient adjusts the proportion of the two losses; Boundary loss formula: L(S1,G)=L BCE (S1,G)+L_IoU(S1,G) Among them, E is the target boundary label; The joint loss formula is: Among them, i is a hyperparameter used to adjust the weight of multi-level supervision.

9. A multi-scale salient target detection method for electromechanical equipment of a hydroelectric unit according to claim 8, characterized in that: The step 4 comprises: Step 4.1: Use the training set in step 2 to train the constructed salient object detection model, then use the validation set in step 2 to verify the training results, and continuously adjust the parameters of the improved network training through back propagation to select the optimal weight file. When the overall loss value reaches the minimum and converges, the network model reaches the optimal value. Step 4.2: Input the test set in step 2 into the trained salient object detection model, and use the mean absolute error, F β The value and S value indicators are used to test the model performance.

10. A multi-scale salient target detection method for electromechanical equipment of a hydropower unit according to claim 4, characterized in that: The implementation of the SE branch includes: Perform global average pooling on the feature map to generate channel descriptors; Modeling inter-channel dependencies through two fully connected layers; The channel weights are generated by the Sigmoid function and multiplied channel by channel with the original feature map.