Substation target detection method and device based on substation target detection model, computer equipment and storage medium

Through the automated detection method based on the substation object detection model, through multiple feature map enhancement and iterative training, the problem of low accuracy of substation object detection is solved, and efficient and accurate automated detection is achieved.

CN120495636APending Publication Date: 2025-08-15MAINTENANCE & TEST CENTRE CSG EHV POWER TRANSMISSION CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510636084.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, substation target detection relies on manual detection, and subjective factors lead to low detection accuracy and it is difficult to effectively identify targets such as small animals.

Method used

Through the method based on the substation object detection model, the sample image feature map is obtained and multiple enhancement processing is performed, and the model is iteratively trained with the predicted probability and loss value to achieve automated object detection.

Benefits of technology

It improves the accuracy of substation target detection, avoids subjective errors in manual detection, and ensures the accuracy of detection and automated processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495636A_ABST
    Figure CN120495636A_ABST
Patent Text Reader

Abstract

The invention relates to a substation target detection method and device based on a substation target detection model, computer equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a sample image associated with a to-be-analyzed transformer substation; extracting a plurality of feature maps corresponding to the sample image, obtaining an initial enhanced feature map corresponding to each feature map, performing re-enhancement processing on each initial enhanced feature map to obtain a target enhanced feature map corresponding to each feature map, and obtaining a target enhanced feature map corresponding to each feature map according to the target enhanced feature map; determining a prediction target detection result corresponding to the sample image and a corresponding prediction probability; according to the prediction probability, obtaining a target loss value, and carrying out iterative training on a to-be-trained transformer substation target detection model to obtain a trained transformer substation target detection model; and inputting a to-be-analyzed image into the trained transformer substation target detection model to obtain a target detection result of the to-be-analyzed image. By adopting the method, the detection accuracy of the transformer substation target can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of power grid technology, and in particular to a substation target detection method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on a substation target detection model. Background Art

[0002] In the power grid sector, substations, as key components of the power system, are crucial for their safe operation. However, inadequate isolation and prevention measures within substations allow small animals (such as rats and snakes) to easily enter and disrupt equipment operation, potentially leading to serious safety incidents, including electrical short circuits, equipment failures, and even grid collapse. Therefore, detecting targets (such as small animals) within substation areas is crucial to ensuring safe and stable grid operation.

[0003] Traditionally, manual detection is used to detect targets within a substation area. However, this manual detection method is subjective and prone to errors, resulting in low detection accuracy for substation targets. Summary of the Invention

[0004] Based on this, it is necessary to provide a substation target detection method, device, computer equipment, computer-readable storage medium and computer program product based on a substation target detection model that can improve the detection accuracy of substation targets in response to the above technical problems.

[0005] In a first aspect, the present application provides a substation target detection method based on a substation target detection model, comprising:

[0006] Acquire a sample image associated with the substation to be analyzed;

[0007] Extracting multiple feature maps corresponding to the sample image using a substation target detection model to be trained, obtaining an initial enhanced feature map corresponding to each feature map, performing further enhancement processing on the initial enhanced feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map, and determining a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result based on the target enhanced feature map;

[0008] Obtaining a target loss value according to the predicted probability, and iteratively training the substation target detection model to be trained according to the target loss value to obtain a trained substation target detection model;

[0009] Acquire an image to be analyzed associated with the substation to be analyzed, and input the image to be analyzed into the trained substation target detection model to obtain a target detection result of the image to be analyzed; the target detection result is used to indicate whether there is a target smaller than a preset size in the substation to be analyzed.

[0010] In one embodiment, obtaining an initial enhanced feature map corresponding to each feature map includes:

[0011] Performing multiple feature extraction processes on each feature map to obtain multiple processed feature maps corresponding to each feature map;

[0012] Performing splicing processing on the multiple processed feature maps corresponding to each feature map to obtain a spliced feature map corresponding to each feature map;

[0013] Performing feature extraction processing on the spliced feature map corresponding to each feature map to obtain a processed spliced feature map corresponding to each feature map;

[0014] Each feature map and the processed spliced feature map corresponding to each feature map are fused to obtain an initial enhanced feature map corresponding to each feature map.

[0015] In one embodiment, the further enhancing the initial enhanced feature map corresponding to each feature map to obtain the target enhanced feature map corresponding to each feature map includes:

[0016] Obtaining a derivative feature map of the initial enhanced feature map corresponding to each feature map;

[0017] The initial enhanced feature map corresponding to each feature map and the derived feature map are fused to obtain a target enhanced feature map corresponding to each feature map.

[0018] In one embodiment, determining the predicted target detection result corresponding to the sample image and the predicted probability corresponding to the predicted target detection result based on the target enhancement feature map includes:

[0019] According to the weight corresponding to the target enhancement feature map, the target enhancement feature map is fused to obtain a fused target enhancement feature map;

[0020] Classification processing is performed on the fused target enhanced feature map to obtain a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result.

[0021] In one embodiment, obtaining a target loss value according to the predicted probability includes:

[0022] Obtaining a weight balancing coefficient corresponding to the sample image;

[0023] According to the predicted probability and the weight balancing coefficient, the target correspondence is queried to obtain the loss value corresponding to the predicted probability and the weight balancing coefficient as the target loss value; the target correspondence is used to represent the correspondence between the predicted probability, the weight balancing coefficient and the loss value.

[0024] In one embodiment, after obtaining an image to be analyzed associated with the substation to be analyzed and inputting the image to be analyzed into the trained substation object detection model to obtain an object detection result of the image to be analyzed, the method further includes:

[0025] If the target detection result satisfies a preset target detection result, determining a target category corresponding to the target detection result;

[0026] generating early warning information corresponding to the target category;

[0027] The warning information is sent to a target terminal associated with the substation to be analyzed.

[0028] In a second aspect, the present application further provides a substation target detection device based on a substation target detection model, comprising:

[0029] A sample acquisition module, used to acquire sample images associated with the substation to be analyzed;

[0030] a model processing module, configured to extract multiple feature maps corresponding to the sample image using a substation target detection model to be trained, obtain an initial enhanced feature map corresponding to each feature map, perform further enhancement processing on the initial enhanced feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map, and determine a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result based on the target enhanced feature map;

[0031] A model training module is used to obtain a target loss value according to the predicted probability, and iteratively train the substation target detection model to be trained according to the target loss value to obtain a trained substation target detection model;

[0032] The target detection module is used to obtain an image to be analyzed associated with the substation to be analyzed, and input the image to be analyzed into the trained substation target detection model to obtain a target detection result of the image to be analyzed; the target detection result is used to indicate whether there is a target smaller than a preset size in the substation to be analyzed.

[0033] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0034] Acquire a sample image associated with the substation to be analyzed;

[0035] Extracting multiple feature maps corresponding to the sample image using a substation target detection model to be trained, obtaining an initial enhanced feature map corresponding to each feature map, performing further enhancement processing on the initial enhanced feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map, and determining a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result based on the target enhanced feature map;

[0036] Obtaining a target loss value according to the predicted probability, and iteratively training the substation target detection model to be trained according to the target loss value to obtain a trained substation target detection model;

[0037] Acquire an image to be analyzed associated with the substation to be analyzed, and input the image to be analyzed into the trained substation target detection model to obtain a target detection result of the image to be analyzed; the target detection result is used to indicate whether there is a target smaller than a preset size in the substation to be analyzed.

[0038] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0039] Acquire a sample image associated with the substation to be analyzed;

[0040] Extracting multiple feature maps corresponding to the sample image using a substation target detection model to be trained, obtaining an initial enhanced feature map corresponding to each feature map, performing further enhancement processing on the initial enhanced feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map, and determining a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result based on the target enhanced feature map;

[0041] Obtaining a target loss value according to the predicted probability, and iteratively training the substation target detection model to be trained according to the target loss value to obtain a trained substation target detection model;

[0042] Acquire an image to be analyzed associated with the substation to be analyzed, and input the image to be analyzed into the trained substation target detection model to obtain a target detection result of the image to be analyzed; the target detection result is used to indicate whether there is a target smaller than a preset size in the substation to be analyzed.

[0043] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0044] Acquire a sample image associated with the substation to be analyzed;

[0045] Extracting multiple feature maps corresponding to the sample image using a substation target detection model to be trained, obtaining an initial enhanced feature map corresponding to each feature map, performing further enhancement processing on the initial enhanced feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map, and determining a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result based on the target enhanced feature map;

[0046] Obtaining a target loss value according to the predicted probability, and iteratively training the substation target detection model to be trained according to the target loss value to obtain a trained substation target detection model;

[0047] Acquire an image to be analyzed associated with the substation to be analyzed, and input the image to be analyzed into the trained substation target detection model to obtain a target detection result of the image to be analyzed; the target detection result is used to indicate whether there is a target smaller than a preset size in the substation to be analyzed.

[0048] The above-mentioned substation target detection method, device, computer equipment, storage medium and computer program product based on the substation target detection model first obtains a sample image associated with the substation to be analyzed, and then uses the substation target detection model to be trained to extract multiple feature maps corresponding to the sample image, and obtains the initial enhanced feature map corresponding to each feature map, and performs further enhancement processing on the initial enhanced feature map corresponding to each feature map to obtain the target enhanced feature map corresponding to each feature map. According to the target enhanced feature map, the predicted target detection result corresponding to the sample image and the predicted probability corresponding to the predicted target detection result are determined. Then, according to the predicted probability, the target loss value is obtained, and according to the target loss value, the substation target detection model to be trained is iteratively trained to obtain a trained substation target detection model. Then, the image to be analyzed associated with the substation to be analyzed is obtained, and the image to be analyzed is input into the trained substation target detection model to obtain the target detection result of the image to be analyzed. In this way, when detecting targets in the substation area, the substation target detection model is pre-trained by using sample images associated with the substation to be analyzed, so that in actual applications, after obtaining the image to be analyzed associated with the substation to be analyzed, the target detection result of the image to be analyzed can be predicted, which is beneficial to improving the detection accuracy of substation targets; moreover, the entire process does not require human intervention, avoiding the subjective factors and easy errors in manual detection, which leads to the defect of low detection accuracy of substation targets, and further improves the detection accuracy of substation targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0050] Figure 1 1 is a flow chart of a substation target detection method based on a substation target detection model in one embodiment;

[0051] Figure 2 Schematic diagram of the structure of a receptive field enhancement module in one embodiment;

[0052] Figure 3 2. A schematic diagram of the structure of multi-scale feature fusion based on a receptive field enhancement module in one embodiment;

[0053] Figure 4 1 is a flow chart of a substation target detection method based on a substation target detection model in another embodiment;

[0054] Figure 5 Schematic diagram of the structure of the YOLOX-S network in one embodiment;

[0055] Figure 6 A schematic diagram of the structure of a YOLOX-S network in another embodiment;

[0056] Figure 7 is a structural block diagram of a substation target detection device based on a substation target detection model in one embodiment;

[0057] Figure 8 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0059] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0060] In an exemplary embodiment, Figure 1 As shown, a substation target detection method based on a substation target detection model is provided. This embodiment uses the method applied to a server as an example for illustration; it is understandable that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, and tablets; the server can be implemented as an independent server or a server cluster consisting of multiple servers. In this embodiment, the method includes the following steps:

[0061] Step S101: Acquire a sample image associated with a substation to be analyzed.

[0062] The substation to be analyzed refers to the substation that requires target detection.

[0063] The sample images are images used to train the substation target detection model.

[0064] Exemplarily, in response to a model training instruction for a target detection model for a substation to be trained, the server selects a candidate substation corresponding to the model training instruction from a plurality of candidate substations as the substation to be analyzed; then, the server selects a candidate camera device whose distance to the substation to be analyzed is within a preset distance threshold from a plurality of candidate cameras as the target camera device; then, the server obtains images taken by the target camera device and uses these images as sample images associated with the substation to be analyzed.

[0065] In step S102, a plurality of feature maps corresponding to the sample image are extracted through the substation target detection model to be trained, and an initial enhanced feature map corresponding to each feature map is obtained. The initial enhanced feature map corresponding to each feature map is further enhanced to obtain a target enhanced feature map corresponding to each feature map. According to the target enhanced feature map, the predicted target detection result corresponding to the sample image and the predicted probability corresponding to the predicted target detection result are determined.

[0066] The substation target detection model refers to a network model that can use the image to be analyzed to obtain the target detection result corresponding to the image to be analyzed, such as the YOLOX-S (You Only Look Once X-Small, single detector X-Small version) model.

[0067] Among them, the feature map refers to the intermediate representation of the sample image after a series of convolution and other operations.

[0068] Among them, the initial enhanced feature map refers to the feature map after the receptive field enhancement process.

[0069] The target enhanced feature map refers to the initial enhanced feature map after further enhancement processing.

[0070] The predicted target detection result refers to the predicted value corresponding to the target detection result of the sample image.

[0071] Among them, the prediction probability is used to characterize the possibility that the substation target detection model determines that the predicted target detection result is correct.

[0072] Exemplarily, the server performs denoising processing on the sample image to obtain a processed sample image; then, the server inputs the processed sample image into the substation target detection model to be trained, extracts multiple feature maps corresponding to the processed sample image through the substation target detection model to be trained, and obtains the initial enhanced feature map corresponding to each feature map, performs further enhancement processing on the initial enhanced feature map corresponding to each feature map, and obtains the target enhanced feature map corresponding to each feature map, and determines the predicted target detection result corresponding to the processed sample image and the predicted probability corresponding to the predicted target detection result based on the target enhanced feature map.

[0073] In step S103 , a target loss value is obtained according to the predicted probability, and the substation target detection model to be trained is iteratively trained according to the target loss value to obtain a trained substation target detection model.

[0074] The target loss value refers to the loss function value corresponding to the iterative training of the substation object detection model to be trained. In practical scenarios, the target loss value refers to the focal loss, a loss function used to address class imbalance in tasks such as object detection.

[0075] Exemplarily, the server queries the correspondence between the predicted probability and the loss value based on the predicted probability, and obtains the loss value corresponding to the predicted probability as the target loss value; then, the server adjusts the model parameters of the substation target detection model to be trained according to the target loss value; then, the server re-trains the substation target detection model after the model parameters are adjusted until the target loss value obtained by the trained substation target detection model is less than the loss value threshold, then stops training, and uses the trained substation target detection model as the trained substation target detection model.

[0076] Step S104: Acquire an image to be analyzed that is associated with the substation to be analyzed, and input the image to be analyzed into a trained substation target detection model to obtain a target detection result for the image to be analyzed; the target detection result is used to indicate whether there is a target smaller than a preset size in the substation to be analyzed.

[0077] The image to be analyzed refers to an image that requires target detection.

[0078] The target detection result is used to indicate whether there is a target with a size smaller than a preset size in the substation to be analyzed.

[0079] The preset size refers to a preset size threshold. It should be noted that the preset size may be determined according to circumstances.

[0080] Among them, targets smaller than a preset size refer to small animals, such as mice, snakes, etc.

[0081] Exemplarily, in response to a target detection instruction of the substation to be analyzed, the server obtains an image to be analyzed associated with the substation to be analyzed from a database; then, the server inputs the image to be analyzed into a trained substation target detection model, and determines the sizes of all targets in the substation to be analyzed through the trained substation target detection model; then, the server determines the target detection result of the image to be analyzed based on the sizes of all targets in the substation to be analyzed and preset sizes.

[0082] In the above-mentioned substation target detection method based on the substation target detection model, a sample image associated with the substation to be analyzed is first obtained, and then the substation target detection model to be trained is used to extract multiple feature maps corresponding to the sample image, and the initial enhanced feature map corresponding to each feature map is obtained. The initial enhanced feature map corresponding to each feature map is further enhanced to obtain the target enhanced feature map corresponding to each feature map. According to the target enhanced feature map, the predicted target detection result corresponding to the sample image and the predicted probability corresponding to the predicted target detection result are determined. Then, according to the predicted probability, the target loss value is obtained, and according to the target loss value, the substation target detection model to be trained is iteratively trained to obtain the trained substation target detection model. Then, the image to be analyzed associated with the substation to be analyzed is obtained, and the image to be analyzed is input into the trained substation target detection model to obtain the target detection result of the image to be analyzed. In this way, when detecting targets in the substation area, the substation target detection model is pre-trained by using sample images associated with the substation to be analyzed, so that in actual applications, after obtaining the image to be analyzed associated with the substation to be analyzed, the target detection result of the image to be analyzed can be predicted, which is beneficial to improving the detection accuracy of substation targets; moreover, the entire process does not require human intervention, avoiding the subjective factors and easy errors in manual detection, which leads to the defect of low detection accuracy of substation targets, and further improves the detection accuracy of substation targets.

[0083] In an exemplary embodiment, the above-mentioned step S102, obtaining the initial enhanced feature map corresponding to each feature map, specifically includes the following contents: performing multiple feature extraction processes on each feature map to obtain multiple processed feature maps corresponding to each feature map; performing splicing processes on the multiple processed feature maps corresponding to each feature map to obtain a spliced feature map corresponding to each feature map; performing feature extraction processes on the spliced feature maps corresponding to each feature map to obtain a processed spliced feature map corresponding to each feature map; performing fusion processes on each feature map and the processed spliced feature maps corresponding to each feature map to obtain the initial enhanced feature map corresponding to each feature map.

[0084] Among them, multiple feature extraction processes refer to convolution processes with different expansion rates, such as convolution processes with an expansion rate of 1, convolution processes with an expansion rate of 3, and convolution processes with an expansion rate of 5.

[0085] The processed feature map refers to a feature map that has been subjected to multiple feature extraction processes.

[0086] The spliced feature map refers to a feature map obtained by splicing multiple processed feature maps corresponding to each feature map.

[0087] The feature extraction process refers to a convolution process with a size of 1*1.

[0088] The processed spliced feature map refers to a spliced feature map obtained by performing feature extraction processing on the spliced feature map corresponding to each feature map.

[0089] Exemplarily, the server performs multiple feature extraction processes on each feature map to obtain each feature map after multiple feature extraction processes as multiple processed feature maps corresponding to each feature map; then, the server performs splicing process on the multiple processed feature maps corresponding to each feature map to obtain multiple processed feature maps after splicing process as the spliced feature map corresponding to each feature map; then, the server performs feature extraction process on the spliced feature map corresponding to each feature map to obtain the spliced feature map after feature extraction process as the processed spliced feature map corresponding to each feature map; finally, the server performs fusion process on each feature map and the processed spliced feature map corresponding to each feature map respectively to obtain the processed spliced feature map after fusion process as the initial enhanced feature map corresponding to each feature map.

[0090] For example, refer to Figure 2 , the server performs convolution processing with an expansion rate of 1, a convolution processing with an expansion rate of 3, and a convolution processing with an expansion rate of 5 on the feature map to obtain multiple processed feature maps of different sizes of receptive fields corresponding to each feature map; then, the server performs BN (BatchNormalization) batch processing and ReLU (Rectified Linear Unit) activation function processing on the multiple processed feature maps corresponding to each feature map, and integrates them through the Concat function to obtain the spliced feature map corresponding to each feature map; then, the server performs convolution processing with a size of 1*1 on the spliced feature map corresponding to each feature map to obtain feature map ', as the processed spliced feature map corresponding to each feature map; finally, the server performs fusion processing on each feature map and the processed spliced feature map corresponding to each feature map to obtain feature map '', as the initial enhanced feature map corresponding to each feature map.

[0091] Furthermore, the server can obtain the initial enhanced feature map corresponding to each feature map through the following formula:

[0092] , formula (1)

[0093] , formula (2)

[0094] Among them, DC1 3×3 , DC3 3×3 , DC5 3×3They refer to convolution processing with an expansion rate of 1, a convolution processing with an expansion rate of 3, and a convolution processing with an expansion rate of 5; C 1×1 Refers to the convolution processing of size 1*1; F' refers to the processed spliced feature map corresponding to each feature map; F refers to each feature map; F'' refers to the initial enhanced feature map corresponding to each feature map.

[0095] In this embodiment, by performing multiple feature extraction processes on each feature map, more image features can be mined from different angles and levels, thereby making the information contained in the feature map richer and more detailed; moreover, multiple processed feature maps are spliced and then feature extraction is performed again, further integrating features at different levels and angles, thereby capturing more complex feature relationships and enhancing the expressiveness of features.

[0096] In an exemplary embodiment, the above step S102 performs further enhancement processing on the initial enhanced feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map, specifically including the following contents: obtaining a derivative feature map of the initial enhanced feature map corresponding to each feature map; fusing the initial enhanced feature map and the derivative feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map.

[0097] The derived feature map refers to a feature map obtained by processing the initial enhanced feature map corresponding to each feature map.

[0098] Exemplarily, the server obtains a derivative feature map of the initial enhanced feature map corresponding to each feature map; then, the server fuses the initial enhanced feature map and the derivative feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map.

[0099] For example, refer to Figure 3 , the feature map C5 is processed as follows Figure 2 After the receptive field enhancement processing shown in FIG, the initial enhanced feature map C5'' corresponding to the feature map C5 is obtained; then, the initial enhanced feature map corresponding to the feature map C5 is up-sampled to obtain the initial enhanced feature map corresponding to the processed feature map C5; then, the initial enhanced feature map corresponding to the processed feature map C5 and the feature map C4 are fused to obtain a fused feature map, and the fused feature map is processed as follows Figure 2 After the receptive field enhancement processing shown in FIG, the initial enhanced feature map C4'' corresponding to the feature map C4 is obtained; then, the initial enhanced feature map C4'' corresponding to the feature map C4 and the feature map C3 are fused to obtain a fused feature map, and the fused feature map is processed as follows Figure 2After the receptive field enhancement processing shown, the initial enhanced feature map C3'' corresponding to the feature map C3 is obtained; then, the initial enhanced feature map C3'' is down-sampled to obtain the feature map C4''; then, the feature map C4''' is down-sampled, and the feature map obtained by the processing is fused with the initial enhanced feature map C5'' to obtain the target enhanced feature map C5, as the target enhanced feature map corresponding to the feature map C5; then, the target enhanced feature map C5 is up-sampled, and the feature map obtained by the processing is fused with the initial enhanced feature map C4'' and the feature map C4''' to obtain the target enhanced feature map C4, as the target enhanced feature map corresponding to the feature map C4; then, the target enhanced feature map C4 is up-sampled, and the feature map obtained by the processing is fused with the initial enhanced feature map C3'' to obtain the target enhanced feature map C3, as the target enhanced feature map corresponding to the feature map C3.

[0100] Furthermore, the server determines each feature map based on the extraction order of multiple feature maps corresponding to the sample image (for example, the size of the feature maps is from small to large). The specific contents are as follows: according to the extraction order of multiple feature maps corresponding to the sample image, the first feature map is used as the current feature map, the current feature map is enhanced to obtain the enhanced current feature map, the enhanced current feature map and the next feature map corresponding to the current feature map are fused to obtain a fused feature map, and the fused feature map is used as the new current feature map, and the step of enhancing the current feature map to obtain the enhanced current feature map is jumped to, until the number corresponding to the fused feature maps is equal to the preset number, and the first feature map and all the fused feature maps are used as each feature map.

[0101] In this embodiment, by obtaining a derivative feature map of the initial enhanced feature map corresponding to each feature map, the diversity of features can be further expanded, thereby compensating for the shortcomings of a single feature, thereby obtaining a more comprehensive and accurate feature representation, and providing a basis for subsequent data processing.

[0102] In an exemplary embodiment, the above step S102 determines the predicted target detection result corresponding to the sample image and the predicted probability corresponding to the predicted target detection result based on the target enhancement feature map, specifically including the following contents: fusing the target enhancement feature map according to the weight corresponding to the target enhancement feature map to obtain a fused target enhancement feature map; classifying the fused target enhancement feature map to obtain the predicted target detection result corresponding to the sample image and the predicted probability corresponding to the predicted target detection result.

[0103] Among them, the weight is used to represent the importance of the target enhanced feature map.

[0104] The fused target enhanced feature map refers to a feature map obtained by fusing the target enhanced feature map.

[0105] Exemplarily, the server determines the correlation between the target enhancement feature maps, and determines the weights corresponding to the target enhancement feature maps based on the correlation between the target enhancement feature maps; for example, a higher weight is assigned to a target enhancement feature map with a lower correlation with other target enhancement feature maps, and a lower weight is assigned to a target enhancement feature map with a higher correlation with other target enhancement feature maps; then, the server fuses the target enhancement feature maps according to the weights corresponding to the target enhancement feature maps to obtain a fused target enhancement feature map; then, the server classifies the fused target enhancement feature map through the FC (Fully Connected layer) layer to obtain a predicted target detection result corresponding to the sample image, and a predicted probability corresponding to the predicted target detection result.

[0106] In this embodiment, fusion processing is performed according to the weights corresponding to the target enhancement feature maps, so that different target enhancement feature maps can be integrated more accurately, thereby obtaining a more representative and accurate fused target enhancement feature map, which is convenient for subsequent classification processing; moreover, classification processing is performed on the fused target enhancement feature map, which can effectively classify and identify the targets in the image, which is convenient for subsequent model training processing.

[0107] In an exemplary embodiment, the above step S103 obtains the target loss value based on the predicted probability, which specifically includes the following contents: obtaining the weight balance coefficient corresponding to the sample image; querying the target correspondence relationship based on the predicted probability and the weight balance coefficient, and obtaining the loss value corresponding to the predicted probability and the weight balance coefficient as the target loss value; the target correspondence relationship is used to represent the correspondence between the predicted probability, the weight balance coefficient and the loss value.

[0108] Among them, the weight balance coefficient is used to represent the ratio of positive and negative samples in the sample image.

[0109] Among them, the target correspondence is used to express the correspondence between the prediction probability, weight balance coefficient and loss value.

[0110] Exemplarily, the server obtains the ratio of positive and negative samples in the sample image, and determines the weight balance coefficient corresponding to the sample image based on the ratio of positive and negative samples in the sample image; then, the server constructs a correspondence between the predicted probability, the weight balance coefficient and the loss value as the target correspondence; then, the server queries the target correspondence based on the predicted probability and the weight balance coefficient, and obtains the loss value corresponding to the predicted probability and the weight balance coefficient as the target loss value.

[0111] For example, the server can obtain the target loss value through the following formula:

[0112] , formula (3)

[0113] in, Refers to the weight coefficient that balances the ratio of positive and negative samples. The positive sample is , the negative sample is 1- , ranging from 0 to 1; It refers to the predicted probability; It refers to the weight balance coefficient between difficult-to-classify samples and easy-to-classify samples, which is used to reduce the focus on easy-to-classify samples and is usually set to 2 or 5.

[0114] In this embodiment, by obtaining the weight balance coefficient corresponding to the sample image, targeted adjustments can be made to different situations, which is conducive to improving the accuracy of determining the target loss value; moreover, by combining the predicted probability and the weight balance coefficient, a more accurate target loss value can be obtained, further improving the accuracy of determining the target loss value.

[0115] In an exemplary embodiment, the above-mentioned step S104, after acquiring the image to be analyzed associated with the substation to be analyzed and inputting the image to be analyzed into the trained substation target detection model to obtain the target detection result of the image to be analyzed, specifically includes the following contents: when the target detection result meets the preset target detection result, determining the target category corresponding to the target detection result; generating early warning information corresponding to the target category; and sending the early warning information to the target terminal associated with the substation to be analyzed.

[0116] The preset target detection result refers to a preset target detection result, for example, there is a target smaller than a preset size in the substation to be analyzed. It should be noted that the preset target detection result may be determined according to the circumstances.

[0117] Among them, the target category refers to the type information of the target, such as mouse, snake, etc.

[0118] Among them, early warning information refers to the warning information corresponding to the target category.

[0119] The target terminal refers to the terminal associated with the substation to be analyzed, for example, the mobile phone or computer of the staff in the substation to be analyzed.

[0120] Exemplarily, when the target detection result meets the preset target detection result, the server determines the target category corresponding to the target detection result; then, based on the target category corresponding to the target detection result, according to the preset warning information template, the server generates warning information corresponding to the target category; then, the server determines the terminal associated with the substation to be analyzed as the target terminal, and sends the warning information to the target terminal.

[0121] In this embodiment, by sending the warning information of the target category corresponding to the target detection result to the target terminal associated with the substation to be analyzed in a timely manner, it is ensured that the relevant personnel can receive the notification in the first time, thereby improving the timeliness of problem handling and reducing the losses caused by potential risks.

[0122] In an exemplary embodiment, Figure 4 As shown, another substation target detection method based on the substation target detection model is provided. Taking the application of this method to a server as an example, the method includes the following steps:

[0123] Step S401: Acquire a sample image associated with the substation to be analyzed.

[0124] Step S402: extract multiple feature maps corresponding to the sample image through the substation target detection model to be trained.

[0125] In step S403, the substation target detection model to be trained is used to perform multiple feature extraction processes on each feature map to obtain multiple processed feature maps corresponding to each feature map; the multiple processed feature maps corresponding to each feature map are spliced to obtain a spliced feature map corresponding to each feature map; the spliced feature map corresponding to each feature map is subjected to feature extraction processes to obtain a processed spliced feature map corresponding to each feature map; and each feature map and the processed spliced feature map corresponding to each feature map are fused to obtain an initial enhanced feature map corresponding to each feature map.

[0126] In step S404, a derivative feature map of the initial enhanced feature map corresponding to each feature map is obtained through the substation target detection model to be trained; the initial enhanced feature map and the derivative feature map corresponding to each feature map are fused to obtain a target enhanced feature map corresponding to each feature map.

[0127] In step S405, the target enhancement feature map is fused according to the weight corresponding to the target enhancement feature map through the substation target detection model to be trained to obtain a fused target enhancement feature map; the fused target enhancement feature map is classified to obtain the predicted target detection result corresponding to the sample image and the predicted probability corresponding to the predicted target detection result.

[0128] Step S406, obtain the weight balance coefficient corresponding to the sample image; query the target correspondence based on the predicted probability and the weight balance coefficient, and obtain the loss value corresponding to the predicted probability and the weight balance coefficient as the target loss value; the target correspondence is used to represent the correspondence between the predicted probability, the weight balance coefficient and the loss value.

[0129] Step S407: Iteratively train the substation target detection model to be trained according to the target loss value to obtain a trained substation target detection model.

[0130] Step S408: Acquire an image to be analyzed that is associated with the substation to be analyzed, and input the image to be analyzed into the trained substation target detection model to obtain a target detection result for the image to be analyzed; the target detection result is used to indicate whether there is a target smaller than a preset size in the substation to be analyzed.

[0131] In the above-mentioned substation target detection method based on the substation target detection model, when detecting targets in the substation area, the substation target detection model is pre-trained by using sample images associated with the substation to be analyzed, so that in actual applications, after obtaining the image to be analyzed associated with the substation to be analyzed, the target detection result of the image to be analyzed can be predicted, which is beneficial to improving the detection accuracy of the substation target; moreover, the entire process does not require human intervention, avoiding the subjective factors and easy errors in the manual detection method, which leads to the defect of low detection accuracy of substation targets, and further improves the detection accuracy of substation targets.

[0132] In an exemplary embodiment, in order to more clearly illustrate the substation target detection method based on the substation target detection model provided by the embodiment of the present application, the substation target detection method based on the substation target detection model is specifically described below with a specific embodiment. In one embodiment, the present application also provides a substation small animal recognition method based on a small target detection algorithm. When detecting targets in the substation area, a sample image associated with the substation to be analyzed is first obtained, and then a substation target detection model to be trained is used to extract multiple feature maps corresponding to the sample image, and an initial enhanced feature map corresponding to each feature map is obtained. The initial enhanced feature map corresponding to each feature map is further enhanced to obtain a target enhanced feature map corresponding to each feature map. According to the target enhanced feature map, a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result are determined. Then, according to the predicted probability, a target loss value is obtained, and according to the target loss value, the substation target detection model to be trained is iteratively trained to obtain a trained substation target detection model. Then, an image to be analyzed associated with the substation to be analyzed is obtained, and the image to be analyzed is input into the trained substation target detection model to obtain a target detection result for the image to be analyzed. Specifically include the following:

[0133] (1) Aiming at the problem that the YOLOX-S feature pyramid cannot extract sufficient feature information of small animals and is difficult to detect multi-scale targets, a multi-scale feature fusion network based on receptive field enhancement is proposed. The receptive field is enhanced by dilated convolution to improve the contextual information of small animal target objects.

[0134] (2) Improve the CSPDarknet53 (Cross Stage Partial Darknet-53) backbone network in YOLOX-S to increase the model calculation speed by reducing the amount of calculation.

[0135] (3) By applying Focal Loss, the imbalance between positive and negative samples is effectively eliminated, thereby improving the detection rate of small animals.

[0136] The details are as follows:

[0137] 1. Multi-scale feature fusion network based on receptive field enhancement technology.

[0138] Taking the YOLOX algorithm as an example, its derivative model YOLOX-S has the advantages of fewer model parameters and faster detection speed. The network structure is as follows: Figure 5 and Figure 6As shown in the figure, it includes the input end, backbone network, Neck, and prediction end. The backbone network uses the core feature extraction algorithm CSPDarknet (Cross Stage Partial Darknet). The Neck module FPN (Feature Pyramid Networks) + PAN (Path Aggregation Network) architecture integrates feature maps of various sizes. The FPN architecture starts from the top and passes higher semantic information to the bottom, thereby providing higher semantic accuracy. The PAN architecture starts from the bottom and passes lower position information to the top, achieving effective fusion of the two feature maps, thereby achieving accurate prediction of images of various sizes.

[0139] Some explanations in the picture:

[0140] Conv (Convolutional) represents the convolutional layer, BN represents the batch normalization layer, SiLU (Sigmoid-weighted Linear Unit) represents the SiLU activation function, Res (Residual) represents the residual structure, Concat represents the concatenation module, and slice represents the slice.

[0141] (1) Backbone network: a convolutional neural network that aggregates and forms image feature maps at different image granularities.

[0142] (2) Neck: A series of network layers that mix and combine image features and pass them to the prediction end.

[0143] (3) Prediction end: predict image features, generate bounding boxes and predict categories.

[0144] (4) CBS (Convolutional, Batch Normalization, Sigmoid-weighted LinearUnit) represents the composition of Conv+BN+SiLU.

[0145] (5) Res represents the residual structure, which is used to build a deep network.

[0146] (6) CSP1_X consists of CBS module, Res module and Concat module.

[0147] (7) CSP2_X consists of CBS module and Concat module.

[0148] (8) Focus (feature focusing) first concats multiple slice results and then sends them to the CBS module.

[0149] (9) SPP (Spatial Pyramid Pooling) consists of CBS, maximum pooling layer, and Concat module, and uses 1×1, 5×5, 9×9 and 13×13 maximum pooling methods to perform multi-scale feature fusion.

[0150] The YOLOX-S algorithm has a low recognition accuracy when applied to small target detection and recognition. The main reasons are: (1) Small targets are relatively small and are greatly affected by background interference during detection. It is difficult to distinguish between positive and negative samples, resulting in sample imbalance. (2) YOLOX-S uses FPN+PAN technology in the Neck module. The FPN structure ignores high-level feature information, and the network pays too much attention to low-level features. The PAN structure adds a bottom-up pathway based on FPN. However, as the network layer deepens, the functional channels will continue to decrease, the feature map will also experience information loss, and the feature maps of other layers will contain less relevant context information.

[0151] Small animal monitoring images from substations contain few pixels, and the feature maps extracted by the detection algorithm after passing through a multi-layer convolutional neural network can suffer from information loss, resulting in reduced detection accuracy. Furthermore, in complex substation scenes, the large variations in the scale of target objects also affect detection accuracy. Therefore, this embodiment proposes the introduction of a Receptive Field Enhancement Module (RFEM) into the feature fusion network structure. This aims to increase the network's receptive field, thereby improving its ability to capture global contextual information and enhancing the network's feature fusion capabilities.

[0152] The receptive field refers to the area of the input image covered by a single output feature of a convolutional neural network (i.e., the activation of a single neuron). The calculation formula is as follows:

[0153] , formula (4)

[0154] Where: is the receptive field size of the nth layer, is the size of the convolution kernel in the nth layer, Represents the convolution step size of the i-th layer.

[0155] Enhancing the receptive field is very effective in improving target detection accuracy. Deep learning networks usually enhance the receptive field by continuously increasing the network depth. However, as the network deepens, the resolution of the feature map decreases, resulting in the loss of information contained in the extracted feature map. Dilated convolution can effectively solve this problem. It can not only increase the receptive field range and improve the contextual information of small targets, but also retain the high definition of the image and provide richer feature information. The formula for calculating the receptive field after introducing dilated convolution is as follows:

[0156] , formula (5)

[0157] , formula (6)

[0158] Where: is the receptive field size of the nth layer, d is the convolution step size, and k is the convolution kernel size.

[0159] Since small animals occupy fewer pixels in an image, information loss may occur in the feature maps extracted after multi-layer convolution operations, resulting in decreased detection accuracy. To address this issue, this embodiment proposes a feature fusion network based on receptive field enhancement technology. The feature layer can obtain a larger receptive field while retaining more local detail feature information, thereby improving the algorithm's sensitivity to small targets. The structure of RFEM is as follows: Figure 2 shown.

[0160] Figure 2 The following steps are included: (1) The feature map F in the figure is subjected to different scale expansion convolution to obtain receptive fields of different sizes, and then processed by BN batch processing and ReLU activation function, integrated by Concat function and 1×1 convolution operation to obtain the feature map ; (2) Finally, it is fused with the feature map F to obtain the feature map The three parallel dilated convolution kernels have the same size of 3×3, and their dilation rates are 1, 3, and 5 respectively.

[0161] Steps (1) and (2) can be expressed by the following formulas:

[0162] , formula (1)

[0163] , formula (2)

[0164] Where: represents a general convolution operation with a convolution kernel of n. Represents a dilated convolution operation with a convolution kernel of n and a dilation rate of m.

[0165] like Figure 3This multi-scale feature fusion architecture is based on the Receptive Field Enhancement Module (RFEM). Three feature maps of different sizes fuse high-level features with low-level features from top to bottom, with the receptive field expanded before and after each fusion step. The fused feature maps enhanced by the receptive field enhancement technology contain a richer multi-scale feature representation. While widening the receptive field, the new feature maps retain channel information while adding multi-scale contextual information, significantly improving the performance of small and multi-scale object detection algorithms.

[0166] Note: If Figure 3 As shown in the figure, after multiple convolutions through the backbone network, the image is converted into a feature map {C1, C2, C3, C4, C5}. The top three feature maps of different sizes, C3, C4, and C5, are input into the multi-scale feature fusion network for deeper fusion.

[0167] (1) From top to bottom, high-level features (smaller size) are organically integrated with low-level features (larger size), and each fusion is enhanced by expanding the receptive field.

[0168] (2) The 13×13 feature map C5 is first obtained by receptive field enhancement (RFEM) , and then upsample (double the size) to get a 26×26 feature map , the feature map is then fused with the 26×26 C4 feature map (fusion refers to the element-wise addition of feature maps at corresponding positions) to obtain C4'.

[0169] (3) Repeat step (2) to get the final feature map (Size 52×52).

[0170] (4) Downsample the C3' feature map (reduce its size by two times) and then fuse it.

[0171] (5) Repeat step (4) to obtain C5''.

[0172] (6) Perform the feature fusion operation from top to bottom again, and obtain C3'' from C5''.

[0173] 2. Improvement method of YOLOX-S backbone network based on ResNet-50 (Residual Network with 50 layers).

[0174] The YOLOX-S backbone network uses CSPDarknet53, which results in certain drawbacks in YOLOX-S's detection of small objects: gradient duplication occurs between individual neuron nodes during backpropagation, reducing the model's overall performance. This embodiment improves the YOLOX-S backbone network, not only resolving the gradient duplication problem but also reducing the number of model parameters. This significantly helps improve the model's ability to recognize small objects and reduce model recognition time.

[0175] The residual network structure has become an important cornerstone structure in the field of computer vision because of its simplicity and practicality. Its width and depth are very easy to expand and modify. As the network depth increases, there is no need to worry too much about network degradation. As long as there is enough training data, better performance can be achieved. Moreover, the residual structure can improve the performance of the model while reducing the number of parameters. Therefore, this embodiment uses the ResNet-50 residual network, which has a model computational complexity far less than CSPDarknet53, to replace the YOLOX-S backbone network. Using ResNet-50 as the backbone network, an efficient feature extraction network is constructed, which can improve its performance on small target objects. While ensuring accuracy, it helps to reduce the number of parameters of the entire model and the amount of computation during training, thereby improving detection speed.

[0176] Residual network is a commonly used deep learning model, which mainly consists of the following steps:

[0177] (1) Input image: The network receives the input image.

[0178] (2) Convolutional layer: The input image first passes through one or more convolutional layers for feature extraction.

[0179] (3) Residual Block: Each residual block contains several layers (such as convolutional layers, batch normalization layers, and activation layers). The outputs of these layers are added to the input to form a residual connection. This allows the network to learn the residual (i.e., the difference) between the input and output, rather than directly learning the output.

[0180] (4) Activation function: A nonlinear activation function, such as ReLU, is usually applied inside or after the residual block to increase the nonlinear expression ability of the network.

[0181] (5) Batch Normalization: In order to improve the stability of training and accelerate convergence, the residual block usually includes a batch normalization layer.

[0182] (6) Skip connection: In the residual block, the input can directly skip some layers and be added to the following layers. This skip connection is a key feature of the residual network. It allows the gradient to flow directly through the network, alleviating the gradient disappearance problem.

[0183] (7) Repeated residual blocks: Residual networks increase the depth of the network by stacking multiple residual blocks. The deeper the network, the richer the feature hierarchy, and the network can abstract higher-level concepts from the original data.

[0184] (8) The output of the residual network is a feature map with rich information, which is used as the input of the Neck end of the YOLOX-S network.

[0185] 3. Imbalanced sample processing method based on Focal Loss loss function.

[0186] YOLOX-S's loss function primarily consists of the position loss function (GIoU_Loss, Generalized Intersection over Union Loss), the confidence loss function, and the category loss function (all based on the mean binary cross entropy loss, or BCE_Loss). Because small animals are relatively small, they are significantly affected by background noise during target detection, making it difficult to distinguish between positive and negative samples. The number of negative samples far outnumbers the number of positive samples, leading to sample imbalance, hindering the model's learning of positive samples and unstable performance. Conventional metrics for evaluating the performance of target detection models include precision, recall, average precision (AP), mean average precision (mAP), and frames per second (FPS), which measures the speed of the target detection algorithm. Higher values are preferred for all four metrics.

[0187] , formula (7)

[0188] , formula (8)

[0189] Where: TP (True Positives) represents the number of positive samples that are correctly identified, FP (False Positives) represents the number of positive samples that are incorrectly identified, and FN (False Negatives) represents the number of positive samples that are incorrectly identified as negative samples.

[0190] As can be seen from the Recall calculation formula, when the samples are unbalanced, it is easy for positive samples to be misidentified as negative samples (small animals are misidentified), resulting in a small TP and a large FN, low model recall and performance. To solve this problem, this example uses Focal Loss instead of confidence loss and category loss BCE_Loss to balance the model's sensitivity to positive and negative samples. The calculation formula is as follows:

[0191]

[0192] Where: Is the weight coefficient for balancing the ratio of positive and negative samples, and the positive sample is , the negative sample is 1- , ranging from 0 to 1. is the predicted probability. It is the weight balance coefficient between difficult-to-classify samples and easy-to-classify samples, which is used to reduce the focus on easy-to-classify samples and is usually set to 2 or 5.

[0193] In this way, Focal Loss encourages the model to focus on small target samples that are difficult to classify, thereby improving the model's detection ability for minority categories.

[0194] In the above embodiment, when detecting targets within a substation area, the substation target detection model is pre-trained using sample images associated with the substation to be analyzed. This facilitates prediction of target detection results for the image to be analyzed after obtaining the image to be analyzed, thereby improving the detection accuracy of substation targets. Furthermore, the entire process requires no human intervention, avoiding the subjective factors and errors that often occur in manual detection, which can lead to low substation target detection accuracy. This further improves the detection accuracy of substation targets. Furthermore, this embodiment proposes a substation small animal target detection algorithm based on the conventional target detection algorithm YOLOX-S. This algorithm proposes a recognition algorithm suitable for substation small animal target detection: First, to address the issues of insufficient small animal feature information extracted by the YOLOX-S feature pyramid and the difficulty in detecting multi-scale targets, a multi-scale feature fusion network based on receptive field enhancement is proposed. This network utilizes dilated convolution to enhance the receptive field and improve the contextual information of small animal targets. Second, the CSPDarknet53 backbone network in YOLOX-S is improved to reduce the computational complexity and improve the model's computational speed. Third, by applying Focal Loss, the imbalance between positive and negative samples is effectively eliminated, improving the detection rate of small animals. This embodiment effectively overcomes the problem that conventional target detection algorithms are not suitable for small animal identification in substations, improving recognition accuracy and speed. Through small animal identification and early warning, small animal intrusion can be detected in a timely manner and notified to operation and maintenance personnel to deal with it, significantly reducing the risk of small animal intrusion and ensuring the safe and stable operation of substation equipment.

[0195] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0196] Based on the same inventive concept, an embodiment of the present application further provides a substation target detection device based on a substation target detection model for implementing the substation target detection method based on a substation target detection model. The implementation solution provided by the device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more embodiments of the substation target detection device based on a substation target detection model provided below can be found in the above limitations of the substation target detection method based on a substation target detection model, and will not be repeated here.

[0197] In an exemplary embodiment, Figure 7 As shown, a substation target detection device based on a substation target detection model is provided, comprising: a sample acquisition module 701, a model processing module 702, a model training module 703 and a target detection module 704, wherein:

[0198] The sample acquisition module 701 is used to acquire sample images associated with the substation to be analyzed.

[0199] The model processing module 702 is used to extract multiple feature maps corresponding to the sample image through the substation target detection model to be trained, and obtain the initial enhanced feature map corresponding to each feature map, and further enhance the initial enhanced feature map corresponding to each feature map to obtain the target enhanced feature map corresponding to each feature map, and determine the predicted target detection result corresponding to the sample image and the predicted probability corresponding to the predicted target detection result based on the target enhanced feature map.

[0200] The model training module 703 is used to obtain a target loss value according to the predicted probability, and iteratively train the substation target detection model to be trained according to the target loss value to obtain a trained substation target detection model.

[0201] The target detection module 704 is used to obtain the image to be analyzed associated with the substation to be analyzed, and input the image to be analyzed into the trained substation target detection model to obtain the target detection result of the image to be analyzed; the target detection result is used to indicate whether there is a target smaller than a preset size in the substation to be analyzed.

[0202] In an exemplary embodiment, the model processing module 702 is further used to perform multiple feature extraction processes on each feature map to obtain multiple processed feature maps corresponding to each feature map; perform splicing processes on the multiple processed feature maps corresponding to each feature map to obtain a spliced feature map corresponding to each feature map; perform feature extraction processes on the spliced feature map corresponding to each feature map to obtain a processed spliced feature map corresponding to each feature map; and perform fusion processes on each feature map and the processed spliced feature map corresponding to each feature map to obtain an initial enhanced feature map corresponding to each feature map.

[0203] In an exemplary embodiment, the model processing module 702 is also used to obtain a derivative feature map of the initial enhanced feature map corresponding to each feature map; and fuse the initial enhanced feature map and the derivative feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map.

[0204] In an exemplary embodiment, the model processing module 702 is also used to fuse the target enhancement feature map according to the weight corresponding to the target enhancement feature map to obtain a fused target enhancement feature map; and to classify the fused target enhancement feature map to obtain a predicted target detection result corresponding to the sample image, and a predicted probability corresponding to the predicted target detection result.

[0205] In an exemplary embodiment, the model training module 703 is also used to obtain the weight balance coefficient corresponding to the sample image; based on the predicted probability and the weight balance coefficient, the target correspondence is queried to obtain the loss value corresponding to the predicted probability and the weight balance coefficient as the target loss value; the target correspondence is used to represent the correspondence between the predicted probability, the weight balance coefficient and the loss value.

[0206] In an exemplary embodiment, the substation target detection device based on the substation target detection model also includes an information sending module, which is used to determine the target category corresponding to the target detection result when the target detection result meets the preset target detection result; generate early warning information corresponding to the target category; and send the early warning information to the target terminal associated with the substation to be analyzed.

[0207] Each module in the aforementioned substation target detection device based on the substation target detection model can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0208] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data such as sample images and feature maps. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a substation target detection method based on a substation target detection model is implemented.

[0209] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0210] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0211] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0212] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0213] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0214] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0215] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A substation target detection method based on a substation target detection model, characterized in that: The method comprises: Acquire a sample image associated with the substation to be analyzed; Extracting multiple feature maps corresponding to the sample image using a substation target detection model to be trained, obtaining an initial enhanced feature map corresponding to each feature map, performing further enhancement processing on the initial enhanced feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map, and determining a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result based on the target enhanced feature map; Obtaining a target loss value according to the predicted probability, and iteratively training the substation target detection model to be trained according to the target loss value to obtain a trained substation target detection model; Acquire an image to be analyzed associated with the substation to be analyzed, and input the image to be analyzed into the trained substation target detection model to obtain a target detection result of the image to be analyzed; the target detection result is used to indicate whether there is a target smaller than a preset size in the substation to be analyzed.

2. The method according to claim 1, characterized in that The obtaining of the initial enhanced feature map corresponding to each feature map includes: Performing multiple feature extraction processes on each feature map to obtain multiple processed feature maps corresponding to each feature map; Performing splicing processing on the multiple processed feature maps corresponding to each feature map to obtain a spliced feature map corresponding to each feature map; Performing feature extraction processing on the spliced feature map corresponding to each feature map to obtain a processed spliced feature map corresponding to each feature map; Each feature map and the processed spliced feature map corresponding to each feature map are fused to obtain an initial enhanced feature map corresponding to each feature map.

3. The method according to claim 1, characterized in that The further enhancing the initial enhanced feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map includes: Obtaining a derivative feature map of the initial enhanced feature map corresponding to each feature map; The initial enhanced feature map corresponding to each feature map and the derived feature map are fused to obtain a target enhanced feature map corresponding to each feature map.

4. The method according to claim 1, wherein Determining, based on the target enhancement feature map, a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result includes: According to the weight corresponding to the target enhancement feature map, the target enhancement feature map is fused to obtain a fused target enhancement feature map; Classification processing is performed on the fused target enhanced feature map to obtain a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result.

5. The method according to claim 1, wherein Obtaining a target loss value according to the predicted probability includes: Obtaining a weight balance coefficient corresponding to the sample image; According to the predicted probability and the weight balancing coefficient, the target correspondence is queried to obtain the loss value corresponding to the predicted probability and the weight balancing coefficient as the target loss value; the target correspondence is used to represent the correspondence between the predicted probability, the weight balancing coefficient and the loss value.

6. The method according to any one of claims 1 to 5, characterized in that After acquiring an image to be analyzed associated with the substation to be analyzed and inputting the image to be analyzed into the trained substation object detection model to obtain an object detection result of the image to be analyzed, the method further includes: If the target detection result satisfies a preset target detection result, determining a target category corresponding to the target detection result; generating early warning information corresponding to the target category; The warning information is sent to a target terminal associated with the substation to be analyzed.

7. A substation target detection device based on a substation target detection model, characterized in that: The device comprises: A sample acquisition module, used to acquire sample images associated with the substation to be analyzed; a model processing module, configured to extract multiple feature maps corresponding to the sample image using a substation target detection model to be trained, obtain an initial enhanced feature map corresponding to each feature map, perform further enhancement processing on the initial enhanced feature map corresponding to each feature map to obtain a target enhanced feature map corresponding to each feature map, and determine a predicted target detection result corresponding to the sample image and a predicted probability corresponding to the predicted target detection result based on the target enhanced feature map; A model training module is used to obtain a target loss value according to the predicted probability, and iteratively train the substation target detection model to be trained according to the target loss value to obtain a trained substation target detection model; The target detection module is used to obtain an image to be analyzed associated with the substation to be analyzed, and input the image to be analyzed into the trained substation target detection model to obtain a target detection result of the image to be analyzed; the target detection result is used to indicate whether there is a target smaller than a preset size in the substation to be analyzed.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.