A power grid operation scene road foreign matter intrusion detection method based on deep learning

By combining a multi-scale road scene perception enhancement module and a cross-scale rich feature fusion module with a multi-core attention enhancement module, the problem of multi-scale and complex background detection of foreign objects in power grid operation scenarios is solved, and high-precision and robust identification of foreign objects is achieved.

CN121305467BActive Publication Date: 2026-04-14JIANGSU HAOYUAN TECHNOLOGY DEVELOPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU HAOYUAN TECHNOLOGY DEVELOPMENT CO LTD
Filing Date
2025-11-03
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify multi-scale, complex backgrounds, and irregular foreign objects in power grid operation scenarios, leading to unstable detection and decreased accuracy. In particular, traditional models are ill-suited to adapting to the variability of foreign objects, background interference, and limitations in feature extraction from non-rigid targets in open and variable outdoor environments.

Method used

The system employs a multi-scale road scene perception enhancement module, a cross-scale rich feature fusion module, and a multi-kernel attention enhancement module. Through multi-scale dilated convolution, multi-kernel parallel attention, and deep and shallow feature fusion, combined with deformable convolution and reparameterizable convolution, it extracts the geometric and structural features of the target, enhances the adaptability to foreign objects, suppresses interference from complex backgrounds, and improves detection accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and robustness of foreign object detection in power grid operation scenarios, and can more accurately capture the characteristics of foreign objects of different sizes, shapes and postures. It solves the problem of detection instability caused by the variability of foreign objects and complex backgrounds in existing technologies, and improves the salience of targets and feature discrimination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121305467B_ABST
    Figure CN121305467B_ABST
Patent Text Reader

Abstract

The application discloses the technical field of power grid safety operation and maintenance, and particularly relates to a power grid operation scene road foreign matter intrusion detection method based on deep learning, which comprises the following steps: firstly, an original power grid operation scene road image is cropped into a specified size, and then input into a multi-scale road scene perception enhancement module MRPE to obtain a shallow detail feature map F2, a middle layer structure feature map F4 and a deep layer semantic feature map F6; then, the F2, the F4 and the F6 are input into a cross-scale rich feature fusion module CS-RFF to obtain a deep layer prediction feature map F12, a middle layer prediction feature map F15 and a full-scale feature map F17; finally, the F12, the F15 and the F17 are input into a multi-core attention enhancement module MAEM to obtain an image with a detection result; based on this, when dealing with diversified intrusion targets, the application can more accurately capture the foreign matter features of different sizes, shapes and postures appearing on the road in the power grid operation scene, and solves the problem of unstable detection caused by the variability of road foreign matters in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grid safety operation and maintenance technology, specifically to a method for detecting foreign object intrusion in power grid operation scenarios based on deep learning. Background Technology

[0002] As a vital national infrastructure, the safe and stable operation of the power grid is of paramount importance. With the increase in construction and maintenance activities, the difficulty of safety management at on-site operations has significantly increased. Particularly on critical roads in power grid operation scenarios, foreign object intrusion has become a major risk factor affecting operational safety. For example, fallen trees may block inspection routes, falling plastic bags may become entangled in equipment, and scattered building debris or metal objects may threaten vehicle passage and the safety of workers. Traditional manual inspection methods suffer from problems such as delayed response and limited coverage, making it difficult to detect potential hazards in a timely manner. Therefore, there is an urgent need for an intelligent detection method that can achieve real-time monitoring and accurate identification. However, current traditional methods, which mainly rely on manual inspections, are no longer sufficient to meet the safety management needs of power grid operation sites.

[0003] Therefore, developing a method for detecting foreign object intrusion on critical roads in power grid operation scenarios that can integrate deep learning visual perception and achieve automated unmanned monitoring and real-time intelligent early warning has become an urgent need to ensure the safety of power grid operations in the new era.

[0004] In power grid operation scenarios, ensuring the safety of access roads and promptly detecting and removing foreign objects are crucial. Related foreign object detection technologies have evolved from traditional image processing methods to deep learning-based intelligent sensing systems to adapt to the complexity and high reliability requirements of the power grid operating environment.

[0005] Early detection technologies primarily relied on traditional image processing, such as inter-frame differencing or background modeling, to identify newly added obstacles on inspection roads or in substation areas. Some solutions also employed sensors like radar waves, utilizing their echo characteristics for detection. However, these traditional methods generally suffer from limitations such as sensitivity to environmental changes like lighting and weather, and difficulty adapting to complex outdoor environments. Their stability and accuracy are severely affected during nighttime repairs or in inclement weather.

[0006] To overcome the limitations of traditional detection methods, current technologies are shifting towards three-dimensional (3D) sensing. LiDAR (Light Detection and Ranging) has become a core sensor in advanced all-weather detection systems due to its ability to provide high-precision 3D point cloud data and its immunity to lighting conditions. Various deep learning methods have been developed for processing LiDAR point cloud data, including voxelization (such as PointPillars) and feature learning directly on the original point cloud (such as PointNet / PointNet++). These methods can more accurately detect and locate foreign objects such as fallen trees, rolled stones, or abandoned tools.

[0007] Furthermore, in practical applications for mobile platforms such as power grid inspection vehicles and emergency repair vehicles, the computational complexity, model size, and real-time inference speed of detection algorithms are crucial. Therefore, lightweight network design has become a current research focus. Existing technologies reduce model parameters and computational load by drawing on efficient network architectures such as GhostNet and MobileNet. For example, some studies have proposed lightweight models like RoadNetV2, which optimize the network structure to reduce floating-point computation, making it suitable for deployment in embedded devices and mobile terminals. Other methods combine YOLOv5 with GhostNet, significantly reducing the number of model parameters and improving detection speed (FPS) while maintaining detection accuracy, thus meeting the real-time detection needs of foreign objects on roads in power grid operation scenarios.

[0008] However, existing technologies face several key technical bottlenecks when dealing with target recognition tasks in complex visual scenes, especially in open and variable outdoor environments. These bottlenecks severely limit the accuracy and robustness of the detection models, mainly in the following aspects:

[0009] 1. The challenge of adaptability to multi-scale targets: In real power grid operation scenarios, the relative size of intruding objects and the background varies greatly, ranging from small stones in the distance to large obstacles nearby. This makes it difficult for traditional models to learn discriminative features at a single scale, easily leading to missed detections of overly small targets and misidentification of overly large targets, constituting a core obstacle to improving detection accuracy.

[0010] 2. Interference and camouflage issues in complex backgrounds: The background of roads at power grid operation sites is extremely complex. Foreign objects on the road are often interfered with or even camouflaged by changes in lighting, shadows, similar textures, or partial occlusion. Existing models struggle to effectively separate targets from the chaotic background during feature extraction, leading to feature confusion and a large number of misjudgments.

[0011] 3. Limitations of feature extraction for non-rigid and irregular targets: Non-rigid foreign objects with twisted and irregular shapes, such as plastic bags, bird nests, and scattered cables, are extremely common on roads in power grid operation sites. Traditional convolutional neural networks have inherent limitations in capturing their irregular contours and detailed features due to their inherent rectangular receptive field and regular sampling characteristics, resulting in insufficient feature extraction and a significant decrease in recognition ability.

[0012] Based on the above, a method for detecting foreign object intrusion on roads in power grid operation scenarios based on deep learning is invented. Summary of the Invention

[0013] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0014] A deep learning-based method for detecting foreign object intrusion on roads in power grid operation scenarios includes the following specific steps:

[0015] S1. The original road image of the power grid operation scene is cropped to a specified size and then input into the multi-scale road scene perception enhancement module MRPE to obtain shallow detail feature map F2, mid-level structure feature map F4, and deep semantic feature map F6.

[0016] S2, input F2, F4, and F6 into the cross-scale rich feature fusion module CS-RFF to obtain deep prediction feature map F12, mid-level prediction feature map F15, and full-scale feature map F17;

[0017] S3, input F12, F15, and F17 into the Multi-core Attention Enhancement Module (MAEM) to obtain an image with the detection results;

[0018] S4. Construct a foreign object intrusion detection model for power grid operation scenarios. The model consists of a multi-scale road scene perception enhancement module, a cross-scale rich feature fusion module, and a multi-core attention enhancement module. After the power grid operation scenario foreign object intrusion detection model is trained, it is applied to the actual power grid operation scenario. Finally, the foreign object intrusion detection results under the power grid operation road scenario are output, including the bounding box coordinates of the foreign object, the foreign object category label, and the confidence score.

[0019] As a preferred embodiment of the deep learning-based road foreign object intrusion detection method for power grid operation scenarios described in this invention, the specific steps of S1 are as follows:

[0020] S11, input the original power grid operation scene road image into the CBS submodule to obtain the initial transition feature map F1; then input F1 into the MARF submodule to obtain the shallow detail feature map F2;

[0021] S12, input the shallow detail feature map F2 into the CBS submodule to obtain the intermediate transition feature map F3; then input F3 into the MARF submodule to obtain the middle structure feature map F4;

[0022] S13, input F4 into the CBS submodule to obtain the deep transition feature map F5; then input F5 into the MARF submodule to obtain the deep semantic feature map F6.

[0023] As a preferred embodiment of the deep learning-based road foreign object intrusion detection method for power grid operation scenarios described in this invention, the execution process of the multi-scale context residual fusion submodule MARF is as follows:

[0024] First, the initial transition feature map F1 is processed in parallel in the following three branches:

[0025] In branch one, to capture contextual information under different receptive fields, F1 is first processed by a 3×3 dilated convolutional layer with an dilation rate of 1 to obtain a standard-scale contextual feature map A1; simultaneously, F1 is processed by a 3×3 dilated convolutional layer with an dilation rate of 3 to obtain a medium-scale contextual feature map A2; finally, F1 is processed by a 3×3 dilated convolutional layer with an dilation rate of 5 to obtain a large-scale contextual feature map A3.

[0026] In branch two, F1 first establishes the dependencies between feature channels through the dual-path context-aware submodule CCAM, and outputs a context-aware weighted feature map A4. Then, A4 is sequentially processed through a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation function to perform feature transformation and non-linear activation, resulting in channel transformation feature map A5, normalized feature map A6, and non-linear activation feature map A7. Finally, an upsampling operation is performed on A7 to restore its spatial resolution, and an upsampled attention feature map A8 is output.

[0027] Then, A1, A2, A3 and A8 are concatenated to achieve information interaction and obtain the spatial-channel concatenated feature map A9; A9 is fed into a 1×1 convolutional layer for feature dimensionality reduction to obtain the channel integrated feature map A10;

[0028] Finally, F1 and A10 in branch three are added element-wise through a residual connection, introducing residual gain while preserving the original feature information, thus obtaining the final shallow detail feature map F2.

[0029] As a preferred embodiment of the deep learning-based road foreign object intrusion detection method for power grid operation scenarios described in this invention, the execution process of the dual-path context-aware submodule CCAM is as follows:

[0030] In branch one, the initial transition feature map F1 is passed through a 3×3 depthwise separable convolution to efficiently extract spatial texture information, outputting a spatial texture feature map B3; then B3 is passed through the ReLU activation function to enhance the nonlinear expression, resulting in a nonlinear spatial feature map B4; B4 is then passed through a 1×1 convolutional layer for feature extraction, outputting a refined spatial feature map B5.

[0031] In branch two, the initial transition feature map F1 is first transformed by a 1×1 convolutional layer to output the weighted branch feature map B1; then B1 is fed into the SE module to learn and generate the attention weights between channels, outputting the channel attention weight feature map B2; B2 and B5 are multiplied element-wise to achieve dynamic modulation of spatial features by channel weights, resulting in the channel-modulated feature map B6; B6 is normalized by the Softmax function to obtain the normalized attention feature map B7, which is then deeply integrated through a 1×1 convolutional layer to output the deeply integrated attention feature map B8;

[0032] In branch three, F1 and B8 are added element-wise through a residual connection, and the final output is a context-aware weighted feature map A4.

[0033] As a preferred embodiment of the deep learning-based road foreign object intrusion detection method for power grid operation scenarios described in this invention, the specific steps of S2 are as follows:

[0034] In S21, in branch one, F6 is fed into the AMFM submodule of the adaptive multi-stream fusion module for feature extraction to obtain the refined deep semantic feature map F9; then F9 is transformed and integrated through a 1×1 convolutional layer to obtain F12.

[0035] In branch S22, F4 is fed into the AMFM submodule to obtain the refined mid-level structural feature map F8. F8 is simultaneously fed into two different convolutional layers to extract features at different scales: in the upper path, F8 passes through a 1×1 convolutional layer for channel transformation and information integration to obtain the mid-level local information feature map F10; in the lower path, F8 passes through a 3×3 convolutional layer to extract richer local spatial features to obtain the mid-level spatial context feature map F11. F12 from branch one undergoes an upsampling operation to obtain the upsampled deep feature map F13. Subsequently, F10, F11, and F13 are concatenated to obtain the cross-level fusion feature map F14. F14 is passed through a BN+SiLU module for normalization and non-linear activation to generate the mid-level prediction feature map F15. F15 undergoes an upsampling operation to generate the upsampled fusion feature map F16.

[0036] In branch S23, the shallow detail feature map F2 is sent to the AMFM submodule for enhancement, and the refined shallow detail feature map F7 is obtained after processing. F7 and F16 are then spliced ​​together to finally generate the full-scale fused feature map F17.

[0037] As a preferred embodiment of the deep learning-based road foreign object intrusion detection method for power grid operation scenarios described in this invention, the execution process of the adaptive multi-stream fusion submodule AMFM is as follows:

[0038] In branch one, F2 first obtains a preliminary convolutional feature map C1 through a 3×3 standard convolution; then, C1 is passed through a 3×3 reparameterized convolution to obtain a reparameterized convolutional feature map C2; and finally, C2 is passed through a 3×3 deformable convolution to obtain a multi-scale contextual feature map C3.

[0039] In branch two, F2 enters the adaptive channel attention enhancement submodule ACAE to aggregate contextual information and outputs the calibration weight feature map C4; then C3 and C4 are multiplied element-wise to output the attention weighted feature map C5.

[0040] C5 undergoes feature transformation via a 1×1 convolution, outputting a multi-branch fused feature map C6; C6 is then processed by a 3×3 deformable convolution to obtain a refined parallel feature map C7.

[0041] Finally, C6, C7 and F2 from branch 3 are added element-wise through residual connection to fuse the feature information extracted from different branches, and finally output the processing result of the entire module, refining the shallow detail feature map F7.

[0042] As a preferred embodiment of the deep learning-based road foreign object intrusion detection method for power grid operation scenarios described in this invention, the execution process of the adaptive channel attention enhancement submodule ACAE is as follows:

[0043] In branch one, F2 first obtains a pointwise convolutional feature map D1 through a 1×1 convolutional layer; then, it obtains a scale-corrected feature map D2 through batch normalization; then, it obtains a non-linear feature map D3 through the ReLU activation function; subsequently, D3 is fed into a 3×3 convolutional layer to obtain a depth-expanded feature map D4; and then, it obtains a normalized spatial feature map D5 through batch normalization; finally, it is processed through the ReLU activation function to output the backbone convolutional feature map D6.

[0044] In branch 2, F2 first passes through a 3×3 convolutional layer to output the weighted branch convolutional feature map D7; then D7 is fed into the ECA submodule to efficiently capture cross-channel interaction information and generate the attention activation feature map D8; finally, the channel attention weight feature map D9 is generated through the Sigmoid function.

[0045] Multiply D6 and D9 element-wise to adaptively weight the channel dimensions of the spatial feature map, resulting in a channel-weighted feature map D10. D10 is then passed through a 1×1 convolutional layer for channel adjustment, outputting a refined weighted feature map D11.

[0046] In branch three, F2 performs channel adjustment through a 3×3 convolutional layer to obtain the parallel convolutional feature map D12;

[0047] Finally, D11 and D12 are added element-wise through a residual connection to output the processing result of the entire module and calibrate the weight feature map C4.

[0048] As a preferred embodiment of the deep learning-based road foreign object intrusion detection method for power grid operation scenarios described in this invention, the specific steps of S3 are as follows:

[0049] S31, after upsampling F12, we obtain the deep scale-aligned feature map F18; we stitch F18 and F15 together to output the cross-layer cascaded feature map F19; we upsample F19 again to obtain the high-resolution aggregated feature map F20; we stitch F20 and F17 together to obtain the panoramic context feature map F21.

[0050] S32, F21 enters a 3×3 convolutional layer, outputting the initial context convolutional feature map F22; then it enters another 3×3 convolutional layer for processing, obtaining the multi-path distribution feature map F23; F23 is simultaneously fed into three SE-Conv modules with channel attention mechanisms and different kernel sizes: in branch one, F23 enters a 3×3 SE-Conv convolutional kernel for processing, obtaining the local receptive field feature map F24; in branch two, F23 enters a 5×5 SE-Conv convolutional kernel... The Conv convolutional kernel is used to process the medium receptive field feature map F25. In branch three, F23 is processed by a 7×7 SE-Conv convolutional kernel to obtain the global receptive field feature map F26. F24, F25, and F26 are concatenated to output the mixed-scale feature map F27. F27 is then processed by a 1×1 convolutional layer for channel adjustment to output the cross-channel reconstruction feature map F28. F28 and F22 are added element-wise through residual connection to obtain the reconstructed representation feature map F29.

[0051] S33, after processing F29 with the ReLU activation function, the residual enhanced context feature map F30 is obtained; F30 is then passed through a 3×3 convolutional layer for final feature transformation, outputting the decoded information feature map F31, and finally F31 is sent to the detection head to generate an image with foreign object intrusion detection results.

[0052] Compared with existing technologies:

[0053] When dealing with diverse intrusion targets, this invention can more accurately capture the features of foreign objects of different sizes, shapes, and postures appearing on roads in power grid operation scenarios, solving the problem of detection instability caused by the variability of road foreign objects in existing technologies. Addressing the complex and variable road environment in power grid operation scenarios, this invention introduces a dual attention mechanism of space and channel, effectively suppressing interference from complex backgrounds such as shadows and reflections, significantly improving the saliency and feature discrimination of intrusion targets. To achieve reliable identification of foreign objects of different distances and sizes on roads, this invention utilizes multi-scale dilated convolution, multi-kernel parallel attention, and deep and shallow feature fusion to fully integrate contextual information from different receptive fields, solving the problem of decreased detection accuracy caused by insufficient feature representation in existing technologies. Furthermore, to cope with various irregular foreign object intrusions that may occur in power grid operation scenarios, this invention combines deformable convolution and reparameterizable convolution in the AMFM module, enabling more efficient extraction of target geometric and structural features, enhancing adaptability to distorted or non-rigid targets, thereby significantly improving target detection accuracy and robustness in complex scenarios. Attached Figure Description

[0054] Figure 1 This is a schematic diagram of the process of the present invention;

[0055] Figure 2 This is a structural diagram of the multi-scale road scene perception enhancement module of the present invention;

[0056] Figure 3 This is a structural diagram of the multi-scale contextual residual fusion submodule of the present invention;

[0057] Figure 4 This is a structural diagram of the dual-path context-aware submodule of the present invention;

[0058] Figure 5 This is a structural diagram of the cross-scale rich feature fusion module of the present invention;

[0059] Figure 6 This is a structural diagram of the adaptive multi-stream fusion submodule of the present invention;

[0060] Figure 7 This is a structural diagram of the adaptive channel attention enhancement submodule of the present invention;

[0061] Figure 8 This is a structural diagram of the multi-core attention enhancement module of the present invention;

[0062] Figure 9 This is a diagram showing the overall structure of the road foreign object intrusion detection model for power grid operation scenarios according to the present invention.

[0063] Figure 10 This is an image of foreign objects on the road during a power grid operation scenario according to the present invention;

[0064] Figure 11 This is an image of a road with foreign objects in a power grid operation scenario, showing the detection results of the present invention. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0066] This invention provides a deep learning-based method for detecting foreign object intrusion on roads during power grid operations. Please refer to [link / reference]. Figures 1-11 The specific steps are as follows:

[0067] S1 involves cropping the original road image of the power grid operation scene to a specified size and then inputting it into the Multi-scale Road Perception Enhancement Module (MRPE) to obtain shallow detail feature map F2, mid-level structural feature map F4, and deep semantic feature map F6. The structure of the multi-scale road scene perception enhancement module is as follows: Figure 2 As shown.

[0068] The specific steps of S1 are as follows:

[0069] S11, input the original power grid operation scene road image into the CBS submodule to obtain the initial transition feature map F1; then input F1 into the MARF submodule to obtain the shallow detail feature map F2.

[0070] This invention designs a multi-scale contextual residual fusion submodule, MARF (Multi-scale AwareResidual Fusion module). The structure diagram of the MARF submodule is shown below. Figure 3 As shown.

[0071] The execution process of the MARF submodule is as follows:

[0072] First, the initial transition feature map F1 is processed in parallel in the following three branches:

[0073] In branch one, to capture contextual information under different receptive fields, F1 is first processed by a 3×3 dilated convolutional layer with a dilation rate of 1 to obtain a standard-scale contextual feature map A1; simultaneously, F1 is processed by a 3×3 dilated convolutional layer with a dilation rate of 3 to obtain a medium-scale contextual feature map A2; finally, F1 is processed by a 3×3 dilated convolutional layer with a dilation rate of 5 to obtain a large-scale contextual feature map A3.

[0074] In branch two, F1 first establishes the dependencies between feature channels through the cross-contextual attention module (CCAM), outputting a context-aware weighted feature map A4. Then, A4 is sequentially processed through a 3×3 convolutional layer, a batch normalization (BN) layer, and a ReLU activation function for feature transformation and non-linear activation, resulting in channel transformation feature map A5, normalized feature map A6, and non-linear activation feature map A7. Finally, an upsampling operation is performed on A7 to restore its spatial resolution, outputting an upsampled attention feature map A8.

[0075] Then, A1, A2, A3 and A8 are concatenated to achieve information interaction and obtain the spatial-channel concatenated feature map A9; A9 is fed into a 1×1 convolutional layer for feature dimensionality reduction to obtain the channel integrated feature map A10;

[0076] Finally, F1 and A10 in branch three are added element-wise through a residual connection, introducing residual gain while preserving the original feature information, thus obtaining the final shallow detail feature map F2.

[0077] The process of applying MARF to process F1 to obtain F2 is shown in equations (1), (2), and (3) below:

[0078]

[0079]

[0080]

[0081] in:

[0082] The CBS submodule (Conv–BN–SiLU Module) is a very common basic unit in deep learning networks. It can realize the complete process from feature extraction to standardization and then to nonlinear mapping, and is usually used as a basic building block in the network.

[0083] Batch normalization (BN) is a method for normalizing the input of each layer when training a deep neural network. It stabilizes the data distribution by adjusting the mean of features to 0 and the variance to 1, and then restores the expressive power through learnable scaling and offset parameters.

[0084] This invention designs a dual-path context-aware residual fusion module (CCAM module), the structure of which is shown in the figure below. Figure 4 As shown.

[0085] The execution process of the dual-path context-aware submodule is as follows:

[0086] In branch one, the initial transition feature map F1 is passed through a 3×3 depthwise separable convolution (DWConv) to efficiently extract spatial texture information, outputting a spatial texture feature map B3; then B3 is passed through the ReLU activation function to enhance the nonlinear expression, resulting in a nonlinear spatial feature map B4; B4 is then passed through a 1×1 convolutional layer for feature extraction, outputting a refined spatial feature map B5.

[0087] In branch two, the initial transition feature map F1 is first transformed by a 1×1 convolutional layer to output the weighted branch feature map B1; then B1 is fed into the SE module to learn and generate the attention weights between channels, outputting the channel attention weight feature map B2; B2 and B5 are multiplied element-wise to achieve dynamic modulation of spatial features by channel weights, resulting in the channel-modulated feature map B6; B6 is normalized by the Softmax function to obtain the normalized attention feature map B7, which is then deeply integrated through a 1×1 convolutional layer to output the deeply integrated attention feature map B8;

[0088] In branch three, F1 and B8 are added element-wise through a residual connection, and the final output is a context-aware weighted feature map A4.

[0089] The process of using CCAM to process F1 to obtain A4 is shown in equations (4) and (5) below:

[0090]

[0091] Among them, SE is a channel attention mechanism that compresses the feature map through global average pooling to obtain channel-level statistical information, then learns the importance weights of each channel through a fully connected network, and applies the weights to the original features to enhance key channels and suppress redundant channels.

[0092] S12, input the shallow detail feature map F2 into the CBS submodule to obtain the intermediate transition feature map F3; then input F3 into the MARF submodule to obtain the middle structure feature map F4.

[0093] S13, input F4 into the CBS submodule to obtain the deep transition feature map F5; then input F5 into the MARF submodule to obtain the deep semantic feature map F6.

[0094] S1 Example:

[0095] The original power grid operation scene image is cropped to 512×512 pixels with 3 channels and input to the multi-scale road scene perception enhancement module. After passing through a CBS submodule, it becomes the initial transition feature map F1, which is 256×256 pixels with 64 channels. F1 is then input to the MARF submodule for deep feature enhancement, resulting in the shallow detail feature map F2, which retains its original size of 256×256 pixels with 64 channels. F2 is then input to the second CBS submodule, resulting in the intermediate transition feature map F3, which is 128×128 pixels with 128 channels. F3 is then input to the MARF submodule for processing, resulting in the mid-level structure feature map F4, which is also 128×128 pixels with 128 channels. Finally, F4 is input to the third CBS submodule, resulting in the deep transition feature map F5, which is 64×64 pixels with 256 channels. Finally, F5 is input into the MARF submodule to obtain the deep semantic feature map F6, which is 64×64 pixels in size and has 256 channels.

[0096] In the MARF submodule, the initial transition feature map F1, with a size of 256×256 pixels and 64 channels, is processed in parallel. In branch one, F1 is first processed by a 3×3 dilated convolutional layer with a dilation rate of 1 to obtain a standard-scale context feature map A1, with a size of 256×256 pixels and 16 channels; simultaneously, F1 is processed by a 3×3 dilated convolutional layer with a dilation rate of 3 to obtain a medium-scale context feature map A2, with a size of 256×256 pixels and 16 channels; finally, F1 is processed by a 3×3 dilated convolutional layer with a dilation rate of 5 to obtain a large-scale context feature map A3, with a size of 256×256 pixels and 16 channels.

[0097] In branch two, F1 enters the dual-path context-aware submodule (CCAM) for feature weighting, resulting in a context-aware weighted feature map A4, with a size of 256×256 pixels and 64 channels. A4 is then passed through a 3×3 convolutional layer with a stride of 2 to obtain a channel-transformed feature map A5, with a size of 128×128 pixels and 16 channels. Batch normalization (BN) is then performed to obtain a normalized feature map A6, with a size of 128×128 pixels and 16 channels. Subsequently, a ReLU activation function is applied to obtain a non-linear activation feature map A7, maintaining its size of 128×128 pixels and 16 channels. Finally, an upsampling operation restores A7 to its original spatial dimensions, resulting in a 256×256 pixel, 16-channel feature map. The channel-upsampled attention feature map A8 is generated. Then, a concatenation operation (Concat) is used to concatenate A1, A2, A3, and A8 along the channel dimension to obtain a spatial-channel concatenated feature map A9, which is 256×256 pixels and has 64 channels. A9 is input into a 1×1 convolutional layer for information integration and channel preservation, outputting a channel-integrated feature map A10, which is 256×256 pixels and has 64 channels. Finally, element-wise addition is performed through a residual connection, adding F1 from branch three to A10 element-wise through the residual connection to obtain the final shallow detail feature map F2, which is 256×256 pixels and has 64 channels.

[0098] In the CCAM submodule, the initial transition feature map F1, with a size of 256×256 pixels and 64 channels, is processed in parallel to enhance its representational power. First, F1 enters branch one and passes through a 3×3 depthwise separable convolution (DWConv) to obtain a depthwise convolution feature map B3, with a size of 256×256 pixels and 64 channels. B3 is then processed by the ReLU activation function to obtain the content branch feature map B4, with a size of 2526×56 pixels and 64 channels. Next, B4 is convolved with a 1×1 kernel to generate the weighted feature map B5, with a size of 256×256 pixels and 64 channels. Simultaneously, F1 enters branch two, where it undergoes a convolution operation with a 1×1 kernel to obtain the weight branch feature map B1, which is 256×256 pixels in size and has 64 channels. B1 is then input into the SE (Squeeze-and-Excitation) module for processing, resulting in the channel attention weight feature map B2, which is also 256×256 pixels in size and has 64 channels. Subsequently, B5 and B2 are multiplied element-wise to form the channel enhancement feature map B6, which is 256×256 pixels in size and has 64 channels. B6 is then processed using the Softmax function to obtain the spatial attention feature map B7, which is also 256×256 pixels in size and has 64 channels. The features are then enhanced again using a 1×1 convolutional kernel to obtain the attention branch feature map B8, which is 256×256 pixels in size and has 64 channels. Finally, the initial transition feature map F1 of the three branch inputs is added element-wise through residual connections to obtain the final context-aware weighted feature map A4, which is 256×256 pixels in size and has 64 channels.

[0099] S2, input F2, F4, and F6 into the Cross-Scale Rich Feature Fusion Module (CS-RFF) to obtain the deep prediction feature map F12, the mid-level prediction feature map F15, and the full-scale feature map F17; the structure diagram of the CS-RFF module is shown below. Figure 5 As shown.

[0100] The CS-RFF module's role is to simultaneously capture the fine edge features of small foreign objects and accurately understand the overall outline of large foreign objects in road monitoring within power grid operation scenarios by fusing low-level high-resolution paths with high-level semantic paths. The CS-RFF module can adapt to targets of different scales and maintain robust recognition capabilities for twisted, rotated, and non-rigid intrusive foreign objects.

[0101] The specific steps of S2 are as follows:

[0102] S21, in Figure 5In branch one, F6 is fed into the AMFM submodule of the adaptive multi-stream fusion module for feature extraction, resulting in a refined deep semantic feature map F9; then F9 is transformed and integrated through a 1×1 convolutional layer to obtain F12.

[0103] This invention designs an Adaptive Multi-stream Fusion Module (AMFM), the structure of which is shown in the diagram below. Figure 6 As shown.

[0104] The execution process of the AMFM submodule is as follows:

[0105] In branch one, F2 first obtains a preliminary convolutional feature map C1 through a 3×3 standard convolution; then, C1 is passed through a 3×3 reparameterized convolution (RepConv) to obtain a reparameterized convolutional feature map C2; and finally, C2 is passed through a 3×3 deformable convolution (DConv) to obtain a multi-scale contextual feature map C3.

[0106] In branch two, F2 enters the Adaptive Channel Attention Enhancement Module (ACAE) to aggregate contextual information and outputs the calibration weight feature map C4; then C3 and C4 are multiplied element-wise to output the attention-weighted feature map C5.

[0107] C5 undergoes feature transformation via a 1×1 convolution, outputting a multi-branch fused feature map C6; C6 is then processed by a 3×3 deformable convolution to obtain a refined parallel feature map C7.

[0108] Finally, C6, C7 and F2 from branch 3 are added element-wise through residual connection to fuse the feature information extracted from different branches, and finally output the processing result of the entire module, refining the shallow detail feature map F7.

[0109] The process of applying AMFM to process F2 to obtain F7 is shown in equations (6) and (7) below:

[0110] (6)

[0111]

[0112] This invention designs an Adaptive Channel Attention Enhancement Module (ACAE), the structure of which is shown in the diagram below. Figure 7 As shown.

[0113] The execution process of the adaptive channel attention enhancement submodule is as follows:

[0114] In branch one, F2 first obtains a pointwise convolutional feature map D1 through a 1×1 convolutional layer; then it obtains a scale-corrected feature map D2 through batch normalization (BN); then it obtains a non-linear feature map D3 through the ReLU activation function; subsequently, D3 is fed into a 3×3 convolutional layer to obtain a depth-expanded feature map D4; and then it obtains a normalized spatial feature map D5 through batch normalization (BN); finally, it is processed through the ReLU activation function to output the backbone convolutional feature map D6.

[0115] In branch 2, F2 first passes through a 3×3 convolutional layer to output the weighted branch convolutional feature map D7; then D7 is fed into the ECA submodule to efficiently capture cross-channel interaction information and generate the attention activation feature map D8; finally, the channel attention weight feature map D9 is generated through the Sigmoid function.

[0116] Multiply D6 and D9 element-wise to adaptively weight the channel dimensions of the spatial feature map, resulting in a channel-weighted feature map D10. D10 is then passed through a 1×1 convolutional layer for channel adjustment, outputting a refined weighted feature map D11.

[0117] In branch three, F2 performs channel adjustment through a 3×3 convolutional layer to obtain the parallel convolutional feature map D12;

[0118] Finally, D11 and D12 are added element-wise through a residual connection to output the processing result of the entire module and calibrate the weight feature map C4.

[0119] Among them, the ECA submodule (Efficient Channel Attention) is an efficient channel attention mechanism. It obtains channel information through global average pooling and uses one-dimensional convolution to model the dependency relationship between adjacent channels in the channel dimension, thereby reducing the number of parameters and computational overhead while avoiding feature compression loss.

[0120] The process of applying ACAE to process F2 to obtain C4 is shown in equations (8) and (9) below:

[0121]

[0122]

[0123] S22, in Figure 5In branch two, F4 is fed into the AMFM submodule to obtain the refined mid-level structural feature map F8. F8 is simultaneously fed into two different convolutional layers to extract features at different scales: in the upper path, F8 passes through a 1×1 convolutional layer for channel transformation and information integration to obtain the mid-level local information feature map F10; in the lower path, F8 passes through a 3×3 convolutional layer to extract richer local spatial features to obtain the mid-level spatial context feature map F11. F12 from branch one undergoes an upsampling operation to obtain the upsampled deep feature map F13. Subsequently, F10, F11, and F13 are concatenated to obtain the cross-level fusion feature map F14. F14 is passed through a BN+SiLU module for normalization and non-linear activation to generate the mid-level prediction feature map F15. F15 then undergoes an upsampling operation to generate the upsampled fusion feature map F16.

[0124] S23, in Figure 5 In branch three, the shallow detail feature map F2 is fed into the AMFM submodule for enhancement, and the refined shallow detail feature map F7 is obtained after processing; F7 and F16 are concatted to finally generate the full-scale fused feature map F17.

[0125] S2 Implementation Example:

[0126] In the multi-scale feature fusion module, F2, F4, and F6 undergo bottom-up information integration and enhancement. In branch one, the deep semantic feature map F6, with a size of 64×64 pixels and 256 channels, is processed through the AMFM submodule to obtain a refined deep semantic feature map F9, also with a size of 64×64 pixels and 256 channels. F9 is then processed through a 1×1... The convolutional kernels are convolved to adjust the channels, resulting in a deep predictive feature map F12, which is 64×64 pixels and has 64 channels. Upsampling is then performed on F12 to obtain an upsampled deep feature map F13, which is 128×128 pixels and has 64 channels. In branch two, the 128×128 pixel, 128 channel mid-level structure feature map F4 is enhanced through the AMFM submodule to obtain a refined mid-level structure feature map F8, which is 128×128 pixels and has 128 channels. F8 is then input in parallel into a 1×1 convolution and a 3×1 convolution. A 3×3 convolution is used to obtain the mid-level local information feature map F10, which is 128×128 pixels and has 64 channels, through a 1×1 convolution. A 3×3 convolution is then used to obtain the mid-level spatial context F11, which is also 128×128 pixels and has 64 channels. F10, F11, and F13 are concatenated along the channel dimension to form the cross-level fused feature map F14, which is 128×128 pixels and has 192 channels. Batch normalization (BN) and SiLU activation are applied to F14 to obtain the mid-level prediction feature map F15, which is 128×128 pixels. Pixels, 192 channels; Upsample F15 to obtain an upsampled fusion feature map F16 with a size of 256×256 pixels and 192 channels; In branch three, the shallow detail feature map F2 with a size of 256×256 pixels and 64 channels is processed through the AMFM submodule to obtain a refined shallow detail feature map F7 with a size of 256×256 pixels and 64 channels; F7 and F16 are concatenated in the channel dimension to obtain a full-scale fusion feature map F17 with a size of 256×256 pixels and 256 channels.

[0127] In the AMFM submodule, the shallow detail feature map F2, with a size of 256×256 pixels and 64 channels, is processed. In branch one, F2 is passed through a regular 3×3 convolution (Conv) to obtain the initial convolutional feature map C1, with a size of 256×256 pixels and 128 channels. C1 is then passed through a reparameterized 3×3 convolution (RepConv) to obtain the reparameterized convolutional feature map C2, with a size of 256×256 pixels and 128 channels. C2 is then passed through a deformable 3×3 convolution (DConv) to obtain the multi-scale context feature map C3, with a size of 256×256 pixels and 128 channels. In branch two, F2 enters the ACAE submodule to generate the calibration weight feature map C4, with a size of 256×256 pixels and 128 channels. C3 and C4 are multiplied element-wise to generate the attention-weighted feature map C5, with a size of 256×256 pixels. The first branch has 128 channels. C5 is input into a 1×1 convolution, outputting a multi-branch fusion feature map C6, which has 256×256 pixels and 64 channels. C6 is then fed into a 3×3 deformable convolution for adaptive feature alignment and enhancement, resulting in a parallel dilated convolution feature map C7, which has 256×256 pixels and 64 channels. Finally, C6, C7, and F2 from branch three are added element-wise through a residual connection to obtain a refined shallow detail feature map F7, which has 256×256 pixels and 64 channels.

[0128] In the ACAE submodule, the shallow detail feature map F2, with a size of 256×256 pixels and 64 channels, is processed. In branch one, F2 is convolved with a 1×1 kernel to obtain a pointwise convolutional feature map D1, with a size of 256×256 pixels and 64 channels. Batch normalization (BN) is applied to D1 to obtain a scale-corrected feature map D2, with a size of 256×256 pixels and 64 channels. ReLU activation is applied to D2 to obtain a non-linear feature map D3, with a size of 256×256 pixels and 64 channels. A 3×3 kernel is then applied to D3 to obtain a depth expansion. Feature map D4 has a size of 256×256 pixels and 128 channels. Batch normalization is applied to D4 to obtain normalized spatial feature map D5, which has a size of 256×256 pixels and 128 channels. ReLU activation is applied to D5 to obtain backbone convolution feature map D6, which has a size of 256×256 pixels and 128 channels. In branch two, F2 goes through a 3×3 convolution expansion channel to obtain weight branch convolution feature map D7, which has a size of 256×256 pixels and 128 channels. Inputting D7 into the ECA submodule for channel information interaction yields the attention-incentivized feature map D8, which is 256×256 pixels and has 128 channels. D8 is then processed using the Sigmoid activation function to generate the channel attention weight feature map D9, also 256×256 pixels and has 128 channels. D6 and D9 are multiplied element-wise to generate the channel weighted feature map D10, which is 256×256 pixels and has 128 channels. D10 is then processed through a 1×1 convolution kernel to obtain the refined weighted feature map D11, which is 256×256 pixels and has 128 channels. In branch three, F2 is processed through a 3×3 convolution kernel to obtain the parallel convolution feature map D12, which is 256×256 pixels and has 128 channels. D11 and D12 are then added element-wise through a residual connection to obtain the calibration weight feature map C4, which is 256×256 pixels and has 128 channels.

[0129] S3, inputs F12, F15, and F17 into the Multi-core Attention Enhancement Module (MAEM) to obtain an image with the detection results. The MAEM module structure diagram is shown below. Figure 8 As shown.

[0130] The MAEM module utilizes its parallel multi-size attention convolutions to adaptively enhance key feature information of foreign objects and effectively suppress complex background noise in road monitoring within power grid operation scenarios. This module not only significantly improves the recognition accuracy of targets at different scales but also maintains the robustness of model detection even when targets and backgrounds are highly mixed.

[0131] The specific steps of S3 are as follows:

[0132] S31, after upsampling F12, we obtain the deep scale-aligned feature map F18; concatenate F18 and F15 to output the cross-layer cascaded feature map F19; upsample F19 again to obtain the high-resolution aggregated feature map F20; concatenate F20 and F17 to obtain the panoramic context feature map F21.

[0133] S32, F21 enters a 3×3 convolutional layer, outputting the initial context convolutional feature map F22; then it enters another 3×3 convolutional layer for processing, obtaining the multi-path distribution feature map F23; F23 is simultaneously fed into three SE-Conv modules with channel attention mechanisms and different kernel sizes: in branch one, F23 enters a 3×3 SE-Conv convolutional kernel for processing, obtaining the local receptive field feature map F24; in branch two, F23 enters a 5×5 SE-Conv convolutional kernel... The Conv convolutional kernel is used to process the medium receptive field feature map F25. In branch three, F23 is processed by a 7×7 SE-Conv convolutional kernel to obtain the global receptive field feature map F26. F24, F25, and F26 are concatenated to output the mixed-scale feature map F27. F27 is then processed by a 1×1 convolutional layer for channel adjustment to output the cross-channel reconstruction feature map F28. F28 and F22 are added element-wise through residual connection to obtain the reconstructed representation feature map F29.

[0134] S33. After processing F29 with the ReLU activation function, the residual enhanced context feature map F30 is obtained. F30 is then passed through a 3×3 convolutional layer for final feature transformation, outputting the decoded information feature map F31. F31 is then sent to the detection head to generate an image with foreign object intrusion detection results.

[0135] The foreign object intrusion detection results include the bounding box coordinates of the foreign object, the foreign object category label, and the confidence score.

[0136] The Head (detection head) is located in the final stage of the model. Its main function is to transform the deep features extracted by the backbone network and feature fusion module into specific detection results. It typically includes a classification branch to determine the target category and a regression branch to predict the target's position and size in the image. In some tasks, it can also be extended to output target confidence, key points, or segmentation masks. The Head is the core part of the network that directly generates detection or segmentation results.

[0137] Example:

[0138] The deep predictive feature map F12, with a size of 64×64 pixels and 64 channels, is upsampled once to obtain a deep scale-aligned feature map F18 with a size of 128×128 pixels and 64 channels. F18 and F15 (192 channels) are concatenated along the channel dimension to output a cross-layer cascaded feature map F19, with a size of 128×128 pixels and 256 channels. F19 is upsampled again to obtain a high-resolution aggregated feature map F20 with a size of 256×256 pixels and 256 channels. F20 and F17 are concatenated along the channel dimension to obtain a panoramic context feature map F21, with a size of 256×256 pixels and 512 channels.

[0139] Then, F21 is subjected to a 3×3 convolution (Conv) operation to reduce the channel dimension, resulting in the initial context convolutional feature map F22, which is 256×256 pixels and has 256 channels. F22 is then subjected to another 3×3 convolution operation to obtain the multi-path distribution feature map F23, which is 256×256 pixels and has 256 channels. F23 is fed into three branches in parallel for processing. In branch one, a 3×3 SE-Conv convolution operation is applied to F23 to generate a local receptive field feature map F24, which has a size of 256×256 pixels and 64 channels. In branch two, a 5×5 SE-Conv convolution operation is applied to F23 to generate a medium receptive field feature map F25, which has a size of 256×256 pixels and 64 channels. In branch three, a 7×7 SE-Conv convolution operation is applied to F23 to generate a global receptive field feature map F26, which has a size of 256×256 pixels and 64 channels.

[0140] F24, F25, and F26 are concatenated along the channel dimension to output a mixed-scale feature map F27, which is 256×256 pixels and has 192 channels. F27 is then subjected to a 1×1 convolution operation to increase its channel dimension, resulting in a cross-channel reconstructed feature map F28, which is 256×256 pixels and has 256 channels. F28 and F22 are then added element-wise through a residual connection to obtain a corrected reconstructed feature map F29, which is 256×256 pixels and has 256 channels. F29 is then processed with a ReLU activation function to obtain a residual-enhanced context feature map F30, which is 256×256 pixels and has 256 channels. Finally, F30 is subjected to a 3×3 convolution for final feature optimization, resulting in the decoded information feature map F31, which is 256×256 pixels and has 256 channels. The F31 sensor is fed into the subsequent head to generate the final image with the foreign object bounding box coordinates, foreign object category label, and confidence score.

[0141] S4. Construct a foreign object intrusion detection model for power grid operation scenarios. This model consists of a multi-scale road scene perception enhancement module, a cross-scale rich feature fusion module, and a multi-core attention enhancement module. After training, the model is applied to actual power grid operation scenarios, ultimately outputting foreign object intrusion detection results, including the bounding box coordinates of the foreign object, the object category label, and the confidence score. The overall structure of the power grid operation scenario foreign object intrusion detection model is as follows: Figure 9 As shown.

[0142] The multi-scale road scene perception enhancement module can adaptively capture the features of road intrusion targets of various sizes in power grid operation scenarios, solving the problem of missed or false detections caused by the diversity of target scales. Simultaneously, facing the complex background environment of power grid operation sites, this module can enhance the expression of effective foreign object features and suppress background noise interference, solving the problem of foreign object targets being easily confused with the background. This ensures that the method described in this invention can achieve high-precision and highly robust foreign object intrusion detection in practical applications.

[0143] Multi-scale road scene perception enhancement module: Input the original power grid operation scene road image into the multi-scale road scene perception enhancement module to obtain shallow detail feature map F2, mid-level structural feature map F4, and deep semantic feature map F6.

[0144] Cross-scale rich feature fusion module: Input F2, F4, and F6 into the cross-scale rich feature fusion module to obtain deep prediction feature map F12, mid-level prediction feature map F15, and full-scale feature map F17.

[0145] Multi-core attention enhancement module: Input F12, F15, and F17 into the multi-core attention enhancement module to obtain an image with detection results.

[0146] Example:

[0147] The system detects foreign objects in road images of power grid operation scenes, identifying water pipes and waste bags as foreign objects. The collected images are as follows: Figure 10 As shown.

[0148] Will Figure 11 As input to the road foreign object intrusion detection model for power grid operation scenarios, the output image contains the detection results, such as... Figure 11 As shown. In Figure 11 In the image, a rectangle is used to identify road debris in power grid operation scenarios, and the category of the debris and its confidence score are displayed above the rectangle.

[0149] in:

[0150] 1. Foreign object type: water pipe, confidence level 0.89, coordinates (25, 705, 830, 155).

[0151] 2. Foreign object type: plastic bag, confidence level 0.86, coordinates (940, 690, 260, 240).

[0152] Although the present invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the disclosed embodiments can be combined with each other in any manner. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for detecting foreign object intrusion on roads in power grid operation scenarios based on deep learning, characterized in that, The specific steps are as follows: S1. The original road image of the power grid operation scene is cropped to a specified size and then input into the multi-scale road scene perception enhancement module MRPE to obtain the shallow detail feature map F2, the mid-level structure feature map F4, and the deep semantic feature map F6. The specific steps are as follows: S11, input the original power grid operation scene road image into the CBS submodule to obtain the initial transition feature map F1; then input F1 into the MARF submodule to obtain the shallow detail feature map F2; S12, input the shallow detail feature map F2 into the CBS submodule to obtain the intermediate transition feature map F3; then input F3 into the MARF submodule to obtain the middle structure feature map F4; S13, input F4 into the CBS submodule to obtain the deep transition feature map F5; then input F5 into the MARF submodule to obtain the deep semantic feature map F6; S2, input F2, F4, and F6 into the cross-scale rich feature fusion module CS-RFF to obtain deep prediction feature map F12, mid-level prediction feature map F15, and full-scale feature map F17; S3. Input F12, F15, and F17 into the Multi-Kernel Attention Enhancement Module (MAEM) to obtain an image with the detection results. The specific steps are as follows: S31, after upsampling F12, we obtain the deep scale-aligned feature map F18; we stitch F18 and F15 together to output the cross-layer cascaded feature map F19; we upsample F19 again to obtain the high-resolution aggregated feature map F20; we stitch F20 and F17 together to obtain the panoramic context feature map F21. S32, F21 enters a 3×3 convolutional layer, outputting the initial context convolutional feature map F22; then it enters another 3×3 convolutional layer for processing, obtaining the multi-path distribution feature map F23; F23 is simultaneously fed into three SE-Conv modules with channel attention mechanisms and different kernel sizes: in branch one, F23 enters a 3×3 SE-Conv convolutional kernel for processing, obtaining the local receptive field feature map F24; in branch two, F23 enters a 5×5 SE-Conv convolutional kernel... The Conv convolutional kernel is used to process the medium receptive field feature map F25. In branch three, F23 is processed by a 7×7 SE-Conv convolutional kernel to obtain the global receptive field feature map F26. F24, F25, and F26 are concatenated to output the mixed-scale feature map F27. F27 is then processed by a 1×1 convolutional layer for channel adjustment to output the cross-channel reconstruction feature map F28. F28 and F22 are added element-wise through residual connection to obtain the reconstructed representation feature map F29. S33, after processing F29 with the ReLU activation function, the residual enhanced context feature map F30 is obtained; F30 is then passed through a 3×3 convolutional layer for final feature transformation, outputting the decoded information feature map F31, and finally F31 is sent to the detection head to generate an image with foreign object intrusion detection results; S4. Construct a foreign object intrusion detection model for power grid operation scenarios. The model consists of a multi-scale road scene perception enhancement module, a cross-scale rich feature fusion module, and a multi-core attention enhancement module. After the power grid operation scenario foreign object intrusion detection model is trained, it is applied to the actual power grid operation scenario. Finally, the foreign object intrusion detection results under the power grid operation road scenario are output, including the bounding box coordinates of the foreign object, the foreign object category label, and the confidence score.

2. The method for detecting foreign object intrusion in power grid operation scenarios based on deep learning according to claim 1, characterized in that, The execution process of the multi-scale context residual fusion submodule MARF is as follows: First, the initial transition feature map F1 is processed in parallel in the following three branches: In branch one, to capture contextual information under different receptive fields, F1 is first processed by a 3×3 dilated convolutional layer with an dilation rate of 1 to obtain a standard-scale contextual feature map A1; simultaneously, F1 is processed by a 3×3 dilated convolutional layer with an dilation rate of 3 to obtain a medium-scale contextual feature map A2; finally, F1 is processed by a 3×3 dilated convolutional layer with an dilation rate of 5 to obtain a large-scale contextual feature map A3. In branch two, F1 first establishes the dependencies between feature channels through the dual-path context-aware submodule CCAM, and outputs a context-aware weighted feature map A4. Then, A4 is sequentially processed through a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation function to perform feature transformation and non-linear activation, resulting in channel transformation feature map A5, normalized feature map A6, and non-linear activation feature map A7. Finally, an upsampling operation is performed on A7 to restore its spatial resolution, and an upsampled attention feature map A8 is output. Then, A1, A2, A3 and A8 are concatenated to achieve information interaction and obtain the spatial-channel concatenated feature map A9; A9 is fed into a 1×1 convolutional layer for feature dimensionality reduction to obtain the channel integrated feature map A10; Finally, F1 and A10 in branch three are added element-wise through a residual connection, introducing residual gain while preserving the original feature information, thus obtaining the final shallow detail feature map F2.

3. The method for detecting foreign object intrusion in a power grid operation scenario based on deep learning according to claim 2, characterized in that, The execution process of the dual-path context-aware submodule CCAM is as follows: In branch one, the initial transition feature map F1 is passed through a 3×3 depthwise separable convolution to efficiently extract spatial texture information, outputting a spatial texture feature map B3. B3 is then activated by ReLU to enhance the nonlinear representation, resulting in a nonlinear spatial feature map B4. B4 is then processed by a 1×1 convolutional layer to refine the features, outputting a refined spatial feature map B5. In branch two, the initial transition feature map F1 is first transformed by a 1×1 convolutional layer to output the weighted branch feature map B1; then B1 is fed into the SE module to learn and generate the attention weights between channels, outputting the channel attention weight feature map B2; B2 and B5 are multiplied element-wise to achieve dynamic modulation of spatial features by channel weights, resulting in the channel-modulated feature map B6; B6 is normalized by the Softmax function to obtain the normalized attention feature map B7, which is then deeply integrated through a 1×1 convolutional layer to output the deeply integrated attention feature map B8; In branch three, F1 and B8 are added element-wise through a residual connection, and the final output is a context-aware weighted feature map A4.

4. The method for detecting foreign object intrusion in a power grid operation scenario based on deep learning according to claim 1, characterized in that, The specific steps of S2 are as follows: In S21, in branch one, F6 is fed into the adaptive multi-stream fusion submodule AMFM for feature extraction to obtain the refined deep semantic feature map F9; then F9 is transformed and integrated through a 1×1 convolutional layer to obtain F12. S22, in branch two, F4 is sent to the AMFM submodule to obtain the refined middle layer structure feature map F8; F8 is fed into two different convolutional layers simultaneously to extract features at different scales: In the upper path, F8 passes through a 1×1 convolutional layer to perform channel transformation and information integration, resulting in the mid-level local information feature map F10; in the lower path, F8 passes through a 3×3 convolutional layer to extract richer local spatial features, resulting in the mid-level spatial context feature map F11. After upsampling, F12 in branch 1 is used to obtain the upsampled deep feature map F13; then, F10, F11 and F13 are concatenated to obtain the cross-level fused feature map F14. F14 is normalized and nonlinearly activated by a BN+SiLU module to generate the mid-level prediction feature map F15; F15 is then subjected to an upsampling operation to generate the upsampled fused feature map F16. In branch S23, the shallow detail feature map F2 is sent to the AMFM submodule for enhancement, and the refined shallow detail feature map F7 is obtained after processing. F7 and F16 are then spliced ​​together to finally generate the full-scale fused feature map F17.

5. The method for detecting foreign object intrusion on roads in power grid operation scenarios based on deep learning according to claim 4, characterized in that, The execution process of the adaptive multi-stream fusion submodule AMFM is as follows: In branch one, F2 first obtains a preliminary convolutional feature map C1 through a 3×3 standard convolution; then, C1 is passed through a 3×3 reparameterized convolution to obtain a reparameterized convolutional feature map C2; and finally, C2 is passed through a 3×3 deformable convolution to obtain a multi-scale contextual feature map C3. In branch two, F2 enters the adaptive channel attention enhancement submodule ACAE to aggregate contextual information and outputs the calibration weight feature map C4; then C3 and C4 are multiplied element-wise to output the attention weight feature map C5; C5 undergoes feature transformation via a 1×1 convolution, outputting a multi-branch fused feature map C6; C6 is then processed by a 3×3 deformable convolution to obtain a refined parallel feature map C7. Finally, C6, C7 and F2 from branch 3 are added element-wise through residual connection to fuse the feature information extracted from different branches, and finally output the processing result of the entire module, refining the shallow detail feature map F7.

6. The method for detecting foreign object intrusion on roads in power grid operation scenarios based on deep learning according to claim 5, characterized in that, The execution process of the Adaptive Channel Attention Enhancement (ACAE) submodule is as follows: In branch one, F2 first obtains a pointwise convolutional feature map D1 through a 1×1 convolutional layer; then, it obtains a scale-corrected feature map D2 through batch normalization; then, it obtains a non-linear feature map D3 through the ReLU activation function; subsequently, D3 is fed into a 3×3 convolutional layer to obtain a depth-expanded feature map D4; and then, it obtains a normalized spatial feature map D5 through batch normalization; finally, it is processed through the ReLU activation function to output the backbone convolutional feature map D6. In branch 2, F2 first passes through a 3×3 convolutional layer to output the weighted branch convolutional feature map D7; then D7 is fed into the ECA submodule to efficiently capture cross-channel interaction information and generate the attention activation feature map D8; finally, the channel attention weight feature map D9 is generated through the Sigmoid function. Multiply D6 and D9 element-wise to adaptively weight the channel dimensions of the spatial feature map, resulting in a channel-weighted feature map D10. D10 is then passed through a 1×1 convolutional layer for channel adjustment, outputting a refined weighted feature map D11. In branch three, F2 performs channel adjustment through a 3×3 convolutional layer to obtain the parallel convolutional feature map D12; Finally, D11 and D12 are added element-wise through a residual connection to output the processing result of the entire module and calibrate the weight feature map C4.

Citation Information

Patent Citations

  • Road abnormity alarm method and system in power transmission channel scene, and medium

    CN117809178A

  • Track foreign matter intrusion detection method based on driving video image

    CN118015575A