Oil and gas pipeline intrusion detection method and system based on bionic vision

By integrating a foveal attention mechanism and distributed learning to optimize bounding box regression, a biomimetic vision-based method for detecting intrusions into oil and gas pipelines was developed. This method addresses the issue of insufficient accuracy in detecting small targets in complex backgrounds and achieves high-precision and real-time intrusion detection.

CN120894571APending Publication Date: 2025-11-04CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511069391.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing pipeline intrusion detection methods are prone to false alarms in complex backgrounds, struggle to maintain detection accuracy for targets of different scales, have a high false negative rate for small targets, are susceptible to noise interference with traditional feature extraction methods, and lack sufficient localization accuracy with bounding box regression methods.

Method used

A biomimetic vision-based method for detecting intrusions into oil and gas pipelines is adopted. This method integrates a feature extraction backbone network based on the foveal attention mechanism, a feature fusion neck network, and a target detection network. It simulates the foveal vision mechanism of the human eye and improves detection accuracy and real-time performance by optimizing bounding box regression through multi-scale feature fusion and distribution learning.

Benefits of technology

It achieves high-precision detection of small targets, occlusions, and complex background scenes, improving detection accuracy and real-time performance. The foveal attention mechanism improves the detection accuracy of small targets by 3.2% and the bounding box regression accuracy by 5.8%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894571A_ABST
    Figure CN120894571A_ABST
Patent Text Reader

Abstract

The invention provides an oil and gas pipeline intrusion detection method and system based on bionic vision, relates to the technical field of computer vision and artificial intelligence safety monitoring, and aims to solve the problem that the positioning precision is insufficient due to scale variability of an intrusion target, complexity of a background environment and proneness to noise interference. According to the method, collected oil and gas pipeline related images are input into an improved feature extraction backbone network for feature extraction, and multi-scale key features are obtained; inputting the multi-scale key features into a feature fusion neck network for multi-scale feature fusion and enhancement to obtain hierarchical feature representation; and inputting the hierarchical feature representation into a target detection network for target detection to obtain a target detection result. The defects in the prior art are overcome, and the target positioning precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computer vision and artificial intelligence safety monitoring, and particularly relates to an oil and gas pipeline intrusion detection method and system based on biomimetic vision. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] Pipeline intrusion detection technology has become an important means of pipeline safety protection due to its wide coverage, fast response speed and relatively low cost.

[0004] At present, pipeline intrusion detection methods mainly include traditional image processing methods and deep learning methods. Among them, traditional methods such as background difference and optical flow method are prone to produce a large number of false positives in complex background, and are sensitive to light changes and weather conditions.

[0005] In addition, although deep learning methods have advantages in feature extraction and pattern recognition, existing general object detection algorithms still have certain deficiencies in the pipeline intrusion detection scene: due to the large scale of the target, the existing detection algorithm is difficult to balance the detection accuracy of different scale targets, and the small target miss detection rate is high; due to the complex background along the pipeline, including vegetation changes, terrain undulations, light differences and other interference factors, traditional feature extraction methods are easily disturbed by noise and are difficult to extract key features with strong discriminability; the existing bounding box regression method usually assumes that the error obeys a single distribution, but the actual scene often presents a multi-modal distribution, resulting in insufficient positioning accuracy. SUMMARY

[0006] In order to overcome the deficiencies of the prior art, the present application provides an oil and gas pipeline intrusion detection method and system based on biomimetic vision. In view of the characteristics of variable target scale, complex background and high positioning accuracy requirement in oil and gas pipeline intrusion detection, a feature extraction backbone network integrated with a foveal attention mechanism, a feature fusion neck network and a target detection network based on focal distribution learning are used to simulate the foveal vision mechanism of the human eye and optimize the bounding box regression distribution, realizing high-precision detection of small targets, occlusions and complex background scenes, and effectively improving the detection accuracy and real-time performance.

[0007] To achieve the above purpose, one or more embodiments of the present application provide the following technical solutions: The first aspect of the present application discloses an oil and gas pipeline intrusion detection method based on biomimetic vision, comprising: inputting the collected oil and gas pipeline related images into an improved feature extraction backbone network for feature extraction to obtain multi-scale key features; The multi-scale key features are input to a feature fusion neck network for multi-scale feature fusion and enhancement to obtain hierarchical feature representation. The hierarchical feature representation is input to a target detection network for target detection to obtain a target detection result. The improved feature extraction backbone network comprises a plurality of cascaded feature extraction modules, a foveal attention mechanism biomimetic visual perception module, and an enhanced dynamic controllable feature pyramid network connected in sequence. The oil and gas pipeline related images are sequentially input to the five cascaded feature extraction modules for feature extraction to obtain corresponding level features. The fifth level feature is input to the foveal attention mechanism biomimetic visual perception module, and feature extraction is performed through the global perception branch, the small-scale center perception branch, and the large-scale center perception branch, respectively, and then the features are fused to obtain the fused features. The fused features are input to the enhanced dynamic controllable feature pyramid network for feature enhancement and fusion to obtain multi-scale key features.

[0008] As a further technical solution, the oil and gas pipeline related images are sequentially input to the five cascaded feature extraction modules for feature extraction to obtain corresponding level features, and the specific process is as follows: The oil and gas pipeline related images are input to the first level feature extraction module, and feature extraction is performed through the convolution layer to obtain the primary feature. The primary feature is sequentially input to the second, third, fourth, and fifth level feature extraction modules, and feature fusion and strengthening are performed through multiple deep convolution fusion blocks to obtain the second, third, fourth, and fifth level features, respectively.

[0009] As a further technical solution, the second, third, fourth, and fifth level feature extraction modules each comprise a convolution layer and an efficient feature fusion module, and the primary feature is input to the second level feature extraction module, and the specific process is as follows: The primary feature is input to the convolution layer for feature extraction to obtain the second convolution feature. The second convolution feature is input to the efficient feature fusion module for splicing and fusion to obtain the second level feature.

[0010] As a further technical solution, the fifth level feature is input to the foveal attention mechanism biomimetic visual perception module, and the specific process is as follows: The fifth level feature is input to the global perception branch and compressed through adaptive average pooling to obtain a global context representation. The fifth level feature is input to the small-scale center perception branch and compressed through adaptive average pooling to obtain a center detail representation. The fifth-level feature is input into a large-scale center perception branch, compressed through adaptive average pooling, and a center representation is obtained; The global context representation, the center detail representation, and the center representation are concatenated and fused; The concatenated and fused feature is subjected to convolution and dimension reduction processing, and attention weights are generated through a convolution layer and an activation function; Based on the attention weights, the fused feature is obtained through adaptive residual connection fusion.

[0011] As a further technical solution, the fused feature is input into an enhanced dynamic controllable feature pyramid network for feature enhancement and fusion, wherein the enhanced dynamic controllable feature pyramid network includes a cross-path fusion module and an SE attention enhancement module, and the specific process is as follows: The fused feature is input into the cross-path fusion module, and multi-scale complementary enhanced features are obtained through bidirectional fusion from top to bottom and from bottom to top; The multi-scale complementary enhanced features are input into the SE attention enhancement module for feature enhancement, and multi-scale key features are obtained.

[0012] As a further technical solution, the fused feature is input into the cross-path fusion module, and the specific process is as follows: When passing through the top-down path, high-level semantic features are fused with low-level detail features through upsampling to obtain spliced and fused features; When passing through the bottom-up path, low-level detail features are fused with high-level features through downsampling to obtain added and fused features; The spliced and fused features and the added and fused features are enhanced through an activation function to obtain multi-scale complementary enhanced features.

[0013] As a further technical solution, the multi-scale key features are input into a feature fusion neck network for multi-scale feature fusion and enhancement, wherein the feature fusion neck network includes a first upsampling fusion module, a first neck feature processing module, a second upsampling fusion module, a second neck feature processing module, a first convolution fusion module, a third neck feature processing module, a second convolution fusion module, and a fourth neck feature processing module connected in sequence, and the specific process is as follows: The multi-scale key features and the fourth-level feature are input into the first upsampling fusion module for feature fusion, and first neck fusion features are obtained; The first neck fusion features are subjected to feature strengthening through the first neck feature processing module, and first processed features are obtained; The first processed features and the third-level feature are input into the second upsampling fusion module for feature fusion, and second neck fusion features are obtained; The second neck fusion feature is processed by a second neck feature processing module to strengthen the feature, and a first hierarchical feature representation is obtained; The first hierarchical feature representation and the first neck fusion feature are input into a first convolution fusion module for feature fusion, and a third neck fusion feature is obtained; The third neck fusion feature is processed by a third neck feature processing module to strengthen the feature, and a second hierarchical feature representation is obtained; The second hierarchical feature representation and the multi-scale key feature are input into a second convolution fusion module for feature fusion, and a fourth neck fusion feature is obtained; The fourth neck fusion feature is processed by a fourth neck feature processing module to strengthen the feature, and a third hierarchical feature representation is obtained.

[0014] As a further technical solution, the hierarchical feature representation is input into a target detection network for target detection, wherein the target detection network includes three parallel detection heads, and the specific process is as follows: The first hierarchical feature representation, the second hierarchical feature representation and the third hierarchical feature representation are input into the first detection head, the second detection head and the third detection head respectively, and target categories, target bounding box positions and prediction confidence are obtained; The target categories are optimized by binary cross-entropy loss, and the bounding box positions are optimized by focus distribution form learning loss; According to the optimization result, the final target detection result is obtained.

[0015] The second aspect discloses an oil and gas pipeline intrusion detection system based on biomimetic vision, comprising: A feature extraction backbone network module is used for inputting the collected oil and gas pipeline related images into an improved feature extraction backbone network for feature extraction, and obtaining multi-scale key features; A feature fusion neck network module is used for inputting the multi-scale key features into a feature fusion neck network for multi-scale feature fusion and enhancement, and obtaining a hierarchical feature representation; wherein the improved feature extraction backbone network includes a plurality of cascade feature extraction modules, a foveal attention mechanism biomimetic vision perception module and an enhanced dynamic controllable feature pyramid network connected in sequence, and in the improved feature extraction backbone network, the specific process is as follows: The oil and gas pipeline related images are input into the five cascade feature extraction modules in sequence for feature extraction, and corresponding level features are obtained; the fifth level feature is input into the foveal attention mechanism biomimetic vision perception module, and feature extraction is performed through a global perception branch, a small-scale center perception branch and a large-scale center perception branch respectively, and fusion is performed, and a fusion feature is obtained; the fusion feature is input into the enhanced dynamic controllable feature pyramid network for feature enhancement and fusion, and multi-scale key features are obtained; The target detection network module is configured to input the hierarchical feature representation into a target detection network to perform target detection and obtain a target detection result.

[0016] A third aspect of the present application provides a computer device comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor implements the steps in the method according to the first aspect of the present application when executing the program.

[0017] The above one or more technical solutions have the following beneficial effects: In the embodiment, in view of the variable target size, complex background and low positioning accuracy in oil and gas pipeline intrusion detection, an oil and gas pipeline intrusion target prediction model based on bionic visual perception is constructed, high-precision detection of small targets, occlusion and complex background scenes is realized, and the detection accuracy and real-time performance are effectively improved. Specifically, by constructing a bionic visual perception module based on a foveal attention mechanism, the visual characteristics of the fovea of the human eye are simulated, and adaptive perception of different scale intrusion targets is realized. At the same time, an enhanced dynamic controllable feature pyramid network is constructed, multi-scale feature fusion and channel attention mechanism are used to generate hierarchical feature representation with strong discriminability. The focal distribution learning loss function is used to convert the bounding box regression into a distribution learning problem, and the high-frequency consistency constraint and smoothness regularization are combined to improve the target positioning accuracy.

[0018] In the embodiment, by simulating the foveal visual mechanism of the human eye, a small-scale center perception branch is constructed, effectively solving the shortcomings of traditional attention mechanisms in small target detection. Compared with the traditional attention method, the foveal attention mechanism of the embodiment improves the small target detection accuracy by 3.2%.

[0019] In the embodiment, the bounding box regression problem is converted into a distribution learning problem, combined with the focal loss, high-frequency consistency and smoothness constraint, and the target positioning accuracy is significantly improved. The positioning accuracy (IoU>0.75) of the embodiment is improved by 5.8% compared with the traditional method.

[0020] The advantages of the additional aspects of the application will be partially given in the following description, partially will become obvious from the following description, or will be understood by the practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0021] The drawings accompanying the specification of the present application form a part thereof, serve to provide further understanding of the present application, and together with the description of the exemplary embodiments of the present application and the explanation thereof, to explain the present application, and do not constitute an improper limitation of the present application.

[0022] Figure 1 A flowchart of an oil and gas pipeline intrusion detection method based on bionic vision according to Embodiment One of the present application; Figure 2This is a schematic diagram of the foveal attention mechanism according to Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the enhanced dynamic controllable feature pyramid network of Embodiment 1 of the present invention; Figure 4 This is a flowchart illustrating the calculation of the focus distribution learning loss function in Embodiment 1 of the present invention; Figure 5 This is a three-stage schematic diagram of the progressive training strategy according to Embodiment 1 of the present invention; Figure 6 This is a diagram showing the detection results of Embodiment 1 of the present invention. Detailed Implementation

[0023] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0024] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0025] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0026] Example 1 This embodiment discloses an intrusion detection method for oil and gas pipelines based on biomimetic vision.

[0027] To more clearly illustrate this embodiment, a biomimetic vision-based oil and gas pipeline intrusion detection process can be specifically described as follows: This embodiment provides a biomimetic vision-based method for detecting intrusions into oil and gas pipelines, including: S1. Input the collected images related to oil and gas pipelines into the improved feature extraction backbone network for feature extraction to obtain multi-scale key features; S2. Input the multi-scale key features into the feature fusion neck network to perform multi-scale feature fusion and enhancement, and obtain a hierarchical feature representation; S3. Input the hierarchical feature representation into the object detection network to perform object detection and obtain the object detection result.

[0028] like Figures 1-3 As shown, in step S1, the collected images related to oil and gas pipelines are input into the improved feature extraction backbone network for feature extraction to obtain multi-scale key features.

[0029] S1-1. Construct a biomimetic vision-based target detection model for oil and gas pipeline intrusion.

[0030] In view of the problems of variable target size, complex background and insufficient positioning accuracy in oil and gas pipeline intrusion detection, an oil and gas pipeline intrusion target detection model based on bionic vision is constructed.

[0031] The target detection model comprises a feature extraction backbone network integrating a bionic vision perception module, a feature fusion neck network, and a target detection network based on distribution form learning.

[0032] The feature extraction backbone network introduces a foveal attention mechanism to extract bionic vision features of different scale intrusion targets in the input image; the feature fusion neck network is used to perform multi-scale fusion on features at different levels to generate hierarchical feature representation adaptive to the pipeline scene; and the target detection network adopts a focal distribution learning loss function to improve the positioning accuracy of the intrusion target and output the detection result.

[0033] S1-2, input the collected oil and gas pipeline related images into the improved feature extraction backbone network for feature extraction.

[0034] In this embodiment, the improved feature extraction backbone network comprises a plurality of cascaded feature extraction modules, a bionic vision perception module of foveal attention mechanism, and an enhanced dynamic controllable feature pyramid network connected in sequence.

[0035] In the improved feature extraction backbone network, the specific process is as follows: S1-2-1, input the oil and gas pipeline related images into the five cascaded feature extraction modules in sequence for feature extraction to obtain corresponding level features.

[0036] The specific process is as follows: (1) input the oil and gas pipeline related images into the first level feature extraction module, perform feature extraction through the convolution layer to obtain the primary features.

[0037] In this embodiment, the oil and gas pipeline related images are oil and gas pipeline along line pictures taken by a camera, i.e. original image data of the oil and gas pipeline monitoring scene.

[0038] The first level feature extraction module comprises a standard convolution layer, a batch normalization layer and an activation function, which are used for primary feature extraction.

[0039] The "convolution kernel" scans the image to extract the most basic features (such as edges, simple textures, etc.), and converts the original image into primary features.

[0040] (2) input the primary features into the second, third, fourth and fifth level feature extraction modules in sequence, perform feature fusion and strengthening through a plurality of deep convolution fusion blocks to obtain the second, third, fourth and fifth level features respectively.

[0041] The second to fourth level modules adopt a deep separable convolution structure, and feature maps are down-sampled level by level through different step size settings. The specific process is as follows: 1) Perform a deep convolution operation on the input feature map (primary feature), and the formula is as follows: (1) wherein, is a deep convolution kernel, and K is the size of the convolution kernel.

[0042] 2) The deep convolution output is fused through point-by-point convolution, and the formula is as follows: (2) wherein, is a point-by-point convolution kernel.

[0043] 3) Set different step sizes to control down-sampling Different step sizes s are set to realize the down-sampling of the feature map level by level: (3) 4) Batch normalization and activation After the deep separable convolution, batch normalization and an activation function are connected: (4) wherein, is a batch normalization operation, is an activation function.

[0044] The second, third, fourth and fifth level feature extraction modules each include a convolution layer and an efficient feature fusion module. The primary feature is input to the second level feature extraction module, and the specific process is as follows: 1) The primary feature is input to the convolution layer for feature extraction, and a second convolution feature is obtained.

[0045] The primary feature is subjected to a convolution operation to compress the size and extract the feature, such as combining edge features into more complex shape features.

[0046] 2) The second convolution feature is input to the efficient feature fusion module for splicing and fusion to obtain a second level feature.

[0047] The efficient feature fusion module splices and fuses the features output by different convolution layers, while the model retains multi-scale information (small details and large outlines) to avoid feature loss.

[0048] Multi-level feature extraction is adopted, and C2f and Con appear multiple times alternately, the image features are filtered and strengthened layer by layer, and complex semantics are extracted from the primary feature. This provides a prerequisite for constructing a feature description of an intrusion target.

[0049] Multi-scale feature maps are obtained through the second, third, fourth and fifth level feature extraction modules wherein B is batch size, C i C is the number of channels, H is the size of the feature map.

[0050] S1-2-2, the fifth-level feature is input into the central fovea attention mechanism biomimetic visual perception module, feature extraction is performed through a global perception branch, a small-scale central perception branch and a large-scale central perception branch respectively, and fusion is performed to obtain a fused feature.

[0051] As shown in Figure 2 , in the present embodiment, the specific process is as follows: (1) The fifth-level feature is input into the global perception branch, and a global context representation is obtained through adaptive average pooling compression.

[0052] The global perception branch adopts adaptive average pooling compression, and the formula is as follows: (5) wherein, C represents the number of channels, H is the image height, and W is the image width. The feature map is compressed to 1x1 size through adaptive average pooling, and the global context information is captured .

[0053] (2) The fifth-level feature is input into the small-scale central perception branch, and a central detail representation is obtained through adaptive average pooling compression.

[0054] The small-scale central perception branch has the formula as follows: (6) wherein, is the feature of the center 1 / 4 region of the feature map.

[0055] The center 1 / 4 region of the feature map is subjected to 4x4 adaptive average pooling, the center 1 / 4 region of the feature map is extracted, the high-resolution visual region of the fovea centralis of the human eye is simulated, and a central detail representation is obtained .

[0056] (3) The fifth-level feature is input into the large-scale central perception branch, and a central representation is obtained through adaptive average pooling compression.

[0057] The large-scale central perception branch has the formula as follows: (7) The center 7 / 8 region of the feature map is subjected to 8x8 adaptive average pooling, and the center 1 / 2 region of the feature map is extracted. The peripheral visual region of the fovea centralis of the human eye is simulated, and a central representation is obtained .

[0058] (4) Cascade fusion of global context representation, center detail representation and center representation.

[0059] The outputs of the three branches are cascade fused, and the formula is: (8).

[0060] (5) The cascade fused features are convolved and dimensionally reduced, and attention weights are generated through the convolution layer and the activation function.

[0061] The outputs of the three branches are concatenated in the channel dimension, passed through a convolution layer with a channel compression ratio of c / / 12, a SiLU activation function and a channel restoration convolution layer, and finally a spatial attention weight map is generated through a Sigmoid function.

[0062] The attention weights are generated through the convolution layer and the activation function, and the formula is: (9).

[0063] (6) Based on the attention weights, fusion is performed through adaptive residual connection to obtain fused features.

[0064] The final output is calculated through adaptive residual connection. Feature enhancement is achieved, and the formula is: (10) Where, a is a learnable parameter, and the initial value is set to 0.1, and represents element-wise multiplication.

[0065] Adaptive residual connection realizes feature enhancement through formula (10).

[0066] After the above steps, adaptive perception of different scale intrusion targets is realized, and the shortcomings of traditional attention mechanisms in small target detection are effectively solved.

[0067] S1-2-3, the fused features are input into the enhanced dynamic controllable feature pyramid network for feature enhancement and fusion to obtain multi-scale key features.

[0068] As shown in Figure 3 In this embodiment, the enhanced dynamic controllable feature pyramid network (Enhanced DCFP) adopts a multi-scale feature fusion architecture and integrates an SE attention mechanism to realize dynamic and controllable feature enhancement. The enhanced dynamic controllable feature pyramid network includes a cross-path fusion module and an SE attention enhancement module. The specific process is as follows: (1) The fused features are input into the cross-path fusion module, and multi-scale complementary enhanced features are obtained through bidirectional fusion from top to bottom and from bottom to top.

[0069] The specific process is as follows: 1) When top-down path transmission, high-level semantic features are fused with low-level detailed features through upsampling to obtain splicing fusion features and transmit semantic information.

[0070] In this embodiment, when top-down transmission, the P5 high semantic features output by the backbone network are compressed to 1 / 4 of the original channel number by 1*1 convolution, restored to the P2 feature map resolution by upsampling, and fused with the P2 high-resolution features by channel dimension splicing to obtain spliced fusion features.

[0071] Specifically, top-down feature propagation is performed to generate enhanced features, and the formula is: , , , , (11) wherein, represents a 2x upsampling operation, represents element-wise addition, represents a 1*1 convolution operation for channel number adjustment.

[0072] Starting from a high-level feature map, high-resolution features are transmitted to a low level through an upsampling operation. High-level features contain rich semantic information and can provide high-level understanding of target categories and positions. Finally, high-level semantic features are fused with low-level features through upsampling to transmit semantic information.

[0073] 2) When bottom-up path transmission, low-level detailed features are fused with high-level features through downsampling to obtain addition fusion features and retain detailed information.

[0074] In this embodiment, when bottom-up transmission, the P2 high-resolution features output by the backbone network are downsampled by 3*3 convolution with a step of 2, and are fused with the P5 low-resolution features by element-wise addition to obtain addition fusion features.

[0075] Specifically, bottom-up feature propagation is performed to generate enhanced features, and the formula is: , , , , (12) wherein, represents a 2x downsampling operation, represents a 3*3 convolution operation.

[0076] From the low layer feature map, the detailed features are transmitted to the high layer through the downsampling operation. The low layer features contain rich spatial detail information such as edge, texture and local features. Through the downsampling and fusion with the high layer features, the detailed information is transmitted upward, and the positioning accuracy of the high layer features is enhanced.

[0077] 3) The spliced fusion features and the added fusion features are enhanced by an activation function to obtain multi-scale complementary enhanced features.

[0078] In this embodiment, the final features of P2, P3, P4 and P5 layers are obtained respectively , that is, the complementary enhanced features of P2, P3, P4 and P5 are obtained respectively.

[0079] The enhanced dynamic controllable feature pyramid network adopts a bidirectional fusion structure to obtain multi-scale complementary enhanced features. The enhanced dynamic controllable feature pyramid network adopts a top-down and bottom-up bidirectional feature fusion strategy, and integrates a channel attention mechanism to weight the importance of the fusion features at the channel level. The network outputs four feature maps of different scales.

[0080] (2) The multi-scale complementary enhanced features are input into the SE attention enhancement module for feature enhancement to obtain multi-scale key features.

[0081] The specific formula is: (13) wherein, represents the SE attention module.

[0082] The enhanced dynamic controllable feature pyramid network outputs four feature maps of different scales, corresponding to 4 times, 8 times, 16 times and 32 times downsampling rates, to adapt to the detection needs of targets of different scales.

[0083] Through the above steps, the deficiencies of the traditional single-path feature pyramid network in semantic information transmission and detail preservation are solved, and higher quality multi-scale key feature representation is provided for subsequent target detection.

[0084] As shown in FIG. 2, in step S2, the multi-scale key features are input into the feature fusion neck network for multi-scale feature fusion and enhancement to obtain hierarchical feature representation. Figure 1

[0085] The feature fusion neck network includes a first upsampling fusion module, a first neck feature processing module, a second upsampling fusion module, a second neck feature processing module, a first convolution fusion module, a third neck feature processing module, a second convolution fusion module and a fourth neck feature processing module connected in sequence, and the specific process is as follows: ​(1) Input the multi-scale key features and the fourth-level features into the first upsampling fusion module to perform feature fusion and obtain the first neck fusion features.

[0086] (2) The first neck fusion feature is enhanced by the first neck feature processing module to obtain the first processed feature.

[0087] (3) Input the first processing feature and the third level feature into the second upsampling fusion module to perform feature fusion and obtain the second neck fusion feature.

[0088] The second neck feature processing module enhances the fused features of the second neck to obtain the first hierarchical feature representation.

[0089] (4) Input the first hierarchical feature representation and the first neck fusion feature into the first convolutional fusion module to perform feature fusion and obtain the third neck fusion feature.

[0090] (5) The third neck fusion feature is enhanced by the third neck feature processing module to obtain the second-level feature representation.

[0091] (6) Input the second-level feature representation and multi-scale key features into the second convolutional fusion module for feature fusion to obtain the fourth neck fusion feature.

[0092] (7) The fourth neck fusion feature is enhanced by the fourth neck feature processing module to obtain the third-level feature representation.

[0093] Through the above steps, the three hierarchical feature representations cover the full-scale detection needs from small targets to large targets; the features at different levels complement each other to form a complete feature representation system; compared with a single feature representation, multi-scale hierarchical features significantly improve detection accuracy, especially the ability to detect small targets; multi-scale feature fusion improves the model's adaptability to complex backgrounds and lighting changes.

[0094] like Figure 1 As shown, in step S3, the hierarchical feature representation is input into the target detection network to perform target detection and obtain the target detection result.

[0095] The target detection network consists of three parallel detection heads, and the specific process is as follows: S3-1. Input the first-level feature representation, the second-level feature representation, and the third-level feature representation into the first detection head, the second detection head, and the third detection head, respectively, to obtain the target category, the target bounding box position, and the prediction confidence.

[0096] In this embodiment, the first hierarchical feature representation is input to the first detection head for medium-scale target detection.

[0097] The second hierarchical feature representation is input to a second detection head for small-scale target detection.

[0098] The third hierarchical feature representation is input to a third detection head for micro-scale target detection.

[0099] In this embodiment, the specific process is as follows: (1) Convert the four coordinate values of the bounding box into discrete distribution prediction to obtain a prediction tensor.

[0100] Convert the four coordinate values of the bounding box into discrete distribution prediction to obtain a prediction tensor , where R = 16 is the regression maximum value.

[0101] (2) Perform softmax normalization processing on the predicted bounding box, respectively.

[0102] The normalization formula is: (14) wherein, , represents the position index, represents the data point; k represents the class index, represents the possible output; R represents the number of categories, defines the summation range; P dist represents the original input value (log score); P soft represents the output probability value.

[0103] (3) Calculate the prediction confidence The formula is: (15) wherein, target value, aggregated result.

[0104] S3-2, optimize the target category through binary cross-entropy loss, and optimize the bounding box position through focus distribution form learning loss.

[0105] As shown in Figure 4 , in this embodiment, the focus distribution learning loss function includes three core components: a focus distribution loss component, which enhances the attention of the model to difficult-to-detect intrusion targets by introducing a modulation factor 1-p t ) γ , wherein p t is the prediction confidence, and γ is the focus parameter, with a default value of 2.0; a high-frequency distribution consistency loss component, which ensures that the gradient change pattern of the bounding box regression is consistent with the real target by calculating the first-order difference of the prediction distribution and the target distribution; and a distribution smoothness constraint component, which prevents excessive oscillation of the bounding box prediction by minimizing the variance of the prediction distribution.

[0106] (1) Calculate binary cross-entropy loss.

[0107] The formula is: (16) wherein, predicted probability value, is a stability coefficient, is a cross-entropy loss value.

[0108] (2) Calculate the focal point distribution learning loss.

[0109] 1) Focal point distribution loss component.

[0110] Calculate the application focal point modulation, the formula is: (17) wherein, focusing parameters, focal point loss weight.

[0111] Focal point distribution loss component, the formula is: (18) wherein, is the focal point loss, (1 / N) is the normalization factor.

[0112] 2) High-frequency distribution consistency loss component.

[0113] Calculate the gradient of the predicted distribution, the formula is: (19) wherein, probability tensor value, difference result.

[0114] Calculate the gradient of the target distribution, the formula is: (20) wherein, output gradient value, forward difference.

[0115] Use mean square error to measure consistency, high-frequency distribution consistency loss, the formula is: (21) wherein, loss function, N normalization coefficient, difference measure.

[0116] 3) Distribution smoothness constraint component.

[0117] Calculate the variance of the predicted distribution, the formula is: (22) where, is the predicted variance for position (i,j), R is the total number of classes, k is the summation index, predicted probability value for position (i,j) and class k, predicted probability mean value for position (i,j).

[0118] The target distribution variance is calculated, and the formula is: (23) where, , predicted variance for position, predicted probability mean value for position.

[0119] The smoothness constraint formula is: (24) where, smoothness loss value.

[0120] 4) Calculate the total loss of the focal point distribution learning loss function.

[0121] (25) where, a is the high-frequency loss weight, the default value is 1.2; β is the smoothness loss weight, the default value is 0.8; total loss value, focal loss, weight coefficient of high-fidelity loss, high-fidelity loss, weight coefficient of smoothness loss.

[0122] (3) According to the total loss of the focal point distribution learning loss function, through the adaptive learning mechanism, the optimization result is obtained.

[0123] The weight parameter is dynamically adjusted through the adaptive learning mechanism, and the formula is: (26) where, is the learnable parameter, total loss of the focal point distribution learning loss function.

[0124] After the above steps, the boundary box regression optimization of distribution shape learning converts the boundary box regression problem into a distribution learning problem, combined with the focal loss, high-frequency consistency and smoothness constraint, which significantly improves the target positioning accuracy. Experimental results show that the positioning accuracy (IoU>0.75) is improved by 5.8% compared with traditional methods.

[0125] S3-3, according to the optimization result, the final target detection result is obtained.

[0126] Assuming that the model predicts the left upper corner x-coordinate of a target's bounding box, instead of directly outputting a numerical value, it outputs a probability distribution of length 16 (such as [0.01, 0.02,..., 0.15, 0.20, 0.10,...], summing to 1), where each position represents the probability of the coordinate falling within that interval. By weighted averaging the distribution, the probability distribution is restored to a specific coordinate value. During training, the model not only uses cross-entropy (focal loss) to make the predicted distribution closer to the true distribution, but also uses high-frequency consistency loss and distribution smoothness regularization to ensure that the distribution is both accurate and smooth, avoiding abnormal spikes or excessive dispersion. For the left upper x, left upper y, right lower x, and right lower y of each target, the model outputs a distribution, which is decoded into the final bounding box coordinates.

[0127] As shown in Figure 5 , S4, the target detection model is trained.

[0128] On the standard COCO dataset and the self-built pipeline intrusion dataset, training is performed according to steps S1 to S3. The experimental environment uses an NVIDIA RTX A5000 GPU, and training is performed for 150 rounds.

[0129] An incremental training strategy is adopted, dividing the 150 rounds of training into three stages: The first stage (1-45 rounds) accounts for the first 30% of the total training rounds, and strong data augmentation is used, including mosaic=1.0, mixup=0.2, and random erasing=0.4, with the learning rate increasing linearly from 0.001 to 0.01.

[0130] The second stage (46-105 rounds) accounts for the middle 40% of the total training rounds, and moderate data augmentation is used, with mosaic=0.8, mixup=0.1, and random erasing=0.2, and the learning rate remains at 0.01.

[0131] The third stage (106-150 rounds) accounts for the last 30% of the total training rounds, and data augmentation is turned off to focus on model convergence. The learning rate is decayed from 0.01 to 0.0001.

[0132] Performance improvement of the incremental training strategy: Through three-stage incremental training and intelligent learning rate scheduling, the model training process is optimized, and the convergence speed and final performance are improved. Experiments show that after adopting the incremental training strategy, the model convergence speed is improved by 25%, and the final mAP is improved by 2.1%.

[0133] After training, the model is converted to ONNX format and deployed in the pipeline monitoring system to realize real-time intrusion detection.

[0134] S5, the target detection model is verified and ablation experiment is carried out.

[0135] To verify the effectiveness of the method, a large number of comparative experiments are carried out on the standard COCO dataset and the self-built pipeline intrusion dataset, and the method is compared with YOLOv8n (baseline), YOLOv5s and YOLOv7-tiny. The specific performance comparison is shown in Table 1.

[0136] Table 1 Comparison of performance of different methods

[0137] As can be seen from Table 1, the method reaches the optimal level in various performance indicators: mAP50-95 reaches 94.38%, which is 1.01 percentage points higher than the baseline method; the precision reaches 99.33%, and the recall rate reaches 97.57%; while maintaining high detection accuracy, the inference speed still reaches 155 FPS, meeting the real-time detection requirements.

[0138] This embodiment verifies the contribution of different components to the detection performance. Through ablation experiment, the effectiveness of each module is analyzed: Baseline model: standard YOLOv8n network, mAP50-95 is 93.37%.

[0139] Add foveal attention: mAP50-95 is improved to 94.15%, which is improved by 0.78 percentage points.

[0140] Add FDFL loss: due to the characteristics of the dataset, the performance remains basically unchanged.

[0141] Complete MTD2 model: combined with all components, mAP50-95 reaches 94.38%, which is 1.01 percentage points higher than the baseline.

[0142] The results show that the foveal attention mechanism is the main contributor to performance improvement, and the FDFL loss function can further optimize the positioning accuracy on specific datasets.

[0143] Embodiment two The embodiment discloses an oil and gas pipeline intrusion detection system based on bionic vision, comprising: A feature extraction backbone network module is used to input the collected oil and gas pipeline related images into an improved feature extraction backbone network for feature extraction to obtain multi-scale key features. The feature fusion neck network module is configured to input the multi-scale key features into the feature fusion neck network for multi-scale feature fusion and enhancement to obtain hierarchical feature representation. The improved feature extraction backbone network comprises a plurality of cascaded feature extraction modules, a foveal attention mechanism biomimetic visual perception module and an enhanced dynamic controllable feature pyramid network connected in sequence. In the improved feature extraction backbone network, the specific process is as follows: The oil and gas pipeline related images are sequentially input into the five cascaded feature extraction modules for feature extraction to obtain corresponding level features. The fifth level features are input into the foveal attention mechanism biomimetic visual perception module, and feature extraction is performed through the global perception branch, the small-scale center perception branch and the large-scale center perception branch respectively, and then the fusion features are obtained by fusion. The fusion features are input into the enhanced dynamic controllable feature pyramid network for feature enhancement and fusion to obtain multi-scale key features. The target detection network module is configured to input the hierarchical feature representation into the target detection network for target detection to obtain a target detection result.

[0144] The method steps in the first embodiment are implemented based on the arrhythmia auxiliary diagnosis system providing fused spatio-temporal features and physiological constraints. The foveal vision-based oil and gas pipeline intrusion detection system in the embodiment is deployed in an environment covering a 50-kilometer pipeline section, and 25 monitoring points are set, each of which is configured with a binocular camera.

[0145] Detection performance: continuous operation for 6 months, cumulative processing of 2.4 million images, detection accuracy of 94.4%, false positive rate of 2.3%, and false negative rate of 1.3%.

[0146] Response efficiency: average detection delay of 85 milliseconds, meeting the real-time monitoring requirements; automatic early warning response time of 3 seconds, significantly improving the safety protection efficiency.

[0147] Economic benefits: compared with traditional manual inspection, labor cost is reduced by 65%; by timely discovering and preventing illegal construction, 3 potential pipeline safety accidents are avoided, and pipeline operation safety is ensured.

[0148] Embodiment three The purpose of the embodiment is to provide a computer device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above method.

[0149] Embodiment four The purpose of the embodiment is to provide a computer readable storage medium.

[0150] A computer readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to perform the steps of the above method.

[0151] Those skilled in the art should understand that the modules or steps of the present application described above can be realized by a general computer device, or alternatively, they can be realized by program codes executable by a computing device, so that they can be stored in a storage device and executed by a computing device, or they can be respectively manufactured into individual integrated circuit modules, or a plurality of modules or steps among them can be manufactured into a single integrated circuit module. The present application is not limited to any specific combination of hardware and software.

[0152] The specific embodiments of the present application described above in conjunction with the accompanying drawings are not intended to limit the protection scope of the present application. Those skilled in the art should understand that various modifications or variations made on the basis of the technical solutions of the present application without creative labor are still within the protection scope of the present application.

Claims

1. A method for oil and gas pipeline intrusion detection based on biomimetic vision, characterized in that, The method comprises the following steps: inputting the collected oil and gas pipeline related images into an improved feature extraction backbone network for feature extraction to obtain multi-scale key features; inputting the multi-scale key features into a feature fusion neck network for multi-scale feature fusion and enhancement to obtain hierarchical feature representation; inputting the hierarchical feature representation into a target detection network for target detection to obtain target detection results; wherein the improved feature extraction backbone network comprises a plurality of cascaded feature extraction modules, a foveal attention mechanism biomimetic visual perception module and an enhanced dynamic controllable feature pyramid network connected in sequence, and in the improved feature extraction backbone network, the specific process is as follows: inputting the oil and gas pipeline related images into the five cascaded feature extraction modules in sequence for feature extraction to obtain corresponding level features; inputting the fifth level feature into the foveal attention mechanism biomimetic visual perception module, performing feature extraction through a global perception branch, a small-scale center perception branch and a large-scale center perception branch respectively, and performing fusion to obtain fusion features; inputting the fusion features into the enhanced dynamic controllable feature pyramid network for feature enhancement and fusion to obtain multi-scale key features.

2. The biomimetic vision based intrusion detection method for oil and gas pipelines of claim 1, wherein, The specific process of inputting the oil and gas pipeline related images into the five cascaded feature extraction modules in sequence for feature extraction to obtain corresponding level features is as follows: inputting the oil and gas pipeline related images into the first level feature extraction module, performing feature extraction through a convolution layer to obtain primary features; inputting the primary features into the second, third, fourth and fifth level feature extraction modules in sequence, performing feature fusion and strengthening through a plurality of deep convolution fusion blocks to obtain the second, third, fourth and fifth level features respectively.

3. A biomimetic vision based intrusion detection method for oil and gas pipelines as claimed in claim 2, wherein, The second, third, fourth and fifth level feature extraction modules all comprise a convolution layer and an efficient feature fusion module, and the specific process of inputting the primary features into the second level feature extraction module is as follows: inputting the primary features into the convolution layer for feature extraction to obtain second convolution features; inputting the second convolution features into the efficient feature fusion module for splicing and fusion to obtain the second level features.

4. The biomimetic vision based intrusion detection method for oil and gas pipelines as claimed in claim 1, wherein, The specific process of inputting the fifth level feature into the foveal attention mechanism biomimetic visual perception module is as follows: inputting the fifth level feature into the global perception branch, performing adaptive average pooling compression to obtain global context representation; inputting the fifth level feature into the small-scale center perception branch, performing adaptive average pooling compression to obtain center detail representation; inputting the fifth level feature into the large-scale center perception branch, performing adaptive average pooling compression to obtain center representation; concatenating the global context representation, the center detail representation and the center representation; performing convolution and dimension reduction processing on the concatenated fusion features, and generating attention weights through a convolution layer and an activation function; based on the attention weights, performing fusion through an adaptive residual connection to obtain fusion features.

5. The biomimetic vision based intrusion detection method for oil and gas pipelines as claimed in claim 1, wherein, The specific process of inputting the fusion features into the enhanced dynamic controllable feature pyramid network for feature enhancement and fusion is as follows: inputting the fusion features into the cross-path fusion module, performing bidirectional fusion from top to bottom and from bottom to top to obtain multi-scale complementary enhanced features; The multi-scale key features are input into the SE attention enhancement module for feature enhancement to obtain multi-scale complementary enhanced features.

6. A biomimetic vision based intrusion detection method for oil and gas pipelines as claimed in claim 5, wherein, The fusion features are input into the cross-path fusion module for top-down and bottom-up fusion, and the specific process is as follows: In the top-down path transmission, the high-level semantic features are fused with the low-level detail features through upsampling to obtain splicing fusion features; In the bottom-up path transmission, the low-level detail features are fused with the high-level features through downsampling to obtain addition fusion features; The splicing fusion features and the addition fusion features are enhanced through an activation function to obtain multi-scale complementary enhanced features.

7. The biomimetic vision based intrusion detection method for oil and gas pipelines as claimed in claim 1, wherein, The multi-scale key features are input into the feature fusion neck network for multi-scale feature fusion and enhancement, wherein the feature fusion neck network comprises a first upsampling fusion module, a first neck feature processing module, a second upsampling fusion module, a second neck feature processing module, a first convolution fusion module, a third neck feature processing module, a second convolution fusion module, and a fourth neck feature processing module connected in sequence, and the specific process is as follows: The multi-scale key features and the fourth-level features are input into the first upsampling fusion module for feature fusion to obtain first neck fusion features; The first neck fusion features are processed by the first neck feature processing module to obtain first processing features; The first processing features and the third-level features are input into the second upsampling fusion module for feature fusion to obtain second neck fusion features; The second neck fusion features are processed by the second neck feature processing module to obtain first hierarchical feature representation; The first hierarchical feature representation and the first neck fusion features are input into the first convolution fusion module for feature fusion to obtain third neck fusion features; The third neck fusion features are processed by the third neck feature processing module to obtain second hierarchical feature representation; The second hierarchical feature representation and the multi-scale key features are input into the second convolution fusion module for feature fusion to obtain fourth neck fusion features; The fourth neck fusion features are processed by the fourth neck feature processing module to obtain third hierarchical feature representation.

8. The biomimetic vision based intrusion detection method for oil and gas pipelines as claimed in claim 1, wherein, The hierarchical feature representation is input into the target detection network for target detection, wherein the target detection network comprises three parallel detection heads, and the specific process is as follows: The first hierarchical feature representation, the second hierarchical feature representation, and the third hierarchical feature representation are input into the first detection head, the second detection head, and the third detection head respectively to obtain target categories, target bounding box positions, and prediction confidence; The target categories are optimized by binary cross-entropy loss, and the bounding box positions are optimized by focal distribution morphology learning loss; According to the optimization results, the final target detection results are obtained.

9. A biomimetic vision based intrusion detection system for oil and gas pipelines characterized in that, The feature extraction backbone network module is used to input the collected oil and gas pipeline related images into the improved feature extraction backbone network for feature extraction to obtain multi-scale key features. ​ The feature fusion neck network module is used for inputting the multi-scale key features into the feature fusion neck network for multi-scale feature fusion and enhancement to obtain hierarchical feature representation. The oil and gas pipeline related images are sequentially input into the five cascaded feature extraction modules for feature extraction to obtain corresponding level features; the fifth level feature is input into the foveal attention mechanism biomimetic visual perception module, feature extraction is performed through the global perception branch, the small-scale center perception branch and the large-scale center perception branch respectively, and fusion is performed to obtain the fusion feature; the fusion feature is input into the enhanced dynamic controllable feature pyramid network for feature enhancement and fusion to obtain the multi-scale key features. The target detection network module is used for inputting the hierarchical feature representation into the target detection network for target detection to obtain the target detection result.

10. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the method of any one of claims 1-8 when executing the program.