Multi-modal fusion oil and gas pipeline perimeter security and protection method and system and storage medium
Through the fusion of radar and camera data, a multi-modal fusion oil and gas pipeline perimeter security method is constructed, and information interaction is enhanced by using MaSA, RCS-OSA and MLIA modules, solving the problems of high false alarm rate and insufficient identification accuracy in the existing technology, and achieving high accuracy intrusion detection in complex environments.
Patent Information
- Application Number
- CN202510855469.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The lack of a high-accuracy oil and gas pipeline perimeter security method based on multimodal fusion of radar and image/video in the prior art, resulting in high false alarm rates and insufficient accuracy in identifying intrusion events in complex environments.
The point cloud data of the perimeter environmental of the oil and gas pipeline is collected through radar, combined with the camera to collect image data, perform feature extraction and fusion, and build a target detection model based on YOLO, introduce MaSA, RCS-OSA and MLIA modules to enhance information interaction capabilities, and optimize model training using DynamicFocaler-IoU loss function.
In complex scenarios such as low light and occlusion, the false alarm rate is significantly reduced, the target detection accuracy is improved, and the intrusion incidents are ensured in a timely manner.
Smart Images

Figure CN120375141A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of oil and gas pipeline security, and in particular, to a multi-modal fusion-based perimeter security method, system, and storage medium for oil and gas pipelines. Background Art
[0002] The perimeters of oil and gas pipelines often span complex terrains such as deserts and mountains. With the continuous progress of perimeter security technologies for oil and gas pipelines, perimeter security systems are gradually enhancing their functions and can quickly identify and respond to various intrusion behaviors. However, most devices are interfered by environmental factors, resulting in a high false alarm rate, and often rely on a single detection method when identifying intrusion events.
[0003] Currently, some devices can only rely on radar detection. Although they can provide long-distance detection and the ability to penetrate obstacles, they have limitations in target recognition and classification. Some devices only rely on cameras for visual recognition. Although they can provide rich color and shape information, their performance is affected under harsh weather and lighting conditions.
[0004] In recent years, there have been many studies on oil and gas pipelines. For example, the prior art with the document number CN 113206978 A discloses an intelligent monitoring and early warning system and method for the security of oil and gas pipeline stations, which automatically monitors the entire pipeline in real time from multiple aspects by setting up a video intelligent recognition module, a perimeter optical fiber security module, a personnel positioning and health monitoring module, and a pipeline monitoring module. There is a lack of relevant research on the perimeter security of oil and gas pipelines with high accuracy and low false alarm rate based on the multi-modal fusion of radar and image / video in the prior art. Summary of the Invention
[0005] The technical problem to be solved by the present invention is: There is a lack of a high-accuracy perimeter security method for oil and gas pipelines based on the multi-modal fusion of radar and image / video in the prior art.
[0006] The technical solution adopted by the present invention to solve the above technical problem: The present invention provides a multi-modal fusion-based perimeter security method for oil and gas pipelines, including the following steps: Step S100: Collect point cloud data of the environment of the perimeter of the oil and gas pipeline through radar, determine suspicious targets according to the radar point cloud data, and enable the camera to track the targets and collect image data of the targets; Step S200: Perform dimensionality reduction processing on the point cloud data, extract features from the dimensionality-reduced point cloud data and the image data respectively, and then perform feature fusion to obtain multi-modal fusion data; Step S300: Construct a YOLO-based object detection model. The MaSA attention module and the RCS-OSA module are introduced into the Head network of the model. The MaSA attention module is used to introduce an explicit spatial prior in the visual backbone network by constructing a two-dimensional bidirectional spatial attenuation matrix, enabling the model to focus on target features. The RCS-OSA module is used for multi-scale feature extraction and feature aggregation based on the dynamic structure reparameterization ability and the cross-channel interaction mechanism of ShuffleNet. The MLIA attention module is also introduced into the backbone network of the model. The MLIA attention module is used to calculate the local importance of pixels based on integrating local importance learning and the channel gating mechanism, enabling the module to have second-order information interaction ability. The YOLO-based object detection model uses a loss function based on DynamicFocaler-IoU to help the model better learn to extract features from medium-difficulty samples. Step S400: Use the YOLO-based object detection model to detect intrusion targets in the multi-modal fusion data.
[0007] Further, in step S100, the process of making the camera track the target and collect the image data of the target is as follows: Dual radars are used to collect the point cloud data of the perimeter environment of the oil and gas pipeline, and a camera is connected to each radar. First, calculate the distances of the target from radar 1 and radar 2, and then calculate the angles of the target relative to the cameras: ; where a is the distance of the target from radar 1, b is the distance of the target from radar 2, c is the distance between the two radars, α and β are the angles between camera 1 and camera 2 and the target respectively, and are the initial angles of camera 1 and camera 2 respectively, and are the rotation angles of camera 1 and camera 2 respectively.
[0008] Further, in S200, the process of reducing the dimensionality of the point cloud data is as follows: The mapping transformation neural network is specifically used to reduce the dimensionality of the point cloud data. The mapping transformation neural network is based on the BP neural network and adds a convolutional layer and a residual module. The convolutional layer is used to extract the features of the radar data, and the residual module is used to combine the deep features and the shallow features to reduce the increase in the number of parameters and the performance degradation caused by the increase in the network depth.
[0009] Further, the mapping transformation neural network uses the mean square error as the loss function: ; Among them E is the average error of the data, n is the number of data, is the correct value of the th data in the data, is the predicted value given by the neural network.
[0010] Furthermore, the functional implementation process of the MaSA module is as follows: ; Among them Q , K and V respectively represent the query, key and value matrices, is the spatial attenuation matrix, is the two-dimensional spatial attenuation matrix, is the attenuation coefficient, is the two-dimensional coordinate of the token in the image, T is the transpose, and Softmax is the normalization process, is the element-wise multiplication.
[0011] Furthermore, the functional implementation process of the RCS-OSA module is as follows: RepVGG uses a multi-branch structure during training and is merged into a single 3x3 convolution during inference through equivalent conversion: ; Among them Identity is the identity mapping, Conv lx1 and Conv 3x3 are convolution operations; Through structural re-parameterization, the multi-branch is converted into a single 3x3 RepConv : ; Rearrange the channel order to promote information interaction between different channel groups: ; Among them W is the channel rearrangement weight matrix, and cross-channel information flow is achieved through grouped convolution; Stack multiple RCS modules and aggregate features in the last layer: ; Among them represents the output of the th OSA sub-module, and Concat represents channel concatenation.
[0012] Furthermore, the functional implementation process of the MLIA module is as follows: The functional implementation process of the local importance of the module is as follows: ; Among them, is the local importance value of the pixel x , R represents the neighborhood centered on x , and w is a learnable weight for refining the importance of measurement; The implementation mechanism of the gating mechanism of the module is as follows: ; Among them, and are the sigmoid activation and bilinear interpolation operations respectively; is the first channel map of the input feature.
[0013] Furthermore, the loss function based on DynamicFocaler-IoU is as follows: ; Among them, IoU is the intersection over union, and IoU focaler is the reconstructed Focaler-IoU value, l is the lower threshold, and u is the upper threshold; is a non-linear function introduced to differentially weight the prediction error, epoch is the total number of training times, and eps is the actual number of training times.
[0014] The present invention provides a multi-modal fusion oil and gas pipeline perimeter security system, which has program modules corresponding to the steps of the method described in any one of the above technical solutions, and executes the steps in the above multi-modal fusion oil and gas pipeline perimeter security method when running.
[0015] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program is configured to implement the steps in the multi-modal fusion oil and gas pipeline perimeter security method described in any one of the above technical solutions when called by a processor.
[0016] Compared with the prior art, the beneficial effects of the present invention are: A multi-modal data fusion method for oil and gas pipelines will be adopted. Radar with strong penetration and not affected by light and weather is used to identify suspicious targets. Combining with the high-resolution images provided by cameras, through data fusion recognition of visual features and point cloud features, the false alarm rate of the device is greatly reduced. It can still maintain high target detection accuracy in complex scenarios such as low light and occlusion, effectively identify all intrusion events and respond in a timely manner. The YOLO-based target detection model of the present invention introduces an explicit spatial prior in the visual backbone network by introducing the MaSA module, introduces the RCS-OSA module to cascade between features at different levels to enhance information flow, and introduces the MLIA module to integrate local importance learning and channel gating mechanisms, enabling the module to have the synergistic effect of second-order information interaction ability, comprehensively improving the target detection accuracy of the model, and providing a more reliable technical solution for the perimeter security of oil and gas pipelines. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a structural diagram of the perimeter security monitoring system for oil and gas pipelines with multi-modal information fusion in the embodiment of the present invention; Figure 2 It is a schematic diagram of the radar guiding the camera to rotate in the embodiment of the present invention; Figure 3 It is a schematic diagram of accurate pixel tracking of the camera in the embodiment of the present invention; Figure 4 It is a schematic diagram of the structure of the mapping transformation neural network in the embodiment of the present invention; Figure 5 It is a diagram of the spatial attenuation matrix in the MaSA module in the embodiment of the present invention; Figure 6 It is a schematic diagram of the structure of the RCS-OSA module in the embodiment of the present invention; Figure 7 It is a loss curve diagram of the model in the embodiment of the present invention after adding various loss functions; Figure 8 It is a data fusion schematic diagram of the perimeter security monitoring system for oil and gas pipelines with multi-modal information fusion in the embodiment of the present invention; Figure 9 It is a technical roadmap of the perimeter security monitoring system for oil and gas pipelines with multi-modal information fusion in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To enable those skilled in the art to better understand the solution of the present invention, the exemplary embodiments or examples of the present invention will be described below in conjunction with the accompanying drawings. Obviously, the described embodiments or examples are only a part of the embodiments or examples of the present invention, rather than all of them. All other embodiments or examples obtained by those of ordinary skill in the art based on the embodiments or examples in the present invention without creative work shall fall within the scope of protection of the present invention.
[0019] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings.
[0020] Specific Embodiment 1: The present invention provides a multi-modal fusion oil and gas pipeline perimeter security method, as Figure 8 and 9 shown, including the following steps: Step S100: Collect point cloud data of the oil and gas pipeline perimeter environment through radar, determine suspicious targets based on the radar point cloud data, calculate the angle of the target relative to the camera, and enable the camera to track the target and collect image data of the target; Step S200: Perform dimensionality reduction processing on the point cloud data, extract features from the dimensionality-reduced point cloud data and image data respectively, and then perform feature fusion to obtain multi-modal fusion data; Step S300: Build a target detection model based on YOLO. The Head network of the model introduces the MaSA attention module and the RCS-OSA module. The MaSA attention module is used to construct a two-dimensional bidirectional spatial attenuation matrix based on the Manhattan distance, introduce an explicit spatial prior in the visual backbone network, and enable the model to focus on target features; the RCS-OSA module is used to reduce the computational complexity based on the dynamic structure reparameterization ability and the cross-channel interaction mechanism of ShuffleNet, and perform multi-scale feature extraction and feature aggregation; the MLIA attention module is also introduced in the backbone network of the model. The MLIA attention module is used to calculate the local importance of pixels based on the integration of local importance learning and the channel gating mechanism, so that the module has second-order information interaction ability; The model uses a loss function based on DynamicFocaler-IoU to help the model better learn to extract features from medium-difficulty samples; Step S400: Detect intrusion targets in the multi-modal fusion data based on the YOLO-based target detection model.
[0021] Specific Embodiment 2: As Figure 2As shown, the calculation of the angle of the target relative to the camera in step S100 is specifically as follows: The point cloud data of the perimeter of the oil and gas pipeline is collected by radar 1 and radar 2, the distances of the target from two different radars are calculated, and further the angles of the target relative to the two cameras are calculated: ; Among them, a is the distance of the target from radar 1, b is the distance of the target from radar 2, c is the distance between the two radars, and β are the angles between camera 1 and camera 2 and the target respectively, and are the initial angles of camera 1 and camera 2 respectively, and are the angles that camera 1 and camera 2 should rotate respectively.
[0022] As Figure 3 shown, after positioning the target angle by radar, when the distance is less than 200 pixels, the camera tracks at a relatively small speed to make the target tracking more stable and solve the overshoot problem. The tracking time is dynamically adjusted according to the distance to be moved. Other parts of this implementation scheme are the same as those of the specific implementation scheme one.
[0023] Specific implementation scheme three: The dimensionality reduction processing of the point cloud data described in S200 is specifically carried out by using a mapping transformation neural network to reduce the dimensionality of the point cloud data. The mapping transformation neural network is based on a BP neural network and adds a convolutional layer and a residual module; As Figure 4 shown, the input layer of the model is set to 7, corresponding to seven state information of the object detected by the millimeter-wave radar, including: longitudinal distance, lateral distance, longitudinal speed, lateral speed, category, length, and width.
[0024] The convolutional layer adds a one-dimensional convolutional operation to extract the features of the radar data.
[0025] The residual module is used to combine the deep features with the shallow features to reduce the increase in the number of parameters and the performance degradation caused by the increase in the network depth.
[0026] The output layer is set to 4, representing the pixel coordinates of the upper left and lower right corners of the bounding box.
[0027] Through the combination of the convolutional operation and the residual module, the model can effectively extract features and perform information fusion. Other parts of this implementation scheme are the same as those of the specific implementation scheme two.
[0028] Specific implementation plan four: The mapping transformation neural network uses the mean square error (MSE) as the loss function to measure the difference between the network output and the label: ; where E is the average error of this batch of data, n is the number of this batch of data, is the correct value of the th data in this batch of data, is the predicted value given by the neural network. Other parts of this implementation plan are the same as those of specific implementation plan three.
[0029] Specific implementation plan five: The functional implementation process of the Manhattan Self-Attention (MaSA) module is as follows: ; where Q , K and V represent the query, key, and value matrices respectively, which are used to calculate the attention weights, is a two-dimensional spatial attenuation matrix, which calculates the attenuation factor of the attention weights based on the Manhattan distance, is the attenuation coefficient, which is used to control the attenuation speed. is the two-dimensional coordinate of the token in the image, T is the transpose, and Softmax is the normalization process, is the element-wise multiplication.
[0030] As Figure 5 shown, the MaSA module extends the temporal attenuation mechanism in RetNet to the spatial domain, that is, constructs a two-dimensional bidirectional spatial attenuation matrix based on the Manhattan distance, thereby introducing an explicit spatial prior in the visual backbone network. Other parts of this implementation plan are the same as those of specific implementation plan four.
[0031] Specific implementation plan six: The functional implementation process of the RCS-OSA module is as follows: RepVGG uses a multi-branch structure during training and is merged into a single 3x3 convolution through equivalent conversion during inference. The formula is as follows: ; where Identity is the identity mapping, Conv lx1 and Conv 3x3 are convolution operations.
[0032] Through structural reparameterization, the multi-branch is converted into a single 3x3 RepConv : ; Channel shuffle rearranges the channel order to facilitate information interaction between different channel groups: ; Among them W is the channel rearrangement weight matrix, which realizes cross-channel information flow through grouped convolution; Stack multiple RCS modules and aggregate features in the last layer: ; Among them represents the output of the -th OSA sub-module, and Concat represents channel concatenation.
[0033] This implementation scheme realizes a multi-branch topology structure in the training stage and a simplified single-branch structure in the inference stage through the RCS module, so as to improve the richness of feature information and the inference speed. The aggregation of multi-scale features is realized through the OSA module, reducing network fragmentation and computational complexity. It is realized that in each branch of OSA, multiple RCS modules are stacked, and the feature extraction ability is enhanced by repeated stacking.
[0034] Such as Figure 6 shown, the RCS-OSA module (RepVGG / RepConv ShuffleNet One-Shot Aggregation) is a lightweight feature enhancement unit designed for object detection tasks. Its technical architecture integrates the dynamic structure reparameterization ability of RepVGG / RepConv and the cross-channel interaction mechanism of ShuffleNet. Through the RCS-OSA module, cascading can be effectively carried out between features at different levels to enhance information flow. By adopting an extensible multi-branch structure during training to capture diverse features, and merging them into a single computational path through structure reparameterization during the inference stage, the optimization of computational efficiency is achieved. Other parts of this implementation scheme are the same as those of the fifth specific implementation scheme.
[0035] Specific implementation scheme seven: The MLIA (Mix Local Importance-based Attention) module integrates local importance learning and channel gating mechanism, and has second-order information interaction ability. The functional implementation process of the MLIA module is as follows: The functional implementation process of the local importance of the module is as follows: ; Among them, is the local importance value of the pixel x , R represents the neighborhood centered on x , wis a learnable weight for refining the importance of measurements; To recalibrate local importance and avoid artifacts caused by strided convolution and bilinear interpolation, the implementation mechanism of the gating mechanism of the module is as follows: ; where and are sigmoid activation and bilinear interpolation operations respectively; is the first channel map of the input feature, which is used to simplify the gating unit. Other aspects of this implementation are the same as those of Specific Implementation Six.
[0036] Specific Implementation Eight: The loss function based on DynamicFocaler-IoU constructs the IoU loss using a linear interval mapping method: ; where IoU is the intersection over union, IoU focaler is the reconstructed Focaler-IoU value, l is the lower threshold, u is the upper threshold, and by adjusting the values of l and u, the Focaler-IoU is controlled to focus on regression samples in different ranges. is a non-linear function introduced to differentially weight the prediction error, epoch is the total number of training times, and eps is the actual number of training times.
[0037] The loss function based on DynamicFocaler-IoU adjusts the loss according to the value of the intersection over union. When IoU is less than a lower threshold l, the loss is 0; when IoU is greater than an upper threshold u the loss is 1; and when IoU is between l and u the loss is not a function that linearly increases according to the IoU value, but a non-linear function that transforms according to . The method of the present invention allows the loss function to be sensitive to the IoU value within a certain range, enabling the model to focus more on samples with a medium overlap degree between the predicted bounding box and the ground truth bounding box, which helps the model better learn to extract features from medium-difficulty samples rather than just focusing on the easiest or most difficult samples. Other aspects of this implementation are the same as those of Specific Implementation Seven.
[0038] To verify the effectiveness of the Dynamic Focaler-IoU loss function of the present invention, training was carried out under the same conditions, and the effects were compared with common loss functions CIoU, EIoU, DIoU, GIoU, SIoU, WIoU, MPDIoU, ShapeIoU, FocalerIoU. The results are as Figure 7As shown, the results indicate that as the number of training iterations changes, the focus of the model also changes, enabling it to be more focused on samples where the overlap between the predicted bounding boxes and the ground truth bounding boxes is moderate.
[0039] A multi-modal fusion oil and gas pipeline perimeter security method (algorithm) proposed by the present invention is the underlying technical core of the present invention, and various products can be derived based on this algorithm.
[0040] Based on the method proposed by the present invention, a multi-modal fusion oil and gas pipeline perimeter security system is developed using a programming language. This system has program modules corresponding to the steps of the above technical solution and executes the steps in the above multi-modal fusion oil and gas pipeline perimeter security method when running. As Figure 1 shown, the system is based on the "intelligent pipeline" architecture system and follows the design principles of intelligence, panoramic view, and integration for system framework design. The system consists of three main levels. The sensing layer is used to collect environmental information around the oil and gas pipeline. The transmission layer is used to ensure that the collected data can be transmitted to the data processing center in real time and accurately. The application layer is used to store and preliminarily process the transmitted data, realize intelligent analysis and comprehensive application of the data, and provide decision support. The system also includes an alarm processing module for emergency situations and a storage module for collecting relevant information data to ensure rapid response in case of emergency and save key data for subsequent analysis.
[0041] The computer program of the developed system (software) is stored on a computer-readable storage medium. The computer program is configured to implement the steps of the above multi-modal fusion oil and gas pipeline perimeter security method when called by a processor. That is, the present invention is materialized on a carrier to become a computer program product.
[0042] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0043] The computing programs (also referred to as programs, software, software applications, or code) in the present invention include machine instructions for a programmable processor, and these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., magnetic disk, optical disk, memory, programmable logic device PLD) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0044] The beneficial effects of the present invention will be described below in conjunction with specific embodiments. Embodiment
[0045] The network model of the present invention is compared with the existing YOLOv5 - YOLOv12. A dataset is constructed, including five key security elements: personnel, vehicles, drones, safety helmets, and safety suits, a comprehensive dataset containing complex environmental features and security semantic information, with a total of 5000 photos; the ratio of the training set, test set, and validation set is 7:2:1. The experimental configuration is as follows: the operating system is Windows 11; the processor is 12th Gen Intel(R) Core(TM) i9 - 12900H 2.50GHz; the GPU is NVIDIA GeForce GTX 3060; the CUDA version is 12.1; the deep learning framework is PyTorch 2.5.1; the scripting language is Python 3.10.16. During the entire training phase, a batch size of 16 images is used, and the input dimension of each image is 640×640. The training process spans 100 epochs. Experiments are conducted while keeping the configuration environment and initial training settings unchanged. The results are shown in Table 1. It can be seen that the object detection accuracy and model performance of the network model of the present invention have been significantly improved.
[0046] Table 1 ; Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present disclosure, and these changes and modifications will all fall within the protection scope of the present invention.
Claims
1. A perimeter security method for oil and gas pipelines with multimodal fusion, characterized in that, It includes the following steps: Step S100: Collect the point cloud data of the perimeter environment of the oil and gas pipeline through the radar, determine the suspicious targets based on the radar point cloud data, and enable the camera to track the targets and collect the image data of the targets; Step S200: Perform dimensionality reduction processing on the point cloud data, extract features from the dimensionality-reduced point cloud data and the image data respectively, and then perform feature fusion to obtain multi-modal fusion data; Step S300: Build a target detection model based on YOLO. The Head network of the model introduces the MaSA attention module and the RCS-OSA module. The MaSA attention module is used to introduce explicit spatial priors in the visual backbone network by constructing a two-dimensional bidirectional spatial attenuation matrix, enabling the model to focus on target features; The RCS-OSA module is used for multi-scale feature extraction and feature aggregation based on the dynamic structure reparameterization ability and the cross-channel interaction mechanism of ShuffleNet; The MLIA attention module is also introduced in the backbone network of the model. The MLIA attention module is used to calculate the local importance of pixels based on the integration of local importance learning and the channel gating mechanism, enabling the module to have second-order information interaction capabilities; The YOLO-based target detection model uses a loss function based on DynamicFocaler-IoU to help the model better learn to extract features from medium-difficulty samples; Step S400: Detect the intrusion targets in the multi-modal fusion data using the YOLO-based target detection model.
2. The multi-modal fusion oil and gas pipeline perimeter security method according to claim 1, characterized in that In step S100, enabling the camera to track the targets and collect the image data of the targets is specifically as follows: Use dual radars to collect the point cloud data of the perimeter environment of the oil and gas pipeline, and a camera is connected to each radar; First, calculate the distances of the target from radar 1 and radar 2, and further calculate the angle of the target relative to the camera: ; Among them, a is the distance of the target from Radar 1, b is the distance of the target from Radar 2, c is the distance between the two radars, α and β are the angles between the target and Camera 1 and Camera 2 respectively, and are the initial angles of Camera 1 and Camera 2 respectively, and are the rotation angles of Camera 1 and Camera 2 respectively.
3. The perimeter security method for oil and gas pipelines with multi-modal fusion according to claim 2, characterized in that In S200, performing dimensionality reduction processing on the point cloud data specifically uses a mapping transformation neural network to perform dimensionality reduction processing on the point cloud data. The mapping transformation neural network is based on the BP neural network and adds a convolutional layer and a residual module. The convolutional layer is used to extract the features of the radar data, and the residual module is used to combine the deep features and the shallow features to reduce the increase in the number of parameters and the performance degradation caused by the increase in the network depth.
4. The multi-modal fusion-based perimeter security method for oil and gas pipelines according to claim 3, characterized in that, The mapping transformation neural network uses the mean square error as the loss function: ; where E is the average error of the data, n is the number of the data, is the correct value of the th data in the data, is the predicted value given by the neural network.
5. The multi-modal fusion-based perimeter security method for oil and gas pipelines according to claim 4, characterized in that, The functional implementation process of the MaSA attention module is: ; Among them Q , K and V represent query, key, and value matrices respectively, is the spatial attenuation matrix, is the two-dimensional spatial attenuation matrix, is the attenuation coefficient, is the two-dimensional coordinate of the token in the image, T is the transpose, and Softmax is the normalization process, is the element-wise multiplication.
6. The multi-modal fusion-based perimeter security method for oil and gas pipelines according to claim 5, characterized in that, The functional implementation process of the RCS-OSA module is: RepVGG uses a multi-branch structure during training and is merged into a single 3x3 convolution through equivalent conversion during inference: ; where Identity is the identity mapping, Conv lx1 and Conv 3x3 are convolutional operations; Convert multi-branch to single-path 3x3 through structural reparameterization RepConv : ; Rearrange the channel order to promote information interaction between different channel groups: ; Among them W is the channel rearrangement weight matrix, which realizes cross-channel information flow through grouped convolution; Stack multiple RCS modules and aggregate features in the last layer: ; Among them represents the output of the th OSA sub-module, and Concat represents channel concatenation.
7. The multi-modal fusion-based perimeter security method for oil and gas pipelines according to claim 6, wherein, The functional implementation process of the MLIA attention module is: The functional implementation process of the local importance of the module is: ; Among them, is the local importance value of the pixel x . R Represents the neighborhood centered on x , w is a learnable weight used to refine the importance of the measurement; The implementation mechanism of the gating mechanism of the module is: ; Among them, and are sigmoid activation and bilinear interpolation operations respectively; is the first channel map of the input feature.
8. The multi-modal fusion-based perimeter security method for oil and gas pipelines according to claim 7, wherein, The loss function based on DynamicFocaler-IoU is: ; Among them, IoU is the intersection over union, and IoU focaler is the reconstructed Focaler-IoU value, l is the lower threshold, and u is the upper threshold; is the introduced non-linear function to differentially weight the prediction error, epoch is the total number of training times, and eps is the actual number of training times.
9. A perimeter security system for oil and gas pipelines with multi-modal fusion, characterized in that, The system has program modules corresponding to the steps of the method described in any one of claims 1 to 8 above, and when running, executes the steps in the above multi-modal fusion oil and gas pipeline perimeter security method.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps in the multi-modal fusion oil and gas pipeline perimeter security method described in any one of claims 1 to 8 when called by a processor.
Citation Information
Patent Citations
Intelligent monitoring and early warning system and method for security and protection of oil and gas pipeline station
CN113206978A
Three-dimensional target detection method based on multi-modal fusion and deep attention mechanism
CN116612468A
Multi-modal data fusion method, system and equipment for oil and gas pipeline and medium
CN118194227A
Method for detecting infrared ship target based on improved yolov7
US20250078541A1
Cited By
Heterogeneous data fusion method and device for automatic driving scene perception and medium
CN121982462A
Heterogeneous data fusion method, device and medium for automatic driving scene perception
CN121982462B