Underground power optical cable multi-scene vibration area positioning method based on PLGS-YOLO model

By improving the PLGS-YOLO model of YOLOv11n, combined with SPPFDMSCA, PLMSAM, GSConv and VoVGSCSP modules, the accuracy and real-time problems of multi-scene vibration area positioning of underground power cables are solved, and efficient vibration area detection and scene recognition are achieved.

CN120388167AActive Publication Date: 2025-07-29NORTHEAST DIANLI UNIVERSITY

Patent Information

Application Number
CN202510521787.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-29
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing vibration area positioning method of underground power cables is difficult to effectively balance performance, real-time and model lightweight in multiple scenarios, and is disturbed by background noise or non-target area signals, resulting in the positioning accuracy not meeting the requirements.

Method used

Using the PLGS-YOLO model, by improving YOLOv11n, SPPFDMSCA and PLMSAM modules are introduced to enhance feature extraction, combined with GSConv and VoVGSCSP optimized feature fusion, adapting to vibration signal processing in buried and downhole scenarios.

Benefits of technology

It significantly improves the positioning accuracy of multiple scene vibration areas of underground power cables, while reducing the computational complexity and memory usage, and is suitable for environments with limited computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388167A_ABST
    Figure CN120388167A_ABST
Patent Text Reader

Abstract

The invention discloses an underground power optical cable multi-scene vibration area positioning method based on a PLGS-YOLO model, relates to the technical field of deep learning models and power optical cable monitoring, and solves the problem that a model adopted for vibration area positioning in an existing underground power optical cable cannot meet application requirements due to interference of background noise or non-target area signals. The method is used for judging the laying scene of the vibration event and positioning the corresponding signal area. YOLOv11n is optimized, SPPFDMSCA and PLMSAM are utilized, and feature extraction efficiency is enhanced through a multi-scale attention mechanism; the GSConv combines the advantages of the traditional convolution and the depth separable convolution, so that the calculation complexity is reduced, and meanwhile, the multi-channel characteristics are reserved; the VoVGSCSP optimizes feature fusion, improves the detection precision and the model efficiency, and remarkably reduces the calculation complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning models and power optical cable monitoring, and particularly relates to a method for locating multi-scene vibration regions of underground power optical cables based on a YOLO model (YOLO with Progressive Lightweight Grouped-Shuffle Feature Fusion and Spatial Attention, abbreviated as PLGS-YOLO) that integrates progressive lightweight feature extraction and grouped shuffle spatial attention. Background Technique

[0002] As an important communication device in the power system, underground power optical cables have become a key infrastructure for ensuring urban power supply and the efficient operation of the power grid. Underground power optical cables have multiple laying methods such as buried, underground, and cable trench. Among them, buried and underground laying often face the risks of natural and man-made damage. For example, during buried laying, construction activities may cause the optical cable to be compressed, displaced, or broken; underground laying may encounter illegal excavation or foreign object impact, affecting the transmission quality of the optical fiber. How to effectively manage and maintain underground power optical cables to ensure their smooth communication has become an urgent problem to be solved.

[0003] The Φ-OTDR fiber sensing technology has the advantages of high sensitivity, long-distance detection, high reliability, and low cost. This technology uses the optical fiber as a medium and carrier to sense various vibrations occurring along the optical fiber and collect vibration information in real time. It is possible to analyze man-made damage events by collecting its vibration signals. This technology is widely used in safety detection fields such as pipeline safety warning, perimeter security, railway transportation, and submarine cables. A part of unused optical fiber is reserved in the power optical cable for replacement when the main optical fiber fails. Therefore, when an abnormal vibration event has not yet threatened the underground power optical cable, the unused internal optical fiber can be used as a sensing medium for vibration signals. It is very suitable for application in underground power optical cable monitoring tasks. Therefore, the Φ-OTDR fiber sensing technology is used to collect vibration signals in the surrounding environment of underground power optical cables, and relevant research methods are used to locate the regions of man-made events that may cause damage.

[0004] Currently, mainly the single-stage method of the You Only Look Once (YOLO) model is used to locate the regions corresponding to vibration events in the spatio-temporal map of the signals collected by the Φ-OTDR system. The input of the YOLO model is the spatio-temporal map of vibration signals, which converts time series data into images. However, this method only measures the quality of the method by accuracy and real-time performance. The laying scenarios of underground power optical cables are complex, there are many types of vibration events, and they are affected by different background noises or signals in non-target regions. The deployment difficulty of the proposed method also needs to be fully considered.

[0005] At present, the vibration regions of Φ-OTDR are mainly divided into two methods based on signal processing technology and based on target detection models. In the actual environment, the interference noise is large, which will submerge the useful signals and is not conducive to the vibration regions. Based on the above problems, a series of positioning methods using signal processing technology have been proposed in recent years. Wu et al. used a simulated annealing algorithm with an adaptive annealing threshold for wavelet denoising, and then calculated the spatial gradient of the grayscale image using a two-dimensional edge detection method for axial positioning. Huang et al. proposed a regional axial positioning method based on overlapping phase cross-correlation. According to the linear relationship between the vibration signals of axial points on the link over time, the autocorrelation coefficient of the vibration signal matrix at adjacent times was calculated, and the maximum value in the link was taken as the position of the vibration.

[0006] The positioning methods based on signal processing technology have good positioning effects on high-frequency and large-amplitude vibration regions, but have poor positioning effects on low-frequency and small-amplitude events and vibration regions that are buried deep or far away. For this reason, a series of positioning methods based on target detection models have been proposed. The target detection models have high precision and real-time performance. Xu et al. proposed a YOLOv3 multi-class vibration detection model for positioning and real-time detection of the vibration regions of intrusion events. Wang et al. proposed a Φ-OTDR perimeter security event vibration region positioning method based on CBAM-YOLOv8. However, both types of positioning methods collect vibration signals and perform positioning in specific laying scenarios, while the laying scenario of a complete underground power cable is complex, and the positioning problems in different laying scenarios of the same cable need to be considered.

[0007] To solve the problem of vibration region positioning of underground power cables in multiple scenarios, the present invention proposes a PLGS-YOLO model based on YOLOv11n, which locates the vibration region while judging the laying scenario of the vibration region. This method comprehensively considers the buried and underground laying scenarios where human damage to underground power cables is relatively frequent, as well as the scale, accuracy, and real-time performance of the model adopted by the method. It can be directly applied to real underground power cables to prevent potential damage behaviors, better protect the safety of underground cables, and has important practical significance and application prospects. Summary of the Invention

[0008] The present invention aims to solve the problems that the vibration region positioning in existing underground power cables mainly focuses on a single laying scenario, and it is difficult to effectively balance the performance, real-time performance, and model lightweight of such methods, which limits their application in multiple laying scenarios, and the models adopted are interfered by background noise or signals in non-target regions, resulting in the models not meeting the application requirements; and provides a multi-scenario vibration region positioning method for underground power cables based on the PLGS-YOLO model.

[0009] Underground power optical cable multi-scenario vibration area positioning method based on the PLGS-YOLO model, which is implemented by the following steps:

[0010] Step 1: Preprocess the acquired vibration signal and convert the preprocessed signal into a spatio-temporal map;

[0011] Step 2: Vibration area positioning;

[0012] By constructing a PLGS-YOLO model, training the PLGS-YOLO model, and using the trained PLGS-YOLO model to locate the vibration area in the spatio-temporal map obtained in Step 1, a located spatio-temporal map and the corresponding laying scenario type are generated;

[0013] The PLGS-YOLO model improves the YOLOv11n network. In the backbone network, the SPPF module and the C2PSA module are respectively replaced by the SPPFDMSCA module and the PLMSAM module; the convolutional layer in the neck network is replaced by the GSConv module, and the C3k2 module is replaced by the VoVGSCSP module;

[0014] The DMSCA module is introduced into the SPPFDMSCA module to enhance the global feature extraction ability of the backbone network;

[0015] The LMSCA module is used in the PLMSAM module to obtain a deep feature map;

[0016] Input the spatio-temporal map into the YOLOv11n backbone network, perform feature enhancement through the dynamic multi-scale convolutional attention mechanism in the SPPFDMSCA module, and input the obtained global feature map into the PLMSAM module;

[0017] Perform multi-scale convolutional operations on the global feature map through the LMSCA module in the PLMSAM module and then input it into the neck network;

[0018] The GSConv module and the VoVGSCSP module in the neck network perform feature fusion on the feature maps after multi-scale convolutional operations, and finally obtain the located spatio-temporal map and the corresponding laying scenario type through the detection head.

[0019] Advantages of the present invention:

[0020] Deploy the Φ-OTDR system in a certain communication station, select the underground power optical cable line with two scenarios of buried and underground where human threats are relatively frequent, use the optical cable as the vibration sensing medium, and design a vibration event experiment to simulate real threats. Preprocess the collected original signal to obtain a self-built positioning model dataset.

[0021] The method of the present invention proposes the PLGS-YOLO model to achieve the positioning of vibration regions in two laying scenarios: underground burial and underground well. Based on the YOLOv11n model, by introducing a Spatial Pyramid Pooling Fusion with Dynamic Multi-Scale Convolution Attention (abbreviated as SPPFDMSCA) structure and a Progressive Lightweight Module based on Partial Self-Attention Mechanism (abbreviated as PLMSAM), the efficiency of vibration feature extraction is enhanced. In addition, by combining Grouped Shuffle Convolution (abbreviated as GSConv) and a One-Shot Aggregation-Based Grouped Shuffle Cross Stage Partial Block (abbreviated as VoVGSCSP), the feature fusion process is optimized. By comparing with YOLO series lightweight models (including models already applied to vibration region positioning), it is proved that PLGS-YOLO significantly reduces the computational complexity and memory occupancy while improving the positioning accuracy. Description of the Drawings

[0022] Figure 1 It is a flowchart of the method for multi-scenario vibration region positioning of underground power optical cables based on the PLGS-YOLO model described in the present invention;

[0023] Figure 2 It is a principle block diagram of the PLGS-YOLO model in the method for multi-scenario vibration region positioning of underground power optical cables based on the PLGS-YOLO model described in the present invention;

[0024] Figure 3 It is a schematic diagram of the SPPFDMSCA module;

[0025] Figure 4 It is a schematic diagram of the DMSCA module;

[0026] Figure 5 It is a schematic diagram of the PLMSAM module;

[0027] Figure 6 It is a schematic diagram of the LMSCA module;

[0028] Figure 7 It is a schematic diagram of the GSConv module;

[0029] Figure 8 It is the schematic diagram of the VoVGSCSP module;

[0030] Figure 9 It is the deployment effect diagram of the experimental scenario and the Ф-OTDR system;

[0031] Figure 10 In (a), (b), (c), and (d) of it, they are respectively the φ-OTDR signal effect diagrams of the vibration events of foreign object falling, excavation, knocking, and electric drill and hammer;

[0032] Figure 11 In (a), (b), (c), and (d) of it, they are respectively the denoised φ-OTDR signal effect diagrams of the vibration events of foreign object falling, excavation, knocking, and electric drill and hammer;

[0033] Figure 12 It is the spatio-temporal diagram for locating the underground vibration area. Among them, (a), (b), (c), (d), (e), (f), (g), and (h) are respectively the spatio-temporal diagrams of background noise, motorcycle, trolley, excavation, pickaxe, shot put, walking, and electric drill and hammer;

[0034] Figure 13 It is the spatio-temporal diagram for locating the underground vibration area. Among them, (a), (b), (c), (d), (e), (f), and (g) are respectively the spatio-temporal diagrams of background noise, foreign object falling, electric drill, dragging, climbing, knocking, and walking;

[0035] Figure 14 It is the performance curve effect diagram of models such as PLGS-YOLO on the validation set. Among them, (a), (b), (c), (d), (e), and (f) are respectively the curve effect diagrams of Boxloss, Clsloss, Precision, Recall, mAP@50, and mAP@50-95;

[0036] Figure 15 It is the visualization effect diagram of the detection results using the underground power optical cable multi-scenario vibration area location method based on the PLGS-YOLO model described in the present invention. Among them, (a), (b), (c), (d), (e), (f), (g), and (h) are respectively the visualization effect diagrams of underground motorcycle, underground knocking, underground excavation, underground trolley, underground drill, underground walking, underground knocking, and underground foreign object falling. Specific implementation manners

[0037] Specific implementation manner 1. In combination with Figures 1 to 8Describe this embodiment, a multi-scenario vibration area localization method for underground power optical cables based on the PLGS-YOLO model. This method aims at multi-scenario vibration events of underground power optical cables. By improving YOLOv11n, a multi-scenario vibration area localization method for underground power optical cables based on the PLGS-YOLO model is constructed. The overall process of this method includes two stages, as Figure 1 shown, where the green line represents the vibration signal preprocessing stage, and the red line represents the vibration area localization stage. The specific steps of this method are as follows:

[0038] Step 1: Vibration signal preprocessing. Denoise the Φ-OTDR signal through high-pass and low-pass filtering to obtain the preprocessed spatio-temporal signal; convert the spatio-temporal signal into a spatio-temporal graph;

[0039] In the processing of vibration signals of underground power optical cables in multi-laying scenarios, the signals collected by the Φ-OTDR system often contain high-frequency internal machine noise and low-frequency environmental noise interference. These noises will cover up the effective features of the vibration signals and affect the positioning performance of subsequent methods. Therefore, in this embodiment, a combination of low-pass filtering and high-pass filtering is used to preprocess the vibration signals, removing irrelevant noises in the frequency domain and improving the signal quality. At the same time, high-pass filtering and low-pass filtering can meet the real-time requirements of the system for signal preprocessing due to their short processing time.

[0040] Step 1-1: High-pass filtering;

[0041] High-pass filtering (High-Pass Filter, HPF) allows high-frequency components in the signal to pass through while suppressing low-frequency components. In the Φ-OTDR signal, low-frequency noise usually comes from the environmental noise of the system operation. These low-frequency interferences may cover up the detailed features in the vibration event. By high-pass filtering, these low-frequency noises can be removed, thus highlighting the key change information in the vibration signal. The transfer function of the high-pass filter is defined as:

[0042]

[0043] In the formula, H HPF (f) represents the transfer function of the high-pass filter, f is the frequency of the input signal, and f c1 is the cut-off frequency of the high-pass filter.

[0044] The output signal of the high-pass filter can be expressed as:

[0045]

[0046] Among them, x[k] is the original input signal, h HPF[n - k] is the impulse response of the high-pass filter at time step n, N represents the length of the filter, n represents the current time step, and the filter calculates the corresponding output signal y HPF [n].

[0047] Step 1-2: Low-pass filtering;

[0048] Low-pass filtering allows the low-frequency components in the signal to pass through while suppressing high-frequency noise. In Φ-OTDR signals, high-frequency noise is usually caused by electromagnetic interference inside the machine or random fluctuations of tiny vibration signals. By low-pass filtering, these high-frequency components can be removed, and the main trend information in the optical cable vibration signal can be retained, thus more effectively reflecting the spatial distribution of vibrations. The transfer function of the low-pass filter is defined as:

[0049]

[0050] where H LPF (f) represents the transfer function of the low-pass filter, and f c2 is the cut-off frequency of the low-pass filter. The output signal of the low-pass fil

[0051] ter can be expressed as:

[0052]

[0053] where h LPF [n - k] is the impulse response of the low-pass filter at time step n, and the filter calculates the corresponding output signal y LPF [n].

[0054] Step 1-3: Combining high-pass and low-pass filtering;

[0055] In the processing of underground power optical cable vibration signals, low-pass filtering is used to retain the main trend of the vibration signal, such as vibration events with large amplitudes like foreign objects falling and excavation; high-pass filtering is used to extract the detailed changes in the vibration signal, such as event characteristics caused by high-frequency vibrations like knocking or electric drills and hammers. By combining low-pass and high-pass filtering, the signal-to-noise ratio of the signal can be significantly improved, optimizing the performance of the proposed method in the vibration event localization task.

[0056] First, the method of high-pass filtering is used to remove low-frequency noise, which is defined as:

[0057]

[0058] where x[m] is the original Φ-OTDR signal from the Φ-OTDR system before filtering, and h HPF [l - m] is the impulse response of the high-pass filter in the Φ-OTDR system at time step l, and y HPF[l] is the output signal after high-pass filtering of the original Φ-OTDR signal at time step l.

[0059] Secondly, a low-pass filtering process is adopted to smooth the high-frequency components. The combined filtering equation that finally includes both high-pass and low-pass filtering is as follows:

[0060]

[0061] where h LPF [n - l] is the impulse response of the low-pass filter in the Φ-OTDR system at time step n, and y filtered [n] is the final output signal after high-pass and low-pass filtering.

[0062] In this embodiment, the combined filtering process not only retains the main trend information of the low frequency but also extracts the important detail features of the high frequency, providing high-quality input data for subsequent positioning. The filter parameters are determined according to the characteristics of the dataset to optimize the filtering process for the given vibration signal.

[0063] Step 2: Vibration area positioning;

[0064] By constructing a PLGS-YOLO model and training the PLGS-YOLO model, the trained PLGS-YOLO model is used to locate the vibration area in the spatio-temporal map obtained in step one, generating the located spatio-temporal map and the corresponding laying scenario type;

[0065] In this embodiment, the method for locating the vibration area of the underground power optical cable in multiple laying scenarios needs to locate the vibration area and at the same time detect the laying scenario corresponding to the vibration area, with high requirements for real-time performance and accuracy. The YOLO series of models have shown good real-time performance and accuracy in object detection and can effectively solve the problems of vibration area positioning and scenario detection. Among them, YOLOv11n has significantly improved real-time performance and accuracy compared with the previous YOLO versions, but the computational complexity and the number of parameters are still relatively high, making it not suitable for deployment in an environment with limited computing power.

[0066] Therefore, this embodiment proposes a vibration region localization method based on the PLGS-YOLO model. Aiming at the differences in signal intensity and distribution in the underground power cable buried and underground scenarios, the YOLOv11n backbone network and neck network are optimized to meet the requirements of real-time performance, accuracy and lightweight, and better locate the regional vibration signals and detect the laying scenarios. By introducing the SPPFDMSCA module to replace the SPPF module, the global feature extraction ability of the backbone network is enhanced. The PLMSAM module is introduced to replace the C2PSA module to enhance the deep feature extraction ability of the backbone network. The GSConv module and VoVGSCSP module are used to balance the model calculation complexity and detection accuracy, and the feature fusion quality of the neck network is optimized. The overall structure of the model is as Figure 2 shown.

[0067] Step 2-1: Construct the PLGS-YOLO model. The PLGS-YOLO model improves the YOLOv11n model. In the backbone network, the SPPF module and C2PSA module are respectively replaced by the SPPFDMSCA module and PLMSAM module; in the neck network, the convolutional layer is replaced by the GSConv module, and the C3k2 module is replaced by the VoVGSCSP module;

[0068] In the SPPFDMSCA module, a Dynamic Multi-Scale Convolution Attention (DMSCA) module is introduced to enhance the global feature extraction ability of the backbone network;

[0069] In the PLMSAM module, the LMSCA module is adopted to obtain the deep feature map;

[0070] In this embodiment, three different scales of features are extracted through the backbone network. The backbone network consists of convolutional layer 1, convolutional layer 2, C3k2 module 1, convolutional layer 3, C3k2 module 2, convolutional layer 4, C3k2 module 3, convolutional layer 5, C3k2 module 4, SPPFDMSCA module and PLMSAM module;

[0071] The spatio-temporal map described in step 1 is input into the backbone network, and the processed feature map is output to the neck network and convolutional layer 4 through convolutional layer 1, convolutional layer 2, C3k2 module 1, convolutional layer 3 and C3k2 module 2 in sequence;

[0072] The feature map output by the convolutional layer 4 is output to the neck network and convolutional layer 5 through the C3k2 module 3;

[0073] The feature map output by the convolutional layer 5 is output to the upsampling module 1 and the Concat module 4 of the neck network through the C3k2 module 4, the SPPFDMSCA module, and the PLMSAM module;

[0074] In YOLOv11n, the original SPPF module uses multi-scale max pooling operations to fuse spatial information with different receptive fields. However, this structure has limitations when processing underground optical cable vibration signals. Vibration signals have obvious scale differences in space, with different signal coverage ranges and feature complexities for different vibration events, and different propagation modes in laying scenarios such as underground burial and underground wells. Traditional pooling operations may lose important spatial details while compressing information, making it difficult to fully express the characteristics of complex vibration regions. To improve the model's ability to model multi-scale vibration signals, in this embodiment, the SPPF module in the backbone network is replaced with the SPPFDMSCA module, and the structure is as Figure 3 shown. On the basis of retaining the original multi-scale structure, this module introduces the DMSCA module, enhances the feature expression ability through parallel convolution and attention fusion, further optimizes the spatial context modeling effect, enhances the global feature extraction ability of the backbone network, thereby more accurately extracting key vibration region features and improving the vibration region localization performance.

[0075] As Figure 3 shown, in this embodiment, the SPPFDMSCA module includes a convolutional layer 1, a max pooling layer 1, a max pooling layer 2, a max pooling layer 3, a Concat module, a convolutional layer 2, and a dynamic multi-scale convolutional attention mechanism module (DMSCA module);

[0076] The SPPFDMSCA module receives the high-semantic feature map after the C3k2 module 4 as input. First, through the multi-layer residual connections of the convolutional layer 1, the max pooling layer 1, the max pooling layer 2, the max pooling layer 3, and the Concat module, and then through the convolutional layer 2 and the DMSCA module for feature enhancement, and finally inputs the obtained global feature map into the PLMSAM module.

[0077] As Figure 4 shown, in the DMSCA module, the input is a global feature map with a size of DMSCA Input ∈R N×C×H×W where N represents the batch size, C represents the number of input channels, and H and W represent the height and width of the feature map respectively. Sigmoid represents the activation function, Softmax represents the normalization function, Concat represents the channel concatenation operation, · represents matrix multiplication, Conv represents the convolution operation, and reshape represents the adjustment of the channel dimension of the feature map.

[0078] First, the module introduces multi-scale convolution operations, and uses convolution kernels of 3×3, 5×5, and 7×7 respectively to perform parallel convolution extraction on the input features Figure X and splice and fuse the features extracted at the three scales in the channel dimension to obtain multi-scale features Figure X ms . This operation can capture vibration features under different receptive fields and enhance the model's perception ability of multi-type vibration regions. The specific calculation process is as follows:

[0079] X ms = Concat(Conv 3×3 (DMSCA Input ), Conv 5×5 (DMSCA Input ), Conv 7×7 (DMSCA Input ))

[0080] Subsequently, the multi-scale features Figure X ms are used to generate an attention map A through a 1×1 convolution and activated by a Sigmoid function, so that the attention map focuses on the key response regions. At the same time, X ms generates key-value pairs K and V respectively after a 1×1 convolution for constructing a context modeling module. K is normalized by a Softmax function and multiplied by V to calculate the context-enhanced feature Y. The specific calculation process is as follows:

[0081] A = Sigmoid(Conv 1×1 (X ms ))

[0082] K = Softmax(reshape(Conv 1×1 (X ms ))

[0083] V = reshape(Conv 1×1 (X ms ))

[0084] Y = Conv 1×1 (V·K)

[0085] Finally, the attention map A and the context-enhanced feature Y are added to the input feature map DMSCA Input to obtain the output global feature map DMSCA output of the module, thus realizing the joint modeling and enhancement of multi-scale and context information. The specific calculation process is as follows:

[0086] DMSCA output = DMSCA Input + A·Y

[0087] The DMSCA module significantly improves the model's ability to model and perceive complex vibration patterns by integrating multi-scale spatial receptive fields and context attention mechanisms, which helps to localize the refined vibration regions.

[0088] The PLMSAM module receives the global feature map DMSCA output and further compresses redundancy and enhances the perception of key regions through an improved lightweight multi-scale attention mechanism module (LMSAM module) to obtain a deep feature map.

[0089] The C2PSA module in the backbone network of the original YOLOv11n extracts deep features of the vibration spatio-temporal map through multiple PSAs and cross connections and is used for feature fusion in the neck network. However, in the processing of vibration signals of underground power optical cables, for example, foreign objects falling and lead balls produce similar spatio-temporal map features under different laying scenarios. This similarity increases the difficulty of the model in detecting laying scenarios, especially when the vibration events have similar frequencies and time-domain variations. To solve the above problems, in this embodiment, a PLMSAM module is designed to replace the C2PSA module for deep feature extraction, thereby enhancing the deep feature extraction ability of the backbone network. The PLMSAM module follows the framework structure of PSA but replaces the Attention module with the LMSAM module. The module structure is as Figure 5 shown, enabling it to more effectively capture key information in the vibration spatio-temporal map while reducing model parameters and computational complexity, significantly improving the model's adaptability to complex vibration signals of underground power optical cables, and providing high-quality input for feature fusion in the subsequent neck network.

[0090] As Figure 5 shown, in this embodiment, the PLMSAM module includes a batch normalization layer 1, an LMSAM module, a batch normalization layer 2, and a feed-forward neural network module (FFN module). The LMSAM module includes a spatial attention module (SpatialAttentionModule, abbreviated as SAM module) and a lightweight multi-scale channel attention module (LMSCA module);

[0091] The PLMSAM module receives the global feature map after the SPPFDMSCA module as input. First, through the residual connection structure composed of the batch normalization layer 1 and the LMSAM module, a multi-branch fusion feature map is obtained. Then, through the residual connection structure composed of the batch normalization layer 2 and the FFN module, a deep feature map is obtained. Finally, the deep feature map obtained by the PLMSAM module is input into the upsampling module 1 and the Concat module 4 of the neck network.

[0092] The pooling operation performed by the channel attention module calculation part on the input feature map converts it into pixels, resulting in a large amount of spatial information loss. Although the convolutional block attention module (CBAM module) alleviates this problem by combining the channel attention module (Channel Attention Module, abbreviated as CAM module) and the spatial attention module (Spatial Attention Module, abbreviated as SAM module), CBAM mainly optimizes features by weighting channels and spaces on the feature map, making it difficult to effectively capture the relationships between targets of different scales. Therefore, in this embodiment, an improvement is made to address this problem, and the LMSAM module is proposed. Among them, the channel attention module of the CBAM module is replaced by the LMSCA module. This mechanism can better capture the key features in the signal through multi-scale convolution, enhance the model's perception ability of complex vibration signals, and thus significantly improve the accurate positioning of the same vibration event in different laying scenarios. The LMSCA structure is as shown in Figure 6 shown.

[0093] The LMSCA structure includes three parts: a 5×5 convolutional layer, a multi-branch convolutional layer, a 1×1 convolutional layer, and an average pooling layer. Among them, the multi-branch convolution includes a 1×7 convolutional layer, a 1×11 convolutional layer, and a 1×11 convolutional layer. The 5×5 convolutional layer obtains local features, the multi-branch convolutional layer captures the multi-scale target relationships of different channels, and the 1×1 convolutional layer weights the original feature map to enhance the expression ability of the feature map. The calculation process is as follows:

[0094] M = Conv 5×5 (LMSCA Input )

[0095] M1 = Conv 1×7 (M)

[0096] M2 = Conv 1×11 (M)

[0097] M3 = Conv 1×21 (M)

[0098]

[0099] Among them, LMSCA Input represents the input feature map, Conv 1×i represents the two-dimensional convolution with different convolutional kernel sizes, M1, M2, and M3 represent three branches, and each branch performs convolution processing on the input using convolutional kernels of different sizes. LMSCA Output represents the deep feature map obtained after the LMSCA Input is processed by the average pooling layer and the weights of the 1×1 convolution.

[0100] Step 2-2: Perform feature fusion in the neck network. This neck network replaces all convolutional layers in the original YOLOv11n neck network with GSConv modules and all C3k2 modules with VoVGSCSP modules. The output of each module serves as the input for the subsequent module in turn, forming a feature transfer path from the backbone to the detection head to support the model in achieving multi-scale vibration region detection.

[0101] The convolutional layers and C3k2 modules in the YOLOv11n neck network process all channel information through global aggregation. Although they can extract rich feature information, due to their high computational complexity and the number of parameters, this increases the computational burden of the model in complex environments. In the two laying scenarios of underground burial and underground well, there are obvious differences in the characteristics of the vibration signals of underground power optical cables. The signals in the underground well usually have strong local variability, while the buried signals are more affected by environmental factors (such as soil medium, optical cable casing, etc.) and are relatively smooth. Therefore, the model not only needs to extract rich multi-channel features but also effectively process the spatio-temporal differences of these signals. To solve the above problems, in this embodiment, all convolutional layers in the YOLOv11n neck network are replaced with GSConv module 1 and GSConv module 2.

[0102] Compared with traditional convolutions, depthwise separable convolutions are outstanding in reducing computational complexity. However, their disadvantage is that they partially ignore the information relationship between channels, resulting in the loss of feature information. To combine the advantages of depthwise separable convolutions and traditional convolutions, GSConv performs channel fusion on the global feature information obtained by traditional convolutions and the channel feature information obtained by depthwise separable convolutions, reducing the amount of computation while effectively retaining multi-channel information.

[0103] In this embodiment, the neck network includes GSConv modules (GSConv module 1 and GSConv module 2), VoVGSCSP modules (VoVGSCSP module 1, VoVGSCSP module 2, VoVGSCSP module 3, and VoVGSCSP module 4), upsampling modules (upsampling module 1 and upsampling module 2), and Concat modules (Concat module 1, Concat module 2, Concat module 3, and Concat module 4);

[0104] The feature map output by the C3k2 module 2 in the backbone network reaches the Concat module 2 in the neck network. The feature map after passing through the Concat module 2 is input into the VoVGSCSP module 2. The output of the VoVGSCSP module 2 is a compressed shallow-layer high-resolution feature map, which is input into the Concat module 3 through the GSConv module 1 and is simultaneously input into the detection head 1 of the head network for shallow-layer vibration region localization and laying scenario prediction;

[0105] The feature map output by the C3k2 module 3 in the backbone network is sent to the Concat module 1 in the neck network. After passing through the Concat module 1, it is input into the VoVGSCSP module 1. The feature map output by the VoVGSCSP module 1 is input into the upsampling module 2 and the Concat module 3. After being upsampled by the upsampling module 2, it is input into the Concat module 2. After passing through the Concat module 3, it is input into the VoVGSCSP module 3, and the output of the VoVGSCSP module 3 is the compressed middle-layer high-resolution feature map; the high-resolution feature map is respectively input into the Concat module 4 through the GSConv module 2 and at the same time input into the detection head 2 of the head network for middle-layer vibration area localization and laying scenario prediction;

[0106] The feature map output by the PLMSAM module in the backbone network is sent to the upsampling module 1 and the Concat module 4 in the neck network. The feature map is upsampled by the upsampling module 1 and then input into the Concat module 1. After passing through the Concat module 4, it is input into the VoVGSCSP module 4, and the output of the VoVGSCSP module 4 is the deep fusion feature map, which is input into the detection head 3 in the head network for deep vibration area localization and laying scenario prediction.

[0107] In this embodiment, the input of the GSConv module 1 is the feature map from the VoVGSCSP module 2, and the output is the compressed shallow-layer high-resolution feature map for use by the Concat module 3; the input of the GSConv module 2 is the feature map of the VoVGSCSP module 3, and the output is the compressed middle-layer high-resolution feature map for use by the Concat module 4;

[0108] The input of the VoVGSCSP module 1 is the concatenated middle and deep feature map obtained by the Concat module 1, and the output is the middle-layer fusion feature map for use by the upsampling 2 and the Concat module 3; the input of the VoVGSCSP module 2 is the concatenated shallow and middle feature map obtained by the Concat module 2, and the output is the shallow-layer fusion feature map for use by the detection head 1 and the GSConv module 1; the input of the VoVGSCSP module 3 is the concatenated shallow and middle feature map obtained by the Concat module 3, and the output is the middle-layer fusion feature map for use by the detection head 2 and the GSConv module 2; the input of the VoVGSCSP module 4 is the concatenated middle and deep feature map obtained by the Concat module 4, and the output is the deep fusion feature map for use by the detection head 3.

[0109] As Figure 7As shown in the figure, in this embodiment, the GSConv module includes a convolutional layer Conv, a depthwise separable convolutional layer DWConv, a Concat module, and a channel shuffle module Shuffle. Among them, Conv consists of a traditional convolution, a batch normalization operation, and an activation function, and DWConv consists of a depthwise separable convolution, a batch normalization operation, and an activation function. Suppose an input feature map with C1 channels is first processed by a traditional convolution to generate a feature map with C2 / 2 channels, and then processed by a depthwise separable convolution to generate another feature map with C2 / 2 channels. Without changing the number of channels, the result of the depth convolution is concatenated with the result of the depthwise separable convolution, and finally, the feature information from the traditional convolution and the depthwise separable convolution is integrated through the channel fusion mechanism of the Shuffle module to generate an output feature map with C2 channels.

[0110] In order to make full use of the function of GSConv in the neck network of YOLOv11n, in this embodiment, without changing the original input and output sizes of C3k2, all C3k2 modules are replaced with VoVGSCSP module 1, VoVGSCSP module 2, VoVGSCSP module 3, and VoVGSCSP module 4. This module adopts a single-stage aggregation strategy and designs an efficient cross-stage network module. By using a residual structure to integrate traditional convolution and GSConv, while maintaining good accuracy, it reduces the inference time complexity and the model parameter count. The structure of VoVGSCSP is as Figure 8 shown.

[0111] In this embodiment, each VoVGSCSP module is composed of convolutional layer 1, convolutional layer 2, convolutional layer 3, a GSConv module branch, a Concat module, and convolutional layer 5. Among them, the GSConv module branch includes GSConv module 1 and GSConv module 2.

[0112] The first part of the C1-channel feature map is obtained by performing shallow feature extraction through convolutional layer 1 to obtain a C1 / 2-channel feature map. The C1 / 2-channel feature map is respectively subjected to deep feature extraction through convolutional layer 2 and the GSConv module branch, and then deep feature maps with C2 / 2 channels are obtained through residual addition.

[0113] The second part of the C1-channel feature map is directly subjected to feature extraction through convolutional layer 3 to obtain a C2 / 2-channel backbone feature map.

[0114] Finally, through the Concat module and convolutional layer 5, the C2 / 2-channel deep feature map and the C2 / 2-channel backbone feature map are subjected to channel splicing and feature extraction to obtain a C2-channel fusion feature map.

[0115] Embodiment 2. Combining Figures 9 to 15 This embodiment will be described in combination with Embodiment 1. This embodiment is an experimental example of the multi-scenario vibration area positioning method for underground power optical cables based on the PLGS-YOLO model described in Embodiment 1:

[0116] 1. Experimental scheme design and deployment;

[0117] In this embodiment, the experimental object is an underground power optical cable line with multiple scenarios in a communication station in Jilin. The total length is 1.98 km, of which the total length of the optical cable within the communication station is 0.42 km, the cumulative length of the buried optical cable outside the communication station is about 1.5 km, and the length of the optical cable in each well is about 0.03 km. In the laying of underground power optical cables, the optical cables in the buried area are protected by sleeves, while the optical cables in the well and within the station area are not. The red line segments in the figure represent the sleeve protection structure outside the optical cable. In order to accurately collect the vibration signals of the specified buried area and the optical cable in the well, first, a 1-km-long G652D single-mode optical fiber is connected to the Ф-OTDR system. The other end of the single-mode optical fiber is connected to one of the cores of the underground power optical cable. The pulse width of the acquisition card is set to 32 ns, the frequency is set to 1000 Hz, and the spatial resolution is 2 m. The scenario and instrument deployment are as Figure 9 shown.

[0118] Fifteen different types of vibration event experiments are carried out in the specified areas of the buried and well scenarios of the optical cable line, as Figure 12 shown. Among them, the buried scenario includes background noise, motorcycle, trolley, excavation, pickaxe, shot put, walking, electric drill and hammer; the well scenario includes background noise, foreign object falling, electric drill, dragging, climbing, knocking, walking. The specific event descriptions are shown in Tables 1 and Tables 1 is the description of the specific vibration events in the buried area, and Table 2 is the description of the specific vibration events in the well.

[0119] Table 1

[0120] Event type Detailed description Background noise Natural state without any disturbance Motorcycle A person rides a motorcycle back and forth through the buried collection area Wheelbarrow Push a single wheelbarrow carrying heavy objects back and forth through the buried collection area Digging Use a single shovel to dig in the buried collection area every 3 seconds Pickaxe Use a single pickaxe to strike the ground forcefully in the buried collection area Shot put A person drops a 5 kg shot put onto the ground every 3 s Walking Two people walk back and forth in opposite directions at a constant speed through the buried collection area Electric drill and hammer drill Use a single multi-functional electric drill to continuously drill the stone slab in the buried collection area

[0121] Table 2

[0122] Event type Detailed description Background noise There is no disturbance on the ground and no one goes down the well Foreign object falling A person drops a hammer or wrench from the well using a rope Electric drill A person uses an electric drill to continuously drill the stone in the vertical direction of the optical cable Pulling and dragging A person uses the sharp end of a hammer to pull up and then lower the optical cable every 2 s Climbing A person crawls back and forth on a ladder Knocking A person uses a hammer to strike different areas of the optical cable in the well every 1 s Walking A person walks back and forth at a constant speed along the line

[0123] 2. Data preprocessing and dataset creation;

[0124] The Ф-OTDR system is used to collect the vibration signals in the specified test area, and at the same time, manually record the event types, spatial positions and placement scenarios of different vibrations. The Ф-OTDR signals output by the Ф-OTDR system are multiple 1024×781 MAT files.

[0125] The Φ-OTDR signal is processed by a high-pass filter (the high-pass filter coefficient b2 is preset in the instrument, with a cut-off frequency of 2 Hz, a passband of 5 Hz, and a stopband attenuation of -80 dB). Then it is processed by a low-pass filter (the low-pass filter coefficient b3 is preset in the instrument, with a cut-off frequency of 150 Hz, a passband of 100 Hz, and a stopband attenuation of -80 dB), and finally the preprocessed spatio-temporal signal is obtained.

[0126] By combining low-pass and high-pass filtering, their effects on different vibration events are compared, such as the falling of foreign objects and excavation, where low-pass filtering is required to maintain the trend of the main signal, and knocking or using an electric drill or hammer drill, where high-pass filtering is needed to extract high-frequency details. Figure 10 、 Figure 11 These visualizations in

[0127] more clearly show the impact of filtering on signal quality and feature retention.

[0128] The spatio-temporal signal is converted into a spatio-temporal image of 1024×781. Based on the manually recorded spatial positions and laying scenarios, Labelme is used to label the laying scenario types and vibration regions to form a positioning dataset. In the positioning dataset, background noise refers to the signal without vibration events, and such samples are not labeled with laying scenario types or vibration regions. Here, 1024 represents the length of each data sample in the time dimension, and 781 represents the number of spatial sampling points.

[0128] The positioning dataset is randomly divided into a training set, a validation set, and a test set at a ratio of 7:2:1. The detailed information of the dataset is shown in Table 3. Table 3 is the description of vibration events in multiple laying scenarios.

[0129] Table 3

[0130]

[0131]

[0132] In Figure 12 、 13 , the green box represents the underground laying scenario, and the red box represents the underground mine laying scenario. Regardless of which laying scenario the vibration event occurs in, the vibration signal will spread to the surrounding areas. However, the vibration diffusion characteristics in the underground mine scenario are more significant, the signal intensity is more concentrated, and the vibration region is easier to distinguish; while in the underground laying scenario, the signal propagation is affected by the casing and soil layer medium, the diffusion range is smaller, the vibration intensity is weaker, and it is more difficult to distinguish the signal characteristics.

[0133] 3. Parameter settings;

[0134] In this embodiment, PyTorch 1.11.0 is selected as the deep learning framework. All experiments are conducted on a workstation equipped with an NVIDIA GeForce RTX 2080 GPU with 64GB of memory, and the CUDA version is 11.3. For all localization models in this experiment, the batch size is set to 8 and the number of epochs is set to 100.

[0135] 4. Evaluation metrics;

[0136] When evaluating the model in the localization method, in this embodiment, the scale of the model is measured by Parameters and GFLOPs. Parameters represents the spatial complexity of the model, while GFLOPs (floating point operations per second) is used to describe the computational complexity. The test time reflects the average localization time of the model.

[0137] To evaluate the performance of vibration region localization, Precision, Recall, mAP@50, and mAP@50-95 are used as key metrics, and each metric reflects different aspects of the localization process:

[0138] Precision focuses on the accuracy of vibration region prediction. A higher Precision can ensure that the detected vibration regions truly correspond to the actual vibration sources, reducing false detections of background noise or environmental interference as vibration events.

[0139] Recall emphasizes the model's ability to capture all existing vibration regions. A higher Recall means that the system can effectively detect weak or small-scale vibrations, which is crucial for detecting subtle disturbances in underground power cables.

[0140] The mean average precision at an IoU threshold of 0.5 (mAP@50) evaluates the correctness of vibration region localization. When allowing a moderate overlap with the ground truth (Intersection over Union IoU = 50%), this metric is particularly applicable to practical applications where a rough localization of vibration events is sufficient for early warning and monitoring.

[0141] The mean average precision at IoU thresholds from 0.5 to 0.95 (mAP@50-95) provides a more rigorous evaluation by assessing the localization accuracy at multiple IoU thresholds. This reflects the model's performance in precisely delineating the boundaries of vibration regions, which is crucial for differentiating overlapping or closely spaced vibration events.

[0142] The bounding box regression loss (Boxloss) is used to measure the deviation between the bounding box predicted by the model and the true annotation box. In the vibration area localization, the bounding box corresponds to the vibration occurrence area predicted by the model. The smaller the Boxloss, the more accurately the model locates the actual vibration area range.

[0143] The class classification loss (Clsloss) measures whether the model's judgment of the vibration area class (i.e., different laying scenarios or vibration event types) is accurate. In the vibration area localization, Clsloss reflects the model's ability to identify whether the area belongs to laying scenarios such as underground or underground.

[0144] To verify the effectiveness of the PLGS-YOLO model described in this embodiment, ablation experiments and comparative experiments were carried out, and the analysis was performed based on the results of the test set in the localization dataset. In the ablation experiments, by gradually introducing the PLMSAM module, GSConv module, and VoVGSCSP module (abbreviated as GSVO), and the SPPFDMSCA module, the improvement effects of each module were verified. In the comparative experiments, PLGS-YOLO was comprehensively compared with lightweight models of the YOLO series (such as YOLOv3-tiny, YOLOv5-tiny, YOLOv8n, etc.). The results showed that the PLGS-YOLO model achieved the best performance in terms of detection accuracy, computational complexity, and real-time performance.

[0145] First, an ablation experiment was carried out using the PLGS-YOLO model;

[0146] Based on YOLOv11n, combined with PLMSAM, GSConv, VoVGSCSP, and SPPFDMSCA, ablation experiments were carried out, where GSVO represents GSConv and VoVGSCSP. The results of the PLGS-YOLO model in different ablation experiments are shown in Table 4.

[0147] Table 4

[0148]

[0149] First, the PLMSAM module was used to improve C2PSA of the backbone network. The experimental results show that the parameters of the model are reduced, the computational load remains basically unchanged, Precison is improved from 93.99% to 94.30%, mAP@50 is improved from 97.27% to 97.34%, mAP@50-95 is improved from 59.94% to 60.53%, Recall is improved from 94.05% to 94.72%, and the test time is reduced from 6.896 ms to 6.790 ms. PLMSAM achieved the improvement and lightweighting of YOLOv11n performance by improving the PSA module and discarding the skip connection structure of the original model. Next, on this basis, Conv in the neck network was replaced with GSConv, and C3k2 was replaced with VoVGSCSP. The experimental results show that GSConv and VoVGSCSP fully extract the deep features of PLMSAM. Compared with only improving PLMSAM in YOLOv11n, the parameters, computational load, and test time of the model are reduced, and Precison, mAP@50, mAP@50-95, and Recall are improved. However, compared with YOLOv11n, Recall still decreased by 0.12%. The SPPF of YOLOv11n will lose some information of the deep features, which is not conducive to feature fusion in the neck network. On the basis of improving PLMSAM and GSVO in YOLOv11n, SPPF was replaced with SPPFDMSCA. Experiments show that although some indicators are not as good as the other ablation models, all performance indicators of PLGS-YOLO are better than the original YOLOv11n model. Precison is improved from 93.99% to 94.46%, Recall is improved from 94.05% to 94.55%, mAP@50 is improved from 97.27% to 97.51%, mAP@50-95 is improved from 59.94% to 61.27%, and the computational load and test time of the model are reduced.

[0150] Then, the performance analysis of the PLGS-YOLO model;

[0151] Mainstream lightweight models in the YOLO series were selected to verify the effectiveness of PLGS-YOLO. These models include YOLOv3-tiny, YOLOv5-tiny, YOLOv7-tiny, YOLOv8n, YOLOv8n+CBAM, YOLOv9n, YOLOv10n, and YOLOv11n. The experiments were carried out under the same conditions. Among them, YOLOv3-tiny and YOLOv8n+CBAM have been applied to the vibration positioning task, and the remaining models directly transplanted the original technical framework to vibration image processing for comparative analysis. All the experiments were carried out under the same conditions, and the results are shown in Table 5, which is a comparison of different object detection models.

[0152] Table 5

[0153]

[0154]

[0155] The simplified backbone network of YOLOv5 is CSPDarknet-tiny, which significantly reduces the number of layers and residual blocks. YOLOv7-tiny only retains the basic spatial pyramid pooling module of the SPPCSPC module, reducing the levels of feature fusion. Although the YOLOv5-tiny and YOLOv7-tiny models meet the real-time requirements for underground power optical cable scene detection, their detection accuracy is significantly lower than our improved model. Compared with YOLOv5-tiny, the Precision of PLGS-YOLO increased by 1.77%, Recall increased by 1.67%, mAP@50 increased by 1.13%, and mAP@50-95 increased by 2.01%. Compared with YOLOv7-tiny, the Precision of PLGS-YOLO increased by 3.97%, Recall increased by 5.74%, mAP@50 increased by 4.19%, and mAP@50-95 increased by 7.70%.

[0156] The network structure of the YOLOv3-tiny model is relatively larger and more complex, and the fixed parameters of the anchor frame cannot fully adapt to multi-scale vibration signals, requiring more computing resources for detection and having poor real-time performance. Compared with PLGS-YOLO, the GFLOPS of the YOLOv3-tiny model increased by 12.60, parameters increased by 9.55M, and the test time increased by 1.18ms.

[0157] The SPPF of YOLOv8n and YOLOv8n-CBAM obtains deep features through multi-scale pooling, but features of different scales will contain duplicate information, causing feature redundancy. Directly using them for feature fusion in the neck network will affect the detection accuracy. Compared with PLGS-YOLO, the Precision of YOLOv8n decreased by 1.41%, Recall decreased by 0.51%, mAP@50 decreased by 0.27%, and mAP@50-95 decreased by 1.09%. The Precision of YOLOv8n-CBAM decreased by 2.50%, Recall decreased by 0.69%, mAP@50 decreased by 0.75%, and mAP@50-95 decreased by 1.53%.

[0158] YOLOv9n uses GELAN as part of the backbone network, replacing traditional convolutional layers and residual blocks to improve the feature extraction ability. However, GELAN adopts a multi-branch structure, which will significantly increase the computational complexity and testing time of the model. Compared with PLGS-YOLO, the GFLOPS of the YOLOv9n model increases by 5.70, and the testing time increases by 3.46 ms.

[0159] YOLOv10n uses PSA to further extract the deep features of SPPF, focusing on the key information of the features, reducing the interference of redundant information, and improving the feature fusion efficiency of the neck network. However, the internal use of the multi-head attention mechanism in PSA will cause some local detailed information to be lost, which is not conducive to the accurate positioning of vibration signals and increases the number of data to be detected. Compared with PLGS-YOLO, the Precision of YOLOv10n decreases by 2.99%, Recall decreases by 1.55%, mAP@50 decreases by 1.53%, and mAP@50-95 decreases by 4.02%.

[0160] By comparing the Precision, Recall, mAP@50, and mAP@50-95 of YOLOv3-tiny, YOLOv8n, YOLOv8n+CBAM, YOLOv9n, and YOLOv10n, the detection accuracy of the PLGS-YOLO model is the best. The GFLOPS, parameters, and testing time of PLGS-YOLO are the lowest, meeting the real-time requirement and having low deployment difficulty. In summary, PLGS-YOLO shows excellent overall performance in the scenario detection and positioning of underground power cable vibration signals collected by the -OTDR system.

[0161] As Figure 14 shown, in order to further evaluate the performance of PGSDYOLO in vibration area positioning, its bounding box regression loss (Boxloss), class classification loss (Clsloss), Precision, Recall, mean average precision (mAP@50) at IoU threshold of 0.5, and mean average precision (mAP@50-95) at IoU threshold from 0.5 to 0.95 on the validation set are compared with existing lightweight YOLO models. As the number of iterations increases, PLGS-YOLO can always obtain higher performance.

[0162] In this embodiment, in the underground power cable The samples collected by -OTDR in real time are used for laying scenario detection and positioning. PLGS-YOLO correctly labels the laying scenario types in each image. The green box represents the buried laying scenario, and the red box represents the underground laying scenario. It can accurately locate different types of vibration signal regions. Some visualization results are as follows Figure 15 .

[0163] The buried motorcycle signal appears as continuous high-frequency energy stripes in the spatio-temporal diagram, with high signal intensity and a wide range. The detection box completely covers the high-frequency signal region without being interfered by other noises. The buried knocking signal is irregularly distributed, with multiple short-time high-energy shock waves and less background noise. Multiple detection boxes respectively label each high-energy signal region to ensure complete signal coverage. The buried excavation signal appears as intermittent strong signals in the spatio-temporal diagram, with a frequency similar to that of the background noise but higher energy. The model successfully filters out the interference of the background signal and accurately frames the vibration signal region. The buried trolley signal appears as low-frequency, continuous energy stripes, with a relatively uniform distribution and low signal intensity. The detection box completely labels the low-frequency signal region without missing important information.

[0164] The underground drilling signal is concentrated in the middle of the image, appearing as continuous high-frequency energy stripes with a moderate distribution range. The detection box accurately labels the signal region and avoids background interference. The underground walking signal is intermittent, with sparse stripes but obvious energy and weak background noise. The detection box completely labels all intermittent signals without missing the main region. The underground knocking signal is scattered, appearing as multiple short-time, high-energy signal peaks, and the interval region is background noise. The model covers all high-energy signal regions through multiple detection boxes without being interfered by noise. The underground falling object signal appears as continuous high-energy stripes in the spatio-temporal diagram, with the signal gradually decaying and having dynamic characteristics. The detection box covers the entire signal path from high energy to low energy, and the signal region is complete.

[0165] In summary, PLGS-YOLO shows excellent vibration region positioning and scenario detection performance in both buried and underground laying scenarios.

[0166] The method for locating the vibration region of underground power optical cables in multiple laying scenarios described in this embodiment uses PLGS-YOLO to detect the laying scenario and extract the regional vibration signal. The rationality of the model is verified from the aspects of model real-time performance, deployment difficulty, detection accuracy, and positioning accuracy. PLGS-YOLO replaces the modules at the corresponding positions with PLMSAM, GSConv, VoVGSCSP, and SPPFDMSCA on the basis of YOLOv11n. The improved PLGS-YOLO is the best in terms of Precision, mAP@50, mAP@50-95, and Recall, and the number of parameters, computational complexity, and test time of this model are all relatively low.

[0167] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0168] The above-described embodiments only represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent shall be subject to the appended claims.

Claims

1. A method for locating multi-scene vibration regions of underground power optical cables based on the PLGS-YOLO model, characterized in that: This method is implemented by the following steps: Step 1: Preprocess the acquired vibration signal and convert the preprocessed signal into a spatio-temporal map; Step 2: Locate the vibration area; By constructing a PLGS-YOLO model, training the PLGS-YOLO model, and using the trained PLGS-YOLO model to locate the vibration area in the spatio-temporal map obtained in Step 1, a located spatio-temporal map and the corresponding laying scenario type are generated; The PLGS-YOLO model improves the YOLOv11n network. In the backbone network, the SPPF module and the C2PSA module are respectively replaced by the SPPFDMSCA module and the PLMSAM module; The convolutional layer in the neck network is replaced by the GSConv module, and the C3k2 module is replaced by the VoVGSCSP module; The DMSCA module is introduced into the SPPFDMSCA module to enhance the global feature extraction ability of the backbone network; The LMSCA module is used in the PLMSAM module to obtain a deep feature map; Input the spatio-temporal map into the YOLOv11n backbone network, perform feature enhancement through the dynamic multi-scale convolutional attention mechanism in the SPPFDMSCA module, and input the obtained global feature map into the PLMSAM module; After performing multi-scale convolutional operations on the global feature map through the LMSCA module in the PLMSAM module, input it into the neck network; The feature maps after multi-scale convolutional operations of the GSConv module and the VoVGSCSP module in the neck network are fused, and finally, a located spatio-temporal map and the corresponding laying scenario type are obtained through the detection head.

2. The method for locating multi-scenario vibration regions of underground power optical cables based on the PLGS-YOLO model according to claim 1, wherein: In Step 1, the acquired Φ-OTDR vibration signal is denoised by high-pass and low-pass filtering to obtain a preprocessed spatio-temporal signal; the spatio-temporal signal is converted into a spatio-temporal map.

3. The method for locating multi-scenario vibration regions of underground power optical cables based on the PLGS-YOLO model according to claim 1, wherein: The SPPFDMSCA module further includes a first convolutional layer, multiple max-pooling layers, a Concat module, and a second convolutional layer; The feature map output by the convolutional layer and the C3k2 module in the backbone network undergoes operations such as multi-layer residual connections through the first convolutional layer, multiple max-pooling layers, and the Concat module in the SPPFDMSCA module, and after passing through the second convolutional layer, is feature-enhanced by the DMSCA module to output a global feature map.

4. The method for locating multi-scenario vibration regions of underground power optical cables based on the PLGS-YOLO model according to claim 3, wherein: In the DMSCA module, multi-scale convolution operations are introduced. Convolution kernels of 3×3, 5×5, and 7×7 are respectively used to perform parallel convolution extraction on the input feature map X, and the features extracted at the three scales are concatenated and fused in the channel dimension to obtain the multi-scale feature map X ms ; Take the multi-scale feature map X ms Generate an attention map A through a 1×1 convolution and activate it through a Sigmoid function; meanwhile, the multi-scale feature map X ms After passing through a 1×1 convolution, generate key-value pairs K and V respectively, normalize K through a Softmax function, and multiply it with V to calculate the context-enhanced feature Y; Add the attention map A and the context-enhanced feature Y to the input feature map X to obtain the output global feature map.

5. The method for locating multi-scenario vibration regions of underground power optical cables based on the PLGS-YOLO model according to claim 1, wherein: The PLMSAM module includes a batch normalization layer 1, an LMSAM module, a batch normalization layer 2, and an FFN module. The LMSAM module includes a spatial attention module and an LMSCA module; The PLMSAM module receives the global feature map after the SPPFDMSCA module as input. First, through the residual connection structure composed of the batch normalization layer 1 and the LMSAM module, a multi-branch fusion feature map is obtained, then through the residual connection structure composed of the batch normalization layer 2 and the FFN module, a deep feature map is obtained. Finally, the deep feature map obtained by the PLMSAM module is input into the upsampling module 1 and the Concat module 4 of the neck network.

6. The method for locating multi-scene vibration regions of underground power optical cables based on the PLGS-YOLO model according to claim 5, characterized in that: The LMSCA module includes a 5×5 convolutional layer, a multi-branch convolutional layer, a 1×1 convolutional layer, and an average pooling layer. Local features are obtained through the 5×5 convolutional layer, the multi-branch convolutional layer captures multi-scale object relationships in different channels, and the original feature map is weighted by the average pooling layer and the 1×1 convolutional layer to obtain a weighted deep feature map.

7. The method for locating multi-scenario vibration regions of underground power optical cables based on the PLGS-YOLO model according to claim 3, wherein: The neck network includes multiple upsampling modules, multiple Concat modules, multiple GSConv modules, and multiple VoVGSCSP modules. The feature map output by the C3k2 module 2 in the backbone network goes to the second Concat module in the neck network. The feature map after the second Concat module is input into the second VoVGSCSP module, and the output is a compressed shallow-layer high-resolution feature map, which is input into the third Concat module through the first GSConv module. At the same time, it is input into the detection head 1 of the head network for shallow-layer vibration area localization and laying scenario prediction. The feature map output by the C3k2 module 3 in the backbone network is transmitted to the first Concat module in the neck network. After the first Concat module, it is input into the first VoVGSCSP module. The feature map output by the first VoVGSCSP module is input into the second upsampling module and the third Concat module. After upsampling by the second upsampling module, it is input into the second Concat module. After the third Concat module, it is input into the third VoVGSCSP module, and the output of the third VoVGSCSP module is a compressed middle-layer high-resolution feature map. The high-resolution feature map is respectively input into the fourth Concat module through the second GSConv module and input into the detection head 2 of the head network for middle-layer vibration area localization and laying scenario prediction. The feature map output by the PLMSAM module in the backbone network goes to the first upsampling module and the fourth Concat module in the neck network. The feature map is upsampled by the first upsampling module and then input into the first Concat module. It is input into the fourth VoVGSCSP module through the fourth Concat module, and the output of the fourth VoVGSCSP module is a deep fusion feature map, which is input into the detection head 3 in the head network for deep-layer vibration area localization and laying scenario prediction.

8. The method for multi-scenario vibration area positioning of underground power optical cables based on the PLGS-YOLO model according to claim 7, characterized in that: Each GSConv module consists of a convolutional layer Conv, a depthwise separable convolutional layer DWConv, a Concat module, and a Shuffle module. It is set that the feature map with C1 channels generates a feature map with C2 / 2 channels through the convolutional layer. At the same time, the feature map with C2 / 2 channels is processed through the depthwise separable convolutional layer to generate another feature map with C2 / 2 channels. The feature map with C2 / 2 channels generated by the convolutional layer and another feature map with C2 / 2 channels generated by the depthwise separable convolutional layer are connected through the Concat module and integrated through the channel fusion mechanism of the Shuffle module to finally generate an output feature map with C2 channels.

9. The method for locating multi-scenario vibration regions of underground power optical cables based on the PLGS-YOLO model according to claim 7, wherein: Each VoVGSCSP module consists of four convolutional layers, two GSConv modules, and a Concat module; It is set that a part of the feature map of the C1 channel is subjected to shallow feature extraction through the first convolutional layer to obtain the C1 / 2 channel feature map. The C1 / 2 channel feature map is respectively subjected to deep feature extraction through the second convolutional layer and two GSConv modules, and then the deep feature map of the C2 / 2 channel is obtained through residual addition; Another part of the C1 channel feature map is subjected to feature extraction through the third convolutional layer to obtain the backbone feature map of the C2 / 2 channel; Through the Concat module and the fourth convolutional layer, the deep feature map of the C2 / 2 channel and the backbone feature map of the C2 / 2 channel are subjected to channel splicing and feature extraction to obtain the fused feature map of the C2 channel.

Citation Information

Patent Citations

  • Sturgeon fry length category detection method

    CN118135317A

Cited By

  • Cable defect detection system

    CN120801369A