A vibration signal positioning and identification method for a phi-otdr system

CN119106356BActive Publication Date: 2026-09-29NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411208211.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-09-29
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

比如《Mixed Intrusion Events Recognition Based on Group ConvolutionalNeural Networks in DAS System》采用分类的方法解决了一个样本中具有多个振动事件的分类问题,但却无法将每个振动事件精确定位,《PIG Tracking Utilizing Fiber OpticDistributed Vibration Sensor and YOLO》,精确定位了振动事件的位置,却没有实现分类任务

Benefits of technology

[0027]第一,本发明的面向Φ-OTDR系统的振动信号定位与识别方法,提出的基于注意力机制的特征融合模块,使模型集中注意力于不易区分的信号的定位上,增强了定位的准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119106356B_ABST
    Figure CN119106356B_ABST
Patent Text Reader

Abstract

The application discloses a vibration signal positioning and identification method for a Phi-OTDR system, first acquires a distributed optical fiber sensing signal, then divides the signal, labels a signal category and a signal occurrence position, and obtains a time-space matrix sample data set; divides the time-space matrix sample data set to obtain a training set and a verification set; constructs an optical fiber vibration signal positioning and classification model according to the size of the sample, inputs the training set into the model to train the model, inputs the verification set into the trained model to verify the model, then adjusts a model hyperparameter to adapt to a field condition, and finally completes classification and positioning of an external vibration signal. The application can improve the intelligence of the Phi-OTDR system, effectively improve Phi-OTDR processing efficiency and reduce artificial burden.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of fiber optic sensing technology, specifically relating to a method for vibration signal localization and identification for Φ-OTDR systems. Background Technology

[0002] Φ-OTDR technology has been used for acoustic and vibration measurement in the Internet of Things (IoT) to build smart cities. With its long measurement distance and continuous, interference-resistant detection method, it can measure vibration events at any location along the fiber optic cable, achieving high-fidelity reconstruction of the amplitude and phase of external vibrations. Therefore, it has been widely applied in perimeter security and power system monitoring, oil and gas exploration, and pipeline leak detection, providing a high-sensitivity, high-capacity, and low-cost long-distance, all-weather acoustic / vibration dynamic detection solution for urban IoT.

[0003] In distributed fiber optic acoustic field sensing applications, fully automated signal localization and recognition are two crucial tasks affecting the system's intelligence level. Most signal processing methods are either localization or recognition methods, or simply a combination of both. For example, the paper "Mixed Intrusion Events Recognition Based on Group Convolutional Neural Networks in DAS System" uses a classification method to solve the problem of classifying multiple vibration events in a sample, but it cannot accurately locate each vibration event. "PIG Tracking Utilizing Fiber Optic Distributed Vibration Sensor and YOLO" accurately locates the vibration events but fails to achieve the classification task. Furthermore, due to the diversity of intrusion methods and the complexity of environmental noise, vibration signal localization and recognition suffer from high interference alarms, making the aforementioned methods inefficient in achieving localization and recognition tasks. Summary of the Invention

[0004] Technical problem solved: This invention proposes a vibration signal localization and identification method for Φ-OTDR systems, which can simultaneously locate and identify external vibration signals, effectively improving the intelligence and processing efficiency of Φ-OTDR systems and reducing manual workload.

[0005] Technical solution:

[0006] A method for vibration signal localization and identification for a Φ-OTDR system, the method comprising the following steps:

[0007] Step 1: Use the Φ-OTDR system to collect external vibration signals and label the collected vibration signals with their category and location;

[0008] Step 2: Divide the labeled vibration signal data into fixed sizes and construct a spatiotemporal matrix sample dataset. The spatiotemporal matrix sample dataset contains two parts: a validation set and a training set. All samples in the dataset have a spatiotemporal matrix of the same size, and each sample has a category label and a vibration location label.

[0009] Step 3: Construct a fiber optic vibration signal localization and classification model based on sample size. This model comprises four parts: a backbone network, a feature extraction and attention fusion network, a detection network, and a post-processing structure. The backbone network uses multiple group convolutions stacked together, followed by multiple residual modules connected in series to perform preliminary feature extraction and compression on the input vibration signal, obtaining multi-layer feature maps with varying degrees of compression, which are then output to the feature extraction and attention fusion network. The feature extraction and attention fusion network uses an FPN structure to further extract features from the multi-layer feature maps, and combines the lowest-level features extracted by the backbone network with the FPN structure. The feature maps extracted from different levels by the N-structure are fused to output a one-dimensional multi-layer compressed feature map, which is then sent to the detection network as a prediction feature map. The detection network includes a first detection head that predicts localization adjustment parameters and a second detection head that predicts classification probabilities, respectively. The number of each detection head is the same as the number of layers in the multi-level feature map. The post-processing structure applies the localization adjustment parameters to a preset length to obtain the signal position, filters low-probability prediction results, sets a confidence threshold, and filters out detection boxes below the threshold. Typically, in practical applications, the threshold is set to 0.75, so detection boxes below 0.75 are discarded. Prediction results exceeding the input sample size are cropped, and non-maximum suppression (NMS) is applied to prediction results of the same category to eliminate redundant detection boxes.

[0010] Step 4: Train the fiber optic vibration signal localization and classification model using the training set. After training, validate the model using the validation set and optimize the model hyperparameters based on the validation results.

[0011] Furthermore, in step 1, the vibration signal categories include at least five types: touching optical fiber, walking, manual digging, excavator digging, and climbing.

[0012] Furthermore, in step 2, the backbone network uses multiple group convolutional blocks (G-BLOCK) to perform preliminary feature extraction and compression on the input vibration signal, obtaining multi-layer feature maps with different degrees of compression. This process includes the following steps:

[0013] The input vibration signal is divided into n groups, and features are extracted from each group by the convolution kernels of n group convolutional blocks (G-BLOCK). Each G-BLOCK convolution kernel performs convolution on only one of the input groups. The group convolution stage is described as follows:

[0014]

[0015] In the formula, O is the output, f is a nonlinear activation function that introduces a nonlinear factor to increase the nonlinear fitting ability of the model, C, w_i, b_i, x_Parti represent the number of convolution kernels, convolution weights, convolution biases, and the input of the i-th group, respectively.

[0016] G-BLOCK is used for initial feature extraction to learn the vibration characteristics at different vibration locations, and features are extracted in the time dimension while maintaining spatial dimension invariance.

[0017] An FP-Block consisting of four stacked RES-Blocks is added to generate the original feature pyramid, thereby improving the correlation between the outputs of the group convolution and obtaining predicted feature maps with different receptive fields. Signals with a larger vibration range are predicted using small feature maps with a large receptive field, while signals with a narrower vibration range are predicted using large feature maps with a smaller receptive field.

[0018] Further, in step 2, the process of sending the one-dimensional multi-layer compressed feature map output by the feature extraction and attention fusion network as a predicted feature map to the detection network includes the following steps:

[0019] The feature maps of each layer are processed by 2DConv, and then the feature maps of the upper layer are upsampled and added to the similar feature maps of the lower layer, so that the target location and high semantic information are fused together.

[0020] Multiple g-conv layers are added, and each g-conv layer compresses each fused feature map. All output feature maps are compressed in the time dimension, and each two-dimensional feature map is converted into a one-dimensional feature map.

[0021] In the Fusion-A module, all compressed one-dimensional feature maps are fused with the original 2D feature map of the lowest layer to enhance the signal localization effect. The Fusion-A module contains a g-conv layer for compressing the original feature map and two 1×1 convolutional layers for adjusting the number of channels. The one-dimensional multi-layer compressed feature map is obtained by multiplying the outputs of the two 1×1 convolutional layers, and is output as the predicted feature map to the detection network.

[0022] Further, in step 2, the detection network includes a first detection head that predicts positioning adjustment parameters and a second detection head that predicts classification probabilities, respectively predicting the positioning adjustment parameters and the classification probabilities.

[0023] A set of default length values ​​is established for each prediction point in the predicted feature map, and the default length depends on the receptive field of the feature map; each point in the predicted feature map is used to make a prediction on an anchor box with three default lengths.

[0024] For each prediction point, the anchor box adjustment parameters and prediction probability are generated using two 1D convs of size 3. The number of output channels of the 1D convs used for classification and localization are the number of object categories and the number of adjustment parameters, respectively.

[0025] Furthermore, in step 4, during the training process of the fiber optic vibration signal localization and classification model, loss functions are set for the localization and classification tasks respectively. The total loss function is a linear superposition of the localization task loss function and the classification task loss function. The stochastic gradient descent method is used to iteratively train the fiber optic vibration signal localization and classification model.

[0026] Beneficial effects:

[0027] First, the vibration signal localization and identification method for Φ-OTDR systems proposed in this invention uses a feature fusion module based on an attention mechanism, which enables the model to focus its attention on the localization of signals that are difficult to distinguish, thereby enhancing the accuracy of localization.

[0028] Second, the vibration signal localization and identification method for Φ-OTDR system of the present invention uses group convolution for initial feature extraction, learns the characteristics of vibration changes at different vibration positions, realizes feature extraction in the time dimension, thereby realizing data processing at high speed and efficiency, and extracts features from the initial unbalanced data after completion.

[0029] Third, the vibration signal localization and identification method for Φ-OTDR systems of the present invention can simultaneously locate and identify external vibration signals, thereby improving the intelligence level and processing efficiency of Φ-OTDR systems. Attached Figure Description

[0030] Figure 1 This is a block diagram of the fiber optic vibration signal localization and classification model of the present invention;

[0031] Figure 2 This is a parameter diagram of the fiber optic vibration signal localization and classification model of the present invention;

[0032] Figure 3 These are the mAP curve and loss curve of the verification results of this invention;

[0033] Figure 4 This is the F1 score curve, which is the verification result of this invention. Detailed Implementation

[0034] The following embodiments are provided to enable those skilled in the art to more fully understand the present invention, but do not limit the invention in any way.

[0035] This invention discloses a method for vibration signal localization and identification for Φ-OTDR systems, the method comprising the following steps:

[0036] Step 1: Use the Φ-OTDR system to collect external vibration signals and label the collected vibration signals with their category and location;

[0037] Step 2: Divide the labeled vibration signal data into fixed sizes to construct a spatiotemporal matrix sample dataset, which includes a validation set and a training set.

[0038] Step 3: Construct a fiber optic vibration signal localization and classification model based on sample size. This model comprises four parts: a backbone network, a feature extraction and attention fusion network, a detection network, and a post-processing structure. The backbone network uses multiple group convolutions stacked together, followed by multiple residual modules connected in series to perform preliminary feature extraction and compression on the input vibration signal, obtaining multi-layer feature maps with varying degrees of compression, which are then output to the feature extraction and attention fusion network. The feature extraction and attention fusion network uses an FPN structure to further extract features from the multi-layer feature maps, and combines the lowest-level features extracted by the backbone network with the FPN structure. The feature maps extracted from different levels by the N-structure are fused to output a one-dimensional multi-layer compressed feature map, which is then sent to the detection network as a prediction feature map. The detection network includes a first detection head that predicts localization adjustment parameters and a second detection head that predicts classification probabilities, respectively. The number of each detection head is the same as the number of layers in the multi-level feature map. The post-processing structure applies the localization adjustment parameters to a preset length to obtain the signal position, filters low-probability prediction results, sets a confidence threshold, and filters out detection boxes below the threshold. Typically, in practical applications, the threshold is set to 0.75, so detection boxes below 0.75 are discarded. Prediction results exceeding the input sample size are cropped, and non-maximum suppression (NMS) is applied to prediction results of the same category to eliminate redundant detection boxes.

[0039] Step 4: Train the fiber optic vibration signal localization and classification model using the training set. After training, validate the model using the validation set and optimize the model hyperparameters based on the validation results.

[0040] The Φ-OTDR system of this invention senses external vibrations based on acoustic phase sensing. Laser light emitted from a narrow-linewidth distributed feedback fiber laser is modulated into pulsed light by an acousto-optic modulator, amplified by an erbium-doped fiber amplifier, and then enters the sensing fiber. The backscattered Rayleigh signal from the fiber enters the coupler via a circulator. Two Faraday rotator mirrors are connected to the two ports on the other side of the coupler, forming a Michelson interferometer. The incident light is split into two beams, reflected by the Faraday rotator mirrors, and interferes at the coupler. The interference signal is received by the detector and finally collected by a host computer for data processing. The host computer is used to acquire and save vibration signals in real time and uses the described vibration signal localization and identification method to locate and identify external vibration signals.

[0041] Taking the application of pipeline external damage alarm detection as an example, the vibration signal localization and identification method in this example includes the following steps:

[0042] Step 1: First, the Φ-OTDR system needs to be connected to the sensing optical cable of the monitoring pipeline to collect a certain number of vibration signals near the pipeline.

[0043] Step 2: Divide the labeled data into equal-sized segments to construct a spatiotemporal matrix sample dataset. This dataset includes both a validation set and a training set. The temporal length of the spatiotemporal matrix samples should ideally encompass the complete vibration signal or reflect the characteristics of the vibration signal changing over time. Similarly, the spatial length of the spatiotemporal matrix samples should reflect the characteristics of the vibration signal changing over space.

[0044] Step 3, see overall network architecture. Figure 1 A model for locating and classifying fiber optic vibration signals was constructed based on the sample size. The model mainly consists of four parts: a backbone network, a feature extraction and attention fusion network, a detection network, and post-processing.

[0045] The backbone network employs multiple group convolutions stacked together, followed by multiple residual modules cascaded to extract four levels of feature maps. In the backbone network, due to prior conditions, the input is divided into n groups (P1, P2, P3... Pi... Pn). Next, n convolutional kernels (C1, C2, C3... Ci... Cn) are used to extract features from each group, with Ci convolving only the i-th part of the input group. The group convolution stage is described below:

[0046]

[0047] In the formula, O is the output, and f is a nonlinear activation function that introduces a nonlinear factor, which can increase the model's nonlinear fitting ability. Due to the imbalance between the temporal and spatial dimensions of the vibration data, G-BLOCK is used for initial feature extraction. This method can learn the vibration characteristics at different vibration locations and achieve temporal feature extraction while maintaining spatial dimension invariance. After group convolution, the original unbalanced input becomes a balanced feature map of size (100, 100). To improve the correlation between the outputs of group convolution, predictive feature maps with different receptive fields are obtained, and an FP-Block consisting of four stacked RES-Blocks is added to generate the original feature pyramid. The feature maps used to predict signals with different vibration ranges are of different sizes. Signals with larger vibration ranges are predicted using small feature maps with large receptive fields, while signals with narrower vibration ranges are predicted using large feature maps with smaller receptive fields.

[0048] The feature extraction and attention fusion network first employs a Feature Pyramid Network (FPN) structure to initially extract feature maps at different levels. Then, an attention-based fusion module is added for further feature fusion. To improve the accuracy of vibration signal classification and localization, feature fusion is performed between feature pyramids, and a Feature Attention Fusion-A block is added to further enhance the focus on vibration signal localization. The inputs to the attention-based fusion module are the lowest-level features extracted by the backbone network and feature maps at different levels of the FPN structure. The number of attention-based fusion modules is the same as the number of layers in the multi-level feature maps. For the predicted feature map, the semantic information of the low-level features is relatively scarce, but the target location is precise. The semantic information of the high-level features is rich, but the target location is coarse. By combining these features from different levels, the target location and high semantic information are fused together. In this process, 2DConv is used to process the feature map of each layer, then the feature map of the upper layer is upsampled and added to the similar feature map of the lower layer. In addition, to effectively reduce the computational cost, g-conv is added to the feature extraction, and the feature map is compressed. All output feature maps are compressed along the time dimension, converting each 2D feature map into a 1D feature map. Finally, in the Fusion-A module, all compressed feature maps are fused with the original 2D feature map of the lowest layer feature map to enhance the signal localization effect. The Fusion-A module contains g-conv for compressing the feature map of the original feature map and two 1×1 convolutions for adjusting the number of channels. The predicted feature map is obtained by dot-multiplying the two outputs.

[0049] The detection network consists of detection heads that predict localization adjustment parameters and detection heads that predict classification probabilities, with the number of each type of detection head being the same as the number of layers in the multi-level feature map. In the detection network, the network generates adjustment parameters and predicted probabilities for anchor boxes for each predicted point. Both tasks are accomplished using two 1D convs with a size of 3. The output channels for classification and localization are 6 and 2, respectively, representing the number of object classes and the adjustment parameters. Each point in the predicted feature map is used to predict anchor boxes with default lengths of three lengths. First, to simplify the location prediction task, a set of default lengths is established for each predicted point in the predicted feature map. The default lengths depend on the receptive field of the feature map.

[0050] The detection network consists of detection heads that predict localization and adjust parameters and detection heads that predict classification probabilities. The number of each type of detection head is the same as the number of layers in the multi-level feature map.

[0051] The post-processing mainly performs four tasks: applying the prediction adjustment parameters to the default length; filtering low-probability prediction results; pruning prediction results that exceed the input sample size; and performing NMS processing on prediction results of the same category.

[0052] During the training process of the model, loss functions are set separately for localization and classification tasks. The total loss function is a linear superposition of the localization loss function and the classification task loss function. Iterative training is performed using stochastic gradient descent.

[0053] The final network architecture details are as follows Figure 2 exhibit.

[0054] Step 4: Train the fiber optic vibration signal localization and classification model using the training dataset. After training, validate the model using the validation set. Based on the actual dataset, the model is selected to output a maximum of 10 predicted targets during the prediction process, ranked from highest to lowest prediction probability. The number of correct matches (TP, TruePositive), the number of missed detections (FN, False Negative), and the number of false (repeated) detections (FP) are calculated by comparing the predicted targets with the real sample labels. Two factors affecting the matching are the prediction probability threshold and the IOU (Intersection over Union) threshold between the predicted length and the actual length. Assuming the predicted category is correct, a match is considered successful only if both the prediction probability threshold and the IOU between the predicted length and the actual length exceed a certain threshold. This is TP. Precision and recall can then be calculated. In pipeline leak monitoring, these are related to the Nuisance Alarm Rate (NAR) and the Missing Alarm Rate (MAR), respectively.

[0055]

[0056]

[0057] In practical applications, the model must consider both the false positive rate and the false negative rate. Therefore, F1socre can be obtained by combining percision and recall.

[0058]

[0059] Similarly, plotting probability and recall on the x and y axes yields the PR curve, and the area under the curve is the AP (Average Precision). The average AP across all classes is then calculated as the map (Mean Average Precision). Assuming the IOU threshold for a successful match is set to 0.5, and the prediction probability thresholds are 100 values ​​ranging from 0 to 1 with intervals of 0.01, throughout the training process, we use a validation set to validate and evaluate the trained network after each training iteration, obtaining the mAP under different conditions. Then, we sum the mAP values ​​across the 100 probability thresholds and average them. Finally, we obtain the curve of training loss versus the averaged mAP. Figure 3 As shown.

[0060] The curve shows that the map accuracy can reach as high as 99.3%, which is a high value, indicating that the proposed method has high precision and recall, while the false positive rate and false negative rate are very low.

[0061] Furthermore, the preset condition for TP is IOU = 0.5, and the F1 score is plotted for different prediction probability thresholds, such as... Figure 4 As shown, the highest F1 score is 99% when the predicted probability is 0.621. Therefore, during model post-processing, predicted targets with a probability below 0.621 can be filtered out. This reduces the need for NMS operations, speeds up prediction, and improves the overall prediction performance of the method.

[0062] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A method for vibration signal localization and identification in a Φ-OTDR system, characterized in that, The vibration signal localization and identification method includes the following steps: Step 1: Use the Φ-OTDR system to collect external vibration signals and label the collected vibration signals with their category and location; Step 2: Divide the labeled vibration signal data into fixed sizes and construct a spatiotemporal matrix sample dataset. The spatiotemporal matrix sample dataset contains two parts: a validation set and a training set. All samples in the dataset have a spatiotemporal matrix of the same size, and each sample has a category label and a vibration location label. Step 3: Construct an optical fiber vibration signal localization and classification model based on sample size. The optical fiber vibration signal localization and classification model includes four parts: a backbone network, a feature extraction and attention fusion network, a detection network, and a post-processing structure. The backbone network uses multiple group convolutions stacked together, and multiple residual modules are connected in series after the group convolutions to perform preliminary feature extraction and compression on the input vibration signal, obtaining multi-layer feature maps with different degrees of compression, and outputting them to the feature extraction and attention fusion network. The feature extraction and attention fusion network uses an FPN structure to further extract features from the multi-layer feature maps, and fuses the lowest-level features extracted by the backbone network with the different-level feature maps extracted by the FPN structure, outputting a one-dimensional multi-layer compressed feature map as a prediction feature map and sending it to the detection network. The detection network includes a first detection head that predicts the localization adjustment parameters and a second detection head that predicts the classification probability, predicting the localization adjustment parameters and the classification probability, respectively. The number of each detection head is the same as the number of layers in the multi-level feature map. The post-processing structure applies the localization adjustment parameters to a preset length to obtain the signal position, filters low-probability prediction results, prunes prediction results that exceed the input sample size, and performs non-maximum suppression processing on prediction results of the same category. Step 4: Train the fiber optic vibration signal localization and classification model using the training set. After training, validate the model using the validation set and optimize the model hyperparameters based on the validation results. In step 2, the process of sending the one-dimensional multi-layer compressed feature map output by the feature extraction and attention fusion network as a predicted feature map to the detection network includes the following steps: The feature maps of each layer are processed by 2DConv, and then the feature maps of the upper layer are upsampled and added to the similar feature maps of the lower layer, so that the target location and high semantic information are fused together. Multiple g-conv layers are added, and each g-conv layer compresses each fused feature map. All output feature maps are compressed in the time dimension, and each two-dimensional feature map is converted into a one-dimensional feature map. In the Fusion-A module, all compressed one-dimensional feature maps are fused with the original 2D feature map of the lowest layer to enhance the signal localization effect. The Fusion-A module contains a g-conv layer for compressing the original feature map and two 1×1 convolutional layers for adjusting the number of channels. The one-dimensional multi-layer compressed feature map is obtained by multiplying the outputs of the two 1×1 convolutional layers, and is output as the predicted feature map to the detection network.

2. The vibration signal localization and identification method for a Φ-OTDR system according to claim 1, characterized in that, In step 1, the vibration signal categories include at least five types: touching optical fiber, walking, manual digging, excavator digging, and climbing.

3. The vibration signal localization and identification method for a Φ-OTDR system according to claim 1, characterized in that, In step 2, the backbone network uses multiple group convolutional blocks (G-BLOCK) to perform preliminary feature extraction and compression on the input vibration signal, obtaining multi-layer feature maps with different degrees of compression. The process includes the following steps: The input vibration signal is divided into n groups, and features are extracted from each group by the convolutional kernels of n group convolutional blocks (G-BLOCK). Each group convolutional block (G-BLOCK) performs convolution on only one of the corresponding groups of input to learn the vibration characteristics at different vibration locations, while maintaining spatial invariance and extracting features in the temporal dimension. The group convolution stage is described as follows: , In the formula, It is the output. It is a nonlinear activation function that introduces nonlinear factors to increase the nonlinear fitting ability of the model. C, w_i, b_i and x_Parti represent the number of convolution kernels, convolution weights, convolution bias and input of the i-th group, respectively. An FP-Block consisting of four stacked RES-Blocks is added to generate the original feature pyramid, thereby improving the correlation between the outputs of the group convolution and obtaining predicted feature maps with different receptive fields. Signals with a larger vibration range are predicted using small feature maps with a large receptive field, while signals with a narrower vibration range are predicted using large feature maps with a smaller receptive field.

4. The vibration signal localization and identification method for a Φ-OTDR system according to claim 1, characterized in that, In step 2, the detection network includes a first detection head that predicts localization adjustment parameters and a second detection head that predicts classification probabilities. The process of predicting localization adjustment parameters and classification probabilities respectively includes the following steps: A set of default length values ​​is established for each prediction point in the predicted feature map, and the default length values ​​depend on the receptive field of the feature map; each point in the predicted feature map is used to make a prediction on an anchor box with three default lengths. For each prediction point, the anchor box adjustment parameters and prediction probability are generated using two 1D convs of size 3. The number of output channels of the 1D convs used for classification and localization are the number of object categories and the number of localization adjustment parameters, respectively.

5. The vibration signal localization and identification method for a Φ-OTDR system according to claim 1, characterized in that, In step 4, during the training process of the fiber optic vibration signal localization and classification model, loss functions are set for the localization and classification tasks respectively. The total loss function is a linear superposition of the localization task loss function and the classification task loss function. The stochastic gradient descent method is used to iteratively train the fiber optic vibration signal localization and classification model.