Vibration space-time diagram event identification method based on distributed optical fiber sensing

By converting the optical fiber vibration signal into a two-dimensional spatiotemporal map and introducing attention modules and dynamic loss functions into the deep learning model, the problems of noise interference and event types in the event recognition of optical fiber vibration signal are solved, and high-precision and robust event recognition are achieved.

CN120123869APending Publication Date: 2025-06-10ANHUI JIALUE INFORMATION TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510192675.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing fiber vibration signal event recognition methods face noise interference, event type diversity and spatiotemporal map similarity when dealing with complex scenarios, resulting in a decrease in recognition accuracy and it is difficult to make full use of spatiotemporal correlation to improve recognition performance.

Method used

The vibration spatiotemporal graph event recognition method based on distributed fiber perception is adopted. By converting one-dimensional time series data into two-dimensional spatiotemporal graphs, and introducing channel attention modules and spatiotemporal attention modules into convolutional neural networks, combining dynamic loss functions, the model training and recognition performance is optimized.

Benefits of technology

It significantly improves the accuracy and robustness of event recognition, enhances the model's adaptability to complex backgrounds, reduces the false positive rate, and improves the system's stability to continuous events, and is suitable for a variety of fiber types and monitoring scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123869A_ABST
    Figure CN120123869A_ABST
Patent Text Reader

Abstract

The invention discloses a vibration space-time diagram event identification method based on distributed optical fiber sensing. According to the invention, through combination of spatial-temporal feature fusion and a deep learning technology, the accuracy and robustness of event identification are significantly improved. The method comprises the following steps of: firstly, integrating space-time dimension information of an optical fiber vibration signal into two-dimensional characteristics by a space-time diagram, and providing richer context association for a model; through introduction of the channel attention and space-time attention module, key features can be adaptively focused, noise interference can be suppressed, and the adaptability of the model to a complex background is enhanced. Meanwhile, the dynamic loss function optimizes the training process of different scale targets by dynamically adjusting the loss weights of the bounding box and the mask, and solves the problem of performance reduction caused by target scale difference in the traditional method. The method has strong generalization, can adapt to various optical fiber types and monitoring scenes through data enhancement and adaptive parameter adjustment, and reduces dependence on data annotation and scene matching in actual deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent optical fiber sensing, and particularly relates to a method for identifying vibration spatio-temporal map events based on distributed optical fiber sensing. Background Technique

[0002] With the rapid development of distributed optical fiber sensing technology, its applications in fields such as perimeter security, structural health monitoring, and oil and gas pipeline leakage detection are becoming more and more extensive. As an important technical branch, Distributed Optical Fiber Vibration Sensing (DOFVS) can realize real-time monitoring of various events such as intrusion events, mechanical construction interference, and traffic interference by detecting vibration signals along the optical fiber.

[0003] In practical applications, optical fiber vibration signals are usually collected in the form of time series. These signals not only contain the spatial distribution information of events but also the time dynamic information of the occurrence of events. To more intuitively analyze and process this information, researchers have proposed a method of converting one-dimensional time series signals into two-dimensional spatio-temporal maps. The two-dimensional spatio-temporal map is a visualization and analysis tool that combines the time dimension and the spatial dimension of optical fiber vibration signals. The rows represent time points, and the columns represent spatial points, which can clearly show the change rules of vibration events in time and space. In recent years, with the development of deep learning technology, researchers have begun to explore the application of deep learning models such as Convolutional Neural Networks (CNNs) in the event recognition of optical fiber vibration signals. For example, by inputting the two-dimensional spatio-temporal map into a convolutional neural network, the spatio-temporal features in the signal can be automatically extracted, thereby achieving high-precision classification of different events.

[0004] However, the existing methods still face challenges in dealing with complex actual scenarios: problems such as noise interference in vibration signals, diversity of event types, and similarity of different events on the spatio-temporal map may all lead to a decrease in the recognition accuracy. At the same time, how to make full use of the spatio-temporal correlation of vibration signals to further improve the performance of event recognition is still an urgent problem to be solved. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for identifying vibration spatio-temporal map events based on distributed optical fiber sensing to solve the above-mentioned problems.

[0006] The technical solution adopted by the present invention is as follows: A method for identifying vibration spatio-temporal map events based on distributed optical fiber sensing, the method comprising the following steps:

[0007] S1: Data collection and collation. First, use a distributed optical fiber sensing system to collect vibration signals along the optical fiber in real time, obtain high-precision time series data, and convert the one-dimensional time series data into a two-dimensional spatio-temporal map, where the rows represent the time dimension and the columns represent the spatial dimension;

[0008] S2: Define a vibration image recognition network. Use VGG as the basic framework of the recognition network, and add a channel attention module and a spatio-temporal attention module after each convolutional layer of the VGG network;

[0009] S3: Define the dynamic scale loss and define the bounding box scale loss and the bounding box position loss:

[0010]

[0011] where IoU represents the intersection over union of the predicted bounding box and the ground truth bounding box, ∈(.) is the Euclidean distance, b p and b gt are the centroids of the predicted bounding box B p and the target bounding box B gt respectively, and c is the diagonal length of the two bounding boxes.

[0012] Define the mask scale loss and the mask position loss as follows:

[0013]

[0014] where M p and M gt are the predicted result pixel set and the ground truth pixel set of the target respectively, d p and d gt are the distances between the average pixels of the predicted result and the ground truth and the origin in polar coordinates respectively, θ p and θ gt are the average angles of the predicted result and the ground truth pixels in polar coordinates respectively, and ω describes the difference between M p and M gt and is a weight parameter used to adjust the influence of each part in the loss function

[0015] S4: Train the model based on the recognition model and loss function. First, initialize the vibration image recognition network. The network parameters are given initial values by using random initialization or loading pre-trained weights. The use of pre-trained weights can significantly accelerate the training speed and improve the convergence performance of the model, especially in the case of limited data volume. Subsequently, build a training environment based on the deep learning framework, configure a high-performance GPU to accelerate the training process, and set key training hyperparameters, including the learning rate, optimizer (such as Adam or SGD), number of training epochs, and batch size, etc. The selection of these hyperparameters has an important impact on the training effect and convergence speed of the model;

[0016] S5: Use the trained model for event recognition. The model quickly extracts key features in the signal through its convolutional layer, channel attention module, and spatio-temporal attention module, and classifies and recognizes events. The model outputs the predicted probability of each event category. The system determines the event type according to the set threshold (such as the probability is greater than 0.5). If the predicted probability exceeds the threshold, it is determined that the corresponding event has occurred, and the corresponding alarm or response mechanism is triggered

[0017] In a preferred embodiment, in step S1, during the acquisition process, key information such as the time, location, and type of the event occurrence is synchronously recorded, providing a basis for subsequent data annotation. To ensure data quality, the collected vibration signals need to be preprocessed, including filtering and denoising, to remove environmental noise and interference signals and retain the effective features related to the event.

[0018] In a preferred embodiment, in step S1, converting the one-dimensional time series data into a two-dimensional spatio-temporal graph can intuitively display the distribution law of vibration events in time and space, providing richer feature information for event recognition. Subsequently, the spatio-temporal graph is annotated in detail. According to the event type (such as intrusion, construction interference, etc.), the event occurrence area and category are clearly marked to ensure the accuracy and consistency of the annotation, providing high-quality supervision information for model training.

[0019] In a preferred embodiment, in step S1, to improve the generalization ability and robustness of the model, the data set needs to be reasonably divided into a training set, a validation set, and a test set, and the ratio is recommended to be 7:2:1. Each data set needs to be representative in terms of event type and distribution to avoid overfitting of the model. In addition, data augmentation operations can be performed on the training data, such as translation on the time axis, flipping in space, etc., to further increase data diversity and improve the adaptability of the model to different scenarios.

[0020] In a preferred embodiment, in step S2, the channel attention module is constructed by utilizing the inter-channel relationship of the feature map. Given an intermediate feature map F with dimensions C×H×W l, where l is a positive integer, C, H, W, and F l respectively represent the number of channels, height, and width of F l . First, a global max pooling layer is applied to compress its spatial dimension. Then, two one-dimensional convolutional layers are used to generate a channel attention map ChaAtt(F l ) with dimensions C×H×W, and its formula is as follows:

[0021]

[0022] where σ represents the sigmoid function, f is the ReLU function, represents the feature map obtained through the global max pooling operation, and W 1 and W 2 are the first and second convolutional kernels respectively, with dimensions k×1×1. Padding operations are used in each convolutional layer to make the output size equal to C. When l is equal to 1, 2, and 3, k is set to 3, 5, and 7 respectively.

[0023] After obtaining ChaAtt(F l ), it is applied to refine the original feature map F l , and we get:

[0024]

[0025] where represents element-wise multiplication, enabling the values in ChaAtt(F l ) to be extended along the spatio-temporal dimension.

[0026] In a preferred embodiment, in step S2, the spatio-temporal attention module is constructed by utilizing the relationship between the time and space axes of the feature map. First, a 1×1 convolutional layer is used to aggregate information along the channel direction of F l to generate a feature map M l with dimensions H×W. Then, two 2D convolutional layers are applied to derive and obtain a spatio-temporal attention map STAtt(F l ), and its formula is

[0027] STAtt(F l ) = σ(Q 1 * f(Q 2 * M l ))

[0028] where Q 1 and Q 2They are the first and second convolutional kernels respectively, with a dimension of k×k. In addition, padding operations are used in each convolutional layer to avoid changes in spatio-temporal size. Since the spatio-temporal size decreases as l increases, when l equals 1, 2, and 3, k is set to 7, 5, and 3 respectively. After that, we get:

[0029]

[0030] During the multiplication process, the values in STAtt(F l ) are expanded along the channel dimension.

[0031] Finally, is the output result of the double attention migration and is output as the input to the next layer.

[0032] In a preferred embodiment, in step S3, ω takes the value of:

[0033] 0.5: A common value, indicating that the ratio of intersection to union has equal weight in the loss function.

[0034] 0.3 - 0.7: The values within this range can be adjusted according to experimental results to find the best balance point.

[0035] 1.0: This indicates full emphasis on the ratio of intersection to union, but may cause the model to overly focus on reducing misclassified pixels while ignoring correctly classified pixels.

[0036] In a preferred embodiment, in step S, the formula for calculating the ratio between the original image and the current feature map is:

[0037]

[0038] where w o , h o are the width and height of the original image, and w c , h c are the width and height of the current feature map. R OC The purpose is to determine the true target size because the target size changes when the model scales the image or subsamples the feature map.

[0039] Calculate the influence coefficients β B and β M as follows:

[0040]

[0041] where B gtmax and M gtmaxThe value range of is from 1 to 100. The influence coefficient of the loss is based on the area of the current target box, and its range is limited within δ.

[0042] δ is a parameter used to limit the upper bound of the loss influence coefficient β B and β M . This parameter ensures that even when the area of the target box is very large, the influence coefficient will not exceed a specific threshold, thus avoiding excessive impact on the loss function.

[0043] In a preferred embodiment, in step S4, during the training process, the preprocessed and labeled two-dimensional spatio-temporal map data is input into the recognition network batch by batch. The network extracts features layer by layer through convolutional layers, channel attention modules, and spatio-temporal attention modules, and outputs the prediction probabilities for each event category. Subsequently, according to the prediction results and the ground truth labels, the bounding box dynamic loss and the mask dynamic loss are calculated respectively. The dynamic loss function can adaptively adjust the loss weights according to the scale and position of the target, ensuring that the model can effectively learn in different scales and complex scenarios, thereby improving the classification performance of the model for different event types.

[0044] The calculated total loss value is propagated back to each layer of the network through the backpropagation algorithm. The optimizer adjusts the network parameters according to the set learning rate and gradient information to minimize the loss value. During the training process, a learning rate decay strategy is adopted, and the learning rate is gradually decreased as the number of training rounds increases to ensure that the model converges quickly in the early stage and is finely adjusted in the later stage to avoid overfitting. In addition, an early stopping mechanism can be introduced. When the performance on the validation set does not improve significantly for consecutive multiple rounds, the training is terminated in advance to further prevent overfitting.

[0045] After each training round, the model is evaluated using the validation set, and performance metrics such as accuracy, recall, and F1-score are calculated. According to the performance of the validation set, the learning rate or hyperparameters are dynamically adjusted to optimize the generalization ability of the model. When the model achieves satisfactory performance on the validation set, the model parameters are saved, and the model is comprehensively tested using the test set to evaluate its performance on unseen data, ensuring the reliability and stability of the model in practical applications.

[0046] In a preferred embodiment, in step S5, in the actual monitoring scenario, the distributed optical fiber sensing system continuously collects the vibration signals along the optical fiber and transmits these signals to the event recognition system in real time. The system first preprocesses the collected vibration signals, including filtering and denoising and normalization operations, to remove environmental interference and extract effective features related to events. Subsequently, the preprocessed signals are converted into two-dimensional spatio-temporal maps, forming an input format consistent with the training stage, providing a clear spatio-temporal feature representation for model recognition.

[0047] To improve the accuracy and reliability of recognition, the system also introduces a post - processing mechanism. For example, the recognition results within consecutive time windows are smoothed to avoid false alarms caused by individual misjudgments. At the same time, combined with the spatio - temporal continuity of events, the recognition results are verified to further reduce the false alarm rate. In addition, the system has the ability of self - learning, which can dynamically update the model parameters according to new recognition results to adapt to environmental changes and new event types, ensuring that the model maintains high performance during long - term operation.

[0048] In summary, due to the adoption of the above - mentioned technical solutions, the beneficial effects of the present invention are as follows:

[0049] 1. In the present invention, the time and space information of fiber optic vibration are combined to form a two - dimensional spatio - temporal map, providing a richer feature basis for event recognition and improving recognition accuracy. Channel and spatio - temporal attention are introduced to focus on important features, suppress redundant information, and enhance the model's adaptability to complex backgrounds. Dynamic loss function: Dynamically adjusts the loss weights according to event features, optimizes model training, and improves the classification performance for different events. Strong generalization ability: The method is applicable to various fiber optic types and monitoring scenarios, with good adaptability and broad application prospects.

[0050] 2. In the present invention, through the combination of spatio - temporal feature fusion and deep learning technology, the accuracy and robustness of event recognition are significantly improved. First of all, the spatio - temporal map integrates the spatio - temporal dimension information of fiber optic vibration signals into two - dimensional features, providing richer context associations for the model. The introduction of channel attention and spatio - temporal attention modules can adaptively focus on key features and suppress noise interference, enhancing the model's adaptability to complex backgrounds. At the same time, the dynamic loss function optimizes the training process of different - scale targets by dynamically adjusting the loss weights of bounding boxes and masks, solving the problem of performance degradation caused by target scale differences in traditional methods. In addition, this method has strong generalization ability. Through data augmentation and adaptive parameter adjustment, it can be adapted to various fiber optic types and monitoring scenarios, reducing the dependence on data annotation and scene matching in actual deployment.

[0051] 3. In the present invention, the combination of dynamic loss function and post - processing mechanism effectively reduces the false alarm rate and improves the stability of the system for continuous events. Through real - time signal processing and model inference, the system can quickly respond to events such as intrusion and illegal excavation, providing real - time early warnings for security monitoring and infrastructure protection. In addition, the self - learning ability of the model supports dynamic parameter updates, which can adapt to environmental changes and new event types, ensuring continuous high efficiency in long - term applications and providing strong technical support for the implementation of distributed fiber optic sensing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1Schematic diagram of the process principle of the present invention;

[0053] Figure 2 Schematic diagram of the recognition test result in the present invention. Specific implementation manners

[0054] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0055] Embodiment:

[0056] Referring to Figure 1-2 , a vibration spatio-temporal map event recognition method based on distributed optical fiber sensing, comprising the following steps:

[0057] Step 1: Data collection and collation

[0058] First, use a distributed optical fiber sensing system to collect vibration signals along the optical fiber in real time to obtain high-precision time series data. During the collection process, key information such as the time, location, and type of events is synchronously recorded to provide a basis for subsequent data annotation. To ensure data quality, the collected vibration signals need to be preprocessed, including filtering and denoising, to remove environmental noise and interference signals and retain the effective features related to events.

[0059] Next, convert the one-dimensional time series data into a two-dimensional spatio-temporal map, where the rows represent the time dimension and the columns represent the space dimension. This conversion can intuitively display the distribution law of vibration events in time and space and provide richer feature information for event recognition. Subsequently, the spatio-temporal map is carefully annotated, and the event occurrence areas and categories are clearly marked according to the event type (such as intrusion, construction interference, etc.) to ensure the accuracy and consistency of the annotation and provide high-quality supervision information for model training.

[0060] To improve the generalization ability and robustness of the model, the data set needs to be reasonably divided into a training set, a validation set, and a test set, and the ratio is preferably 7:2:1. Each data set needs to be representative in terms of event type and distribution to avoid overfitting of the model. In addition, data augmentation operations can also be performed on the training data, such as translation on the time axis, flipping in space, etc., to further increase data diversity and improve the adaptability of the model to different scenarios.

[0061] Step 2: Define the vibration image recognition network

[0062] Adopt VGG as the basic framework of the recognition network, and add a channel attention module and a spatio-temporal attention module after each convolutional layer of the VGG network.

[0063] The channel attention module is constructed by leveraging the relationships among the channels of the feature map. Given an intermediate feature map F of dimension C×H×W l , where l is a positive integer, C, H, W, and F l represent the number of channels, height, and width of F l , respectively. First, a global max pooling layer is applied to compress its spatial dimension. Then, two one-dimensional convolutional layers are used to generate a channel attention map ChaAtt(F l ) of dimension C×H×W, and its formula is as follows:

[0064]

[0065] where σ represents the sigmoid function, f is the ReLU function, represents the feature map obtained through the global max pooling operation, and W 1 and W 2 are the first and second convolutional kernels, respectively, with dimensions k×1×1. Padding operations are used in each convolutional layer to make the output size equal to C. When l equals 1, 2, and 3, k is set to 3, 5, and 7, respectively.

[0066] After obtaining ChaAtt(F l ), it is applied to refine the original feature map F l , resulting in:

[0067]

[0068] where represents element-wise multiplication, enabling the values in ChaAtt(F l ) to be extended along the spatio-temporal dimensions.

[0069] The spatio-temporal attention module is constructed by leveraging the relationships between the temporal and spatial axes of the feature map. First, a 1×1 convolutional layer is used to aggregate information along the channel direction of F l , generating a feature map M l of H×W. Then, two 2D convolutional layers are applied to derive and obtain a spatio-temporal attention map STAtt(F l ), and its formula is

[0070] STAtt(F l ) = σ(Q 1 * f(Q 2 * M l ))

[0071] where Q 1 and Q 2They are the first and second convolutional kernels respectively, with a dimension of k×k. In addition, padding operations are used in each convolutional layer to avoid changes in spatio-temporal size. Since the spatio-temporal size decreases as l increases, when l equals 1, 2, and 3, k is set to 7, 5, and 3 respectively. After that, we get:

[0072]

[0073] During the multiplication process, the values in STAtt(F l ) are expanded along the channel dimension.

[0074] Finally, is the output result of the dual attention migration and is output as the input to the next layer.

[0075] Step 3: Define the dynamic scale loss

[0076] Define the bounding box scale loss and the bounding box position loss:

[0077]

[0078] where IoU represents the intersection over union of the predicted bounding box and the ground truth bounding box, ∈(.) is the Euclidean distance, b p and b gt are the centroids of the predicted bounding box B p and the target bounding box B gt , and c is the diagonal length of the two bounding boxes.

[0079] Define the mask scale loss and the mask position loss as follows:

[0080]

[0081] where M p and M gt are the predicted result pixel set and the ground truth pixel set of the target, d p and d gt are the distances between the average pixels of the predicted result and the ground truth and the origin in polar coordinates, θ p and θ gt are the average angles of the predicted result and the ground truth pixels in polar coordinates, and ω describes the difference between M p and M gt , and it is a weight parameter used to adjust the influence of each part in the loss function.

[0082] The value range of ω:

[0083] 0.5: A common value indicating that the ratio of the intersection and union has equal weights in the loss function.

[0084] 0.3 - 0.7: Values within this range can be adjusted according to experimental results to find the optimal balance point.

[0085] 1.0: This indicates complete emphasis on the ratio of the intersection and union, but may cause the model to overly focus on reducing misclassified pixels while neglecting correctly classified pixels.

[0086] Calculate the ratio between the original image and the current feature map:

[0087]

[0088] where w o , h o are the width and height of the original image, and w c , h c are the width and height of the current feature map. The purpose of R OC is to determine the true target size because when the model scales the image or subsamples the feature map, the target size changes.

[0089] Calculate the influence coefficients β B and β M as follows:

[0090]

[0091] where B gtmax and M gtmax range from 1 to 100. The influence coefficients of the loss are based on the area of the current target box, and their range is limited within δ.

[0092] δ is a parameter used to limit the upper bound of the loss influence coefficients β B and β M This parameter ensures that even when the area of the target box is very large, the influence coefficients do not exceed a specific threshold, thus avoiding excessive impact on the loss function.

[0093] Considerations for choosing δ:

[0094] 1. Data characteristics: The value of δ should be chosen according to the characteristics of the dataset. If the area of the target boxes in the dataset varies greatly, a larger value of δ may be needed to adapt to this variation.

[0095] 2. Model stability: A smaller value of δ can provide a more stable training process because it limits the impact of a single sample on the loss function. However, if δ is too small, it may limit the model's ability to learn from large target boxes.

[0096] 3. Experimental adjustment: The optimal value of δ is usually determined through experiments. The best setting can be found by adjusting the value of δ and observing the performance of the model on the validation set.

[0097] Common values of δ:

[0098] 0.1 to 0.5: These values are usually used to provide moderate constraints while allowing the model to have a certain degree of adaptability to large target boxes.

[0099] 1.0: This value is relatively large and is suitable for datasets with large variations in the area of target boxes, or when the model needs to have better recognition ability for large target boxes.

[0100] Smaller values: If it is found that the model is too sensitive to certain samples, the value of δ can be tried to be reduced to reduce the impact of these samples on the loss function.

[0101] Practical suggestions:

[0102] Initial setting: Experiments can be started from a medium value (such as 0.5).

[0103] Adjustment strategy: Adjust the value of δ according to the performance of the model on the validation set. If the model is too sensitive to certain samples, try reducing δ; if the model has insufficient recognition ability for large target boxes, try increasing δ.

[0104] Cross-validation: Use cross-validation to evaluate the impact of different δ values on the model performance, so as to select the optimal setting.

[0105] Calculate the dynamic loss of the bounding box

[0106]

[0107] where and are respectively and influence factors of.

[0108] Calculate the dynamic loss of the mask

[0109]

[0110] where and are influence factors of.

[0111] Step 4: Train the model based on the recognition model and loss function

[0112] In the model training stage, first, initialize the vibration image recognition network. The initial values of the network parameters are assigned by using random initialization or loading pre-trained weights. The use of pre-trained weights can significantly accelerate the training speed and improve the convergence performance of the model, especially when the amount of data is limited. Subsequently, build a training environment based on a deep learning framework (such as PyTorch or TensorFlow), configure a high-performance GPU to accelerate the training process, and set key training hyperparameters, including the learning rate, optimizer (such as Adam or SGD), number of training epochs, and batch size, etc. The selection of these hyperparameters has an important impact on the training effect and convergence speed of the model.

[0113] During the training process, the preprocessed and annotated two-dimensional spatio-temporal map data is input into the recognition network batch by batch. The network extracts features layer by layer through convolutional layers, channel attention modules, and spatio-temporal attention modules, and outputs the predicted probabilities for each event category. Subsequently, according to the prediction results and the ground truth labels, the bounding box dynamic loss and the mask dynamic loss are calculated respectively. The dynamic loss function can adaptively adjust the loss weights according to the scale and position of the target, ensuring that the model can effectively learn in different scales and complex scenarios, thereby improving the classification performance of the model for different event types.

[0114] The calculated total loss value is propagated back to each layer of the network through the backpropagation algorithm. The optimizer adjusts the network parameters according to the set learning rate and gradient information to minimize the loss value. During the training process, a learning rate decay strategy is adopted, and the learning rate is gradually decreased as the number of training epochs increases to ensure that the model converges quickly in the early stage and makes fine adjustments in the later stage to avoid overfitting. In addition, an early stopping mechanism can be introduced. When the performance on the validation set does not improve significantly for several consecutive epochs, the training is terminated in advance to further prevent overfitting.

[0115] After each training epoch ends, use the validation set to evaluate the model, and calculate performance metrics such as accuracy, recall, and F1 score. According to the performance of the validation set, dynamically adjust the learning rate or hyperparameters to optimize the generalization ability of the model. When the model reaches satisfactory performance on the validation set, save the model parameters, and use the test set to comprehensively test the model to evaluate its performance on unseen data, ensuring the reliability and stability of the model in practical applications.

[0116] Step 5: Use the trained model for event recognition

[0117] In an actual monitoring scenario, the distributed fiber optic sensing system continuously collects vibration signals along the fiber optic cable and transmits these signals to the event recognition system in real time. The system first preprocesses the collected vibration signals, including filtering and denoising and normalization operations, to remove environmental interference and extract effective features related to events. Subsequently, the preprocessed signals are converted into two-dimensional spatio-temporal maps, forming an input format consistent with the training phase, providing a clear spatio-temporal feature representation for model recognition.

[0118] The generated two-dimensional spatio-temporal map is input into the trained vibration image recognition model. The model quickly extracts key features in the signal through its convolutional layer, channel attention module, and spatio-temporal attention module, and classifies and recognizes events. The model outputs the predicted probability of each event category. The system determines the event type according to the set threshold (such as the probability being greater than 0.5). If the predicted probability exceeds the threshold, it is determined that the corresponding event has occurred, and the corresponding alarm or response mechanism is triggered.

[0119] To improve the accuracy and reliability of recognition, the system also introduces a post-processing mechanism. For example, the recognition results within continuous time windows are smoothed to avoid false alarms caused by single misjudgments. At the same time, combining the spatio-temporal continuity of events, the recognition results are verified to further reduce the false alarm rate. In addition, the system also has the ability of self-learning, which can dynamically update the model parameters according to new recognition results to adapt to environmental changes and new event types, ensuring that the model maintains high performance during long-term operation.

[0120] Vibration spatio-temporal images were collected in the experimental scenario of excavator theft. The present invention was applied to identify theft events. From Figure 1 the results, the recognition results of the present invention are highly consistent with manual annotation, and can identify events missed by manual annotation.

[0121] In the present invention, the time and space information of fiber optic vibration are combined to form a two-dimensional spatio-temporal map, providing a richer feature basis for event recognition and improving recognition accuracy. Channel and spatio-temporal attention are introduced to focus on important features, suppress redundant information, and enhance the model's adaptability to complex backgrounds. Dynamic loss function: Dynamically adjust the loss weight according to event features, optimize model training, and improve the classification performance for different events. Strong generalization ability: The method is applicable to various fiber optic types and monitoring scenarios, with good adaptability and broad application prospects.

[0122] In the present invention, by combining spatio-temporal feature fusion with deep learning technology, the accuracy and robustness of event recognition are significantly improved. First, the spatio-temporal graph integrates the spatio-temporal dimensional information of the fiber optic vibration signal into two-dimensional features, providing the model with richer context associations; the introduction of the channel attention and spatio-temporal attention modules can adaptively focus on key features and suppress noise interference, enhancing the model's adaptability to complex backgrounds. At the same time, the dynamic loss function optimizes the training process of targets at different scales by dynamically adjusting the loss weights of the bounding boxes and masks, solving the problem of performance degradation caused by target scale differences in traditional methods. In addition, this method has strong generalization ability. Through data augmentation and adaptive parameter adjustment, it can adapt to various fiber types and monitoring scenarios, reducing the dependence on data annotation and scene matching in actual deployment.

[0123] In the present invention, the combination of the dynamic loss function and the post-processing mechanism effectively reduces the false alarm rate and improves the stability of the system for continuous events. Through real-time signal processing and model inference, the system can quickly respond to events such as intrusion and illegal excavation, providing real-time warnings for security monitoring and infrastructure protection. In addition, the self-learning ability of the model supports dynamic parameter updates, which can adapt to environmental changes and new event types, ensuring continuous high efficiency in long-term applications and providing strong technical support for the implementation of distributed fiber optic sensing technology.

[0124] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0125] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A vibration spatiotemporal event recognition method based on distributed optical fiber sensing, characterized in that: The method comprises the following steps: S1: Data collection and organization. First, the distributed optical fiber sensing system is used to collect vibration signals along the optical fiber in real time to obtain high-precision time series data. The one-dimensional time series data is converted into a two-dimensional space-time diagram, where the middle row represents the time dimension and the column represents the space dimension. S2: Define the vibration image recognition network, use VGG as the basic framework of the recognition network, and add a channel attention module and a spatiotemporal attention module after each convolutional layer of the VGG network; S3: Define dynamic scale loss, define bounding box scale loss and the bounding box position loss: Among them, IoU represents the intersection over union ratio of the predicted bounding box and the true bounding box, ∈(.) is the Euclidean distance, and b p and b gt is the predicted bounding box B p and the target bounding box B gt The centroid of , c is the diagonal length of the two bounding boxes; Defining mask scale loss and mask position loss as follows: Among them, M p and M gt is the target’s predicted pixel set and target’s true pixel set, d p and d gt is the distance between the predicted result and the target true value average pixel and the origin in polar coordinates, θ p and θ gt is the average angle between the predicted result and the target true value pixel in polar coordinates, ω describes the M p and M gt The difference between is a weight parameter used to adjust the influence of each part in the loss function; S4: Based on the recognition model and loss function training model, the vibration image recognition network is first initialized, and the network parameters are given initial values ​​by random initialization or loading pre-trained weights; the use of pre-trained weights can significantly speed up the training speed and improve the convergence performance of the model, especially when the amount of data is limited; then, a training environment based on a deep learning framework is built, a high-performance GPU is configured to accelerate the training process, and key training hyperparameters are set, including learning rate, optimizer, number of training rounds, and batch size; S5: Use the trained model for event recognition. The model quickly extracts key features from the signal and classifies and recognizes events through its convolutional layer, channel attention module, and spatiotemporal attention module. The model outputs the predicted probability of each event category. The system determines the event type based on the set threshold. If the predicted probability exceeds the threshold, it is determined that the corresponding event has occurred and the corresponding alarm or response mechanism is triggered.

2. A vibration spatiotemporal graph event recognition method based on distributed optical fiber sensing as claimed in claim 1, characterized in that: In step S1, during the collection process, key information such as the time, location and type of the event is synchronously recorded to provide a basis for subsequent data annotation; in order to ensure data quality, the collected vibration signal needs to be preprocessed, including filtering and denoising, to remove environmental noise and interference signals and retain valid features related to the event.

3. A vibration spatiotemporal graph event recognition method based on distributed optical fiber sensing as claimed in claim 1, characterized in that: In step S1, converting the one-dimensional time series data into a two-dimensional space-time diagram can intuitively display the distribution pattern of vibration events in time and space, providing richer feature information for event identification; then the space-time diagram is annotated in detail, and the event occurrence area and category are clearly marked according to the event type to ensure the accuracy and consistency of the annotation, providing high-quality supervision information for model training.

4. A vibration spatiotemporal graph event recognition method based on distributed optical fiber sensing as claimed in claim 1, characterized in that: In step S1, in order to improve the generalization ability and robustness of the model, the data set needs to be reasonably divided into a training set, a validation set, and a test set, with a recommended ratio of 7:2:1; each data set needs to be representative in terms of event type and distribution to avoid overfitting of the model.

5. The vibration spatiotemporal graph event recognition method based on distributed optical fiber sensing according to claim 1, characterized in that: In step S2, the channel attention module is constructed by utilizing the inter-channel relationship of the feature map; given an intermediate feature map F with a dimension of C×H×W l , where l is a positive integer, C, H, W and F l Respectively represent F l The number of channels, height, and width of the network are firstly applied with a global maximum pooling layer to compress its spatial dimensions; then, two one-dimensional convolutional layers are used to generate a channel attention map ChaAtt(F l ), the formula is as follows: Where σ represents the sigmoid function, f is the ReLU function, represents the feature map obtained by the global maximum pooling operation, W1 and W2 are the first and second convolution kernels, respectively, with dimensions of k×1×1; padding operations are used in each convolution layer to make the output size equal to C; when l is equal to 1, 2, and 3, k is set to 3, 5, and 7, respectively; After obtaining ChaAtt(F l ), and then applied to refine the original feature map F l ,get: in represents element-wise multiplication, so that ChaAtt(F l ) is expanded along the space-time dimension.

6. A vibration spatiotemporal event recognition method based on distributed optical fiber sensing as claimed in claim 1, characterized in that: In step S2, the spatiotemporal attention module is constructed by utilizing the relationship between the time and space axes of the feature graph; first, a 1×1 convolutional layer is used along F l Aggregate information in the channel direction to generate a H×W feature map M l ; Then, two 2D convolutional layers are applied to derive and obtain the H×W spatiotemporal attention map STAtt(F l ), whose formula is STAtt(F l )=σ(Q1*f(Q2*M l )) Where Q1 and Q2 are the first and second convolution kernels, respectively, with dimensions of k×k. In addition, padding operations are used in each convolution layer to avoid changes in the spatiotemporal size. Since the spatiotemporal size decreases with the increase of l, k is set to 7, 5, and 3 when l is equal to 1, 2, and 3, respectively. After that, we get: During the multiplication process, STAtt(F l ) are expanded along the channel dimension; at last, It is the output result of the dual attention migration and is output to the input of the next layer.

7. A vibration spatiotemporal graph event recognition method based on distributed optical fiber sensing as claimed in claim 1, characterized in that: In step S3, the value of ω is: 0.5: A commonly used value, indicating that the ratio of intersection and union has equal weight in the loss function; 0.3-0.7: The value in this range can be adjusted according to experimental results to find the best balance point; 1.0: This places full emphasis on the ratio of intersection and union, but may cause the model to focus too much on reducing misclassified pixels and ignore correctly classified pixels.

8. The vibration spatiotemporal graph event recognition method based on distributed optical fiber sensing according to claim 1, characterized in that: In step S3, the formula for calculating the ratio between the original image and the current feature map is: where w o ,h o is the width and height of the original image, w c ,h c is the width and height of the current feature map; R OC The purpose is to determine the true object size, because the object size changes when the model scales the image or subsamples the feature map; Calculate the influence coefficient β of the bounding box and mask B and β M The calculation of is as follows: Among them B gtmax and M gtmax The value range of is 1 to 100; the impact coefficient of the loss is based on the area of ​​the current target box, and its range is limited to δ; δ is a coefficient used to limit the loss effect β B and β M The upper limit parameter of This parameter ensures that even if the area of ​​the target box is very large, the influence coefficient will not exceed a certain threshold, thereby avoiding excessive impact on the loss function.

9. A vibration spatiotemporal graph event recognition method based on distributed optical fiber sensing as claimed in claim 1, characterized in that: In step S4, during the training process, the preprocessed and annotated two-dimensional spatiotemporal graph data is input into the recognition network batch by batch; The network extracts features layer by layer through convolutional layers, channel attention modules and spatiotemporal attention modules, and outputs the predicted probability for each event category; then, the bounding box dynamic loss and mask dynamic loss are calculated respectively based on the predicted results and the true labels; the dynamic loss function can adaptively adjust the loss weight according to the scale and position of the target, ensuring that the model can effectively learn at different scales and complex scenarios, thereby improving the model's classification performance for different event types.

10. The vibration spatiotemporal graph event recognition method based on distributed optical fiber sensing according to claim 1, characterized in that: In step S5, in an actual monitoring scenario, the distributed optical fiber sensing system continuously collects vibration signals along the optical fiber and transmits these signals to the event recognition system in real time; The system first preprocesses the collected vibration signals, including filtering, denoising and normalization operations, to remove environmental interference and extract effective features related to the event; The preprocessed signal is then converted into a two-dimensional space-time graph to form an input format consistent with the training phase, providing a clear space-time feature representation for model recognition.

Citation Information

Cited By

  • Double-branch sound-vibration fusion event identification and positioning method based on DAS and AI

    CN120508911A

  • A dual-branch acoustic-vibration fusion event recognition and positioning method based on DAS and AI

    CN120508911B

  • Distributed optical fiber sensing human activity vibration signal monitoring method

    CN121256719A