A method and apparatus for acoustic sensing event detection based on distributed underground optical fiber.
By introducing clue information from geographical and geological features and a clue masking network, combined with a cue learning method, and optimizing the acoustic Swing Transformer neural network model, efficient and accurate detection of acoustic sensing events is achieved, solving the problem of insufficient detection performance in existing technologies, especially in terms of accuracy in event category and spatiotemporal localization.
Patent Information
- Application Number
- CN202510050018.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-01-13
AI Technical Summary
Existing distributed fiber optic sensing systems have performance limitations in event detection, especially when using information from the external deployment environment, they cannot achieve accurate detection and lack precise detection of event type and time of occurrence.
We employ a pre-trained acoustic Swin Transformer neural network model, combined with geographical and geological features as cue information. Through cue masking networks and cue learning methods, we perform acoustic sensing event detection. We train the model using spatiotemporal data and cue information, and only adjust the cue vectors and task-related network structures to predict the acoustic sensing event categories and timestamps.
It improves the accuracy of acoustic event category detection and the precision of spatiotemporal distribution localization, reduces the number of model training parameters, enhances detection performance, and supports lightweight deployment and long-distance continuous detection.
Smart Images

Figure CN120008725B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of distributed fiber optic acoustic sensing technology, specifically relating to an acoustic sensing event detection method and device based on distributed buried optical fiber. Background Technology
[0002] As a novel acoustic wave detection technology, distributed fiber optic acoustic wave sensing technology utilizes large-area distributed optical fibers as sensing units. By detecting subtle bending and twisting of the fiber caused by physical events in the surrounding environment, as well as phase changes in the backscattered Rayleigh light within the fiber, it achieves real-time localization and sensing of external vibration or sound sources. Distributed fiber optic acoustic wave sensing technology features high sensitivity, high information richness, and multi-dimensional signal characteristics, and has been widely applied in perimeter security, water / oil and gas pipeline monitoring, and earthquake monitoring.
[0003] Chinese patent document CN114857504A discloses a pipeline safety monitoring method based on distributed optical fiber sensors and deep learning. The method includes building an oil and gas pipeline simulation platform, collecting data, performing data preprocessing based on wavelet denoising and normalization, using a convolutional neural network as a feature extractor and a support vector machine as a classifier, establishing a joint training model of the convolutional neural network and support vector machine, performing oil and gas pipeline safety identification, classifying the pipelines according to the output digital labels, and realizing pipeline safety monitoring.
[0004] Chinese patent document CN114139583A discloses a method and system for detecting abnormal events on highways. This method utilizes a distributed optical fiber sensing system installed in the highway guardrail to collect multiple parallel optical fiber signals, obtaining signals from several points to be measured in space. The points to be analyzed are then selected using a threshold method. Based on this, feature extraction is performed on the optical fiber signal of each monitoring point, and features are fused with those of adjacent points. These features are then input into a model to obtain a label for the abnormal event.
[0005] Chinese patent document CN115622626A discloses a distributed acoustic wave sensing speech information recognition system and method. The scheme involves the design of a distributed sensing fiber optic system, which uses sensing fiber optics to receive speech signals, uses a laser emitting unit to emit narrowband laser signals, uses a circulator to detect the speech signals, and finally uses an acquisition unit and an acquisition circulator to receive backscattered signals. Finally, a convolutional recurrent network unit is used to perform complex domain mapping on the scattered signals to reconstruct the speech signals.
[0006] Chinese patent document CN113295259A discloses a distributed optical fiber sensing system, which includes an optical fiber sensing module, a display module, a communication module, and a data processing module to realize intelligent detection and intelligent control of the target to be detected.
[0007] The aforementioned existing technologies primarily focus on the modular functional implementation of distributed optical fiber systems, with limited research on core technologies for event detection based on spatiotemporal data acquired by these systems, including event category and timing detection. Furthermore, the external deployment environment information of distributed optical fibers is not fully utilized, and even with improved performance designs, accurate detection remains elusive. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a method and device for acoustic sensing event detection based on distributed buried optical fibers. This method involves collecting spatiotemporal data of the phase of backscattered Rayleigh light from distributed buried optical fibers. Based on the input spatiotemporal data, an acoustic sensing event neural network model is used to detect acoustic sensing events. This model utilizes a cue learning method, adjusting only the cue vectors and task-related network structures during task training. The acoustic sensing event neural network model is based on a pre-trained acoustic Swin Transformer neural network model, incorporating geographical and geological features as cue information. A cue mask network is constructed, and the cue mask vectors corresponding to the cue information are obtained by inputting into the cue mask network. Both the cue mask vectors and cue vectors are simultaneously input into the acoustic sensing event neural network model for model training and acoustic sensing event detection. This invention results in more accurate model training and more precise detection of acoustic event categories and spatiotemporal distribution localization.
[0009] This invention provides an acoustic sensing event detection method based on distributed buried optical fiber, comprising:
[0010] Spatiotemporal data of the phase of backscattered Rayleigh light from distributed buried optical fibers were collected.
[0011] Based on the input spatiotemporal data, acoustic sensing events are detected using an acoustic sensing event neural network model. The acoustic sensing event neural network model uses a cue learning method, and during task training, only the cue vector and the task-related network structure are adjusted.
[0012] Among them, the acoustic sensing event neural network model is a pre-trained acoustic Swin Transformer neural network model. Geographic and geological features are introduced as clue information. At the same time, a clue mask network is constructed. The clue mask network is input to obtain the clue mask vector corresponding to the clue information. The clue mask vector and the cue vector are simultaneously input into the acoustic sensing event neural network model for model training and acoustic sensing event detection.
[0013] Preferably, the input spatiotemporal data is fused with acoustic sensing event category and timestamp information to construct strong and weak labels for the spatiotemporal data. The weak labels have acoustic sensing event category labels, and the strong labels have both acoustic sensing event category and timestamp information labels. During the network structure adjustment process, strong and weak labels are introduced together for model optimization training.
[0014] Preferably, the spatiotemporal data with weak labels is data with acoustic sensing event category labels. For a spatiotemporal data segment, the event category label is represented by a one-dimensional vector, where i is the event category index in the one-dimensional vector. When the element at index i is 1, it indicates that there is an acoustic sensing event of the i-th event category in the entire spatiotemporal data segment. Conversely, when the element at index i is 0, it indicates that there is no acoustic event of the i-th event category in the entire data segment.
[0015] Preferably, the strongly labeled spatiotemporal data is spatiotemporal data with acoustic sensing event category and timestamp information labels. For a certain spatiotemporal data segment, the event category label and frame-level timestamp label are represented by a two-dimensional matrix, where j is the frame number index in the two-dimensional matrix and i is the event category index in the two-dimensional matrix. When the element at index (i,j) is 1, it indicates that there is an event of the i-th category in the j-th frame of the entire spatiotemporal data segment. Conversely, when the element is 0, it indicates that there is no event of the i-th category in the j-th frame of the entire spatiotemporal data segment.
[0016] Preferably, the spatiotemporal data are used to extract acoustic features using a log-Mel-time spectrum before being input into the acoustic sensing event neural network model.
[0017] Preferably, before the spatiotemporal data is input into the acoustic sensing event neural network model, acoustic features are extracted using the three-dimensional log-Mel-time spectrum, including the log-Mel-time spectrum, the first-order difference of the log-Mel-time spectrum coefficients, and the second-order difference of the log-Mel-time spectrum. The specific steps are as follows:
[0018] Acoustic time-domain signals are obtained from spatiotemporal data, and short-time Fourier transform is performed to obtain the amplitude spectrum of the short-time Fourier transform.
[0019] The obtained short-time Fourier amplitude spectrum is input into multiple Mel filters to obtain the Mel short-time Fourier transform amplitude spectrum.
[0020] Taking the logarithm of the obtained Mel-time Fourier transform amplitude spectrum yields the log-Melt-time spectrum, the specific process of which is expressed by the following formula:
[0021]
[0022] The first and second differences of the log-Mel-time spectrum are obtained by performing first and second differences on the log-Mel-time spectrum coefficients of the acoustic features.
[0023] in, These are the values of the Mel spectrum at frequency f and time t. It is the acoustic characteristic response value of the i-th Mel filter at frequency f. is the amplitude of the frequency domain signal corresponding to the i-th filter at time t. The summation operation represents the accumulation of the acoustic characteristic responses of all Mel filters, where n is the total number of Mel filters.
[0024] Preferably, the surrounding geographical information and geological and soil information are used as clue information. The clue information is input into a clue masking network to obtain the clue mask vector corresponding to the clue information. The clue masking network includes an embedding layer and a Transformer Decoder.
[0025] Preferably, the cue masking network includes an encoder, an embedding layer, three one-dimensional convolutional layers, a self-attention module, a cross-attention module, two feedforward neural networks, and a skip connection layer. The three one-dimensional convolutional layers are a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, and a third one-dimensional convolutional layer, respectively. The two feedforward neural networks are a first feedforward neural network and a second feedforward neural network, respectively. The input of the embedding layer includes the encoder output and geographic and geological features. The output of the embedding layer is connected to the first one-dimensional convolutional layer and the second one-dimensional convolutional layer, respectively. The first one-dimensional convolutional layer is sequentially connected to the self-attention module and the first feedforward neural network. The second one-dimensional convolutional layer is sequentially connected to the cross-attention module, the second feedforward neural network, and the third one-dimensional convolutional layer, respectively. The output of the first feedforward neural network is also connected to the cross-attention module. The output of the third one-dimensional convolutional layer is connected to the skip connection layer.
[0026] Preferably, inputting the clue information into the clue masking network to obtain the clue mask vector corresponding to the clue information includes the following steps:
[0027] First, the encoder output is represented as... and clues The inputs are fed into the embedding layer, where multiplication and query integration are performed to obtain the updated encoded representation. : ;
[0028] Then, the encoded representation is processed by the first one-dimensional convolutional layer and the second one-dimensional convolutional layer, respectively. and Projected onto decoder dimension N d The projection encoding representations are obtained respectively. and ;in R Represent real numbers, N d Indicates the encoder dimension. T Indicates the time dimension;
[0029] Secondly, the projection encoding representation and Input the self-attention module and the cross-attention module respectively, calculate the decoded representation, and obtain the target mask in the projection decoder space;
[0030] Next, the target mask is projected back to the encoder dimension using a third one-dimensional convolutional layer. N e ,get ;
[0031] Finally, skipping the connection layer is used to compute the final mask. : .
[0032] Preferably, the acoustic sensing event neural network model includes a pre-trained audio Swing Transformer model and a Token-Semantic CNN network, with the pre-trained audio Swing Transformer model and the Token-Semantic CNN network connected; the pre-trained audio Swing Transformer model includes multiple layers of Swing Transformer blocks.
[0033] Preferably, the acoustic sensing event neural network model utilizes a cue learning method, adjusting only the cue vector and the task-related network structure during task training. Specifically, this includes:
[0034] A set of learnable cue vectors is introduced into the input of each Swing Transformer block of the pre-trained audio Swing Transformer model. The Swing Transformer model formula for the cue vectors is expressed as:
[0035]
[0036] Where _ indicates that the output cue vector of the i-th layer is empty, E i P represents the feature vector of the i-th layer of the output. i-1L represents the cue vector of the (i-1)th layer. i This represents the Swing Transformer block at layer i, where N is the total number of layers in the Swing Transformer model, and i is the layer number of the Swing Transformer model.
[0037] Original input After passing through 4 layers of Swing Transformer Blocks, we obtain ;
[0038] Before inputting into the Token-Semantic CNN, Perform restoration Where T represents the time dimension, F represents the frequency dimension, and p represents the kernel size of the convolution in the encoder through which the acoustic features pass ( ). p p The patch-embed CNN encoding network, where C represents the size of the temporal window dimension in the acoustic signal;
[0039] By convolution kernel size A CNN network with padding size (1,0) is used to aggregate information in the time and frequency dimensions, and then interpolation is performed to obtain the output of the timestamp granularity distribution of events.
[0040] In this process, only the cue vectors and the task-related network structure are trained;
[0041] The Adam optimizer is used, and the loss function is jointly calculated using strong and weak labels to update the model. Two prediction results are output: acoustic sensing event category prediction and timestamp prediction. The loss of the two prediction results is calculated separately using the loss function.
[0042] The system outputs acoustic sensor event detection results. Based on the acoustic sensor event category prediction results and the threshold method, the event category is judged. When the event category prediction value is higher than the preset threshold, it indicates that the event category has occurred. Based on the timestamp prediction value, the system infers the time of the event activity of the category and maps it to a specific valid event.
[0043] Preferably, different loss functions are used to calculate the losses for the two prediction results separately;
[0044] For event category prediction, binary cross-entropy loss is used; for timestamp prediction, smoothed absolute error loss is used. The global loss function L(X) is a weighted sum of the two loss functions, as shown in the following formula:
[0045]
[0046] Where BCE is the binary cross-entropy loss and SmoothL1 loss is the smoothing absolute error loss. For event category predictions, For event category labels, For timestamp predictions, For timestamp tags, It's a hyperparameter.
[0047] The present invention also provides an acoustic sensing event detection device based on distributed buried optical fiber, comprising: a laser, an acousto-optic modulator, an erbium-doped fiber amplifier, a circulator, a phase demodulation unit, a photodetector, a spatiotemporal data acquisition card, and a distributed buried optical fiber acoustic sensing event detection and processing system.
[0048] The laser is sequentially connected to an acousto-optic modulator, an erbium-doped fiber amplifier, and a circulator; the circulator is connected to a distributed underground fiber optic cable; the circulator is also sequentially connected to a phase demodulation unit, a photodetector, a spatiotemporal data acquisition card, and a distributed underground fiber optic acoustic sensing event detection and processing system; the laser is also connected to the distributed underground fiber optic acoustic sensing event detection and processing system.
[0049] The spatiotemporal data acquisition card is used to acquire spatiotemporal data;
[0050] The distributed buried fiber acoustic sensing event detection and processing system utilizes the aforementioned distributed buried fiber acoustic sensing event detection method to detect acoustic sensing events based on the input spatiotemporal data and an acoustic sensing event neural network model.
[0051] Compared with the prior art, the present invention has at least the following beneficial effects:
[0052] (1) This invention uses the external geographical and geological environment of the buried optical fiber as clue information, and combines acoustic sensing event category prediction and event occurrence time prediction loss to assist in training the acoustic sensing event neural network model, so that the model can detect acoustic event categories and locate spatiotemporal distribution more accurately and obtain better detection performance.
[0053] (2) The present invention obtains the mask vector corresponding to the clue information through the clue mask network, which is used to guide the training of the acoustic sensing event neural network model, thereby making the model training more accurate and the detection performance more accurate;
[0054] (3) This invention only adjusts the cue vector and the task-related network structure, without adjusting the structure of the backbone network, which can effectively reduce the number of parameters for model training, improve model training time, and make the model easier to deploy and implement in a lightweight manner.
[0055] (4) The present invention uses distributed acoustic sensing, which can provide continuous, long-distance detection data. Attached Figure Description
[0056] Figure 1 This is a schematic diagram of an acoustic sensing event detection device based on distributed underground optical fiber, according to an embodiment of the present invention.
[0057] Figure 2 This is a schematic diagram of the acoustic feature extraction process according to an embodiment of the present invention;
[0058] Figure 3 This is a schematic diagram illustrating how a clue masking network obtains a clue information mask vector according to an embodiment of the present invention.
[0059] Figure 4 This is a detailed flowchart of an acoustic sensing event detection model according to an embodiment of the present invention, and a schematic diagram showing that the training process only adjusts the cue vector and the task-related network structure.
[0060] Figure 5 This is an embodiment of the acoustic sensing event detection method based on distributed buried optical fiber, which is an embodiment of the present invention.
[0061] In the figure, 111-laser, 112-acousto-optic modulator, 113-erbium-doped fiber amplifier, 114-circulator, 115-distributed buried fiber, 116-phase demodulation unit, 117-photodetector, 118-spatiotemporal data acquisition card, and 119-distributed buried fiber acoustic sensing event detection and processing system. Detailed Implementation
[0062] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0063] This invention provides an acoustic sensing event detection method based on distributed buried optical fiber, comprising:
[0064] Spatiotemporal data of the phase of backscattered Rayleigh light from distributed buried optical fibers were collected.
[0065] Based on the input spatiotemporal data, acoustic sensing events are detected using an acoustic sensing event neural network model. The acoustic sensing event neural network model uses a cue learning method, and during task training, only the cue vector and the task-related network structure are adjusted.
[0066] The acoustic sensing events in this invention are a generalized category of sensing events, which can be oil pipeline route monitoring (such as leakage, non-leakage, etc.) or monitoring events in other application fields;
[0067] Among them, the acoustic sensing event neural network model is a pre-trained acoustic Swin Transformer neural network model. Geographic and geological features are introduced as clue information. At the same time, a clue mask network is constructed. The clue mask network is input to obtain the clue mask vector corresponding to the clue information. The clue mask vector and the cue vector are simultaneously input into the acoustic sensing event neural network model for model training and acoustic sensing event detection.
[0068] According to a specific embodiment of the present invention, the input spatiotemporal data is fused with acoustic sensing event category and timestamp information to construct strong and weak labels for the spatiotemporal data. The weak label has an acoustic sensing event category label, and the strong label has both acoustic sensing event category and timestamp information labels. During the network structure adjustment process, the strong and weak labels are introduced together for model optimization training.
[0069] According to a specific embodiment of the present invention, the spatiotemporal data with weak labels is data with acoustic sensing event category labels. For a spatiotemporal data segment, the event category label is represented by a one-dimensional vector, where i is the event category index in the one-dimensional vector. When the element at index i is 1, it indicates that there is an acoustic sensing event of the i-th event category in the entire spatiotemporal data segment. Conversely, when the element at index i is 0, it indicates that there is no acoustic event of the i-th event category in the entire data segment.
[0070] According to a specific embodiment of the present invention, the strongly labeled spatiotemporal data is spatiotemporal data with acoustic sensing event category and timestamp information labels. For a certain spatiotemporal data segment, the event category label and frame-level timestamp label are represented by a two-dimensional matrix, where j is the frame number index in the two-dimensional matrix and i is the event category index in the two-dimensional matrix. When the element at index (i,j) is 1, it indicates that the j-th frame in the entire spatiotemporal data segment contains an event of the i-th category. Conversely, when the element is 0, it indicates that the j-th frame in the entire spatiotemporal data segment does not contain an event of the i-th category.
[0071] According to one specific embodiment of the present invention, spatiotemporal data is used to extract acoustic features using log-Mel-time spectra before being input into the acoustic sensing event neural network model.
[0072] According to a specific embodiment of the present invention, before the spatiotemporal data is input into the acoustic sensing event neural network model, acoustic features are extracted using a three-dimensional log-Mehr-time spectrum, comprising the log-Mehr-time spectrum, the first-order difference of the log-Mehr-time spectrum coefficients, and the second-order difference of the log-Mehr-time spectrum. The specific steps are as follows:
[0073] Acoustic time-domain signals are obtained from spatiotemporal data, and short-time Fourier transform is performed to obtain the amplitude spectrum of the short-time Fourier transform.
[0074] The obtained short-time Fourier amplitude spectrum is input into multiple Mel filters to obtain the Mel short-time Fourier transform amplitude spectrum.
[0075] Taking the logarithm of the obtained Mel-time Fourier transform amplitude spectrum yields the log-Melt-time spectrum, the specific process of which is expressed by the following formula:
[0076]
[0077] The first and second differences of the log-Mel-time spectrum are obtained by performing first and second differences on the log-Mel-time spectrum coefficients of the acoustic features.
[0078] in, These are the values of the Mel spectrum at frequency f and time t. It is the acoustic characteristic response value of the i-th Mel filter at frequency f. is the amplitude of the frequency domain signal corresponding to the i-th filter at time t. The summation operation represents the accumulation of the acoustic characteristic responses of all Mel filters, where n is the total number of Mel filters.
[0079] According to a specific embodiment of the present invention, the geographical information and geological and soil information of the surrounding environment are used as clue information. The clue information is input into a clue masking network to obtain the clue mask vector corresponding to the clue information. The clue masking network includes an embedding layer and a Transformer Decoder.
[0080] According to a specific embodiment of the present invention, the cue masking network includes an encoder, an embedding layer, three one-dimensional convolutional layers, a self-attention module, a cross-attention module, two feedforward neural networks, and a skip connection layer. The three one-dimensional convolutional layers are a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, and a third one-dimensional convolutional layer, respectively. The two feedforward neural networks are a first feedforward neural network and a second feedforward neural network, respectively. The input of the embedding layer includes the encoder output and geographic and geological features. The output of the embedding layer is connected to the first one-dimensional convolutional layer and the second one-dimensional convolutional layer, respectively. The first one-dimensional convolutional layer is sequentially connected to the self-attention module and the first feedforward neural network. The second one-dimensional convolutional layer is sequentially connected to the cross-attention module, the second feedforward neural network, and the third one-dimensional convolutional layer, respectively. The output of the first feedforward neural network is also connected to the cross-attention module. The output of the third one-dimensional convolutional layer is connected to the skip connection layer.
[0081] According to a specific embodiment of the present invention, inputting clue information into a clue masking network to obtain the mask vector corresponding to the clue information includes the following steps:
[0082] First, the encoded output of the encoder is represented as and clues The inputs are fed into the embedding layer, where multiplication and query integration are performed (i.e., multiplication is performed first, followed by a table lookup based on the multiplication), resulting in the encoded representation. : ;
[0083] Then, the encoded representation is processed by the first one-dimensional convolutional layer and the second one-dimensional convolutional layer, respectively. and Projected onto decoder dimension ≤ The projection encoding representations are obtained respectively. and ;in R Represent real numbers, N d Indicates the encoder dimension. T Indicates the time dimension;
[0084] Secondly, the projection encoding representation and Input the self-attention module and the cross-attention module respectively, calculate the decoded representation, and obtain the target mask in the projection decoder space;
[0085] Next, the target mask is projected back to the encoder dimension using a third one-dimensional convolutional layer, resulting in... ;
[0086] Finally, skipping the connection layer is used to compute the final mask. : .
[0087] According to a specific embodiment of the present invention, the acoustic sensing event neural network model includes a pre-trained audio Swing Transformer model and a Token-Semantic CNN network, and the pre-trained audio Swing Transformer model and the Token-Semantic CNN network are connected; the pre-trained audio Swing Transformer model includes multiple layers of Swing Transformer blocks.
[0088] According to a specific embodiment of the present invention, the acoustic sensing event neural network model utilizes a cue learning method, which, during task training, adjusts only the cue vector and the task-related network structure, specifically including:
[0089] A set of learnable cue vectors is introduced into the input of each Swing Transformer block of the pre-trained audio Swing Transformer model. The Swing Transformer model formula for the cue vectors is expressed as:
[0090]
[0091] Here, _ indicates that the cue vector of the i-th layer is empty. The underscore indicates that the output only contains the feature vector of the i-th layer, but not the cue vector of the i-th layer. That is, the cue vector is fine-tunable and needs to be trained. The feature vector is the output of the previous Swin Transformer; E i P represents the feature vector of the i-th layer of the output. i-1 L represents the cue vector of the (i-1)th layer. i This represents the Swing Transformer block at layer i, where N is the total number of layers in the Swing Transformer model, and i is the layer number of the Swing Transformer model.
[0092] Original input After passing through 4 layers of Swing Transformer Blocks, we obtain ;
[0093] Before inputting into the Token-Semantic CNN, Perform restoration Where T represents the time dimension, F represents the frequency dimension, and p represents the kernel size of the convolution in the encoder through which the acoustic features pass ( ). p p The patch-embed CNN encoding network, where C represents the size of the temporal window dimension in the acoustic signal;
[0094] By convolution kernel size A CNN network with padding size (1,0) is used to aggregate information in the time and frequency dimensions, and then interpolation is performed to obtain the output of the timestamp granularity distribution of events.
[0095] In this process, only the cue vectors and the task-related network structure are trained;
[0096] The Adam optimizer is used, and the loss function is jointly calculated using strong and weak labels to update the model. Two prediction results are output: acoustic sensing event category prediction and timestamp prediction. The loss of the two prediction results is calculated separately using the loss function.
[0097] The system outputs acoustic sensor event detection results. Based on the acoustic sensor event category prediction results and the threshold method, the event category is judged. When the event category prediction value is higher than the preset threshold, it indicates that the event category has occurred. Based on the timestamp prediction value, the system infers the time of the event activity of the category and maps it to a specific valid event.
[0098] According to a specific embodiment of the present invention, different loss functions are used to calculate the losses of the two prediction results separately;
[0099] For event category prediction, binary cross-entropy loss is used; for timestamp prediction, smoothed absolute error loss is used. The global loss function L(X) is a weighted sum of the two loss functions, as shown in the following formula:
[0100]
[0101] Where BCE is the binary cross-entropy loss and SmoothL1 loss is the smoothing absolute error loss. For event category predictions, For event category labels, For timestamp predictions, For timestamp tags, It's a hyperparameter.
[0102] The present invention also provides an acoustic sensing event detection device based on distributed buried optical fiber, comprising: a laser 111, an acousto-optic modulator 112, an erbium-doped fiber amplifier 113, a circulator 114, a phase demodulation unit 116, a photodetector 117, a spatiotemporal data acquisition card 118, and a distributed buried optical fiber acoustic sensing event detection and processing system 119.
[0103] The laser 111 is sequentially connected to the acousto-optic modulator 112, the erbium-doped fiber amplifier 113, and the circulator 114; the circulator 114 is connected to the distributed buried fiber optic cable 115; the circulator 114 is also sequentially connected to the phase demodulation unit 116, the photodetector 117, the spatiotemporal data acquisition card 118, and the distributed buried fiber optic acoustic sensing event detection and processing system 119; the laser 111 is also connected to the distributed buried fiber optic acoustic sensing event detection and processing system 119.
[0104] The spatiotemporal data acquisition card 118 is used to acquire spatiotemporal data;
[0105] The distributed buried fiber acoustic sensing event detection and processing system 119 utilizes the aforementioned distributed buried fiber acoustic sensing event detection method to detect acoustic sensing events based on the input spatiotemporal data and an acoustic sensing event neural network model.
[0106] The optical input in circulator 114 all originates from erbium-doped fiber amplifier 113. This portion of light propagates forward in the fiber and also undergoes backscattering. (Therefore, the distributed buried fiber optic cable 115 is not an input to circulator 114, but rather an optical transmission device). Thus, the output is actually a superimposed optical signal received at the time of reception. Spatiotemporal data acquisition card 118 analyzes the characteristics of the backscattered portion from the received optical signal to perform acoustic sensing event detection.
[0107] The laser 111 outputs a clean light signal, while the spatiotemporal data acquisition card 118 collects a mixed signal superimposed with backscattered light. During the training process, the distributed buried fiber optic acoustic sensing event detection and processing system 119 uses both of these signals to perform training.
[0108] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above description or related technical or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. A method for acoustic sensing event detection based on distributed buried optical fiber, characterized in that, include: Spatiotemporal data of the phase of backscattered Rayleigh light from distributed buried optical fibers were collected. Based on the input spatiotemporal data, acoustic sensing events are detected using an acoustic sensing event neural network model. The acoustic sensing event neural network model uses a cue learning method, and during task training, only the cue vector and the task-related network structure are adjusted. Among them, the acoustic sensing event neural network model is a pre-trained acoustic Swin Transformer neural network model. Geographic and geological features are introduced as clue information. At the same time, a clue mask network is constructed. The clue mask network is input to obtain the clue mask vector corresponding to the clue information. The clue mask vector and the cue vector are simultaneously input into the acoustic sensing event neural network model for model training and acoustic sensing event detection. The input spatiotemporal data is fused with acoustic sensing event category and timestamp information to construct strong and weak labels for the spatiotemporal data. The weak labels have acoustic sensing event category labels, and the strong labels have both acoustic sensing event category and timestamp information labels. During the network structure adjustment process, strong and weak labels are introduced together for model optimization training.
2. The acoustic sensing event detection method based on distributed buried optical fiber according to claim 1, characterized in that, The spatiotemporal data are used to extract acoustic features using a log-Mel-time spectrum before being input into the acoustic sensing event neural network model.
3. The acoustic sensing event detection method based on distributed buried optical fiber according to claim 2, characterized in that, Before being input into the acoustic sensing event neural network model, the spatiotemporal data are used to extract acoustic features using the log-Melbourne spectrum, the first difference of the log-Melbourne spectrum coefficients, the second difference of the log-Melbourne spectrum, and the three-dimensional log-Melbourne spectrum. The specific steps are as follows: Acoustic time-domain signals are obtained from spatiotemporal data, and short-time Fourier transform is performed to obtain the amplitude spectrum of the short-time Fourier transform. The obtained short-time Fourier amplitude spectrum is input into multiple Mel filters to obtain the Mel short-time Fourier transform amplitude spectrum. Taking the logarithm of the obtained Mel-time Fourier transform amplitude spectrum yields the log-Melt-time spectrum, the specific process of which is expressed by the following formula: The first and second differences of the log-Mel-time spectrum are obtained by performing first and second differences on the log-Mel-time spectrum coefficients of the acoustic features. in, These are the values of the Mel spectrum at frequency f and time t. It is the acoustic characteristic response value of the i-th Mel filter at frequency f. is the amplitude of the frequency domain signal corresponding to the i-th filter at time t. The summation operation represents the accumulation of the acoustic characteristic responses of all Mel filters, where n is the total number of Mel filters.
4. The acoustic sensing event detection method based on distributed buried optical fiber according to claim 1, characterized in that, Using the geographical and geological information of the surrounding environment as clue information, the clue information is input into the clue masking network to obtain the clue mask vector corresponding to the clue information. The clue masking network includes an embedding layer and a TransformerDecoder.
5. The acoustic sensing event detection method based on distributed buried optical fiber according to claim 4, characterized in that, The cue masking network includes an encoder, an embedding layer, three one-dimensional convolutional layers, a self-attention module, a cross-attention module, two feedforward neural networks, and a skip connection layer. The three one-dimensional convolutional layers are a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, and a third one-dimensional convolutional layer. The two feedforward neural networks are a first feedforward neural network and a second feedforward neural network. The input of the embedding layer includes the encoder output and geographic and geological features. The output of the embedding layer is connected to the first one-dimensional convolutional layer and the second one-dimensional convolutional layer. The first one-dimensional convolutional layer is sequentially connected to the self-attention module and the first feedforward neural network. The second one-dimensional convolutional layer is sequentially connected to the cross-attention module, the second feedforward neural network, and the third one-dimensional convolutional layer. The output of the first feedforward neural network is also connected to the cross-attention module. The output of the third one-dimensional convolutional layer is connected to the skip connection layer.
6. The acoustic sensing event detection method based on distributed buried optical fiber according to claim 1, characterized in that, The acoustic sensing event neural network model includes a pre-trained audio Swing Transformer model and a Token-Semantic CNN network, with connections between the pre-trained audio Swing Transformer model and the Token-Semantic CNN network; the pre-trained audio Swing Transformer model includes multi-layer Swing Transformer blocks.
7. The acoustic sensing event detection method based on distributed buried optical fiber according to claim 6, characterized in that, The acoustic sensing event neural network model utilizes a cue learning method, adjusting only the cue vectors and task-related network structures during task training. Specifically, this includes: A set of learnable cue vectors is introduced into the input of each Swing Transformer block of the pre-trained audio Swing Transformer model. The Swing Transformer model formula for the cue vectors is expressed as: Where _ indicates that the output cue vector of the i-th layer is empty, E i P represents the feature vector of the i-th layer of the output. i-1 L represents the cue vector of the (i-1)th layer. i This represents the Swing Transformer block at layer i, where N is the total number of layers in the Swing Transformer model, and i is the layer number of the Swing Transformer model. Original input After passing through 4 layers of Swing Transformer Blocks, we obtain ; Before inputting into the Token-Semantic CNN, Perform restoration Where T represents the time dimension, F represents the frequency dimension, and p represents the kernel size of the convolution in the encoder through which the acoustic features pass ( ). p p The patch-embed CNN encoding network, where C represents the size of the temporal window dimension in the acoustic signal; By convolution kernel size ( A CNN network with padding size (1,0) is used to aggregate information in the time and frequency dimensions, and then interpolation is performed to obtain the output of the timestamp granularity distribution of events. In this process, only the cue vectors and the task-related network structure are trained; The Adam optimizer is used, and the loss function is jointly calculated using strong and weak labels to update the model. Two prediction results are output: acoustic sensing event category prediction and timestamp prediction. The loss of the two prediction results is calculated separately using the loss function. The system outputs acoustic sensor event detection results. Based on the acoustic sensor event category prediction results and the threshold method, the event category is judged. When the event category prediction value is higher than the preset threshold, it indicates that the event category has occurred. Based on the timestamp prediction value, the activity time of the event of the event category is inferred and mapped to a specific valid event.
8. The acoustic sensing event detection method based on distributed buried optical fiber according to claim 7, characterized in that, Different loss functions are used to calculate the loss for the two prediction results separately; For event category prediction, binary cross-entropy loss is used; for timestamp prediction, smoothed absolute error loss is used. The global loss function L(X) is a weighted sum of the two loss functions, as shown in the following formula: Where BCE is the binary cross-entropy loss and SmoothL1 loss is the smoothing absolute error loss. For event category predictions, For event category labels, For timestamp predictions, For timestamp tags, It's a hyperparameter.
9. An acoustic sensing event detection device based on distributed buried optical fiber, characterized in that, include: Lasers, acousto-optic modulators, erbium-doped fiber amplifiers, circulators, phase demodulation units, photodetectors, spatiotemporal data acquisition cards, and distributed buried fiber optic acoustic sensing event detection and processing systems; The laser is sequentially connected to an acousto-optic modulator, an erbium-doped fiber amplifier, and a circulator; the circulator is connected to a distributed underground fiber optic cable; the circulator is also sequentially connected to a phase demodulation unit, a photodetector, a spatiotemporal data acquisition card, and a distributed underground fiber optic acoustic sensing event detection and processing system; the laser is also connected to the distributed underground fiber optic acoustic sensing event detection and processing system. The spatiotemporal data acquisition card is used to acquire spatiotemporal data; The distributed buried optical fiber acoustic sensing event detection and processing system utilizes the acoustic sensing event detection method based on distributed buried optical fiber as described in any one of claims 1-8, and detects acoustic sensing events using an acoustic sensing event neural network model based on the input spatiotemporal data.
Citation Information
Patent Citations
Distributed optical fiber sensing system
CN113295259A
Expressway abnormal event detection method and system
CN114139583A
Pipeline safety monitoring method based on distributed optical fiber sensor and deep learning
CN114857504A
Distributed sound wave sensing voice information recognition system and method
CN115622626A
Intelligent optical fiber distributed acoustic wave sensing system based on AI chip and method
CN110487391A