Distributed optical fiber sound wave sensing data feature extraction method and device
The Transformer network model pre-trained by mask reconstruction task is used to improve the encoding and decoding network using the attention mechanism, which solves the problem of distributed fiber sensor dependence on label data, and realizes efficient waterfall map feature extraction and good generalization capabilities.
Patent Information
- Application Number
- CN202510333900.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
Existing distributed fiber sensors need to rely on a large amount of labeled data when extracting feature, resulting in low training efficiency and insufficient generalization capabilities, and the inability to effectively extract high-quality waterfall chart features.
The Transformer network model pre-trained by mask reconstruction task is used to improve the encoding and decoding network using the attention mechanism, and pre-training is performed by label-free data to extract waterfall chart features and reduce dependence on labels.
The training efficiency of the model and the generalization ability of feature extraction are improved, and it can migrate and apply in different tasks to extract high-quality waterfall chart features.
Smart Images

Figure CN120256936A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed optical fiber sensing, and in particular to a method and device for extracting distributed optical fiber acoustic wave sensing data features. Background Art
[0002] Distributed fiber optic sensing is a fiber optic sensing technology used to collect vibration signals. Compared with electrical vibration sensors, it has the advantages of high resolution, small size, light weight, no-electric operation and adaptability to harsh environments. It is widely used in earthquake monitoring, oil and gas pipeline monitoring, railway monitoring and other fields. However, the signal (waterfall chart) generated by this type of sensor is a two-dimensional signal, which contains a time dimension and a space dimension. Since distributed fiber optic sensors are similar to microphone arrays, the data features they generate are more similar to one-dimensional audio signals rather than easy-to-understand two-dimensional image features. Therefore, waterfall chart data is more abstract to humans, which poses a challenge to the subsequent design of feature extraction algorithms. In addition, distributed fiber optic sensors generate a large amount of data (due to their relatively high sampling rate and continuous measurement spatial resolution). The huge amount of data makes it more difficult for people to understand their data features, further exacerbating the difficulty of designing feature extraction algorithms. More and more researchers have paid attention to the characteristics of waterfall charts and began to use deep learning technology to extract waterfall chart features. According to the degree of annotation of the original waterfall chart, feature learning algorithms based on deep learning can be divided into three categories: supervised learning (all training data are annotated), semi-supervised learning (partial training data are annotated) and unsupervised learning (training data does not need to be annotated). Supervised learning methods all use labeled waterfall charts to learn their features. Since the data labeling process is often time-consuming and laborious, the labeled data sets are usually limited in amount and the data is carefully selected, and the features of the waterfall charts are usually incomplete or fragmented. This situation may make the model unable to capture the full picture of the data during the learning process, resulting in insufficient generalization ability. In addition, relying on limited labeled data, the model may be affected by task-dependent biases, resulting in poor performance when processing unseen data. The problem of limited labeled data is solved by introducing additional unlabeled data through semi-supervised learning methods. Although the semi-supervised method improves the performance of the supervised learning model, the training process is still heavily dependent on labeled data and specific tasks, resulting in the impact of task-dependent bias in the process of learning waterfall chart features. The unsupervised learning method is a training method based entirely on unlabeled data, which helps to learn comprehensive and task-independent data features. In the prior art, a small-scale unsupervised convolutional network UNet is used for feature extraction. Although label training is avoided, the network of this method has insufficient generalization ability for different tasks and cannot obtain good features; or, a transformer architecture is used for graph feature extraction, such as the Chinese patent application "CN119360036A", which provides a method for multi-view feature extraction using a lightweight Transformer-CNN fusion network, extracting global and local features respectively and performing feature fusion, which improves the accuracy of feature recognition and pre-trains the model according to task requirements to improve the generalization of the model, but it still needs to rely on labels during pre-training.
[0003] Therefore, it is a technical problem to be solved to provide a method that can both extract high-quality waterfall chart features and reduce the dependence of model training on labels. Summary of the Invention
[0004] The purpose of the present invention is to overcome the disadvantages of the existing technology, that is, when extracting waterfall chart features, the model pre-training used depends on a large amount of labeled data and the quality of the extracted features is unstable, and to provide a distributed optical fiber acoustic wave sensing data feature extraction method and device.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] According to the first aspect of the present invention, a distributed optical fiber acoustic wave sensing data feature extraction method is provided. The method uses a Transformer network model pre-trained by a masked reconstruction task to perform distributed optical fiber acoustic wave sensing data feature extraction. The Transformer network model is improved by using an attention mechanism and includes an encoding network and a decoding network with the same structure. Both the encoding network and the decoding network include a linear layer, a position encoding layer, and a multi-level cascaded improved Transformer encoder. The pre-training includes:
[0007] Obtain a pre-training data set, cut each waterfall chart in the pre-training data set into multiple waterfall chart blocks of the same size, and the shape of the waterfall chart block is: space × time;
[0008] Perform independent masking processing on each cut waterfall chart according to a uniform distribution. The waterfall chart blocks not masked are called visible waterfall chart blocks;
[0009] Stack the visible waterfall chart blocks of the same waterfall chart to obtain a first 3D tensor with a shape of: the number of waterfall chart blocks × space × time; perform dimension processing on the first 3D tensor and then use the Transformer network to generate a reconstructed waterfall chart, and calculate the difference between the reconstructed waterfall chart and the corresponding original waterfall chart. Based on the difference, construct a loss function for pre-training.
[0010] As a preferred technical solution, the method of dimension processing is: flatten the spatial dimension and the time dimension of the first 3D tensor to obtain a 2D tensor with a shape of: the number of waterfall chart blocks × the number of waterfall chart sampling points; add a dimension to the 2D tensor to expand it into a second 3D tensor with a shape of: the number of waterfall chart blocks × the number of waterfall chart sampling points × the processing dimension.
[0011] As a preferred technical solution, the encoding network performs the following operations:
[0012] After expanding the processing dimension of the second 3D tensor using the linear layer described above, trainable standard sine-cosine positional information is added using a positional encoding layer to obtain first intermediate data;
[0013] Based on the first intermediate data, it is processed using an improved Transformer encoder. The improved Transformer encoder includes multiple cascaded networks, and each network layer includes a multi-head attention mechanism. Each head of the attention mechanism outputs a second intermediate data. All the intermediate data are concatenated in the processing dimension to obtain an output. Among them, the input of the first-layer network is the first intermediate data, the input of the Nth-layer network is the output of the (N - 1)th-layer network, and the output of the last-layer network is the third intermediate data;
[0014] The dimension of the third intermediate data is processed using a linear layer to obtain first waterfall plot data with the same shape and size as the first intermediate data.
[0015] As a preferred technical solution, the following operations are performed in each head of the attention mechanism to obtain the second intermediate data, including:
[0016] Calculate a query matrix, a key matrix, and a value matrix based on the first intermediate data;
[0017] After performing matrix multiplication on the query matrix and the key matrix, perform normalization processing to obtain a weight matrix;
[0018] Multiply the weight matrix and the value matrix to obtain the intermediate data.
[0019] As a preferred technical solution, the first intermediate data, the output of each network layer, and the second intermediate data have the same shape and size; the number of waterfall plot blocks and the number of waterfall plot sampling points of the third intermediate data are the same as those of the second intermediate data, its processing dimension is an integer multiple of the second intermediate data, and this integer multiple is the number of heads of the multi-head attention mechanism.
[0020] As a preferred technical solution, the decoding network performs the following steps:
[0021] After expanding the processing dimension of the first waterfall plot data using a linear layer, it is restored using a trainable 3D tensor to obtain first data;
[0022] Add trainable standard sine-cosine positional information to the first data using a positional encoding layer to obtain second data;
[0023] Process the second data based on the improved transformer encoder, where the improved transformer encoder includes a multi-level cascaded network, and each layer of the network includes a multi-head attention mechanism. Each head of the attention mechanism outputs a third data, and all the third data are concatenated in the processing dimension to obtain an output. The input of the first layer network is the second data, the input of the Nth layer network is the output of the (N - 1)th layer network, and the output of the last layer network is the fourth data.
[0024] Use a linear layer to process the dimension of the fourth data to make its shape and size the same as the second data, and map the processing dimension of the dimension-processed fourth data to 1 to obtain a reconstructed waterfall diagram.
[0025] As a preferred technical solution, the number of sampling points and the processing dimension of the waterfall diagram of the trainable 3D tensor are the same as those of the first data after the processing dimension is expanded, and the number of its waterfall diagrams is 1.
[0026] As a preferred technical solution, the expression of the loss function is:
[0027]
[0028] where represents the missing waterfall diagram block after masking the ith waterfall diagram; represents the reconstruction of the missing waterfall diagram block of the ith waterfall diagram; ‖·‖2 represents the calculation of the second norm; Mean(·) represents the average over the entire training data set.
[0029] As a preferred technical solution, the method includes:
[0030] Cut the target waterfall diagram into multiple waterfall diagram blocks of the same size, and the shape of each waterfall diagram block is space × time;
[0031] Stack and flatten the multiple waterfall diagram blocks into a 3D tensor;
[0032] Input the 3D tensor into the encoding network in the pre-trained Transformer network model to extract the distributed fiber optic acoustic sensing data features.
[0033] According to the second aspect of the present invention, there is provided a distributed fiber optic acoustic sensing data feature extraction device, including a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the above method is implemented.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1) The present invention pre - trains the improved Transformer network model according to the difficulty of the task using the mask reconstruction task. This not only avoids simply learning a simple interpolation reconstruction algorithm, forces the model to learn valuable features from limited information, but also can reduce the data volume of the waterfall plot while retaining important features, improving the training efficiency of the model. Moreover, the entire pre - training process does not use data labels for training. Therefore, the extracted features have better generalization compared to the features extracted by other dedicated networks and can be transferred to various different tasks, including but not limited to classification, data denoising, waterfall plot phase unfolding, etc.
[0036] 2) The present invention improves the traditional Transformer structure using the attention mechanism for the 2D data mode of the waterfall plot. After cutting the waterfall plot into multiple time series of the same length, the attention mechanism is used to analyze the similarity of its time series to obtain accurate spatio - temporal information. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of the pre - training of the Transformer network model of the present invention;
[0038] Figure 2 It is a flowchart of the encoding network of the present invention performing tasks;
[0039] Figure 3 It is a flowchart of the decoding network of the present invention performing tasks;
[0040] Figure 4 It is a flowchart of the feature extraction of the present invention;
[0041] Figure 5 It is a schematic diagram of the results of the comparative experiment in Example 2 of the present invention;
[0042] Figure 6 It is a schematic diagram of the confusion matrix of the Transformer network model of the present invention;
[0043] Figure 7 It is a schematic diagram of the confusion matrix of the initial model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Unless otherwise defined, technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one", "the" and the like involved in this application do not indicate a limitation in quantity and may represent a singular or plural number. The terms "comprising", "including", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include steps or units not listed, or may further include other steps or units inherent to these processes, methods, products or devices. The words such as "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates an "or" relationship between the associated objects before and after. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0046] Example 1
[0047] In recent years, with the rapid rise and development of large models in the field of artificial intelligence, networks using self-supervised and Transformer architectures have become mainstream, replacing the original neural network processing methods in various signal processing fields and achieving excellent performance. For example, the Generative Pre-trained Transformer (GPT) well-known in the field of natural language processing and the Masked Autoencoder (MAE) in the field of image processing. Transformers trained in this self-supervised mode have extremely strong feature extraction and generalization capabilities. However, currently, in the data processing of distributed fiber optic sensing, there is no Transformer model suitable for self-supervised training of waterfall diagrams. Therefore, the present invention extends the training method of MAE to the data processing of waterfall diagrams, and uses a Transformer improved by an attention mechanism as the basic model architecture to train a brand-new feature extraction model, enabling the model to learn higher-quality waterfall diagram features and greatly reducing the data volume of waterfall diagrams during the training process, thus accelerating the training process. Specifically, the Transformer network model adopted by the present invention is improved using an attention mechanism, including an encoding network and a decoding network with the same structure. Both the encoding network and the decoding network include a linear layer, a position encoding layer, and multiple cascaded improved Transformer encoders.
[0048] Specifically, the pre-training process of the Transformer network model is as Figure 1 shown, including:
[0049] S1. Obtain a pre-training dataset, and cut each waterfall diagram in the pre-training dataset into multiple waterfall diagram blocks of the same size, and the shape of the waterfall diagram block is: space × time.
[0050] S11. Use an existing DAS device to lay out optical fibers to collect waterfall diagram data, screen the collected waterfall diagrams, retain the waterfall diagrams where there are events in the time dimension, delete the waterfall diagrams where there are no events in a period of time, cut the processed waterfall diagrams into the same shape and size, such as space × time = 12 × 10000, and then normalize the data to obtain the unlabeled data for pre-training.
[0051] Or, directly use the existing publicly available distributed vibration sensor waterfall diagram dataset (its source is: Cao X, Su Y, Jin Z, et al. An open dataset of events with two classification models as baselines. Results in Optics, 2023, 10: 100372.) As the training dataset, each data in this dataset is normalized so that the shape and size of each data in the dataset is space × time = 12 × 10000.
[0052] S12. Refer to Figure 1 In the part of [[ ]], each waterfall chart in the dataset is cut into waterfall chart blocks of the required size for the task. In this embodiment, the set cutting size is 1 × 624, and the redundant data points at both ends are discarded, and a total of 192 waterfall chart blocks of the same size are obtained.
[0053] S2. Perform independent masking processing on each cut waterfall chart according to a uniform distribution. The waterfall chart blocks not masked are called visible waterfall chart blocks.
[0054] Specifically, only 50% of the waterfall chart blocks are retained, which are called visible waterfall chart blocks, a total of 96; 50% of the waterfall chart blocks are discarded, which are called masked waterfall chart blocks, a total of 96. And the masking process of each waterfall chart data is independent of each other, without interference, and is completely carried out according to a uniform distribution.
[0055] S3. Stack the visible waterfall chart blocks of the same waterfall chart to obtain the first 3D tensor. After dimension processing of the first 3D tensor, use the Transformer network to generate a reconstructed waterfall chart, and calculate the difference between the reconstructed waterfall chart and the corresponding original waterfall chart. Based on the difference, construct a loss function for pre-training.
[0056] S31. 3D tensor dimension processing.
[0057] S311. Stack the 96 waterfall chart blocks obtained in step S2 to obtain the first 3D tensor, and its shape and size are: the number of waterfall chart blocks × space × time = 96 × 1 × 624.
[0058] S312. Flatten the space dimension and time dimension of the first 3D tensor to obtain a 2D tensor, and its shape is: the number of waterfall chart blocks × the number of waterfall chart sampling points = 96 × 624.
[0059] S313. Add a dimension to the 2D tensor to expand it into the second 3D tensor, and its shape is: the number of waterfall chart blocks × the number of waterfall chart sampling points × the processing dimension = 96 × 624 × 1.
[0060] S32. Encoding network feature extraction.
[0061] Input the second 3D tensor obtained in step S31 into the encoding network of the Transformer network model for feature extraction, and its process is as Figure 2As shown, it includes:
[0062] S321: Use a linear layer to expand the processing dimension of the second 3D tensor by 576 from 1, and the size of the output data is 96×624×576. Then use a positional encoding layer to add trainable standard sine and cosine positional information to obtain the first intermediate data.
[0063] S322: Based on the first intermediate data, use an improved transformer encoder for processing, analyze the similarity between visible waterfall plot blocks, and extract the corresponding waterfall plot features.
[0064] The improved transformer encoder includes multiple cascaded networks and each layer of the network includes a multi-head attention mechanism; in each layer of the improved transformer encoder, each head of the attention mechanism outputs a second intermediate data, and all the second intermediate data are concatenated in the said processing dimension to obtain the output.
[0065] Among them, the input of the first layer of the improved transformer encoder is the first intermediate data, the input of the Nth layer of the improved transformer encoder is the output of the (N - 1)th layer of the improved transformer encoder, and the output of the last layer of the improved transformer encoder is the third intermediate data.
[0066] The first intermediate data, the output of each layer of the network, and the second intermediate data have the same shape and size; the number of waterfall plot blocks and the number of waterfall plot sampling points of the third intermediate data are the same as those of the second intermediate data, its processing dimension is an integer multiple of the second intermediate data, and this integer multiple is the number of heads of the multi-head attention mechanism. Specifically, the method for obtaining the second intermediate data includes:
[0067] i. Calculate a query matrix, a key matrix, and a value matrix based on the first intermediate data.
[0068] ii. After performing matrix multiplication on the query matrix and the key matrix, perform normalization processing to obtain a weight matrix.
[0069] iii. Multiply the weight matrix and the value matrix to obtain the intermediate data.
[0070] In this embodiment, it is set that there are 6 completely identical layers of networks, and taking the eight-head attention mechanism as an example, where the attention mechanism is a self-attention mechanism, the size of the second intermediate data output by each self-attention mechanism is 96×624×576. The size of the output obtained by concatenating the 8 second intermediate data according to the processing dimension is 96×624×(576×8), that is, 96×624×4608.
[0071] S323: Use a linear layer to perform dimensional processing on the third intermediate data for the processing dimension, and generate the first waterfall plot data with the same shape and size as the first intermediate data, and its size is 96×624×576.
[0072] S33. Decode network feature extraction.
[0073] The decoding network has the same structure as the encoding network, and its processing process is also similar. However, before adding the standard sine-cosine position information, the output of the linear layer needs to be restored. The process is as Figure 3 shown, including:
[0074] S331. After expanding the processing dimension of the first waterfall chart data from 576 to 384 using the linear layer, a 3D tensor with the same size as the visible waterfall chart block is used for restoration. In this embodiment, the size of the filled 3D tensor is set to 1×624×384. This 3D tensor is not the masked waterfall chart block. It only indicates that a waterfall chart block is discarded at this position during masking. This operation can restore the size of the first data output by the linear mapping to the same 192×624×384 as the original waterfall chart.
[0075] S332. Add trainable standard sine-cosine position information to the first data using the position encoding layer to obtain the second data.
[0076] S333. Process the second data using the improved transformer encoder to obtain the fourth data, whose size is 192×624×384. The processing process is referred to step S322 and will not be elaborated here.
[0077] S334. Use the linear layer to map the processing dimension of the fourth data from 384 to 1 to obtain the reconstructed waterfall chart, whose size is 192×624×1.
[0078] S34. Difference calculation.
[0079] Calculate the difference between the reconstructed waterfall chart and the original waterfall chart, and train the network based on this difference. In this embodiment, the mean squared error (MSE) is selected as the reconstruction loss function. The expression is:
[0080]
[0081] where, represents the missing waterfall chart block after compression of the i-th waterfall chart; represents the reconstruction of the missing waterfall chart block of the i-th waterfall chart; ‖·‖2 represents the calculation of the second norm; Mean(·) represents the average over the entire training data set. Moreover, when evaluating the difference between the reconstructed waterfall chart and the original waterfall chart, only the reconstruction loss of the missing waterfall chart block is considered, and the influence on the visible waterfall chart block during the model operation process is not considered. Any gradient descent algorithm and its iterative optimization method can be selected as the optimization algorithm.
[0082] Use the Transformer network model trained through steps S1 to S3 to extract features from the collected distributed fiber optic acoustic sensing data. The process is as Figure 4 shown, including:
[0083] A1. Cut the target waterfall diagram according to the pre-trained cutting method, and the shape of each waterfall diagram block is space × time, and its size is 1 × 624 in this embodiment.
[0084] A2. Stack and flatten multiple waterfall diagram blocks into a 3D tensor, and the size of this 3D tensor is related to the task requirements.
[0085] A3. Input the 3D tensor into the encoding network in the pre-trained Transformer network model to extract the features of the distributed fiber optic acoustic sensing data. In this process, the feature size can be set according to the task requirements.
[0086] This embodiment also provides a device for extracting features of distributed fiber optic acoustic sensing data. The device includes a central processing unit (CPU), which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0087] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0088] The processing unit executes the various methods and processes described above, such as methods S1 to S3 and methods A1 to A3. For example, in some embodiments, methods S1 to S3 and methods A1 to A3 can be implemented as computer software programs, which are tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 to S3 and methods A1 to A3 described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute methods S1 to S3 and methods A1 to A3 in any other appropriate manner (for example, by means of firmware).
[0089] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0090] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0091] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0092] Embodiment 2
[0093] To verify the feasibility and superiority of the method provided by the present invention, the method provided by the present invention is now compared with the method selected by the principal component analysis method in the publicly available DAS classification dataset (Cao X, et al. An open dataset of Compare the experimental results with those of two classification models as baselines. (Results in Optics, 2023, 10: 100372.) The dataset contains a total of six different types of events, namely: background noise (0), excavation (1), knocking (2), watering (3), shaking the railing (4), walking (5). It is divided into a training set and a test set according to a ratio of 4:1, with more than 12,000 and more than 2,000 samples respectively.
[0094] First, pre-train the model provided by the present invention on the training set. During the training process, only waterfall chart data is used, and their label information is not used. Optimize the training of the Transformer network model, set the learning rate to 0.0001, and decay the learning rate according to the cosine function with the number of iterations. After 500 rounds of training, good results are achieved.
[0095] Test the trained Transformer network model and the principal component analysis method on the test set. For the method provided by the present invention, extract the data features in this test set, and the obtained feature size is 196×624×576 and perform the following operations:
[0096] a. Average the 3D features along the dimension of the number of waterfall chart blocks, and remove the dimension of the number of waterfall chart blocks to obtain 2D average features;
[0097] b. Use t-distributed stochastic neighbor embedding (t-SNE) to reduce the dimensionality of the average features to 2D to obtain a 1×1 matrix, representing a point in the plane. The result of dimensionality reduction does not depend on any label information, and only retains the proximity relationship or similarity degree of the features in the high-dimensional space.
[0098] C. Reduce the dimensionality of the data features of each waterfall chart in the public data test set using the t-SNE method and plot them in Figure 5 For better visual effects, use label information to color the corresponding two-dimensional points after dimensionality reduction, that is, the same color is used for the same type of event, and different colors are used for different types of events.
[0099] Use principal component analysis (PCA) to replace the pre-trained model to extract the features of the waterfall chart, and use the exact same t-SNE algorithm to visualize these features and plot them in Figure 5 in.
[0100] From Figure 5It can be clearly seen that the features extracted by the method adopted in the present invention show clear and obvious dispersion from various types of waterfall plots, while the features from the same type of waterfall plot tend to cluster together. Although there is still dispersion within certain event types (such as walking), this dispersion is mainly due to the rough annotation when establishing the dataset. More specifically, the "walking" label includes two activities, including walking and running, which is consistent with the two clustering features in Figure 5 ; after the features generated by PCA are dimensionally reduced, there is no obvious feature separation between different types of events, nor is there obvious clustering within the same event type. Its feature extraction ability is significantly lower than that of the method of the present invention.
[0101] Example 3
[0102] In this embodiment, the generalization ability of the model is verified in the actual anti-external damage application of a certain city. Specifically, the sensing optical fiber is buried in a U-shaped groove, and a DAS device with a sampling rate of 2kSa / s and a spatial resolution of 10 meters is used to capture vibrations within a range of 70 meters. We created a dataset containing eight different event types: background noise (0), roller compactor driving (1), roller compactor compaction (2), excavator excavation (3), excavator driving (4), electric drill excavation (5), fully loaded forklift driving (6), and unloaded forklift driving (7). The dataset contains approximately 1,400 elements, and the number of elements of each type is 42 (0), 100 (1), 300 (2), 332 (3), 334 (4), 38 (5), 54 (6), 186 (7) respectively. The size of each waterfall plot is set to 7×9984 (spatial×time). The waterfall plots of each type are divided into training data and test data in a ratio of 4:1, thereby creating a training set and a test set.
[0103] First, the pre-trained Transformer network model provided by the present invention is fine-tuned using the training set. Specifically, an additional trainable linear layer is added to the tail of the trained encoding network. This linear layer maps the features to the probability of each class on the training set by optimizing the cross-entropy loss, and further fine-tunes 150 rounds on the training set of this data. After the fine-tuning is completed, a classification task is performed on the test set, and the observed classification accuracy is 92.80%. Since the accuracy exceeds 90%, this is very valuable for the actual application of this network.
[0104] To demonstrate the advantages of the model of the present invention, a classification model exactly the same as it is built. This model uses the initial parameters and has not been pre-trained, and is called the initial model. The initial model is supervised and trained on the same training set until convergence, and its classification accuracy on the test set is only 74.24%, which is 18.56% lower than that of the pre-trained model of the present invention.
[0105] Draw the confusion matrices of the pre-trained network model and the initial model, as Figure 6 and Figure 7 shown. The horizontal axis of the confusion matrix represents the labels predicted by the model, and the vertical axis represents the true labels of the data (usually manually calibrated). Therefore, the (i, j) element in the matrix represents the data for which the model predicts the j-th type of event and the actual label is the i-th type of event. Most confusion matrices are square matrices. The diagonal elements in the square matrix are the samples predicted correctly, and the non-diagonal elements represent the samples predicted incorrectly. The prediction accuracy of the model can be obtained by dividing the total number of data on the diagonal by the total number of data in the confusion matrix.
[0106] From Figure 6 the confusion matrix shown, it is not difficult to find that the prediction errors of the pre-trained network model (the method provided by the present invention) in background noise (0) and full-load forklift operation (6) are much smaller than those of the initial model. This indicates that the negative impact brought by limited training samples can be alleviated through the Transformer network model. In addition, the confusion between excavator digging (3) and excavator driving (4) in the initial model has been improved in the results of the pre-trained network model.
[0107] As mentioned above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for extracting characteristics of distributed fiber optic acoustic sensing data, characterized in that, The method uses a Transformer network model pre-trained by a masked reconstruction task to extract features from distributed fiber optic acoustic sensing data. The Transformer network model is improved using an attention mechanism and includes an encoding network and a decoding network with the same structure. Both the encoding network and the decoding network include a linear layer, a position encoding layer, and a multi-level cascaded improved Transformer encoder. The pre-training includes: Obtain a pre-training dataset, and cut each waterfall diagram in the pre-training dataset into multiple waterfall diagram blocks of the same size, and the shape of the waterfall diagram block is: space × time. Perform independent masking processing on each cut waterfall diagram according to a uniform distribution. The unmasked waterfall diagram blocks are called visible waterfall diagram blocks. Stack the visible waterfall diagram blocks of the same waterfall diagram to obtain a first 3D tensor with the shape: number of waterfall diagram blocks × space × time. After dimension processing on the first 3D tensor, use the Transformer network to generate a reconstructed waterfall diagram, and calculate the difference between the reconstructed waterfall diagram and the corresponding original waterfall diagram. Based on the difference, construct a loss function for pre-training.
2. The method for extracting characteristics of distributed fiber optic acoustic sensing data according to claim 1, wherein The method of dimension processing is: flatten the space dimension and time dimension of the first 3D tensor to obtain a 2D tensor with the shape: number of waterfall diagram blocks × number of waterfall diagram sampling points; add a dimension to the 2D tensor to expand it into a second 3D tensor with the shape: number of waterfall diagram blocks × number of waterfall diagram sampling points × processing dimension.
3. A method for extracting characteristics of distributed fiber optic acoustic sensing data according to claim 2, characterized in that The encoding network performs the following operations: After expanding the processing dimension of the second 3D tensor using the linear layer, add trainable standard sine and cosine position information using the position encoding layer to obtain first intermediate data. Process the first intermediate data using the improved Transformer encoder. The improved Transformer encoder includes a multi-level cascaded network, and each layer of the network includes a multi-head attention mechanism. Each head of the attention mechanism outputs a second intermediate data. Concatenate all the second intermediate data in the processing dimension to obtain an output. The input of the first layer network is the first intermediate data, the input of the Nth layer network is the output of the (N - 1)th layer network, and the output of the last layer network is the third intermediate data. Use the linear layer to perform dimension processing on the third intermediate data to obtain first waterfall diagram data with the same shape and size as the first intermediate data.
4. A method for extracting characteristics of distributed fiber optic acoustic sensing data according to claim 3, characterized in that In each head of the attention mechanism, perform the following operations to obtain the second intermediate data, including: Calculate the query matrix, key matrix, and value matrix based on the first intermediate data. After performing matrix multiplication on the query matrix and the key matrix, perform normalization processing to obtain a weight matrix. Multiply the weight matrix and the value matrix to obtain intermediate data.
5. A method for extracting characteristics of distributed fiber optic acoustic sensing data according to claim 3, characterized in that, The shapes and sizes of the first intermediate data, the output of each layer of the network, and the second intermediate data are the same; the number of waterfall diagram blocks and the number of waterfall diagram sampling points of the third intermediate data are the same as those of the second intermediate data, and its processing dimension is an integer multiple of the second intermediate data, and this integer multiple is the number of heads of the multi-head attention mechanism.
6. A method for extracting characteristics of distributed fiber optic acoustic sensing data according to claim 3, characterized in that The decoding network performs the following steps: After expanding the processing dimension of the first waterfall diagram data using a linear layer, it is restored using a trainable 3D tensor to obtain the first data; Using a position encoding layer to add trainable standard sine-cosine position information to the first data to obtain the second data; Based on the second data, it is processed using an improved Transformer encoder. The improved Transformer encoder includes a multi-level cascaded network, and each layer of the network includes a multi-head attention mechanism. Each head of the attention mechanism outputs a third data. All the third data are concatenated in the processing dimension to obtain an output. The input of the first layer network is the second data, the input of the Nth layer network is the output of the N - 1th layer network, and the output of the last layer network is the fourth data; Using a linear layer to process the dimension of the fourth data to make its shape and size the same as the second data, and mapping the processing dimension of the dimension-processed fourth data to 1 to obtain a reconstructed waterfall diagram.
7. A method for extracting characteristics of distributed fiber optic acoustic sensing data according to claim 6, characterized in that The number of sampling points and the processing dimension of the trainable 3D tensor for the waterfall diagram are the same as those of the first data after the processing dimension is expanded, and the number of waterfall diagrams is 1.
8. A method for extracting characteristics of distributed fiber optic acoustic sensing data according to claim 1, characterized in that, The expression of the loss function is: Among them, represents the missing waterfall plot block after masking the i-th waterfall plot; represents the reconstruction of the missing waterfall plot block of the i-th waterfall plot; ‖·‖2 represents the calculation of the second norm; Mean(·) represents the average over the entire training data set.
9. A method for extracting characteristics of distributed fiber optic acoustic sensing data according to claim 1, characterized in that, The method includes: Cutting the target waterfall diagram into multiple waterfall diagram blocks of the same size, and the shape of each waterfall diagram block is space × time; Stacking and flattening the multiple waterfall diagram blocks into a 3D tensor; Inputting the 3D tensor into the encoding network in the pre-trained Transformer network model to extract the distributed fiber optic acoustic sensing data features.
10. A distributed optical fiber acoustic wave sensing data feature extraction device, comprising a memory and a processor, wherein a computer program is stored on the memory, and is characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method for carrying out multi-view feature extraction by utilizing lightweight Transform-CNN (Convolutional Neural Network) fusion network
CN119360036A
Cited By
Optical fiber sound wave event identification method, equipment and medium
CN120724138A
Optical fiber sound wave sensing data feature extraction method and device and computer storage medium
CN121479281A
Method, apparatus and computer storage medium for feature extraction of fiber optic acoustic wave sensing data
CN121479281B