Brillouin Temperature Extraction Method and System Based on Self-Attention Recurrent Neural Network
By adopting a self-attention recurrent neural network method in BOTDA technology and optimizing the gating mechanism with meta-learning algorithm, the problem of noise sensitivity, fitting dependence on initial parameters and long calculation time in traditional technology is solved, and more efficient and accurate Brillouin temperature extraction is achieved.
Patent Information
- Application Number
- CN202510336074.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Traditional Brillouin Optical Time Domain Analysis (BOTDA) technology is sensitive to noise during temperature extraction. The fitting results rely on initial parameter selection and have a long calculation time, which limits the feasibility of real-time applications. The static gating mechanism cannot adapt to the dynamic changes in information flow requirements, affecting the extraction efficiency.
The self-attention recurrent neural network is used to optimize the task adaptive weights of the transfer gate and the conversion gate in combination with the meta-learning algorithm, so that the model can dynamically adjust the information flow and adapt to complex environments, thereby improving the prediction accuracy and real-time performance of Brillouin temperature.
It improves the accuracy and real-time performance of Brillouin temperature extraction, reduces the time of curve fitting, enhances the adaptability and robustness of the model, is suitable for real-time temperature monitoring, and reduces measurement errors caused by noise interference.
Smart Images

Figure CN119848440B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distributed optical fiber sensing, and particularly relates to a Brillouin temperature extraction method and system based on a self-attention recurrent neural network. Background Art
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] BOTDA refers to a technology for optical fiber sensing based on the Brillouin scattering principle, and its full name is Brillouin Optical Time Domain Analysis. The Brillouin optical time domain analysis system is a typical Brillouin distributed optical fiber sensing system, which usually consists of a laser, a coupler, an electro-optic modulator, an arbitrary signal generator, an erbium-doped fiber amplifier, a grating filter, a circulator, a polarization scrambler, a photodetector, and a data acquisition unit. Based on the Brillouin scattering principle, BOTDA extracts the temperature and strain information of the target area by analyzing the scattering signal during the optical fiber transmission process. It is widely used in geological monitoring, environmental monitoring, intelligent buildings and other fields, and is applied to monitor the temperature and strain of infrastructure such as oil and gas pipelines, bridges, and tunnels.
[0004] The traditional method uses the Lorentz curve fitting method to obtain the Brillouin frequency shift, finds the center frequency corresponding to the maximum value of the measured Brillouin gain spectrum as the Brillouin frequency shift, and obtains the temperature through the linear relationship formula between the Brillouin frequency shift and the temperature. However, the traditional Lorentz curve fitting method is very sensitive to noise in the data. Environmental noise or signal interference may lead to inaccurate fitting results, thus affecting the accuracy of temperature measurement. The fitting result strongly depends on the selection of initial parameters. If the initial value is set improperly, it may lead to fitting failure or result deviation, especially in the case of complex data. The fitting process requires a long calculation time, especially when dealing with big data, and the efficiency is low, which limits the feasibility of real-time applications. With the development of machine learning, the method of combining machine learning with the Lorentz curve has also been used for the extraction of Brillouin temperature. Among them, the Recurrent Highway Networks (RHN) network is used for Brillouin temperature extraction. However, the gating mechanism of the RHN network is static and has various limitations. The static gating cannot adapt to the dynamic changes of the information flow requirements. The information flow requirements are different between different tasks, but the fixed gating mechanism may lead to a decline in the performance of some tasks. In multi-task learning, the static gating mechanism can only select the shared optimization between tasks, and it is difficult to quickly adjust the characteristics of each task, thus making the overall extraction time too long and affecting the efficiency of Brillouin temperature extraction. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a Brillouin temperature extraction method and system based on a self-attention recurrent neural network. In the BOTDA temperature extraction task, after introducing meta-learning, the gating mechanism of the RHN can dynamically adjust the information flow according to different time steps or tasks, enabling the model to better adapt to complex environments, thereby improving the prediction and extraction accuracy and real-time performance of Brillouin temperature, and avoiding problems such as long time for curve fitting, poor real-time performance, and poor adaptability.
[0006] According to some embodiments, the first solution of the present invention provides a Brillouin temperature extraction method based on a self-attention recurrent neural network, adopting the following technical solution:
[0007] The Brillouin temperature extraction method based on a self-attention recurrent neural network includes:
[0008] Using a pre-built Brillouin optical time domain analysis system to collect real Brillouin gain spectrum data and perform preprocessing;
[0009] Based on the preprocessed Brillouin gain spectrum data, using a trained self-attention recurrent neural network model to obtain the corresponding Brillouin temperature;
[0010] Among them, the training process of the self-attention recurrent neural network model is specifically:
[0011] Simulating ideal Brillouin gain spectrum data and converting it into an input sequence to obtain a sample data set;
[0012] Setting an initial hidden state, and using a meta-learning algorithm to optimize the task-adaptive weights of the transfer gate and the conversion gate at each time step;
[0013] Using the optimized conversion gate to determine the conversion gate output of each layer at each time step, and using the optimized transfer gate to determine the candidate hidden state of each layer at each time step;
[0014] Based on the hidden state of the previous time step, the conversion gate output of the current time step, and the candidate hidden state to update the hidden state within the current time step, which is used as the final output of the model, that is, the Brillouin temperature;
[0015] Iteratively optimizing and adjusting the task-adaptive weights and biases of the transfer gate and the conversion gate, and training with the goal of minimizing the loss to obtain a trained self-attention recurrent neural network model.
[0016] Furthermore, the Brillouin optical time domain analysis system includes a narrowband laser, and the narrowband laser is sequentially connected to a polarization-maintaining fiber coupler, a semiconductor optical amplifier, a first erbium-doped fiber amplifier, a first fiber circulator, a second wavelength division multiplexer, a sensing fiber, a first wavelength division multiplexer, an optical fiber isolator, a polarization scrambler, and a dual-drive electro-optic phase modulator;
[0017] The first erbium-doped fiber amplifier is also connected to an acquisition card through a semiconductor optical amplifier frequency-sweeping module;
[0018] The first optical fiber circulator is also connected to an optical fiber grating filter and a second erbium-doped fiber amplifier through a second optical fiber circulator, and the second erbium-doped fiber amplifier is connected to the acquisition card through an optoelectronic probe; the optical fiber grating filter is also connected to a temperature control device;
[0019] The dual-side electro-optic phase modulator is also connected to the acquisition card through a frequency-sweeping module, the dual-side electro-optic phase modulator is also connected to a polarization-maintaining fiber coupler, and the acquisition card is connected to a computer.
[0020] Further, the self-attention recurrent neural network model includes an input layer, a hidden layer, a self-attention layer, and an output layer;
[0021] A highway layer is added to the hidden layer, and the highway layer includes two gating units, a carry gate and a transform gate.
[0022] Further, after the self-attention layer obtains the output of the hidden layer, it calculates the similarity between positions of the input sequence to assign attention weights, and performs weighted summation on important features in the input sequence to obtain the output of the self-attention layer;
[0023] The output of the output layer is the result of fusing the output of the self-attention layer and the output of the hidden layer.
[0024] Further, the task adaptive weights of the carry gate and the transform gate are optimized using a meta-learning algorithm in each time step, specifically:
[0025] In each time step, multiple training task data are randomly sampled, where the training task data includes training data and validation data;
[0026] For each training task, model training is performed on its training data, and gradient updates are performed based on the loss function, learning rate, and parameters of the model of the training task, thereby updating the meta-parameters of the gating mechanism;
[0027] Based on the updated meta-parameters of the gating mechanism, the task adaptive weights of the carry gate and the transform gate are optimized.
[0028] Further, the meta-learning algorithm includes inner-loop optimization and outer-loop optimization;
[0029] The inner-loop optimization is to optimize the parameters of the gating mechanism for each task separately, and use a standard time series loss function as the task-level loss function for the inner-loop optimization;
[0030] The outer loop optimization optimizes the meta-parameters of the gating mechanism through meta-gradient update, and uses the average loss of all tasks as the meta-loss function for outer loop optimization.
[0031] According to some embodiments, the second solution of the present invention provides a Brillouin temperature extraction system based on a self-attention recurrent neural network, adopting the following technical solution:
[0032] The Brillouin temperature extraction system based on a self-attention recurrent neural network includes:
[0033] A data acquisition module, configured to collect real Brillouin gain spectrum data and perform preprocessing by using a pre-built Brillouin optical time domain analysis system;
[0034] A temperature extraction module, configured to obtain the corresponding Brillouin temperature by using a trained self-attention recurrent neural network model based on the preprocessed Brillouin gain spectrum data;
[0035] Among them, the training process of the self-attention recurrent neural network model is specifically as follows:
[0036] Simulate ideal Brillouin gain spectrum data and convert it into an input sequence to obtain a sample data set;
[0037] Set the initial hidden state, and optimize the task adaptive weights of the transmission gate and the conversion gate at each time step by using a meta-learning algorithm;
[0038] Use the optimized conversion gate to determine the conversion gate output of each layer at each time step, and use the optimized transmission gate to determine the candidate hidden state of each layer at each time step;
[0039] Update the hidden state within the current time step based on the hidden state of the previous time step, the conversion gate output of the current time step, and the candidate hidden state, as the final output of the model, that is, the Brillouin temperature;
[0040] Iteratively optimize and adjust the task adaptive weights and biases of the transmission gate and the conversion gate, and perform training with the goal of minimizing the loss to obtain a trained self-attention recurrent neural network model.
[0041] According to some embodiments, the third solution of the present invention provides a computer-readable storage medium.
[0042] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the Brillouin temperature extraction method based on a self-attention recurrent neural network as described in the first aspect above.
[0043] According to some embodiments, the fourth solution of the present invention provides a computer device.
[0044] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the Brillouin temperature extraction method based on the self-attention recurrent neural network described in the first aspect above.
[0045] According to some embodiments, a fifth aspect of the present invention provides a computer program product or a computer program.
[0046] The present invention provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the Brillouin temperature extraction method based on the self-attention recurrent neural network described in the first aspect above.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0048] The present invention proposes a Brillouin temperature extraction method based on a self-attention recurrent neural network, which combines the ideas of recurrent neural network (RNN) and Highway Network, can solve the problem of long-term dependence, and has the ability to enhance the training depth of the recurrent network and process deep data features, can well train a deep network, learn deep feature relationships, solve the problems of gradient disappearance and gradient explosion, and improve the accuracy and real-time performance of temperature extraction; at the same time, combining the self-attention mechanism can better capture global and local features, improve the accuracy of temperature measurement, and can more effectively analyze complex optical fiber signals; the network model can quickly respond to dynamic changes in the measured environment, improve the sensitivity to temperature changes, improve the signal processing speed through efficient feature extraction and analysis, is applicable to real-time temperature monitoring, and more accurate signal processing helps to reduce measurement errors caused by noise and interference. The meta-learning algorithm is introduced, especially its application in the gating mechanism. By learning task-specific gating parameters, the network can adjust the information flow according to the characteristics of the BOTDA signal, ensuring that the network can effectively extract useful information from the input data while avoiding the interference of noise. In the practical application of BOTDA, various different tasks may be encountered, such as different temperature ranges. After introducing meta-learning, the network can adaptively adjust the gating mechanism according to task embeddings or context information, so that the training process of each task can obtain optimal gating parameters, thereby improving the performance of each task in a multi-task learning environment. The BOTDA sensor data may vary with environmental changes or different devices. Traditional models may perform poorly on new tasks, while the self-attention recurrent neural network model introduced with meta-learning can improve the temperature extraction performance of the model in different situations by learning how to quickly adapt to new tasks or data distributions. In the BOTDA temperature extraction task, after introducing meta-learning, the gating mechanism of the self-attention recurrent neural network model can dynamically adjust the information flow according to different time steps or tasks, enabling the model to better adapt to complex environments and improve prediction accuracy and robustness. For example, on new sensor data, meta-learning enables the network to quickly adjust its gating mechanism to better adapt to the characteristics of the new input data.
[0049] In practical applications, BOTDA technology may face different environmental conditions, different fiber optic layouts, different measurement configurations, etc. These conditions may affect the accuracy of temperature measurement. Multi-task learning can be used to model different measurement scenarios, thereby improving the generalization ability of the model under different conditions. The Brillouin frequency shift in BOTDA is affected by temperature and may have differences in frequency bands or resolutions. Multi-task learning can improve the accuracy of temperature extraction by learning different frequency bands or data with different resolutions. The gating mechanism in the self-attention recurrent neural network model is a static structure. Introducing a meta-learning algorithm during the training process of this network model can help dynamically adjust the gating strategy, solve the new problem of dynamic information flow, and enable the network to adaptively adjust the information transmission method according to different inputs, tasks, time steps, and other dynamic conditions. Through meta-learning, the network model can learn how to adjust its gating mechanism according to the characteristics of the current task or input, thereby achieving more efficient learning and adaptability during the information flow transmission process; it can share the initialization parameters among multiple tasks and optimize these parameters through meta-training, enabling the network to quickly adapt to new tasks; when facing multi-task learning or new tasks, it does not need to be trained from scratch, but can quickly adjust the model through a small number of gradient updates, reducing the training time and improving the performance. After introducing meta-learning, the training process of the network model becomes more complex and hierarchical. In the meta-training stage, the network optimizes the initialization parameters of the model by learning on multiple tasks, aiming to enable the model to quickly adapt to new tasks. In the meta-testing stage, the model quickly adapts to unseen tasks through a small number of gradient updates to verify the effect of meta-learning. Through such a training process, the self-attention recurrent neural network model can, when facing new tasks, not be trained from scratch, but be quickly fine-tuned based on an optimized initialization parameter to achieve better results. This enables the model to have stronger adaptability and generalization ability in multi-task learning and task transfer. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0051] Figure 1 is a flowchart of a Brillouin temperature extraction method based on a self-attention recurrent neural network in an embodiment of the invention;
[0052] Figure 2 is a schematic diagram of the overall architecture of the self-attention recurrent neural network model in an embodiment of the invention;
[0053] Figure 3 is a schematic diagram of the gating mechanism structure of the self-attention recurrent neural network model in an embodiment of the invention;
[0054] Figure 4 It is a schematic structural diagram of the self-attention layer in an embodiment of the present invention;
[0055] Figure 5 It is a structural diagram of the Brillouin optical time domain analysis system in an embodiment of the present invention;
[0056] Reference numerals:
[0057] NB Laser - narrowband laser; PM - OC - polarization-maintaining fiber coupler; SOA - semiconductor optical amplifier; EDFA1 - first erbium-doped fiber amplifier; EDFA2 - second erbium-doped fiber amplifier; SOA driver - SOA frequency-sweeping module; PD - photoelectric probe; FBG - fiber Bragg grating filter; TEC - temperature control device; WDM1 - first wavelength division multiplexer; WDM2 - second wavelength division multiplexer; FUT - sensing fiber; ISO - fiber isolator; Polarization scrambler - polarization scrambler; EOM - dual-side electro-optic phase modulator; EOM driver - EOM frequency-sweeping module; Circulator1 - first fiber circulator; Circulator2 - second fiber circulator. Detailed implementation manners
[0058] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0059] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0060] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0061] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0062] Embodiment 1
[0063] This embodiment provides a Brillouin temperature extraction method based on a self-attention recurrent neural network. This embodiment takes the application of this method to a server as an example. It can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, web servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. In this embodiment, the method includes the following steps:
[0064] Use the pre-built Brillouin optical time domain analysis system to collect real Brillouin gain spectrum data and perform preprocessing;
[0065] Based on the preprocessed Brillouin gain spectrum data, use the trained self-attention recurrent neural network model to obtain the corresponding Brillouin temperature;
[0066] Among them, the training process of the self-attention recurrent neural network model is specifically as follows:
[0067] Simulate ideal Brillouin gain spectrum data and convert it into an input sequence to obtain a sample data set;
[0068] Set the initial hidden state, and use the meta-learning algorithm to optimize the task adaptive weights of the transfer gate and the conversion gate at each time step;
[0069] Use the optimized conversion gate to determine the conversion gate output of each layer at each time step, and use the optimized transfer gate to determine the candidate hidden state of each layer at each time step;
[0070] Based on the hidden state of the previous time step, the conversion gate output of the current time step, and the candidate hidden state, update the hidden state within the current time step as the final output of the model, that is, the Brillouin temperature;
[0071] Iteratively optimize and adjust the task adaptive weights and biases of the transfer gate and the conversion gate, and train with the goal of minimizing the loss to obtain the trained self-attention recurrent neural network model.
[0072] As Figure 1 shown, the method described in this embodiment specifically includes the following steps:
[0073] S1. Generate the Brillouin gain spectrum by the pseudo-Voigt function, obtain the Brillouin gain spectrum data to form a training set; the pseudo-Voigt curve function is a combination of the Lorentz curve and the Gaussian curve, and the formula of the pseudo-Voigt curve function is as follows:
[0074] (1);
[0075] where is the peak gain, is the Brillouin frequency shift (BFS), is the full width at half maximum of the Brillouin gain (BGS), is the proportion of the Lorentz component.
[0076] S2. Construct a self-attention recurrent neural network model (SA-RHN), and input the training set obtained in S1 into the SA-RHN for training to obtain a trained network model;
[0077] As Figure 2 shown, based on the temperature extraction technology of Brillouin optical time domain analysis (BOTDA), a self-attention recurrent neural network model (SA-RHN) based on the RHN (Recurrent Highway Network) integrated with the self-attention mechanism is proposed. Among them, the self-attention recurrent neural network model includes an input layer, a hidden layer, a self-attention layer, and an output layer, and the output layer is a fully connected layer (Fully Connected Layer). The activation functions ReLU and Tanh are applied to the recurrent unit and each gate.
[0078] A highway layer is added to the hidden layer, and the highway layer includes two gating units, namely a carry gate and a transform gate. The RHN controls the information flow through a high-dimensional transformation layer. At each time step, the RHN calculates a transform gate and a carry gate, and these two gates jointly determine how the old state is updated and when the new state starts; after obtaining the output of the hidden layer, the self-attention layer calculates the attention weights by calculating the similarity between the positions of the input sequence, and performs weighted summation on the important features in the input sequence to obtain the output of the self-attention layer; the output of the output layer is the result of the fusion of the output of the self-attention layer and the output of the hidden layer, and can be output through methods such as weighted summation, concatenation, or residual connection.
[0079] In step S2, the training process of the SA-RHN network model includes:
[0080] 1) Receive the input sequence including preprocessing, and convert the original input into a format suitable for the network.
[0081] 2) Set the initial hidden state Usually a zero vector or a small random value.
[0082] 3) For each time step and for each layer, perform the following operations:
[0083] (2);
[0084] In the formula, is the output of the transformation gate, is the sigmoid function, which is used to map the result to the interval [0, 1], is the task adaptive weight of the transformation gate optimized by meta-learning, is the weight matrix, is the transformation gate bias vector, is the layer's hidden state at time step , is the layer's hidden state at the previous time step , which is another input for the current layer to combine the information flow in time.
[0085] (3);
[0086] Among them, is the hyperbolic tangent function, and the output is in the interval [-1, 1], is the task adaptive weight of the transfer gate optimized by meta-learning, is the weight matrix, is the transfer gate bias vector, is the candidate hidden state at the current time step, is the layer's hidden state at the previous time step , representing the historical information of the current layer in the time series. is the layer's hidden state at time step . This is one of the inputs calculated by the current layer and contains the information of the previous layer.
[0087] 4) Update the hidden state
[0088] (4);
[0089] Among them, ⊙ represents element-wise multiplication, and the transfer gate (C) and the transformation gate (T) are used to update the hidden state.
[0090] 5) Apply the above update formula (4) multiple times in the same time step, and the multi-layer design further extracts features.
[0091] Finally, extract the output from the last layer or the specified layer for subsequent processing.
[0092] 6) Calculate the loss function, commonly cross-entropy or mean squared error. Use the backpropagation algorithm, combined with optimization methods (Adam, SGD), to adjust the weights and biases of each layer to minimize the loss.
[0093] Use the model-agnostic meta-learning (MAML) algorithm to train the gating parameters of the RHN's gating mechanism. In the RHN, the gating mechanism (transmission gate and transformation gate) is static, that is, these gating parameters are fixed as a set of task-specific weights during training. However, the static gating mechanism has obvious limitations, especially in task environments or multi-task scenarios that require handling dynamic changes. By introducing meta-learning, the gating mechanism of the RHN can be made dynamic, thus solving a series of new problems in dynamic information flow control. In the RHN network, the model focuses only on a specific task, while meta-learning aims to "learn" a learning strategy that can be generalized to different tasks through multi-task learning. During the training process, the learning rate and other hyperparameters are dynamically adjusted, enabling the network to quickly adapt to different inputs or tasks and adaptively adjust according to the distribution of different tasks or data. After introducing meta-learning, the initialization of the model is no longer random but is optimized through learning on multiple tasks. In this way, the network can quickly adapt to new tasks, reducing the time and complexity of training from scratch. Meta-learning helps dynamically learn the gating mechanism (weights of the transmission gate and transformation gate), enabling them to be adjusted according to different input sequences, contexts, or historical information. This process occurs during the gating calculation at each time step, especially in the weight update of the transmission gate and transformation gate, allowing the gating mechanism to adjust more quickly when encountering new data without having to retrain the entire network completely. In the meta-training stage, the model learns how to quickly adapt to new tasks through multiple tasks. Each task has its own training set and validation set. In the meta-testing stage, the model needs to face new tasks that did not appear in the meta-training process. This stage mainly tests whether the model can utilize the previously learned initialization to quickly adapt to new tasks. It is necessary to design a strategy for dynamically adjusting gating parameters (such as the transmission gate and transformation gate) so that the gating mechanism can optimize the transmission of information flow according to the dynamic changes of tasks and inputs.
[0094] The process of the meta-learning algorithm is carried out at each time step during the training of the self-attention recurrent neural network model (SA-RHN). The model can update its own parameters in real-time according to the feedback, improving its learning ability on new tasks. The implementation process and calculation details are as follows:
[0095] Define the dynamic gating mechanism
[0096] As Figure 3 shown, the gating mechanism of RHN includes: a transmission gate - determining the passing ratio of the current input information. A conversion gate - determining the influence of the hidden state at the previous time step on the current state. The task of meta-learning is to dynamically generate and task embedding according to the input and so that the gating mechanism can adapt to the current task or input. Dynamic gating generation formula Traditional and are generated with fixed weights: , . After introducing meta-learning, and can be dynamically adjusted during generation:
[0097] (5);
[0098] (6);
[0099] wherein, is the task adaptive weight of the conversion gate optimized by meta-learning, is the task adaptive weight of the transmission gate optimized by meta-learning, is the optimized conversion gate, is the optimized transmission gate, is the conversion gate bias vector, is the transmission gate bias vector, is the task embedding, representing the global information of the task (dynamically updated by meta-learning), is the activation function Sigmoid .
[0100] Inner and outer loop training
[0101] Meta-learning algorithms usually include two optimization layers: inner loop optimization (Task-specific Training) and outer loop optimization (Meta-Training). The goal of the inner loop is to optimize the parameters of the gating mechanism for each task separately, so that the network performs best on the specific data of the current task. For each task , according to the task data (training set and validation set) adjust the parameters of the gating mechanism and , and the standard time series loss function (mean square error in temperature prediction) is used during the training process:
[0102] (7);
[0103] in, It means Task No. The true value of the time step, It means Task No. The predicted value of time steps, N represents the total number of time steps in the current task (i.e. the length of the time series);
[0104] The goal of the outer loop is to optimize the meta parameters of the gating mechanism through meta gradient updates so that the model can quickly adapt to new tasks. The meta-optimization loss is the average loss of all tasks:
[0105] (8);
[0106] is the meta-optimization loss of the outer loop, where is the total number of tasks, used for averaging;
[0107] The optimization process calculates the meta-gradient through back-propagation, updates the meta-parameters of the gating mechanism, and obtains the updated meta-parameters :
[0108] (9);
[0109] in, is the learning rate of the outer loop meta-learning, which controls the step size of the meta-parameter update. Formula (9) shows that in meta-learning, the model needs to optimize its initial parameters based on the performance of the task’s validation set. , so the gradient calculation here uses the validation loss, which optimizes the average performance of all tasks on the validation set so that the initial parameters of the model Ability to adapt quickly to new tasks.
[0110] Loss function design
[0111] Task-level loss: ;
[0112] Yuan loss: ;
[0113] Total loss: ;
[0114] The RHN network training process, which goes through the meta-training and meta-testing phases, is as follows:
[0115] Meta-training phase:
[0116] Task Sampling: Randomly sample multiple tasks from the task distribution , each task contains training data and validation data. The training data is used for in-task training, and the validation data is used for in-task evaluation.
[0117] In-task Training: For each task perform model training on its training data. At this time, the model will perform several gradient updates according to the data of the task. Through such in-task training, the parameters of the model will be optimized on this task. The loss function of the task , then the gradient update process on task is as follows:
[0118] (10);
[0119] wherein, is the original parameter of the model, is the new parameter obtained after in-task training of task , is the learning rate, is the loss function of task , and formula (10) optimizes the loss for a single task.
[0120] In-task Evaluation (Meta-Validation): Evaluate the model parameters after in-task training on the validation set of each task Calculate the validation loss of task :
[0121] (11);
[0122] Meta-Update:
[0123] (12);
[0124] wherein, is the learning rate of meta-learning in the inner loop, which controls the step size of task-specific parameter updates, is the validation loss on task , and the sum of losses for multiple tasks in formula (12) is used for different optimization strategies, affecting the overall adaptability between tasks.
[0125] Meta-Testing Phase
[0126] Task Selection: Select a new task from the task set , this task did not appear in meta-training. The goal is to test the model's ability to quickly adapt to unseen tasks.
[0127] Fast Adaptation: Using the initialization parameters learned in the meta-training phase, directly apply this initialization to the training data of the new task. Perform a small number of gradient updates (usually a few times) on the new task, aiming to quickly adapt to the new task. At this time, the model adjusts its parameters in small steps. Assume that on the new task the loss function is The update process is as follows:
[0128] (13);
[0129] where represents the model parameters after being updated with a small amount of training data, represents the meta-parameters of the current model, represents the learning rate, represents based on the current parameters The loss function on the new task calculated, because in the fast adaptation phase, the model uses the training data to perform a small number of gradient updates in order to better adapt to the new task.
[0130] Since the parameters initialized by using the meta-learning algorithm have been used, the network can quickly adapt to the new task through a small number of updates.
[0131] Evaluation: Evaluate the performance of the model on the validation set of the new task to check the generalization ability of the model after a small number of gradient updates. The goal is to verify whether the model initialization obtained through meta-learning can help the model quickly adapt to the new task.
[0132] As Figure 4 shown, in step S2, the calculation process of the self-attention mechanism includes:
[0133] Prepare the input: Receive n inputs and prepare for calculation;
[0134] Inputs: Q - Query, K - Key, V - Value;
[0135] Initialize weights: Initialize weights for the query, key, and value.
[0136] Derive Key, Query, and Value: Calculate the key, query, and value for each element.
[0137] Calculate the attention score: Calculate the attention score through the similarity between the query and the key. Calculate softmax : Pass the attention score through softmaxThe function performs normalization processing.
[0138] Matrix multiplication (MatMul) MatMul(Q, K^T): First, perform matrix multiplication on the query vector Q and the transpose of the key vector K to generate an attention score matrix. Each element of this score matrix represents the similarity between the query vector and the key vector.
[0139] Scale (scaling) After matrix multiplication, the score matrix is scaled. Generally, the scaling factor is , where is the dimension of the key vector. The purpose of scaling is to avoid the score value being too large when the dimension is large, resulting in too small a gradient of the Softmax function.
[0140] Mask (optional) Masking: In some cases (such as processing autoregressive models), to prevent the model from using future information, a mask operation is performed on future time steps, so that when calculating attention, the model cannot see the values at future moments. This is achieved by adding a very large negative value to certain positions in the score matrix.
[0141] Softmax: Apply the Softmax function to convert the scaled score matrix into a probability distribution. The result of this step is the attention weight of each word to other words.
[0142] Weighted sum: Multiply the normalized attention scores by the values and add the results to obtain the output.
[0143] Matrix multiplication (MatMul) MatMul(Softmax Output, V ): Finally, perform matrix multiplication on the attention weights of the Softmax output and the value vector V to obtain the final weighted sum representation.
[0144] Attention calculation formula in key-value pair form:
[0145] (14).
[0146] S3. Build a Brillouin optical time domain analysis system (BOTDA), collect real Brillouin gain spectrum data, and input it into the SA-RHN network model trained in S2 to obtain the temperature corresponding to the gain spectrum.
[0147] Such as Figure 5As shown in the figure, the Brillouin optical time domain analysis system includes a narrowband laser, which is sequentially connected to a polarization maintaining fiber coupler, a semiconductor optical amplifier, a first erbium-doped fiber amplifier, a first fiber circulator, a second wavelength division multiplexer, a sensing fiber, a first wavelength division multiplexer, an optical fiber isolator, a polarization scrambler, and a dual-side electro-optic phase modulator;
[0148] The first erbium-doped fiber amplifier is also connected to an acquisition card through a semiconductor optical amplifier frequency sweeping module;
[0149] The first fiber circulator is also connected to an optical fiber grating filter and a second erbium-doped fiber amplifier through a second fiber circulator, and the second erbium-doped fiber amplifier is connected to the acquisition card through an optical and electrical probe; the optical fiber grating filter is also connected to a temperature control device;
[0150] The dual-side electro-optic phase modulator is also connected to the acquisition card through a frequency sweeping module, the dual-side electro-optic phase modulator is also connected to a polarization maintaining fiber coupler, and the acquisition card is connected to a computer.
[0151] In step S3, the Brillouin optical time domain analysis system includes a narrowband laser, a semiconductor optical amplifier, an erbium-doped fiber amplifier, an optical fiber grating filter, a polarization scrambler, an optical and electrical probe, a dual-side electro-optic phase modulator, a frequency sweeping module, a sensing fiber, an optical fiber grating filter, a polarization maintaining fiber coupler, an optical fiber isolator, a fiber circulator, a wavelength division multiplexer, a distributed Raman amplifier, a computer, and an acquisition card. After building the system, a real Brillouin gain spectrum data set is collected.
[0152] This embodiment proposes a Brillouin temperature extraction technology based on the SA-RHN network, which collects Brillouin gain spectrum data and extracts the corresponding temperature, improving the accuracy and real-time performance of temperature extraction in BOTDA. The RHN integrates the self-attention mechanism network (SA-RHN), which combines the ideas of the recurrent neural network (RNN) and the highway network, can solve the problem of long-term dependence, and has the ability to enhance the training depth of the recurrent network and process deep data features. It can well train deep networks, learn deep feature relationships, solve the problems of gradient disappearance and gradient explosion, and improve the accuracy and real-time performance of temperature extraction. Introducing the meta-learning algorithm RHN, when facing multi-task learning or new tasks, it does not need to be trained from scratch, but can quickly adjust the model through a small number of gradient updates, reducing the training time and improving the performance. Introducing the meta-learning algorithm in RHN, especially its application in the gating mechanism, can significantly improve the performance of the model on multiple tasks. Through MAML, RHN can learn a shared gating mechanism initialization, enabling the model to quickly adapt to task requirements through a small number of updates when facing new tasks. This method can not only improve the adaptability and generalization ability of the model, but also reduce the data requirements during the training process, thus performing well in multiple practical tasks.
[0153] Embodiment 2
[0154] This embodiment provides a Brillouin temperature extraction system based on a self-attention recurrent neural network, including:
[0155] A data acquisition module configured to collect real Brillouin gain spectrum data and perform preprocessing using a pre-built Brillouin optical time domain analysis system;
[0156] A temperature extraction module configured to obtain the corresponding Brillouin temperature using a trained self-attention recurrent neural network model based on the preprocessed Brillouin gain spectrum data;
[0157] Wherein, the training process of the self-attention recurrent neural network model is specifically:
[0158] Simulate ideal Brillouin gain spectrum data and convert it into an input sequence to obtain a sample data set;
[0159] Set the initial hidden state, and optimize the task-adaptive weights of the transfer gate and the transformation gate using the meta-learning algorithm at each time step;
[0160] Use the optimized transformation gate to determine the transformation gate output of each layer at each time step, and use the optimized transfer gate to determine the candidate hidden state of each layer at each time step;
[0161] Update the hidden state within the current time step based on the hidden state of the previous time step, the output of the transition gate at the current time step, and the candidate hidden state, and use it as the final output of the model, that is, the Brillouin temperature;
[0162] Iteratively optimize and adjust the task adaptive weights and biases of the transmission gate and the transition gate, and train with the goal of minimizing the loss to obtain a trained self-attention recurrent neural network model.
[0163] The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the first embodiment above. It should be noted that the above modules can be executed in a computer system such as a set of computer-executable instructions as part of the system.
[0164] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0165] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0166] Embodiment III
[0167] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the Brillouin temperature extraction method based on a self-attention recurrent neural network as described in the first embodiment above.
[0168] Embodiment IV
[0169] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the Brillouin temperature extraction method based on a self-attention recurrent neural network as described in the first embodiment above.
[0170] Embodiment V
[0171] This embodiment provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, which are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the Brillouin temperature extraction method based on a self-attention recurrent neural network as described in the first embodiment above.
[0172] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.
[0173] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0174] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0176] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above-mentioned embodiment methods, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0177] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.
Claims
1. A Brillouin temperature extraction method based on a self-attention recurrent neural network, characterized in that: include: Using the pre-built Brillouin optical time domain analysis system, the real Brillouin gain spectrum data is collected and pre-processed; Based on the preprocessed Brillouin gain spectrum data, the corresponding Brillouin temperature is obtained using a trained self-attention recurrent neural network model; the self-attention recurrent neural network model includes an input layer, a hidden layer, a self-attention layer, and an output layer; a highway layer is added to the hidden layer, and the highway layer includes two gating units, a transfer gate and a conversion gate; After the self-attention layer obtains the output of the hidden layer, it allocates attention weights by calculating the similarity between each position of the input sequence, and weights the important features in the input sequence to obtain the output of the self-attention layer; the output of the output layer is the result of the fusion of the output of the self-attention layer and the output of the hidden layer; The training process of the self-attention recurrent neural network model is specifically as follows: Simulate the ideal Brillouin gain spectrum data and convert it into an input sequence to obtain a sample data set; Set the initial hidden state and use the meta-learning algorithm to optimize the task-adaptive weights of the transfer gate and the transformation gate at each time step; The optimized transition gate is used to determine the transition gate output of each layer at each time step, and the optimized transfer gate is used to determine the candidate hidden state of each layer at each time step; Based on the hidden state of the previous time step, the output of the conversion gate of the current time step, and the candidate hidden state, the hidden state in the current time step is updated as the final output of the model, i.e., the Brillouin temperature; The task adaptive weights and biases of the transfer gate and the conversion gate are adjusted by iterative optimization, and training is performed with the goal of minimizing the loss to obtain a trained self-attention recurrent neural network model.
2. The Brillouin temperature extraction method based on self-attention recurrent neural network according to claim 1, characterized in that: The Brillouin optical time domain analysis system comprises a narrowband laser, which is sequentially connected to a polarization-maintaining fiber coupler, a semiconductor optical amplifier, a first erbium-doped fiber amplifier, a first fiber circulator, a second wavelength division multiplexer, a sensing fiber, a first wavelength division multiplexer, a fiber isolator, a polarization scrambler and a double-sideband electro-optical phase modulator; The first erbium-doped fiber amplifier is also connected to the acquisition card via a semiconductor optical amplifier frequency sweep module; The first fiber optic circulator is also connected to a fiber grating filter and a second erbium-doped fiber amplifier through a second fiber optic circulator, and the second erbium-doped fiber amplifier is connected to the acquisition card through a photoelectric probe; the fiber grating filter is also connected to a temperature control device; The double-sideband electro-optical phase modulator is also connected to an acquisition card through a frequency scanning module, the double-sideband electro-optical phase modulator is also connected to a polarization-maintaining optical fiber coupler, and the acquisition card is connected to a computer.
3. The Brillouin temperature extraction method based on self-attention recurrent neural network according to claim 1, characterized in that: The meta-learning algorithm is used to optimize the task adaptive weights of the transfer gate and the conversion gate in each time step, specifically: Randomly sample multiple training task data in each time step, where the training task data includes training data and verification data; Perform model training on the training data of each training task, perform gradient updates based on the loss function, learning rate, and model parameters of the training task, and then update the meta-parameters of the gating mechanism; The task-adaptive weights of transfer gates and transformation gates are optimized based on the updated meta-parameters of the gating mechanism.
4. The Brillouin temperature extraction method based on self-attention recurrent neural network according to claim 3, characterized in that: The meta-learning algorithm includes inner loop optimization and outer loop optimization; The inner loop optimization is to optimize the parameters of the gating mechanism for each task separately, and use the standard time series loss function as the task-level loss function of the inner loop optimization; The outer loop optimization is to optimize the meta-parameters of the gating mechanism through meta-gradient update, and the average loss of all tasks is used as the meta-loss function of the outer loop optimization.
5. Brillouin temperature extraction system based on self-attention recurrent neural network, characterized in that: include: A data acquisition module is configured to collect and pre-process real Brillouin gain spectrum data using a pre-built Brillouin optical time domain analysis system; The temperature extraction module is configured to obtain the corresponding Brillouin temperature based on the preprocessed Brillouin gain spectrum data using a trained self-attention recurrent neural network model; the self-attention recurrent neural network model includes an input layer, a hidden layer, a self-attention layer, and an output layer; a highway layer is added to the hidden layer, and the highway layer includes two gating units, a transfer gate and a conversion gate; After the self-attention layer obtains the output of the hidden layer, it allocates attention weights by calculating the similarity between each position of the input sequence, and weights the important features in the input sequence to obtain the output of the self-attention layer; the output of the output layer is the result of the fusion of the output of the self-attention layer and the output of the hidden layer; The training process of the self-attention recurrent neural network model is specifically as follows: Simulate the ideal Brillouin gain spectrum data and convert it into an input sequence to obtain a sample data set; Set the initial hidden state and use the meta-learning algorithm to optimize the task-adaptive weights of the transfer gate and the transformation gate at each time step; The optimized transition gate is used to determine the transition gate output of each layer at each time step, and the optimized transfer gate is used to determine the candidate hidden state of each layer at each time step; Based on the hidden state of the previous time step, the output of the conversion gate of the current time step, and the candidate hidden state, the hidden state in the current time step is updated as the final output of the model, i.e., the Brillouin temperature; The task adaptive weights and biases of the transfer gate and the conversion gate are adjusted by iterative optimization, and training is performed with the goal of minimizing the loss to obtain a trained self-attention recurrent neural network model.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the Brillouin temperature extraction method based on a self-attention recurrent neural network as described in any one of claims 1 to 4 are implemented.
7. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the Brillouin temperature extraction method based on a self-attention recurrent neural network as described in any one of claims 1 to 4 are implemented.
8. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the steps in the Brillouin temperature extraction method based on a self-attention recurrent neural network as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Brillouin frequency shift extraction method and device and storable medium
CN118152891A
Method and apparatus for implementing at least a first recurrent unit of a recurrent optical neural network
WO2024156808A1