Method for training machine learning model

Through a two-step training method, the general model is trained using label-free fiber sensing data and combined with a small amount of real-value data to adapt, the model accuracy and applicability problems in fiber sensing data are solved, and efficient model training and task adaptation are achieved.

CN120283240APending Publication Date: 2025-07-08SENSONIC GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081218.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-12-15
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, machine learning models have low accuracy due to the limited amount of real value data in fiber sensing data training, and the differences between different fiber sensing systems lead to limited model applicability.

Method used

A two-step training method is adopted, first using label-free fiber sensing data to train a general machine learning model, and then through a combination of transfer learning and supervised/unsupervised learning, a small amount of real-value data is used to adapt the model to perform specific tasks.

Benefits of technology

It improves the accuracy and applicability of machine learning models in fiber-optic sensing data, reduces the computational cost of training different tasks, and enhances the flexibility and versatility of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283240A_ABST
    Figure CN120283240A_ABST
Patent Text Reader

Abstract

A machine learning method in which, in a first step, a first (generic) machine learning model is trained using a first training data set comprising tagless fiber optic sensing data. Then, in a second step, a transfer learning process is applied to adapt or fine-tune the first machine learning model to a more specific application (e.g., to perform a particular type of detection or classification). Since a large amount of fiber optic sensing data is available, the first machine learning model may provide a generic machine learning model that has a high level of versatility and is highly adaptable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for training a machine learning model for use with data representing fiber optic sensing measurements, such as distributed acoustic sensing measurements. Background Art

[0002] Fiber optic based sensing systems typically use a laser source that emits laser pulses into a length of optical fiber. The light scattered within the fiber from the pulses is then detected at an optical detector such that the scattered light can be analyzed to determine various properties of the optical fiber. In particular, properties such as strain and / or temperature at various positions along the optical fiber can be determined based on the received scattered light.

[0003] For example, distributed acoustic sensing (DAS) enables the detection of an acoustic field (e.g., vibrations) in the environment of an optical fiber, where the optical fiber acts as a distributed transducer for the measurement. More specifically, the acoustic field can modulate the strain along the length of the optical fiber, thereby creating a modulation of the length and refractive index along the optical fiber, which in turn affects the scattering of the laser pulses along the optical fiber. Thus, the scattered light provides a measurement of the acoustic field in the environment of the optical fiber. To perform a DAS measurement, laser pulses can be launched along the optical fiber and then the scattered (e.g., backscattered) signal is recorded as a function of time. The reception time of the scattered signal enables the determination of the position of the corresponding scattering site in the optical fiber (e.g., by comparing the reception time of the scattered signal with the time when the laser pulse was launched into the optical fiber), such that the strain at the scattering site can be estimated. In this way, by observing the scattered signals received over time, the strain and thus the acoustic field can be monitored along the entire length of the optical fiber. Such a method is referred to as the optical time domain reflectometer (OTDR) method. In practice, optical fibers with lengths of 10 km to 50 km or longer can be monitored using DAS and OTDR techniques.

[0004] While DAS measurements rely on Rayleigh scattering at the scattering sites of laser pulses in an optical fiber, other scattering mechanisms can also be used to perform fiber optic sensing measurements. For example, Raman scattering of laser light in an optical fiber can be measured to determine the temperature along the length of the optical fiber using distributed temperature sensing (DTS) techniques. As another example, Brillouin scattering of laser light in an optical fiber can be measured to determine strain and / or temperature using distributed strain and temperature sensing (DSTS) techniques. An overview of known fiber optic sensing techniques is provided in the book "An Introduction to Distributed Optical Fibre Sensors" by A.H. Hartog (CRC Press, 2017).

[0005] Given the ability of fiber optic sensing technologies to continuously monitor optical fibers over significant distances, these technologies have been applied across a range of different fields. For example, fiber optic sensing systems can be used to monitor train tracks, for instance, in cases where an optical fiber is laid along the length of a train track. Other possible applications include monitoring pipelines, power lines, or road traffic, among others. Machine learning techniques and classification algorithms can be applied to the measurement data collected by fiber optic sensing systems in order to classify the data and / or enable event detection. As an example, in the case of using a fiber optic system to monitor a train track, the data obtained from the fiber optic sensing system can be analyzed to detect whether a train derailment or a defect in the track has occurred. As another example, in the case of using a fiber optic sensing system to monitor a pipeline, the data from the fiber optic sensing system can be analyzed to detect whether a leak has occurred in the pipeline.

[0006] The present invention has been designed in view of the above considerations. Summary of the Invention

[0007] The inventors have found that a difficulty with applying machine learning techniques to fiber optic sensing data is that there may typically be only a relatively small amount of ground truth data available for training a machine learning model. For example, in the case of using a fiber optic sensing system to monitor a train track, a user may wish to use a machine learning model to detect when a train derailment or a defect in the track occurs based on measurement data received from the fiber optic sensing system. However, there may be only a relatively small amount of measurement data corresponding to the actual occurrence of a train derailment or a track defect available for use in training the machine learning model. Additionally, differences between different train tracks (e.g., caused by different site conditions (such as different soil conditions), different track structures, different trains, and / or different fiber optic laying conditions) can mean that a model developed for one set of train tracks may not be applicable to another set of train tracks. Thus, machine learning models trained using conventional techniques for use with fiber optic sensing data can have a relatively low accuracy due to the limited amount of ground truth data available for training.

[0008] Although there may be only a limited amount of ground truth data available for use in machine learning, fiber optic systems typically generate large amounts of measurement data due to their use in continuously monitoring large lengths of optical fiber. For example, some fiber optic sensing systems can routinely collect approximately 100 megabytes of measurement data per second, but most of this collected data is currently not used for machine learning purposes.

[0009] The present invention is based on the recognition that a large amount of untagged data obtainable from a fiber optic sensing system can be used in training a machine learning model to improve the accuracy of the machine learning model. Most generally, the present invention provides a machine learning method in which, in a first step, a first (general) machine learning model is trained using a first training data set comprising untagged fiber optic sensing data. Then, in a second step, a transfer learning process is applied to adapt or fine-tune the first machine learning model to a more specific application (e.g., to perform a specific type of detection or classification). Since a large amount of fiber optic sensing data is available, the first machine learning model can provide a general machine learning model that has a high level of generality and is highly adaptable. Thus, for example, once the first machine learning model is trained using an untagged training data set, a relatively small amount of ground truth data can be used to adapt the first machine learning model to make the model suitable for a specific task. The inventors have found that, for example, such a two-step training approach can produce a machine learning model with higher accuracy compared to machine learning methods that mainly rely on a limited amount of ground truth data.

[0010] According to a first aspect of the present invention, there is provided a method of training a machine learning model for use with data representative of fiber optic sensing measurements, the method comprising: training a first machine learning model using a first training data set, wherein the first training data set comprises data representative of fiber optic sensing measurements and wherein the first training data set is untagged; and training a second machine learning model using a second training data set, the second machine learning model comprising at least a portion of the trained first machine learning model, wherein the second training data set comprises data representative of fiber optic sensing measurements, and the size of the second training data set is less than or equal to the first training data set; wherein the second machine learning model is configured to generate an output based on received data representative of fiber optic sensing measurements.

[0011] A machine learning model trained using the method of the present invention is configured for use with data representative of fiber optic sensing measurements. Herein, the fiber optic sensing measurements can correspond to any type of optical sensing measurement performed on an optical fiber. The fiber optic sensing measurements can be distributed sensing measurements. Examples of fiber optic sensing measurements that can be used with the method of the present invention include (but are not limited to) OTDR measurements, DAS measurements, DTS measurements, and / or DSTS measurements.

[0012] As used herein, data representing fiber optic sensing measurements can refer to any type of data that can be obtained from fiber optic sensing measurements. For example, the data can include information related to a detected signal (i.e., a scattered signal from the optical fiber detected during the measurement), such as the amplitude of the signal, the time at which the signal was detected, the location in the optical fiber of the scattering site corresponding to the detected signal, the frequency of the signal, the phase of the signal, and / or the polarization of the signal. The data can additionally or alternatively include information about the optical fiber obtained from the detected signal, such as strain and / or temperature within the optical fiber, e.g., as a function of time and / or position along the optical fiber.

[0013] A first machine learning model is trained using a first training dataset that is unlabeled. In other words, the first training dataset does not include any labels for training the first machine learning model. Thus, the data representing fiber optic sensing measurements in the first training dataset can not include any labels that categorize or classify the data for training purposes. In this way, the first training dataset can be established by collecting data from one or more fiber optic sensing measurement systems without the user having to label the data. This enables the use of a large amount of data that can be obtained from fiber optic sensing measurement systems, as constructing the first training dataset does not require a time-intensive labeling process.

[0014] The data in the first training dataset can correspond to (raw) data obtained directly from one or more fiber measurement systems. In some cases, the data received from one or more fiber measurement systems can undergo data preparation steps (e.g., preprocessing steps) before being included in the first training dataset, e.g., to put the data in a format suitable for the training process.

[0015] Thus, the method can further include the steps of receiving data representing fiber optic sensing measurements from one or more fiber optic sensing systems and including the received data in the first training dataset.

[0016] The first machine learning model can be trained using any suitable machine learning technique that takes unlabeled data as input. For example, as discussed in more detail below, the first machine learning model can be trained using self-supervised and / or unsupervised machine learning techniques. The first machine learning model can be any suitable type of machine learning model, such as an artificial neural network (ANN), e.g., a convolutional neural network (CNN) and / or a transformer network.

[0017] A second machine learning model is then trained using a second training dataset. The second machine learning model includes at least a portion of the trained first machine learning model. Thus, the training of the second machine learning model can be performed after the training of the first machine learning model is completed.

[0018] The second machine learning model including at least a part of the trained first machine learning model can mean that part or all of the first machine learning model is included in the second machine learning model. For example, one or more layers of the trained first machine model can be included in the second machine learning model. Thus, the features learned by the first machine learning model using the first training dataset are included in the second machine learning model, which is then trained using the second training dataset.

[0019] At least a part of the first machine learning model can be incorporated into the second machine learning model in various ways. For example, in some cases, the second machine learning model can include one or more (possibly all) layers of the first machine learning model, with one or more further layers added to the second machine learning model. One or more further layers can then take the output from one or more layers of the first machine learning model as input. Alternatively, all or part of the first machine learning model can be used as the second machine learning model, where training of the second machine learning model results in adjusted parameters of the first machine learning model.

[0020] The second training dataset includes data representing fiber optic sensing measurements and has a size less than or equal to that of the first training dataset. In other words, the number of data elements in the second training dataset can be less than or equal to the number of data elements in the first training dataset.

[0021] The second training dataset can be different from the first training dataset. For example, compared to the first training dataset, the second training dataset can correspond to a different set of fiber optic sensing measurements. In some cases, the second training dataset can correspond to a subset of the data in the first training dataset. Alternatively, in some cases, the second training dataset can include the same data as the first training dataset.

[0022] The second training dataset can be adapted to the task to be performed by the second machine learning model. Thus, while the first training dataset is unlabeled and does not cater to any specific task, the second training dataset can be adapted to train the second machine learning model to perform a specific task. As an example, the second training dataset can include ground truth data (e.g., labeled data) for training the second machine learning model. For example, in the case where it is desired to train the second machine learning model as a classifier, the second training dataset can include labeled data for classifier training. Various examples of the types of the second training dataset are provided below.

[0023] The data in the first training dataset and the second dataset can be obtained from one or more fiber optic sensing systems. The data in the first training dataset and the second dataset can be obtained from the same and / or different fiber optic sensing systems. The fiber optic sensing systems used to collect the data in the first training dataset and the second training dataset can include any suitable type of fiber optic sensing system, e.g., capable of performing the fiber optic sensing measurements mentioned above. Additionally or alternatively, the first training dataset and / or the second training dataset can include simulated (e.g., computer-generated) fiber optic sensing data.

[0024] Each fiber optic sensing system in the one or more fiber optic sensing systems can be arranged to perform the same type of fiber optic sensing measurement. In other words, the data in the first training dataset and the data in the second training dataset can represent the same type of fiber optic sensing measurement.

[0025] Each fiber optic sensing system in the one or more fiber optic sensing systems can be arranged to monitor (observe) the same type of system or environment. Examples of the types of systems or environments that can be monitored using fiber optic sensing systems include (but are not limited to) train (railway) tracks, pipelines, roads, fences, boundaries (or perimeters), power transmission lines, and underwater infrastructure. Thus, for example, in the case of using multiple fiber optic sensing systems to collect the data in the first training dataset and / or the second training dataset, each fiber optic sensing system in the fiber optic sensing systems can be arranged to monitor (corresponding) a set of train tracks (or any other suitable type of system or environment).

[0026] The training of the second machine learning model can be performed using any suitable machine learning technique. For example, as discussed in more detail below, the second machine learning model can be trained using supervised and / or unsupervised machine learning techniques. The type of training used for the second machine learning model can be selected based on the task to be performed by the second machine learning model (e.g., event detection, event prediction, data classification, anomaly detection, etc.).

[0027] The second machine learning model is configured to generate an output based on the received data representing fiber optic sensing measurements. Thus, after training the second machine learning model, the second machine learning model can receive new (e.g., previously unseen) fiber optic sensing data as input and, in response, generate an output. The type of output generated by the second machine learning model can depend on the type of model and training used.

[0028] By way of example, the second machine learning model can be trained for event detection, in which case the second machine learning model can be configured to output an indication of whether an event has been detected based on the received data. As another example, the second machine learning model can be trained as a classifier, in which case the second machine learning model can be configured to output information indicative of a classification based on the received data. As a further example, the second machine learning model can be trained as an anomaly detector, in which case the second machine learning model can be configured to output an indication of whether an anomaly has been detected based on the received data.

[0029] As yet a further example, the second machine learning model can be configured to predict properties or variables of a system being monitored using an optical fiber system. For example, in the case of receiving optical fiber sensing data from an optical fiber sensing system arranged to monitor a set of train tracks, the machine learning model can be configured to predict the position and / or speed of a train on the train tracks based on the received data.

[0030] The second machine learning model is configured to generate an output based on the received data obtained from the first optical fiber sensing system, where the second training dataset includes data obtained from the first optical fiber sensing system. In this way, the second machine learning model can be specifically trained to work with data from the first optical fiber sensing system.

[0031] Based on the above discussion, training the first machine learning model using the unlabeled first training dataset enables the use of a large amount of optical fiber sensing data to train the first machine learning model. Thus, the first machine learning model can provide a general (generalized) model of the optical fiber sensing data. For example, as a result of being trained using the first training dataset, the first machine learning model may have learned a feature map in the latent space of the optical fiber sensing data. Then, the step of training the second machine learning model can act as a transfer learning step, where parts of the trained first machine learning model are reused in the context of a more specific task.

[0032] Advantageously, by reusing at least a part of the first machine learning model in the second machine learning model, effective training of the second machine learning model can still be achieved even with a relatively small second training dataset. This is because the second machine learning model includes features of the general first machine learning model trained using a potentially very large first training dataset, such that the insights into the optical fiber sensing data obtained by the first machine learning model can help improve the accuracy of the second machine learning model. Thus, the present invention enables effective training of a machine learning model for use with optical fiber sensing data even when scarce ground truth data is available.

[0033] A further benefit of the two-step training process of the present invention is that a general first machine learning model obtained by training using a first training dataset can be used to solve a variety of different problems by adapting a second training dataset for training a second machine learning model. For example, once the first machine learning model is trained, the first machine learning model can be used in a variety of different second machine learning models, which can be adapted to perform different specific tasks. In other words, multiple different second machine learning models can reuse the same first machine learning model, such that the first machine learning model only needs to be trained once. Thus, once the first machine learning model is trained, the computational cost of developing multiple different second machine learning models can be relatively low.

[0034] The method can include performing a first machine learning process for training a first machine learning model and a second machine learning process for training a second machine learning model, where the second machine learning process is different from the first machine learning process. In other words, different machine learning processes can be applied to train the first machine learning model and the second machine learning model. This can enable the first machine learning model to be trained as a general model for fiber optic sensing data, while the second machine learning model is trained for a specific task. The first machine learning process and the second machine learning process can, for example, rely on different training (learning) algorithms. The first machine learning process and the second machine learning process can correspond to different types (or techniques) of machine learning.

[0035] The second machine learning model can be trained using a supervised learning process, and the second training dataset can include labeled data used in the supervised learning process. In this way, the second machine learning model can be trained via supervised learning using the labeled data in the second training dataset for a specific task. The labeling of the data in the second dataset can be adapted to the type of task for which the second machine learning model is trained. For example, supervised learning can be used to train the second machine learning model as a classifier, an event detection model, an event prediction model, an anomaly detector, and / or a model for predicting properties or variables of a system being monitored using a fiber optic system. Any suitable type of supervised machine learning technique can be used to train the second machine learning model.

[0036] In the case of using supervised learning to train the second machine learning model, the second training dataset can be smaller than the first training dataset (e.g., having fewer data elements than the first training dataset). The labeled data in the second training dataset can be referred to as ground truth data. Advantageously, due to the larger first training dataset used to train the first machine learning model, training the second machine learning model may only require a relatively small number of labeled data elements.

[0037] The tagged data may include data elements that are tagged to indicate attributes of the data elements; and the output generated by the second machine learning model may indicate the attributes of the received data. In this way, the second machine learning model may learn to identify attributes in the fiber optic sensing data based on the tagged data, such that the second machine learning model may then determine or predict corresponding attributes in new (e.g., previously unseen) fiber optic sensing data.

[0038] The attributes may correspond to predefined conditions. Thus, each data element may be tagged to indicate whether the data element corresponds to a predefined condition (e.g., whether the predefined condition occurred when the data element was recorded). Then, the output generated by the second machine learning model may indicate the probability that the received data corresponds to the predefined condition. In other words, the second machine learning model may learn to identify and / or predict predefined conditions from the received fiber optic sensing data. The predefined conditions may correspond to any condition or event along the fiber optic (or along the system monitored by the fiber optic) that a user may wish to detect or predict. For example, in the example given above where a fiber optic sensing system is used to monitor a train track, the predefined conditions may correspond to a train derailment, a defect occurring in the train track, or a person or animal crossing the train track.

[0039] Alternatively, the attributes may correspond to values of predefined parameters of the fiber optic sensing system and / or the system monitored by the fiber optic sensing system. For example, each data element may be tagged with the value of a predefined parameter. In this way, the second machine learning model may be trained to predict the values of the predefined parameters based on the received fiber optic sensing data. For example, in the example given above where a fiber optic sensing system is used to monitor a train track, the attributes may correspond to the values of the speed and / or position of a train traveling along the train track.

[0040] The second machine learning model may be trained using an unsupervised learning process, and the second training dataset may include untagged data used in the unsupervised learning process. This may enable the second machine learning model to be trained to perform various tasks without the need for tagged data. For example, the second machine learning model may be trained as a clustering algorithm using any suitable clustering technique. An example application where the second machine learning model is configured as a clustering algorithm is where the second training dataset includes untagged data from a track infrastructure having a double-track line. Then, the second machine learning model may be trained to cluster the data into two separate clusters, e.g., one cluster per track.

[0041] In some cases, both supervised and unsupervised processes can be combined for training a second machine learning model. In such cases, the second training dataset can include labeled data elements (used in supervised learning) and unlabeled data elements (used in unsupervised learning).

[0042] The second machine learning model can be configured as an anomaly detector. This can, for example, enable the second machine learning model to detect when a system monitored by an optical fiber sensing system is not performing as expected. The anomaly detector can be trained using supervised learning with the labeled data elements in the second training dataset. For example, the data elements can be labeled to indicate whether they correspond to normal behavior of the system or whether they correspond to anomalous (unconventional) behavior of the system. Alternatively, the anomaly detector can be trained using unsupervised training with the unlabeled data elements in the second training dataset. For example, the anomaly detector can be trained to determine what constitutes normal and anomalous behavior of the system.

[0043] The first training dataset can be collected using two or more different optical fiber sensing systems. In other words, the data in the first training dataset can be obtained from multiple different physical optical fiber sensing systems. This can be used to increase the level of generality of the first machine learning model, as the data in the first training dataset is not limited to a single physical optical fiber sensing system. This can make the first machine learning model more flexible and adaptable, thereby facilitating its use in a wider range of scenarios. In particular, this can enable the use of data from multiple different systems to train a single general first machine learning model, which can then be used to solve problems using input data obtained from different physical optical fiber sensing systems. For example, there can be variations across different optical fiber sensing systems, which can affect the data generated by different optical fiber sensing systems. Thus, including data from multiple optical fiber sensing systems in the first training dataset can enable the first machine learning model to learn and account for the differences between different optical fiber sensing systems. As an example, this can facilitate the inclusion of the same first machine learning model into multiple different second machine learning models, each of the multiple different machine learning models being configured for use with data received from a different respective optical fiber sensing system. More generally, increasing the number of optical fiber sensing systems contributing to the first training dataset can increase the level of generality of the first machine learning model, which in turn enhances its flexibility and adaptability.

[0044] Herein, an optical fiber sensing system can also be referred to as a measurement system.

[0045] Two or more different fiber optic sensing systems can be located at different geographical locations. This can be used to further improve the level of generality of the first machine learning model, which can enable the first machine learning model to process data received from a greater variety of fiber optic sensing systems. By including data obtained from different geographical locations in the first training dataset, the first machine learning model can learn and account for variations across different geographical locations. This can facilitate the use of the same first machine learning model in the context of data received from systems located at different geographical locations around the world. In particular, the inventors have found that for many applications, data obtained from fiber optic sensing systems located at different geographical locations can vary, which may be due to differences in local conditions. For example, in the case of using fiber optic sensing systems to monitor train tracks at different geographical locations, differences in soil conditions, fiber optic laying conditions, trains used on the tracks, and other factors can contribute to variations in the fiber optic sensing data obtained at different geographical locations.

[0046] Compared to the first training dataset, the second training dataset can be collected using fewer fiber optic sensing systems. In this way, the first machine learning model can be trained to have a greater level of generality, while the training of the second machine learning model can be more specific to the particular fiber optic sensing system of interest. This can help improve the accuracy of the second machine learning model, which can benefit from the features of the first machine learning model that has learned based on a larger number of fiber optic sensing systems. Additionally, based on the discussion above, using data from a larger number of fiber optic sensing systems to train the first machine learning model can enable effective training even when there is scarce training data available for the particular fiber optic sensing system of interest. In other words, the second machine learning model can apply the features learned by the first machine learning model from a larger number of fiber optic sensing systems in the context of the particular fiber optic sensing system of interest.

[0047] In the case where the first training dataset has been collected using two or more fiber optic sensing systems located at different geographical locations, the second training dataset can be collected using fiber optic sensing systems located at a smaller number of geographical locations.

[0048] Data in the first training dataset can be obtained from a first group of fiber optic sensing systems, and data in the second training dataset can be obtained from a second group of fiber optic sensing systems, where the first group is larger than the second group. The first group of fiber optic sensing systems can include one or more of the fiber optic sensing systems in the second group. In other words, the first training dataset and the second training dataset can include data from one or more of the same fiber optic sensing systems. Alternatively, the first group of fiber optic sensing systems may not include any of the fiber optic sensing systems from the first group, i.e., the first training dataset and the second training dataset can be obtained from completely different groups of fiber optic sensing systems.

[0049] In an alternative scenario, data from the same number of fiber optic sensing systems can be used in the first training dataset and the second training dataset. For example, the same fiber optic sensing systems can be used for both the first training dataset and the second training dataset. For example, this may be the case when the first training dataset and the second training dataset have the same size.

[0050] The second machine learning model can include a trained first machine learning model and a decision model, where the decision model is connected to the output of the first machine learning model. In other words, the decision model can be added to the trained first machine learning model to form the second machine learning model. Such a structure can facilitate adapting the first machine learning model to a desired task because an appropriate decision model can be added to the first machine learning model without having to otherwise modify the first machine learning model. Thus, this can provide a flexible and adaptable way to train the second machine learning model. The decision model can be arranged to receive (as its input) one or more outputs from the first machine learning model (e.g., from one or more layers of the first machine learning model) and generate an output based on the received one or more inputs. For example, the first machine learning model can be configured to map input fiber optic sensing data to an output in a latent space. Then, the decision model can be configured to receive the output in the latent space and map the output in the latent space to a desired output (e.g., classification, event detection, event prediction, anomaly detection, etc.).

[0051] The decision model can be any suitable type of machine learning model. Examples of suitable types of machine learning models include artificial neural networks (ANNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), or transformer networks. Other types of machine learning models can also be used, such as support vector machines (SVMs), logistic regression models, or k-means clustering models.

[0052] Once the decision model has been connected to the output of the trained first machine learning model to form a second machine learning model, the second training dataset can be used to train the second machine learning model as discussed above. In such a case, training the second machine learning model can involve adjusting the parameters of the first machine learning model and / or the decision model based on the second training dataset.

[0053] The second machine learning model can include the trained first machine learning model and two or more decision models, each of the two or more decision models being connected to the output of the first machine learning model, and each of the two or more decision models can be trained using a corresponding second training dataset. In other words, the trained first machine learning model can be combined with two or more decision models, each of the two or more decision models being trainable using a corresponding training dataset. Each decision model can be configured to perform a respective task, i.e., each decision model can be configured to generate a respective output (e.g., classification, event detection, event prediction, anomaly detection, etc.) in response to receiving an input from the first machine learning model. Such a structure can leverage the level of generality of the first machine learning model, enabling a wide variety of tasks to be performed using different decision models connected to the same first machine learning model. Additionally, since the same first machine learning model is used with multiple decision models, the first machine learning model may only need to be trained once, thereby reducing the amount of time and computational resources required for training (e.g., as compared to a scenario where a completely separate model is trained for each task). This can also avoid having to replicate the first machine learning model to perform different tasks, as a single instance of the first machine learning model can provide its output to multiple different decision models.

[0054] The first machine learning model can include multiple outputs (e.g., from multiple layers in the first machine learning model). Each of the two or more decision models can be connected to one or more of the multiple outputs from the first machine learning model. In some cases, the two or more decision models can be connected to different (respective) inputs from the first machine learning model. In other cases, the two or more decision models can be connected to the same output from the first machine learning model.

[0055] The respective second training datasets for each of two or more decision models can be adapted to the tasks to be performed by that decision model. Any suitable training process can be performed to train each of the two or more decision models. Each of the two or more decision models can be trained independently, i.e., such that the training of a given decision model using its respective second training dataset does not affect the parameters (features) in any of the other two or more decision models.

[0056] Each of the two or more decision models can be in the form of a suitable type of machine learning model as described above, such as an artificial neural network (ANN), a convolutional neural network (CNN), a recurrent neural network (RNN), and / or a transformer network. Other types of machine learning models can also be used, such as a support vector machine (SVM), a logistic regression model, or a k-means clustering model.

[0057] During the training of the second machine learning model, the trained first machine learning model can be kept fixed. In other words, during the training of the second machine learning model, the parameters (e.g., features, coefficients) of the first machine learning model can be kept fixed (unchanged). For example, only the parameters in the decision model connected to the first machine learning model can be adjusted during the training of the second machine learning model, while the parameters in the first machine learning model remain fixed. This can simplify the training of the second machine learning model by avoiding the retraining of the first machine learning model, which is typically larger and more complex than any decision model connected to the first machine learning model. Once the first machine learning model has been trained, this can also facilitate the reuse of the first machine learning model, as the level of generality of the first machine learning model can still be maintained even after the second machine learning model has been trained.

[0058] Alternatively, during the training of the second machine learning model, one or more parameters of the trained first machine learning model can be adjusted. In this way, the parameters (e.g., features, coefficients, weights, and / or biases) of the first machine learning model can be adapted based on the second training dataset. This can enable the relatively general first machine learning model to be specifically adapted to a particular task using the second training dataset. This can help improve the accuracy of the second machine learning model. In such cases where the parameters of the first machine learning model are adjusted, the second machine learning model can consist of at least a portion of the first machine learning model, and the second machine learning model is adjusted as part of the training using the second training dataset. In other ways, similar to the example above, the second machine learning model can include the first machine learning model connected to a decision model, where the parameters of both the first machine learning model and the decision model are adjusted during the training using the second training dataset.

[0059] The first machine learning model can be trained using the first training dataset via a self-supervised learning process. Self-supervised learning can enable the effective training of the first machine learning model using the unlabeled first training dataset and can provide the first machine learning model with a detailed understanding of the data.

[0060] The self-supervised learning process can include partially masking the data in the first training dataset and training the first machine learning model to reconstruct the data in the first training dataset based on the partially masked data. The inventors have found that this technique can provide an effective and reliable training for the first machine learning model. In particular, by training the first machine learning model to reconstruct (generate) the partially masked data, the first machine learning model can effectively learn to identify and predict the characteristics of the fiber optic sensing data. Thus, the first machine learning model is trained to infer the missing information based on the context provided by the remaining (i.e., unmasked) data, thereby enhancing the ability of the first machine learning model to handle incomplete or noisy datasets (including real-world fiber optic sensing data that typically can have a low signal-to-noise ratio).

[0061] The first machine learning model can include the encoder portion of an autoencoder (or a masked autoencoder). Thus, during the self-supervised learning process, the masked autoencoder can be trained using the first training dataset. The encoder portion of the masked autoencoder can then be used as the first machine learning model.

[0062] Partial masking of the data in the first training dataset can, for example, involve applying one or more randomly generated masks to the data. The self-supervised learning process can be an iterative process in which the machine learning model is iteratively adjusted to reduce (minimize) the difference between the data reconstructed by the first machine learning model and the corresponding data in the first training dataset. For example, each iteration can involve randomly generating one or more masks, applying the one or more masks to the data in the first training dataset, using the first machine learning model to reconstruct the data masked by the one or more masks, and comparing the reconstructed data with the corresponding original (unmasked) data in the first training dataset.

[0063] The self-supervised learning process can include providing the unmasked portion of the first dataset to the first machine learning model and training the first machine learning model to reconstruct (i.e., generate) the data in the first training dataset based on the unmasked portion. In other words, using only the unmasked portion of the first training dataset, the first machine learning model can reconstruct the entire dataset, including the masked portion. Here, the unmasked portion of the first training dataset can correspond to the portion of the first training dataset that is not covered by the mask - i.e., remains visible after the partial masking of the data in the first training dataset. For example, compared to a conventional autoencoder method that provides the entire dataset as input, providing only the unmasked portion of the first training dataset to the machine learning model can enhance the computational efficiency of self-supervised learning. For example, only a relatively small unmasked portion (subset) of the first training dataset (e.g., 25%) can be used, which can result in a significant reduction in training time and computational resources. This self-supervised learning that uses only the unmasked portion of the first training dataset can be implemented by a masked autoencoder.

[0064] As an example, 50% or more of the data in the first training dataset can be masked. Thus, the unmasked portion of the first training dataset (for training the first machine learning model) can represent 50% or less of the data in the first training dataset. As a further example, 75% or more of the data in the first training dataset can be masked, i.e., 25% or less of the data in the first training dataset is unmasked.

[0065] The first machine learning model can include an encoder (e.g., the encoder portion of a masked autoencoder) configured to take the unmasked portion of the first training dataset as input and generate a representation of the input in the latent space. The latent space can have an equal or higher dimension compared to the input. In contrast to conventional autoencoder methods that typically compress the input into a lower-dimensional latent space, using a latent space with an equal or higher dimension compared to the input facilitates the generation of the missing (i.e., masked) portion of the first training dataset. In some cases, the latent space can have a lower dimension compared to the input.

[0066] The self-supervised learning process may further include reconstructing the data in the first training dataset using a decoder configured to decode the latent space representation generated by the encoder. The encoder and the decoder may together form an autoencoder, e.g., a masked autoencoder. According to the above, the encoder may be configured to operate only on the unmasked portion of the first training dataset. In contrast, the decoder may take as input the latent space representation generated by the autoencoder and information indicating the masked portion of the first training dataset, such that the decoder may reconstruct the data from the first training dataset.

[0067] The encoder and the decoder may have an asymmetric architecture, where the encoder has a higher complexity (e.g., is larger) compared to the decoder. For example, the first machine learning model (encoder) may have a larger number of parameters compared to the decoder. This is in contrast to a conventional encoder-decoder architecture that is symmetric in nature. This architecture allows the masked portion to be excluded from the input provided to the encoder, which may alternatively be handled by a more lightweight decoder.

[0068] The first training dataset may further include data indicating the environmental conditions associated with the recording (collection) of the data in the first training dataset, and / or the second training dataset may further include data indicating the environmental conditions associated with the recording (collection) of the data in the second training dataset. The use of the data indicating the environmental conditions in the first training dataset may enable the first machine learning model to correlate the features in the fiber optic sensing data with the environmental conditions, which may improve the accuracy of the first machine learning model. Similarly, the use of the data indicating the environmental conditions in the second training dataset may enable the second machine learning model to correlate the features in the fiber optic sensing data with the environmental conditions, which may improve the accuracy of the second machine learning model. As an example, each entry of the fiber optic sensing data in the first training dataset and / or the second training dataset may be associated with corresponding environmental data indicating the environmental conditions at the time and location corresponding to the recording of the fiber optic sensing data.

[0069] The data indicating the environmental conditions may relate to any type of environmental conditions. Examples of the types of environmental conditions include weather conditions (e.g., temperature, cloud cover, humidity, wind speed, etc.), date, time of day, geographical location, etc. The data indicating the environmental conditions may be obtained using any suitable sensor, which may be separate from the fiber optic sensing system used to obtain the fiber optic sensing data.

[0070] In cases where data indicating environmental conditions is used to train the first machine learning model and / or the second machine learning model, the second machine learning model can be configured to generate an output based on the received data representing the fiber optic measurement and the data representing the environmental conditions associated with the fiber optic sensing measurement. In this way, the second machine learning model can take into account the environmental conditions, which can improve the accuracy of the output of the second machine learning model.

[0071] According to a second aspect of the present invention, there is provided a method of analyzing data representing a fiber optic sensing measurement, the method comprising: receiving, by a machine learning model, data representing a fiber optic sensing measurement, wherein the machine learning model has been trained using the method according to any one of the preceding claims; and generating, in response to receiving the data, an output by the machine learning model.

[0072] The method of the second aspect of the present invention corresponds to the use of a machine learning model that has been trained according to the method of the first aspect of the present invention. Accordingly, any features described above with respect to the first aspect may be shared with the second aspect.

[0073] According to the discussion above with respect to the first aspect, the output may include one or more of classification, event detection, event prediction, anomaly detection, etc.

[0074] In some cases, the output generated by the machine learning model may indicate the probability that the received data corresponds to a predetermined condition, and / or the output generated by the machine learning model may indicate an attribute of the data.

[0075] According to a third aspect of the present invention, there is provided a computer-implemented system for training a machine learning model, the system being configured to: train a first machine learning model using a first training dataset, wherein the first training dataset includes data representing a fiber optic sensing measurement, and wherein the first training dataset is unlabeled; and train a second machine learning model using a second training dataset, the second machine learning model comprising at least a portion of the trained first machine learning model, wherein the second training dataset includes data representing a fiber optic sensing measurement, the size of the second training dataset is less than or equal to the first training dataset, and wherein the second machine learning model is configured to generate an output based on the received data representing the fiber optic sensing measurement.

[0076] The system of the third aspect of the present invention can be used to perform the method of the first aspect of the present invention. Accordingly, any of the features described above with respect to the first aspect may be shared with the third aspect.

[0077] A computer-implemented system may include any suitable computing device or network of computing devices configured to perform the steps. The computer-implemented system may include a storage medium storing executable instructions and a processing system configured to execute the instructions, wherein execution of the instructions causes the processing system to perform the steps (i.e., the steps of the method of the first aspect).

[0078] In a more general aspect, the present invention may provide a computer storage medium storing executable instructions, wherein when the instructions are executed by a processing system, the instructions are arranged to cause the processing system to perform the steps of the method of the first aspect.

[0079] According to a fourth aspect of the present invention, there is provided a computer-implemented system for analyzing data representative of fiber optic sensing measurements, the system including a machine learning model trained using the method of the first aspect of the present invention.

[0080] The computer-implemented system of the fourth aspect of the present invention may be used to perform the method of the second aspect of the present invention. Accordingly, any features described above with respect to the first aspect and / or the second aspect may be shared with the fourth aspect of the present invention.

[0081] The computer-implemented system of the fourth aspect may include a storage medium in which the machine learning model is stored. The storage medium may be implemented using one or more of physical storage devices (e.g., an array).

[0082] The present invention includes combinations of the described aspects and preferred features, unless such combinations are clearly impermissible or expressly avoided. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Embodiments illustrating the principles of the present invention will now be discussed with reference to the drawings, in which:

[0084] Figure 1 is a diagram showing a method according to an embodiment of the present invention;

[0085] Figure 2a is a schematic diagram of a system according to an embodiment of the present invention;

[0086] Figure 2b is a schematic diagram of a system according to an embodiment of the present invention;

[0087] Figure 3a is a schematic diagram of an optical fiber system that may be used with the system of the present invention;

[0088] Figure 3b shows an example of data that may be obtained using Figure 3a the system of;

[0089] Figure 3cshows another example of data that can be obtained using Figure 3a a system;

[0090] Figure 4a is a diagram showing an example process for training a first machine learning model in an embodiment of the present invention;

[0091] Figure 4b shows an example of data corresponding to different steps in the Figure 4a method;

[0092] Figure 4c is a diagram showing an example process for training a first machine learning model in an embodiment of the present invention;

[0093] Figure 5 is a diagram showing an example process for training a second machine learning model in an embodiment of the present invention; and

[0094] Figure 6 is a diagram showing an example process for training a second machine learning model in an embodiment of the present invention. DETAILED DESCRIPTION

[0095] Aspects and embodiments of the present invention will now be discussed with reference to the accompanying drawings. Those skilled in the art will appreciate other aspects and embodiments.

[0096] Figure 1 shows a method 100 for training a machine learning model according to an embodiment of the present invention, while Figure 2a shows a schematic diagram of a system 200 that can be used to implement method 100. For illustrative purposes, method 100 will be described in the context of system 200, but method 100 is not limited to being implemented using system 200.

[0097] In a first step 102, method 100 involves training a first machine learning model using a first training dataset, where the first training dataset includes data representing fiber optic sensing measurements and where the first training dataset is unlabeled. Moving on to Figure 2a, the fiber optic sensing data in the first training dataset can be obtained using one or more fiber optic sensing systems 202a - 202c. In the example shown, the fiber optic sensing data is collected from three separate fiber optic sensing systems 202a - 202c. However, other implementations can use data from a different number of systems. Each of the fiber optic sensing systems 202a - 202c is configured to perform the same type of fiber optic sensing measurement, e.g., such as DAS, DTS, or DSTS. Additionally, each of the fiber optic sensing systems 202a - 202c is arranged in the same type of system or environment. In this way, the fiber optic sensing data collected from systems 202a - 202c can correspond to the same type of data. For example, each of the fiber optic sensing systems 202a - 202c can be arranged to monitor a corresponding set of train tracks, i.e., each of the fiber optic sensing systems 202a - 202c can include an optical fiber extending along the length of the corresponding track. Of course, the present invention is not limited to monitoring train tracks, and other types of systems or environments that can be monitored using systems 202a - 202c include pipelines, roads, fences, boundaries (or perimeters), power transmission lines, and underwater infrastructure. Thus, each of the fiber optic sensing systems 202a - 202c can be located at different geographical locations, i.e., corresponding to the respective systems or environments monitored by the fiber optic sensing systems 202a - 202c.

[0098] The fiber optic sensing data received from the fiber optic sensing systems 202a - 202c can correspond to raw measurement data and can take various forms depending on the specific type of measurement performed using systems 202a - 202c. For example, the data received from each of the systems 202a - 202c contains information related to the signal detected by that system, such as the amplitude of the signal, the time at which the signal was detected, the position in the optical fiber of the scattering site corresponding to the detected signal, the frequency of the signal, the phase of the signal, and / or the polarization of the signal. The data can additionally or alternatively include information about the optical fiber obtained from the detected signal, such as strain and / or temperature within the optical fiber, e.g., as a function of time and / or position along the optical fiber. Examples of fiber optic sensing systems and fiber optic sensing data are described in more detail below with respect to Figure 3a and Figure 3b The system 200 can be configured to store the data received from the fiber optic sensing systems 202a - 202c in the first database 204. The first database 204 can be implemented using any suitable storage system.

[0099] The raw fiber optic sensing data in the first database 204 can then be provided to a first data preparation (e.g., preprocessing) module 206, which is configured to prepare the fiber optic sensing data for use in training a first machine learning model. Depending on the particular training process and machine learning model to be used, this can include a variety of different steps. In some cases, the first data preparation module 206 can be configured to sample, for example, the raw fiber optic sensing data from the fiber optic sensing systems 202a - 202c at a predetermined sampling rate to reduce the amount of processing power required to process the data. The first data preparation module 206 can also be configured to convert the raw data in the first database 204 into a predetermined format (or representation), where the predetermined format is adapted to the input of the first machine learning model. Other types of preprocessing (such as noise reduction and / or smoothing algorithms) can also be applied to the raw data. The prepared data output by the first data preparation module 206 can be stored in the first training database 208. In other words, the first training database 208 can store the data prepared by the first data preparation module 206 based on the raw data received from the fiber optic sensing systems 202a - 202c. The data stored in the first training database 208 can be considered the first training dataset. Thus, the first training dataset can include a plurality of data elements derived from the raw data obtained from the fiber optic sensing systems 202a - 202c, which represent the fiber optic sensing measurements performed by the systems 202a - 202c. The first training dataset is unlabeled, such that no ground truth labeling needs to be performed on the data in the first training dataset.

[0100] The system 200 further includes a first training module 210, which is configured to train a first machine learning model 212 using the unlabeled first training dataset from the first training database 208. The first machine learning model 212 can be any suitable type of machine learning model, e.g., an artificial neural network (ANN), such as a convolutional neural network (CNN) or a transformer network. The first training module 210 can be configured to train the first machine learning model 212 using any suitable machine learning process that can take the unlabeled first dataset from the first training database 208 as input. For example, the training process performed by the first training module 210 can involve self-supervised or unsupervised learning techniques. An example of a self-supervised learning technique that can be used to train the first machine learning model is described below with respect to FIG. 4.

[0101] The trained first machine learning model 212 generated from the training process performed by the first training module 210 (i.e., generated from step 102 of method 100) can be considered a general model, which can have a relatively high level of generality due to the use of fiber optic sensing data from several different fiber optic sensing systems 202a - 202c that can be located at several different geographical locations (positions). Additionally, since the training of the first machine learning model 212 does not require labeling of the fiber optic sensing data, the first training dataset can include a relatively large amount of data, which can be used to enhance the generality level of the model. The generality level of the first machine learning model can be increased by increasing the amount of data included in the first training dataset - for example, by using data collected over a longer time period and / or by using data collected from a larger number of fiber optic sensing systems.

[0102] Returning to Figure 1 , in the second step 104 of method 100, a second machine learning model is trained. The second machine learning model includes the trained first machine learning model, i.e., the first machine learning model trained in step 102 (e.g., the first machine learning model 212 trained by the first training module 210). A second training dataset is used to train the second machine learning model. Depending on the type of training to be performed, the second training dataset can be labeled or unlabeled. In Figure 2a the example shown, the fiber optic sensing system 202d is used to obtain the second training dataset. The fiber optic sensing system 202d is configured to perform the same type of measurements as the fiber optic sensing systems 202a - 202c and is arranged to monitor the same type of system or environment as systems 202a - 202c. In some cases, the fiber optic sensing system 202d can correspond to one of the systems 202a - 202c used for the first training dataset, i.e., the same fiber optic sensing system can contribute to both the first training dataset and the second training dataset. Alternatively, the fiber optic sensing system 202d used for the second training dataset may not be included in the fiber optic sensing systems 202a - 202c used for the first training dataset. Although Figure 2a a single fiber optic sensing system 202d is shown for collecting the second training dataset, other embodiments may use a different number of systems to collect the second training dataset. The number of systems used to collect the second training dataset can generally be less than the number of systems used to collect the first training dataset, but in some cases, the same number of systems can be used for both training datasets.

[0103] The fiber optic sensing data received from the fiber optic system 202d may correspond to the raw measurement data and may have the same form as the data received from the systems 202a - 202c described above. The system 200 may be configured to store the data received from the fiber optic sensing system 202d in a second database 214, which may have a configuration similar to that of the first database 204 mentioned above. Then, a second data preparation module 216 is configured to prepare the fiber optic sensing data from the fiber optic sensing system 202d for use in training a second machine learning model. The second data preparation module 216 may perform steps similar to those performed by the first data preparation module 206 described above. In particular, the second data preparation module 216 may place the raw data received from the fiber optic sensing system 202d in a format adapted to the input of the second machine learning model. In some cases, the same data format may be used for both the first and second machine learning models, such that the first data preparation module 206 and the second data preparation module 216 may perform the same data preparation steps. The prepared data output by the second data preparation module 216 may be stored in a second training database 218. In other words, the second training database 218 may store the data prepared by the second data preparation module 216 based on the raw data received from the fiber optic sensing system 202d. The data stored in the second training database 218 may be considered a second training dataset. Thus, the second training dataset may include a plurality of data elements obtained based on the raw data acquired from the fiber optic sensing system 202d, and the plurality of data elements represent the fiber optic sensing measurements performed by the system 202d.

[0104] In some cases, the second training dataset may be labeled. For example, the labeling step may be performed as part of the data preparation steps performed by the second data preparation module 216. The second data preparation module 216 or another part of the system 200 may provide a user interface that is arranged to enable a user to label the data elements in the second training dataset. Then, the second training database 218 may store the second training dataset and a set of labels associated with the data elements in the second training dataset. Each label may indicate an attribute and / or a predefined condition corresponding to its associated data element. As an example, a data element may be labeled to indicate that a predefined condition or event has occurred in the system or environment monitored by the corresponding fiber optic sensing system when the corresponding measurement was performed. As another example, a data element may be labeled to indicate the nature of the system or environment monitored by the corresponding fiber optic sensing system when the corresponding measurement was performed.

[0105] System 200 further includes a second training module 220 configured to train a second machine learning model 222. The second machine learning model 222 includes at least a portion of the first machine learning model 212. For example, as Figure 2a shown, the second machine learning model 222 may include a decision model 224 connected to the output of the first machine learning model 212. The decision model 224 may be any suitable type of machine learning model, e.g., ANN, CNN, RNN, and / or transformer network. The decision model 224 may also be an SVM, a logistic regression model, or a k-means clustering model. The second training module 220 may be configured to train the second machine learning model 222 using any suitable machine learning process. In cases where the second training dataset includes labeled data elements, the second training module 220 may employ supervised learning techniques to train the second machine learning model 222. In particular, the second machine learning module 220 may treat the labeled data elements as ground truth data for supervised learning of the second machine learning model 222. Examples of second machine learning models that may be trained via the second training module 220 are described below with respect to Figure 5 .

[0106] In Figure 2a 's example, the processing of data received from the fiber optic sensing system 202d is shown separately from the data received from the fiber optic sensing systems 202a - 202c. However, this is not necessarily the case in practice. For example, in cases where a first fiber optic sensing system (e.g., system 202a) contributes to both the first training dataset and the second training dataset, the data from the fiber optic sensing system 202a may only need to be stored and prepared once. In other words, the data elements prepared by the first preparation module 206 based on the measurement data received from the fiber optic sensing system 202 may be used in both the first training dataset and the second training dataset. Thus, in some embodiments, the functions described above with respect to the second database 214, the second preparation module 216, and the second training database 218 may be implemented via the first database 204, the first preparation module 206, and the first training database, respectively, as shown by the dashed arrows in Figure 2a .

[0107] In Figure 2aIn the example, the first data preparation module 206 is used to prepare the data in the first training dataset, and the second data preparation module 216 is used to prepare the data in the second training dataset. However, in other embodiments, there may be no data preparation step. For example, the raw data stored in the first database 204 can be used as the first training dataset, that is, the first training module 210 can use the data stored in the first database 204 to train the first machine learning model 212. Similarly, the raw data stored in the second database 214 can be used as the second training dataset, that is, the second training module 220 can use the data stored in the second database 214 to train the second machine learning model 222.

[0108] It should be noted that Figure 2a The components of the system 200 depicted in (including various databases and modules) are intended to reflect the functions of the system 200 and may not necessarily represent the physical configuration of the system 200. The various modules and databases of the system 200 can be implemented using any suitable computer system (including a network of computer systems). For example, the various modules of the system 200 can be implemented as software modules on a computer system that is configured to execute the software modules to perform the tasks described herein. The various databases can be implemented using any suitable computer storage (memory) system.

[0109] After being trained by the second training module 220, the second machine learning model 222 is configured to generate an output based on the received data representing the fiber optic sensing measurement. Figure 2b An example use of the trained second machine learning model 222 is depicted in, which shows a schematic diagram of a system 226 according to an embodiment of the present invention. The system 226 includes a preparation module 228 that is configured to receive measurement data from the fiber optic sensing system 202e. In some cases, the fiber optic sensing system 202e can be the same system whose data was used to train the second machine learning model 222. For example, in an example where the system 202d was used to train the second machine learning model 222 Figure 2a the system 202e can correspond to (be the same as) the system 202d. In other cases, the system 202e can be different from the system that was used to train the second machine learning model 222.

[0110] The fiber optic sensing system 202e is configured to perform the same type of measurements as the systems 202a - 202d described above and is arranged to monitor the same type of systems or environments. Thus, the data received by the preparation module 228 from the system 202e can be raw data in the same form as the data received from the systems 202a - 202d as described above. The data received from the system 202e can correspond to data that the second machine learning model 222 has not seen before, i.e., the received data can correspond to new data that has not been used for the first training dataset or the second training dataset. In some cases, the preparation module 228 can be configured to receive a live (e.g., real - time) data stream from the fiber optic sensing system 202e. The preparation module 228 is configured to prepare the received raw data for input into the second machine learning model 222. Thus, the preparation module 228 can perform the same steps as the second preparation module 216 described above, such as data sampling, placing the data in a predetermined format, and / or applying noise reduction or smoothing algorithms to the data.

[0111] The preparation module 228 is configured to provide the prepared data to the input of the second machine learning model 222, which, in response, generates an output 230 based on the input data. In this way, the second machine learning model 222 can generate an output substantially in real - time based on the measurement data received from the fiber optic sensing system 202e. This can facilitate the monitoring of systems using the fiber optic sensing system 202e. The output generated by the second machine learning model 222 can depend on the specific task for which the second machine learning model 222 has been trained. As an example, in the case where the second training dataset has been labeled to train the second machine learning model 222 as an event detection model or an event prediction model, the output 230 generated by the second machine learning model 222 can include an indication of whether an event has occurred and / or the likelihood of the event occurring. As another example, in the case where the second training dataset has been labeled to train the second machine learning model 222 as a classifier, the output 230 generated by the second machine learning model 222 can include a classification of the received data. As a further example, in the case where the second machine learning model 222 is trained as an anomaly detector, the output 230 can indicate whether the system or environment being monitored is experiencing normal or anomalous behavior.

[0112] Figure 3aFIG. shows a schematic diagram of an optical fiber sensing system 300 that can be used to obtain optical fiber sensing data for use in the methods or systems of the present invention. For example, system 300 may correspond to one or more of the systems 202a - 202e described above. System 300 is an OTDR - based system that can be adapted to perform various optical fiber sensing measurements, such as DAS, DTS, or DSTS. The system includes a laser source 302, the output of which is split along two paths. The first path includes a pulse generation stage 304 that is configured to receive light from the laser source 302 and generate a pulsed optical signal. The pulsed signal is transmitted via an optical circulator 306 to an optical fiber 308. As the pulsed optical signal travels along the optical fiber 308, a portion of the signal is backscattered along the optical fiber 308. The backscattered signal from the optical fiber 308 that reaches the optical circulator 306 is then transmitted to a detection stage 310. A second path connected to the laser source 302 provides a local oscillator signal, and this second path may be directly connected to the detection stage 310. The detection stage 310 includes suitable detection components for detecting the backscattered signal and the local oscillator signal and for generating an output signal based on the received signals. In some cases, the detection stage 310 may be configured to interfere the backscattered signal with the local oscillator signal and detect the interference of the backscattered signal and the local oscillator signal. For example, this may occur in the case of performing DAS measurements.

[0113] Figure 3b FIG. shows an example of data representing an optical fiber measurement that can be obtained, for example, from system 300. In particular, Figure 3b the data in corresponds to a DAS measurement, which may indicate the acoustic environment along the optical fiber 308. Figure 3b A graph shows the magnitude of the power spectral density (PSD) of the DAS measurement signal within a certain frequency range (about 10 Hz to 60 Hz) as a function of time and distance along the optical fiber 308. The distance along the optical fiber 308 can be determined, for example, based on the reception time of the backscattered signal. Figure 3c A graph shows the magnitude of the PSD at a specific moment as a function of the distance along the optical fiber 308. In other words, Figure 3c the graph of corresponds to a slice (sample) along the distance axis of the graph from Figure 3b In other words, the graph of corresponds to a slice (sample) along the distance axis of the graph from

[0114] Figure 3b and Figure 3cAn example representation of DAS data is shown. Such a representation of DAS data can be generated, for example, by one of the preparation modules described above based on the raw data received from the fiber optic sensing system 300. Then, this data representation can be used in one of the training data sets and / or input into the second machine learning model. Of course, other possible representations of the measurement data are possible and can be used with the present invention. For example, other types of charts or graphs can be used to represent the data, or the data can be provided in the form of a numerical table.

[0115] Figure 4a 、 Figure 4b and Figure 4c Process 400 is shown that can be used to train the first machine learning model. Process 400 can be implemented, for example, as part of step 102 of method 100, and / or by the first training module 210 described above. Process 400 is a self-supervised learning process performed using an unlabeled first training data set. Figure 4b A data element 402 in the first training data set is shown. The data element 402 corresponds to a representation (e.g., an image) of a sample of fiber optic sensing data obtained from a fiber optic sensing system (e.g., one of systems 202a - 202c). The data element 402 may have been prepared, for example, by the first data preparation module 206 and stored in the first training database 208 as part of the first training data set. The data representation for the data element 402 can correspond to the data representation described above with respect to Figure 3b i.e., it can represent the amplitude of the DAS signal as a function of time and distance along the fiber. The first training data set can include multiple such data elements, each representing a corresponding sample of fiber optic sensing data.

[0116] First, process 400 involves applying a randomly generated mask 404 to the data element 402 to partially mask the data in the data element. The mask 404 can be arranged such that a predetermined proportion of the data element 402 is covered, but the covered regions of the data element 402 can be random. As an example, the data element 402 can be divided into multiple small blocks, and a certain number of randomly selected small blocks are masked. Then, the partially masked data element 402 is provided as an input to the first machine learning model, which is trained to predict or reconstruct the data element 402. Thus, the first machine learning model can be configured as the encoder part of an autoencoder (e.g., a masked autoencoder) trained to reconstruct the masked data.

[0117] More specifically, as Figure 4aAs shown, a masked data element 402 is provided as an input to a first machine learning model 406, which is configured as an encoder (the first machine learning model 406 may correspond to the model 212 described above). The first machine learning model 406 then encodes the masked data element 402 into a latent space 408. The decoder 410 is then configured to decode the representation in the latent space 408 output by the first machine learning model 406 to obtain a series of image patches 412, which are then reconstructed into a predicted data element 414. The first machine learning model 406 (encoder) and the decoder 410 form a masked autoencoder. The predicted data element 414 can then be compared with the original (unmasked) data element 402 to refine the first machine learning model 406. Many iterations of this training process can be performed using each of the data elements in the first training dataset, each time applying a randomly generated mask to the data element and then using the first machine learning model 406 and the decoder 410 to reconstruct the data element. The first machine learning model 406 can be trained in this way until a desired level of accuracy of the predicted data element 414 is achieved. The encoder (i.e., the first machine learning model 406) may include a vision transformer (ViT) network. Such a ViT may be particularly applicable in cases where fiber optic sensing data is represented as a graph or an image (e.g., as shown in Figure 3b ).

[0118] Once the first machine learning model 406 has been trained, it can be used to form a second machine learning model 502, as shown in Figure 5 . The second machine learning model 502 includes the first machine learning model 406 (i.e., the encoder from process 400) and a decision model 502, which is connected to the output of the first machine learning model 406. In particular, the decision model 504 is arranged to take as input the representation in the latent space 408 output by the first machine learning model 406. Then, based on the latent space representation received from the first machine learning model 406, the decision model 504 is configured to generate an output 506. The decision model 504 may correspond to, for example, an ANN (such as a CNN).

[0119] Figure 4c A more detailed example of the training process 400 is shown in. The features discussed above with respect to Figure 4a and Figure 4b similarly apply to Figure 4cExample. As discussed above, the randomly generated mask 404 is applied to the data element 402. In particular, as shown, the data element 402 is divided into a plurality of small blocks (or portions) of the same size, and a certain number of randomly selected small blocks are masked. A predetermined portion of the data element 402 can be masked. For example, 50% or more (e.g., 75%) of the data element 402 can be masked. Then, as Figure 4c shown, the remaining unmasked (i.e., visible) small blocks 405 from the data element are provided as the input to the first machine learning model 406, which is implemented as the encoder part of a masked autoencoder. Then, the first machine learning model 406 encodes the unmasked small blocks 405 into the latent space 408. The latent space 408 can have a dimension equal to or higher than the dimension of the input unmasked small blocks 405. Using a latent space 408 with a higher dimension compared to the input unmasked small blocks 405 can facilitate the generation of missing data from the masked small blocks. Here, the latent space dimension can correspond to the number of visible (unmasked) small blocks multiplied by the number of features embedded in each small block, plus additional features from the classification (CLS) token (if used). For example, if the data element 402 is divided into 256 small blocks, and 75% are masked (leaving 64 visible small blocks), and each small block embedding has a dimension of 768 (a common size in the ViT architecture), then the latent space dimension size will be 64×768 = 49152. If the CLS token is used, its feature dimension (usually also 768) will be added to this 64*768 + 768 = 49920. It should be noted that such a latent space dimension is used for the self-supervised learning process. However, after training the first machine learning model 406, when using the first machine learning model 406 to form the second machine learning model, it is not necessarily required to use the full latent space. For example, depending on the task to be performed by the decision model 504, the decision model 504 can take only a part (subset) of the latent space representation generated by the first machine learning model 406 as the input. For example, the decision model 504 can use only the CLS token (e.g., dimension 768), and / or perform average pooling on the small blocks in the latent space representation. Such average pooling can involve averaging the small block dimensions, leaving only the embedded dimension (e.g., 768).

[0120] Return to Figure 4cIn process 400, decoder 410 is configured to decode the representation in latent space 408 output by first machine learning model 406. Additionally, decoder 410 takes as input the representation of the masked patches from original data element 402. In other words, decoder 410 takes as input the combination 409 of the masked patches and the representation in latent space 408 of the unmasked patches. Thus, first machine learning model 406 takes only the unmasked patches 405 as input, while decoder 410 takes both the latent space representation 408 and the masked patches as input.

[0121] The autoencoder formed by first machine learning model (encoder 406) and decoder 410 can have an asymmetric architecture, where first machine learning model 406 can be larger and more complex than decoder 410. For example, first machine learning model 406 can have a larger number of parameters compared to decoder 410. This enables the training of a more complex first machine learning model 406 while improving the overall computational efficiency of the training process. This asymmetric architecture achieves an improvement in computational efficiency. In particular, since only a relatively small unmasked portion of original data element 402 is provided to first machine learning model 406 (e.g., only 25% of the patches when 75% are masked), using a larger and more complex encoder as first machine learning model 406 can reduce the computational load and memory usage. This allows for faster training and facilitates the handling of larger data elements. The more complex model of first machine learning model 406 allows for a deeper learning of the characteristics of the fiber sensing data. In contrast, since the role of decoder 410 is to reconstruct original data element 402 based on the latent space representation, this task is not as complex as the feature extraction and representation learning performed by first machine learning model 406. Thus, decoder 410 can be implemented using a smaller and more lightweight model. Using a more lightweight decoder 410 allows more computational resources to be dedicated to the resource-intensive tasks performed by first machine learning model 406 and further helps to improve the balance of the training dynamics. In particular, this asymmetric architecture encourages deep and meaningful learning in first machine learning model 406 because it can be independent of an overly powerful decoder to "fix" poor representations. The smaller decoder can also avoid the problem of overfitting, because a larger decoder with more parameters can potentially memorize specific details of the first training dataset rather than learning to reconstruct the data based on the representation generated by first machine learning model 406.

[0122] The representation of the masked patches (hereinafter referred to as masked tokens) included in the combined input 409 can include details of the patch positions in the original data elements 402, such that the decoder 410 can reconstruct the data elements 402. The decoder 410 is configured to decode the combined input 409 to obtain a series of image patches 412, which are then reconstructed into the generated data elements 414. The generated data elements 414 can then be compared with the original (unmasked) data elements 402 to refine the first machine learning model 406. The quality of the generated data elements 414, and in particular the patches corresponding to the masked patches in the original data elements 402, provides a measure of how well the autoencoder has learned to understand and represent the data. As discussed with respect to Figure 4a The process, Figure 4c can be iteratively repeated, each time applying a randomly generated mask to the data elements 402 in order to gradually train the first machine learning model 406. More specifically, the masked tokens serve as placeholders for the patches in the data elements 402 that have been masked (e.g., removed or hidden) during the masking process. These tokens are used to indicate the positions of the masked patches when the data is passed through the decoder 410. The masked tokens are designed to have the same dimensions as the encoded visible patches. This ensures that when the combined representation 409 (including both the encoded visible patches 408 and the masked tokens) is fed into the decoder 410, all components of the input have compatible dimensions to allow for the efficient operation of the masked autoencoder architecture. When the decoder 410 attempts to reconstruct the original data elements 402, it uses information from both the encoded unmasked patches 408 and the position information provided by the masked tokens to generate the missing parts of the data elements 402. In some cases, a single learned representation (e.g., a trainable vector) is used for all masked tokens, which can be optimized during the training process.

[0123] Figure 5 Figure 500 shows a process 500 for training a second machine learning model 502. The training of the second machine learning model 502 is performed using a second training dataset. The data elements 508 are provided to the input of the second machine learning model 502, which can correspond to the input of the first machine learning model 406. Similar to the data elements 402 described above, the data elements 508 are representations (e.g., images) of samples of fiber optic sensing data obtained from a fiber optic sensing system (e.g., system 202d). The data elements 508 may have been prepared, for example, by a second data preparation module 216 and stored in a second training database 218 as part of the second training dataset. The data representation for the data elements 508 from the second training dataset is the same as that for the data elements 402 from the first training dataset, e.g., as described above with respect to Figure 3bThe described data representation corresponds. However, unlike in process 400, data element 508 is not masked, i.e., the complete data element 508 is provided to the second machine learning model 502. According to the above, after the data element 508 is input into the second machine learning model 502, the decision model 504 generates an output 506. The output 506 can be used as feedback during the training process to refine (adjust) the second machine learning model 502. Process 500 can be iteratively executed multiple times using each data element in the data elements of the second training dataset to optimize the accuracy of the output 506 provided by the model. In some cases, the first machine learning model 406 can remain fixed during the process of training the second machine learning model 502. In other words, only the parameters in the decision model 504 can be adjusted as part of the training process 500, while the parameters in the first machine learning model 406 remain fixed.

[0124] As described above, the data elements in the second training dataset can be labeled. Therefore, process 500 can involve a supervised learning process, where the labeled data elements are used for the supervised learning of the second machine learning model. In other words, the labeled data elements can act as the ground truth data for training the second machine learning model 506. For example, in the case where the second machine learning model 502 is trained as a classifier or an event detection or prediction model, supervised learning techniques can be used. For example, the data elements in the second training dataset can be labeled to classify the data and / or indicate whether the data element corresponds to a predetermined event. In such cases, the decision model 504 can be a logistic regression model for solving classification problems. Supervised learning techniques can also be used to predict the properties or variables (e.g., attributes) of the system or environment monitored by the fiber optic sensing system. For example, the data elements in the second training dataset can be labeled to indicate the value of the property or variable corresponding to the time when the data was recorded. Supervised learning techniques can further be used for anomaly detection. In such cases, the data elements in the second training dataset can be labeled to indicate whether the data element corresponds to the normal behavior or the abnormal behavior of the system or environment monitored by the fiber optic sensing system. Examples of suitable supervised machine learning models that can be used for the decision model 504 include ANN, CNN, transformer networks, and SVM.

[0125] Alternatively, in the case where the data elements in the second training dataset are unlabeled, an unsupervised learning process can be performed. For example, unsupervised learning can be used to train the second machine learning model 502 as an anomaly detector to predict whether the received data corresponds to normal behavior or anomalous behavior. For example, nearest neighbor anomaly detection can be used. Nearest neighbor anomaly detection techniques rely on the assumption that "normal" data instances occur in dense neighborhoods, while anomalies occur far from their nearest neighbors. Thus, the distance or similarity metric between two data instances can be defined and calculated in various ways, such as the Euclidean distance.

[0126] Figure 6 A process 600 for training the second machine learning model 602 is shown. The process 600 is based on the same principle as the process 500 described above, except that the second machine learning model 602 includes a plurality of decision models 604a - 604n. Thus, the second machine learning model 502 includes a single decision model 504 connected to the output of the first machine learning model 406, while the second machine learning model 602 includes a plurality of decision models 604a - 604n connected to the output of the first machine learning model 406. Each of the decision models 604a - 604n is configured to generate a corresponding output 606a based on the input received from the first machine learning model 406. In the example shown, each of the decision models 604a - 604n is arranged to take the same output from the latent space 408 in the first machine learning model 406 as its input. However, in other examples, the decision models 604a - 604n can be arranged to receive the outputs from different layers in the first machine learning model 406.

[0127] Each of the decision models 604a - 604n in the decision model can be configured to perform a corresponding task. For example, the first decision model 604a can be configured as a classifier, the second decision model 604 can be configured as an event detector, and another decision model 604n can be configured as an anomaly detector. Of course, depending on the type and number of tasks to be performed, a different number of decision models can be connected to the first machine learning model 406. Each of the decision models 604a - 604n can be independently trained using a corresponding second training dataset. Thus, during the training of one of the decision models in the decision model (e.g., model 604a), the first machine learning model 406 and the other decision models (e.g., models 604b - 604n) can remain fixed. In this way, each of the decision models 604a - 604n can be sequentially trained using data elements 608 from the corresponding second training dataset for each decision model. Each corresponding second training dataset can be adapted to the type of training and tasks to be performed by the corresponding decision model. For example, the second training dataset for a particular decision model can be labeled or unlabeled depending on whether the training of the model is to be supervised. Once the second machine learning model 602 has been trained, it can take the data elements as input and provide multiple outputs 606a - 606n, i.e., one output from each of the decision models 604a - 604n.

[0128] Features disclosed in the foregoing description or in the appended claims or in the drawings, expressed in their specific form or in means for performing the disclosed function or in a method or process for obtaining the disclosed result, may, where appropriate, be used individually or in any combination of such features in their diverse forms to implement the present invention.

[0129] Although the present invention has been described in connection with the exemplary embodiments described above, many equivalent modifications and variations will be apparent to those skilled in the art when the present disclosure is given. Therefore, the exemplary embodiments set forth above of the present invention are considered illustrative and not restrictive. Various changes may be made to the described embodiments without departing from the spirit and scope of the present invention.

[0130] To avoid any doubt, any theoretical explanations provided herein are for the purpose of enhancing the reader's understanding. The inventors do not wish to be bound by any of these theoretical explanations.

[0131] Any section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0132] Throughout the specification, including the appended claims, unless the context requires otherwise, the words "comprise" and "include" and variations thereof will be understood to imply the inclusion of the stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.

[0133] It must be noted that, unless the context clearly indicates otherwise, as used in this specification and the appended claims, the singular forms "a", "an" and "the" include plural referents. Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, another embodiment includes from one particular value and / or to another particular value. Similarly, when values are expressed as approximations by use of the antecedent "about", it will be understood that the particular value forms another embodiment. The term "about" with respect to a numerical value is optional and means, for example, + / - 10%.

[0134] The present invention may be defined in the following terms:

[0135] 1. A method of training a machine learning model for use with data representing fiber optic sensing measurements, the method comprising:

[0136] training a first machine learning model using a first training dataset, wherein the first training dataset includes data representing fiber optic sensing measurements, and wherein the first training dataset is unlabeled; and

[0137] training a second machine learning model using a second training dataset, the second machine learning model comprising at least a portion of the trained first machine learning model, wherein the second training dataset includes data representing fiber optic sensing measurements, and the size of the second training dataset is less than or equal to the first training dataset;

[0138] wherein the second machine learning model is configured to generate an output based on received data representing fiber optic sensing measurements.

[0139] 2. The method according to clause 1, wherein the second machine learning model is trained using a supervised learning process, and wherein the second training dataset includes labeled data used in the supervised learning process.

[0140] 3. The method according to clause 2, wherein:

[0141] the labeled data includes data elements that are labeled to indicate an attribute of the data elements; and

[0142] the output generated by the second machine learning model indicates an attribute of the received data.

[0143] 4. The method according to any one of the preceding clauses, wherein an unsupervised learning process is used to train the second machine learning model, and wherein the second training dataset includes unlabeled data used in the unsupervised learning process.

[0144] 5. The method according to any one of the preceding clauses, wherein the second machine learning model is configured as an anomaly detector.

[0145] 6. The method according to any one of the preceding clauses, wherein the first training dataset is collected using two or more different fiber optic sensing systems.

[0146] 7. The method according to clause 6, wherein the two or more different fiber optic sensing systems are located at different geographical locations; and / or

[0147] wherein, compared with the first training dataset, the second training dataset is collected using fewer fiber optic sensing systems.

[0148] 8. The method according to any one of the preceding clauses, wherein the second machine learning model includes the trained first machine learning model and a decision model, and the decision model is connected to the output of the first machine learning model.

[0149] 9. The method according to any one of clauses 1 to 7, wherein:

[0150] the second machine learning model includes the trained first machine learning model and two or more decision models, and each of the two or more decision models is connected to the output of the first machine learning model; and

[0151] each of the two or more decision models is trained using a corresponding second training dataset.

[0152] 10. The method according to any one of the preceding clauses, wherein during the training of the second machine learning model, the trained first machine learning model remains fixed.

[0153] 11. The method according to any one of clauses 1 to 9, wherein during the training of the second machine learning model, one or more parameters of the trained first machine learning model are adjusted.

[0154] 12. The method according to any one of the preceding clauses, wherein the first machine learning model is trained using the first training dataset via a self-supervised learning process; and optionally

[0155] Wherein, the self-supervised learning process includes partially masking the data in the first training dataset and training the first machine learning model to reconstruct the data in the first training dataset based on the partially masked data.

[0156] 13. The method according to any one of the preceding clauses, wherein the first training dataset further includes data indicating environmental conditions associated with the recording of the data in the first training dataset, and / or the second training dataset further includes data indicating environmental conditions associated with the recording of the data in the second training dataset.

[0157] 14. A method for analyzing data representing fiber optic sensing measurements, the method comprising:

[0158] Receiving, by a machine learning model, data representing fiber optic sensing measurements, wherein the machine learning model has been trained using the method according to any one of the preceding claims; and

[0159] Generating, in response to receiving the data, an output by the machine learning model.

[0160] 15. The method according to clause 14, wherein the output generated by the machine learning model indicates the probability that the received data corresponds to a predetermined condition, and / or the output generated by the machine learning model indicates an attribute of the data.

[0161] 16. A computer-implemented system for training a machine learning model, the system being configured to:

[0162] Train a first machine learning model using a first training dataset, wherein the first training dataset includes data representing fiber optic sensing measurements, and wherein the first training dataset is unlabeled; and

[0163] Train a second machine learning model using a second training dataset, the second machine learning model comprising at least a portion of the trained first machine learning model, wherein the second training dataset includes data representing fiber optic sensing measurements, the size of the second training dataset is less than or equal to the first training dataset, and wherein the second machine learning model is configured to generate an output based on received data representing fiber optic sensing measurements.

[0164] 17. A computer-implemented system for analyzing data representing fiber optic sensing measurements, the system including a machine learning model that has been trained using the method according to any one of clauses 1 to 13.

Claims

1. A method of training a machine learning model for use with data representing fiber optic sensing measurements, the method comprising: training a first machine learning model using a first training dataset, wherein the first training dataset includes data representing fiber optic sensing measurements and wherein the first training dataset is unlabeled; and training a second machine learning model using a second training dataset, the second machine learning model comprising at least a portion of the trained first machine learning model, wherein the second training dataset includes data representing fiber optic sensing measurements and the size of the second training dataset is less than or equal to the first training dataset; wherein the second machine learning model is configured to generate an output based on received data representing fiber optic sensing measurements; wherein the first machine learning model is trained using the first training dataset via a self-supervised learning process; and wherein the self-supervised learning process includes partially masking data in the first training dataset and training the first machine learning model to reconstruct the data in the first training dataset based on the partially masked data.

2. The method according to claim 1, wherein, training the second machine learning model using a supervised learning process, and wherein the second training dataset includes labeled data used in the supervised learning process.

3. The method according to claim 2, wherein: the labeled data includes data elements labeled to indicate an attribute of the data elements; and the output generated by the second machine learning model indicates an attribute of the received data.

4. The method according to any one of the preceding claims, wherein, training the second machine learning model using an unsupervised learning process, and wherein the second training dataset includes unlabeled data used in the unsupervised learning process.

5. The method according to any one of the preceding claims, wherein, the second machine learning model is configured as an anomaly detector.

6. The method according to any one of the preceding claims, wherein, the first training dataset is collected using two or more different fiber optic sensing systems.

7. The method according to claim 6, wherein, the two or more different fiber optic sensing systems are located at different geographical locations; and / or wherein, compared to the first training dataset, the second training dataset is collected using fewer fiber optic sensing systems.

8. The method according to any one of the preceding claims, wherein, the second machine learning model includes the trained first machine learning model and a decision model connected to an output of the first machine learning model.

9. The method according to any one of claims 1 to 7, wherein: the second machine learning model includes the trained first machine learning model and two or more decision models, each of the two or more decision models being connected to an output of the first machine learning model; and each of the two or more decision models is configured to perform a respective task and is trained using a respective second training dataset.

10. The method according to any one of the preceding claims, wherein, during training of the second machine learning model, the trained first machine learning model remains fixed.

11. The method according to any one of claims 1 to 9, wherein during training of the second machine learning model, one or more parameters of the trained first machine learning model are adjusted.

12. The method according to any one of the preceding claims, wherein, The self-supervised learning process includes providing the unmasked portion of the first data set to the first machine learning model and training the first machine learning model to reconstruct the data in the first training data set based on the unmasked portion.

13. The method according to any one of the preceding claims, wherein, The first training data set further includes data indicating environmental conditions associated with the recording of the data in the first training data set, and / or the second training data set further includes data indicating environmental conditions associated with the recording of the data in the second training data set.

14. A method of analyzing data representative of fiber optic sensing measurements, the method comprising: receiving, by a machine learning model, data representative of fiber optic sensing measurements, wherein the machine learning model has been trained using the method according to any one of the preceding claims; and generating, in response to receiving the data, an output by the machine learning model.

15. The method according to claim 14, wherein, The output generated by the machine learning model indicates the probability that the received data corresponds to a predetermined condition, and / or the output generated by the machine learning model indicates an attribute of the data.

16. A computer-implemented system for training a machine learning model, the system being configured to: Use a first training data set to train a first machine learning model, where, The first training data set includes data representative of fiber optic sensing measurements, and wherein the first training data set is unlabeled; and training a second machine learning model using a second training data set, the second machine learning model including at least a portion of the trained first machine learning model, wherein the second training data set includes data representative of fiber optic sensing measurements, the size of the second training data set is less than or equal to the first training data set, and wherein the second machine learning model is configured to generate an output based on the received data representative of fiber optic sensing measurements; wherein the system is further configured to: train the first machine learning model using the first training data set via a self-supervised learning process; and the self-supervised learning process includes partially masking the data in the first training data set and training the first machine learning model to reconstruct the data in the first training data set based on the partially masked data.

17. A computer-implemented system for analyzing data representative of fiber optic sensing measurements, the system including a machine learning model that has been trained using the method according to any one of claims 1 to 13.