Methods and systems for pipeline anomaly detection

The method uses fiber optic sensors and meters to transform and cluster pipeline data into encoded signatures, calculating anomaly scores to accurately detect anomalies and reduce false alarms, improving pipeline safety and efficiency.

WO2025137766A1PCT designated stage expired Publication Date: 2025-07-03HIFI ENGINEERING INC

Patent Information

Application Number
PCT/CA2024/051727
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-23
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing pipeline monitoring systems face challenges in accurately detecting anomalies while minimizing false positives, particularly in complex pipelines with diverse materials and conditions, and manual monitoring is cumbersome and ineffective.

Method used

A method involving fiber optic sensors and meters to acquire operational data, transform it into encoded signatures, cluster these signatures in a feature space, and calculate anomaly scores using baseline clusters to identify anomalies of interest, with adaptive updating of baseline clusters and thresholds to accommodate changing conditions.

Benefits of technology

Enhances pipeline safety and operational accuracy by reducing false alarms and adaptively distinguishing between common patterns and anomalies, providing timely alerts for potential hazards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2024051727_03072025_PF_FP_ABST
    Figure CA2024051727_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems for detecting anomalies in pipeline operational data. The methods and systems include acquiring operational data, transforming the data into encoded signatures via a feature extraction model, clustering the signatures in a feature space, and calculating an anomaly score based on a comparison to baseline clusters. By comparing the signatures to known event clusters and scoring against a threshold, anomalies of interest in pipeline operations can be efficiently identified, while false detections of anomalies can be greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR PIPELINE ANOMALY DETECTIONTECHNICAL FIELD

[0001] The present disclosure is directed to the field of pipeline monitoring, and more particularly to methods and systems for detecting anomalies of interest in pipeline operations using encoded data signatures.BACKGROUND

[0002] Pipeline systems are vital infrastructure that connect production sites to consumption points over vast distances. Their role in transporting essential commodities such as oil and gas underscores the importance of their constant and efficient monitoring. Ensuring the safety, operational efficiency and integrity of these systems is paramount.

[0003] A major challenge in pipeline management is the detection of anomalies. Undetected anomalies can have serious consequences, including environmental hazards, significant economic losses, and potential safety hazards. Given the vast expanse of many pipeline systems, manual monitoring is not only cumbersome, but often ineffective.

[0004] To address this challenge, the industry has turned to technology, ushering in the era of Computational Pipeline Monitoring (CPM). CPM leak detection systems use computational methods to analyze flow and pressure data to identify potential leaks or anomalies within the pipeline. These systems have been instrumental in automating the detection process, providing continuous monitoring and faster response times.

[0005] However, as pipelines become more complex and the materials transported become more diverse, the requirements for monitoring systems are evolving. In order to accurately monitor pipelines while reducing and ideally eliminating false positives, improved systems and methods are desirable.SUMMARY

[0006] According to a first aspect, there is provided a method, comprising: acquiring a first set of operational data from one or more fiber optic sensors deployed along a pipeline over a first time period; acquiring a second set of operational data from one or more meters deployed along the pipeline over the first time period; transforming, by a feature extraction model, theacquired first and second sets of operational data into first and second encoded signatures, respectively, wherein each signature of the first set of encoded signatures or each signature of the second set of encoded signatures contains a plurality of features; clustering the first and second encoded signatures in a feature space based on the features of the first and second encoded signatures; for at least one of the clustered first encoded signatures, calculating a first anomaly score by comparing the at least one of the clustered first encoded signatures with baseline clusters associated with known events; for at least one of the clustered second encoded signatures, calculating a second anomaly score by comparing the at least one of the clustered second encoded signatures with the baseline clusters; and identifying an anomaly of interest in response to both the first and second anomaly score meeting a threshold value.

[0007] In some embodiments, the method may further comprise storing the first and second sets of operational data over time in a buffer, and the first and second sets of operational data may be acquired from the buffer.

[0008] In some embodiments, the threshold value may comprise a first threshold value and a second threshold value, and the identifying of the anomaly of interest may further comprise: identifying the anomaly of interest in response to the first anomaly score meeting the first threshold value and the second anomaly score meeting the second threshold value.

[0009] In some embodiments, the one or more fiber optic sensors may be segmented into a plurality of monitoring channels, and acquiring the first set of operational data may comprise acquiring at least one of acoustic data, acoustic based flow rate estimate, strain based pressure estimate, differential temperature estimate, absolute temperature estimate, or transfer function based flow rate and pressure estimate for at least one of the plurality of monitoring channels.

[0010] In some embodiments, the one or more meters may comprise at least one of a temperature sensor, a pressure meter, a flow meter, a fluid density meter, or a viscosity meter.

[0011] In some embodiments, the transforming may comprise extracting the plurality of features from the acquired first and second sets of operational data by the feature extraction model utilizing a Convolutional Neural Network.

[0012] In some embodiments, the first and second encoded signatures may be clustered using k-means clustering or Manhattan distance-based clustering, which assigns a respectiveencoded signature into one of the baseline clusters by minimizing a distance between the respective encoded signature and centroids of the baseline clusters in the feature space; and the anomaly score may be calculated based on the distance between a respective clustered signature and a centroid of a baseline cluster associated with the known events.

[0013] In some embodiments, the first and second encoded signatures may be clustered using cosine similarity -based clustering, which assigns a respective encoded signature into one of the baseline clusters by minimizing an angle between the respective encoded signature and centroids of the baseline clusters in the feature space; and the anomaly score may be calculated based on the angle between a respective clustered signature and a centroid of a baseline cluster associated with the known events.

[0014] In some embodiments, the first and second encoded signatures may be clustered using one of adaptive Gaussian mixture model, hierarchical clustering, mean shift, or spectral clustering; and the anomaly score may be calculated based on a placement of the at least one signature of the clustered first and second encoded signatures in a designated cluster of the baseline clusters in the feature space and a frequency of occurrences of signatures within the designated cluster.

[0015] In some embodiments, the method may further comprise dynamically updating the baseline clusters in response to the first and second encoded signatures being introduced based on features thereof

[0016] In some embodiments, the feature extraction model may be trained using a training process comprising: a) providing a training dataset representative of operational data and associated training signatures; b) feeding the training dataset to the feature extraction model; c) adjusting parameters of the feature extraction model based on differences between signatures encoded by the feature extraction model and the training signatures in the training dataset; and d) iterating b) and c) until a predetermined accuracy or a number of iterations is achieved.

[0017] In some embodiments, the training dataset may comprise normal operational data and data associated with known historical anomalies of interest.

[0018] In some embodiments, the baseline clusters may be obtained by a baselining process comprising: acquiring a set of baseline operational data detected over a second time period; transforming the baseline operational data into baseline encoded signatures using the feature extraction model; and clustering the baseline encoded signatures in the feature space to generate the baseline clusters.

[0019] In some embodiments, the baselining process may further comprise updating the baseline clusters based on the first and second operational data acquired over a third time period to accommodate changing operational conditions of the pipeline.

[0020] In some embodiments, the baselining process may further comprise, over time: updating the baseline clusters and the threshold value based on the first and second operational data from a sliding time window; and removing the baseline encoded signatures that are encoded prior to the sliding time window from the feature space.

[0021] In some embodiments, the feature extraction model may be selected from one of autoencoder, statistical feature extraction, principle component analysis (PCA), t-distributed Stochastic Neighbor Embedding (t-SNE), or an embedding from a pre-trained neural network.

[0022] In some embodiments, the anomaly of interest may comprise at least one of a leak, an intrusion, or a turbulence.

[0023] According to a second aspect, there is provided a method, comprising: acquiring operational data from one or more detectors deployed along a pipeline over a first time period, the one or more detectors comprising a fiber optic sensor deployed along the pipeline; transforming, by an feature extraction model, the acquired operational data into encoded signatures, wherein each of the encoded signatures contains a plurality of features; clustering the encoded signatures in a feature space using the features of the encoded signatures; for at least one of the clustered signatures, calculating an anomaly score by comparing the at least one of the clustered signatures with baseline clusters associated with known events; and identifying an anomaly of interest in response to the anomaly score meeting a threshold value.

[0024] In some embodiments, the one or more detectors may further comprise one or more non-optical meters deployed along the pipeline.

[0025] In some embodiments, the identifying may further comprise identifying the anomaly of interest in response to both the anomaly score corresponding to the operational data acquired from the fiber optic sensor and the anomaly score corresponding to the operational data acquired from the one or more non-optical meters meeting the threshold value.

[0026] In some embodiments, the one or more non-optical meters may comprise at least one of a temperature sensor, a pressure meter, a flow meter, a fluid density meter, or a viscosity meter.

[0027] In some embodiments, the calculating may further comprises calculating respective anomaly scores for all of the clustered signatures.

[0028] In some embodiments, the method may further comprise storing the operational data over time in a buffer, wherein the operational data is acquired from the buffer.

[0029] In some embodiments, the fiber optic sensor may be segmented into a plurality of monitoring channels, and acquiring the operational data may further comprise acquiring at least one of acoustic data, acoustic based flow rate estimate, strain based pressure estimate, differential temperature estimate, absolute temperature estimate, or transfer function based flow rate and pressure estimate for at least one of the plurality of monitoring channels.

[0030] In some embodiments, the transforming may further comprise extracting the plurality of features from the acquired operational data by the feature extraction model utilizing a Convolutional Neural Network.

[0031] In some embodiments, the encoded signatures may be clustered using k-means clustering or Manhattan distance-based clustering, which assigns a respective encoded signature into one of the baseline clusters by minimizing a distance between the respective encoded signature and centroids of the baseline clusters in the feature space; and the anomaly score may be calculated based on the distance between a respective clustered signature and a centroid of a baseline cluster associated with the known events.

[0032] In some embodiments, the encoded signatures may be clustered using cosine similarity-based clustering, which assigns a respective encoded signature into one of thebaseline clusters by minimizing an angle between the respective encoded signature and centroids of the baseline clusters in the feature space; and the anomaly score may be calculated using the angle between a respective clustered signature and a centroid of a baseline cluster associated with the known events.

[0033] In some embodiments, the encoded signatures may be clustered using one of adaptive Gaussian mixture model, hierarchical clustering, mean shift, or spectral clustering; and the anomaly score may be calculated based on a placement of the at least one of the clustered signatures in a designated cluster of the baseline clusters in the feature space and a frequency of occurrences of signatures within the designated cluster.

[0034] In some embodiments, the method may further comprise dynamically updating the baseline clusters in response to the encoded signatures being clustered using features thereof.

[0035] In some embodiments, the feature extraction model may be trained using a training process comprising: a) providing a training dataset representative of operational data and associated training signatures; b) feeding the training dataset to the feature extraction model; c) adjusting parameters of the feature extraction model based on differences between signatures encoded by the feature extraction model and the training signatures in the training dataset; and d) iterating b) and c) until a predetermined accuracy or a number of iterations is achieved.

[0036] In some embodiments, the training dataset may comprise normal operational data and data associated with known historical anomalies of interest.

[0037] In some embodiments, the baseline clusters may be obtained by a baselining process comprising: acquiring a set of baseline operational data detected over a second time period; transforming the baseline operational data into baseline encoded signatures using thefeature extraction model; and clustering the baseline encoded signatures in the feature space to generate the baseline clusters.

[0038] In some embodiments, the baselining process may further comprise updating the baseline clusters based on the operational data acquired over a third time period to accommodate changing operational conditions of the pipeline.

[0039] In some embodiments, the baselining process may further comprise, over time: updating the baseline clusters and the threshold value based on the operational data from a sliding time window; and removing the baseline encoded signatures that are encoded prior to the sliding time window from the feature space.

[0040] In some embodiments, the feature extraction model may be selected from one of autoencoder, statistical feature extraction, principle component analysis (PCA), t-distributed Stochastic Neighbor Embedding (t-SNE), or an embedding from a pre-trained neural network.

[0041] In some embodiments, the anomaly of interest may comprise at least one of a leak, an intrusion, or a turbulence.

[0042] According to a third aspect, there is provided a system, comprising: at least one pipeline; an anomaly detection apparatus; one or more fiber optic sensors deployed along one of the at least one pipeline; one or more meters deployed along the one of the at least one pipeline or along another one of the at least one pipeline; wherein the anomaly detection apparatus is configured to perform: a) acquire a first set of operational data from the one or more fiber optic sensors; b) acquire a second set of operational data from the one or more meters; c) transform, by a feature extraction model, the acquired first set of operational data into first encoded signatures, wherein each signature of the first set of encoded signatures contains a first plurality of features;d) transform, by a feature extraction model, the acquired second set of operational data into second encoded signatures, wherein each signature of the second set of encoded signatures contains a second plurality of features; e) cluster the first encoded signatures in a first feature space based on the features of the first encoded signatures; f) cluster the second encoded signatures in a second feature space based on the features of the second encoded signatures; g) for at least one of the clustered first encoded signatures, calculate a first anomaly score by comparing the at least one of the clustered first encoded signatures with baseline clusters associated with known events; h) for at least one of the clustered second encoded signatures, calculate a second anomaly score by comparing the at least one of the clustered second encoded signatures with the baseline clusters; and i) identify an anomaly of interest in response to both the first and second anomaly score meeting a threshold value.

[0043] In some embodiments, the anomaly detection apparatus may comprise a first anomaly detection device and a second anomaly detection device, and the first anomaly detection device may be configured to perform a), c), e), and g), and the second anomaly detection device may be configured to perform b), d), f), and h).

[0044] In some embodiments, the anomaly detection apparatus may further comprise a server configured to communicatively couple with the first anomaly detection device and the second anomaly detection device, and the server may be configured to receive the first anomaly score and the second anomaly score from the first anomaly detection device and the second anomaly detection device, respectively, and perform i).

[0045] According to a fourth aspect, there is provided a system, comprising: a pipeline; an anomaly detection apparatus; one or more detectors deployed along the pipeline, the one or more detectors comprising a fiber optic sensor deployed along the pipeline; wherein the anomaly detection apparatus is configured to perform:a) acquire operational data from the one or more detectors; b) transform, by a feature extraction model, the acquired operational data into encoded signatures, wherein each of the encoded signatures contains a plurality of features; c) cluster the encoded signatures in a feature space using the features of the encoded signatures; d) for at least one of the clustered encoded signatures, calculate an anomaly score by comparing the at least one of the clustered encoded signatures with baseline clusters associated with known events; and e) identify an anomaly of interest in response to the anomaly score meeting a threshold value.

[0046] In some embodiments, the one or more detectors may further comprise one or more non-optical meters deployed along the pipeline.

[0047] In some embodiments, the anomaly of interest may be identified in response to both the anomaly score corresponding to the operational data acquired from the fiber optic sensor and the anomaly score corresponding to the operational data acquired from the one or more non-optical meters meeting the threshold value.

[0048] In some embodiments, the one or more non-optical meters may comprise at least one of a temperature sensor, a pressure meter, a flow meter, a fluid density meter, or a viscosity meter.

[0049] In some embodiments, the anomaly detection apparatus may further comprise a server configured to communicatively couple with the anomaly detection device, and wherein the server is configured to receive the anomaly score from the anomaly detection device and perform e).

[0050] According to a fifth aspect, there is provided a non-transitory computer readable medium have stored thereon computer program code that is executable by at least one processor and that, when executed by the at least one processor, causes the at least one processor to perform the methods described herein.

[0051] This disclosure integrates data from one or more detectors to provide a holistic approach to pipeline monitoring. By maintaining a record of recurring events at specific locations or in specific channels, the system astutely distinguishes between common operational patterns and anomalies. Advanced clustering algorithms respond adaptively to new and unique signatures, ensuring that only rare and potentially harmful events trigger alarms. In addition, another advantage of this disclosure is its resistance to noise and highly active channels, which significantly reduces false alarms. By coupling this robust memory mechanism with dual data sources, the aspects of this disclosure improve existing anomaly monitoring techniques, enhancing both pipeline safety and operational accuracy.

[0052] This summary does not necessarily describe the full scope of all aspects. Other aspects, features and advantages will become apparent to those of ordinary skill in the art upon review of the following description of specific embodiments.BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In the accompanying drawings, which illustrate one or more example embodiments:

[0054] FIG. 1 is a block diagram of a system for fiber optic sensing using backscattering occurring in an optical fiber, according to an embodiment.

[0055] FIG. 2 is an example diagram showing the deployment of detectors along a pipeline, according to an embodiment.

[0056] FIG. 3 is a block diagram of a system for anomaly detection based on various inputs, according to an embodiment.

[0057] FIG. 4 is a flowchart of a training path, a baselining path, and an execution path for anomaly detection, according to an embodiment.

[0058] FIG. 5 is an example feature space with different signatures distributed, according to an embodiment.

[0059] FIG. 6 is an example diagram visually indicating anomaly detection for a channel based on an anomaly score, according to an embodiment.

[0060] FIGS. 7A and 7B are example bar graphs indicating the change of an anomaly score over time, according to an embodiment.DETAILED DESCRIPTION

[0061] Embodiments herein utilize detectors deployed along a pipeline to obtain measurements at various locations of the pipeline, such that fluid flow conditions within the pipeline can be monitored in real time. These detectors not only enhance the safety and efficiency of pipeline operations, but also provide operators with actionable insights to prevent potential pipeline problems and ensure smooth fluid transportation.

[0062] One type of detector is a variety of meters placed in or on a pipeline to monitor changes in parameters such as temperature, pressure, and flow. These meters may be an integral part of Computational Pipeline Monitoring (CPM) systems, which convert real-time measurements into valuable data that can be used to detect anomalies or inefficiencies in pipeline operations. CPM uses the data obtained from these types of detectors to provide an understanding of pipeline health, allowing deviations from standard operating parameters to be addressed in a timely manner.

[0063] Another type of detector is an optical fiber (also referred to as a fiber optic sensor) placed along a pipeline to measure changes in temperature and strain, for example, with high spatial resolution. This optical sensing approach can take advantage of three types of scattering: Rayleigh, Brillouin and Raman. While Rayleigh scattering results from variations in the fiber's refractive index, Brillouin scattering is induced by acoustic waves generated by light-fiber interactions and provides insight into temperature and strain changes. Raman scattering, resulting from the interaction of light with the fiber's vibrational modes, provides nuanced data on chemical composition, temperature and strain. Fiber Bragg gratings (FBGs) may enhance this technique by reflecting Raman-scattered light for high-resolution detection. However, the inherent properties of Raman scattering itself do not require reflectors. Distributed Acoustic Sensing (DAS) is a technology that leverages the entire length of the optical fiber as a sensor. It enables the continuous monitoring of acoustic vibrations along the pipeline, making it possible to detect events such as leaks, turbulences, intrusions, or equipment malfunctions. DAS works by analyzing the backscattered light from the fiber caused by strain- induced changes, allowing for real-time, multi-dimensional monitoring of the pipeline's conditions. By combining optical fibers with traditional meter-based detectors andincorporating FBG, DAS, Raman scatering techniques, etc., operators can deploy a comprehensive multi-dimensional monitoring approach. This approach enhances reliability and accuracy in capturing various aspects of pipeline conditions.

[0064] Referring now to FIG. 1, an embodiment of a system 100 for fiber optic sensing using backscatering occurring in an optical fiber is shown. This embodiment illustrates a working environment for simultaneously measuring temperature and strain using Raman scattering, as described herein. The system 100 comprises an optical fiber 112, an interrogator 106 optically coupled to the optical fiber 112, and a signal processing device (controller) 118 in communication with the interrogator 106. Although not shown in FIG. 1, the interrogator 106 includes an optical source, an optical receiver, and an optical circulator. The optical circulator directs light pulses from the optical source to the optical fiber 112 and directs light pulses received by the interrogator 106 from the optical fiber 112 to the optical receiver.

[0065] The optical fiber 112 comprises one or more fiber optic strands, each of which is made of silica (amorphous SiCh). The optical fiber strands are doped with a rare earth compound (such as germanium, praseodymium, or erbium oxides) to alter their refractive indices, although in alternative embodiments the optical fiber strands may be un-doped. Single mode and multimode optical fiber strands are commercially available from, for example, Coming® Optical Fiber. Examples of optical fibers include ClearCurve™ fibers (bend insensitive), SMF28 series single mode fibers such as SMF-28 ULL fibers or SMF-28e fibers, and InfiniCor® series multimode fibers.

[0066] The interrogator 106 generates sense and reference pulses and outputs the reference pulse after the sense pulse. The pulses are transmited along the optical fiber 112, which includes an FBG 114. The FBG 114 acts as a reflector as discussed above, but it should be understood that the reflector, such as the FBG 114 shown in FIG. 1, is not necessary for Raman scattering to occur, and the system 100 can function without the FBG 114.

[0067] In this embodiment, the interrogator 106 emits laser light toward the optical fiber 112, and the FBG 114 partially reflects the light back toward the interrogator 106. An optical receiver (not shown) detects the backscatered light and generates an output signal. The signal processing device (controller) 118 is communicatively coupled to the interrogator 106 to receive the output signal. The signal processing device 118 includes a processor 102 and a non-transitory computer readable medium 104 communicatively coupled to each other. Aninput device 110 and a display 108 interact with the processor 102. The computer readable medium 104 has statements and instructions encoded thereon to cause the processor 102 to perform any appropriate signal processing methods on the output signal. For example, if the fiber segment 116 is laid adjacent to a region of interest that simultaneously experiences vibration at a rate below 20 Hz and acoustics at a rate above 20 Hz, the fiber segment 116 will experience similar strain and the output signal will comprise a superposition of signals representative of that vibration and those acoustics. To isolate the vibration portion of the output signal from the acoustics portion of the output signal, the processor 102 may apply a low pass filter with a cutoff frequency of 20 Hz to the output signal. Similarly, the processor 102 may apply a high pass filter with a cutoff frequency of 20 Hz to isolate the acoustic portion of the output signal from the vibration portion. The processor 102 may also apply more complex signal processing methods to the output signal; example methods include those described in PCT application PCT / CA2012 / 000018 (publication number WO 2013 / 102252), the entirety of which is hereby incorporated by reference.

[0068] FIG. 2 illustrates an example arrangement of detectors along a pipeline, identified as 210. This arrangement plays a role in ensuring comprehensive, real-time monitoring of fluid flow and other parameters within the pipeline. The detectors employed in this embodiment serve a dual purpose: to capture accurate data about the conditions of the pipeline, and to alert operators to any anomalies or deviations that may compromise the integrity of the pipeline.

[0069] One or more optical fibers, referred to as 220, are laid out along the length of the conduit 210. These fibers may be segmented into N different channels, labeled from 220i to 220N. Each of these channels is responsible for monitoring a specific section of pipeline 210, ensuring granularity and specificity in data collection. In the case of FBGs, a channel emerges as a sensing region delineated by two FBGs. For example, a channel is the length of optical fiber between two specific sensing points or regions, often delineated by a pair of FBGs or other types of reflectors. Strain-induced changes or variations in temperature along this channel alter the reflected or backscattered light from the FBGs, which can be analyzed to detect changes in the fiber’s properties. Numerous pairs of FBGs scattered along the fiber delineate multiple channels, each representing a discrete sensing region. These regions, in turn, map to specific real-world locations along the pipeline. For example, in FIG. 2, Channel 1 (220i) may monitor the 0m to 5m segment, while Channel 2 (22O2) may monitor the 5m to 10m segment.If an anomaly of interest is detected in “channel 2”, operators can pinpoint the exact section of pipeline that needs to be further inspected.

[0070] It should be understood that, although FBGs are described as an example to explain what constitutes a channel, the system or method disclosed herein is not limited to FBGs. Other sensing techniques, such as distributed acoustic sensing (DAS), may also be used to define sensing regions along the optical fiber.

[0071] In the example shown in FIG. 2, conventional (non-optical) meters are placed at selective locations along the pipeline 210. At a first designated location, a first set of meters, including a first temperature sensor 231, a first pressure meter 232, and a first flow meter 233, may be installed. Further downstream, a second set of meters, including a second temperature sensor 241, a second pressure meter 242, and a second flow meter 243, may be installed. These meters, working in conjunction with the optical fiber 220, form a multi -pronged data collection mechanism that captures changes in temperature, flow, pressure, and other desirable parameters.

[0072] In one embodiment, the deployment of non-optical meters along the pipeline 210 is configured to facilitate data collection and monitoring efficiency. The first set of meters, including the first temperature sensor 231, first pressure meter 232, and first flow meter 233, may be positioned at a first location in relation to a pipeline. The first temperature sensor 231 may be installed either on the outer surface of the pipeline to measure the environment temperature or within the pipeline to directly monitor the fluid temperature. The first pressure meter 232 and the first flow meter 233 may be installed within the pipeline to measure fluid pressure and flow rate, respectively. The second set of meters, comprising the second temperature sensor 241, second pressure meter 242, and second flow meter 243, may be positioned at a second location in relation to a pipeline, which is different from the first location. The second temperature sensor 241 may also be installed on the outer surface or within the pipeline, while the second pressure meter 242 and second flow meter 243 may be installed within the pipeline. This arrangement allows monitoring of flow, pressure, and temperature conditions at multiple locations. Additional sets of meters may be deployed at intervals corresponding to the number of optical fiber channels or other selected locations along the pipeline. For example, in an arrangement where the optical fiber 220 is segmented into N channels, each channel may correspond to a discrete sensing region, and non-optical metersmay be deployed within or adjacent to some or all of these regions to provide complementary data.

[0073] A notable challenge in pipeline monitoring is the presence of “active channels”. These are channels located near areas of increased acoustic activity, such as near pumps or pipe bends, where flow dynamics create constant acoustics. Such channels often become hotspots for false alarms, especially when using signature-based algorithms that rely on supervised training. The anomaly detection (AD) system used herein helps to avoid this problem. After a learning phase, the AD system recognizes these regular acoustic patterns by default, reducing false positives. Over time, the system adapts to consider consistent acoustic signatures from these channels as the “norm”, providing more accurate monitoring results. The implementation of the system is described in more detail below.

[0074] FIG. 3 provides a visual representation of a system designed for enhanced anomaly detection in pipelines according to an embodiment described herein. Central to this figure is block 310, the anomaly detection process, which acts as the central processing unit of the system. This block receives input data from various detectors, sensors, or gauges deployed along the pipeline to collect a comprehensive set of real-time data.

[0075] In the context of pipeline monitoring and anomaly detection, operational data refers to the real-time or near real-time data generated and collected by sensors, detectors, gauges, fibers, and / or meters deployed along the pipeline. This data provides insight into the ongoing conditions and performance metrics of the pipeline system. It includes parameters such as flow, pressure, temperature, acoustic activity, strain, vibration, and / or other relevant metrics. Operational data serves as a source of information for assessing the condition of the pipeline, enabling timely decision making and anomaly detection.

[0076] Operational data according to embodiments of the present disclosure is obtained from non-optical detectors and / or optical fibers.Operational data sourced from non-optical fiber detectors• Flow meters: The first flow meter 233 and the second flow meter 243, as shown in FIG. 2, are instruments designed to measure the amount of fluid flowing through a section of pipeline. These meters may be constructed using various technologies, including ultrasonic, turbine, or electromagnetic sensing methods.Their exact locations along the pipeline are determined based on actual needs, and are typically located at intervals or at key junctions to ensure accurate fluid flow readings. It should be appreciated that more than two flow meters may be used in other embodiments.• Pressure meters: The first pressure meter 232 and the second pressure meter 242, as shown in FIG. 2, measure the fluid pressure within the pipeline. These meters may include piezoelectric or capacitive sensors to detect pressure changes. Like flow meters, they are placed at specific intervals or points of interest along the pipeline where the measurement of pressure data is desired. It should be appreciated that more than two pressure meters may be used in other embodiments.• Temperature sensors: Both the first temperature sensor 231 and the second temperature sensor 241, as shown in FIG. 2, provide real-time temperature readings of the fluid. They may be resistance temperature detectors (RTDs) or thermocouples, which provide high accuracy and fast response times. The temperature sensors are placed along the pipeline to monitor the thermal dynamics of the fluid being transported. It should be appreciated that more than two temperature meters may be used in other embodiments.Operational data sourced from optical fibers• Differential and absolute temperature estimates: By analyzing backscattered light and / or back-reflected interferometry, the fiber provides estimates of temperature changes (differential) and absolute temperature values over segments of the pipeline, as described for example in PCT application PCT / CA2016 / 050560 (publication number WO 2016 / 183677), the entirety of which is hereby incorporated by reference. This helps in understanding localized heat fluctuations that may indicate anomalies.• Acoustic data: The fiber also senses vibration within an acoustic frequency range (such as from 20 to 20,000 Hz) along its length. The data generated from the vibrations can be used to identify anomalous activities, including leaks, turbulences, or intrusions. For example, a leak may produce a distinctiveacoustic pattern within the acoustic frequency range, which can be detected and analyzed by the optical fiber to facilitate the identification of the leak.• Acoustic-based flow estimates: By analyzing acoustic patterns, the system can infer the flow characteristics of the fluid in the pipeline. The data complements the direct measurements from non-optical flow meters.• Strain-based pressure estimates: An optical fiber can detect strain changes in the pipe walls due to internal pressure changes. This strain affects the optical fiber’s light properties, especially the phase difference between the reflected pulses, which allows for pressure estimation. This non-intrusive method complements traditional pressure meters and increases the accuracy in detecting anomalies throughout the pipeline.• Transfer function-based flow and pressure estimates: The system may also use transfer functions to predict flow and pressure conditions based on the interaction of multiple data sources, providing an additional layer of analytical depth.

[0077] Block 310 is configured to process multi-dimensional data obtained from different inputs to assess the operational integrity of the pipeline. By correlating and analyzing the data in real time, block 310 can make an informed prediction about potential anomalies of interest to facilitate timely intervention and reduce potential risks. The methods and mechanisms employed for anomaly prediction using the obtained data will be described in greater detail below with reference to FIGS. 4-6.

[0078] It should be understood that the embodiments shown in FIGS. 2 and 3 are illustrative configurations of detectors, sensors, fibers, gauges and / or meters for purposes of explanation and should not be construed as limiting. The actual number, type, and arrangement of these components may vary based on specific implementation requirements or preferences. While the inclusion of multiple detectors can improve accuracy, it is conceivable that effective anomaly detection may still be achieved using operational data from a single detector, such as an optical fiber, to lower the possibility of false alarms. As the technology evolves and operational requirements dictate, modifications and variations to the configurations presented are possible.

[0079] FIG. 4 illustrates a flowchart of an anomaly identifying method including a training path 420, a baselining path 430, and an execution path 410, according to an embodiment.

[0080] The execution path 410 is configured to function as a real-time monitoring component of the method, assessing the operational health of the pipeline. At block 411, operational data is acquired from a detector deployed along a pipeline over a first time period. As described above, the operational data refers to the real-time or near real-time data generated and collected by at least an optical fiber and optionally one or more non-optical detectors deployed along the pipeline. The operational data may include parameters such as temperature, strain, or flow rate. The method may collect the operational data for the first time period, such as the most recent 30 seconds, 60 seconds, 90 seconds, or any other predetermined duration. The operational data may be collected or sampled at a predetermined frequency, such as 5 Hz, 10 Hz, 15 Hz, or any other predetermined frequencies.

[0081] Optionally, the operational data may be stored over time in a buffer (not shown), and the first period of operational data may be acquired from the buffer as needed. Alternatively, the first period of operational data may not be pre-stored in a buffer, but instead may be retrieved in real time directly from the detectors as they acquire and transmit the data. In such a setup, a continuous stream of operational data flows from the detectors, allowing for immediate processing and analysis. In the case of real-time data retrieval, the first time period can be construed as a short duration, often defined by the immediacy of the data processing requirements or the frequency with which the detectors collect and transmit data. In some implementations, a hybrid model may be used where the data of immediate interest is streamed in real time for immediate action, while the remainder is buffered for in-depth analysis at a later time or for maintaining historical records.

[0082] After obtaining the initial operational data, the method may optionally proceed to block 412, which involves selecting active channel(s). Due to the dynamic nature of pipeline operations, not every channel on the fiber exhibits consistent activity. Therefore, the method identifies and prioritizes certain channels that frequently exhibit significant activity, referred to herein as “active channels”. Prioritizing these active channels allows the method to focus on regions where anomalous events are of interest or more likely to occur. It should be noted that, depending on the specific setup and computing capacity, an optical fiber may have only a single dedicated channel or may support simultaneous monitoring of multiple channels.

[0083] After channel selection, the operational data is transformed using a feature extraction model at block 413. This feature extraction model, which may be pre-trained by an external third party or trained through the training path 420, is a specialized neural network configured to encode the collected operational data into a set of signatures. These signatures act as compact representations of the original data, capturing aspects of interest or relevance for efficient subsequent analysis. Each data point collected by a detector may generate a unique signature comprising multiple features. In this context, a “signature” refers to a compressed numerical representation of the original data point, while a “feature” refers to an individual characteristic or attribute represented by the signature. The combination of features represent the essence of the data in a way that facilitates effective anomaly detection.

[0084] It should be appreciated that the feature extraction model in this example may be an autoencoder, which is a specialized neural network chosen to encode operational data into compact signatures, but it is not the only method for achieving this goal within the scope of the present disclosure. Various other feature extraction models may be used to accomplish the same purpose. These alternative models include techniques such as statistical feature extraction, principal component analysis (PCA), t-distributed stochastic neighbor embedding (t-SNE), or the use of embeddings derived from pre-trained neural networks. Statistical feature extraction captures essential statistical properties of the data, while PCA reduces dimensionality by identifying orthogonal components that capture significant variance. t-SNE is used to visualize high-dimensional data in lower-dimensional spaces while preserving relative distances. Pre-trained neural network embeddings use the knowledge encoded in the neural network weights during pre-training to generate meaningful feature representations. Each of these models may offer distinct advantages and can be selected based on specific requirements or data characteristics, serving as viable alternatives to the autoencoder for creating signatures that facilitate anomaly detection. Therefore, it should be understood that the present disclosure is not intended to limit the type of the feature extraction model to the autoencoder.

[0085] In the context of the present disclosure, an autoencoder is an algorithmic structure configured for unsupervised learning of efficient encodings. The main operation of an autoencoder is to approximate the identity function with the goal of outputting a representation, which may be referred to as the “encoded” form, that closely matches its input, but in a more compact form. This structure is segmented into two main components: an encodermodule that reduces the dimensionality of the input data, and a corresponding decoder module that reconstructs the input data from this condensed representation. This condensed intermediate representation is referred to herein as an “encoded signature”.

[0086] An intrinsic feature of autoencoders, particularly relevant to the methods of the present disclosure, is their ability to detect and encapsulate salient features of the input data during the encoding phase. This facilitates the generation of encoded signatures that are compact and loaded with the most relevant attributes of the original data, thereby optimizing them for subsequent analysis processes.

[0087] To further elaborate on the operation of the autoencoder within the method disclosed herein, the autoencoder may incorporate a convolutional neural network (CNN) architecture. CNNs are specialized neural network architectures that are particularly effective at processing structured grid data, such as spatio-temporal matrices. The convolutional layers of the CNN apply specific filters to systematically scan the input data to identify and extract patterns at varying levels of complexity. As the data progresses through successive convolutional layers, the extracted patterns are transformed and consolidated into features that collectively form the encoded signature. This encoded signature serves as a compact representation of the original data for efficient anomaly detection and further processing.

[0088] In the transformation phase of the described system, as operational data passes through the CNN-integrated autoencoder, the convolutional layers iteratively extract a spectrum of features from the data. Through this extraction process facilitated by the convolutional operations, the method ensures that the resulting signatures are informationally robust, making them well-suited for subsequent anomaly detection processes. The training of the CNN-integrated autoencoder may involve using a training dataset that includes operational data associated with known events or normal conditions, as described in relation to a training path 420 shown in FIG. 4, which is described below. The model can be iteratively trained by minimizing the difference between the input data and its reconstruction to update weights and biases of the autoencoder for accurate feature extraction.

[0089] In accordance with the present disclosure, the term “feature space” refers to a multidimensional domain for data clustering. Each dimension within the feature corresponds to a distinct feature represented within a signature. For example, there is a one-to-one relationship between individual features and the components of a feature space vector, meaningthat each dimension of the feature space vector represents a single feature. However, in certain scenarios, the feature space may also include composite features. This can occur, for instance, when employing techniques such as PCA, where multiple original features are combined to create a more compact representation of the data within the feature space. This approach facilitates efficient data analysis and clustering while preserving the essential characteristics of the original data.

[0090] FIG. 5 illustrates the arrangement of data points within a two-dimensional coordinate system, where the abscissa (X-axis) represents one predefined feature and the ordinate (Y -axis) represents another predefined feature. When signatures include more than two distinguishable features, the spatial representation of feature space will require expansion into higher dimensions beyond the two-dimensional representation. Notably, in FIG. 5, individual events correlate with distinct distributions of signatures along the X and Y axes. These events, including but not limited to acoustics, intrusions, leaks, turbulences, and instances of low-quality data, are delineated within their respective regions in this plane. When a new measurement is acquired and then encoded by the autoencoder, the resulting signature is positioned within or mapped to a defined region within the feature space. This positioning or mapping facilitates categorization and association of the signature with a relevant event or anomaly.

[0091] It should be understood that computational methods can adeptly process and analyze data in higher dimensions. Using sophisticated algorithms, signatures derived from operational data may be positioned or mapped within a multidimensional feature space based on their feature values. This positioning or mapping allows similar signatures to be closely clustered, while divergent or anomalous signatures can be clearly separated. Over time, the plotted signatures reveal discernible patterns that aid in anomaly detection and trend identification. This computational approach provides a granular understanding of data relationships within the feature space.

[0092] For the purpose of effective data categorization within the present system, clustering methods, including but not limited to k-means clustering and Gaussian Mixture Models (GMM), may be utilized. These methods may facilitate the orderly grouping of signatures based on inherent feature similarities, resulting in distinct clusters or regions within the aforementioned feature space. By building these clusters using historical, training, ongoing, or real-time data sets, an operational framework is established for anomaly detection. When anew signature is generated and integrated into the predefined feature space, its spatial relationship to the pre-established clusters is determined. This relational positioning, in turn, enables the method to correlate the signature with known events, increasing the accuracy and effectiveness of anomaly detection.

[0093] In the context of the present disclosure, the clustering algorithm may be k-means clustering, a technique recognized for its efficiency and simplicity. This algorithm requires the pre-specification of the number of clusters, denoted “k”, and attempts to minimize the intracluster Euclidean distance between signatures. Once the cluster centroids have been defined based on the input data (for example, historical long-term data), newly derived signatures can be quickly assigned to the nearest cluster using Euclidean distance as a metric.

[0094] However, it should be appreciated that the clustering process can also be achieved using alternative distance metrics, such as Manhattan distance and cosine similarity, which are also effective. Manhattan distance measures the absolute differences along each dimension of the feature vectors, providing an alternative way to assess similarity. Cosine similarity, on the other hand, focuses on the direction of the vectors in feature space, making it suitable for capturing angular differences. These alternative metrics, like Euclidean distance, enable the grouping of signatures into clusters based on their respective distances, ensuring efficient and accurate clustering for subsequent analysis.

[0095] The process begins by encoding a new data point to produce its corresponding signature. When projected into feature space, the proximity of this signature to previously established cluster centroids is determined. For example, if the signature is close to the centroid of a cluster for an anomaly of interest (such as a leak, an intrusion, or a turbulence), the method determines that an anomaly of interest is occurring. On the other hand, if the signature deviates significantly from the centroid of a particular cluster - especially the cluster for the anomaly of interest - the method knows that the corresponding data point is less likely to be classified as an anomaly of interest, such as a leak, an intrusion, or a turbulence in the flow, thus increasing the specificity and sensitivity of the system’s anomaly detection capabilities.

[0096] In the context of the present disclosure, the GMM may serve as an alternative tool for clustering signatures within the intended feature space. When signatures are processed, distinct and informative representations of the original data is generated. These transformedsignatures exhibit an intrinsic property: those that share analogous features will gravitate toward a common region within the feature space.

[0097] Adaptive GMM enhances the dynamism of clustering by allowing the creation of new clusters in response to the introduction of new signatures that deviate from the recognized patterns. This adaptability facilitates continuous refinement of the model and accommodates evolving data trends specific to individual channels. If a newly encoded signature matches a previously observed event, it is assimilated into the existing cluster designated for that event. In the process, the cluster’s characteristics may be adjusted to account for the new data, and an internal counter keeps track of its usage.

[0098] Based on this clustering mechanism, the method may also determine an “anomaly score”. This score may be inversely proportional to the frequency of occurrence of a signature and serves as a probabilistic measure of its novelty. In other words, if a particular signature repeats itself, after a while it will be considered normal and not anomalous. If an event signature, previously unidentified for a particular channel, manifests itself repeatedly and meets a predetermined threshold - whether as an upper bound or a lower bound depending on the scoring criteria - an alert mechanism is activated. This methodological approach, which leverages the strengths of the Adaptive GMM, provides the method with a balance of sensitivity and specificity in anomaly detection. It should be appreciated that other clustering means than the Adaptive GMM may be used, such as hierarchical clustering, mean shift, spectral clustering, etc., offering flexibility and adaptability in selecting the most suitable clustering method for specific application scenarios.

[0099] Referring again to FIG. 4, the method enters a comparative analysis phase at block 414. Here, the encoded signatures are compared with baseline clusters, which are available from an external source or have been obtained through the baselining path 430. The term “baseline cluster” in this context refers to a collection of data points associated with previously identified and categorized events (also referred to as known events). Because these events have been repeatedly observed and characterized in past operations, their signatures have been known or recognized and structured within the feature map. This organized clustering within the feature space enables rapid identification and categorization of new data based on its proximity to these pre-established clusters.

[0100] In light of the above comparison, an anomaly score is calculated - a numerical metric that quantifies the degree of deviation of the newly acquired data from established norms. Specifically, this score can measure the deviation from the centroid of a pre-identified cluster corresponding to a particular event of interest, such as a leak within a particular channel of the pipeline, an intrusion into the pipeline, or a turbulence in the flow within the pipeline. This calculated anomaly score is then subjected to an evaluation at block 415, referred to as “Detect Event (EIG)” - where EIG stands for Event Identification and Grouping.

[0101] Within block 415, the anomaly score is compared with a predetermined threshold to determine whether the detected event or anomaly is consistent with or deviates from expected operational patterns. If the derived score meets this predetermined threshold, the method proceeds to block 416 to activate an alerting mechanism. Such activation serves as an indication of potential irregularities in operational behavior, thereby signaling that the event in question merits further analysis and investigation.

[0102] With respect to the training path 420, its role is to ensure the robustness and accuracy of the system’s real-time or near real-time (i.e., with a latency for data sampled over a short period of time) anomaly detection capabilities. The training may be performed either by an external third party, thus enabling its use as a standalone product, or directly by the user wishing to perform anomaly detection. The training is initiated at block 421, where a training dataset is provided. This training dataset includes operational data systematically associated with corresponding training signatures. While the dataset may be labeled or unlabeled, the data that is known does not contain the anomaly of interest. For example, the dataset may include data collected during a pig run or flow change, i.e., it deviates from normal baseline operating conditions but does not contain the specific anomaly of interest. The sourcing of this dataset may vary; it can be obtained from external repositories or accumulated on site over an extended period of time, thereby covering a wide range of operational states and events. This comprehensive dataset provides a robust foundation for training the feature extraction model, so that the model is calibrated to the specific characteristics and patterns relevant to the operational environment.

[0103] At optional block 422, data may be selected from specific active channels for use in the training process. This channel selection depends on the complexity and structure of the data pipeline. If the optical fiber contains only a single channel or if the model is intended to be trained uniformly across all available channels, this selection step may be omitted.However, in scenarios where multiple channels are present and certain channels exhibit increased activity or criticality, focusing on these active channels during training may improve the accuracy and relevance of the model.

[0104] At block 423, the training dataset is introduced to the feature extraction model, such as an autoencoder. During this phase, the internal parameters of the feature extraction model, such as its weights and biases, are iteratively refined. The refinement of these parameters is guided by the differences between the signatures encoded by the feature extraction model and the training signatures of interest found in the training dataset. This iterative training continues, recalibrating the feature extraction model, until either a predetermined level of accuracy is achieved or a predetermined number of iterations have been completed.

[0105] Upon completion of the sequence at block 424, a trained feature extraction model is obtained. This feature extraction model is trained using both normal operational data and data related to known historical anomalies of interest, such as leaks, intrusions, or turbulences. Once trained, the model is integrated into block 413, as described above with respect to the execution path 410. The trained model is configured to encode input operational data (either for a given number of samples or in real time), enabling the precise detection of anomalies.

[0106] Various feature extraction models may be used in the embodiments described herein. For example, the autoencoder may be useful for feature extraction in pipeline anomaly detection applications due to its ability to leam compact, meaningful representations of operational data while preserving information for anomaly detection. For large training datasets, batch training or distributed computing techniques may be used to ensure efficient processing. In addition, pre-trained models from external sources may be fine-tuned using the training dataset to adapt to site-specific characteristics.

[0107] Moving to the baselining path 430, its role is to provide a contextual basis against which real-time or near real-time (i.e., with a latency for data sampled over a short period of time) operational data can be assessed and evaluated. The initiation of this path begins at block 431 with the acquisition of a specific set of baseline operational data, referred to in FIG. 4 as historical long-term data. The baseline operational data aggregates measurements, estimates, or values from the detector(s) over a predetermined second time period, typicallyseveral days or even weeks (hence “long-term”), encapsulating a wide range of operational events, both normative and anomalous. Through this process, the majority of standard conditions associated with the relevant pipeline channels are effectively cataloged. As a result, their signatures are delineated within the feature space, resulting in distinct clusters corresponding to already recognized events.

[0108] At block 432, there may be an optional step of selecting active channels for the baselining process. The discretion to perform this selection is based on the number of channels available. It should be understood that, as with blocks 412 and 422, this step can be omitted if only one channel is present or if the intent is to perform baselining on all channels.

[0109] Transitioning to block 433, the acquired historical long-term data undergoes a transformation process. Utilizing the feature extraction model, which may be externally sourced or derived from the training path 420 and finalized at block 424, the historical data is transformed into baseline encoded signatures. This encoding process distills the complexities of the data into compact signatures or representations to facilitate subsequent clustering operations.

[0110] Proceeding to block 434, the resulting baseline encoded signatures are subjected to a clustering process within the feature space, culminating in the generation of baseline clusters. These clusters, which are derived by merging signatures that share inherent feature similarities, represent the various operational states and events to which the system has been exposed over time (e.g, over the second time period, usually substantially longer than the first time period). The spatial distribution and density of these baseline clusters within the feature space may be used for the establishment of the threshold as described with respect to block 415. This threshold may also be set or updated manually. Alternatively, the threshold can be determined automatically through event simulation, by selecting thresholds that lead to the simulated event being detected as an anomaly that falls outside of the baseline clusters. This data-driven approach may enhance the adaptability and accuracy of the anomaly detection, allowing it to dynamically adjust to changing operational conditions and detect previously unseen anomalies.

[0111] In addition, there may be an inherent dynamism to the baselining process. The baseline clusters and the threshold are not static entities; they are subject to updates based on operational data collected over a specified third time period. Such periodic updates ensure thatthe baselining process remains attuned to the evolving operational conditions of the pipeline. This dynamic adaptation may be advantageous when operational conditions change to the point where they trigger alarms due to the introduction of new data, which is indicative of anomalies, yet may not necessarily represent events of interest, such as leaks, intrusions, or turbulences. In such scenarios, the utilization of a sliding window approach within the baselining process may be advantageous. With this approach, operational data from a sliding time window is given precedence, ensuring that the baseline clusters and thresholds consistently reflect more recent operational nuances. By doing so, transient anomalies caused by changes in operational conditions that are not of concern are effectively fdtered out, allowing the system to maintain its focus on identifying events that are of interest. An example regarding this approach is discussed below in FIGS. 7A and 7B, illustrating its practical implementation within the system.

[0112] In a particular embodiment, the anomaly identification method may be primarily limited to the execution path 410. For example, in situations where the auto-encoder is provided pre-trained by an external entity (e.g, a machine learning service provider utilizing services from Google™ or Microsoft™) and the baseline clusters (as described with respect to block 434) are also externally sourced (e.g, baseline clusters derived from publicly available industrial datasets, such as those from Kaggle™, or from third-party anomaly detection solutions utilizing services from Siemens™ or GE™ tailored to similar operational environments), the method comprises primarily using the pre-existing and pre-trained autoencoder to extract the signature from operational data collected during the first time period. This extracted signature is then compared with the pre-established set of baseline clusters. Based on this comparison, the anomaly score can be determined.

[0113] In another embodiment, the anomaly identification method may be extended to include the baselining path 430. This is particularly relevant when the baseline clusters are not readily available for use. Under such circumstances, the method integrates the creation of baseline clusters. Historical long-term data is used as a foundation to facilitate the gradual accumulation of signature clusters within the feature space. These clusters are linked to a variety of events observed during the baselining phase, thereby allowing subsequent operational data to be compared with these foundational clusters.

[0114] In yet another embodiment, the anomaly identification process may be extended to include the training path 420. This scenario becomes relevant when a pre-trained featureextraction model is absent or not readily available for immediate use. In such cases, the method requires the training phase of the feature extraction model. Here, a dataset is used to iteratively train the feature extraction model so that it can generate signatures and clusters in the feature space. For example, the dataset may include operational data collected from pipeline detectors, such as flow rate, pressure, and temperature readings recorded during normal operating conditions or controlled tests, such as pig runs or maintenance cycles. The dataset may also include data from diverse pipeline scenarios, such as varying flow regimes or transient conditions, for robust training. The iterative training process optimizes the parameters of the feature extraction model, such as weights and biases, ensuring its effectiveness in accurately encoding operational data and enabling accurate anomaly detection in subsequent applications.

[0115] It should be explicitly clarified that the training path 420 and the baselining path 430 are distinct processes, each contributing individually to the anomaly identification method disclosed herein. In certain implementations, the execution path 410 may operate autonomously when both the feature extraction model and the baseline clusters are obtained preconfigured or from external sources. However, in scenarios where either component is not readily available or requires custom configuration tailored to specific operating conditions, the execution path 410 may follow the processes described in either the training path 420 or the baselining path 430. Further, it is conceivable that in certain operational circumstances, the execution path 410 may sequentially follow both the training path 420 and the baselining path 430 to ensure a comprehensive approach to anomaly identification that includes both training and baselining components.

[0116] Referring to FIG. 6, a representative feature space is shown in which various signatures are distributed according to a particular embodiment. This embodiment illustrates the generation of an anomaly score derived from operational data pertaining to a particular channel, specifically referred to as “Chl9”. Specifically, the clustering in this embodiment employs an adaptive Gaussian Mixture Model (GMM), which results in the accumulation of older data that manifests as distinct geometric representations within the feature space, as exemplified by the various shapes in FIG. 6.

[0117] Following the encoding phase described in execution path 410, as previously discussed in connection with FIG. 4, operational data is transformed into distinct signatures. These signatures, after being subjected to clustering, assume recognizable geometric configurations within the feature space. A subsequent evaluation follows, in which the newlyformed shape representative of the encoded operational data is compared with pre-existing baseline clusters. For example, if the resulting shape closely resembles an existing configuration, such as the diamond-shaped signatures in FIG. 6 within the area delineated by dashed lines and labeled “Chl9”, it can be inferred that its similarity to baseline clusters is high. Consequently, the resulting anomaly score may be attenuated, registering a value as low as 0.1, for example, on a scale of 0 to 1. Such a score indicates the potential normalcy of the data, suggesting a reduced likelihood that the data belongs to an unprecedented or abnormal event.

[0118] Conversely, if the shape is inconsistent with the recognized configurations within the baseline clusters, its dissimilarity is flagged. As an illustrative example, the pentagon-shaped signatures in FIG. 6 stand out as outliers when compared to established baseline clusters within the area delineated by dashed lines and labeled “Chi 9”. Such incongruence results in an escalated anomaly score, potentially rising to values as high as 0.9, for example, on a scale of 0 to 1. Such an elevated score serves as an indicator that the data may belong to an anomalous or previously unobserved event, thereby triggering the activation of an alert mechanism.

[0119] For example, with particular emphasis on the adaptive GMM clustering approach, the calculation of the anomaly score may be performed by keeping a record of the temporal occurrence of each signature across individual channels. This recording facilitates an understanding of the frequency with which particular signatures occur, thereby facilitating the calculation of the anomaly score.

[0120] In this context, the anomaly score may be designed to be proportional to the inverse of the frequency of the signature. This relationship may be succinctly expressed by the following equation:where n represents the frequency of occurrence of the signature.

[0121] By way of illustration, a scenario can be considered where a particular signature manifests with different frequencies, such as 1, 10, 100, and 1000 times. Corresponding to these frequencies, the anomaly scores calculated using the above equation would be 1, 0.5, 0.25, and 0.125, respectively. It is clear that as the frequency of the signature’s appearanceincreases, the anomaly score decreases. This inverse proportionality ensures that frequently observed signatures are assigned a lower anomaly score, emphasizing their routine nature, while rare signatures, potentially indicative of anomalies, are highlighted with a higher score, prompting closer scrutiny.

[0122] It should be understood that the above example is provided solely for illustrative purposes. The specific scoring format and approach described herein represent one possible implementation among numerous alternatives and are not intended to limit or restrict the scope of the disclosure.

[0123] Those skilled in the art will readily appreciate the versatility inherent in the field of anomaly detection. Various alternative scoring formats or computational approaches may be conceived and implemented to meet specific operational requirements or to address unique challenges. Such modifications, alterations, and innovations, while departing from the illustrated example, remain firmly within the scope of the present disclosure and its underlying principles.

[0124] In the context of effective monitoring and management of operational data, particularly in dynamic environments, in one embodiment, the anomaly identification method may additionally include a process for not only detecting anomalies of interest, but also for responding to recent events in a timely manner. Further, there may be a need to prevent the system from remaining overly sensitive to singular events that occurred a significant amount of time ago, allowing for the natural attrition of their relevance over time.

[0125] To meet these requirements, the system may consider the frequency of events over different time frames, such as hourly, daily, weekly, and monthly intervals. By analyzing the frequency of events over these different durations, a more nuanced understanding of their recurrence and potential significance is achieved.

[0126] As a result, the anomaly score is adaptively adjusted based on the frequency of events within these specified time frames. This adaptive approach enables several beneficial behaviors:• For ongoing events that are currently impacting the system, the anomaly score quickly drops below the pre-determined threshold. This rapid decrease ensuresthat repetitive alerts are not triggered for the same ongoing event, preventing alert fatigue and allowing for more focused remediation.• Conversely, for an event that occurred once and was not recent (e.g. , six months ago), its anomaly score gradually increases over time. This ensures that if such an event occurs again, the anomaly detection system will again be primed to recognize it as an outlier. In essence, the system is regaining sensitivity to events that have not recurred for an extended period of time.

[0127] This example may be considered a sliding window approach in which operational data from a sliding time window (such as the most recent six months, as described in the example above) is used to update the baseline clusters and threshold and the baseline encoded signatures encoded prior to the time window may be removed from the baseline clusters.

[0128] For example, as shown in FIGS. 7A and 7B, example bar graphs indicating the change of an anomaly score over time are shown. In FIG. 7A, when something new happens, the anomaly score (shown in y-axis) is high across all time frames initially, reflecting an immediate and strong response to a recent event. As time passes, fewer bars remain, suggesting that the sensitivity diminishes as the event becomes less recent. Conversely, FIG. 7B demonstrates the anomaly scores (shown in y-axis) following the same event after a period of 10 days; the scores are substantially reduced at the hourly, daily, and weekly marks, indicating the system’s adaptive response has already adjusted to the event, reducing the likelihood of redundant alerts. However, at the monthly interval, the scores are comparably high, maintaining a level of alertness for potential recurrence of less frequent events. This visual representation underscores the system’s capacity to adjust the anomaly scores over different time frames, embodying a sliding window approach that balances between immediate response and long-term sensitivity adjustments.

[0129] In this example, the above mechanisms may ensure that the anomaly detection system remains both responsive to current events and adaptive to the evolving nature of the operational environment. They strike a balance between providing timely alerts for new anomalies and reducing redundant alerts for known, ongoing events. It should be understood that this adaptive approach is merely illustrative, and those skilled in the art could readily derive or employ other methods or strategies consistent with the principles outlined herein.

[0130] In a particular embodiment, the operational data comprises a first set of data derived solely from one or more fiber optic sensors. These fiber optic based sensors are located along the length of the pipeline, either externally or internally. By segmenting each of the one or more optical fibers into one or more distinct channels, the first data set is subjected to a comparative analysis against the baseline clusters, as previously described. This comparison facilitates determining whether the operational data corresponds to any anomalous patterns, as indicated by the resulting anomaly score.

[0131] In an alternative embodiment, the operational data comprises a second set of data derived solely from one or more meters located along the pipeline. This second set of data is subjected to a comparative analysis against the baseline clusters, as described above, to determine whether the operational data exhibits characteristics indicative of an anomaly of interest, which is subsequently quantified by an anomaly score.

[0132] In a further embodiment, the operational data comprises both the first and second data sets. These data sets, which are derived from one or more fiber optic sensors and one or more meters, may either be subjected to a combined comparison or individually evaluated against the baseline clusters. When evaluated separately, each data set provides its own anomaly score. An alert is triggered when both scores meet a common threshold or independent thresholds. This two-tiered approach provides an added layer of security by ensuring that potential anomalies are corroborated by multiple data sources, thereby increasing the reliability and accuracy of the anomaly detection mechanism.

[0133] In the embodiment where the operational data includes both the first and second data sets, multiple setups of the system may be used. For example, a first setup of the system may include one or more separate pipelines. The fiber(s) may be deployed along one of the pipelines, while the meter(s) may be deployed in the same pipeline or in the other of the pipelines. The first data set is provided to a first anomaly detection device on which a first anomaly score is calculated according to the embodiments described herein; the second data set is provided to a second anomaly detection device on which a second anomaly score is calculated according to the embodiments described herein. The above calculation processes for the first data set and the second data set may be performed simultaneously or individually at different times. The calculated scores are then transmitted to a central server where they can be compared to a predetermined threshold or to each other.

[0134] Alternatively, a second setup of the system may include one or more separate pipelines. The fiber(s) may be deployed on one of the pipelines, while the meter(s) may be deployed along the same pipeline or along the other of the pipelines. The first and second data sets are simultaneously provided to a same anomaly detection device which calculates a first anomaly score and a second anomaly score. The calculated scores are then compared to a predetermined threshold or to each other on the same anomaly detection device or on a different server.

[0135] The embodiments have been described above with reference to flow, sequence, and block diagrams of methods, apparatuses, systems, and computer program products. In this regard, the depicted flow, sequence, and block diagrams illustrate the architecture, functionality, and operation of implementations of various embodiments. For instance, each block of the flow and block diagrams and operation in the sequence diagrams may represent a module, segment, or part of code, which comprises one or more executable instructions for implementing the specified action(s). In some alternative embodiments, the action(s) noted in that block or operation may occur out of the order noted in those figures. For example, two blocks or operations shown in succession may, in some embodiments, be executed substantially concurrently, or the blocks or operations may sometimes be executed in the reverse order, depending upon the functionality involved. Some specific examples of the foregoing have been noted above but those noted examples are not necessarily the only examples. Each block of the flow and block diagrams and operation of the sequence diagrams, and combinations of those blocks and operations, may be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0136] The terminology used herein is only for the purpose of describing particular embodiments and is not intended to be limiting. Accordingly, as used herein, the singular forms “a”, “an”, and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and “comprising”, when used in this specification, specify the presence of one or more stated features, integers, steps, operations, elements, and components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and groups. Directional terms such as “top”, “bottom”, “upwards”, “downwards”, “vertically”, and “laterally” are used in the following description for the purpose of providing relative referenceonly, and are not intended to suggest any limitations on how any article is to be positioned during use, or to be mounted in an assembly or relative to an environment. Additionally, the term “connect” and variants of it such as “connected”, “connects”, and “connecting” as used in this description are intended to include indirect and direct connections unless otherwise indicated. For example, if a first device is connected to a second device, that coupling may be through a direct connection or through an indirect connection via other devices and connections. Similarly, if the first device is communicatively connected to the second device, communication may be through a direct connection or through an indirect connection via other devices and connections.

[0137] Phrases such as “at least one of A, B, and C”, “at least one of A, B, or C”, “one or more of A, B, and C”, and “A, B, and / or C” are intended to include both a single item from the enumerated list of items (i.e., only A, only B, or only C) and multiple items from the list (i.e., A and B, B and C, A and C, and A, B, and C). Accordingly, the phrases “at least one of’, “one or more of’, and similar phrases when used in conjunction with a list are not meant to require that each item of the list be present, although each item of the list may be present.

[0138] It is contemplated that any part of any aspect or embodiment discussed in this specification can be implemented or combined with any part of any other aspect or embodiment discussed in this specification, so long as such those parts are not mutually exclusive with each other.

[0139] While every effort has been made to provide a detailed and accurate description of the disclosure herein, it should be noted that the scope of the disclosure is not limited to the exact configurations and embodiments described. The description provided is intended to illustrate the principles of the disclosure and not to limit the disclosure to the specific embodiments illustrated. It is intended that the scope of the disclosure be defined by the appended claims, their equivalents, and their potential applications in other fields.

Claims

CLAIMS1. A method comprising: acquiring operational data from one or more detectors deployed along a pipeline over a first time period, the one or more detectors comprising a fiber optic sensor deployed along the pipeline; transforming, by an feature extraction model, the acquired operational data into encoded signatures, wherein each of the encoded signatures contains a plurality of features; clustering the encoded signatures in a feature space using the features of the encoded signatures; for at least one of the clustered signatures, calculating an anomaly score by comparing the at least one of the clustered signatures with baseline clusters associated with known events; and identifying an anomaly of interest in response to the anomaly score meeting a threshold value.

2. The method of claim 1 , wherein the one or more detectors further comprise one or more non-optical meters deployed along the pipeline.

3. The method of claim 2, wherein the identifying further comprises identifying the anomaly of interest in response to both the anomaly score corresponding to the operational data acquired from the fiber optic sensor and the anomaly score corresponding to the operational data acquired from the one or more non-optical meters meeting the threshold value.

4. The method of claim 2, wherein the one or more non-optical meters comprises at least one of a temperature sensor, a pressure meter, a flow meter, a fluid density meter, or a viscosity meter.

5. The method of any one of claims 1 to 4, wherein the calculating further comprises calculating respective anomaly scores for all of the clustered signatures.

6. The method of any one of claims 1 to 5, further comprising storing the operational data over time in a buffer, wherein the operational data is acquired from the buffer.

7. The method of any one of claims 1 to 6, wherein the fiber optic sensor is segmented into a plurality of monitoring channels, and wherein acquiring the operational data further comprises acquiring at least one of acoustic data, acoustic based flow rate estimate, strain based pressure estimate, differential temperature estimate, absolute temperature estimate, or transfer function based flow rate and pressure estimate for at least one of the plurality of monitoring channels.

8. The method of any one of claims 1 to 7, wherein the transforming further comprises extracting the plurality of features from the acquired operational data by the feature extraction model utilizing a Convolutional Neural Network.

9. The method of any one of claims 1 to 8, wherein: the encoded signatures are clustered using k-means clustering or Manhattan distancebased clustering, which assigns a respective encoded signature into one of the baseline clusters by minimizing a distance between the respective encoded signature and centroids of the baseline clusters in the feature space; and the anomaly score is calculated using the distance between a respective clustered signature and a centroid of a baseline cluster associated with the known events.

10. The method of any one of claims 1 to 8, wherein: the encoded signatures are clustered using cosine similarity-based clustering, which assigns a respective encoded signature into one of the baseline clusters by minimizing an angle between the respective encoded signature and centroids of the baseline clusters in the feature space; and the anomaly score is calculated using the angle between a respective clustered signature and a centroid of a baseline cluster associated with the known events.

11. The method of any one of claims 1 to 10, wherein: the encoded signatures are clustered using one of an adaptive Gaussian mixture model, hierarchical clustering, mean shift, or spectral clustering; andthe anomaly score is calculated using a placement of the at least one of the clustered signatures in a designated cluster of the baseline clusters in the feature space and a frequency of occurrences of signatures within the designated cluster.

12. The method of any one of claims 1 to 11, further comprising dynamically updating the baseline clusters in response to the encoded signatures being clustered using features thereof.

13. The method of any one of claims 1 to 12, wherein the feature extraction model is trained using a training process comprising: a) providing a training dataset representative of operational data and associated training signatures; b) feeding the training dataset to the feature extraction model; c) adjusting parameters of the feature extraction model using differences between signatures encoded by the feature extraction model and the training signatures in the training dataset; and d) iterating b) and c) until a predetermined accuracy or a number of iterations is achieved.

14. The method of claim 13, wherein the training dataset comprises normal operational data and data associated with known historical anomalies of interest.

15. The method of any one of claims 1 to 14, wherein the baseline clusters are obtained by a baselining process comprising: acquiring a set of baseline operational data detected over a second time period; transforming the baseline operational data into baseline encoded signatures using the feature extraction model; and clustering the baseline encoded signatures in the feature space to generate the baseline clusters.

16. The method of claim 15, wherein the baselining process further comprises updating the baseline clusters using the operational data acquired over a third time period to accommodate changing operational conditions of the pipeline.

17. The method of claim 15, wherein the baselining process further comprises, over time: updating the baseline clusters and the threshold value using the operational data from a sliding time window; and removing the baseline encoded signatures that are encoded prior to the sliding time window from the feature space.

18. The method of any one of claims 1 to 17, wherein the feature extraction model is selected from one of autoencoder, statistical feature extraction, principle component analysis (PCA), t-distributed Stochastic Neighbor Embedding (t-SNE), or an embedding from a pre-trained neural network.

19. The method of any one of claims 1 to 18, wherein the anomaly of interest comprises at least one of a leak, an intrusion, or a turbulence.

20. A system, comprising: a pipeline; an anomaly detection apparatus; one or more detectors deployed along the pipeline, the one or more detectors comprising a fiber optic sensor deployed along the pipeline; wherein the anomaly detection apparatus is configured to perform: a) acquire operational data from the one or more detectors; b) transform, by a feature extraction model, the acquired operational data into encoded signatures, wherein each of the encoded signatures contains a plurality of features; c) cluster the encoded signatures in a feature space using the features of the encoded signatures;d) for at least one of the clustered encoded signatures, calculate an anomaly score by comparing the at least one of the clustered encoded signatures with baseline clusters associated with known events; and e) identify an anomaly of interest in response to the anomaly score meeting a threshold value.

21. The system of claim 20, wherein the one or more detectors further comprise one or more non-optical meters deployed along the pipeline.

22. The system of claim 21, wherein the anomaly of interest is identified in response to both the anomaly score corresponding to the operational data acquired from the fiber optic sensor and the anomaly score corresponding to the operational data acquired from the one or more non-optical meters meeting the threshold value.

23. The system of claim 21, wherein the one or more non-optical meters comprises at least one of a temperature sensor, a pressure meter, a flow meter, a fluid density meter, or a viscosity meter.

24. The system of claim 20, wherein the anomaly detection apparatus further comprises a server configured to communicatively couple with the anomaly detection device, and wherein the server is configured to receive the anomaly score from the anomaly detection device and perform e).

25. A non- transitory computer readable medium have stored thereon computer program code that is executable by at least one processor and that, when executed by the at least one processor, causes the at least one processor to perform: acquiring operational data from one or more detectors deployed along a pipeline over a first time period, the one or more detectors comprising a fiber optic sensor deployed along the pipeline; transforming, by an feature extraction model, the acquired operational data into encoded signatures, wherein each of the encoded signatures contains a plurality of features;clustering the encoded signatures in a feature space using the features of the encoded signatures; for at least one of the clustered signatures, calculating an anomaly score by comparing the at least one of the clustered signatures with baseline clusters associated with known events; and identifying an anomaly of interest in response to the anomaly score meeting a threshold value.

Citation Information

Patent Citations

  • Pipeline constriction detection

    CA2947915A1

  • Method for detecting anomalies in time series data produced by devices of an infrastructure in a network

    US11831527B2

Cited By

  • Water supply network pressure prediction model construction method, pressure prediction method and device

    CN121765692A

  • Gas pressure regulator outlet pressure set value determination method based on online monitoring data

    CN122064147A