Time-frequency fusion self-coding oil pipeline abnormal data reconstruction method and system

Through the time-frequency fusion autocoding method, the time-frequency domain information of oil pipeline data is extracted and fused, which solves the data abnormality caused by sensor aging and environmental dirt, and realizes the improvement of automatic data reconstruction and wall thickness prediction accuracy.

CN119988929APending Publication Date: 2025-05-13BEIJING INSTITUTE OF PETROCHEMICAL TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411953146.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the measurement of wall thickness of oil pipelines, sensor aging and environmental fouling lead to the generation of data outliers, affecting the accuracy of wall thickness prediction, and the existing data reconstruction methods are complex and difficult to adapt to data drift.

Method used

The time-frequency fusion autocoding method is adopted, and the local and global frequency domain information is extracted through the time-frequency domain data distribution analysis module and the global frequency domain feature learning module, and fused with the time-frequency data, and input the time-frequency autocoding reconstruction model for data reconstruction.

Benefits of technology

Automatic data reconstruction without artificial location outliers is realized, which enhances the model's adaptability to long-term timing data and generalizes data in special regions, and improves the accuracy of wall thickness prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988929A_ABST
    Figure CN119988929A_ABST
Patent Text Reader

Abstract

The invention provides a time-frequency fusion self-coding oil pipeline abnormal data reconstruction method and system, and the method comprises the steps: collecting and preprocessing oil pipeline parameter time sequence data, and obtaining a 2D time sequence data segment and a complete 2D time sequence data set; inputting the 2D time sequence data fragments into a time-frequency domain data distribution analysis module, and training and extracting local frequency domain information FH * 1; inputting a complete 2D time sequence data set into a global frequency domain feature learning module, training and extracting global frequency domain information # imgabs0 # of frequency domain data in a plurality of long-period sequences, adding and fusing FH * 1, # imgabs1 # and 2D time sequence data fragments, obtaining local-global fusion data fragments # imgabs2 #, inputting # imgabs3 # into a time sequence self-encoding reconstruction model, and training and reconstructing to generate a whole time sequence data segment. According to the method, the abnormal data of the pipeline can be reconstructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of abnormal data reconstruction, and in particular to a method and system for reconstructing abnormal data of a time-frequency fusion self-encoding oil pipeline. Background Art

[0002] The oil pipeline is a pipeline system specially used for transporting oil and gas, with key parameters such as internal pressure, temperature, wall thickness, etc. The wall thickness parameter is related to the oil pipeline's pressure bearing capacity, corrosion resistance, thermal efficiency, service life and other important performances. Therefore, timely measurement and acquisition of the oil pipeline's wall thickness parameters is crucial for the updated design and operation and maintenance of the oil pipeline.

[0003] At present, there are many ways to measure the wall thickness of oil pipelines, including physical cutting measurement, radiometric measurement, thickness gauge measurement, and inference prediction through other pipeline parameters. Physical cutting measurement will cause irreversible damage to the pipeline; radiometric measurement involves radioactive substances, which are harmful to the environment; and although thickness gauge measurement is non-destructive and does not involve radioactive substances, the equipment cost is expensive and requires additional personnel for equipment maintenance. Pipeline parameters such as internal pressure and temperature can be measured by sensors with lower costs, and they have a more obvious influence on wall thickness parameters. Therefore, collecting other pipeline parameters and inferring and predicting wall thickness through them is an emerging method for measuring the wall thickness of oil pipelines. However, due to the aging of electronic devices during long-term use, the performance of sensors will gradually decline, and dirt, oil or corrosive substances are easily accumulated inside and outside the pipeline. These attachments may hinder the normal operation of the sensor, resulting in inaccurate measurement of pipeline parameters of oil pipelines and prone to abnormal values. Traditional machine learning and deep learning prediction methods are based on data-driven, and abnormal values ​​in the data will have a serious impact on these methods. Therefore, it is necessary to detect and reconstruct abnormal values ​​in the collected parameter data.

[0004] Common data reconstruction methods can reconstruct data, but they require manual experience to judge outliers and locate them before reconstruction, which is complicated. In the field of deep learning, a generative approach can be used to automatically locate and reconstruct outlier data without manual work. However, most oil pipeline data are long-period time series data, which often involve complex distribution patterns and dynamic changes in distribution (i.e., data drift). Generative models are usually expected to be trained and applied on relatively stable data distributions. Data drift may lead to a decrease in the accuracy of generative model reconstruction. In addition, some new oil pipelines are laid in special geological areas (such as frozen soil and loess layers). Sensors in these areas are easily affected by low temperatures and wet soil erosion, causing them to be damaged in the early stages of use. The amount of data collected by these sensors in the early stages is not enough to reach the scale of long-period series. Models usually need to be trained on large-scale time series data to achieve better results, but it is difficult to improve their generalization ability on small-scale time series data. Summary of the invention

[0005] In order to solve the technical problems existing in the above-mentioned prior art, the present invention provides a method for reconstructing abnormal data of a time-frequency fusion self-encoding oil pipeline, and the technical solution is as follows:

[0006] On the one hand, a method for reconstructing abnormal data of an oil pipeline using a time-frequency fusion autoencoding method is provided, the method comprising:

[0007] S1, collecting and preprocessing the oil pipeline parameter time series data to obtain 2D time series data segments and a complete 2D time series data set;

[0008] S2, inputting the 2D time series data fragment into a time-frequency domain data distribution analysis module, and training the time-frequency domain data distribution analysis module to extract local frequency domain information F H×1 ;

[0009] S3, inputting the complete 2D time series data set into a global frequency domain feature learning module, and training the global frequency domain feature learning module to extract global frequency domain information of frequency domain data in multiple long-period sequences

[0010] S4, the local frequency domain information F H×1 , global frequency domain information Add and fuse with the 2D time series data fragment to obtain the local-global fused data fragment

[0011] S5, the local-global fusion data segment Input a time series autoencoder reconstruction model, and train the time series autoencoder reconstruction model to reconstruct and generate a whole segment of time series data;

[0012] S6. Reconstruct the abnormal data of the oil pipeline to be reconstructed using the trained overall abnormal data reconstruction model composed of the time-frequency domain data distribution analysis module, the global frequency domain feature learning module, and the time series autoencoder reconstruction model.

[0013] Optionally, the S1 specifically includes:

[0014] S11, collecting pipeline parameter time series data obtained by detecting the oil pipeline with various sensors, wherein the pipeline parameter time series data includes long-period time series data of the long-time pipeline laying and a small amount of time series data of special geological areas;

[0015] S12, performing mean normalization on the pipeline parameter time series data to eliminate the dimensional differences between different indicators so that they can be compared on the same scale;

[0016] S13, arranging the time series data of multiple parameter indicators in parallel, for long-period time series data, using a sliding window method, setting the window size and step size and traversing these long-period time series data, dividing them into multiple time period data, so that they are presented as 2D time series data segments, which are used in the time-frequency domain data distribution analysis module, and the complete 2D time series data set before segmentation is used in the global frequency domain feature learning module. For a small amount of time series data in special geological areas, they are directly arranged in parallel to form 2D time series data segments;

[0017] S14. Divide the 2D time series data set into a training set and a test set, wherein the training set requires normal data corresponding to the outlier part as labels for model training outlier reconstruction.

[0018] Optionally, the time-frequency domain data distribution analysis module converts a 2D time series data segment X H×W Each row of indicator data in is Fourier transformed to convert it from time domain data to frequency domain data segment Y H×L ;

[0019] Then, a multi-layer perceptron is used to perform preliminary feature extraction on the frequency domain data;

[0020] Since the generative model needs to specify a certain pipeline parameter indicator data as the generation target during use, two densely connected layers are used for further feature extraction to distinguish the specified indicator from other indicators. The densely connected layer is a stack of multiple linear layers and activation functions, and compared with the multi-layer perceptron, there is no hidden layer for deep learning feature extraction. Its function is to perform nonlinear transformation on the input data, map and extract it as a feature vector, and change its own matrix shape to facilitate the subsequent frequency domain feature fusion calculation. The first densely connected layer performs complete feature mapping extraction on the preliminary frequency domain features of the input to obtain The second dense connection layer is a selection dense connection layer. According to the user's specification, only the frequency domain features of the row corresponding to the selected parameter index are mapped and extracted to obtain Select the features extracted from the densely connected layer mapping Perform the transpose operation to obtain In order to integrate the extracted frequency domain features into the time series data fragment, the transposed matrix and Multiply to obtain the local frequency domain information F H×1 .

[0021] Optionally, the global frequency domain feature learning module converts a complete 2D time series data set Perform Fourier transform to convert it into long-period frequency domain data

[0022] Then the long period frequency domain data With the local frequency domain information F H×1 After addition and fusion, it is sent to the feature extraction layer composed of a multi-layer perceptron and a dense connection layer for deep frequency domain information extraction to obtain the global frequency domain information.

[0023] Optionally, the global frequency domain feature learning module also includes a probability freezing trainer to prevent the global frequency domain feature learning module from overfitting on a certain time series data when learning a lot of complete long-period time series data, resulting in poor generalization ability. The probability freezing trainer resets the weights of certain linear layers in the feature extraction layer to 0 according to the loss value returned by the loss function, so that the transmitted features of the model will not change after passing through these layers during forward propagation, and these layers will not update parameters during back propagation. Moreover, the smaller the returned loss value, the greater the probability that the feature extraction layer is frozen. Since the loss function calculation results are all between 0 and 1, the following formula is used to link the loss function and the freezing probability:

[0024]

[0025] Where P m is a freezing probability vector with only 0 and non-zero values. After multiplying with the feature extraction layer matrix, the corresponding layer is frozen. L is the loss function value. The larger the L value, the higher the P m The more 0 values ​​there are in , m is the length of the feature extraction layer.

[0026] Optionally, the temporal self-encoding reconstruction model comprises a temporal encoder and a temporal decoder;

[0027] The timing encoder splices the sequence number of the pipeline parameter timing data specified by the user into the local-global fusion data segment and the local-global fusion data fragment Each column of data is divided into time step data X, and then one-dimensional convolution is used to extract features of these time step data X in turn to obtain hidden layer features Y1, and then the hidden layer features Y1 are compressed by one-dimensional convolution to compress these features into a smaller subspace to form hidden layer features Y2. These hidden layer features Y1 and Y2 are feature vectors carrying rich information, which not only carry the normal and abnormal value position information of the input data, but also carry the information of the data distribution characteristics. The hidden layer features Y2 are compressed into latent variables Z by one-dimensional convolution. The latent variable Z is the most implicit feature, representing the abnormal value distribution and characteristic distribution of all data fragments, and is used for subsequent data decoding generation;

[0028] The time series decoder is completely mirror-symmetric with the time series encoder. After the latent variable Z passes through the 1D deconvolution in the time series decoder with different 1D convolution weights from the time series encoder, it is gradually upsampled to obtain new feature vectors Y'1 and Y'2. The feature vector Y'2 is then transformed into time step data corresponding to the input of the reconstructed values ​​X' through 1D deconvolution, and they are combined to obtain the entire segment of time series data reconstructed and generated.

[0029] Optionally, the loss function of the abnormal data reconstruction overall model includes a reconstruction loss function and a regularization term loss function;

[0030] The reconstruction loss function is used to measure the matching degree between the reconstructed data and the normal label data, and its formula is as follows:

[0031]

[0032] where X′ i is the reconstruction value, X i is the corresponding input time step value, n is the length of the input data;

[0033] The reconstruction loss function only compares the numerical difference between the reconstructed data and the real data. Even if the error between the reconstructed data and the real data is small, the data distribution trend may not be consistent. It is relatively rough to use the reconstruction loss function alone for back propagation training. Therefore, the loss function also includes a regularization term loss function, which is used to measure the matching degree between the encoding distribution of the latent variables and the input data distribution, and whether the measurement model is in place for the distribution change analysis of long-period time series data. The formula is as follows:

[0034] L 正则化 =KL(q(Z|X)||p(Z))

[0035]

[0036] Where KL represents KL divergence, q(Z|X) represents the conditional probability distribution of Z when the input is fixed to X, that is, the actual distribution of data encoding, p(Z) is the encoding distribution of the model output, and n is the length of the input data;

[0037] By combining the two loss functions, we can not only evaluate the reconstruction accuracy of the reconstruction model numerically, but also evaluate the reconstruction model's ability to analyze data distribution.

[0038] On the other hand, a time-frequency fusion self-encoding oil pipeline abnormal data reconstruction system is provided, the system comprising:

[0039] The acquisition and preprocessing module is used to acquire and preprocess the oil pipeline parameter time series data to obtain 2D time series data segments and a complete 2D time series data set;

[0040] The local frequency domain information extraction module is used to input the 2D time series data fragment into the time-frequency domain data distribution analysis module, and train the time-frequency domain data distribution analysis module to extract the local frequency domain information F H×1 ;

[0041] A global frequency domain information extraction module is used to input the complete 2D time series data set into a global frequency domain feature learning module, and train the global frequency domain feature learning module to extract the global frequency domain information of the frequency domain data in multiple long-period sequences.

[0042] A fusion module is used to combine the local frequency domain information F H×1 , global frequency domain information Add and fuse with the 2D time series data fragment to obtain the local-global fused data fragment

[0043] A reconstruction module is used to reconstruct the local-global fusion data fragment Input a time series autoencoder reconstruction model, and train the time series autoencoder reconstruction model to reconstruct and generate a whole segment of time series data;

[0044] The data reconstruction module to be reconstructed is used to reconstruct the abnormal data of the oil pipeline to be reconstructed using the trained abnormal data reconstruction overall model composed of the time-frequency domain data distribution analysis module, the global frequency domain feature learning module, and the time series autoencoder reconstruction model.

[0045] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned time-frequency fusion self-encoding oil pipeline abnormal data reconstruction method.

[0046] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned time-frequency fusion self-encoding oil pipeline abnormal data reconstruction method.

[0047] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0048] 1. Compared with traditional data reconstruction methods such as interpolation and filtering, the present invention reconstructs the entire data segment in a generative manner based on the time series autoencoder reconstruction model. It can reconstruct the data without manually locating the position of outliers, eliminating manual operation costs, and can combine other related data to reconstruct the target data more accurately.

[0049] 2. The time-frequency domain data distribution analysis module designed by the present invention is based on the fact that time domain features can reflect the real situation of the data, mean variance and other characteristics, while frequency domain features can reflect the distribution change trend and degree of change of the data. The time domain and frequency domain features of the input long-period time series data are extracted and fused to capture the data distribution changes in the long-period data, so as to adapt the model to the data drift phenomenon and enhance the model's ability to reconstruct outliers in the long-period time series data.

[0050] 3. The global frequency domain feature learning module designed in the present invention extracts and stores global frequency domain information while the model continuously learns the reconstruction of outliers in time series data, and fuses the global information with the local frequency domain information extracted from a small amount of data in special areas to enhance the model's generalized anomaly reconstruction capability on oil pipeline data sets in special areas. The global frequency domain feature learning module also includes a probabilistic freezing trainer to prevent the global frequency domain feature learning module from overfitting on a certain time series data when learning a large number of complete long-period time series data, resulting in poor generalization capability. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0052] Figure 1 This is a flow chart of a method for reconstructing abnormal data of an oil pipeline using a time-frequency fusion self-encoding method provided by an embodiment of the present invention;

[0053] Figure 2 It is an overall block diagram of a method for reconstructing abnormal data of an oil pipeline by time-frequency fusion self-encoding provided by an embodiment of the present invention;

[0054] Figure 3is a flowchart of time series data preprocessing provided by an embodiment of the present invention;

[0055] Figure 4 It is a structural block diagram of a time-frequency domain data distribution analysis module provided by an embodiment of the present invention;

[0056] Figure 5 is a structural block diagram of a global frequency domain feature learning module provided by an embodiment of the present invention;

[0057] Figure 6 is a flow chart of local-global data segment fusion provided by an embodiment of the present invention;

[0058] Figure 7 It is a structural block diagram of a temporal autoencoder reconstruction model provided by an embodiment of the present invention;

[0059] Figure 8 It is a block diagram of a time-frequency fusion self-encoding oil pipeline abnormal data reconstruction system provided by an embodiment of the present invention;

[0060] Fig. 9 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0061] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0062] In order to solve the problem that the aging and damage of sensors in the prediction of the wall thickness of oil pipelines lead to abnormal parameter index data, which seriously affects the accuracy of wall thickness prediction, an embodiment of the present invention designs a time-frequency fusion autoencoder oil pipeline abnormal data reconstruction method: first, a time series autoencoder reconstruction model is designed to allow the multi-dimensional time series data of the oil pipeline to adapt to the input of the autoencoder and directly output the reconstructed value of the entire target performance data, eliminating the operation of manually locating abnormal values, and combining the information of all relevant data to reconstruct the target more accurately and reliably; then, a time-frequency domain data distribution analysis module is designed to extract frequency domain features of long-period data, and fuse these frequency domain information with time domain information to capture long-period The data distribution changes in the data allow the model to adapt to the data drift phenomenon; finally, by designing a global frequency domain feature learning module, the global frequency domain information of the oil pipeline related data is extracted in the training of a large amount of long-period data, and these global information are fused with the local frequency domain information extracted from a small amount of data in special areas to enhance the model's generalization anomaly reconstruction ability on the oil pipeline data set in special areas, and improve the accuracy of the final wall thickness prediction method. In addition, the global frequency domain feature learning module also includes a probability freezing trainer to prevent the global frequency domain feature learning module from overfitting on a certain time series data when learning a lot of complete long-period time series data, resulting in poor generalization ability. A time-frequency fusion self-encoding oil pipeline abnormal data reconstruction method of an embodiment of the present invention can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of this method is shown in FIG. Figure 2 The overall block diagram of the method is shown below. The processing flow may include the following steps:

[0063] S1, collecting and preprocessing the oil pipeline parameter time series data to obtain 2D time series data segments and a complete 2D time series data set;

[0064] Alternatively, if Figure 3 As shown, the S1 specifically includes:

[0065] S11, collecting pipeline parameter time series data (such as temperature, internal pressure, pipeline vibration amplitude, flow rate and other pipeline parameter time series data) obtained by detecting the oil pipeline using various sensors, wherein the pipeline parameter time series data includes long-period time series data of pipelines laid for a long time (which will contain more abnormal values ​​for negative sample training of the model) and a small amount of time series data from special geological areas (which will be used as a small amount of data set samples to improve the generalization ability of the model on various data sets);

[0066] S12, performing mean normalization on the pipeline parameter time series data to eliminate the dimensional differences between different indicators so that they can be compared on the same scale;

[0067] S13, arranging the time series data of multiple parameter indicators in parallel, for long-period time series data, using a sliding window method, setting the window size and step size and traversing these long-period time series data, dividing them into multiple time period data, so that they are presented as 2D time series data segments, which are used in the time-frequency domain data distribution analysis module, and the complete 2D time series data set before segmentation is used in the global frequency domain feature learning module. For a small amount of time series data in special geological areas, they are directly arranged in parallel to form 2D time series data segments;

[0068] S14. Divide the 2D time series data set into a training set and a test set, where the training set requires normal data corresponding to the outlier part as labels for model training outlier reconstruction (the test set does not require labels, but the long-period time series data in the test set needs to be distinguished from the small amount of time series data in special geological areas).

[0069] S2, inputting the 2D time series data fragment into a time-frequency domain data distribution analysis module, and training the time-frequency domain data distribution analysis module to extract local frequency domain information F H×1 ;

[0070] Optional, such as Figure 4 As shown, the time-frequency domain data distribution analysis module converts a 2D time series data segment X H×W Each row of indicator data in is Fourier transformed to convert it from time domain data to frequency domain data segment Y H×L ;

[0071] Then, a multi-layer perceptron is used to perform preliminary feature extraction on the frequency domain data;

[0072] Since the generative model needs to specify a certain pipeline parameter indicator data as the generation target during use, two densely connected layers are used for further feature extraction to distinguish the specified indicator from other indicators. The densely connected layer is a stack of multiple linear layers and activation functions, and compared with the multi-layer perceptron, there is no hidden layer for deep learning feature extraction. Its function is to perform nonlinear transformation on the input data, map and extract it as a feature vector, and change its own matrix shape to facilitate the subsequent frequency domain feature fusion calculation. The first densely connected layer performs complete feature mapping extraction on the preliminary frequency domain features of the input to obtain The second dense connection layer is a selection dense connection layer. According to the user's specification, only the frequency domain features of the row corresponding to the selected parameter index are mapped and extracted to obtain Select the features extracted from the densely connected layer mapping Perform the transpose operation to obtain In order to integrate the extracted frequency domain features into the time series data fragment, the transposed matrix and Multiply to obtain the local frequency domain information FH×1 .

[0073] S3, inputting the complete 2D time series data set into a global frequency domain feature learning module, and training the global frequency domain feature learning module to extract global frequency domain information of frequency domain data in multiple long-period sequences

[0074] Alternatively, if Figure 5 As shown, the global frequency domain feature learning module transforms a complete 2D time series dataset Perform Fourier transform to convert it into long-period frequency domain data

[0075] Then the long period frequency domain data With the local frequency domain information F H×1 After addition and fusion, it is sent to the feature extraction layer composed of a multi-layer perceptron and a dense connection layer for deep frequency domain information extraction to obtain the global frequency domain information.

[0076] Optionally, the global frequency domain feature learning module also includes a probability freezing trainer to prevent the global frequency domain feature learning module from overfitting on a certain time series data when learning a lot of complete long-period time series data, resulting in poor generalization ability. The probability freezing trainer resets the weights of certain linear layers in the feature extraction layer to 0 according to the loss value returned by the loss function, so that the transmitted features of the model will not change after passing through these layers during forward propagation, and these layers will not update parameters during back propagation. Moreover, the smaller the returned loss value, the greater the probability that the feature extraction layer is frozen. Since the loss function calculation results are all between 0 and 1, the following formula is used to link the loss function and the freezing probability:

[0077]

[0078] Where P m is a freezing probability vector with only 0 and non-zero values. After multiplying with the feature extraction layer matrix, the corresponding layer is frozen. L is the loss function value. The larger the L value, the higher the P m The more 0 values ​​there are in , m is the length of the feature extraction layer.

[0079] S4, the local frequency domain information F H×1 , global frequency domain information Add and fuse with the 2D time series data fragment to obtain the local-global fused data fragment like Figure 6 As shown;

[0080] S5, the local-global fusion data segment Input a time series autoencoder reconstruction model, and train the time series autoencoder reconstruction model to reconstruct and generate a whole segment of time series data;

[0081] Alternatively, if Figure 7 As shown, the temporal self-encoding reconstruction model includes a temporal encoder and a temporal decoder;

[0082] The timing encoder splices the serial number of the pipeline parameter timing data specified by the user (such as 1 for temperature, 2 for pressure, and 3 for amplitude) into the local-global fusion data segment and the local-global fusion data fragment Each column of data is divided into time step data X, and then one-dimensional convolution is used to extract features of these time step data X in turn to obtain hidden layer features Y1, and then the hidden layer features Y1 are compressed by one-dimensional convolution to compress these features into a smaller subspace to form hidden layer features Y2. These hidden layer features Y1 and Y2 are feature vectors carrying rich information, which not only carry the normal and abnormal value position information of the input data, but also carry the information of the data distribution characteristics. The hidden layer features Y2 are compressed into latent variables Z by one-dimensional convolution. The latent variable Z is the most implicit feature, representing the abnormal value distribution and characteristic distribution of all data fragments, and is used for subsequent data decoding generation;

[0083] The time series decoder is completely mirror-symmetric with the time series encoder. After the latent variable Z passes through the 1-dimensional deconvolution in the time series decoder with different 1-dimensional convolution weights from the time series encoder, it is gradually upsampled to obtain new feature vectors Y′1 and Y′2. The feature vector Y′2 is then transformed into reconstructed values ​​X′ corresponding to the input time step data through 1-dimensional deconvolution, and they are combined to obtain the reconstructed entire time series data (including the reconstructed data of abnormal data).

[0084] Optionally, the loss function of the abnormal data reconstruction overall model includes a reconstruction loss function and a regularization term loss function;

[0085] The reconstruction loss function is used to measure the matching degree between the reconstructed data and the normal label data, and its formula is as follows:

[0086]

[0087] where X′ i is the reconstruction value, X i is the corresponding input time step value, n is the length of the input data;

[0088] The reconstruction loss function only compares the numerical difference between the reconstructed data and the real data. Even if the error between the reconstructed data and the real data is small, the data distribution trend may not be consistent. It is relatively rough to use the reconstruction loss function alone for back propagation training. Therefore, the loss function also includes a regularization term loss function, which is used to measure the matching degree between the encoding distribution of the latent variables and the input data distribution, and whether the measurement model is in place for the distribution change analysis of long-period time series data. The formula is as follows:

[0089] L 正则化 =KL(q(Z|X)||p(Z))

[0090]

[0091] Where KL represents KL divergence, q(Z|X) represents the conditional probability distribution of Z when the input is fixed to X, that is, the actual distribution of data encoding, p(Z) is the encoding distribution of the model output, and n is the length of the input data;

[0092] By combining the two loss functions, we can not only evaluate the reconstruction accuracy of the reconstruction model numerically, but also evaluate the reconstruction model's ability to analyze data distribution.

[0093] The embodiment of the present invention inputs the long-period time series data in the training set into the overall model for training. After the training is completed, the preliminary outlier reconstruction model weights are obtained, and then the weights are loaded. The model is first tested on the long-period time series data test set to confirm that the outlier reconstruction capability of the model meets the requirements. Then, the model is trained and tested on a small amount of time series data in special geological areas to ensure that the model can also perform effective outlier location reconstruction on small-scale data based on previous training and learning experience.

[0094] S6. Reconstruct the abnormal data of the oil pipeline to be reconstructed using the trained overall abnormal data reconstruction model composed of the time-frequency domain data distribution analysis module, the global frequency domain feature learning module, and the time series autoencoder reconstruction model.

[0095] like Figure 8 As shown, the embodiment of the present invention also provides a time-frequency fusion self-encoding oil pipeline abnormal data reconstruction system, the system comprising:

[0096] The acquisition and preprocessing module 810 is used to acquire and preprocess the oil pipeline parameter time series data to obtain 2D time series data segments and a complete 2D time series data set;

[0097] The local frequency domain information extraction module 820 is used to input the 2D time series data segment into the time-frequency domain data distribution analysis module, and train the time-frequency domain data distribution analysis module to extract the local frequency domain information F H×1 ;

[0098] The global frequency domain information extraction module 830 is used to input the complete 2D time series data set into the global frequency domain feature learning module, and train the global frequency domain feature learning module to extract the global frequency domain information of the frequency domain data in multiple long-period sequences.

[0099] The fusion module 840 is used to combine the local frequency domain information F H×1 , global frequency domain information Add and fuse with the 2D time series data fragment to obtain the local-global fused data fragment

[0100] Reconstruction module 850, used to reconstruct the local-global fusion data segment Input a time series autoencoder reconstruction model, and train the time series autoencoder reconstruction model to reconstruct and generate a whole segment of time series data;

[0101] The data reconstruction module 860 to be reconstructed is used to reconstruct the abnormal data of the oil pipeline to be reconstructed using the trained abnormal data reconstruction overall model composed of the time-frequency domain data distribution analysis module, the global frequency domain feature learning module, and the time series autoencoder reconstruction model.

[0102] The time-frequency fusion self-encoding oil pipeline abnormal data reconstruction system provided in the embodiment of the present invention has a functional structure corresponding to the time-frequency fusion self-encoding oil pipeline abnormal data reconstruction method provided in the embodiment of the present invention, which will not be repeated here.

[0103] Fig. 9 It is a structural diagram of an electronic device 900 provided in an embodiment of the present invention. The electronic device 900 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 901 and one or more memories 902, wherein at least one instruction is stored in the memory 902, and the at least one instruction is loaded and executed by the processor 901 to implement the steps of the above-mentioned time-frequency fusion self-encoding oil pipeline abnormal data reconstruction method.

[0104] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, which can be executed by a processor in a terminal to complete the above-mentioned time-frequency fusion self-encoding oil pipeline abnormal data reconstruction method. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device.

[0105] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0106] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for reconstructing abnormal data of oil pipelines by time-frequency fusion self-encoding, characterized in that: The method comprises: S1, collecting and preprocessing the oil pipeline parameter time series data to obtain 2D time series data segments and a complete 2D time series data set; S2, inputting the 2D time series data fragment into a time-frequency domain data distribution analysis module, and training the time-frequency domain data distribution analysis module to extract local frequency domain information F H×1 ; S3, inputting the complete 2D time series data set into a global frequency domain feature learning module, and training the global frequency domain feature learning module to extract global frequency domain information of frequency domain data in multiple long-period sequences S4, the local frequency domain information F H×1 , global frequency domain information Add and fuse with the 2D time series data fragment to obtain the local-global fused data fragment S5, the local-global fusion data segment Input a time series autoencoder reconstruction model, and train the time series autoencoder reconstruction model to reconstruct and generate a whole segment of time series data; S6. Reconstruct the abnormal data of the oil pipeline to be reconstructed using the trained overall abnormal data reconstruction model composed of the time-frequency domain data distribution analysis module, the global frequency domain feature learning module, and the time series autoencoder reconstruction model.

2. The method according to claim 1, characterized in that The S1 specifically includes: S11, collecting pipeline parameter time series data obtained by detecting the oil pipeline with various sensors, wherein the pipeline parameter time series data includes long-period time series data of the long-time pipeline laying and a small amount of time series data of special geological areas; S12, performing mean normalization on the pipeline parameter time series data to eliminate the dimensional differences between different indicators so that they can be compared on the same scale; S13, arranging the time series data of multiple parameter indicators in parallel, for long-period time series data, using a sliding window method, setting the window size and step size and traversing these long-period time series data, dividing them into multiple time period data, so that they are presented as 2D time series data segments, which are used in the time-frequency domain data distribution analysis module, and the complete 2D time series data set before segmentation is used in the global frequency domain feature learning module. For a small amount of time series data in special geological areas, they are directly arranged in parallel to form 2D time series data segments; S14. Divide the 2D time series data set into a training set and a test set, wherein the training set requires normal data corresponding to the outlier part as labels for model training outlier reconstruction.

3. The method according to claim 1, characterized in that: The time-frequency domain data distribution analysis module converts a 2D time series data segment X H×W Each row of indicator data in is Fourier transformed to convert it from time domain data to frequency domain data segment Y H×L ; Then, a multi-layer perceptron is used to perform preliminary feature extraction on the frequency domain data; Since the generative model needs to specify a certain pipeline parameter indicator data as the generation target during use, two densely connected layers are used for further feature extraction to distinguish the specified indicator from other indicators. The densely connected layer is a stack of multiple linear layers and activation functions, and compared with the multi-layer perceptron, there is no hidden layer for deep learning feature extraction. Its function is to perform nonlinear transformation on the input data, map and extract it as a feature vector, and change its own matrix shape to facilitate the subsequent frequency domain feature fusion calculation. The first densely connected layer performs complete feature mapping extraction on the preliminary frequency domain features of the input to obtain The second dense connection layer is a selection dense connection layer. According to the user's specification, only the frequency domain features of the row corresponding to the selected parameter index are mapped and extracted to obtain Select the features extracted from the densely connected layer mapping Perform the transpose operation to obtain In order to integrate the extracted frequency domain features into the time series data fragment, the transposed matrix and Multiply to obtain the local frequency domain information F H×1 .

4. The method according to claim 1, characterized in that: The global frequency domain feature learning module transforms a complete 2D time series dataset Perform Fourier transform to convert it into long-period frequency domain data Then the long period frequency domain data With the local frequency domain information F H×1 After addition and fusion, it is sent to the feature extraction layer composed of a multi-layer perceptron and a dense connection layer for deep frequency domain information extraction to obtain the global frequency domain information.

5. The method according to claim 4, characterized in that The global frequency domain feature learning module also includes a probability freezing trainer to prevent the global frequency domain feature learning module from overfitting on a certain time series data when learning a lot of complete long-period time series data, resulting in poor generalization ability. The probability freezing trainer resets the weights of some linear layers in the feature extraction layer to 0 according to the loss value returned by the loss function, so that the features transmitted by the model will not change after passing through these layers during forward propagation, and these layers will not update parameters during back propagation. Moreover, the smaller the loss value returned, the greater the probability that the feature extraction layer is frozen. Since the loss function calculation results are all between 0 and 1, the following formula is used to link the loss function and the freezing probability: Where P m is a freezing probability vector with only 0 and non-zero values. After multiplying with the feature extraction layer matrix, the corresponding layer is frozen. L is the loss function value. The larger the L value, the higher the P m The more 0 values ​​there are in , m is the length of the feature extraction layer.

6. The method according to claim 1, characterized in that The temporal self-encoding reconstruction model comprises a temporal encoder and a temporal decoder; The timing encoder splices the sequence number of the pipeline parameter timing data specified by the user into the local-global fusion data segment and the local-global fusion data fragment Each column of data is divided into time step data X, and then one-dimensional convolution is used to extract features of these time step data X in turn to obtain hidden layer features Y1, and then the hidden layer features Y1 are compressed by one-dimensional convolution to compress these features into a smaller subspace to form hidden layer features Y2. These hidden layer features Y1 and Y2 are feature vectors carrying rich information, which not only carry the normal and abnormal value position information of the input data, but also carry the information of the data distribution characteristics. The hidden layer features Y2 are compressed into latent variables Z by one-dimensional convolution. The latent variable Z is the most implicit feature, representing the abnormal value distribution and characteristic distribution of all data fragments, and is used for subsequent data decoding generation; The time series decoder is completely mirror-symmetric with the time series encoder. After the latent variable Z passes through the 1D deconvolution in the time series decoder with different 1D convolution weights from the time series encoder, it is gradually upsampled to obtain new feature vectors Y'1 and Y'2. The feature vector Y'2 is then transformed into time step data corresponding to the input of the reconstructed values ​​X' through 1D deconvolution, and they are combined to obtain the entire segment of time series data reconstructed and generated.

7. The method according to claim 1, characterized in that The loss function of the abnormal data reconstruction overall model includes a reconstruction loss function and a regularization term loss function; The reconstruction loss function is used to measure the matching degree between the reconstructed data and the normal label data, and its formula is as follows: where X' i is the reconstruction value, X i is the corresponding input time step value, n is the length of the input data; The reconstruction loss function only compares the numerical difference between the reconstructed data and the real data. Even if the error between the reconstructed data and the real data is small, the data distribution trend may not be consistent. It is relatively rough to use the reconstruction loss function alone for back propagation training. Therefore, the loss function also includes a regularization term loss function, which is used to measure the matching degree between the encoding distribution of the latent variables and the input data distribution, and whether the measurement model is in place for the distribution change analysis of long-period time series data. The formula is as follows: Y 正则化 =KL(q(Z|X)||p(Z)) Where KL represents KL divergence, q(Z|X) represents the conditional probability distribution of Z when the input is fixed to X, that is, the actual distribution of data encoding, p(Z) is the encoding distribution of the model output, and n is the length of the input data; By combining the two loss functions, we can not only evaluate the reconstruction accuracy of the reconstruction model numerically, but also evaluate the reconstruction model's ability to analyze data distribution.

8. A time-frequency fusion self-encoding oil pipeline abnormal data reconstruction system, characterized in that: The system comprises: The acquisition and preprocessing module is used to acquire and preprocess the oil pipeline parameter time series data to obtain 2D time series data segments and a complete 2D time series data set; The local frequency domain information extraction module is used to input the 2D time series data fragment into the time-frequency domain data distribution analysis module, and train the time-frequency domain data distribution analysis module to extract the local frequency domain information F H×1 ; A global frequency domain information extraction module is used to input the complete 2D time series data set into a global frequency domain feature learning module, and train the global frequency domain feature learning module to extract the global frequency domain information of the frequency domain data in multiple long-period sequences. A fusion module is used to combine the local frequency domain information F H×1 , global frequency domain information Add and fuse with the 2D time series data fragment to obtain the local-global fused data fragment A reconstruction module is used to reconstruct the local-global fusion data fragment Input a time series autoencoder reconstruction model, and train the time series autoencoder reconstruction model to reconstruct and generate a whole segment of time series data; The data reconstruction module to be reconstructed is used to reconstruct the abnormal data of the oil pipeline to be reconstructed using the trained abnormal data reconstruction overall model composed of the time-frequency domain data distribution analysis module, the global frequency domain feature learning module, and the time series autoencoder reconstruction model.

9. An electronic device, comprising a processor and a memory, wherein at least one instruction is stored in the memory, wherein: The at least one instruction is loaded and executed by the processor to implement the time-frequency fusion self-encoding oil pipeline abnormal data reconstruction method as described in any one of claims 1-7.

10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that: The at least one instruction is loaded and executed by the processor to implement the time-frequency fusion self-encoding oil pipeline abnormal data reconstruction method as described in any one of claims 1-7.

Citation Information

Cited By

  • Industrial multi-dimensional time sequence anomaly detection method based on multi-granularity overall period reconstruction

    CN121144699A