DAS signal generation method and device based on autoencoder and conditional diffusion model
By combining autoencoders with conditional diffusion models, high-quality and diverse oil and gas pipeline threat event signals are generated, solving the problems of low identification accuracy and poor generalization ability caused by insufficient sample size, and improving the accuracy and reliability of threat event identification.
Patent Information
- Application Number
- CN202510548536.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-04-28
AI Technical Summary
In existing oil and gas pipeline threat event identification, the lack of sufficient sample size leads to low classification accuracy and poor model generalization ability. In particular, it is difficult to accurately identify multiple types of events in complex scenarios, which affects the practicality and reliability of the monitoring system.
A DAS signal generation method based on autoencoder and conditional diffusion model is adopted. After training with autoencoder, one-dimensional time series signal data is mapped to two-dimensional feature space. Then, the conditional diffusion model is used to generate high-quality and diverse signal data under the guidance of conditional information such as event type, thereby enhancing the richness and representativeness of the dataset.
It significantly improves the accuracy and robustness of threat event identification models, especially in cases where data samples are scarce or unevenly distributed, thereby enhancing the model's performance and generalization ability.
Smart Images

Figure CN120493054B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of oil and gas pipeline monitoring and signal generation, and particularly relates to a DAS signal generation method and device based on a self-encoder and a conditional diffusion model. BACKGROUND
[0002] In the field of time series data processing and signal generation, with the continuous development of deep learning technology, researchers are increasingly focusing on how to improve the performance of models by enhancing the data set, especially in the case of data sample scarcity or unevenness. Traditional data enhancement methods, such as shifting signals, adding noise, scaling, etc., although to some extent can increase the diversity of data, but often cannot effectively capture the actual distribution of data, and is easy to introduce unreasonable samples, affecting the performance of the model, especially when dealing with complex time series signals, the limitations of these methods become particularly obvious.
[0003] In recent years, deep learning methods such as GAN (Generative Adversarial Networks), VAE (Variational Auto Encoder), LSTM (Long Short-Term Memory) have been widely used in signal generation tasks, especially in dealing with the scarcity problem of time series data, which has shown good potential. However, these methods still have problems such as unstable training and mode collapse when facing long sequence generation tasks, making it difficult to guarantee the diversity and authenticity of the generated signal. Diffusion models, as a new generation of generation model, have made remarkable achievements in image generation field with their step-by-step denoising characteristics, and have shown high stability and generation quality in signal generation.
[0004] Existing deep learning-based time series signal generation methods, although to some extent can improve the diversity of data, but mostly have the disadvantages of not enough diversified generated samples and low quality, especially in dealing with complex threat event recognition tasks such as oil and gas pipelines, due to the lack of data samples, the accuracy of threat event recognition and the generalization ability of the model are poor. Therefore, the existing technology relies on simple noise addition or signal disturbance method, and does not effectively consider the actual characteristics of time series signal, resulting in the quality and diversity of enhanced data cannot meet the needs of complex applications.
[0005] Traditional data augmentation methods usually expand the dataset by simple ways such as shifting signals, adding noise and synthesizing signals. Wen et al. summarized the effects of noise addition (e.g., Gaussian noise) and signal perturbation (e.g., shifting, scaling, etc.) on improving model performance, and provided various application examples in time series data augmentation. Kamycki et al. obtained a new time series in the twist space between suboptimal aligned input samples of different lengths. Although these methods can increase the diversity of data to some extent, the generated signals do not fully consider the actual distribution of the data, and are prone to introduce unreasonable samples, which in turn affects the performance of the model. In particular, for complex time series signals, simple augmentation methods often cannot effectively improve the recognition ability of the model.
[0006] In the field of deep learning, methods such as generative adversarial networks, variational autoencoders, RNNs (Recurrent Neural Networks) and long short-term memory networks are widely used in signal generation tasks. Yoon et al. proposed an attention-based generative adversarial network to solve the problem of data scarcity in time series sensor data collected by unmanned aerial vehicles. Li et al. proposed a robust GAN model with adaptive recovery strategy, TSA-GAN, which helped the time series classifier achieve 8.3% to 12.5% performance improvement, far better than the baseline. Liu et al. proposed a data augmentation method based on GAN, Bi-LSTM and time series attention to increase the number of abnormal samples. In GANBATS, Bi-LSTM is introduced to extract time series features, which are then passed to the generator network of GANBATS. GANBATS also modifies the discriminator network, adding an attention mechanism to achieve global attention on time series. Wu et al. proposed a VSG method based on Time CVAE. By incorporating labeled data into the VSG process, the generator ensures that the generated virtual samples are closer to the distribution of real samples. In addition, a time-related module is designed as an important part of the Time CVAE model. Although these methods have shown good ability in generating images and time series signals, they often face the problem of unstable training in long sequence generation tasks, and are prone to mode collapse, i.e., the generated signals lack diversity and authenticity, which affects the quality of the generated signals and the generalization ability of the model.
[0007] As a generative model that has emerged in recent years, the diffusion model has obvious advantages in stability and quality of generated signals due to its step-by-step denoising feature. Compared with traditional generative models, the diffusion model can precisely control the signal generation process by gradually adding noise and denoising process, thus having higher stability and reliability in signal generation tasks. Especially in generating images and high-quality signals, the diffusion model has achieved remarkable results. Dai et al. proposed a time series denoising diffusion probability model (Time DDPM) to construct a soft sensor for limited time series samples. First, LSTM units and one-dimensional convolutional neural networks are introduced into the Time DDPM noise prediction network to mine the spatiotemporal characteristics of the samples. Then, virtual samples are reconstructed step by step in the reverse process to expand the sample space that lacks data. However, although the diffusion model has been widely used in image generation and other fields and has shown strong generation ability, the existing research on one-dimensional time series signal generation tasks is still limited, and the related technology is still in the exploratory stage. Therefore, in the signal data enhancement task, how to effectively use the diffusion model to generate one-dimensional signal data to improve the performance of the threat event identification model is still a technical problem to be solved. SUMMARY
[0008] In order to solve the technical problems of low classification accuracy and poor model generalization ability caused by insufficient sample quantity in existing oil and gas pipeline threat event identification, the existing threat event data is often limited in quantity and uneven in distribution, which is difficult to support the accurate identification of multiple events in complex scenarios by deep learning models, and seriously restricts the practicality and reliability of the monitoring system. The embodiments of the present application provide a DAS (Distributed Acoustic Sensing) signal generation method and device based on an autoencoder and a conditional diffusion model, which automatically synthesizes high-quality and diverse event signal data using a deep generative model, enhances the richness and representativeness of the data set, and thus improves the performance of the classification model. By introducing a conditional diffusion mechanism, the generation process is guided by combining condition information such as event type, so that the generated samples have stronger pertinence and controllability, effectively making up for the problem of sample scarcity or uneven distribution in the original data set. The technical solution is as follows:
[0009] On the one hand, a DAS signal generation method based on an autoencoder and a conditional diffusion model is provided, which is implemented by a DAS signal generation device, and the method comprises:
[0010] S1, constructing an autoencoder, training the autoencoder, and obtaining a trained autoencoder.
[0011] S2, obtaining a threat event signal data set to be enhanced.
[0012] S3, input one-dimensional time sequence signal data in the threat event signal data set to be enhanced into the encoding module of the autoencoder to obtain a two-dimensional signal feature map.
[0013] S4, input the two-dimensional signal feature map into the conditional diffusion model to obtain a generated two-dimensional signal feature map.
[0014] S5, input the generated two-dimensional signal feature map into the decoding module of the autoencoder to obtain new one-dimensional time sequence signal data, and construct an enhanced threat event signal data set according to the new one-dimensional time sequence signal data.
[0015] S6, train the oil and gas pipeline threat event identification model according to the enhanced threat event signal data set to complete the oil and gas pipeline threat event identification task.
[0016] Optionally, the loss function of the autoencoder in S1 includes a time domain loss, a frequency domain loss, and a similarity loss.
[0017] The loss function of the autoencoder is shown in the following formula (1):
[0018] L = L Time + L Freq + 0.3·L Sim (1)
[0019] In the formula, L represents the loss function of the autoencoder, L Time represents the time domain loss, L Freq represents the frequency domain loss, and L Sim represents the signal similarity loss.
[0020] The frequency domain loss is shown in the following formula (2):
[0021] L Freq = ||F(x(t))-F(x(t))||2 (2)
[0022] In the formula, F(·) represents Fourier transform, x(t) represents an input signal, x(t) represents a reconstructed signal, and ||·||2 represents L2 norm.
[0023] The similarity loss is shown in the following formula (3):
[0024]
[0025] Optionally, the encoding module includes a time-frequency graph convolution model, and the time-frequency graph convolution model includes a real part multi-scale convolution layer and an imaginary part multi-scale convolution layer.
[0026] S3 includes:
[0027] S31, input the one-dimensional time sequence signal data to the real part multi-scale convolution layer, perform convolution operation through the multiple different dimension convolution kernels of the real part multi-scale convolution layer, and extract the real part feature information of the signal.
[0028] S32, input the one-dimensional time sequence signal data to the imaginary part multi-scale convolution layer, perform convolution operation through the multiple different dimension convolution kernels of the imaginary part multi-scale convolution layer, and extract the imaginary part feature information of the signal.
[0029] S33, respectively take absolute values of the real part feature information and the imaginary part feature information, and obtain the real part convolution result and the imaginary part convolution result.
[0030] S34, fuse the real part convolution result and the imaginary part convolution result, and obtain the two-dimensional signal feature map.
[0031] Optionally, the convolution operation in S31 comprises:
[0032] The weight initialized by the Fourier transform function is used as the kernel function to perform the convolution operation, as shown in the following formula (4):
[0033] h(j) = H r (j) + iH i (j) (4)
[0034] wherein,
[0035] H r (j) = cos(2pifj + f) (5)
[0036] H i (j) = sin(2pifj + f) (6)
[0037] In the formula, h(j) represents the Fourier transform coefficient corresponding to the frequency point j, j represents the index of the frequency point, H r (j) represents the real part function of the Fourier kernel, H i (j) represents the imaginary part function of the Fourier kernel, i represents the imaginary unit, f represents the frequency, and f represents the phase.
[0038] Optionally, the taking absolute values of the real part feature information and the imaginary part feature information in S33 to obtain the real part convolution result and the imaginary part convolution result comprises:
[0039]
[0040] In the formula, y r [i] represents the real part of the Fourier transform result at position i, i represents the current starting position index, k represents the time step index within the window, x[i+j] represents the value of the signal at the i+k time, and y i[i] represents the imaginary part of the Fourier transform result at position i, H r (j) represents the real part function of the Fourier kernel, H i (j) represents the imaginary part function of the Fourier kernel.
[0041] Optionally, the condition of the conditional diffusion model is the label of the event type and the time step of diffusion.
[0042] The denoising network of the conditional diffusion model is a deep UNet model.
[0043] Optionally, the decoding module includes a plurality of bottleneck blocks and residual blocks.
[0044] In another aspect, a DAS signal generation device based on an autoencoder and a conditional diffusion model is provided, which is applied to a DAS signal generation method based on an autoencoder and a conditional diffusion model, and the device includes:
[0045] The construction module is configured to construct an autoencoder, train the autoencoder, and obtain a trained autoencoder.
[0046] The acquisition module is configured to acquire a threat event signal dataset to be enhanced.
[0047] The encoding module is configured to input one-dimensional time series signal data in the threat event signal dataset to be enhanced into an encoding module of the autoencoder to obtain a two-dimensional signal feature map.
[0048] The generation module is configured to input the two-dimensional signal feature map into a conditional diffusion model to obtain a generated two-dimensional signal feature map.
[0049] The decoding module is configured to input the generated two-dimensional signal feature map into a decoding module of the autoencoder to obtain new one-dimensional time series signal data, and construct an enhanced threat event signal dataset according to the new one-dimensional time series signal data.
[0050] The output module is configured to train an oil and gas pipeline threat event identification model according to the enhanced threat event signal dataset to complete an oil and gas pipeline threat event identification task.
[0051] Optionally, the loss function of the autoencoder includes a time domain loss, a frequency domain loss, and a similarity loss.
[0052] The loss function of the autoencoder is shown in the following formula (1):
[0053] L = L Time + L Freq + 0.3·L Sim (1)
[0054] In the formula, L represents the loss function of the autoencoder, L Time represents the time domain loss, and LFreq represents a frequency domain loss, L Sim represents a signal similarity loss.
[0055] The frequency domain loss is shown in the following formula (2):
[0056] L Freq =||F(x(t))-F(x(t))||2 (2)
[0057] In the formula, F(·) represents a Fourier transform, x(t) represents an input signal, x(t) represents a reconstructed signal, and ||·||2 represents an L2 norm.
[0058] The similarity loss is shown in the following formula (3):
[0059]
[0060] Optionally, the encoding module comprises a time-frequency graph convolution model, and the time-frequency graph convolution model comprises a real part multi-scale convolution layer and an imaginary part multi-scale convolution layer.
[0061] The encoding module is further configured to:
[0062] S31, input one-dimensional time series signal data to the real part multi-scale convolution layer, and perform convolution operation through a plurality of different dimension convolution kernels of the real part multi-scale convolution layer to extract real part feature information of the signal.
[0063] S32, input one-dimensional time series signal data to the imaginary part multi-scale convolution layer, and perform convolution operation through a plurality of different dimension convolution kernels of the imaginary part multi-scale convolution layer to extract imaginary part feature information of the signal.
[0064] S33, respectively take absolute values of the real part feature information and the imaginary part feature information to obtain real part convolution results and imaginary part convolution results.
[0065] S34, fuse the real part convolution results and the imaginary part convolution results to obtain a two-dimensional signal feature map.
[0066] Optionally, the convolution operation comprises:
[0067] The convolution operation is performed by initializing the weight as a kernel function using a Fourier transform function, as shown in the following formula (4):
[0068] h(j)=H r (j)+iH i (j) (4)
[0069] wherein,
[0070] H r (j)=cos(2πfj+Φ) (5)
[0071] H i (j)=sin(2πfj+Φ) (6)
[0072] wherein h(j) represents a Fourier transform coefficient corresponding to a frequency point j, j represents an index of the frequency point, H r (j) represents a real part function of the Fourier kernel, H i (j) represents an imaginary part function of the Fourier kernel, i represents an imaginary unit, f represents a frequency, and Φ represents a phase.
[0073] Optionally, absolute values of the real part feature information and the imaginary part feature information are calculated respectively to obtain a real part convolution result and an imaginary part convolution result, as shown in the following formulas (7) and (8):
[0074]
[0075] wherein y r [i] represents a real part of a Fourier transform result at a position i, i represents a current starting position index, k represents a time step index within a window, x[i+k] represents a value of a signal at the i+k time, y i [i] represents an imaginary part of the Fourier transform result at the position i, H r (j) represents the real part function of the Fourier kernel, H i (j) represents the imaginary part function of the Fourier kernel.
[0076] Optionally, the condition of the conditional diffusion model is a label of an event type and a time step of diffusion.
[0077] The denoising network of the conditional diffusion model is a deep UNet model.
[0078] Optionally, the decoding module includes a plurality of bottleneck blocks and residual blocks.
[0079] In another aspect, a DAS signal generation device is provided, and the DAS signal generation device includes a processor and a memory having computer readable instructions stored thereon, wherein the computer readable instructions, when executed by the processor, implement any one of the above-described DAS signal generation methods based on an autoencoder and a conditional diffusion model.
[0080] In another aspect, a computer readable storage medium is provided, and the storage medium has at least one instruction stored therein, wherein the at least one instruction is loaded and executed by a processor to implement any one of the above-described DAS signal generation methods based on an autoencoder and a conditional diffusion model.
[0081] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:
[0082] In the present application, in order to solve the problem that the prior art relies on simple noise addition or signal disturbance method, the actual characteristics of the time sequence signal are not effectively considered, and the quality and diversity of the enhanced data cannot meet the needs of complex applications. The present application provides a data enhancement method based on diffusion model, which can significantly improve the diversity and representativeness of the data while ensuring the consistency of the signal characteristics, thereby improving the recognition accuracy and robustness of the classification model under the condition of few samples. By introducing the conditional diffusion mechanism, combining the information such as event type, and guiding the generation process, the problems of low quality and insufficient diversity of generated samples in traditional methods are overcome. This method not only can expand the data set, but also can enhance the performance of the threat event recognition model, especially suitable for the case of few samples or uneven distribution, and has wide application prospect. BRIEF DESCRIPTION OF DRAWINGS
[0083] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0084] Figure 1 It is a DAS signal generation method flow chart based on the self-encoder and the conditional diffusion model provided by the embodiment of the present application;
[0085] Figure 2 It is a signal effect diagram generated by the self-encoder provided by the embodiment of the present application;
[0086] Figure 3 It is a general architecture diagram of the time-frequency diagram generation module provided by the embodiment of the present application;
[0087] Figure 4 It is a structure diagram of real part multi-scale convolution provided by the embodiment of the present application;
[0088] Figure 5 It is a denoising network structure diagram of the diffusion model provided by the embodiment of the present application;
[0089] Figure 6 It is a structure diagram of the decoder provided by the embodiment of the present application;
[0090] Figure 7 It is a general architecture diagram of the DAS signal generation model provided by the embodiment of the present application;
[0091] Figure 8 It is an effect diagram of the diffusion model generated image provided by the embodiment of the present application;
[0092] Figure 9It is the original signal graph of the oil and gas pipeline threat event provided by the embodiment of the application;
[0093] Figure 10 It is the newly generated vehicle signal graph provided by the embodiment of the application;
[0094] Figure 11 It is the newly generated artificial operation signal graph provided by the embodiment of the application;
[0095] Figure 12 It is the newly generated mechanical operation signal graph provided by the embodiment of the application;
[0096] Figure 13 It is the training curve graph of the time-frequency graph convolution network after data enhancement provided by the embodiment of the application;
[0097] Figure 14 It is the training curve graph of the original time-frequency graph convolution network provided by the embodiment of the application;
[0098] Figure 15 It is the device block diagram of the DAS signal generation device based on the autoencoder and the conditional diffusion model provided by the embodiment of the application;
[0099] Figure 16 It is the structural schematic diagram of the DAS signal generation device provided by the embodiment of the application. DETAILED DESCRIPTION
[0100] The technical solutions in the application will be described below with reference to the drawings.
[0101] In the embodiments of the application, the words such as "example", "for example" and the like are used to represent an example, illustration or description. Any embodiment or design scheme described as "example" in the application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific manner. In addition, in the embodiments of the application, the meaning expressed by "and / or" can be both, or can be one of the two.
[0102] In the embodiments of the application, "image" and "picture" can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized. "Of", "corresponding" and "corresponding" can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized.
[0103] In the embodiments of the application, sometimes the subscript such as W1 can be written in the form of non-subscript such as W1, and the meanings expressed are consistent when the distinction is not emphasized.
[0104] In order to make the technical problems, technical solutions and advantages to be solved by the present application more clear, the following will be described in detail in combination with the drawings and specific embodiments.
[0105] The embodiment of the present application provides a DAS signal generation method based on a self-encoder and a conditional diffusion model. Figure 1 As shown in the flow chart of the DAS signal generation method based on the self-encoder and the conditional diffusion model, the processing flow of the method can include the following steps:
[0106] S1, a self-encoder is constructed, the self-encoder is trained, and a trained self-encoder is obtained.
[0107] In a feasible implementation manner, the self-encoder in the present application is an unsupervised learning model, and does not need to input a label. The self-encoder mainly includes two steps of encoding and decoding. The encoding process extracts effective representation in a high-dimensional space through analysis of original input data, and the high-dimensional space vectors can be used for other tasks. The decoding process is a reverse reconstruction process, which reconstructs the representation in the high-dimensional space into new data with the same characteristics as the original input data. In addition, the present application also innovatively designs a loss function of the self-encoder, which is composed of three parts of a time domain loss, a frequency domain loss and a similarity loss, and aims to improve the encoding and reconstruction ability of the model, so as to improve the quality and precision of signal generation.
[0108] The loss function of the self-encoder is shown in the following formula (1):
[0109] L=L Time +L Freq +0.3·L Sim (1)
[0110] In the formula, L represents the loss function of the self-encoder, L Time represents the time domain loss, which is the mean absolute error of the input signal and the reconstructed signal, L Freq represents the frequency domain loss, and L Sim represents the signal similarity loss.
[0111] In a one-dimensional signal generation model, the frequency domain loss is a commonly used error function, and is specifically shown in formula 2. The Fourier transform function is used to map the signal into the frequency domain space, and the frequency domain error between the original signal and the reconstructed signal is calculated:
[0112] L Freq =||F(x(t))-F(x(t))||2 (2)
[0113] In the formula, F(·) represents the Fourier transform, x(t) represents the input signal, x(t) represents the reconstructed signal, ||·||2 represents the L2 norm, that is, the Euclidean distance, which is used to measure the error of the signal in the frequency domain.
[0114] Further, the similarity loss is also a loss function in the signal generation model, which continuously adjusts the structure of the generated signal to have a similar structure to the original signal. Here, the cosine similarity function is used for measurement:
[0115]
[0116] After the autoencoder is built, it needs to be pre-trained to ensure that the diffusion model can generate high-quality two-dimensional signal graphs. In this process, the encoder can map the original one-dimensional signal in the two-dimensional time-frequency space, and the decoder reconstructs the one-dimensional signal from the two-dimensional space vector. The generated signal effect is as Figure 2
[0117] S2, obtaining a threat event signal data set to be enhanced.
[0118] S3, inputting one-dimensional time series signal data in the threat event signal data set to be enhanced into an encoding module of the autoencoder to obtain a two-dimensional signal feature map.
[0119] Optionally, the encoding module comprises a time-frequency graph convolution model, and the time-frequency graph convolution model comprises a real part multi-scale convolution layer and an imaginary part multi-scale convolution layer.
[0120] In a feasible implementation manner, the encoder is composed of a time-frequency graph convolution network generating time-frequency graph module, as shown in Figure 3 Through the calculation of the real part convolution layer and the imaginary part convolution layer, the real part and the imaginary part feature maps of the signal are extracted respectively. In order to fully extract the feature information of the signal, the time series signal is first mapped to a high-dimensional space to extract the time domain feature map. Then, the real part feature, the imaginary part feature and the time domain feature are fused by using the channel attention mechanism to generate the required two-dimensional signal feature map, and the feature map is taken as the input image of the diffusion model.
[0121] The above step S3 can comprise steps S31-S34 as follows:
[0122] S31, inputting the one-dimensional time series signal data into the real part multi-scale convolution layer to extract the real part feature information of the signal through the convolution operation of a plurality of different dimension convolution kernels of the real part multi-scale convolution layer.
[0123] In a feasible implementation manner, the core component of the time-frequency graph convolution network encoder in the application is a multi-scale Fourier transform convolution layer, as shown in Figure 4 The layer can respectively extract real part and imaginary part feature maps of the two-dimensional signal by convolution operation on the one-dimensional signal, and then obtain the required time-frequency graph by fusion. Unlike the traditional time-frequency graph acquisition method, the application modifies the kernel function in the one-dimensional convolution operation, so that the convolution network model can autonomously learn the time-frequency characteristics in the signal, thereby generating a time-frequency graph distribution with good performance, as shown in the formula 3. Figure 4 The structure of the model adopts a multi-scale convolution design strategy, and the convolution operation is performed by using multiple convolution kernels of different dimensions to extract signal features from different receptive fields, and finally fuse these feature information to enhance the feature extraction capability of the model. In the multi-scale Fourier transform convolution layer, the weight is initialized by using the Fourier transform function, and the original kernel function is replaced, as shown in the formula 4. Real part and imaginary part feature information of the signal is extracted by the convolution operation:
[0124] h(j)=H r (j)+iH i (j) (4)
[0125] The definition of the real part and the imaginary part can be selected as the representation of the sine and cosine functions:
[0126] H r (j)=cos(2πfj+Φ) (5)
[0127] H i (j)=sin(2πfj+Φ) (6)
[0128] In the formula, h(j) represents the Fourier transform coefficient corresponding to the frequency point j, j represents the index of the frequency point, H r (j) represents the real part function of the Fourier kernel, H i (j) represents the imaginary part function of the Fourier kernel, i represents the imaginary unit, f represents the frequency, and Φ represents the phase.
[0129] The input signal and the real part and imaginary part convolution operation are combined to obtain the corresponding real part and imaginary part convolution module.
[0130] S32, input the one-dimensional time sequence signal data into the imaginary part multi-scale convolution layer, and perform convolution operation by using multiple convolution kernels of different dimensions in the imaginary part multi-scale convolution layer to extract the imaginary part feature information of the signal.
[0131] In a feasible implementation manner, the imaginary part multi-scale convolution layer is similar to the real part multi-scale convolution layer.
[0132] S33, respectively, the absolute value of the real part feature information and the imaginary part feature information is obtained, and the real part convolution result and the imaginary part convolution result are obtained.
[0133] In a feasible implementation, in order to highlight the strength of the signal, ignore its direction, reduce the influence of the symbol on the feature map, it is necessary to take the absolute value of the real part and the imaginary part respectively, and finally obtain the results of the real part convolution part and the imaginary part convolution:
[0134]
[0135] In the formula, y r [i] represents the real part of the Fourier transform result at position i, i represents the current starting position index, k represents the time step index within the window, x[i+j] represents the value of the signal at the i+k time, y i [i] represents the imaginary part of the Fourier transform result at position i, H r (j) represents the real part function of the Fourier kernel, H i (j) represents the imaginary part function of the Fourier kernel.
[0136] S34, fuse the real part convolution result and the imaginary part convolution result to obtain a two-dimensional signal feature map.
[0137] In a feasible implementation, after completing the multi-scale feature fusion, the model of the application further introduces two one-dimensional convolution layers to perform deep extraction on the fused features, and further improves the feature representation capability of the model. Finally, by performing absolute value operation on the feature map, the real part and the imaginary part feature maps of the signal are obtained respectively. This design fully considers the necessity of multi-scale feature extraction, and combines the advantages of feature fusion and one-dimensional convolution, to provide high-expression feature representation for subsequent tasks, and greatly improves the performance of the model.
[0138] S4, input the two-dimensional signal feature map into a conditional diffusion model to obtain a generated two-dimensional signal feature map.
[0139] Optionally, the condition of the conditional diffusion model is a label of an event type and a time step of diffusion.
[0140] The denoising network of the conditional diffusion model is a deep UNet model.
[0141] In a feasible implementation, the diffusion model of the latent space generates two-dimensional data by simulating data distribution and reverse diffusion, mainly including two stages, a forward diffusion process and a reverse diffusion process. In the forward diffusion process, the input x 0 is gradually damaged into a Gaussian noise vector. Specifically, at the kth step, x k is generated by damaging the previous iteration x k-1 with zero-mean Gaussian noise:
[0142]
[0143] In the formula, q(x kx k-1 represents the conditional probability distribution, N represents the total number of diffusion steps, x k represents the state of the diffusion process at the kth step, β k represents the noise intensity coefficient at the kth step, I represents the unit matrix, and it is ensured that the noise is isotropic.
[0144] This can also be rewritten as:
[0145]
[0146] where, α k = 1-β k , α s represents the probability of an independent event.
[0147] x k can be represented by the formula, x k can be easily recovered from x 0 :
[0148]
[0149] where ε follows a N(0, I) normal distribution.
[0150] Further, in DDPM (Denoising Diffusion Probabilistic Models), the backward denoising is defined as a Markov process. Specifically, in the kth denoising step, x k-1 is generated from x k by sampling from the following normal distribution:
[0151] p θ (x k-1 |x k ) = N(x k-1 ; μ θ (x k , k), ∑ θ (x k , k)) (12)
[0152] where p θ (x k-1 |x k ) represents the conditional probability distribution of the previous state x k when the current state x k-1 is known, and the variance ∑ θ (x k , k) is usually fixed as represents the variance of the current noise state, while the mean μ θ (xk k) parameterized by θ, for noise estimation, the network ε θ Predicted diffusion input x k The noise of x
[0153]
[0154] Where x θ (x k k) is the output of the neural network θ, usually represents the predicted noise or intermediate state, used to adjust the denoising mean μ θ .
[0155] The parameter θ is learned by minimizing the loss:
[0156]
[0157] The conditional diffusion model is an extension of the diffusion model, which aims to introduce some conditions when generating data. The condition added in this invention is the label of the event type and the time step of diffusion, so as to generate data that meets the condition. In the forward diffusion process, the condition y is integrated into the data point, and then the noise is added step by step:
[0158] q(x t |x t-1 ,y) (15)
[0159] In the reverse diffusion process, the condition information y is used to guide the generation of data:
[0160] p θ (x t-1 |x t ,y) = N(x t-1 ; μ θ (x t ,y,t),∑ θ (x t ,y,t)) (16)
[0161] The goal of reverse diffusion is to train an efficient denoising network to recover the signal step by step by estimating the noise term. This invention selects the deep UNet model as the denoising network of the diffusion model, as shown in Figure 5As shown, this structure can efficiently fuse features of different levels, and is particularly suitable for processing image data. The architecture of the model includes down-sampling and up-sampling layers, which can simultaneously retain low-level details and high-level semantic information of the image, thereby enhancing the noise removal capability. In the down-sampling stage, a deep convolutional residual block is used to increase the depth of the network, so as to fully extract the features of the image, and a residual connection is combined to solve the problem of gradient disappearance. In the up-sampling stage, the input vector is dimensionally expanded through the deconvolution operation, and finally the noise is successfully predicted. At the same time, the event type label and the diffusion time step are added in the up-sampling and down-sampling processes, which can more accurately restore the signal and prevent information loss, thereby improving the generation performance of the model.
[0162] S5, input the generated two-dimensional signal feature map into the decoding module of the autoencoder to obtain new one-dimensional time series signal data, and construct an enhanced threat event signal data set according to the new one-dimensional time series signal data.
[0163] Optionally, the decoding module comprises a plurality of bottleneck blocks and residual blocks.
[0164] In a feasible implementation, the decoder decodes the generated two-dimensional signal feature map, extracts features through a series of two-dimensional convolution operations and one-dimensional convolution operations, and finally reconstructs a one-dimensional time series signal similar to the original signal. The model uses a deep network fused with three bottleneck blocks and residual blocks to extract features of the image, as shown in Figure 6 The bottleneck structure compresses the feature map with 1x1 convolution, which can reduce the number of parameters and improve the calculation efficiency of the model. At the same time, 3x3 convolution operations are used to obtain the information of the feature map, and finally restored to the original input image size. In the deep network structure, in order to alleviate the problem of gradient disappearance during gradient descent, a residual block is introduced, and the skip connection is applied in the bottleneck block, that is, the output vector of the bottleneck block is fused with the feature map of the same channel.
[0165] Based on the advantages of the diffusion model, the present application proposes a network architecture that fuses an autoencoder and a conditional diffusion model, aiming to enhance the threat event signal data set. The autoencoder is used to map the one-dimensional time series signal to a two-dimensional feature space, making it more suitable for processing by the diffusion model; the conditional diffusion model generates high-quality signals by gradually removing noise and generating high-quality signals under the guidance of conditional information such as event type, thereby expanding the data set. This method can effectively generate diversified signal samples, improve the performance and robustness of the threat event recognition model, and significantly improve the recognition accuracy, especially in the case of scarce data samples. Figure 7The network architecture used in the application is shown, which shows how the autoencoder and the conditional diffusion model work together to complete the signal generation task and provide high-quality training data for the subsequent threat event identification model.
[0166] Further, the overall architecture of the signal generation model proposed by the application mainly consists of two parts: signal space and latent space. In the signal space part, the self-encoder method is used for deep analysis of the signal. First, the one-dimensional time series signal is mapped to the two-dimensional feature space through the encoder, and then the signal in the two-dimensional space is reconstructed to the one-dimensional signal by using the decoder. In the latent space part, the conditional diffusion model is used to process the latent space of the signal, and the generated two-dimensional space signal is highly similar to the two-dimensional space signal expanded by the encoder. The diffusion model generates signals by gradually adding Gaussian noise and denoising process, where the denoising process is realized by a deep UNet model. Finally, a large number of one-dimensional time series signal data are generated by the model, enhancing the diversity of the data set and improving the generalization ability and recognition accuracy of the threat event identification model. Through this method, the problem of insufficient data samples can be effectively solved, and high-quality signal data for threat event identification are provided.
[0167] S6, training the oil and gas pipeline threat event identification model according to the enhanced threat event signal data set to complete the oil and gas pipeline threat event identification task.
[0168] In order to evaluate the generation effect of the diffusion model, the application uses the quality evaluation index commonly used in the field of image generation-SSIM (Structural Similarity Index). SSIM aims to simulate the perception characteristics of the human visual system and comprehensively evaluate the similarity of two images from multiple dimensions such as brightness, contrast and structural information. Its calculation is shown in formula (17). The closer the SSIM value is to 1, the more similar the generated image is to the original image, and the higher the quality is. In addition, the generation result is shown once every 100 epochs, as shown in Figure 8 By intuitively comparing the effect of the original image and the generated image, it can be clearly observed that the generation ability of the model is continuously improved. The SSIM value of the time-frequency graph generated by the conditional diffusion model gradually increases from 0.025 at 100 epochs to 0.446 at 200 epochs, and finally reaches the best effect at 300 epochs, with an SSIM value as high as 0.917, close to 1. At this time, the generated image is significantly optimized in detail, showing higher similarity and quality, fully proving the generation ability and effectiveness of the model. This shows that the model successfully realizes the high-quality reconstruction of the time-frequency image, laying a good foundation for subsequent research.
[0169]
[0170] where μ x and μ y are the mean of the two images respectively, and are the variance, C1 and C2 are small constants to prevent the denominator from being zero, and σ xy is the covariance.
[0171] The original signal is input into the generation model for training, where each event original signal contains 3000 sample data as shown in Figure 9 . After multiple optimization and training of the model, the best effect is finally achieved. Through the generation model, signals including three kinds of events of vehicles, manual operations and mechanical operations are successfully generated, as shown in Figure 10 to 12 . From the three generated signal graphs, it can be found that the signals of each type of event have unique characteristics and are highly similar in structure to the original signals. This means that the generation model can be used to generate high-quality signals to expand the data set.
[0172] The signal generation model generates 3000 sample data of vehicle signals, manual operation signals and mechanical operation signals respectively. In order to prove the authenticity of the generated signals, 30% of the original data set is used as the test set, and the newly generated data set is combined with the original 70% data set to form a new data set. In order to keep the number of event samples in the data set balanced, the noise signal data is expanded using the interpolation method to the same number as other events, and the final training sample number of the model is 20400.
[0173] In order to verify the enhancement effect of the new generated signal on the performance of the recognition model, the time-frequency graph convolution network is combined with the perception machine (Multilayer Perceptron, MLP) to realize the classification task of threat events. The unexpanded data set is divided into a training set and a test set in a ratio of 7:3. On the training set, 5-fold cross-validation method is used to evaluate the model and select hyperparameters. In this process, the heuristic search algorithm of grid search is used to search for the optimal combination of hyperparameters. After 5-fold cross-validation is completed, the average accuracy, average recall, average F1 score and standard deviation of the model are calculated by combining the results of the 5 validations, as shown in Table 1, to obtain an unbiased estimate of the performance of the model. The accuracy fluctuates very little between each fold (standard deviation is reasonable), showing the stability of the model, and the recall and F1 score are very close to the accuracy, reflecting the characteristics of a high-quality classifier. The optimal hyperparameters of the model are determined by cross-validation, and the optimal hyperparameters used by the final model and the environment are shown in Table 2.
[0174] Table 1 Cross-validation results of the original model
[0175]
[0176]
[0177] Table 2. Hyperparameters and environment of the original model
[0178]
[0179] After the dataset was expanded, the new training set was input into a time-frequency graph convolutional network (a combined model of the encoder and decoder parts of the generative model) for training. The hyperparameters and environment configuration were the same as the original training model, as shown in Table 2. The effectiveness and reliability of data augmentation were illustrated by comparing the performance of the new training model. Figure 13 As shown in the figure, the curves depicting the changes in average loss and average accuracy during model training are illustrated. In the early stages of training, the changes in model loss and accuracy are significant. In the later stages of training, the changes are gradual, eventually converging to an average accuracy of 99.4% with an error of 0.043. Compared to the training performance of the original time-frequency plot model, as shown... Figure 14 As shown, the accuracy of the new model is improved, and it achieves even higher accuracy and faster convergence in the early stages of training. Comparing the performance of the two models demonstrates that the data-augmented model not only improves generalization ability but also accelerates training efficiency. This indicates that the new data augmentation method is successful and can be used to generate high-quality new signals. Of course, a well-defined validation set is still needed. This allows us to detect overfitting during model training and to verify the correctness and reliability of the data augmentation method by testing the original signal dataset.
[0180] The experimental test results show that the newly generated data is highly similar to the original data in the feature distribution, the model maintains stable recognition accuracy on the validation set, and the test accuracy reaches 99.2%, and the classification test effect of the threat event is shown in Tables 3 and 4. It can be found from the tables that compared with the model without data enhancement, the accuracy of the new model is improved by 2.1%, the precision of manual operation and mechanical operation is improved by 5.4% and 2.8% respectively, so that the manual operation and mechanical operation events are more easily distinguished. The precision and recall rate of noise and manual operation type are both 1.000, indicating that the model has almost no error in identifying these events. The F1 score is also maintained at a high level in all types, indicating that the model does a good job in balancing precision and recall, especially for vehicle events (0.985) and mechanical operation events (0.985), the gap between the two indicators is small, meaning that false positives and false negatives are less. Compared with the model trained only using the original data, the generalization ability of the model is significantly improved after training using the fusion of the original data set and the generated data set, especially in the recognition effect of the few-sample class. This shows that the signal generation model plays a role in enhancing the diversity of the data set, and effectively improves the threat event recognition performance of the time-frequency graph convolution network.
[0181] Table 3 Classification data of the original model
[0182]
[0183] Table 4 Classification data of the model after expanding the data set
[0184]
[0185] In the embodiment of the application, in order to solve the problem that the prior art relies on simple noise addition or signal disturbance method, and fails to effectively consider the actual characteristics of time series signals, resulting in that the quality and diversity of the enhanced data cannot meet the needs of complex applications. The application proposes a data enhancement method based on diffusion model, which can significantly improve the diversity and representativeness of the data while ensuring the consistency of the signal characteristics, thereby improving the recognition accuracy and robustness of the classification model under the condition of few samples. By introducing a conditional diffusion mechanism, combined with event type information, the generation process is guided in a targeted manner, overcoming the problems of low quality and insufficient diversity of generated samples in traditional methods. This method not only can expand the data set, but also can enhance the performance of the threat event recognition model, especially suitable for the case of insufficient or uneven distribution of samples, and has wide application prospect.
[0186] Figure 15 It is a DAS signal generation device block diagram based on an autoencoder and a conditional diffusion model according to an exemplary embodiment, which is used for a DAS signal generation method based on an autoencoder and a conditional diffusion model.Figure 15 The apparatus comprises a construction module 310, an acquisition module 320, an encoding module 330, a generation module 340, a decoding module 350, and an output module 360.
[0187] The construction module 310 is configured to construct an autoencoder, train the autoencoder, and obtain a trained autoencoder.
[0188] The acquisition module 320 is configured to acquire a threat event signal data set to be enhanced.
[0189] The encoding module 330 is configured to input one-dimensional time sequence signal data in the threat event signal data set to be enhanced into an encoding module of the autoencoder to obtain a two-dimensional signal feature map.
[0190] The generation module 340 is configured to input the two-dimensional signal feature map into a conditional diffusion model to obtain a generated two-dimensional signal feature map.
[0191] The decoding module 350 is configured to input the generated two-dimensional signal feature map into a decoding module of the autoencoder to obtain new one-dimensional time sequence signal data, and construct an enhanced threat event signal data set according to the new one-dimensional time sequence signal data.
[0192] The output module 360 is configured to train an oil and gas pipeline threat event identification model according to the enhanced threat event signal data set, and complete an oil and gas pipeline threat event identification task.
[0193] In the embodiment of the present application, in order to solve the problem that the prior art relies on simple noise addition or signal disturbance method, and fails to effectively consider the actual characteristics of time sequence signal, resulting in that the quality and diversity of enhanced data cannot meet the needs of complex applications, the present application proposes a data enhancement method based on diffusion model, which can significantly improve the diversity and representativeness of data while ensuring the consistency of signal characteristics, thereby improving the recognition accuracy and robustness of the classification model under the condition of few samples. By introducing a conditional diffusion mechanism, combined with information such as event type, the generation process is guided in a targeted manner, overcoming the problems of low quality and insufficient diversity of generated samples in traditional methods. This method not only can expand the data set, but also can enhance the performance of the threat event identification model, and is particularly suitable for the case where the number of samples is scarce or unevenly distributed, and has a wide application prospect.
[0194] Figure 16 is a structural schematic diagram of a DAS signal generation device provided by the embodiment of the present application, as Figure 16 shown, the DAS signal generation device can comprise the DAS signal generation apparatus based on the autoencoder and the conditional diffusion model shown in Figure 15 Optionally, the DAS signal generation device 410 can comprise the first processor 2001.
[0195] Optionally, the DAS signal generating device 410 can further include a memory 2002 and a transceiver 2003.
[0196] The first processor 2001 is connected with the memory 2002 and the transceiver 2003, for example, through a communication bus.
[0197] The specific implementation of the DAS signal generating device 410 will be described below in conjunction with Figure 16 The specific implementation of the DAS signal generating device 410 will be described below in conjunction with
[0198] The first processor 2001 is the control center of the DAS signal generating device 410, which can be one processor or a plurality of processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present application, such as one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0199] Optionally, the first processor 2001 can execute various functions of the DAS signal generating device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0200] In a specific implementation, as an embodiment, the first processor 2001 can include one or more CPUs, such as the CPU0 and CPU1 shown in FIG. 2. Figure 16
[0201] In a specific implementation, as an embodiment, the DAS signal generating device 410 can also include a plurality of processors, such as the first processor 2001 and the second processor 2004 shown in FIG. 2. Each of these processors can be a single-CPU or a multi-CPU. The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). Figure 16
[0202] The memory 2002 is used to store software programs for implementing the schemes of the present application, and is controlled by the first processor 2001 to execute. The specific implementation can refer to the above method embodiments, and will not be described here.
[0203] Optionally, the memory 2002 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory 2002 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through an interface circuit (not shown in the figure) of the DAS signal generation device 410. The embodiments of the present application do not make specific limitations hereon. Figure 16
[0204] The transceiver 2003 is configured to communicate with a network device or a terminal device.
[0205] Optionally, the transceiver 2003 can include a receiver and a transmitter (not shown separately in the figure). The receiver is configured to implement a receiving function, and the transmitter is configured to implement a transmitting function. Figure 16
[0206] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through an interface circuit (not shown in the figure) of the DAS signal generation device 410. The embodiments of the present application do not make specific limitations hereon. Figure 16 It should be noted that the structure of the DAS signal generation device 410 shown in the figure does not constitute a limitation on the router. The actual knowledge structure identification device can include more or fewer components than those shown in the figure, or combine certain components, or different component arrangements.
[0207] Figure 16 In addition, the technical effects of the DAS signal generation device 410 can refer to the technical effects of the DAS signal generation method based on the autoencoder and the conditional diffusion model described in the above method embodiments, which will not be repeated here.
[0208]
[0209] It should be appreciated that the first processor 2001 in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or can also be any conventional processor.
[0210] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).
[0211] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A DAS signal generation method based on an autoencoder and a conditional diffusion model, characterized in that, The method includes: S1. Construct an autoencoder and train the autoencoder to obtain a trained autoencoder; S2. Obtain the dataset of threat event signals to be enhanced; S3. Input the one-dimensional time-series signal data from the threat event signal dataset to be enhanced into the encoding module of the autoencoder to obtain a two-dimensional signal feature map; S4. Input the two-dimensional signal feature map into the conditional diffusion model to obtain the generated two-dimensional signal feature map; S5. Input the generated two-dimensional signal feature map into the decoding module of the autoencoder to obtain new one-dimensional time-series signal data, and construct an enhanced threat event signal dataset based on the new one-dimensional time-series signal data. S6. Train the oil and gas pipeline threat event identification model based on the enhanced threat event signal dataset to complete the oil and gas pipeline threat event identification task.
2. The DAS signal generation method based on an autoencoder and a conditional diffusion model according to claim 1, characterized in that, The loss function of the autoencoder in S1 includes time domain loss, frequency domain loss, and similarity loss; The loss function of the autoencoder is shown in equation (1) below: L=L Time +L Freq +0.3·L Sim (1) In the formula, L represents the loss function of the autoencoder, L Time L represents the time-domain loss. Freq L represents the frequency domain loss. Sim This represents the signal similarity loss; The frequency domain loss is shown in equation (2) below: L Freq =||F(x(t))-F(x(t))||2 (2) In the formula, F(·) represents the Fourier transform, x(t) represents the input signal, x(t) represents the reconstructed signal, and ||·||2 represents the L2 norm; The similarity loss is shown in equation (3) below:
3. The DAS signal generation method based on an autoencoder and a conditional diffusion model according to claim 1, characterized in that, The encoding module includes a time-frequency graph convolution model, which includes a real part multi-scale convolutional layer and an imaginary part multi-scale convolutional layer. S3 includes: S31. Input one-dimensional time-series signal data into a real part multi-scale convolutional layer, and perform convolution operations through multiple convolution kernels of different dimensions in the real part multi-scale convolutional layer to extract the real part feature information of the signal. S32. Input one-dimensional time-series signal data into the imaginary part multi-scale convolutional layer, and perform convolution operations through multiple convolution kernels of different dimensions in the imaginary part multi-scale convolutional layer to extract the imaginary part feature information of the signal. S33. Take the absolute values of the real part feature information and the imaginary part feature information respectively to obtain the real part convolution result and the imaginary part convolution result; S34. The real part convolution result and the imaginary part convolution result are fused to obtain a two-dimensional signal feature map.
4. The DAS signal generation method based on an autoencoder and a conditional diffusion model according to claim 3, characterized in that, The convolution operation in S31 includes: The weights are initialized using the Fourier transform function and used as the kernel function for convolution operation, as shown in equation (4) below: h(j)=H r (j)+iH i (j) (4) in, H r (j)=cos(2πfj+Φ) (5) H i (j)=sin(2πfj+Φ) (6) In the formula, h(j) represents the Fourier transform coefficients corresponding to frequency point j, j represents the index of the frequency point, and H r (j) represents the real part of the Fourier kernel, H i (j) represents the imaginary part of the Fourier kernel, i represents the imaginary unit, f represents the frequency, and Φ represents the phase.
5. The DAS signal generation method based on an autoencoder and a conditional diffusion model according to claim 3, characterized in that, In step S33, the absolute values of the real part feature information and the imaginary part feature information are calculated respectively to obtain the real part convolution result and the imaginary part convolution result, as shown in equations (7)-(8) below: In the formula, y r [i] represents the real part of the Fourier transform result at position i, i represents the current starting position index, k represents the time step index within the window, x[i+j] represents the value of the signal at time i+k, and y i [i] represents the imaginary part of the Fourier transform result at position i, H r (j) represents the real part of the Fourier kernel, H i (j) represents the imaginary part of the Fourier kernel.
6. The DAS signal generation method based on an autoencoder and a conditional diffusion model according to claim 1, characterized in that, The conditions of the conditional diffusion model are the event type label and the time step of the diffusion. The denoising network for the conditional diffusion model is a deep UNet model.
7. The DAS signal generation method based on an autoencoder and a conditional diffusion model according to claim 1, characterized in that, The decoding module includes multiple bottleneck blocks and residual blocks.
8. A DAS signal generation device based on an autoencoder and a conditional diffusion model, wherein the DAS signal generation device based on the autoencoder and conditional diffusion model is used to implement the DAS signal generation method based on the autoencoder and conditional diffusion model as described in any one of claims 1-7, characterized in that, The device includes: A construction module is used to construct an autoencoder, train the autoencoder, and obtain a trained autoencoder; The acquisition module is used to acquire the dataset of threat event signals to be enhanced. The encoding module is used to input one-dimensional time-series signal data from the threat event signal dataset to be enhanced into the encoding module of the autoencoder to obtain a two-dimensional signal feature map; The generation module is used to input the two-dimensional signal feature map into the conditional diffusion model to obtain the generated two-dimensional signal feature map; The decoding module is used to input the generated two-dimensional signal feature map into the decoding module of the autoencoder to obtain new one-dimensional time-series signal data, and to construct an enhanced threat event signal dataset based on the new one-dimensional time-series signal data. The output module is used to train an oil and gas pipeline threat event identification model based on the enhanced threat event signal dataset to complete the oil and gas pipeline threat event identification task.
9. A DAS signal generation device, characterized in that, The DAS signal generation device includes: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Hybrid electromagnetic threat system identification method and system based on deep learning in strong confrontation environment
CN116383603A
Multi-behavior recommendation method based on diffusion contrast learning
CN118821839A