Oil and gas well yield prediction method and device based on pre-training large language model

Through the oil and gas well output prediction method based on pre-trained large language model, the characteristic sequence is processed using normalization and input embedding layers, and combined with freezing and fine-tuning strategies, the problem of low prediction accuracy of oil and gas well output is solved, achieving higher prediction accuracy and stability.

CN120471196APending Publication Date: 2025-08-12CHINA UNIV OF PETROLEUM (BEIJING)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510352611.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, oil and gas well output prediction accuracy is low, making it difficult to accurately respond to dynamic changes in bottom well pressure and flow.

Method used

The oil and gas well output prediction method based on the pre-trained large language model is adopted. The feature sequence is converted into vectorized standard feature sequences through the normalization module and input embedding layer of the trained oil and gas well output prediction model. The dynamic interaction relationship is learned using the large language model, and the oil and gas well output prediction task is adapted to the oil and gas well output prediction task through freezing and fine-tuning strategies.

Benefits of technology

It improves the accuracy of oil and gas well production forecasting, can accurately respond to dynamic changes, enhances the modeling ability of complex nonlinear and long-term trends, and improves the stability of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471196A_ABST
    Figure CN120471196A_ABST
Patent Text Reader

Abstract

The invention provides an oil and gas well yield prediction method and device based on a pre-trained large language model. The method comprises the following steps: acquiring a feature sequence according to multivariable production dynamic data to be predicted; and inputting the feature sequence into a trained oil and gas well yield prediction model, wherein the model comprises a normalization module, an input embedded layer, a prediction module, an output linear layer and an anti-normalization module. And the prediction module is obtained through freezing and fine tuning strategies by taking the trained large language model as a main network. Freezing and fine tuning strategies enable the prediction module to adapt to characteristics of oil and gas well yield prediction tasks. A vectorized standard feature sequence is obtained through a normalization module and an input embedding layer, and alignment of the feature sequence and a prediction module based on the trained large language model is achieved. And the prediction module carries out prediction according to the vectorized standard feature sequence, and a final oil and gas well yield prediction value is obtained through an output linear layer and an anti-normalization module. And the prediction precision of the oil-gas well yield is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of oil and gas well production, and in particular to a method and device for predicting oil and gas well production based on a pre-trained large language model. Background Art

[0002] Forecasting oil and gas well production is an important basis for assessing gas field reserves and resources. By predicting oil and gas well production, we can understand the gas production capacity of oil and gas wells at different stages in advance, thereby formulating scientific and reasonable production plans.

[0003] Currently, existing technologies use production decline curve analysis to predict oil and gas well production based on the assumptions of constant bottomhole pressure and declining flow rates. By fitting historical production data from oil and gas wells, patterns of production change over time are identified and corresponding production decline models are established, allowing for extrapolated predictions of future well production.

[0004] However, bottom hole pressure and flow are often not constant and deviate from actual production conditions, making it difficult to accurately respond to dynamic changes, resulting in reduced prediction accuracy of oil and gas well production. Summary of the Invention

[0005] The embodiments of the present application provide a method and device for predicting oil and gas well production based on a pre-trained large language model, so as to achieve the effect of improving the prediction accuracy of oil and gas well production.

[0006] In a first aspect, an embodiment of the present application provides an oil and gas well production prediction method based on a pre-trained large language model, comprising: receiving multivariate production dynamic data of an oil and gas well to be predicted sent by a data terminal; obtaining a feature sequence based on the multivariate production dynamic data to be predicted; inputting the feature sequence into a trained oil and gas well production prediction model, so that the trained oil and gas well production prediction model performs the following steps on the feature sequence: a normalization module of the trained oil and gas well production prediction model normalizes the feature sequence to obtain a standard feature sequence; an input embedding layer of the trained oil and gas well production prediction model, The standard feature sequence is vectorized to obtain a vectorized standard feature sequence; the prediction module of the trained oil and gas well production prediction model predicts the future oil and gas well production based on the vectorized standard feature sequence, and outputs the initial oil and gas well production prediction value through the output linear layer of the trained oil and gas well production prediction model; the prediction module is obtained by freezing and fine-tuning strategies with the trained large language model as the backbone network; the denormalization module of the trained oil and gas well production prediction model restores the initial oil and gas well production prediction value to the original numerical range to obtain the final oil and gas well production prediction value.

[0007] In one possible implementation, before receiving the multivariate production dynamic data of the oil and gas wells to be predicted sent by the data terminal, it also includes: constructing a prediction module through a trained large language model; wherein the prediction module includes a position encoding layer, a multi-head attention layer, a feedforward layer and a layer normalization layer; adding a normalization module and an input embedding layer at the input end of the prediction module, and adding an output linear layer and an anti-normalization module at the output end of the prediction module to obtain an initial oil and gas well production prediction model; freezing the parameters of the multi-head attention layer and the parameters of the feedforward layer of the initial oil and gas well production prediction model, and adjusting the parameters of the position encoding layer and the parameters of the layer normalization layer of the initial oil and gas well production prediction model to obtain the oil and gas well production prediction model to be trained.

[0008] In one possible implementation, after obtaining the oil and gas well production prediction model to be trained, the method further includes: receiving historical multivariable production dynamic data of the oil and gas well sent by a data terminal; wherein the historical multivariable production dynamic data is arranged according to the production time of the oil and gas well production; dividing the historical multivariable production dynamic data into a training set and a test set according to the production time; inputting the training set into the oil and gas well production prediction model to be trained, so that the oil and gas well production prediction model to be trained is trained according to the training set through a cosine annealing learning rate scheduling strategy to obtain a trained oil and gas well production prediction model; and inputting the test set into the trained oil and gas well production prediction model, so that the trained oil and gas well production prediction model is verified according to the test set.

[0009] In a possible implementation, obtaining a feature sequence based on the multivariable production dynamic data to be predicted includes: cutting the multivariable production dynamic data to be predicted according to the length of a preset input sequence through a sliding window to obtain a feature sequence.

[0010] In one possible implementation, the input embedding layer of the trained oil and gas well production prediction model vectorizes the standard feature sequence to obtain a vectorized standard feature sequence, and the formula is:

[0011]

[0012] Where, represents the vectorized standard feature sequence, represents the standard feature sequence, represents a real number, represents the dimension of the vector space; represents the batch size of the vectorized standard feature sequence, represents the variable dimension, represents the length of the feature sequence, Represents the feature dimensions required by the prediction module.

[0013] In one possible implementation, the prediction module of the trained oil and gas well production prediction model predicts future oil and gas well production based on the vectorized standard feature sequence. The formula is:

[0014]

[0015] Where, represents the output features of the prediction module, Represents a vectorized standard feature sequence. Accordingly, the output linear layer of the trained oil and gas well production prediction model is used to output an initial oil and gas well production prediction value, including: outputting the feature through the output linear layer of the trained oil and gas well production prediction model to output the initial oil and gas well production prediction value.

[0016] In one possible implementation, the normalization module of the trained oil and gas well production prediction model normalizes the feature sequence to obtain a standard feature sequence, including: the normalization module of the trained oil and gas well production prediction model calculates the mean and variance of the feature sequence, and uses a normalization method to obtain a standard feature sequence based on the mean and variance.

[0017] In a second aspect, an embodiment of the present application provides an oil and gas well production prediction device based on a pre-trained large language model, comprising:

[0018] A receiving module, configured to receive the multivariable production dynamic data of the oil and gas well to be predicted sent by the data terminal;

[0019] An acquisition module is used to obtain a feature sequence based on the multivariate production dynamic data to be predicted;

[0020] An input module is used to input the feature sequence into the trained oil and gas well production prediction model, so that the trained oil and gas well production prediction model performs the following steps on the feature sequence;

[0021] The input module includes:

[0022] Normalization unit, which is used for the normalization module of the trained oil and gas well production prediction model, normalizes the feature sequence to obtain a standard feature sequence;

[0023] The vectorization unit is used for the input embedding layer of the trained oil and gas well production prediction model to vectorize the standard feature sequence to obtain a vectorized standard feature sequence;

[0024] The output unit is used for the prediction module of the trained oil and gas well production prediction model. It predicts the future oil and gas well production based on the vectorized standard feature sequence and outputs the initial oil and gas well production prediction value through the output linear layer of the trained oil and gas well production prediction model. The prediction module is based on the trained large language model as the backbone network and is obtained through the freezing and fine-tuning strategy.

[0025] The acquisition unit is used for the denormalization module of the trained oil and gas well production prediction model to restore the initial oil and gas well production prediction value to the original numerical range to obtain the final oil and gas well production prediction value.

[0026] In a third aspect, an embodiment of the present application provides a server, comprising: a memory, a processor;

[0027] The memory stores computer-executable instructions;

[0028] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0029] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.

[0030] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0031] The embodiments of the present application provide a method and device for predicting oil and gas well production based on a pre-trained large language model. The multivariate production dynamic data is dynamic data that changes with production time and has time series dynamics. A feature sequence of the multivariate production dynamic data to be predicted is obtained. The backbone network of the prediction module is a trained large language model. The feature sequence is converted into a word sequence representation that can be received by the prediction module based on the trained large language model through the normalization module and input embedding layer of the trained oil and gas well production prediction model, thereby obtaining a vectorized standard feature sequence. The prediction module based on the large language model can accurately respond to dynamic changes and improve prediction accuracy by learning the dynamic interaction relationship in the vectorized standard feature sequence. The freezing strategy can retain existing knowledge while fully adapting to the characteristics of the oil and gas well production prediction task during the fine-tuning process, further improving prediction accuracy. In addition, the feature sequence is normalized and vectorized through the normalization module and input embedding layer of the trained oil and gas well production prediction model, so that it is adapted to the prediction module based on the trained large language model, achieving alignment between the feature sequence and the prediction module. In addition, the trained large language model uses massive pre-training knowledge to enable the prediction module to enhance its modeling capabilities for complex nonlinearities and long-term trends, as well as its generalization capabilities and improve stability, thereby further improving the accuracy of oil and gas well production predictions. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0033] Figure 1 A schematic diagram of a scenario of an oil and gas well production prediction method based on a pre-trained large language model provided in an embodiment of the present application;

[0034] Figure 2 A flow chart of a method for predicting oil and gas well production based on a pre-trained large language model provided in an embodiment of the present application;

[0035] Figure 3 A schematic diagram of the framework of the oil and gas well production prediction model provided in the embodiment of the present application;

[0036] Figure 4a The prediction results on the oil and gas well test set provided in the embodiment of the present application;

[0037] Figure 4b The prediction results on the second test set of oil and gas wells provided in the embodiment of this application;

[0038] Figure 4c The prediction results on the three test sets of oil and gas wells provided in the embodiment of this application;

[0039] Figure 4dThe prediction results on the four test sets of oil and gas wells provided in the embodiment of this application;

[0040] Figure 5 A schematic diagram of the structure of an oil and gas well production prediction device based on a pre-trained large language model provided in an embodiment of the present application;

[0041] Figure 6 A schematic diagram of the structure of the server provided in an embodiment of the present application.

[0042] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0043] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0044] Figure 1 A schematic diagram of a scenario of an oil and gas well production prediction method based on a pre-trained large language model provided in an embodiment of the present application, such as Figure 1 As shown, the server provided by this embodiment includes: a receiving device 101, a processor 102 and a display device 103.

[0045] It is understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the oil and gas well production prediction method based on a pre-trained large language model. In other feasible embodiments of this application, the above architecture may include more or fewer components than shown, or may combine or split certain components, or arrange the components differently. The specific configuration can be determined based on the actual application scenario and is not limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0046] In a specific implementation process, the receiving device 101 may be an input / output interface or a communication interface, and may receive the multivariable production dynamic data to be predicted of the oil and gas wells sent by the data terminal.

[0047] The processor 102 may perform a series of processing on the multivariable production dynamic data to be predicted to obtain a predicted value of the oil and gas well production.

[0048] The display device 103 can be used to display the predicted value of the oil and gas well production.

[0049] It should be understood that the above-mentioned processor can be implemented by the processor reading instructions in the memory and executing the instructions, or it can be implemented by a chip circuit.

[0050] In addition, the network architecture and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0051] Predicting oil and gas well production is an important basis for assessing gas field reserves and resources. By predicting oil and gas well production, we can understand in advance the gas production capacity of oil and gas wells at different stages, thereby formulating scientific and reasonable production plans. Currently, existing technologies use production decline curve analysis to predict oil and gas well production based on the assumption of constant bottomhole pressure and declining flow. By fitting historical production data during the oil and gas well production process, the pattern of production changes over time is identified, and a corresponding production decline model is established to extrapolate and predict future oil and gas well production. However, bottomhole pressure and flow are often not constant and deviate from actual production conditions, making it difficult to accurately respond to dynamic changes, resulting in reduced prediction accuracy of oil and gas well production.

[0052] To address the aforementioned technical issues, the present application proposes the following technical concept: Considering that bottomhole pressure and flow are often not constant and deviate from actual production conditions, it is difficult to accurately respond to dynamic changes, resulting in reduced prediction accuracy for oil and gas well production. The inventors devised the idea of using multivariate production dynamic data as the data to be predicted. Multivariate production dynamic data is dynamic data that changes over time. A feature sequence of the multivariate production dynamic data to be predicted is obtained. The backbone network of the prediction module of the trained oil and gas well production prediction model is a trained large language model. The normalization module and input embedding layer of the trained oil and gas well production prediction model convert the feature sequence into a word sequence representation that can be received by the prediction module based on the trained large language model, thereby obtaining a vectorized standard feature sequence. By learning the dynamic interactions in the vectorized standard feature sequence, the prediction module based on the trained large language model can accurately respond to dynamic changes and improve prediction accuracy. Furthermore, a freezing strategy can preserve existing knowledge while fully adapting to the characteristics of the oil and gas well production prediction task during fine-tuning, further improving prediction accuracy. In addition, the alignment of feature sequence and prediction module is achieved through the design of normalization module and input embedding layer module.

[0053] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0054] Figure 2 A flow chart of the oil and gas well production prediction method based on a pre-trained large language model provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the method includes:

[0055] S201: Receive multivariable production dynamic data to be predicted of an oil and gas well sent by a data terminal.

[0056] The multivariable production dynamic data to be predicted is dynamic data that changes over time and has a time series dynamic nature. The multivariable production dynamic data to be predicted includes, but is not limited to, production data, production time, oil pressure, casing pressure, permeability, porosity, and reservoir pressure.

[0057] In this embodiment, the multivariate production dynamic data to be predicted is sorted and preprocessed, including missing value filling and data smoothing.

[0058] Alternatively, the multivariate production performance data to be predicted may be the multivariate production performance data of the oil and gas well for the previous m days, and the production of the oil and gas well for the next n days is predicted based on the multivariate production performance data for the previous m days.

[0059] S202: Acquire a feature sequence based on the multivariable production dynamic data to be predicted.

[0060] Specifically, the multivariate production dynamic data to be predicted is cut according to the length of the preset input sequence by means of a sliding window to obtain a feature sequence.

[0061] In this embodiment, the feature sequence includes data such as production rate, production time, oil pressure, casing pressure, permeability, porosity and reservoir pressure.

[0062] In this embodiment, the size of the sliding window corresponds to the length of the input sequence and the output sequence.

[0063] S203: Input the feature sequence into the trained oil and gas well production prediction model, so that the trained oil and gas well production prediction model performs the following steps on the feature sequence.

[0064] Specifically, step S203 includes S2031 to S2034:

[0065] S2031: The normalization module of the trained oil and gas well production prediction model normalizes the feature sequence to obtain a standard feature sequence.

[0066] Specifically, the normalization module of the trained oil and gas well production prediction model calculates the mean and variance of the feature sequence, and uses the normalization method to obtain the standard feature sequence based on the mean and variance.

[0067] Specifically, the formula for calculating the mean of the feature sequence is:

[0068]

[0069] Where, represents a feature sequence, represents the mean of the feature sequence, represents the length of the feature sequence, Indicates the first data.

[0070] Specifically, the formula for calculating the variance of the feature sequence is:

[0071]

[0072] Where, represents the variance of the feature sequence, represents the length of the feature sequence, Indicates the first data, represents the mean of the feature sequence.

[0073] Specifically, the normalization method is used to obtain the standard feature sequence based on the mean and variance. The formula is:

[0074]

[0075] In this embodiment, Represents a constant close to 0, used to avoid zero division errors in numerical calculations. Each value in is standardized to obtain the standard feature sequence .

[0076] In this embodiment, normalization is used to eliminate dimensional differences in feature sequences. The mean and variance of the feature sequences are calculated and used to standardize the feature sequences, making the input feature sequences more evenly distributed across all dimensions. This facilitates the transfer of pre-trained knowledge in the prediction module based on the trained large language model across different tasks and data.

[0077] S2032: The input embedding layer of the trained oil and gas well production prediction model vectorizes the standard feature sequence to obtain a vectorized standard feature sequence.

[0078] In this embodiment, the processing results of the input embedding layer of the oil and gas well yield prediction model are input into the prediction module of the oil and gas well yield prediction model. The prediction module of the oil and gas well yield prediction model uses the trained large language model as its backbone network. Therefore, the input embedding layer is required to map the standard feature sequence to the dimensional space required by the prediction module based on the trained large language model. In other words, it is necessary to convert it into a word-meta sequence representation that can be accepted by the prediction module based on the trained large language model, thereby achieving vectorization of the standard feature sequence.

[0079] Specifically, the formula for vectorization processing is:

[0080]

[0081] Where, represents the vectorized standard feature sequence, represents the standard feature sequence, represents a real number, represents the dimension of the vector space; represents the batch size of the vectorized standard feature sequence, represents the variable dimension, represents the length of the feature sequence, Represents the feature dimensions required by the prediction module.

[0082] S2033: The prediction module of the trained oil and gas well production prediction model predicts the future oil and gas well production based on the vectorized standard feature sequence, and outputs the initial oil and gas well production prediction value through the output linear layer of the trained oil and gas well production prediction model; the prediction module is obtained by freezing and fine-tuning strategies with the trained large language model as the backbone network.

[0083] In this embodiment, the prediction module's backbone network is a pre-trained large language model. The prediction module includes a positional encoding module, a multi-head attention layer, a feedforward layer, and a layer normalization layer. The parameters of the multi-head attention layer and the feedforward layer are frozen, while the parameters of the positional encoding module and the layer normalization layer are fine-tuned.

[0084] Specifically, the prediction module of the trained oil and gas well production prediction model predicts the future oil and gas well production based on the vectorized standard feature sequence. The formula is:

[0085]

[0086] Where, represents the output features of the prediction module, Represents a vectorized standard feature sequence.

[0087] In this embodiment, the output features pass through the output linear layer of the trained oil and gas well production prediction model to output the initial oil and gas well production prediction value.

[0088] S2034: A denormalization module of the trained oil and gas well production prediction model restores the initial oil and gas well production prediction value to the original numerical range to obtain the final oil and gas well production prediction value.

[0089] In this embodiment, since the data input into the prediction module is normalized and has a value range of 0-1, the initial oil and gas well production prediction value needs to be restored to the original data range to obtain the true oil and gas well production prediction value.

[0090] In summary, multivariate production dynamics data is dynamic data that changes over time and exhibits temporal dynamics. A feature sequence of the multivariate production dynamics data to be predicted is obtained. The prediction module's backbone network is a trained large language model. The normalization module and input embedding layer of the trained oil and gas well production prediction model transform the feature sequence into a word sequence representation that can be accepted by the prediction module based on the trained large language model, resulting in a vectorized standard feature sequence. By learning the dynamic interactions within the vectorized standard feature sequence, the prediction module based on the trained large language model can accurately respond to dynamic changes and improve prediction accuracy. A freezing strategy preserves existing knowledge while fully adapting to the specific characteristics of the oil and gas well production prediction task during fine-tuning, further improving prediction accuracy. Furthermore, the normalization module and input embedding layer of the trained oil and gas well production prediction model normalize and vectorize the feature sequence, enabling it to be aligned with the prediction module based on the trained large language model, achieving alignment between the feature sequence and the prediction module. In addition, the trained large language model uses massive pre-training knowledge to enable the prediction module to enhance its modeling capabilities for complex nonlinearities and long-term trends, as well as its generalization capabilities and improve stability, thereby further improving the accuracy of oil and gas well production predictions.

[0091] Figure 3 This is a schematic diagram of the framework of the oil and gas well production prediction model provided in the embodiment of this application. Based on the above embodiment, in this embodiment, the training process of the oil and gas well production prediction model is introduced, as detailed below:

[0092] S301: Construct a prediction module using the trained large language model; the prediction module includes a position encoding layer, a multi-head attention layer, a feedforward layer, and a layer normalization layer.

[0093] Optionally, the large language model can be GPT-2.

[0094] S302: Add a normalization module and an input embedding layer to the input end of the prediction module, and add an output linear layer and an anti-normalization module to the output end of the prediction module to obtain an initial oil and gas well production prediction model.

[0095] In this embodiment, an input module and an output module are constructed. Figure 3 As shown in Figure 2, a normalization module and an input embedding layer are added at the input end, and an output linear layer and an anti-normalization module are added at the output end.

[0096] In this embodiment, the input to the prediction module based on the trained large language model should be a text sequence. Therefore, by reconstructing the input module and adding a normalization module and input embedding layer to the input of the prediction module, the feature sequence obtained from the multivariate production dynamic data is converted into a word sequence representation that can be received by the prediction module based on the trained large language model, resulting in a vectorized standard feature sequence. An output linear layer and a denormalization module are added to the output of the prediction module to convert the output features of the prediction module into a predicted oil and gas well production value.

[0097] S303: Freeze the parameters of the multi-head attention layer and the feedforward layer of the initial oil and gas well production prediction model, and adjust the parameters of the position encoding layer and the layer normalization layer of the initial oil and gas well production prediction model to obtain the oil and gas well production prediction model to be trained.

[0098] S304: Receive historical multivariable production dynamic data of oil and gas wells sent by the data terminal; wherein the historical multivariable production dynamic data are arranged according to the production time of the oil and gas well output.

[0099] like Figure 3 As shown in the figure, the parameters of the multi-head attention mechanism and the feedforward layer are frozen, the learned knowledge is retained, and the parameters of the position encoding layer and the layer normalization layer are fine-tuned to ensure that they can adapt to the characteristics of the oil and gas well production prediction task.

[0100] In this embodiment, the prediction module carries the rich context information and feature expression capabilities learned on a large-scale corpus.

[0101] S305: Divide the historical multivariate production dynamic data into a training set and a test set according to production time.

[0102] In this example, the training set is segmented according to the length of the preset input sequence using a sliding window approach to generate feature sequences and label sequences. The feature sequence includes data such as production, production time, oil pressure, casing pressure, permeability, porosity, and reservoir pressure, while the label sequence only includes production data.

[0103] Similarly, the test set is cut according to the length of the preset input sequence through a sliding window method to obtain a feature sequence and a label sequence.

[0104] S306: Inputting the training set into the oil and gas well production prediction model to be trained, so that the oil and gas well production prediction model to be trained is trained according to the training set through the cosine annealing learning rate scheduling strategy to obtain a trained oil and gas well production prediction model.

[0105] Specifically, the feature sequence and label sequence corresponding to the test set are input into the oil and gas well production prediction model to be trained, so that the oil and gas well production prediction model to be trained is trained according to the feature sequence and label sequence corresponding to the training set through the cosine annealing learning rate scheduling strategy to obtain a trained oil and gas well production prediction model.

[0106] In this embodiment, the cosine annealing learning rate scheduling strategy is introduced to make the learning rate Dynamically reduce to , thus adapting to different training stages.

[0107] Alternatively, multivariate production dynamics data from 72 carbonate gas wells spanning approximately 10 years can be used. This multivariate production dynamics data includes multiple dimensions: production time, oil pressure, casing pressure, daily production, permeability, porosity, and reservoir pressure. The multivariate production dynamics data is divided into training and test sets according to a preset ratio. Preprocess the multivariate production dynamics data, including missing value filling, data smoothing, and data normalization. Set reasonable input and output lengths. For example, input daily production for one month and output daily production for the next month. Construct the feature sequence and label sequence required for the model using a sliding window approach. Use the oil and gas well production prediction model to predict daily production for the next month.

[0108] Optionally, during training, the Adam optimizer can be used, with mean squared error as the loss function. Model hyperparameters, such as the number of backbone network layers, learning rate, batch size, and number of iterations, can be adjusted as needed to optimize model performance and prediction accuracy. The model's input and output sequence lengths can also be adjusted to accommodate different prediction lengths. The model's prediction performance can be evaluated on the test set using various metrics, such as root mean squared error and mean absolute percentage error.

[0109] S307: Input the test set into the trained oil and gas well production prediction model, so that the trained oil and gas well production prediction model is verified based on the test set.

[0110] In this embodiment, the feature sequence corresponding to the test set is input into the trained oil and gas well production prediction model to obtain the prediction result, which is then verified based on the label sequence corresponding to the test set.

[0111] In summary, adding a normalization module and input embedding layer to the input transforms the data input to the prediction module into a word sequence representation that can be accepted by the prediction module based on the trained large language model. This achieves alignment between the input data and the prediction module based on the trained large language model. Combining a freezing and fine-tuning strategy, the oil and gas well production prediction model is adapted to the characteristics of the oil and gas well production prediction task, eliminating the need for retraining. This reduces the requirements for training sets and computing resources and improves training efficiency. The final oil and gas well production prediction results are obtained after the output linear layer and denormalization module are added to the output. This improves the oil and gas well production prediction model's ability to handle multivariate variables and adapt to complex trends, providing a foundation for oil and gas well production prediction.

[0112] For example, Table 1 lists some parameters of the oil and gas well production prediction model. As shown in Table 1, Backbone represents the backbone network; Freq represents the temporal granularity of the data, with 1D representing daily sampling data; Number of GPT layers represents the number of GPT-2 layers in the model; Model Dimension represents the spatial dimension, which is 12 in GPT-2; Number of heads represents the number of heads in the multi-head attention layer, which is 12 in GPT-2; Sequence length represents the length of the input sequence; Prediction length represents the length of the predicted sequence; Batch size represents the sample batch size for each training session; and Epochs represents the number of training iterations.

[0113] Table 1 Some parameters of the oil and gas well production prediction model

[0114]

[0115] The oil and gas well production prediction model is set with the parameters in Table 1 and predictions are made on a test set of four oil and gas wells. The training set length for each well is 2880 days and the test set length is 720 days. Figure 4a The prediction results on the oil and gas well test set provided in the embodiment of the present application; Figure 4b The prediction results on the second test set of oil and gas wells provided in the embodiment of this application; Figure 4c The prediction results on the three test sets of oil and gas wells provided in the embodiment of this application; Figure 4dPrediction results for the four oil and gas well test sets provided in the examples of this application. The oil and gas well production prediction model based on the trained large language model showed relatively stable prediction results across the entire test set. It was able to maintain high prediction accuracy when faced with complex fluctuations and sudden changes in the data. When predicting long-sequence production data, the model was able to maintain relatively accurate predictions over a wide time span. The mean absolute percentage error on the test set was 4.813%, indicating that the model had high prediction accuracy and good performance.

[0116] Figure 5 A schematic diagram of the structure of an oil and gas well production prediction device based on a pre-trained large language model provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the oil and gas well production prediction device based on the pre-trained large language model provided in this embodiment includes: a receiving module 501, an acquisition module 502 and an input module 503; wherein the input module 503 includes a normalization unit 5031, a vectorization unit 5032, an output unit 5033 and an acquisition unit 5034.

[0117] The receiving module 501 is used to receive the multivariable production dynamic data of the oil and gas well to be predicted sent by the data terminal;

[0118] An acquisition module 502 is used to acquire a feature sequence based on the multivariable production dynamic data to be predicted;

[0119] An input module 503 is used to input the feature sequence into the trained oil and gas well production prediction model, so that the trained oil and gas well production prediction model performs the following steps on the feature sequence;

[0120] The input module 503 includes:

[0121] Normalization unit 5031 is used for normalization of the trained oil and gas well production prediction model, normalizing the feature sequence to obtain a standard feature sequence;

[0122] The vectorization unit 5032 is used for the input embedding layer of the trained oil and gas well production prediction model to vectorize the standard feature sequence to obtain a vectorized standard feature sequence;

[0123] Output unit 5033 is used for the prediction module of the trained oil and gas well production prediction model to predict the future oil and gas well production based on the vectorized standard feature sequence and output the initial oil and gas well production prediction value through the output linear layer of the trained oil and gas well production prediction model. The prediction module is based on the trained large language model as the backbone network and is obtained through the freezing and fine-tuning strategy.

[0124] The acquisition unit 5034 is used for the denormalization module of the trained oil and gas well production prediction model to restore the initial oil and gas well production prediction value to the original value range to obtain the final oil and gas well production prediction value.

[0125] In one possible embodiment, the oil and gas well production prediction device based on a pre-trained large language model also includes: a training module, which is used to construct a prediction module through the trained large language model; wherein the prediction module includes a position encoding layer, a multi-head attention layer, a feedforward layer and a layer normalization layer; a normalization module and an input embedding layer are added to the input end of the prediction module, and an output linear layer and an anti-normalization module are added to the output end of the prediction module to obtain an initial oil and gas well production prediction model; the parameters of the multi-head attention layer and the parameters of the feedforward layer of the initial oil and gas well production prediction model are frozen, and the parameters of the position encoding layer and the parameters of the layer normalization layer of the initial oil and gas well production prediction model are adjusted to obtain the oil and gas well production prediction model to be trained.

[0126] In one possible implementation, the training module is also used to receive historical multivariable production dynamic data of oil and gas wells sent by a data terminal; wherein the historical multivariable production dynamic data is arranged according to the production time of the oil and gas well production; according to the production time, the historical multivariable production dynamic data is divided into a training set and a test set; the training set is input into the oil and gas well production prediction model to be trained, so that the oil and gas well production prediction model to be trained is trained according to the training set through a cosine annealing learning rate scheduling strategy to obtain a trained oil and gas well production prediction model; the test set is input into the trained oil and gas well production prediction model, so that the trained oil and gas well production prediction model is verified according to the test set.

[0127] In a possible implementation, the vectorization unit 5032 is configured to cut the multivariate production dynamic data to be predicted according to the length of a preset input sequence by means of a sliding window to obtain a feature sequence.

[0128] In a possible implementation, the formula of the vectorization unit 5032′ is:

[0129]

[0130] Where, represents the vectorized standard feature sequence, represents the standard feature sequence, represents a real number, represents the dimension of the vector space; represents the batch size of the vectorized standard feature sequence, represents the variable dimension, represents the length of the feature sequence, Represents the feature dimensions required by the prediction module.

[0131] In one possible implementation, the prediction module of the trained oil and gas well production prediction model predicts future oil and gas well production based on the vectorized standard feature sequence. The formula is:

[0132]

[0133] Where, represents the output features of the prediction module, Represents a vectorized standard feature sequence. Accordingly, the output unit 5033 is used to output the features through the output linear layer of the trained oil and gas well production prediction model to output the initial oil and gas well production prediction value.

[0134] In a possible implementation, the normalization unit 5031 is a normalization module for the trained oil and gas well production prediction model, calculates the mean and variance of the feature sequence, and uses a normalization method to obtain a standard feature sequence based on the mean and variance.

[0135] The oil and gas well production prediction device based on the pre-trained large language model provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar, and are not described in detail in this embodiment.

[0136] Figure 6 This is a schematic diagram of the structure of the server provided in the embodiment of the present application. Figure 6 As shown, the server provided in this embodiment includes: at least one processor 601 and a memory 602. Optionally, the server also includes a communication component 603. The processor 601, the memory 602 and the communication component 603 are connected via a bus.

[0137] During the specific implementation process, at least one processor 601 executes the computer-executable instructions stored in the memory 602, so that the at least one processor 601 performs the above method.

[0138] The specific implementation process of the processor 601 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0139] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0140] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0141] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0142] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0143] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0144] The readable storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0145] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0146] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.

[0147] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0148] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0149] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0150] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0151] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.

Claims

1. A method for predicting oil and gas well production based on a pre-trained large language model, characterized in that: Applicable to servers, including: Receiving multivariable production dynamic data to be predicted of oil and gas wells sent by a data terminal; Acquiring a characteristic sequence according to the multivariable production dynamic data to be predicted; The feature sequence is input into a trained oil and gas well production prediction model, so that the trained oil and gas well production prediction model performs the following steps on the feature sequence: The normalization module of the trained oil and gas well production prediction model performs normalization processing on the feature sequence to obtain a standard feature sequence; The input embedding layer of the trained oil and gas well production prediction model performs vectorization processing on the standard feature sequence to obtain a vectorized standard feature sequence; The prediction module of the trained oil and gas well production prediction model predicts future oil and gas well production based on the vectorized standard feature sequence and outputs an initial oil and gas well production prediction value through the output linear layer of the trained oil and gas well production prediction model; wherein the prediction module is obtained by freezing and fine-tuning the trained large language model as the backbone network; The denormalization module of the trained oil and gas well production prediction model restores the initial oil and gas well production prediction value to the original numerical range to obtain the final oil and gas well production prediction value.

2. The method according to claim 1, characterized in that Before receiving the multivariable production dynamic data of the oil and gas well to be predicted sent by the data terminal, the method further includes: Constructing a prediction module using a trained large language model; wherein the prediction module includes a position encoding layer, a multi-head attention layer, a feedforward layer, and a layer normalization layer; Adding a normalization module and an input embedding layer to the input end of the prediction module, and adding an output linear layer and an anti-normalization module to the output end of the prediction module to obtain an initial oil and gas well production prediction model; The parameters of the multi-head attention layer and the parameters of the feedforward layer of the initial oil and gas well production prediction model are frozen, and the parameters of the position encoding layer and the parameters of the layer normalization layer of the initial oil and gas well production prediction model are adjusted to obtain the oil and gas well production prediction model to be trained.

3. The method according to claim 2, characterized in that After obtaining the oil and gas well production prediction model to be trained, the method further includes: Receiving historical multivariable production dynamic data of oil and gas wells sent by a data terminal; wherein the historical multivariable production dynamic data are arranged according to the production time of the oil and gas wells; Dividing the historical multivariate production dynamic data into a training set and a test set according to the production time; Inputting the training set into the oil and gas well production prediction model to be trained, so that the oil and gas well production prediction model to be trained is trained according to the training set through a cosine annealing learning rate scheduling strategy to obtain a trained oil and gas well production prediction model; The test set is input into the trained oil and gas well production prediction model, so that the trained oil and gas well production prediction model is verified based on the test set.

4. The method according to claim 1, wherein The step of obtaining a feature sequence based on the multivariable production dynamic data to be predicted includes: By means of a sliding window, the multivariate production dynamic data to be predicted is cut according to the length of a preset input sequence to obtain a feature sequence.

5. The method according to claim 4, characterized in that The input embedding layer of the trained oil and gas well production prediction model vectorizes the standard feature sequence to obtain a vectorized standard feature sequence, and the formula is: Where, represents the vectorized standard feature sequence, represents the standard feature sequence, represents a real number, represents the dimension of the vector space; represents the batch size of the vectorized standard feature sequence, represents the variable dimension, represents the length of the feature sequence, Represents the feature dimension required by the prediction module.

6. The method according to claim 5, characterized in that The prediction module of the trained oil and gas well production prediction model predicts the future oil and gas well production based on the vectorized standard feature sequence. The formula is: Where, represents the output features of the prediction module, Represents a vectorized standard feature sequence; Accordingly, the output linear layer of the trained oil and gas well production prediction model outputs the initial oil and gas well production prediction value, including: The output features pass through the output linear layer of the trained oil and gas well production prediction model to output an initial oil and gas well production prediction value.

7. The method according to any one of claims 1 to 6, characterized in that The normalization module of the trained oil and gas well production prediction model performs normalization processing on the feature sequence to obtain a standard feature sequence, including: The normalization module of the trained oil and gas well production prediction model calculates the mean and variance of the feature sequence and obtains a standard feature sequence based on the mean and variance using a normalization method.

8. An oil and gas well production prediction device based on a pre-trained large language model, characterized in that: Applicable to servers, including: A receiving module, configured to receive the multivariable production dynamic data of the oil and gas well to be predicted sent by the data terminal; An acquisition module, configured to acquire a feature sequence based on the multivariable production dynamic data to be predicted; An input module, configured to input the feature sequence into a trained oil and gas well production prediction model, so that the trained oil and gas well production prediction model performs the following steps on the feature sequence; Wherein, the input module includes: A normalization unit, which is used for the normalization module of the trained oil and gas well production prediction model, normalizes the feature sequence to obtain a standard feature sequence; A vectorization unit, used for the input embedding layer of the trained oil and gas well production prediction model, vectorizes the standard feature sequence to obtain a vectorized standard feature sequence; An output unit, a prediction module for the trained oil and gas well production prediction model, predicts future oil and gas well production based on the vectorized standard feature sequence, and outputs an initial oil and gas well production prediction value through the output linear layer of the trained oil and gas well production prediction model; wherein the prediction module is obtained by freezing and fine-tuning the trained large language model as the backbone network; The acquisition unit is used for the denormalization module of the trained oil and gas well production prediction model to restore the initial oil and gas well production prediction value to the original numerical range to obtain the final oil and gas well production prediction value.

9. A server, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Gas well effusion classification and prediction method based on frequency channel conversion and self-supervision

    CN118378135A

  • Source load small sample time sequence prediction method, system and device based on pre-trained large language model, and storage medium

    CN118839730A

  • Electric submersible pump well yield prediction method fusing electric submersible pump mechanism and large language model

    CN119378767A

  • Power system time sequence prediction model training method, system and device based on large language model weight fine tuning and storage medium

    CN119670893A