A method and device for predicting the productivity of low-permeability gas wells

Through multi-model stacking architecture and dynamic weight optimization technology, combined with models such as random forests, support vector machines and lightweight neural networks, the problem of insufficient prediction accuracy of low-permeability gas well production capacity is solved, and higher prediction accuracy and robustness are achieved.

CN119849706BActive Publication Date: 2025-06-24YAZHOU BAY INNOVATION RESEARCH INSTITUTE HAINAN TROPICAL OCEAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510322294.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-24
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the production capacity of low-permeability gas wells, and traditional methods are difficult to capture the nonlinear relationship between multivariables, and the existing machine learning methods are relatively single, making it difficult to comprehensively utilize multidimensional data, resulting in limited prediction accuracy.

Method used

A multi-model stacking architecture is adopted, combined with models such as random forests, support vector machines and lightweight neural networks, and the output of the first layer submodel is used as the input of the second layer submodel, and the model performance is optimized using K-fold cross-validation, and the submodel weight is dynamically adjusted to improve prediction accuracy.

Benefits of technology

The accuracy and generalization ability of low-permeability gas well production capacity prediction can better adapt to the heterogeneity and stress-sensitive characteristics of low-permeability gas reservoirs, and enhance the robustness of the prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119849706B_ABST
    Figure CN119849706B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for predicting the productivity of low-permeability gas wells. The method includes obtaining low-permeability gas well data; preprocessing the obtained low-permeability gas well data to form a multi-source heterogeneous data set; constructing a multi-model stacking architecture; using the multi-source heterogeneous data set to train the multi-model stacking architecture to obtain a prediction model; and the prediction model outputs a prediction result. The present invention adopts a multi-model stacking technology. By combining the advantages of models such as random forest, support vector machine, and lightweight neural network, it makes up for the deficiencies of a single model. The introduction of the second-layer sub-model further optimizes the output of the sub-model, making the prediction result more accurate. At the same time, through the model stacking technology, by combining the non-linear expression ability of the deep learning model and the robustness of the traditional machine learning model, it can better adapt to the heterogeneity and stress sensitivity characteristics of low-permeability gas reservoirs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of oil and gas exploration and development, and particularly relates to a method and device for predicting the productivity of low-permeability gas wells. Background Art

[0002] Low-permeability gas reservoirs are widely distributed in global oil and gas resources. However, due to the strong reservoir heterogeneity, low permeability, and significant stress sensitivity of such reservoirs, predicting the productivity of low-permeability gas wells has become a major technical challenge during the development process. Traditional productivity prediction methods are difficult to capture the non-linear relationships between multiple variables and do not fully consider different types of input parameters, resulting in limited prediction accuracy. In recent years, machine learning methods have gradually been widely used. However, existing machine learning methods are relatively single, making it difficult to comprehensively utilize multi-dimensional data, and current fusion models are difficult to dynamically adjust weights, leading to poor model fusion effects and an inability to balance local characteristics and global performance.

[0003] Traditional machine learning productivity prediction methods usually use single models such as support vector machines, random forests, or neural networks. However, due to their limitations, it is difficult to comprehensively utilize the characteristics of multi-dimensional data, and the prediction accuracy is limited. Some current studies attempt to improve the accuracy through model fusion, but the fixed weighted average or voting methods they adopt lack the ability of adaptive adjustment and are difficult to handle the reservoir heterogeneity of low-permeability gas reservoirs. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned deficiencies of the prior art and provide a method and device for predicting the productivity of low-permeability gas wells, which can adapt to different types of data distributions and characteristics, enhance the robustness of the prediction model, and improve the prediction accuracy and generalization ability.

[0005] To achieve the above purpose, the technical solution of the present invention is as follows:

[0006] In a first aspect, the present invention provides a method for predicting the productivity of low-permeability gas wells, including:

[0007] Obtaining low-permeability gas well data;

[0008] Preprocessing the obtained low-permeability gas well data to form a multi-source heterogeneous data set;

[0009] Constructing a multi-model stacking architecture; the multi-model stacking architecture includes a first-layer sub-model and a second-layer sub-model; the first-layer sub-model includes three sub-models: a random forest, a support vector machine, and a lightweight neural network; the output of the first-layer sub-model is used as the input of the second-layer sub-model, and during the inter-layer optimization process, K-fold cross-validation is used to evaluate the performance of the first-layer sub-model, and its prediction results are used as the data source for training the second-layer sub-model;

[0010] Use the multi-source heterogeneous dataset to train the multi-model stacking architecture to obtain a prediction model;

[0011] The prediction model outputs a prediction result.

[0012] Optionally, the preprocessing of the obtained low-permeability gas well data includes:

[0013] Use the KNN method to fill in the missing data in the low-permeability gas well data; the calculation formula of the KNN method is:

[0014]

[0015] Where, are the known values of K neighbors;

[0016] Use the z-score method to detect and remove outliers in the low-permeability gas well data; the calculation formula of the z-score method is:

[0017]

[0018] Where, is the data point, is the mean of the dataset, is the standard deviation of the dataset;

[0019] Adopt correlation analysis and principal component analysis methods to screen important features in the low-permeability gas well data and remove redundant information;

[0020] Normalize or standardize data with different dimensions in the low-permeability gas well data; the formula for the standardization process is:

[0021]

[0022] Where, is the data after standardization; is the original data; is the minimum value in the original data; is the maximum value of the original data.

[0023] Optionally, the training of the multi-model stacking architecture using the multi-source heterogeneous dataset includes three levels: sub-model training, multi-model stacking training, and second-layer sub-model training.

[0024] Optionally, the multi-source heterogeneous dataset is divided into a training set and a test set. Each time, the K-fold cross-validation method is used to divide the training set into K equal parts. When training each time, K - 1 parts of the data are used for training, and the remaining one part is used as the validation set to complete the training of the multi-model stacking architecture.

[0025] Optionally, the multi-model stacked training includes: using the prediction results of the three sub-models of random forest, support vector machine, and lightweight neural network as the input features of the second-layer sub-model, and retaining the feature outputs of the three sub-models of random forest, support vector machine, and lightweight neural network during the training process.

[0026] Optionally, the training of the second-layer sub-model includes: using XGBoost as the second-layer sub-model, and optimizing the number of XGBoost trees, learning rate, and depth parameters through the prediction results of the three sub-models of random forest, support vector machine, and lightweight neural network input.

[0027] Optionally, use the genetic algorithm optimization method to dynamically adjust the weights of the three sub-models of random forest, support vector machine, and lightweight neural network in the second-layer sub-model.

[0028] Optionally, the use of the genetic algorithm optimization method to dynamically adjust the weights of the three sub-models of random forest, support vector machine, and lightweight neural network in the second-layer sub-model includes:

[0029] Based on the initialization population method, generate multiple random weight combinations, and each combination represents the weights of different sub-models in the second-layer sub-model;

[0030] Use the root mean square error and coefficient of determination as fitness functions to evaluate the quality of each weight combination, and the weight combinations that reach the threshold enter the next round;

[0031] Select the weight combinations that reach the threshold for crossover, combine the weight parts of different individuals to generate new weight combinations;

[0032] Introduce mutation by randomly adjusting the weight values of some individuals to maintain the diversity of the population;

[0033] After each crossover and mutation operation, the new weight combinations will undergo fitness evaluation. The weight combinations that meet the requirements are selected into the next generation and gradually optimized until an optimal or approximately optimal weight distribution is found.

[0034] In a second aspect, the present invention provides a low-permeability gas well productivity prediction device, including:

[0035] A data acquisition module for acquiring low-permeability gas well data;

[0036] A data preprocessing module for preprocessing the acquired low-permeability gas well data to form a multi-source heterogeneous data set;

[0037] A multi - model stacking architecture module for constructing a multi - model stacking architecture; the multi - model stacking architecture includes a first - layer sub - model and a second - layer sub - model; the first - layer sub - model includes three sub - models: random forest, support vector machine, and lightweight neural network; the output of the first - layer sub - model serves as the input of the second - layer sub - model, and during the inter - layer optimization process, K - fold cross - validation is used to evaluate the performance of the first - layer sub - model, and its prediction results are used as the data source for training the second - layer sub - model.

[0038] A model training module that uses the multi - source heterogeneous dataset to train the multi - model stacking architecture to obtain a prediction model;

[0039] A result output module that outputs prediction results based on the prediction model.

[0040] Optionally, the low - permeability gas well productivity prediction device further includes:

[0041] A dynamic weight optimization module for dynamically adjusting the weights of the three sub - models of random forest, support vector machine, and lightweight neural network in the second - layer sub - model using a genetic algorithm optimization method.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] The present invention adopts a multi - model stacking technology. By combining the advantages of models such as random forest, support vector machine, and lightweight neural network, it makes up for the deficiencies of a single model. The introduction of the second - layer sub - model further optimizes the output of the sub - model, making the prediction results more accurate. At the same time, through the model stacking technology, by combining the non - linear expression ability of deep learning models and the robustness of traditional machine learning models, it can better adapt to the heterogeneity and stress - sensitivity characteristics of low - permeability gas reservoirs. Description of the Drawings

[0044] Figure 1 It is a flowchart of the low - permeability gas well productivity prediction method provided by the embodiment of the present application;

[0045] Figure 2 It is a schematic diagram of the principle of the low - permeability gas well productivity prediction method provided by the embodiment of the present application;

[0046] Figure 3 It is a result comparison diagram of different single - machine learning productivity prediction models and the low - permeability gas well productivity prediction method based on multi - model stacking and weight self - adaptive optimization of the present application;

[0047] Figure 4 It is a schematic diagram of the composition of the low - permeability gas well productivity prediction device provided by the embodiment of the present application;

[0048] Figure 5 It is a schematic diagram of the composition of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0049] Examples:

[0050] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0051] For ease of understanding of the embodiments of the present application, further explanatory descriptions will be given below with reference to the accompanying drawings and specific embodiments, and the embodiments do not limit the embodiments of the present application.

[0052] Figure 1 is a flowchart of the low-permeability gas well productivity prediction method provided by the embodiments of the present application. As Figure 1 shown, the low-permeability gas well productivity prediction method provided by the embodiments of the present application mainly includes the following steps:

[0053] 110. Obtain low-permeability gas well data;

[0054] 120. Preprocess the obtained low-permeability gas well data to form a multi-source heterogeneous data set;

[0055] 130. Construct a multi-model stacking architecture;

[0056] Specifically, the multi-model stacking architecture includes a first-layer sub-model and a second-layer sub-model; the first-layer sub-model includes three sub-models: random forest, support vector machine, and lightweight neural network; Random Forest (RF): Utilize its powerful non-linear modeling ability to capture the relationships between various parameters and productivity. Support Vector Machine (SVR): Enhance the adaptability of the model to small-sample data sets and the ability to handle complex high-dimensional data; Lightweight Neural Network (LightNN): Enhance the ability to capture complex non-linear feature relationships through deep learning. In this way, by combining the advantages of models such as random forest, support vector machine, and lightweight neural network, the deficiencies of a single model are made up for.

[0057] The output of the first-layer sub-model is used as the input of the second-layer sub-model. And in the inter-layer optimization process, K-fold cross-validation is used to evaluate the performance of the first-layer sub-model, and its prediction results are used as the data source for training the second-layer sub-model. In this way, through the introduction of the second-layer sub-model, the output of the sub-model is further optimized, making the prediction results more accurate.

[0058] 140. Use the multi-source heterogeneous dataset to train the multi-model stacking architecture to obtain a prediction model;

[0059] 150. The prediction model outputs a prediction result.

[0060] Thus, this method adopts the multi-model stacking technology. By combining the advantages of models such as random forest, support vector machine, and lightweight neural network, it makes up for the deficiencies of a single model. The introduction of the second-layer sub-model further optimizes the output of the sub-model, making the prediction result more accurate. At the same time, through the model stacking technology, by combining the non-linear expression ability of the deep learning model and the robustness of the traditional machine learning model, it can better adapt to the heterogeneity and stress sensitivity characteristics of the low-permeability gas reservoir.

[0061] In specific implementation, step 120 above includes: extracting reservoir parameters and production dynamic data from the production data of low-permeability gas wells, and after preprocessing, forming a multi-source heterogeneous dataset in a unified format. The preprocessing includes:

[0062] Using the KNN method to fill in missing data to ensure data integrity, and its calculation formula is:

[0063]

[0064] where are the known values of K neighbors.

[0065] Using the z-score method to detect and remove outliers, and its calculation formula is:

[0066]

[0067] where is the data point, is the mean of the dataset, is the standard deviation of the dataset.

[0068] Taking the cumulative gas production of the gas well in the first year as the target value, and the rest as feature values. Using correlation analysis and principal component analysis (PCA) methods to screen important features, remove redundant information, and reduce the data dimension.

[0069] Pearson correlation coefficient method calculation formula:

[0070]

[0071] where is the correlation coefficient value; is the variable corresponding different numerical values; is the variable 's average value; is the variable corresponding different values; is a variable the average value of; is the number of variables.

[0072] Normalize or standardize data with different dimensions to enhance the stability of subsequent model training. The standardization formula is:

[0073]

[0074] where is the standardized data; is the original data; is the minimum value in the original data; is the maximum value of the original data.

[0075] In a specific implementation, using the multi-source heterogeneous dataset to train the multi-model stacking architecture includes three levels: sub-model training, multi-model stacking training, and second-layer sub-model training, so as to be adaptable to different types of data distributions and features and enhance the robustness of the obtained prediction model.

[0076] In a specific embodiment, the multi-source heterogeneous dataset is divided into a training set and a test set. Each time, the K-fold cross-validation method is used to divide the training set into K equal parts. When training each time, K - 1 parts of the data are used for training, and the remaining one part is used as the validation set to complete the training of the multi-model stacking architecture. In this way, by looping K times, it is ensured that each part of the data is used as the validation set, avoiding the influence of data distribution deviation on the model performance evaluation.

[0077] In a specific embodiment, the multi-model stacking training includes: using the prediction results of the three sub-models of random forest, support vector machine, and lightweight neural network as the input features of the second-layer sub-model, and retaining the feature outputs of the three sub-models during the training process. In this way, by "stacking" the prediction results of the three sub-models of random forest, support vector machine, and lightweight neural network, a new feature set is formed, and the prediction ability is enhanced by combining the outputs of different models.

[0078] In a specific embodiment, the second-layer sub-model training includes: using XGBoost as the second-layer sub-model, and optimizing hyperparameters such as the number of XGBoost trees, learning rate, and depth through the prediction results of the three sub-models of random forest, support vector machine, and lightweight neural network input. Through training, XGBoost can fully integrate the prediction advantages of the three sub-models and improve the overall performance.

[0079] In a preferred embodiment, after each training, the genetic algorithm optimization method is used to dynamically adjust the weights of the three sub-models in the second-layer sub-model. In this way, by adopting the strategy of the weight adaptive optimization method to dynamically adjust the model weights, the changes in complex reservoir heterogeneity can be adapted, and the accuracy and generalization ability of the production prediction of the prediction model can be improved. Specifically, it includes:

[0080] Based on the initialization population method, multiple random weight combinations are generated, and each combination represents the weights of different sub-models in the second-layer sub-model; in this embodiment, the weight combinations of the three sub-models of random forest, support vector machine, and lightweight neural network are [w1, w2, w3], where each w represents the weight of a sub-model.

[0081] The root mean square error RMSE and the coefficient of determination R² are used as fitness functions to evaluate the quality of each weight combination, and the weight combinations with fitness reaching the threshold enter the next round;

[0082] The formula for the root mean square error is:

[0083]

[0084] where, is the number of samples; is the actual value; is the predicted value.

[0085] The formula for the coefficient of determination is:

[0086]

[0087] where, is the number of samples; is the actual value; is the predicted value; is the average value.

[0088] Select the weight combinations with fitness reaching the threshold for crossover, combine the weight parts of different individuals to generate new weight combinations;

[0089] By randomly adjusting the weight values of some individuals, mutation is introduced to maintain the diversity of the population; this process is equivalent to "gene recombination" and can generate more diverse candidate solutions.

[0090] After each crossover and mutation operation, the new weight combinations will undergo fitness evaluation, and the weight combinations that meet the requirements are selected into the next generation and gradually optimized until an optimal or approximately optimal weight distribution is found.

[0091] Through the above steps, when new data is input or reservoir conditions change, the genetic algorithm can be restarted. Through this dynamic adjustment process, the model adaptively adjusts the weights according to the new data and the changing reservoir characteristics, maintaining high-precision prediction.

[0092] In this way, through the above method, the genetic algorithm can optimize the weights of each sub-model in the multi-model stack, adapt to the data changes caused by reservoir heterogeneity, enabling the prediction model to maintain good generalization ability under different geological conditions or production environments, thereby improving the accuracy of production capacity prediction.

[0093] As Figure 3 shown, it is a result comparison diagram of different single machine learning production capacity prediction models and the low-permeability gas well production capacity prediction method based on multi-model stack and weight adaptive optimization in this application. The results show that the low-permeability gas well production capacity prediction method based on multi-model stack and weight adaptive optimization proposed in this invention has significantly improved prediction accuracy compared with traditional single machine learning production capacity prediction models, providing a new method basis for the production capacity prediction work of low-permeability gas wells.

[0094] As Figure 4 shown, the embodiment of this application also provides a low-permeability gas well production capacity prediction device 400, including:

[0095] A data acquisition module 410, configured to acquire low-permeability gas well data;

[0096] A data preprocessing module 420, configured to preprocess the acquired low-permeability gas well data to form a multi-source heterogeneous data set;

[0097] A multi-model stack architecture module 430, configured to construct a multi-model stack architecture; the multi-model stack architecture includes a first-layer sub-model and a second-layer sub-model; the first-layer sub-model includes three sub-models: random forest, support vector machine, and lightweight neural network; the output of the first-layer sub-model is used as the input of the second-layer sub-model, and during the inter-layer optimization process, the performance of the first-layer sub-model is evaluated using K-fold cross-validation, and its prediction results are used as the data source for training the second-layer sub-model.

[0098] A model training module 440, which uses the multi-source heterogeneous data set to train the multi-model stack architecture to obtain a prediction model;

[0099] A result output module 450, which outputs a prediction result based on the prediction model.

[0100] As a preference for the low-permeability gas well production capacity prediction device, this device further includes:

[0101] A dynamic weight optimization module, which is used to dynamically adjust the weights of three sub-models, namely random forest, support vector machine, and lightweight neural network, in the second-layer sub-model by using the genetic algorithm optimization method, including:

[0102] Based on the initialization population method, generate multiple random weight combinations, and each combination represents the weights of different sub-models in the second-layer sub-model;

[0103] Use the root mean square error and coefficient of determination as fitness functions to evaluate the quality of each weight combination, and the weight combinations that reach the threshold of fitness enter the next round;

[0104] Select the weight combinations that reach the threshold of fitness for crossover, combine the weight parts of different individuals to generate new weight combinations;

[0105] Introduce mutation by randomly adjusting the weight values of some individuals to maintain the diversity of the population;

[0106] After each crossover and mutation operation, the new weight combinations will undergo fitness evaluation. The weight combinations that meet the requirements are selected into the next generation and gradually optimized until an optimal or approximately optimal weight distribution is found.

[0107] Optionally, the preprocessing of the obtained low-permeability gas well data includes:

[0108] Use the KNN method to fill in the missing data in the low-permeability gas well data; the calculation formula of the KNN method is:

[0109]

[0110] Among them, are the known values of K neighbors.

[0111] Use the z-score method to detect and remove outliers in the low-permeability gas well data; the calculation formula of the z-score method is:

[0112]

[0113] Among them, is the data point, is the mean of the data set, is the standard deviation of the data set.

[0114] Adopt correlation analysis and principal component analysis methods to screen important features in the low-permeability gas well data and remove redundant information;

[0115] Normalize or standardize the data with different dimensions in the low-permeability gas well data; the formula for the standardization process is:

[0116]

[0117] Among them, is the standardized data; is the original data; is the minimum value in the original data; is the maximum value of the original data.

[0118] Optionally, the training of the multi-model stacking architecture using the multi-source heterogeneous dataset includes three levels: sub-model training, multi-model stacking training, and second-layer sub-model training.

[0119] Optionally, the multi-source heterogeneous dataset is divided into a training set and a test set. Each time, the K-fold cross-validation method is used to divide the training set into K equal parts. During each training, K - 1 parts of the data are used for training, and the remaining one part is used as the validation set to complete the training of the multi-model stacking architecture.

[0120] Optionally, the multi-model stacking training includes: using the prediction results of the three sub-models of random forest, support vector machine, and lightweight neural network as the input features of the second-layer sub-model, and retaining the feature outputs of the three sub-models during the training process.

[0121] Optionally, the second-layer sub-model training includes: using XGBoost as the second-layer sub-model and optimizing the number, learning rate, and depth parameters of XGBoost through the prediction results of the three sub-models of random forest, support vector machine, and lightweight neural network input.

[0122] It should be noted that the low-permeability gas well productivity prediction device provided in the embodiments of the present application can execute the low-permeability gas well productivity prediction method provided in any embodiment of the present application, and has the corresponding functions and beneficial effects of executing the method.

[0123] As Figure 5 shown, the embodiments of the present application provide an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114; the memory 113 is used to store a computer program; when the processor 111 executes the program stored on the memory 113, it implements the steps of the low-permeability gas well productivity prediction method provided in any one of the foregoing method embodiments.

[0124] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the low-permeability gas well productivity prediction method provided in any one of the foregoing method embodiments.

[0125] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0126] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for predicting the productivity of a low permeability gas well, characterized in that: include: Obtain low permeability gas well data; Preprocess the acquired low-permeability gas well data to form a multi-source heterogeneous data set; Build a multi-model stacking architecture; The multi-model stacking architecture includes a first-layer sub-model and a second-layer sub-model; the first-layer sub-model includes three sub-models: random forest, support vector machine and lightweight neural network; the output of the first-layer sub-model is used as the input of the second-layer sub-model, and in the inter-layer optimization process, the performance of the first-layer sub-model is evaluated by K-fold cross validation, and its prediction result is used as the data source for training the second-layer sub-model; The genetic algorithm optimization method is used to dynamically adjust the weights of the three sub-models of random forest, support vector machine and lightweight neural network in the second-layer sub-model, including: based on the initialization population method, multiple random weight combinations are generated, each combination represents the weight of different sub-models in the second-layer sub-model; the root mean square error and determination coefficient are used as fitness functions to evaluate the quality of each weight combination, and the weight combination with the fitness reaching the threshold enters the next round; the weight combination with the fitness reaching the threshold is selected for crossover, and the weight parts of different individuals are combined to generate a new weight combination; by randomly adjusting the weight values ​​of some individuals, mutations are introduced to maintain the diversity of the population; after each crossover and mutation operation, the new weight combination will be evaluated for fitness, and the weight combination that meets the requirements will be selected into the next generation and gradually optimized to finally find an optimal or approximately optimal weight score; The multi-source heterogeneous data set is used to train the multi-model stacking architecture to obtain a prediction model; the multi-source heterogeneous data set is divided into a training set and a test set, and the K-fold cross-validation method is used each time to divide the training set into K equal parts, and K-1 parts of data are used for training in each training, and the remaining part is used as a validation set to complete the training of the multi-model stacking architecture; the training of the multi-model stacking architecture includes: using the prediction results of the three sub-models of random forest, support vector machine and lightweight neural network as the input features of the second-layer sub-model, and retaining the feature outputs of the three sub-models of random forest, support vector machine and lightweight neural network during the training process; the second-layer sub-model training includes: using XGBoost as the second-layer sub-model, and optimizing the number of XGBoost trees, learning rate and depth parameters through the prediction results of the three sub-models of random forest, support vector machine and lightweight neural network input; The prediction model outputs a prediction result.

2. The method for predicting the productivity of a low permeability gas well according to claim 1, characterized in that: The preprocessing of the acquired low permeability gas well data includes: The KNN method is used to fill the missing data in the low permeability gas well data; the calculation formula of the KNN method is: in, are the known values ​​of the K neighbors; The z-score method is used to detect and remove outliers in the low permeability gas well data; the calculation formula of the z-score method is: in, is a data point, is the mean of the data set, is the standard deviation of the data set; Correlation analysis and principal component analysis were used to screen important features in low-permeability gas well data and remove redundant information. The data of different dimensions in the low permeability gas well data are normalized or standardized; the formula for the standardization is: in, is the standardized data; is the original data; is the minimum value in the original data; is the maximum value of the original data.

3. A low permeability gas well productivity prediction device, characterized in that: include: A data acquisition module, used to acquire low permeability gas well data; A data preprocessing module is used to preprocess the acquired low-permeability gas well data to form a multi-source heterogeneous data set; Multi-model stacking architecture module, used to build multi-model stacking architecture; The multi-model stacking architecture includes a first-layer sub-model and a second-layer sub-model; the first-layer sub-model includes three sub-models: random forest, support vector machine and lightweight neural network; the output of the first-layer sub-model is used as the input of the second-layer sub-model, and in the inter-layer optimization process, the output of the first-layer sub-model is used as the input of the second-layer sub-model. The performance of the first-layer sub-model is evaluated by cross-validation, and its prediction results are used as the data source for training the second-layer sub-model; The model training module uses the multi-source heterogeneous data set to train the multi-model stacking architecture to obtain a prediction model; the multi-source heterogeneous data set is divided into a training set and a test set, and each time the multi-source heterogeneous data set is used The fold cross validation method divides the training set into Use equal parts in each training session The data is used for training, and the remaining data is used as a validation set to complete the training of the multi-model stacking architecture; The training of the multi-model stacking architecture includes: using the prediction results of the three sub-models of random forest, support vector machine and lightweight neural network as input features of the second-layer sub-model, and retaining the feature outputs of the three sub-models of random forest, support vector machine and lightweight neural network during the training process; the second-layer sub-model training includes: using XGBoost as the second-layer sub-model, and optimizing the number of XGBoost trees, learning rate and depth parameters through the prediction results of the three sub-models of random forest, support vector machine and lightweight neural network input; A result output module, outputting prediction results based on the prediction model; The dynamic weight optimization module is used to dynamically adjust the weights of the three sub-models of random forest, support vector machine and lightweight neural network in the second-layer sub-model by using the genetic algorithm optimization method, including: based on the initialization population method, generating multiple random weight combinations, each combination represents the weight of different sub-models in the second-layer sub-model; using the root mean square error and determination coefficient as the fitness function to evaluate the quality of each weight combination, and the weight combination with the fitness reaching the threshold enters the next round; selecting the weight combination with the fitness reaching the threshold for crossover, combining the weight parts of different individuals to generate a new weight combination; introducing mutations by randomly adjusting the weight values ​​of certain individuals to maintain the diversity of the population; after each crossover and mutation operation, the new weight combination will undergo fitness evaluation, and the weight combination that meets the requirements will be selected into the next generation and gradually optimized to eventually find an optimal or approximately optimal weight score.

Citation Information

Patent Citations

  • Reasoning acceleration method of large language model based on hybrid neural network structure

    CN117787410A

  • Listed company risk assessment method based on cross validation and Bayesian optimization

    CN119624097A