Petroleum exploration method and system based on artificial intelligence
By adopting artificial intelligence-based methods in oil exploration, including the combination of improved PCA algorithms and DNN models, the problem of insufficient feature extraction and prediction accuracy in the prior art is solved, and more accurate oil capacity prediction and dynamic capacity threshold adjustment are achieved, and exploration efficiency and decision-making accuracy are improved.
Patent Information
- Application Number
- CN202510309998.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The existing petroleum exploration methods and systems have shortcomings in feature extraction, prediction accuracy, dynamic change considerations and capacity threshold adjustment, resulting in prediction deviations and inaccuracies.
Using an artificial intelligence-based method, we collect seismic data, logging data, exploration environment data and international crude oil price data, and feature extraction is performed using improved PCA algorithm, combining DNN models to predict oil thickness, porosity and permeability data, and using regular exploration environment regression model to predict oil production capacity, dynamically adjust the capacity threshold.
It improves the prediction accuracy and efficiency of oil exploration, can more accurately consider the dynamic changes in the exploration environment, dynamically adjust the capacity threshold, reduce the risk of blind exploitation, and avoid resource waste.
Smart Images

Figure CN120197148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and specifically to an oil exploration method and system based on artificial intelligence. Background Art
[0002] Existing oil exploration methods and systems have many problems. First of all, existing feature extraction methods cannot effectively extract deep information hidden in data, especially when facing complex non-linear relationships, which affects the prediction accuracy of the model and the reliability of the results. Secondly, when predicting oil production capacity, it usually relies on simplified physical models or static analysis based on historical data, ignoring the actual impact of dynamic changes in the exploration environment on oil production capacity. Changes in geological and environmental conditions will have an important impact on the permeability, production, etc. of oil reservoirs, but existing models often fail to consider these changing factors, resulting in deviations and inaccuracies in production capacity prediction. In addition, existing oil production capacity threshold calculation methods often rely on fixed historical data and empirical formulas, lacking a flexible dynamic adjustment mechanism. Once there are major changes in the market or environment, such as a sharp drop or rise in oil prices, existing methods cannot adjust the production capacity threshold in a timely manner, which may affect oil extraction decisions.
[0003] In view of this, the present invention proposes an oil exploration method and system based on artificial intelligence to solve the above problems. Summary of the Invention
[0004] To overcome the above defects of the prior art and to achieve the above object, the present invention provides the following technical solution. An oil exploration method based on artificial intelligence includes: Step S1: Collect seismic data, logging data, exploration environment data, and international crude oil price data; Step S2: Perform preliminary processing on the seismic data and logging data to obtain a preliminary data set; based on the preliminary data set, use an improved PCA algorithm for feature extraction to obtain a new feature data set; Step S3: Based on the new feature data set, use a DNN model to predict oil thickness data, porosity data, and permeability data; Step S4: Based on the predicted oil thickness data, porosity data, and permeability data, use a regularized exploration environment regression model to predict oil production capacity data, set a production capacity threshold, compare the predicted oil production capacity data with the production capacity threshold, and determine whether to extract oil.
[0005] Further, the seismic data includes amplitude feature data, wave velocity feature data, radiation feature data, seismic wave frequency feature data, and magnetic force feature data; Well logging data includes resistivity characteristic data, gamma ray characteristic data, acoustic travel time characteristic data, rock density characteristic data, neutron porosity characteristic data, conductivity characteristic data, and natural gamma characteristic data; Exploration environment data includes temperature characteristic data, humidity characteristic data, and air pressure characteristic data.
[0006] Furthermore, the acquisition method of the preliminary data set includes: For seismic data and well logging data, the mean filling method is used for missing value processing, the interpolation method is used for outlier processing, Z-Score is used for standardization processing, and unified timestamp processing is performed to obtain the preliminary data set.
[0007] Furthermore, the specific method of using the improved PCA algorithm for feature extraction based on the preliminary data set to obtain the new feature data set includes: Step A11: For the preliminary data set, use The method to form the preprocessing matrix; Step A12: Construct the covariance matrix for the preprocessing matrix, , where is the covariance matrix, is the transpose of the preprocessing matrix, is the preprocessing matrix, with dimension , n is the total number of samples in the preprocessing matrix, is the total number of features in the preprocessing matrix. One row in the preprocessing matrix represents a sample, The dimension of ; Step A13: Use the singular value decomposition method SVD to decompose the covariance matrix into three matrices. The formula is , is The orthogonal matrix of, containing the left singular vectors; is The diagonal matrix of, where the elements on the diagonal are the singular values; is The orthogonal matrix of, containing the right singular vectors; Step A14: Calculate the variance contribution degree of the principal components through the singular values. The formula is: , where is the th singular value of the principal component, that is, the variance of the th principal component. i is the principal component index, represents the sum of all principal component singular values, is the th variance contribution degree of the principal component; Step A15: Calculate the average variance contribution degree of all principal components. The formula is: , is the average contribution degree of all principal components. Select the k best principal components with contribution degrees greater than the average contribution degree. Step A16: Construct the best principal component matrix from the k best principal components . Project the preprocessing matrix onto the new feature matrix. The calculation formula is: , where is the new feature matrix after projection. Use the pandas.DataFrame method to convert the new feature matrix into a new feature dataset.
[0008] Furthermore, the specific method of using the DNN model to predict the oil thickness data, porosity data, and permeability data based on the obtained new feature dataset includes: Step D11: Input the sample set. The sample set includes groups of samples. Each group of samples includes a new feature dataset and the corresponding oil thickness data, porosity data, and permeability data. Step D12: Set the number of neurons VB in the input layer to be equal to the number of features in the new feature dataset. Use grid search to set the number of hidden layers to OP, and set the number of neurons in each hidden layer to 2*VB. Set the number of neurons in the output layer to 3, which corresponds to the predicted oil thickness data, porosity data, and permeability data. Initialize the model weights and biases, and set the number of iterations. Step D13: In the PR-th iteration, pass the samples from the input layer to the first hidden layer, and use the mixed activation function to calculate the output of the first hidden layer. Pass the output of the first hidden layer to the next hidden layer as the input of the next hidden layer, and repeatedly use the mixed activation function to calculate the output of the hidden layer until the last hidden layer. Pass the output of the last hidden layer to the output layer to obtain the output of the PR-th iteration. The output includes the predicted oil thickness data, predicted porosity data, and output permeability data. Step D14: For the predicted oil thickness data, predicted porosity data, and output permeability data of the PR-th iteration output, use the mean square error as the loss function to calculate the error between the predicted value of the PR-th iteration and the true value, and perform a weighted sum of the errors of the predicted oil thickness data, predicted porosity data, and output permeability data to obtain the total loss function error. Here, the true value refers to the corresponding oil thickness data, porosity data, and permeability data in Step D11. Step D15: Calculate the gradients of the weights and biases for each layer using the chain rule based on the total loss function error, and update the weights and biases for each layer using the Adam optimizer; Step D16: When the set number of iterations is reached, stop the iteration; output the predicted oil thickness data, porosity data, and permeability data.
[0009] Furthermore, the specific manner of calculating the output of the first hidden layer using the hybrid activation function includes: In the DNN model, the activation functions include activation function, activation function, activation function, and activation function; Use the weighted average method to calculate the hybrid activation function. The formula is: , where , , and are the weight coefficients of activation function, activation function, activation function, and activation function respectively, and .
[0010] Furthermore, the specific manner of predicting the oil production capacity data using the regularized exploration environment regression model based on the predicted oil thickness data, porosity data, and permeability data includes: Step ff1: Input groups of oil samples. Each group of oil samples includes a predicted oil thickness feature data, porosity feature data, and permeability feature data, as well as the corresponding oil production capacity data; Step ff2: Set the number of iterations; initialize the intercept , initialize the regression coefficients , and , where , and are the regression coefficients of the input predicted oil thickness data, porosity data, and permeability data respectively; Step ff3: Calculate the predicted value of the oil production capacity data based on the initialized regression coefficients and intercept. The formula is: , where , and Denote the predicted oil thickness data, porosity data, and permeability data of the s-th group of oil samples, where s is the oil sample index. Denote the predicted oil production capacity data of the s-th group of oil samples; Step ff4: Use regularization and exploration environment data to correct the loss function; Step ff5: Calculate the gradients of the loss function with respect to the regression coefficients and intercept through backpropagation; then use the gradient descent method to update the regression coefficients and intercept; Step ff6: Repeat Steps ff3 to ff5 until the set number of iterations is reached and stop to obtain the final regression coefficients and intercept.
[0011] Furthermore, the specific method of using regularization and exploration environment data to correct the loss function includes: Use the mean squared error to calculate the loss function, and the formula is: , where is the oil production capacity data in the s-th group of oil samples; Use regularization and exploration environment data to correct the loss function to obtain the corrected loss function, and the formula is: , where is the regularization parameter, is the L2 norm of the regression coefficients, is the index of the regression coefficients, is the -th regression coefficient, is the correction coefficient of the exploration environment data at the same moment of the s-th oil sample, is the adjustment factor; Among them, the correction coefficient of the exploration environment data is calculated by the weighted average method, , where , and are the weights of the temperature data, humidity data, and air pressure data in the exploration environment data respectively, , and are the temperature characteristic data value, humidity characteristic data value, and air pressure characteristic data obtained at the same moment as the s-th group of oil samples respectively.
[0012] Furthermore, the specific method of setting the production capacity threshold and comparing the predicted oil production capacity data with the production capacity threshold to determine whether to extract oil includes: Calculate the average value and standard deviation based on historical oil production capacity data, and set the initial production capacity threshold according to the average value and standard deviation of historical oil production capacity data; Calculate the price average based on historical international crude oil price data, and correct the initial production capacity threshold through the ratio of the historical international crude oil average price data to the current international crude oil price data to obtain the production capacity threshold; When the oil production capacity data is greater than or equal to the production capacity threshold data, it means that the explored oil can be exploited; when the oil production capacity data is less than the production capacity threshold data, it means that the explored oil cannot be exploited; The formula for the average value of historical oil production capacity data is: , where is the average value of historical oil production capacity, is the total number of historical oil production capacity data, is the index of historical oil production capacity data, is the th historical oil production capacity data; The formula for the standard deviation of historical oil production capacity data is: , is the standard deviation of historical oil production capacity; The formula for the initial production capacity threshold is: , is the initial production capacity threshold, is the adjustment coefficient; The formula for the price average of historical international crude oil price data is: , where is the index of historical international crude oil price, is the th price of historical international crude oil, is the average price of historical international crude oil; The formula for the production capacity threshold is: , where is the current international crude oil price data.
[0013] An oil exploration system based on artificial intelligence, applied to the described oil exploration method based on artificial intelligence, includes: Data acquisition module: Collect seismic data, logging data, exploration environment data, and international crude oil price data; Data processing module: Conduct preliminary processing on seismic data and logging data to obtain a preliminary data set; Based on the preliminary data set, use an improved PCA algorithm for feature extraction to obtain a new feature data set; Data prediction module: For the new feature data set, use a DNN model to predict the oil thickness data, porosity data, and permeability data; Oil exploration module: Based on the predicted oil thickness data, porosity data, and permeability data, use a regularized exploration environment regression model to predict oil production capacity data, set a production capacity threshold, compare the predicted oil production capacity data with the production capacity threshold, and determine whether to extract oil.
[0014] Technical effects and advantages of an oil exploration method and system based on artificial intelligence according to the present invention: Through in-depth analysis of data and model prediction, the present invention realizes the intelligence and high efficiency of oil exploration; First, by collecting seismic data, logging data, exploration environment data, and international crude oil price data, the system can comprehensively obtain multi-dimensional information affecting oil production capacity, improving the accuracy of prediction; Then, through the improved PCA algorithm for feature extraction, it effectively reduces data redundancy, improves the accuracy of feature data, reduces the amount of calculation, while retaining the core information of the data and enhancing the learning ability of the model; Next, based on the deep neural network DNN model, it accurately predicts the oil thickness data, porosity data, and permeability data, providing a basis for subsequent oil production capacity prediction; Finally, combined with the regularized exploration environment regression model, the system can comprehensively consider the influence of the exploration environment and perform accurate oil production capacity prediction; Using historical oil production capacity data and international oil price data to correct the production capacity threshold makes the production capacity prediction more dynamic and flexible, can respond to market changes in real time, and reduces the risk of blind extraction by accurately predicting production capacity and intelligently judging whether to extract, avoiding resource waste caused by over-extraction or inefficient extraction. Description of the Drawings
[0015] Figure 1 Schematic diagram of an oil exploration method based on artificial intelligence according to the present invention; Figure 2 Schematic diagram of an oil exploration system based on artificial intelligence according to the present invention. Detailed Embodiments
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] Embodiment 1
[0018] Please refer to Figure 1 As shown, an oil exploration method based on artificial intelligence in this embodiment includes: Step S1: Collect seismic data, logging data, exploration environment data, and international crude oil price data; Step S2: Conduct preliminary processing based on the seismic data and logging data to obtain a preliminary data set; based on the preliminary data set, use an improved PCA algorithm for feature extraction to obtain a new feature data set; Step S3: Based on the new feature data set, use a DNN model to predict oil thickness data, porosity data, and permeability data; Step S4: Based on the predicted oil thickness data, porosity data, and permeability data, use a regularized exploration environment regression model to predict oil production capacity data, set a production capacity threshold, compare the predicted oil production capacity data with the production capacity threshold, and determine whether to extract oil.
[0019] Seismic data includes amplitude feature data, wave velocity feature data, radiation feature data, seismic wave frequency feature data, and magnetic force feature data; Logging data includes resistivity feature data, gamma ray feature data, acoustic time difference feature data, rock density feature data, neutron porosity feature data, conductivity feature data, and natural gamma feature data; Exploration environment data includes temperature feature data, humidity feature data, and air pressure feature data.
[0020] Install sensors in the drilling hole to obtain data, specifically including: obtain amplitude feature data, wave velocity feature data, radiation feature data, and seismic wave frequency feature data through geophones; obtain magnetic force feature data through magnetometers; Measure the resistivity feature data of the formation through a resistivity logging tool; measure the gamma ray feature data and natural gamma feature data in the formation through a gamma ray logging tool; measure the acoustic time difference feature data in the formation through an acoustic logging tool; measure the rock density feature data through a density logging tool; measure the neutron porosity feature data of the formation through a neutron logging tool; measure the conductivity feature data through a conductivity logging tool; Obtain temperature feature data, humidity feature data, and air pressure feature data by installing temperature sensors, humidity sensors, and air pressure sensors; Obtain international crude oil price data through financial websites or commodity trading platforms.
[0021] The acquisition method of the preliminary data set includes: For seismic data and logging data, use the mean filling method for missing value processing, use the interpolation method for outlier processing, use Z-Score for standardization processing, and perform unified timestamp processing to obtain a preliminary data set.
[0022] Based on the preliminary data set, the specific ways to extract features using the improved PCA algorithm to obtain a new feature data set include: Step A11: For the preliminary data set, use method to form a preprocessing matrix; Step A12: Construct a covariance matrix for the preprocessing matrix, , where, is the covariance matrix, is the transpose of the preprocessing matrix, is the preprocessing matrix, with dimension , n is the total number of samples in the preprocessing matrix, is the total number of features in the preprocessing matrix. One row in the preprocessing matrix represents a sample, has a dimension of ; Step A13: Use the singular value decomposition method SVD to decompose the covariance matrix into three matrices. The formula is , is orthogonal matrix, containing left singular vectors; is diagonal matrix, where the elements on the diagonal are singular values; is orthogonal matrix, containing right singular vectors; Step A14: Calculate the variance contribution degree of the principal components through singular values. The formula is: , where, is the th singular value of the principal component, that is, the variance of the th principal component. i is the principal component index, represents the sum of all principal component singular values, is the th variance contribution degree of the principal component; Step A15: Calculate the average variance contribution degree of all principal components. The formula is: , is the average contribution degree of all principal components. Select k best principal components with contribution degrees greater than the average contribution degree; Step A16: Construct the k best principal components into the best principal component matrix , and project the preprocessing matrix onto the new feature matrix. The calculation formula is: , where, is the new feature matrix after projection. Use the pandas.DataFrame method to convert the new feature matrix into a new feature data set; Using an improved PCA model for feature extraction can effectively reduce the data dimension, retain the most important features, thereby reducing noise and redundant information, and improving the accuracy of subsequent analysis; The PCA model usually relies on calculating the eigenvalues and eigenvectors of the covariance matrix, while the improved PCA model directly decomposes the data matrix through the singular value decomposition method SVD, avoiding the process of solving the covariance matrix and reducing the computational complexity, especially when dealing with large-scale data sets; SVD can efficiently capture the principal components of the data and shows better performance in terms of numerical stability and computational speed; In addition, using the pandas.DataFrame method in the Pandas library to convert the new feature matrix into a data set facilitates subsequent data processing and analysis, making the entire process more efficient and practical.
[0023] Based on the obtained new feature data set, the specific ways to use the DNN model to predict oil thickness data, porosity data, and permeability data include: Step D11: Input the sample set, and the sample set includes groups of samples, and each group of samples includes a new feature data set and the corresponding oil thickness data, porosity data, and permeability data; Step D12: Set the number of neurons VB in the input layer to be equal to the number of features in the new feature data set; Use grid search to set the number of hidden layers to OP, and set 2*VB neurons in each hidden layer; Set the number of neurons in the output layer to 3, that is, corresponding to the predicted oil thickness data, porosity data, and permeability data; Initialize the model weights and biases, and set the number of iterations; Step D13: In the PR-th iteration, transfer the samples from the input layer to the first hidden layer, and use the hybrid activation function to calculate the output of the first hidden layer. The formula is: , where, is the output of the first hidden layer, is the weight of the first hidden layer, is the input sample, is the bias of the first hidden layer, is the hybrid activation function; Transfer the output of the first hidden layer to the next hidden layer as the input of the next hidden layer, and repeatedly use the hybrid activation function to calculate the output of the hidden layer until the last hidden layer ends. The formula is: , where, is the output of the th layer, is the weight of the th layer, is the output of the layer, is the bias of the layer; Pass the output of the last hidden layer to the output layer to obtain the output of the PR-th iteration. The calculation formula is , where is the output of the output layer, and the output includes the predicted oil thickness data, the predicted porosity data, and the output permeability data, is the weight of the output layer, is the bias of the output layer; Step D14: For the predicted oil thickness data, the predicted porosity data, and the output permeability data of the PR-th iteration of the output, use the mean squared error as the loss function to calculate the error between the predicted value of the PR-th iteration and the true value, and perform a weighted sum of the errors of the predicted oil thickness data, the predicted porosity data, and the output permeability data to obtain the total loss function error, where the true value refers to the corresponding oil thickness data, porosity data, and permeability data in Step D11; Step D15: Through the total loss function error, use the chain rule to calculate the gradients of the weights and biases of each layer, and use the Adam optimizer to update the weights and biases of each layer; Step D16: When the set number of iterations is reached, stop the iteration; output the predicted oil thickness data, porosity data, and permeability data; The DNN model can effectively capture complex non-linear relationships, perform well in processing large-scale data, have good generalization ability, and can adapt to different data types and feature distributions. These characteristics enable DNN to provide more accurate and comprehensive prediction results in oil exploration, thereby improving exploration efficiency and the accuracy of decision-making.
[0024] The specific way to calculate the output of the first hidden layer using the mixed activation function includes: In the DNN model, the activation functions include activation function, activation function, activation function, and activation function; Use the weighted average method to calculate the mixed activation function. The formula is: , where , , and are respectively activation function, activation function, Activation function and the weight coefficients of the activation function, and , 、 、 and are each set to 0.25; In the DNN model, a hybrid activation function is used to calculate the output of the hidden layer. By combining the characteristics of multiple activation functions, complex patterns in the input data can be captured more effectively. The hybrid activation function integrates the advantages of different activation functions using the weighted average method, thereby improving the non-linear expression ability of the model. This approach not only enhances the model's adaptability to different features but also improves gradient propagation and reduces the problem of gradient vanishing, ultimately enhancing the model's learning efficiency and prediction accuracy.
[0025] Based on the predicted oil thickness data, porosity data, and permeability data, the specific ways to use the regularized exploration environment regression model to predict oil production data include: Step ff1: Input groups of oil samples, each group of oil samples includes a predicted oil thickness feature data, porosity feature data, and permeability feature data, as well as the corresponding oil production data; Step ff2: Set the number of iterations; initialize the intercept , initialize the regression coefficients 、 and , where 、 and are the regression coefficients of the input predicted oil thickness data, porosity data, and permeability data respectively; Step ff3: Calculate the predicted value of the oil production data based on the initialized regression coefficients and intercept. The formula is: , where 、 and represent the predicted oil thickness data, porosity data, and permeability data of the s-th group of oil samples, s is the oil sample index, represents the predicted oil production data of the s-th group of oil samples; Step ff4: Use regularization and exploration environment data to correct the loss function; Step ff5: Calculate the gradients of the loss function with respect to the regression coefficients and intercept through backpropagation; then use the gradient descent method to update the regression coefficients and intercept; Step ff6: Repeat steps ff3 to ff5 until the set number of iterations is reached and stop to obtain the final regression coefficients and intercept; Predicting oil production data using predicted oil thickness data, porosity data, and permeability data through a regularized exploration environment regression model has higher accuracy and reliability compared to directly predicting oil production data through a new feature dataset. First, it can capture the complex relationship between geological features and production more deeply, and reduce overfitting through regularization, improving the generalization ability of the model. In addition, using the intermediate prediction results as features can better integrate and utilize information from different data sources, making the final production prediction more scientific, not only improving the quality of the prediction, but also providing stronger support for decision-making, and helping to optimize the oil resource development strategy.
[0026] The specific ways to use regularization and exploration environment data to correct the loss function include: Use the mean squared error to calculate the loss function, and the formula is: , where is the oil production data in the s-th group of oil samples; Use regularization and exploration environment data to correct the loss function to obtain the corrected loss function, and the formula is: , where is the regularization parameter, obtained through the empirical rule, and is used to control the weight of the regularization term, is the L2 norm of the regression coefficient, is the index of the regression coefficient, is the -th regression coefficient, is the correction coefficient of the exploration environment data at the same moment of the s-th oil sample, is the adjustment factor, set by the random method; Among them, the correction coefficient of the exploration environment data is calculated by the weighted average method, , where , and are the weights of the temperature data, humidity data, and air pressure data in the exploration environment data, respectively set to , , and are the temperature feature data value, humidity feature data value, and air pressure feature data obtained at the same moment as the s-th group of oil samples, respectively.
[0027] By using the method of regularizing and correcting the loss function with exploration environment data, the robustness and generalization ability of the model are effectively improved; regularization controls the complexity of the regression coefficients to prevent overfitting, and at the same time, combined with the correction coefficient of the exploration environment data, the model can better reflect the impact of the exploration environment on oil production capacity; by integrating temperature, humidity and air pressure exploration environment data through the weighted average method, more comprehensive context information is provided, thus improving the accuracy and reliability of the prediction.
[0028] The specific methods for setting the production capacity threshold and comparing the predicted oil production capacity data with the production capacity threshold to determine whether to extract oil include: Calculating the average value and standard deviation based on historical oil production capacity data, and setting the initial production capacity threshold according to the average value and standard deviation of historical oil production capacity data; Calculating the average price based on historical international crude oil price data, and correcting the initial production capacity threshold through the ratio of historical international crude oil average price data and current international crude oil price data to obtain the production capacity threshold; When the oil production capacity data is greater than or equal to the production capacity threshold data, it means that the explored oil can be extracted; when the oil production capacity data is less than the production capacity threshold data, it means that the explored oil cannot be extracted; The formula for the average value of historical oil production capacity data is: , where, is the average value of historical oil production capacity, is the total number of historical oil production capacity data, is the index of historical oil production capacity data, is the th historical oil production capacity data; The formula for the standard deviation of historical oil production capacity data is: , is the standard deviation of historical oil production capacity; The formula for the initial production capacity threshold is: , is the initial production capacity threshold, is the adjustment coefficient; The formula for the average price of historical international crude oil price data is: , where, is the index of historical international crude oil price, is the th historical international crude oil price, is the average price of historical international crude oil; The formula for the production capacity threshold is: , where, is the current international crude oil price data; Obtaining the production threshold through historical oil production capacity data and international crude oil price data can make decision-making more scientific and data-driven. First, calculating the mean and standard deviation using historical data provides a reasonable basis for setting the production threshold, which helps identify the normal production capacity range. Second, combining the proportion of international crude oil prices to correct the initial threshold makes the threshold more flexible, enabling it to promptly reflect market fluctuations and changes in the economic environment, improving the accuracy of judgments on resource exploitation, optimizing resource allocation, reducing exploitation risks, and thus more effectively supporting oil exploration and development decisions.
[0029] In this embodiment, through in-depth data analysis and model prediction, the intelligence and efficiency of oil exploration are realized. First, by collecting seismic data, logging data, exploration environment data, and international crude oil price data, the system can comprehensively obtain multi-dimensional information affecting oil production capacity, improving the accuracy of prediction. Then, feature extraction is carried out through an improved PCA algorithm, effectively reducing data redundancy, enhancing the accuracy of feature data, reducing the computational amount, while retaining the core information of the data and enhancing the learning ability of the model. Next, based on the deep neural network DNN model, the oil thickness data, porosity data, and permeability data are accurately predicted, providing a basis for subsequent oil production capacity prediction. Finally, combined with the regularized exploration environment regression model, the system can comprehensively consider the impact of the exploration environment and conduct accurate oil production capacity prediction. Using historical oil production capacity data and international oil price data to correct the production threshold makes the production capacity prediction more dynamic and flexible, enabling real-time response to market changes. By accurately predicting production capacity and making intelligent judgments on whether to exploit, the risk of blind exploitation is reduced, and resource waste caused by over-exploitation or inefficient exploitation is avoided.
[0030] Embodiment 2
[0031] Please refer to Figure 2 As shown, for the parts not described in detail in this embodiment, refer to the description content of Embodiment 1. A petroleum exploration system based on artificial intelligence is provided, including: Data acquisition module: Collect seismic data, logging data, exploration environment data, and international crude oil price data. Data processing module: Conduct preliminary processing on seismic data and logging data to obtain a preliminary data set; based on the preliminary data set, use an improved PCA algorithm for feature extraction to obtain a new feature data set. Data prediction module: For the new feature data set, use the DNN model to predict the oil thickness data, porosity data, and permeability data. Oil exploration module: Based on the predicted oil thickness data, porosity data, and permeability data, use a regularized exploration environment regression model to predict oil production capacity data, set a production capacity threshold, compare the predicted oil production capacity data with the production capacity threshold, and determine whether to extract oil.
[0032] Embodiment 3
[0033] This embodiment discloses and provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the operation mode of the above-provided oil exploration method and system based on artificial intelligence.
[0034] Since the electronic device introduced in this embodiment is the electronic device used to implement an oil exploration method and system based on artificial intelligence in an embodiment of the present application, based on the oil exploration method and system based on artificial intelligence introduced in an embodiment of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in an embodiment of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device used in an oil exploration method and system based on artificial intelligence in an embodiment of the present application, it falls within the scope of protection of the present application.
[0035] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0036] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for ordinary technical users in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. An artificial intelligence-based oil exploration method, characterized in that: include: Step S1, collecting seismic data, well logging data, exploration environment data and international crude oil price data; Step S2: performing preliminary processing based on the seismic data and the well logging data to obtain a preliminary data set; based on the preliminary data set, using an improved PCA algorithm to perform feature extraction to obtain a new feature data set; Step S3: Based on the new feature data set, use the DNN model to predict oil thickness data, porosity data, and permeability data; Step S4: Based on the predicted oil thickness data, porosity data and permeability data, a regularized exploration environment regression model is used to predict the oil production capacity data, and a production capacity threshold is set. The predicted oil production capacity data is compared with the production capacity threshold to determine whether to exploit the oil.
2. The artificial intelligence-based oil exploration method according to claim 1, characterized in that: The seismic data includes amplitude characteristic data, wave velocity characteristic data, radiation characteristic data, seismic wave frequency characteristic data and magnetic characteristic data; Well logging data include resistivity characteristic data, gamma ray characteristic data, acoustic wave time difference characteristic data, rock density characteristic data, neutron porosity characteristic data, conductivity characteristic data and natural gamma characteristic data; The exploration environment data includes temperature characteristic data, humidity characteristic data and air pressure characteristic data.
3. The artificial intelligence-based oil exploration method according to claim 2, characterized in that: The method of obtaining the preliminary data set includes: For seismic data and well logging data, the mean filling method is used to handle missing values, the interpolation method is used to handle outliers, the Z-Score is used for standardization, and unified timestamp processing is performed to obtain a preliminary data set.
4. The artificial intelligence-based petroleum exploration method according to claim 3, characterized in that: The specific method of extracting features based on the preliminary data set using the improved PCA algorithm to obtain a new feature data set includes: Step A11: For the preliminary data set, use Methods: Construct preprocessing matrix; Step A12: construct a covariance matrix for the preprocessing matrix. ,in, is the covariance matrix, is the transpose of the preprocessing matrix, is the preprocessing matrix with dimension , n is the total number of preprocessing matrix samples, is the total number of features in the preprocessing matrix. A row in the preprocessing matrix represents a sample. The dimension is ; Step A13: Use singular value decomposition method SVD to decompose the covariance matrix into three matrices, the formula is: , for An orthogonal matrix containing left singular vectors; for A diagonal matrix of , where the elements on the diagonal are singular values; for An orthogonal matrix containing the right singular vectors; Step A14: Calculate the variance contribution of the principal component using the singular value. The formula is: ,in, For the The singular value of the principal component, that is, The variance of the principal components, i is the principal component index, represents the sum of all principal component singular values, For the The variance contribution of each principal component; Step A15: Calculate the average variance contribution of all principal components. The formula is: , is the average contribution of all principal components, and selects the k best principal components whose contribution is greater than the average contribution; Step A16: Construct the k best principal components into the best principal component matrix , the preprocessing matrix Projected to the new feature matrix, the calculation formula is: ,in, The new feature matrix after projection is converted into a new feature dataset using the pandas.DataFrame method.
5. The artificial intelligence-based petroleum exploration method according to claim 4, characterized in that: The specific method of using the DNN model to predict oil thickness data, porosity data and permeability data based on the obtained new feature data set includes: Step D11: Input a sample set, the sample set includes A set of samples, each set of samples includes a new feature data set and corresponding oil thickness data, porosity data, and permeability data; Step D12, set the number of neurons VB in the input layer to be equal to the number of features in the new feature data set; use grid search to set the number of hidden layers to OP, and set 2*VB neurons in each hidden layer; set the number of neurons in the output layer to 3, which corresponds to the predicted oil thickness data, porosity data, and permeability data; initialize the model weights and biases, and set the number of iterations; Step D13, in the PRth iteration, the sample is transferred from the input layer to the first hidden layer, and the output of the first hidden layer is calculated using the mixed activation function; the output of the first hidden layer is transferred to the next hidden layer as the input of the next hidden layer, and the mixed activation function is repeatedly used to calculate the output of the hidden layer until the last hidden layer ends; the output of the last hidden layer is transferred to the output layer to obtain the output of the PRth iteration, which includes the predicted oil thickness data, the predicted porosity data, and the output permeability data; Step D14: for the output PR-th iteration predicted oil thickness data, predicted porosity data and output permeability data, use the mean square error as the loss function to calculate the error between the PR-th iteration predicted value and the true value, and perform weighted summation on the errors of the predicted oil thickness data, predicted porosity data and output permeability data to obtain a total loss function error, wherein the true value refers to the corresponding oil thickness data, porosity data and permeability data in step D11; Step D15: Calculate the gradient of weights and biases of each layer using the chain rule through the total loss function error, and use the Adam optimizer to update the weights and biases of each layer; Step D16: When the set number of iterations is reached, the iteration is stopped; and the predicted oil thickness data, porosity data and permeability data are output.
6. The artificial intelligence-based petroleum exploration method according to claim 5, characterized in that: The specific method of using the hybrid activation function to calculate the output of the first hidden layer includes: In the DNN model, the activation function includes Activation function, Activation function, Activation function and Activation function; The weighted average method is used to calculate the mixed activation function, the formula is: , in, , , and They are Activation function, Activation function, Activation function and The weight coefficient of the activation function, and .
7. The artificial intelligence-based petroleum exploration method according to claim 6, characterized in that: The specific method of using the regularized exploration environment regression model to predict oil production capacity data based on the predicted oil thickness data, porosity data and permeability data includes: Step ff1, input A group of oil samples, each group of oil samples includes a predicted oil thickness characteristic data, a porosity characteristic data and a permeability characteristic data, and corresponding oil production capacity data; Step ff2, set the number of iterations; initialize the intercept , initialize the regression coefficients , and ,in, , and are the regression coefficients of the input predicted oil thickness data, porosity data, and permeability data respectively; Step ff3: Calculate the predicted value of oil production capacity data based on the initialized regression coefficient and intercept. The formula is: ,in, , and represents the predicted oil thickness data, porosity data and permeability data of the sth group of oil samples, s is the oil sample index, represents the predicted oil production capacity data of the sth group of oil samples; Step ff4, using regularization and exploration environment data to correct the loss function; Step ff5, calculate the gradient of the loss function with respect to the regression coefficient and the intercept by back propagation; then use the gradient descent method to update the regression coefficient and the intercept; Step ff6: Repeat steps ff3 to ff5 until the set number of iterations is reached to obtain the final regression coefficient and intercept.
8. The artificial intelligence-based petroleum exploration method according to claim 7, characterized in that: The specific method of using regularization and exploration environment data to correct the loss function includes: The loss function is calculated using the average error, and the formula is: ,in, is the oil production capacity data in the sth group of oil samples; The loss function is corrected using regularization and exploration environment data to obtain the corrected loss function, the formula is: ,in, is the regularization parameter, is the L2 norm of the regression coefficient, is the index of the regression coefficient, For the regression coefficients, is the correction coefficient of the exploration environment data of the sth oil sample at the same time, is the regulating factor; Among them, the correction coefficient of exploration environment data is calculated by weighted average method. ,in, , and It is divided into the weights of temperature data, humidity data and air pressure data in the exploration environment data. , and They are respectively the temperature characteristic data values, humidity characteristic data values and air pressure characteristic data obtained at the same time as the sth group of oil samples.
9. The artificial intelligence-based petroleum exploration method according to claim 8, characterized in that: The specific method of setting the production capacity threshold and comparing the predicted oil production capacity data with the production capacity threshold to determine whether to produce oil includes: Calculate the mean and standard deviation based on historical oil production capacity data, and set the initial production capacity threshold based on the mean and standard deviation of the historical oil production capacity data; The price average is calculated based on the historical international crude oil price data, and the initial capacity threshold is corrected by the ratio of the historical international crude oil average price data to the current international crude oil price data to obtain the capacity threshold; When the oil production capacity data is greater than or equal to the production capacity threshold data, it means that the explored oil can be mined; when the oil production capacity data is less than the production capacity threshold data, it means that the explored oil cannot be mined; The formula for the average of historical oil production data is: ,in, is the average value of historical oil production capacity, is the total number of historical oil production capacity data, is an index of historical oil production data, For the Historical oil production capacity data; The standard deviation formula for historical oil production data is: , is the standard deviation of historical oil production capacity; The formula for the initial capacity threshold is: , is the initial capacity threshold, is the adjustment factor; The price average formula of historical international crude oil price data is: ,in, It is an index of historical international crude oil prices. For the The historical international crude oil prices, is the average price of historical international crude oil; The capacity threshold formula is: ,in, It is the current international original price data.
10. An artificial intelligence-based oil exploration system, used to implement the artificial intelligence-based oil exploration method according to any one of claims 1 to 9, characterized in that: include: Data acquisition module: collects seismic data, well logging data, exploration environment data and international crude oil price data; Data processing module: Perform preliminary processing on seismic data and well logging data to obtain a preliminary data set; based on the preliminary data set, use the improved PCA algorithm to extract features to obtain a new feature data set; Data prediction module: For new feature data sets, use DNN models to predict oil thickness data, porosity data, and permeability data; Oil exploration module: Based on the predicted oil thickness data, porosity data and permeability data, the regularized exploration environment regression model is used to predict the oil production capacity data, and the production capacity threshold is set. The predicted oil production capacity data is compared with the production capacity threshold to determine whether to exploit the oil.
Citation Information
Patent Citations
Prediction model training method, price prediction method, storage medium and electronic equipment
CN112101566A
Geographical property analysis method based on neural network
CN116204831A
Transverse wave time difference prediction method and system based on integrated machine learning
CN118642170A
Prediction method and device based on principal component regression and nonvolatile storage medium
CN118916689A