A distributed filling method for missing values ​​in well logging data based on Spark

By building a distributed computing environment on the Spark platform, using distributed random forest and GBT models to fill the missing value of logging data, the problems of high time cost and low accuracy of missing value filling in the existing technology are solved, and efficient and accurate massive logging data processing is achieved.

CN115268848BActive Publication Date: 2025-05-13CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210855411.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-05-13
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

When processing massive well logging data, the time cost of missing values ​​is high and the accuracy is low, and the stand-alone machine learning method is memory-limited, so it is impossible to effectively process big data scenarios.

Method used

Using a distributed filling method based on Spark, the distributed storage and computing of well logging data is realized by building Hadoop clusters and Spark on Yarn clusters. Missing value prediction was performed using distributed random forest and distributed GBT models, and model parameters were optimized through distributed grid search + k-fold cross-validation and Train-Validation-Split algorithm.

Benefits of technology

It reduces the time complexity of missing values, improves the prediction accuracy of the model, breaks through the single-machine memory limit, and can efficiently process massive well logging data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115268848B_ABST
    Figure CN115268848B_ABST
Patent Text Reader

Abstract

The present invention relates to a distributed filling method for missing values ​​of well logging data based on Spark, and belongs to the field of missing data filling. The distributed filling method for missing values ​​of well logging data based on Spark provided by the present invention realizes distributed storage of well logging data in exploration work by using HDFS as a storage system, as an information source for distributed computing; installs and deploys a Spark cluster, and uses Yarn as a resource management and task scheduling framework; performs secondary preprocessing on well logging data in the data warehouse by building indexes, standardization processing and other methods; predicts missing values ​​of well logging data in exploration work by distributed random forest and distributed GBT models; optimizes distributed prediction filling models by distributed grid search + k-fold cross validation and Train‑Validation‑Split methods. The present invention can provide a solution with higher accuracy and lower time cost for the problem of missing data in well logging, and provides a guarantee for further research, analysis and utilization of well logging data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a distributed filling method for missing values ​​of well logging data based on Spark, and belongs to the field of missing filling of well logging data. Background Art

[0002] Well logging data in exploration work is the data basis for lithology identification, mineralization prediction, data analysis and data mining. Well logging technology mainly uses professional equipment to emit reflectors, electricity, sound waves and other attributes to explore the attribute information of underground geological layers. Experts can further understand the underground stratigraphic structure by analyzing well logging data, and use well logging data to establish a more accurate geological three-dimensional model. However, due to factors such as well diameter expansion, instrument failure and human factors, some well section attribute information is often distorted or missing in actual applications. For cost considerations, people usually use artificial generation methods, interactive graph methods, and multivariate regression methods to fill in the missing value logging curve attribute information. These methods often have problems such as poor effect, low accuracy, high labor cost, and high time cost. Moreover, it is difficult to fully guarantee the integrity of data in the process of massive well logging data collection, transmission, and storage. The integrity of well logging data is destroyed, and the accuracy of applications such as mineralization prediction, lithology identification, intelligent interpretation, and geological three-dimensional models cannot be guaranteed. With the establishment of a distributed storage and management system for exploration big data. Traditional solutions for missing values, such as manual input, deletion, statistical learning (mode, mean, maximum, minimum, etc.) and single-machine machine learning, perform poorly in terms of time cost and accuracy of missing value filling. Manual input usually fills missing values ​​based on experience, but this is subjective and cannot produce an accurate prediction value; the deletion method directly removes the data with missing values. In the case of large dimensions of exploration logging data, if each attribute information has some missing values, all attribute information is analyzed and all missing value data is deleted, resulting in the loss of a large amount of sample data, thereby losing a large amount of useful information and causing a large amount of resource waste; the statistical learning method usually uses similar statistical values ​​such as mode, mean, maximum value to fill missing values, ignoring the correlation and nonlinear relationship between data; the single-machine machine learning method is to train the missing value model on a single node. However, in the big data scenario, the memory limitation of a single node becomes a bottleneck, resulting in the inability to train the model or high time cost. Therefore, using mining technology to mine the correlation between data from the big data distributed storage system to fill the missing values ​​of logging data has become an urgent problem to be solved, and it is also a key step to provide high-quality, high-precision and complete logging data. Summary of the invention

[0003] In order to solve the problems existing in the above-mentioned prior art, the present invention provides a distributed filling method for missing values ​​of well logging data based on Spark. The present invention can provide a solution with higher accuracy and lower time cost for the problem of missing data in well logging, and provides a guarantee for further research, analysis and utilization of well logging data.

[0004] To achieve the above purpose, the technical solution provided by the present invention is: a distributed filling method for missing values ​​of well logging data based on Spark, which is operated according to the following steps:

[0005] (1) Building a storage module: By building a MapReduce parallel computing framework in the server and building a Hadoop cluster within the MapReduce parallel computing framework, the HDFS component in the Hadoop cluster is used to perform distributed storage of well logging data in the exploration work; the HDFS cluster is used to store data in the Hive on Spark well logging data warehouse;

[0006] (2) Build a Spark on Yarn cluster: Optimize the MapReduce parallel computing framework by installing and deploying a Spark cluster, and use Yarn as a resource management and task scheduling framework;

[0007] (3) Secondary preprocessing of well logging data: Secondary preprocessing of well logging data in the data warehouse by building indexes and standardizing them;

[0008] (4) Integrated algorithm model construction: By building distributed random forest and distributed GBT models, missing values ​​of well logging data in exploration work are predicted;

[0009] (5) Model parameter adjustment: By building a distributed prediction filling model optimized by the distributed grid search + k-fold cross-validation model and the Train-Validation-Split algorithm model, and optimizing the parameters of the distributed prediction filling model, the validation error and test accuracy of the distributed prediction filling model meet the design requirements;

[0010] (6) Predict and fill missing values ​​in logging data: Use the optimized distributed prediction and filling model to perform distributed prediction and data filling for missing values ​​in logging data in mineral exploration based on its performance and efficiency.

[0011] In step (1), follow these steps:

[0012] 1) Data transmission: Java programming is used to upload the collected semi-structured data and unstructured data in batches. For structured data, the Sqoop tool is used to extract data and transfer the data to the HDFS component.

[0013] 2) Distributed data storage: Build a Hadoop cluster through servers and use HDFS components to achieve distributed data storage;

[0014] 3) Hive data warehouse: Establish a well logging data warehouse based on Hive On Spark. The well logging data warehouse mainly consists of GODS layer, GDWD layer and GDWT layer;

[0015] 4) Determine the data synchronization strategy: According to the storage form of logging data, the synchronization strategy is divided into full table, incremental table and special table;

[0016] 5) Optimization of Hive data warehouse: Replace the computing engine in the MapReduce parallel computing framework with a Spark cluster, and use the Spark computing engine to improve the efficiency of Hive query and data analysis.

[0017] In step (2), the MapReduce parallel computing framework optimized by the Spark cluster is installed and deployed, and Yarn is used as the resource management and task scheduling framework, where Spark only implements the scheduling task to enable the MapReduce parallel computing framework to achieve the purpose of iteration and adapting to real-time computing.

[0018] In step (3), the well logging data in the data warehouse are preprocessed for secondary processing by building indexes and standardizing the data using the installed and deployed Spark cluster.

[0019] Step (4) comprises at least the following steps:

[0020] 1) Use HDFS components to store uranium exploration and logging data in a distributed manner and use it as a data source for missing value filling;

[0021] 2) Initialize SparkSession, index the non-numeric attributes of uranium exploration and logging data, and standardize the uranium exploration and logging data;

[0022] 3) Using the method of randomly extracting data sets, the standardized uranium exploration logging data are divided into training data sets and test data sets in a ratio of 8:2;

[0023] 4) Uniformly establish vector index values ​​for the input feature labels and output feature labels of the logging data; convert the engineering feature values ​​of the training data set and the test data set into vectors, and complete basic data processing;

[0024] 5) Build a distributed random forest model and a distributed GBT model respectively; for the distributed random forest model, use Scala language iterative programming, adopt the determined feature vector index and label value, use the model's fit operator training data set and transform test data set and build a regression prediction evaluation model; for the distributed GBT model, transform the feature vector index and label value of the distributed GBT model, and use the validation set to verify the model's fit;

[0025] 6) Save the prediction model, prediction data, and statistical values ​​to the HDFS component;

[0026] 7) Use IDEA to package the algorithm model and deploy it to the Spark distributed environment.

[0027] The step (5) is to optimize the distributed prediction filling model by using a distributed grid search + k-fold cross validation model and a Train-Validation-Split algorithm model, wherein the distributed grid search + k-fold cross validation model is suitable for small data sets, and the Train-Validation-Split algorithm model is suitable for massive data sets.

[0028] The step (5) is performed as follows:

[0029] 1) Initialize SparkSession; read logging data from Hive and convert it into a DataFrame data structure, save the input feature labels and output feature labels of the logging data in the DataFrame as objects, and convert the features in the DataFrame into Vectors;

[0030] 2) Split the data set consisting of well logging data into trainData and testData in a ratio of 8:2, and convert the Vector data of trainData into Vector index data;

[0031] 3) Iterative programming is used to increase the hyperparameter grid, where the data format of the hyperparameter grid is {model hyperparameter, Array (hyperparameter value)};

[0032] 4) Set the predicted label value, output the predicted label name and evaluation index;

[0033] 5) For the distributed grid search + k-fold cross validation model, first build the grid search model, then define the pipeline, evaluator and grid model into the distributed grid search + k-fold cross validation model, and use the trainData dataset to train the distributed grid search + k-fold cross validation model;

[0034] For the Train-Validation-Split algorithm model, first define the Train-Validation-Split algorithm model, then define the defined pipeline and evaluato in the Train-Validation-Split algorithm model, use the trainData dataset to train, validate and optimize the Train-Validation-Split algorithm model, and use the testData dataset to evaluate the Train-Validation-Split algorithm model;

[0035] 6) The distributed prediction filling model is optimized by using the layout grid search + k-fold cross validation model and the Train-Validation-Split algorithm model.

[0036] For the Train-Validation-Split optimized distributed prediction filling model, first define the Train-Validation-Split model, then define the defined pipeline and evaluato in the Train-Validation-Split model, use the trainData dataset to train, validate and optimize the Train-Validation-Split distributed prediction filling model, and use the testData dataset to evaluate the Train-Validation-Split optimized distributed prediction filling model.

[0037] According to the above technical scheme, the distributed filling method of missing values ​​of well logging data based on Spark provided by the present invention realizes distributed storage of well logging data in exploration work by using HDFS components as the storage system, as the information source of distributed computing; installs and deploys Spark clusters, and uses Yarn as a resource management and task scheduling framework; performs secondary preprocessing on well logging data in the data warehouse by building indexes, standardization processing and other methods; predicts missing values ​​of well logging data in exploration work through distributed random forest and distributed GBT models implemented in Scala language; implements distributed grid search + k-fold cross validation and Train-Validation-Split method to optimize distributed prediction filling model through Scala language programming; finally, uses the optimized model to perform distributed prediction and data filling of missing values ​​of well logging data in mineral exploration according to its performance and efficiency. Compared with existing technical schemes, this technical scheme has the following advantages:

[0038] (1) Because this technical solution solves the memory overflow phenomenon of the missing value prediction and filling model of massive well logging data in uranium mine exploration work through a distributed prediction and filling model, and the distributed random forest and distributed GBT models perform distributed prediction and filling of missing values ​​of uranium well logging data, this method has lower time complexity and higher model prediction and filling accuracy.

[0039] (2) Since the missing value distributed random forest prediction filling model and the missing value distributed GBT prediction filling model proposed in this technical solution predict and fill the missing values ​​of well logging data in uranium mine exploration, this method breaks through the bottleneck of single-machine memory limitation under massive data. In addition, since it fully utilizes cluster resources and performance and mines potential knowledge and information in the data in a distributed environment, this method improves the accuracy of the missing value prediction filling model and reduces the time cost.

[0040] (3) Since the technical solution proposes a distributed grid search + k-fold cross validation and Train-Validation-Split to compare and optimize the distributed prediction filling model, among which the distributed grid search + k-fold cross validation is suitable for small data sets; Train-Validation-Split is suitable for massive data sets, so this method improves the speed and accuracy of the distributed prediction filling model for missing values ​​of uranium mine logging data. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 Flowchart of the distributed prediction filling model for missing values ​​in well logging data;

[0042] Figure 2 Flowchart of distributed storage of logging data based on Hive data warehouse;

[0043] Figure 3 Flowchart of distributed random forest implementation;

[0044] Figure 4 Flowchart of distributed GBT implementation. Specific implementation methods

[0045] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments, but the protection scope of the present invention is not limited to the following embodiments.

[0046] In the technical solution provided by the present invention, a distributed filling method for missing values ​​of well logging data based on Spark is provided, such as Figure 1 As shown, follow the steps below:

[0047] (1) Constructing a storage module: By building a MapReduce parallel computing framework in the server and building a Hadoop cluster inside the MapReduce parallel computing framework, the HDFS component in the Hadoop cluster is used to distribute the well logging data in the exploration work; the HDFS cluster is used to store data in the Hive on Spaker well logging data warehouse; HDFS does not require high machine performance and can be deployed on cheap machines. Fault tolerance is improved by configuring multiple copies of the data directory. Distributed storage of well logging data in the exploration work is implemented as an information source for distributed computing, making it highly reliable, scalable, efficient, and fault tolerant; wherein, step (1) is as follows Figure 2 As shown, please follow the steps below:

[0048] 1) Data transmission: Java programming is used to upload batch data of collected semi-structured and unstructured data. For structured data, such as well logging data, Sqoop tool is used to extract data and transmit it to HDFS cluster;

[0049] 2) Distributed data storage: Hadoop clusters are built through servers, and HDFS components are used to achieve distributed data storage; Hadoop implements a distributed file system, one of which is HDFS. HDFS has the characteristics of high fault tolerance and is designed to be deployed on low-cost hardware; it also provides high throughput to access application data, and is suitable for applications with very large data sets.

[0050] 3) Hive data warehouse: Establish a well logging data warehouse based on Hive On Spark. The well logging data warehouse is mainly composed of GODS layer, GDWD layer and GDWT layer. The data in the Hive On Spark data warehouse is stored in Hive in the form of table. Users process and analyze data using hql with Hive syntax specification, i.e. hive sql. However, when users submit these hql for execution, the underlying layer will be parsed, optimized and compiled by Hive, and finally run as Spark jobs.

[0051] 4) Data synchronization strategy: According to the storage form of logging data, the synchronization strategy is divided into full table, incremental table and special table;

[0052] 5) Optimization of Hive data warehouse: Replace the computing engine of the MapReduce parallel computing framework with a Spark cluster, and use the Spark computing engine to improve the efficiency of Hive query and data analysis.

[0053] (2) Build a Spark on Yarn cluster: Install and deploy a Spark cluster, and use Yarn as a resource management and task scheduling framework; and in step (2), install and deploy the Spark cluster to optimize the MapReduce parallel computing framework, and use Yarn as a resource management and task scheduling framework. The Spark distributed environment can be deployed on the established distributed storage system for model training. Since MapReduce has problems such as inability to perform iterative operations, frequent disk operations, and inability to adapt to real-time computing, and Spark provides more operators, the models between operators are divided into narrow dependencies and wide dependencies, so the Spark job scheduling mechanism is adopted to replace MapReduce. The specific idea is to use IDEA to package the constructed model into a jar package, and then deploy it to the Spark distributed environment for model training. Spark only implements scheduling tasks to enable MapReduce to achieve iteration and adapt to real-time computing.

[0054] (3) Secondary preprocessing of logging data: Secondary preprocessing of logging data in the data warehouse is performed by building indexes and standardizing the data. In step (3), the well logging data in the data warehouse is secondarily preprocessed by building indexes and standardizing the data using the installed and deployed Spark cluster. In this process, data processing is mainly completed by building indexes and standardizing the data.

[0055] (4) Integrated algorithm model construction: By building distributed random forest and distributed GBT models, such as Figure 3 and Figure 4 As shown, the missing values ​​of the well logging data in the exploration work are predicted; and step (4) at least includes the following steps:

[0056] 1) Establish HDFS distributed storage for uranium exploration and logging data, and use it as a data source for missing value filling;

[0057] 2) Initialize SparkSession, index the non-numeric attributes of uranium exploration and logging data, and standardize the uranium exploration and logging data;

[0058] 3) Using the method of randomly extracting data sets, the standardized uranium exploration logging data are divided into training data sets and test data sets in a ratio of 8:2;

[0059] 4) Unify the input feature labels and output feature labels to establish vector index values; convert the engineering feature values ​​of the training data set and the test data set into vector vectorTransfor, and complete basic data processing;

[0060] 5) Build a distributed random forest model and a distributed GBT model respectively; for the distributed random forest model, use Scala language iterative programming, adopt the determined feature vector index and label value, use the model's fit operator training data set and transform test data set and build a regression prediction evaluation model; for the distributed GBT model, transform the feature vector index and label value of the distributed GBT model, and use the validation set to verify the model's fit;

[0061] When the optimization of the validation error does not exceed the threshold BoostingStrategy, stop training. Build a pipeline to build a machine learning workflow with featureIndexer and the above prediction model and validation set model.

[0062] Then, the model's fit operator is used to train the data set and transform the test data set, and a regression prediction evaluation model is constructed.

[0063] 6) Save the prediction model, prediction data, and statistical values ​​to HDFS;

[0064] 7) Use IDEA to package the algorithm model and then deploy it to the Spark distributed environment.

[0065] (5) Model parameter adjustment: By building a distributed prediction filling model optimized by distributed grid search + k-fold cross validation and Train-Validation-Split method, the parameters of the model are optimized so that the validation error and test accuracy of the model meet the design requirements; the step (5) is to use distributed grid search + k-fold cross validation and Train-Validation-Split to optimize the distributed prediction filling model, wherein distributed grid search + k-fold cross validation is suitable for small data sets, and Train-Validation-Split is suitable for massive data sets.

[0066] 1) Initialize SparkSession; read logging data from Hive and convert it into a DataFrame data structure, save the input feature labels and output feature labels of the logging data in the DataFrame as objects, and convert the features in the DataFrame into Vectors;

[0067] 2) Split the data set consisting of well logging data into trainData and testData in a ratio of 8:2, and convert the Vector data of trainData into Vector index data;

[0068] 3) Iterative programming is used to increase the hyperparameter grid, where the data format of the hyperparameter grid is {model hyperparameter, Array (hyperparameter value)};

[0069] 4) Set the predicted label value, output the predicted label name and evaluation index;

[0070] 5) For the distributed grid search + k-fold cross validation model, first build the grid search model, then define the pipeline, evaluator and grid model into the distributed grid search + k-fold cross validation model, and use the trainData dataset to train the distributed grid search + k-fold cross validation model;

[0071] For the Train-Validation-Split algorithm model, first define the Train-Validation-Split algorithm model, then define the defined pipeline and evaluato in the Train-Validation-Split algorithm model, use the trainData dataset to train, validate and optimize the Train-Validation-Split algorithm model, and use the testData dataset to evaluate the Train-Validation-Split algorithm model;

[0072] 6) The distributed prediction filling model is optimized by using the cloth grid search + k-fold cross validation model and the Train-Validation-Split algorithm model.

[0073] The optimized model is used to perform actual distributed prediction and filling of missing values ​​in well logging data in mineral exploration based on its performance and efficiency. In addition, mse, rmse, mae, and r2 can be used as evaluation indicators to evaluate the distributed prediction and filling model, thereby verifying the effectiveness and feasibility of the distributed random forest prediction and filling model for missing values ​​in uranium well logging data and the distributed GBT prediction and filling model.

Claims

1. A distributed filling method for missing values ​​in well logging data based on Spark, characterized in that Proceed as follows: (1) Building a storage module: By building a MapReduce parallel computing framework in the server and building a Hadoop cluster within the MapReduce parallel computing framework, the HDFS component in the Hadoop cluster is used to perform distributed storage of well logging data in the exploration work; the HDFS cluster is used to store data in the HiveonSpark well logging data warehouse; (2) Build a Spark on Yarn cluster: Optimize the MapReduce parallel computing framework by installing and deploying a Spark cluster, and use Yarn as a resource management and task scheduling framework; (3) Secondary preprocessing of well logging data: Secondary preprocessing of well logging data in the data warehouse by building indexes and standardizing them; (4) Integrated algorithm model construction: By building distributed random forest and distributed GBT models, missing values ​​of well logging data in exploration work are predicted; (5) Model parameter adjustment: By building a distributed prediction filling model optimized by the distributed grid search + k-fold cross-validation model and the Train-Validation-Split algorithm model, and optimizing the parameters of the distributed prediction filling model, the validation error and test accuracy of the distributed prediction filling model meet the design requirements; (6) Predict and fill missing values ​​in logging data: Use the optimized distributed prediction and filling model to perform distributed prediction and data filling for missing values ​​in logging data in mineral exploration based on its performance and efficiency.

2. The distributed filling method for missing values ​​of well logging data based on Spark according to claim 1 is characterized in that: In step (1), follow these steps: 1) Data transmission: Java programming is used to upload the collected semi-structured data and unstructured data in batches. For structured data, the Sqoop tool is used to extract data and transfer the data to the HDFS component. 2) Distributed data storage: Build a Hadoop cluster through servers and use HDFS components to achieve distributed data storage; 3) Hive data warehouse: Establish a well logging data warehouse based on HiveOnSpark. The well logging data warehouse mainly consists of GODS layer, GDWD layer and GDWT layer; 4) Determine the data synchronization strategy: According to the storage form of logging data, the synchronization strategy is divided into full table, incremental table and special table; 5) Optimization of Hive data warehouse: Replace the computing engine in the MapReduce parallel computing framework with a Spark cluster, and use the Spark computing engine to improve the efficiency of Hive query and data analysis.

3. The distributed filling method for missing values ​​of well logging data based on Spark according to claim 1 is characterized in that: In step (2), the MapReduce parallel computing framework optimized by the Spark cluster is installed and deployed, and Yarn is used as the resource management and task scheduling framework, where Spark only implements the scheduling task to enable the MapReduce parallel computing framework to achieve the purpose of iteration and adapting to real-time computing.

4. The distributed filling method for missing values ​​of well logging data based on Spark according to claim 1 is characterized in that: In step (3), the well logging data in the data warehouse are preprocessed for secondary processing by building indexes and standardizing the data using the installed and deployed Spark cluster.

5. The distributed filling method for missing values ​​of well logging data based on Spark according to claim 1 is characterized in that: Step (4) comprises at least the following steps: 1) Use HDFS components to store uranium exploration and logging data in a distributed manner and use it as a data source for missing value filling; 2) Initialize SparkSession, index the non-numeric attributes of uranium exploration and logging data, and standardize the uranium exploration and logging data; 3) Using the random extraction method, the standardized uranium exploration logging data are divided into a training data set and a test data set in a ratio of 8:2; 4) Uniformly establish vector index values ​​for the input feature labels and output feature labels of the logging data; convert the engineering feature values ​​of the training data set and the test data set into vectors, and complete basic data processing; 5) Build a distributed random forest model and a distributed GBT model respectively; for the distributed random forest model, use Scala language iterative programming, adopt the determined feature vector index and label value, use the model's fit operator training data set and transform test data set and build a regression prediction evaluation model; for the distributed GBT model, transform the feature vector index and label value of the distributed GBT model, and use the validation set to verify the model's fit; 6) Save the prediction model, prediction data, and statistical values ​​to the HDFS component; 7) Use IDEA to package the algorithm model and deploy it to the Spark distributed environment.

6. The distributed filling method for missing values ​​of well logging data based on Spark according to claim 1 is characterized in that: The step (5) is to optimize the distributed prediction filling model by using a distributed grid search + k-fold cross validation model and a Train-Validation-Split algorithm model, wherein the distributed grid search + k-fold cross validation model is suitable for small data sets, and the Train-Validation-Split algorithm model is suitable for massive data sets.

7. The distributed filling method for missing values ​​of well logging data based on Spark according to claim 1 is characterized in that: The step (5) is performed as follows: 1) Initialize SparkSession; read logging data from Hive and convert it into a DataFrame data structure, save the input feature labels and output feature labels of the logging data in the DataFrame as objects, and convert the features in the DataFrame into Vectors; 2) Split the data set consisting of well logging data into trainData and testData in a ratio of 8:2, and convert the Vector data of trainData into Vector index data; 3) Iterative programming is used to increase the hyperparameter grid, where the data format of the hyperparameter grid is {model hyperparameter, Array (hyperparameter value)}; 4) Set the predicted label value, output the predicted label name and evaluation index; 5) For the distributed grid search + k-fold cross validation model, first build the grid search model, then define the pipeline, evaluator and grid model into the distributed grid search + k-fold cross validation model, and use the trainData dataset to train the distributed grid search + k-fold cross validation model; For the Train-Validation-Split algorithm model, first define the Train-Validation-Split algorithm model, then define the defined pipeline and evaluato in the Train-Validation-Split algorithm model, use the trainData dataset to train, validate and optimize the Train-Validation-Split algorithm model, and use the testData dataset to evaluate the Train-Validation-Split algorithm model; 6) The distributed prediction filling model is optimized by using the cloth grid search + k-fold cross validation model and the Train-Validation-Split algorithm model.

Citation Information

Patent Citations

  • Logging data complementing method based on Bayesian optimization and auto-encoder

    CN114996625A