A method and related device for filling missing values of drilling sampling and testing data

CN122817657APending Publication Date: 2026-09-25SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610827692.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007]综上所述,从最接近的现有技术来看,已有机器学习插补方法通常是将待处理数据作为普通表格变量进行预测,或者在地学场景中主要依据已有元素数值和少量空间代理变量估计缺失项;已有地质文本挖掘方法则多服务于找矿判据提取、知识组织或预测辅助;上述方案虽然分别提供了算法基础、空间建模思路和文本结构化基础,但仍存在以下不足:其一,缺少面向钻孔取样化验表、以“孔号—样品编号—起止深度”共同确定样段对象的专门补全流程,难以反映钻孔数据沿孔深组织的结构特征;其二,缺少将取样化验表、钻孔分层表、钻孔测斜表和钻孔基本信息表按样段进行统一对齐,并进一步组织为补全样本的技术路径;其三,缺少将元素协同关系、样段空间位置和地质文本语义三类信息联合组织为补全约束的机制,而不仅仅是对若干字段进行简单拼接;其四,现有方法通常偏重预测精度比较,缺少与钻孔数据库修复、结果来源标记、写回状态和人工复核流程相衔接的闭环处理方式;因此,有必要提出一种基于多源约束的钻孔取样化验数据缺失值补全方法,以更好满足矿床勘探数据库修复与后续三维建模、矿体解释和地学分析的需求

Benefits of technology

[0019]在本发明实施例中,实现以钻孔样段为基本处理单元,在不依赖外部异源训练数据的条件下,利用目标钻孔数据集中已有完整样段记录构建补全模型,并联合引入元素协同关系、样段空间位置和地质文本语义信息,对缺失字段进行补全;与现有主要依赖数值统计关系或一般表格预测流程的缺失值补全方法相比,能够更好适配钻孔数据“孔号—样号—起止深度”的组织方式,更充分利用矿区钻孔数据中已有的数值、空间和文本信息,提高缺失值补全结果对目标矿区数据结构和地质背景的针对性与合理性;最后,生成的补全结果可直接写回取样化验表,因而不仅能够完成元素含量字段的缺失值估计,还能够直接服务于钻孔数据库修复以及后续三维建模、矿体解释和地学分析。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817657A_ABST
    Figure CN122817657A_ABST
Patent Text Reader

Abstract

The application discloses a kind of drilling sampling assay data missing value completion method and related device, its method includes: obtaining basic input data, and the content field missing value of target element is as target completion object;Training sample set is constructed, and training sample set is handled to multi-source constraint, and multi-source constraint training data set is formed;Target element missing value completion model is constructed, and multi-source constraint training data set is as target element missing value completion model input, and the real observation value of target element in multi-source constraint training data set is as output mode and is trained to process, and training convergent target element missing value completion model is formed;Basic input data is input into training convergent target element missing value completion model, input content field completion value corresponding to target completion object, and content field completion value is filled into sampling assay table to complete completion.In the embodiment of the application, the completion of the missing value of drilling sampling assay data can be more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for completing missing values ​​in borehole sampling and testing data. Background Technology

[0002] While missing value completion in borehole sampling and testing data is formally a structured data missing value handling problem, its application is significantly constrained by the spatial location of the sample section, mineralization distribution, and geological background, and cannot be simply equated with missing value completion for general tabular data. Existing related technologies can be roughly divided into three categories: one is general missing value completion methods, such as mean substitution, K-nearest neighbor, multiple imputation, and tree model imputation; another is missing value reconstruction and element estimation methods for geochemical, logging, or wellbore data; and the third is information extraction and structuring methods for mineral exploration texts and geological reports. Existing research shows that the effectiveness of missing value completion is usually affected by factors such as data structure, missing mechanism, variable type, sample size, and validation method. In spatial or geoscience data scenarios, it is also necessary to consider spatial autocorrelation, sample distribution, and model generalization issues.

[0003] In the field of general missing value completion, existing methods mainly estimate missing terms through statistical relationships between variables. For example, the missForest method proposed by Stekhoven and Bühlmann has shown that random forests can be used for nonparametric missing value imputation of mixed-type data and can handle complex nonlinear relationships and variable interactions. Recent comparative studies have also consistently shown that there is no universally optimal solution among different imputation methods for all data scenarios, and the performance of the methods is closely related to the data size, missing value ratio, and data distribution. It can be seen that although general missing value completion technology provides a methodological foundation that can be learned from, its design goal is mainly for general tabular data and it has not been specifically organized for borehole samples, which have borehole number-depth structural constraints.

[0004] In geochemical or wellbore data scenarios, machine learning has been used to estimate missing element values ​​or reconstruct incomplete records. For example, a 2025 study compared KNN, MICE, miceforest, MIDAS, and various tree models based on XRF data from 14 wells in the Vaca Muerta Formation in Argentina for missing element completion and trace element estimation in specific depth ranges. Such studies demonstrate that missing element values ​​in geoscientific data can be predicted using machine learning. However, they also reflect that existing methods mainly focus on the mapping relationship between numerical fields and compare algorithm accuracy, while the combined use of spatial location of sample segments, semantic description of stratification, and data organization characteristics within mining areas remains insufficient. Especially under conditions of limited sample size and significant geological heterogeneity, the completion process relying solely on numerical input is easily limited by data distribution and sample representativeness.

[0005] Meanwhile, geoscientific data naturally possess spatial correlation. Existing research indicates that adding spatial proxy variables such as coordinates or distance fields to models like random forests can indeed improve the performance of some spatial interpolation tasks. However, this effect is not universally applicable and depends on the modeling objective, the strength of residual spatial autocorrelation, and the sample distribution. Under certain conditions, spatial proxy variables may even be ineffective or have a counterproductive effect. This suggests that spatial information has significant value in geoscientific modeling, but its use should not be limited to simply adding coordinate variables. Instead, it should be organized in a targeted manner based on specific data objects and application tasks. For borehole sampling and testing data, information such as sample length, sample midpoint depth, relative end borehole depth, and borehole opening position are all spatial constraints that are more closely aligned with sample-level scenarios, and existing general methods lack specific procedures for this.

[0006] Furthermore, research in the field of mineral exploration has demonstrated that geological texts are not limited to manual interpretation; they can be transformed into computable mineral exploration criteria or structured knowledge through text mining and natural language processing. A 2024 study on the Xiyu diamond deposit showed that geological text big data can be used to construct exploration criteria through text mining and further combined with random forest models for 3D mineral exploration prediction, achieving good application results. This indicates that the semantics of geological texts have a realistic basis for entering the modeling process. However, existing related work mainly serves the extraction of mineral exploration criteria, knowledge organization, or prediction assistance, and has not yet formed a sample-level processing method for completing missing values ​​in borehole sampling test tables. In particular, the semantics of geological texts, element synergy, and sample spatial location have not been organized together as constraints for completing missing values.

[0007] In summary, based on the closest existing technologies, current machine learning imputation methods typically treat the data to be processed as ordinary tabular variables for prediction, or in geoscientific scenarios, they mainly estimate missing terms based on existing element values ​​and a small number of spatial proxy variables. Existing geological text mining methods primarily serve mineral exploration criterion extraction, knowledge organization, or prediction assistance. While the aforementioned solutions provide algorithmic foundations, spatial modeling ideas, and text structuring foundations, they still have the following shortcomings: First, they lack a dedicated completion process for borehole sampling and analysis tables, using "bore number—sample number—start and end depth" to jointly determine the sample segment object, making it difficult to reflect the structural characteristics of borehole data organized along the borehole depth; Second, they lack a dedicated process for integrating sampling and analysis tables, The technical approach involves aligning borehole stratification tables, borehole inclination tables, and basic borehole information tables according to sample segments and further organizing them into a complete sample. Thirdly, it lacks a mechanism to jointly organize three types of information—elemental synergy, sample segment spatial location, and geological text semantics—into complete constraints, rather than simply splicing together several fields. Fourthly, existing methods typically emphasize prediction accuracy comparisons and lack a closed-loop processing method that connects with borehole database repair, result source marking, write-back status, and manual review processes. Therefore, it is necessary to propose a multi-source constraint-based method for completing missing values ​​in borehole sampling and testing data to better meet the needs of mineral deposit exploration database repair and subsequent 3D modeling, ore body interpretation, and geoscientific analysis. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a method and related apparatus for completing missing values ​​in borehole sampling and testing data, which can more accurately complete missing values ​​in borehole sampling and testing data.

[0009] To address the aforementioned technical problems, embodiments of the present invention provide a method for completing missing values ​​in borehole sampling and testing data, the method comprising: Extract the sampling test table, borehole basic information table, borehole inclination table and the corresponding layer description information or lithological description information of the sample section from the target borehole database as basic input data, and use the missing values ​​of the content field of at least one continuous target element in the sampling test table as target completion objects. Based on the target completion object, a training sample set is constructed by selecting sample segments with real observation values ​​in the target elements in the target borehole database as candidate samples, and multi-source constraint processing is performed on the candidate samples in the training sample set to form a multi-source constraint training dataset. Based on the target element, a target element missing value completion model is constructed. The multi-source constrained training dataset is used as the input of the target element missing value completion model, and the actual observed values ​​of the target element in the multi-source constrained training dataset are used as the output of the target element missing value completion model for training. The target element missing value completion model is an XGBoost regression model. The basic input data is input into the target element missing value completion model that has been trained and converged. The content field completion value corresponding to the target completion object is input, and the content field completion value is filled into the sampling test table to complete the completion of the missing value of the content field of at least one continuous target element in the sampling test table.

[0010] Optionally, the sampling and analysis table provides the primary key of the sample segment, the sampling start and end depths, and elemental observation information; the borehole basic information table provides the spatial attributes of the borehole opening plane coordinates and the final borehole depth; the borehole inclination table provides borehole attitude information such as borehole angle and borehole dip angle; the stratification description information or lithological description information corresponding to the sample segment provides the underlying category and geological text semantic information corresponding to the sample segment; and the sample segment selection is jointly determined by the borehole number, sample number, sampling start point, and sampling end point; the sample number is used to represent a single sampling record; the borehole number and the sampling start point and sampling end point are used to determine the position of the corresponding sampling record in the borehole and depth range.

[0011] Optionally, the method of constructing a training sample set by selecting sample segments with real observation values ​​from the target elements in the target borehole database based on the target completion object includes: Obtain the target elements in the target completion object, and filter the sample segments with real observation values ​​in the target elements as candidate samples in the target borehole database; Each sample segment in the candidate samples is initially screened according to the following criteria: the primary key information of the sample segment is complete, the target field has real observed values, and it can establish a correspondence with the borehole basic information table, borehole inclination table, stratification description information, or lithology description information. The initial screening of the candidate samples is then obtained. A secondary screening process is performed on the candidate samples after the initial screening. Samples with abnormal data, missing fields, non-standard record formats, non-unique primary keys, misaligned depth ranges, or misalignment with other basic attributes are removed. The candidate sample data after the secondary screening is obtained and used as the training sample set.

[0012] Optionally, the step of performing multi-source constraint processing on candidate samples in the training sample set to form a multi-source constraint training dataset includes: Based on the sampling test table in each candidate sample in the training sample set, the element collaboration relationship features of the corresponding sample segment are constructed and processed to obtain the element collaboration relationship features of each candidate sample in the training sample set. Based on the borehole basic information table in each candidate sample in the training sample set, the spatial location features of the corresponding sample segment are constructed and processed to obtain the sample segment spatial location features of each candidate sample in the training sample set. Based on the hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set, the geological text semantic features of the corresponding sample segments are constructed and processed to obtain the geological text semantic features of each candidate sample in the training sample set. The element collaboration relationship features, sample spatial location features, and geological text semantic features of each candidate sample in the training sample set are associated according to the primary key to form a multi-source constrained training dataset. Each candidate sample in the multi-source constrained training dataset simultaneously contains the element collaboration relationship features representing the combination relationship of elements, the sample spatial location features representing the spatial background, and the geological text semantic features representing the geological environment.

[0013] Optionally, the step of constructing the element-cooperative relationship features of corresponding sample segments based on the sampling test table within each candidate sample in the training sample set to obtain the element-cooperative relationship features of each candidate sample in the training sample set includes: Based on the sampling test table in each candidate sample in the training sample set, the element observation values ​​corresponding to several elements given in the corresponding sample segment are extracted. Based on the element observation values ​​of several elements given in the sample segment, the element coordination relationship feature is established to form the element coordination relationship feature of each candidate sample in the training sample set. The element coordination relationship feature is used to express the combination relationship and numerical correlation between different elements in the sample segment of each candidate sample.

[0014] Optionally, the step of constructing spatial location features of corresponding sample segments based on the borehole basic information table in each candidate sample of the training sample set to obtain the sample segment spatial location features of each candidate sample in the training sample set includes: Based on the borehole basic information table and borehole inclination table in each candidate sample in the training sample set, the spatial attributes of the corresponding sample segment's open plane coordinates, final borehole depth, and borehole attitude are extracted. Based on the spatial attributes of the corresponding sample segment's open plane coordinates, final hole depth, and drilling posture, spatial position features are constructed to obtain the sample segment spatial position features of each candidate sample in the training sample set. The sample segment spatial position features are used to characterize the spatial position difference of the sample segment in the hole and in space, as well as its spatial posture information.

[0015] Optionally, the step of constructing the geological text semantic features of the corresponding sample segments based on the hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set, to obtain the geological text semantic features of each candidate sample in the training sample set, includes: The lithological codes corresponding to the sample segments within each candidate sample in the training sample set are systematically organized, and the main category codes are extracted and grouped into category groups according to prefix or controlled classification rules; simultaneously, The hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set is segmented according to rules in order to extract relevant terms such as mineral name, lithology, color, tectonic phenomenon, occurrence mode, mineralization intensity and alteration characteristics of the sample segments; The relevant terms related to mineral names, lithology, color, tectonic phenomena, occurrence mode, mineralization intensity and alteration characteristics corresponding to the sample segments are semantically vectorized and clustered. Combined with expert review and terminology dictionary or controlled term set, terms that are too general, unclear or difficult to stably participate in modeling are removed to form the core terms of geological text. One-hot encoding is used to encode the category group and the core terms of the geological text into binary features of 0-1, forming the geological text semantic features of each candidate sample in the training sample set.

[0016] In addition, this invention also provides a device for completing missing values ​​in borehole sampling test data, the device comprising: The target completion confirmation module is used to extract the sampling test table, borehole basic information table, borehole inclination table and the corresponding layer description information or lithological description information of the sample section from the target borehole database as basic input data, and to use the missing value of the content field of at least one continuous target element in the sampling test table as the target completion object. Training set construction module: used to construct a training sample set by selecting sample segments with real observation values ​​in the target borehole database based on the target completion object, and to perform multi-source constraint processing on the candidate samples in the training sample set to form a multi-source constraint training dataset; Model training module: used to construct a target element missing value completion model based on the target element, using the multi-source constrained training dataset as the input of the target element missing value completion model, and using the real observed values ​​of the target element in the multi-source constrained training dataset as the output of the target element missing value completion model for training processing, to form a target element missing value completion model that has been trained and converged, and the target element missing value completion model is an XGBoost regression model; The completion module is used to input the basic input data into the target element missing value completion model that has been trained and converged, input the content field completion value corresponding to the target completion object, and fill the content field completion value into the sampling test table to complete the completion of the missing value of the content field of at least one continuous target element in the sampling test table.

[0017] In addition, embodiments of the present invention also provide an electronic device, including a processor and a memory, wherein the processor runs a computer program or code stored in the memory to implement the missing value completion method as described in any of the above.

[0018] In addition, embodiments of the present invention also provide a computer-readable storage medium for storing a computer program or code, which, when executed by a processor, implements the missing value completion method as described above.

[0019] In this embodiment of the invention, borehole samples are used as the basic processing unit. Without relying on external heterogeneous training data, a completion model is constructed using existing complete sample records in the target borehole dataset. Elemental coherence relationships, sample spatial locations, and geological textual semantic information are jointly introduced to complete missing fields. Compared with existing missing value completion methods that mainly rely on numerical statistical relationships or general table prediction processes, this method better adapts to the "bore number—sample number—start and end depth" organization of borehole data, and makes fuller use of existing numerical, spatial, and textual information in the mine borehole data. This improves the relevance and rationality of the missing value completion results to the target mine data structure and geological background. Finally, the generated completion results can be directly written back to the sampling and analysis table. Therefore, it can not only complete the estimation of missing values ​​in the element content field but also directly serve borehole database repair and subsequent 3D modeling, ore body interpretation, and geological analysis. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the missing value completion method for borehole sampling and testing data in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of the missing value completion device for borehole sampling and testing data in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structural composition of the electronic device in an embodiment of the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Example 1, please refer to Figure 1 , Figure 1 This is a flowchart illustrating the missing value completion method for borehole sampling and testing data in an embodiment of the present invention.

[0024] like Figure 1 As shown, a method for completing missing values ​​in borehole sampling test data includes: S101: Extract the sampling test table, borehole basic information table, borehole inclination table and the corresponding layer description information or lithological description information of the sample section from the target borehole database as basic input data, and use the missing values ​​of the content field of at least one continuous target element in the sampling test table as target completion objects. In the specific implementation of this invention, the sampling and analysis table is used to provide the primary key of the sample segment, the sampling start and end depths, and element observation information; the borehole basic information table is used to provide the spatial attributes of the borehole opening plane coordinates and the final borehole depth; the borehole inclination table is used to provide borehole attitude information such as borehole angle and borehole dip angle; the stratification description information or lithological description information corresponding to the sample segment is used to provide the bottom layer category and geological text semantic information corresponding to the sample segment; and the selection of the sample segment is jointly determined by the borehole number, sample number, sampling start point, and sampling end point; the sample number is used to represent a single sampling record; the borehole number and the sampling start point and sampling end point are used to determine the position of the corresponding sampling record in the borehole and depth range.

[0025] Specifically, the first step is to extract sampling and testing tables, borehole basic information tables, borehole inclination tables, and corresponding stratigraphic or lithological description information from the target borehole database as the basic input data required for this embodiment. The sampling and testing tables provide the primary key of the sample section, the sampling start and end depths, and element observation information; the borehole basic information tables provide spatial attributes such as borehole coordinates and final borehole depth; the borehole inclination tables provide borehole attitude information such as azimuth and dip angles; and the stratigraphic or lithological description information provides the stratigraphic category and geological text semantic information corresponding to the sample section.

[0026] Then, the target completion object is determined, that is, the target completion object is the missing value of at least one continuous element content field in the sampling test table; in order to ensure that the completion process is consistent with the borehole data organization method, the sample segment is used as the basic processing unit. The sample segment is preferably determined by the borehole number BHID, sample number ASSAY, sampling start point FROM, and sampling end point TO; among them, ASSAY is used to identify a single sampling record, and BHID and FROM-TO are used to determine the position of the sampling record in the borehole and depth range.

[0027] In this embodiment, taking the Fankou lead-zinc mine borehole dataset as an example, four key elements—Pb, Zn, S, and Ag—are selected as target elements for verification. Among them, Pb and Zn are the main ore-forming elements of the deposit, S is closely related to sulfide mineral assemblage, and Ag, as a common associated element, has certain indicative significance for the strength of mineralization and the relationship of element combinations. The above four types of elements can not only reflect the main characteristics of the Fankou lead-zinc mine sampling and testing data, but are also suitable as the preferred implementation objects of this embodiment. In other mining areas or other borehole databases, this embodiment can also be applied to fill in missing values ​​of other continuous test element fields such as Cu, Au, Fe, and Mn.

[0028] S102: Based on the target completion object, a training sample set is constructed by selecting sample segments with real observation values ​​in the target elements in the target borehole database as candidate samples, and multi-source constraint processing is performed on the candidate samples in the training sample set to form a multi-source constraint training dataset. In the specific implementation of this invention, the method of constructing a training sample set by selecting sample segments with real observation values ​​from the target elements in the target borehole database based on the target completion object includes: obtaining the target elements in the target completion object, and selecting sample segments with real observation values ​​from the target elements in the target borehole database as candidate samples; performing an initial screening process on each sample segment in the candidate samples according to the following conditions: complete sample segment primary key information, target fields with real observation values, and the ability to establish a correspondence with borehole basic information table, borehole inclination table, stratification description information, or lithological description information, to obtain candidate samples after the initial screening; performing a secondary screening process on the candidate samples after the initial screening that have data anomalies, missing fields, non-standard record formats, non-unique sample segment primary keys, incompatible depth intervals, or incompatible alignment with other basic attributes, to obtain candidate sample data after the secondary screening, and using the candidate sample data after the secondary screening as the training sample set.

[0029] Furthermore, the multi-source constraint processing of candidate samples in the training sample set to form a multi-source constraint training dataset includes: constructing the element-cooperative relationship features of corresponding sample segments based on the sampling test table in each candidate sample in the training sample set, to obtain the element-cooperative relationship features of each candidate sample in the training sample set; constructing the spatial location features of corresponding sample segments based on the borehole basic information table in each candidate sample in the training sample set, to obtain the sample segment spatial location features of each candidate sample in the training sample set; constructing the geological text semantic features of corresponding sample segments based on the hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set, to obtain the geological text semantic features of each candidate sample in the training sample set; associating the element-cooperative relationship features, sample segment spatial location features, and geological text semantic features of each candidate sample in the training sample set according to the primary key to form a multi-source constraint training dataset, wherein each candidate sample in the multi-source constraint training dataset simultaneously contains the associated element-cooperative relationship features representing the element combination relationship, the sample segment spatial location features representing the spatial background, and the geological text semantic features representing the geological environment.

[0030] Furthermore, the step of constructing the element coordination relationship features of the corresponding sample segment based on the sampling test table in each candidate sample in the training sample set to obtain the element coordination relationship features of each candidate sample in the training sample set includes: extracting the element observation values ​​corresponding to several elements given in the corresponding sample segment based on the sampling test table in each candidate sample in the training sample set; and performing element coordination relationship feature establishment processing based on the element observation values ​​of several elements given in the sample segment to form the element coordination relationship features of each candidate sample in the training sample set. The element coordination relationship features are used to express the combination relationship and numerical correlation between different elements in the sample segment of each candidate sample.

[0031] Furthermore, the step of constructing spatial position features of corresponding sample segments based on the borehole basic information table in each candidate sample of the training sample set to obtain the sample segment spatial position features of each candidate sample in the training sample set includes: extracting the spatial attributes of the borehole plane coordinates, final borehole depth, and borehole attitude of the corresponding sample segment based on the borehole basic information table and borehole inclination table in each candidate sample of the training sample set; and constructing spatial position features based on the spatial attributes of the borehole plane coordinates, final borehole depth, and borehole attitude of the corresponding sample segment to obtain the sample segment spatial position features of each candidate sample in the training sample set. The sample segment spatial position features are used to characterize the spatial position difference and spatial attitude information of the sample segment in the borehole and in space.

[0032] Furthermore, the step of constructing and processing the geological text semantic features of the corresponding sample segments based on the hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set to obtain the geological text semantic features of each candidate sample in the training sample set includes: organizing the lithological codes corresponding to the sample segments in each candidate sample in the training sample set into a standardized manner, extracting the main category code, and grouping them into category groups according to prefix or controlled classification rules; simultaneously, performing rule-based segmentation on the hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set to extract the geological text semantic features of the sample segments. The relevant terms for mineral names, lithology, color, tectonic phenomena, occurrence mode, mineralization intensity, and alteration characteristics are used to semantically vectorize and cluster the relevant terms for mineral names, lithology, color, tectonic phenomena, occurrence mode, mineralization intensity, and alteration characteristics corresponding to the sample segments. Combined with expert review and terminology dictionaries or controlled term sets, terms that are too general, have unclear meanings, or are difficult to stably participate in modeling are removed to form the core terms of geological text. One-hot encoding is used to encode the category groups and the core terms of geological text into binary features of 0-1 to form the geological text semantic features of each candidate sample in the training sample set.

[0033] Specifically, after determining the target for completion, a comprehensive sample table at the sample segment level is first constructed around the primary key of the sample segment. Using individual sampling records in the sampling and analysis table as the basic unit, information such as elemental content values, LITHCODE (lithological code) and LITHDESC (lithological description) in the borehole stratification table, BGR and DIP in the borehole survey table, and borehole head coordinates and final depth in the borehole basic information table are aligned according to the BHID and FROM-TO depth intervals and uniformly linked to the same sample segment record. For interval-type records in the stratification table, the corresponding stratification code and geological description of the sample segment can be determined based on the inclusion relationship, intersection relationship, or maximum overlap length with the sampling interval. For point-type records in the survey table, the stratification code and geological description of the sample segment can be determined based on the inclusion relationship, intersection relationship, or maximum overlap length with the sampling interval. The azimuth and dip angles are obtained by selecting nearby measuring points, measuring points in the depth segment, or by interpolation to determine the midpoint depth of the sample section. The borehole coordinates and final depth in the basic information table are directly associated using BHID. The element values, spatial background, borehole attitude, and geological description information, which were originally scattered in different business tables, are organized into a comprehensive input object at the same sample scale. This allows the selection of sample sections with real observation values ​​for target elements from the target borehole database as candidate samples for the subsequent construction of the target element missing completion model. Since the missing information and the number of effective samples for different element fields in the sampling and testing table may vary, it is preferable to establish independent training sample sets for different target elements.

[0034] In the process of selecting candidate samples, firstly, sample records with complete primary key information, target fields with real observed values, and effective correspondence with the borehole basic information table, borehole inclination table, and stratification description information are retained; secondly, sample records with severe field missingness, non-standard record format, non-unique primary key, incompatible depth intervals, or inability to match other basic attributes are removed, individually marked, or transferred to the set for review; thirdly, the original primary key, source table, and processing status of the complete sample samples used for model training are retained for subsequent result traceability; through the above processing, a complete sample set of sample segments that can be used for subsequent feature construction and model training is obtained; that is, taking the Fankou lead-zinc mine borehole dataset as an example, independent sample sets are constructed for the four target elements of Pb, Zn, S, and Ag respectively, so that corresponding target element missing value completion models can be built later.

[0035] Then, the observation values ​​of the target elements in the candidate samples are subjected to anomaly identification and screening to reduce the impact of abnormal records and extreme samples on subsequent model training. Since the element content in geological sampling and testing data often has characteristics such as skewed distribution, long tail distribution and local high values, if the original observation values ​​are used directly for modeling, it is easy to cause unstable model fitting or excessive influence of local samples on the overall completion result.

[0036] Outlier identification and processing are preferably performed on the observation values ​​of the target elements. For null values, values ​​below the detection limit, non-numerical records, negative records, obviously erroneous input values, and samples determined to be abnormal after anomaly identification, they can be removed, marked, or uniformly processed according to established rules. Preferably, the element content values ​​can be numerically analyzed first, and abnormal sample segments can be identified in the logarithmic space by combining robust statistical thresholds and upper and lower quantile thresholds. Then, the filtered observation values ​​can be restored to their original dimensions for subsequent modeling. For records below the detection limit, they can be uniformly converted into numerical inputs or excluded from the training samples according to the detection limit label, project rules, or manual review opinions. The purpose is not to weaken the real geological differences in the samples, but to remove records that are obviously affected by input, format conversion, or extreme outliers, thereby improving the stability and usability of the target element training sample set.

[0037] After establishing the training sample set, it is necessary to construct a multi-source constrained training dataset. First, construct the element collaborative relationship features, then construct the sample segment spatial location features, then construct the geological text semantic features, and finally associate the element collaborative relationship features, sample segment spatial location features, and geological text semantic features according to the primary key to form a multi-source constrained training dataset.

[0038] Among them, the element synergy relationship features are used to characterize the symbiotic, associated, and synergistic changes among different chemical elements. In this embodiment, it is the most basic type of input information when completing missing values ​​of target fields. For any target element, the observation values ​​of other chemical elements besides the target element are extracted as input features to express the combination relationship and numerical correlation between different elements within the sample segment. For example, when completing missing values ​​of the Pb field, the observation values ​​of Zn, S, Ag, and other available elements can be extracted as input features; when completing missing values ​​of the Zn field, the observation values ​​of Pb, S, Ag, and other available elements are extracted as input features. For any target element, the model input features do not include the observation value of the target element itself to avoid leakage of target field information. In this way, the estimation process of missing values ​​of target elements no longer relies solely on univariate information, but is inferred based on the existing multi-element synergy relationships within the sample segment. In the preferred embodiment, the element synergy relationship features constitute the basic input part of the target element missing value completion model, and together with the subsequent sample segment spatial location features and geological text semantic features, form a multi-source constrained input dataset.

[0039] In addition to elemental synergies, spatial location features of sample segments are further introduced to characterize the spatial location differences and spatial attitude information of sample segments within and between boreholes. Geological borehole sampling and analysis data are naturally organized along the borehole depth direction. The distribution of the same element in different sample segments is not only affected by the values ​​of other elements, but also closely related to the position, depth range, and spatial attitude of the sample segment within the borehole. Therefore, incorporating spatial location features of sample segments into the completion process helps improve the adaptability of the target field missing value completion results to the actual mineralization environment. The preferred spatial location features of sample segments include: borehole opening plane coordinates, sample segment length, sample segment midpoint depth, sample segment position relative to the final borehole depth, and borehole azimuth and dip angle. Among them, the borehole opening plane coordinates are used to characterize the spatial location of the borehole to which the sample segment belongs in the mining area; the sample segment length, sample segment midpoint depth, and relative final borehole depth position are used to characterize the depth distribution characteristics of the sample segment within the borehole; the azimuth and dip angle are used to characterize the differences in borehole attitude; preferably, the sample segment length can be determined by TO-F. For ROM calculation, the midpoint depth of the sample segment can be calculated using (FROM+TO) / 2, and the relative position of the sample segment to the final hole depth can be calculated using the ratio of the midpoint depth to the final hole depth. When the inclinometer data consists of discrete depth measurement points, the adjacent or interpolated BGR and DIP values ​​can be assigned to the sample segment based on the correspondence between the midpoint depth of the sample segment and the inclinometer measurement points. Continuous variables such as borehole plane coordinates, midpoint depth of the sample segment, and relative final hole depth can be centered, standardized, or normalized according to the data dimensions and sample distribution to characterize the relative planar position and borehole position of the sample segment in the study area and to reduce the influence of different dimensions on the model training process. For periodic angular variables such as azimuth, further trigonometric function transformations such as sine and cosine can be performed to avoid the problem of 0° and 360° being close but with large numerical distances. For dip angle variables, they can be directly used or transformed using corresponding trigonometric functions according to data expression habits to convert borehole attitude information into continuous features suitable for model input.

[0040] Textual data such as borehole stratigraphic descriptions, lithological logging, and mineralization descriptions contain a wealth of geological semantic information that reflects mineral assemblage, structural features, mineralization phenomena, and surrounding rock conditions. Existing missing value completion methods often struggle to directly utilize this unstructured textual information. This embodiment transforms unstructured descriptions into structured semantic features that can participate in modeling by standardizing, semantically merging, terminology filtering, and structurally encoding the stratigraphic codes and geological texts corresponding to the sample segments. Specifically, the LITHCODE (lithological code) corresponding to the sample segment is first standardized and organized, the main category code is extracted and merged into category groups according to prefix or controlled classification rules, and then one-hot encoding is used to form sample segment-level category features. At the same time, the LITHDESC (lithological description) corresponding to the sample segment is segmented according to rules to extract phrases such as mineral name, lithology, color, structural phenomena, occurrence mode, mineralization intensity, and alteration characteristics. This rule-based segmentation is not a general mechanical word segmentation, but rather combines geological description characteristics to split, merge, and standardize parallel phrases, modifiers, compound expressions, and synonymous expressions, while maintaining the correspondence between the segmentation results and the sample segment number, borehole number, and depth interval. In specific implementation, candidate geological phrases can be semantically vectorized and clustered, and combined with expert review, terminology dictionaries, or controlled terminology sets, terms that are too general, have unclear meanings, or are difficult to stably participate in modeling are eliminated, while core terms that can reflect geological backgrounds such as mineralization, structure, lithology, and alteration are retained. Subsequently, whether a sample segment contains the corresponding category or term is encoded as a 0-1 binary feature, or discrete structured features can be formed by term category, frequency of occurrence, or expert grouping. By introducing the semantic features of geological text into the missing value completion process, the target field completion result can be simultaneously constrained by numerical information, spatial information, and geological semantic information, thereby improving the rationality of completion and the consistency of geological interpretation under limited sample size conditions.

[0041] Finally, a multi-source constrained training dataset is formed. After constructing the element collaborative relationship features, sample spatial location features, and geological text semantic features, the above three types of features are associated and integrated according to the sample primary key to form a multi-source constrained training dataset for filling in missing values ​​of target elements. Each sample record in the multi-source constrained training dataset corresponds to a sample segment and simultaneously contains numerical features for characterizing element combination relationships, spatial location features for characterizing spatial background, and text semantic features for characterizing geological environment.

[0042] Different levels of input datasets can be constructed based on the order of feature introduction. For example, a basic dataset containing only element cooperative relationship features can be constructed first. Spatial location features of sample segments can be added to the basic dataset to form a spatial augmentation dataset. Then, stratigraphic category features and geological text semantic features can be added to the spatial augmentation dataset to form a multi-source constrained training dataset. For text augmentation datasets, different sizes of text feature sets can be set according to the frequency of term occurrence or expert screening results to compare the impact of different sizes of text information on the completion effect. Through hierarchical construction, the target element missing value completion model can be trained and applied under different information constraints, and it is also convenient to flexibly select the input structure according to the sample conditions of different target elements.

[0043] S103: Construct a target element missing value completion model based on the target element, using the multi-source constrained training dataset as the input of the target element missing value completion model, and using the real observed values ​​of the target element in the multi-source constrained training dataset as the output of the target element missing value completion model for training, to form a target element missing value completion model that has converged in training, and the target element missing value completion model is an XGBoost regression model; In the specific implementation of this invention, a supervised learning model for missing value completion is established for each target element. A multi-source constraint training dataset is used as the model input, and the actual observed value of the corresponding target element is used as the model output, thereby learning the mapping relationship between multi-source constraint features and target element values. Since the distribution characteristics, sample size, and influencing factors of different elements vary, this embodiment preferably adopts a single-objective modeling approach for different target elements, i.e., predicting only one target element at a time, and using the content values ​​of other elements besides the target element, spatial location features of the sample segment, and semantic features of the geological text as input. For the input features used for training, validation, and actual completion, none of them contain the observed value of the target element itself or derived variables directly calculated from the target element, in order to avoid leakage of target field information and ensure that the model can be implemented in real missing value scenarios.

[0044] The supervised learning model for completing missing values ​​of target elements preferably employs a nonlinear regression model suitable for finite sample tabular data. In a preferred embodiment, a tree model regressor, a gradient boosting tree model, or other supervised learning models capable of handling nonlinear relationships and feature interactions can be used. More preferably, an XGBoost regression model can be used as the model for completing missing values ​​of target elements. During model training, the multi-source constraint features in the complete sample segment are used as input, and the corresponding observation values ​​of the target elements are used as output to obtain the model for completing missing values ​​of the target elements. To verify the model's effectiveness in utilizing different information sources, the model performance of the basic dataset, the spatial augmentation dataset, and the text augmentation dataset can be compared under the same samples and the same occlusion position. In practical applications, appropriate combinations of input features can also be selected based on the sample size, feature completeness, and verification effect of different target elements.

[0045] S104: Input the basic input data into the target element missing value completion model that has been trained and converged, input the content field completion value corresponding to the target completion object, and fill the content field completion value into the sampling test table to complete the completion of the missing value of the content field of at least one continuous target element in the sampling test table.

[0046] In the specific implementation of this invention, target field completion is performed on the target completion object. After the target element missing value completion model is established and trained, the sample segments (basic input data) with missing target fields in the to-be-completed set are input into the corresponding trained and converged target element missing value completion model to obtain the predicted value of the sample segment on the target field, that is, the content field completion value corresponding to the target completion object; the predicted value is used as the completion result of the target field. If multiple target element fields are missing in the same segment, the corresponding model can be called to perform completion for different target elements; the missing element content field values ​​that cannot be directly recovered by deterministic rules or original evidence can be transformed into a completion output with numerical results; the output results are generated with the sample segment as the basic unit, thus maintaining consistency with the organization of the sampling test table.

[0047] After obtaining the content field completion value corresponding to the target completion object, the content field completion value corresponding to the target completion object is associated with the primary key information of the corresponding sample segment to generate a completion result table for database repair. The completion result table preferably includes at least: borehole number BHID, sample number ASSAY, sampling start point FROM, sampling end point TO, target field name, original field status, completion result value, completion method marker, source marker, write-back status, and necessary manual review status information. The completion method marker is used to distinguish different processing sources such as model completion, rule repair, or evidence backfilling; the source marker is used to record the data table, feature source, or model version involved in the completion; and the write-back status and review status are used to distinguish records that have been written back, have not been written back, and are awaiting manual confirmation. Subsequently, according to the sample segment... The primary key writes the completion result back to the target field position of the corresponding record in the sampling test table, forming the completed sampling test dataset. During the write-back process, the original state information is not directly deleted or overwritten; instead, the original missing identifier, completed value, processing method, source marker, and review status are retained. Records with insufficient model confidence, severe input feature loss, inability to uniquely match the primary key of the sample segment, predicted values ​​significantly exceeding a reasonable range, or unsuitable for automatic modification are not directly written back but are retained as records awaiting review. By simultaneously retaining the original field state, completion result, completion method marker, source marker, and write-back status, the write-back step not only achieves the estimation and completion of missing values ​​in the target field but also enables the completion result to directly serve borehole database repair and subsequent 3D modeling, ore body interpretation, and geological analysis.

[0048] To verify the feasibility and effectiveness of this embodiment, randomly occluded portions of the real observation values ​​in the complete sample segment can be used as simulated missing values. After model training, the difference between the predicted values ​​and the original real values ​​can be compared to evaluate the accuracy and stability of the completion process. Furthermore, mean imputation, median imputation, multiple linear regression imputation, K-nearest neighbor imputation, random forest regression, or other models can be set at the same occlusion location as comparison methods, and RMSE, MAE, R², and other indicators can be used for evaluation. Comparisons can also be carried out under different input conditions such as element-based collaborative features only, element-based collaborative features plus spatial location features, and element-based collaborative features plus spatial location features plus geological text semantic features to illustrate the contribution of multi-source constraint information to the completion results. If necessary, feature contribution interpretation methods can be combined to analyze the relative roles of element-based collaborative features, spatial location features, and geological text semantic features in the target element completion process. It should be noted that the above verification steps are effect verification means in the preferred embodiment.

[0049] In this embodiment of the invention, borehole samples are used as the basic processing unit. Without relying on external heterogeneous training data, a completion model is constructed using existing complete sample records in the target borehole dataset. Elemental coherence relationships, sample spatial locations, and geological textual semantic information are jointly introduced to complete missing fields. Compared with existing missing value completion methods that mainly rely on numerical statistical relationships or general table prediction processes, this method better adapts to the "bore number—sample number—start and end depth" organization of borehole data, and makes fuller use of existing numerical, spatial, and textual information in the mine borehole data. This improves the relevance and rationality of the missing value completion results to the target mine data structure and geological background. Finally, the generated completion results can be directly written back to the sampling and analysis table. Therefore, it can not only complete the estimation of missing values ​​in the element content field but also directly serve borehole database repair and subsequent 3D modeling, ore body interpretation, and geological analysis.

[0050] Example 2, please refer to Figure 2 , Figure 2 This is a schematic diagram of the structure of the missing value completion device for borehole sampling and testing data in an embodiment of the present invention.

[0051] like Figure 2 As shown, a device for completing missing values ​​in borehole sampling test data includes: The target completion confirmation module 201 is used to extract the sampling test table, borehole basic information table, borehole inclination table and the corresponding layer description information or lithological description information of the sample section from the target borehole database as basic input data, and to use the missing value of the content field of at least one continuous target element in the sampling test table as the target completion object. In the specific implementation of this invention, the sampling and analysis table is used to provide the primary key of the sample segment, the sampling start and end depths, and element observation information; the borehole basic information table is used to provide the spatial attributes of the borehole opening plane coordinates and the final borehole depth; the borehole inclination table is used to provide borehole attitude information such as borehole angle and borehole dip angle; the stratification description information or lithological description information corresponding to the sample segment is used to provide the bottom layer category and geological text semantic information corresponding to the sample segment; and the selection of the sample segment is jointly determined by the borehole number, sample number, sampling start point, and sampling end point; the sample number is used to represent a single sampling record; the borehole number and the sampling start point and sampling end point are used to determine the position of the corresponding sampling record in the borehole and depth range.

[0052] Specifically, the first step is to extract sampling and testing tables, borehole basic information tables, borehole inclination tables, and corresponding stratigraphic or lithological description information from the target borehole database as the basic input data required for this embodiment. The sampling and testing tables provide the primary key of the sample section, the sampling start and end depths, and element observation information; the borehole basic information tables provide spatial attributes such as borehole coordinates and final borehole depth; the borehole inclination tables provide borehole attitude information such as azimuth and dip angles; and the stratigraphic or lithological description information provides the stratigraphic category and geological text semantic information corresponding to the sample section.

[0053] Then, the target completion object is determined, that is, the target completion object is the missing value of at least one continuous element content field in the sampling test table; in order to ensure that the completion process is consistent with the borehole data organization method, the sample segment is used as the basic processing unit. The sample segment is preferably determined by the borehole number BHID, sample number ASSAY, sampling start point FROM, and sampling end point TO; among them, ASSAY is used to identify a single sampling record, and BHID and FROM-TO are used to determine the position of the sampling record in the borehole and depth range.

[0054] In this embodiment, taking the Fankou lead-zinc mine borehole dataset as an example, four key elements—Pb, Zn, S, and Ag—are selected as target elements for verification. Among them, Pb and Zn are the main ore-forming elements of the deposit, S is closely related to sulfide mineral assemblage, and Ag, as a common associated element, has certain indicative significance for the strength of mineralization and the relationship of element combinations. The above four types of elements can not only reflect the main characteristics of the Fankou lead-zinc mine sampling and testing data, but are also suitable as the preferred implementation objects of this embodiment. In other mining areas or other borehole databases, this embodiment can also be applied to fill in missing values ​​of other continuous test element fields such as Cu, Au, Fe, and Mn.

[0055] Training set construction module 202: is used to construct a training sample set by selecting sample segments with real observation values ​​in the target elements in the target borehole database based on the target completion object, and to perform multi-source constraint processing on the candidate samples in the training sample set to form a multi-source constraint training dataset; In the specific implementation of this invention, the method of constructing a training sample set by selecting sample segments with real observation values ​​from the target elements in the target borehole database based on the target completion object includes: obtaining the target elements in the target completion object, and selecting sample segments with real observation values ​​from the target elements in the target borehole database as candidate samples; performing an initial screening process on each sample segment in the candidate samples according to the following conditions: complete sample segment primary key information, target fields with real observation values, and the ability to establish a correspondence with borehole basic information table, borehole inclination table, stratification description information, or lithological description information, to obtain candidate samples after the initial screening; performing a secondary screening process on the candidate samples after the initial screening that have data anomalies, missing fields, non-standard record formats, non-unique sample segment primary keys, incompatible depth intervals, or incompatible alignment with other basic attributes, to obtain candidate sample data after the secondary screening, and using the candidate sample data after the secondary screening as the training sample set.

[0056] Furthermore, the multi-source constraint processing of candidate samples in the training sample set to form a multi-source constraint training dataset includes: constructing the element-cooperative relationship features of corresponding sample segments based on the sampling test table in each candidate sample in the training sample set, to obtain the element-cooperative relationship features of each candidate sample in the training sample set; constructing the spatial location features of corresponding sample segments based on the borehole basic information table in each candidate sample in the training sample set, to obtain the sample segment spatial location features of each candidate sample in the training sample set; constructing the geological text semantic features of corresponding sample segments based on the hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set, to obtain the geological text semantic features of each candidate sample in the training sample set; associating the element-cooperative relationship features, sample segment spatial location features, and geological text semantic features of each candidate sample in the training sample set according to the primary key to form a multi-source constraint training dataset, wherein each candidate sample in the multi-source constraint training dataset simultaneously contains the associated element-cooperative relationship features representing the element combination relationship, the sample segment spatial location features representing the spatial background, and the geological text semantic features representing the geological environment.

[0057] Furthermore, the step of constructing the element coordination relationship features of the corresponding sample segment based on the sampling test table in each candidate sample in the training sample set to obtain the element coordination relationship features of each candidate sample in the training sample set includes: extracting the element observation values ​​corresponding to several elements given in the corresponding sample segment based on the sampling test table in each candidate sample in the training sample set; and performing element coordination relationship feature establishment processing based on the element observation values ​​of several elements given in the sample segment to form the element coordination relationship features of each candidate sample in the training sample set. The element coordination relationship features are used to express the combination relationship and numerical correlation between different elements in the sample segment of each candidate sample.

[0058] Furthermore, the step of constructing spatial position features of corresponding sample segments based on the borehole basic information table in each candidate sample of the training sample set to obtain the sample segment spatial position features of each candidate sample in the training sample set includes: extracting the spatial attributes of the borehole plane coordinates, final borehole depth, and borehole attitude of the corresponding sample segment based on the borehole basic information table and borehole inclination table in each candidate sample of the training sample set; and constructing spatial position features based on the spatial attributes of the borehole plane coordinates, final borehole depth, and borehole attitude of the corresponding sample segment to obtain the sample segment spatial position features of each candidate sample in the training sample set. The sample segment spatial position features are used to characterize the spatial position difference and spatial attitude information of the sample segment in the borehole and in space.

[0059] Furthermore, the step of constructing and processing the geological text semantic features of the corresponding sample segments based on the hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set to obtain the geological text semantic features of each candidate sample in the training sample set includes: organizing the lithological codes corresponding to the sample segments in each candidate sample in the training sample set into a standardized manner, extracting the main category code, and grouping them into category groups according to prefix or controlled classification rules; simultaneously, performing rule-based segmentation on the hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set to extract the geological text semantic features of the sample segments. The relevant terms for mineral names, lithology, color, tectonic phenomena, occurrence mode, mineralization intensity, and alteration characteristics are used to semantically vectorize and cluster the relevant terms for mineral names, lithology, color, tectonic phenomena, occurrence mode, mineralization intensity, and alteration characteristics corresponding to the sample segments. Combined with expert review and terminology dictionaries or controlled term sets, terms that are too general, have unclear meanings, or are difficult to stably participate in modeling are removed to form the core terms of geological text. One-hot encoding is used to encode the category groups and the core terms of geological text into binary features of 0-1 to form the geological text semantic features of each candidate sample in the training sample set.

[0060] Specifically, after determining the target for completion, a comprehensive sample table at the sample segment level is first constructed around the primary key of the sample segment. Using individual sampling records in the sampling and analysis table as the basic unit, information such as elemental content values, LITHCODE (lithological code) and LITHDESC (lithological description) in the borehole stratification table, BGR and DIP in the borehole survey table, and borehole head coordinates and final depth in the borehole basic information table are aligned according to the BHID and FROM-TO depth intervals and uniformly linked to the same sample segment record. For interval-type records in the stratification table, the corresponding stratification code and geological description of the sample segment can be determined based on the inclusion relationship, intersection relationship, or maximum overlap length with the sampling interval. For point-type records in the survey table, the stratification code and geological description of the sample segment can be determined based on the inclusion relationship, intersection relationship, or maximum overlap length with the sampling interval. The azimuth and dip angles are obtained by selecting nearby measuring points, measuring points in the depth segment, or by interpolation to determine the midpoint depth of the sample section. The borehole coordinates and final depth in the basic information table are directly associated using BHID. The element values, spatial background, borehole attitude, and geological description information, which were originally scattered in different business tables, are organized into a comprehensive input object at the same sample scale. This allows the selection of sample sections with real observation values ​​for target elements from the target borehole database as candidate samples for the subsequent construction of the target element missing completion model. Since the missing information and the number of effective samples for different element fields in the sampling and testing table may vary, it is preferable to establish independent training sample sets for different target elements.

[0061] In the process of selecting candidate samples, firstly, sample records with complete primary key information, target fields with real observed values, and effective correspondence with the borehole basic information table, borehole inclination table, and stratification description information are retained; secondly, sample records with severe field missingness, non-standard record format, non-unique primary key, incompatible depth intervals, or inability to match other basic attributes are removed, individually marked, or transferred to the set for review; thirdly, the original primary key, source table, and processing status of the complete sample samples used for model training are retained for subsequent result traceability; through the above processing, a complete sample set of sample segments that can be used for subsequent feature construction and model training is obtained; that is, taking the Fankou lead-zinc mine borehole dataset as an example, independent sample sets are constructed for the four target elements of Pb, Zn, S, and Ag respectively, so that corresponding target element missing value completion models can be built later.

[0062] Then, the observation values ​​of the target elements in the candidate samples are subjected to anomaly identification and screening to reduce the impact of abnormal records and extreme samples on subsequent model training. Since the element content in geological sampling and testing data often has characteristics such as skewed distribution, long tail distribution and local high values, if the original observation values ​​are used directly for modeling, it is easy to cause unstable model fitting or excessive influence of local samples on the overall completion result.

[0063] Outlier identification and processing are preferably performed on the observation values ​​of the target elements. For null values, values ​​below the detection limit, non-numerical records, negative records, obviously erroneous input values, and samples determined to be abnormal after anomaly identification, they can be removed, marked, or uniformly processed according to established rules. Preferably, the element content values ​​can be numerically analyzed first, and abnormal sample segments can be identified in the logarithmic space by combining robust statistical thresholds and upper and lower quantile thresholds. Then, the filtered observation values ​​can be restored to their original dimensions for subsequent modeling. For records below the detection limit, they can be uniformly converted into numerical inputs or excluded from the training samples according to the detection limit label, project rules, or manual review opinions. The purpose is not to weaken the real geological differences in the samples, but to remove records that are obviously affected by input, format conversion, or extreme outliers, thereby improving the stability and usability of the target element training sample set.

[0064] After establishing the training sample set, it is necessary to construct a multi-source constrained training dataset. First, construct the element collaborative relationship features, then construct the sample segment spatial location features, then construct the geological text semantic features, and finally associate the element collaborative relationship features, sample segment spatial location features, and geological text semantic features according to the primary key to form a multi-source constrained training dataset.

[0065] Among them, the element synergy relationship features are used to characterize the symbiotic, associated, and synergistic changes among different chemical elements. In this embodiment, it is the most basic type of input information when completing missing values ​​of target fields. For any target element, the observation values ​​of other chemical elements besides the target element are extracted as input features to express the combination relationship and numerical correlation between different elements within the sample segment. For example, when completing missing values ​​of the Pb field, the observation values ​​of Zn, S, Ag, and other available elements can be extracted as input features; when completing missing values ​​of the Zn field, the observation values ​​of Pb, S, Ag, and other available elements are extracted as input features. For any target element, the model input features do not include the observation value of the target element itself to avoid leakage of target field information. In this way, the estimation process of missing values ​​of target elements no longer relies solely on univariate information, but is inferred based on the existing multi-element synergy relationships within the sample segment. In the preferred embodiment, the element synergy relationship features constitute the basic input part of the target element missing value completion model, and together with the subsequent sample segment spatial location features and geological text semantic features, form a multi-source constrained input dataset.

[0066] In addition to elemental synergies, spatial location features of sample segments are further introduced to characterize the spatial location differences and spatial attitude information of sample segments within and between boreholes. Geological borehole sampling and analysis data are naturally organized along the borehole depth direction. The distribution of the same element in different sample segments is not only affected by the values ​​of other elements, but also closely related to the position, depth range, and spatial attitude of the sample segment within the borehole. Therefore, incorporating spatial location features of sample segments into the completion process helps improve the adaptability of the target field missing value completion results to the actual mineralization environment. The preferred spatial location features of sample segments include: borehole opening plane coordinates, sample segment length, sample segment midpoint depth, sample segment position relative to the final borehole depth, and borehole azimuth and dip angle. Among them, the borehole opening plane coordinates are used to characterize the spatial location of the borehole to which the sample segment belongs in the mining area; the sample segment length, sample segment midpoint depth, and relative final borehole depth position are used to characterize the depth distribution characteristics of the sample segment within the borehole; the azimuth and dip angle are used to characterize the differences in borehole attitude; preferably, the sample segment length can be determined by TO-F. For ROM calculation, the midpoint depth of the sample segment can be calculated using (FROM+TO) / 2, and the relative position of the sample segment to the final hole depth can be calculated using the ratio of the midpoint depth to the final hole depth. When the inclinometer data consists of discrete depth measurement points, the adjacent or interpolated BGR and DIP values ​​can be assigned to the sample segment based on the correspondence between the midpoint depth of the sample segment and the inclinometer measurement points. Continuous variables such as borehole plane coordinates, midpoint depth of the sample segment, and relative final hole depth can be centered, standardized, or normalized according to the data dimensions and sample distribution to characterize the relative planar position and borehole position of the sample segment in the study area and to reduce the influence of different dimensions on the model training process. For periodic angular variables such as azimuth, further trigonometric function transformations such as sine and cosine can be performed to avoid the problem of 0° and 360° being close but with large numerical distances. For dip angle variables, they can be directly used or transformed using corresponding trigonometric functions according to data expression habits to convert borehole attitude information into continuous features suitable for model input.

[0067] Textual data such as borehole stratigraphic descriptions, lithological logging, and mineralization descriptions contain a wealth of geological semantic information that reflects mineral assemblage, structural features, mineralization phenomena, and surrounding rock conditions. Existing missing value completion methods often struggle to directly utilize this unstructured textual information. This embodiment transforms unstructured descriptions into structured semantic features that can participate in modeling by standardizing, semantically merging, terminology filtering, and structurally encoding the stratigraphic codes and geological texts corresponding to the sample segments. Specifically, the LITHCODE (lithological code) corresponding to the sample segment is first standardized and organized, the main category code is extracted and merged into category groups according to prefix or controlled classification rules, and then one-hot encoding is used to form sample segment-level category features. At the same time, the LITHDESC (lithological description) corresponding to the sample segment is segmented according to rules to extract phrases such as mineral name, lithology, color, structural phenomena, occurrence mode, mineralization intensity, and alteration characteristics. This rule-based segmentation is not a general mechanical word segmentation, but rather combines geological description characteristics to split, merge, and standardize parallel phrases, modifiers, compound expressions, and synonymous expressions, while maintaining the correspondence between the segmentation results and the sample segment number, borehole number, and depth interval. In specific implementation, candidate geological phrases can be semantically vectorized and clustered, and combined with expert review, terminology dictionaries, or controlled terminology sets, terms that are too general, have unclear meanings, or are difficult to stably participate in modeling are eliminated, while core terms that can reflect geological backgrounds such as mineralization, structure, lithology, and alteration are retained. Subsequently, whether a sample segment contains the corresponding category or term is encoded as a 0-1 binary feature, or discrete structured features can be formed by term category, frequency of occurrence, or expert grouping. By introducing the semantic features of geological text into the missing value completion process, the target field completion result can be simultaneously constrained by numerical information, spatial information, and geological semantic information, thereby improving the rationality of completion and the consistency of geological interpretation under limited sample size conditions.

[0068] Finally, a multi-source constrained training dataset is formed. After constructing the element collaborative relationship features, sample spatial location features, and geological text semantic features, the above three types of features are associated and integrated according to the sample primary key to form a multi-source constrained training dataset for filling in missing values ​​of target elements. Each sample record in the multi-source constrained training dataset corresponds to a sample segment and simultaneously contains numerical features for characterizing element combination relationships, spatial location features for characterizing spatial background, and text semantic features for characterizing geological environment.

[0069] Different levels of input datasets can be constructed based on the order of feature introduction. For example, a basic dataset containing only element cooperative relationship features can be constructed first. Spatial location features of sample segments can be added to the basic dataset to form a spatial augmentation dataset. Then, stratigraphic category features and geological text semantic features can be added to the spatial augmentation dataset to form a multi-source constrained training dataset. For text augmentation datasets, different sizes of text feature sets can be set according to the frequency of term occurrence or expert screening results to compare the impact of different sizes of text information on the completion effect. Through hierarchical construction, the target element missing value completion model can be trained and applied under different information constraints, and it is also convenient to flexibly select the input structure according to the sample conditions of different target elements.

[0070] Model training module 203: is used to construct a target element missing value completion model based on the target element, using the multi-source constrained training dataset as the input of the target element missing value completion model, and using the real observed values ​​of the target element in the multi-source constrained training dataset as the output of the target element missing value completion model for training processing, to form a target element missing value completion model that has been trained and converged, wherein the target element missing value completion model is an XGBoost regression model; In the specific implementation of this invention, a supervised learning model for missing value completion is established for each target element. A multi-source constraint training dataset is used as the model input, and the actual observed value of the corresponding target element is used as the model output, thereby learning the mapping relationship between multi-source constraint features and target element values. Since the distribution characteristics, sample size, and influencing factors of different elements vary, this embodiment preferably adopts a single-objective modeling approach for different target elements, i.e., predicting only one target element at a time, and using the content values ​​of other elements besides the target element, spatial location features of the sample segment, and semantic features of the geological text as input. For the input features used for training, validation, and actual completion, none of them contain the observed value of the target element itself or derived variables directly calculated from the target element, in order to avoid leakage of target field information and ensure that the model can be implemented in real missing value scenarios.

[0071] The supervised learning model for completing missing values ​​of target elements preferably employs a nonlinear regression model suitable for finite sample tabular data. In a preferred embodiment, a tree model regressor, a gradient boosting tree model, or other supervised learning models capable of handling nonlinear relationships and feature interactions can be used. More preferably, an XGBoost regression model can be used as the model for completing missing values ​​of target elements. During model training, the multi-source constraint features in the complete sample segment are used as input, and the corresponding observation values ​​of the target elements are used as output to obtain the model for completing missing values ​​of the target elements. To verify the model's effectiveness in utilizing different information sources, the model performance of the basic dataset, the spatial augmentation dataset, and the text augmentation dataset can be compared under the same samples and the same occlusion position. In practical applications, appropriate combinations of input features can also be selected based on the sample size, feature completeness, and verification effect of different target elements.

[0072] Indeed, the completion module 204 is used to input the basic input data into the target element missing value completion model that has been trained and converged, input the content field completion value corresponding to the target completion object, and fill the content field completion value into the sampling test table to complete the completion of the missing value of the content field of at least one continuous target element in the sampling test table.

[0073] In the specific implementation of this invention, target field completion is performed on the target completion object. After the target element missing value completion model is established and trained, the sample segments (basic input data) with missing target fields in the to-be-completed set are input into the corresponding trained and converged target element missing value completion model to obtain the predicted value of the sample segment on the target field, that is, the content field completion value corresponding to the target completion object; the predicted value is used as the completion result of the target field. If multiple target element fields are missing in the same segment, the corresponding model can be called to perform completion for different target elements; the missing element content field values ​​that cannot be directly recovered by deterministic rules or original evidence can be transformed into a completion output with numerical results; the output results are generated with the sample segment as the basic unit, thus maintaining consistency with the organization of the sampling test table.

[0074] After obtaining the content field completion value corresponding to the target completion object, the content field completion value corresponding to the target completion object is associated with the primary key information of the corresponding sample segment to generate a completion result table for database repair. The completion result table preferably includes at least: borehole number BHID, sample number ASSAY, sampling start point FROM, sampling end point TO, target field name, original field status, completion result value, completion method marker, source marker, write-back status, and necessary manual review status information. The completion method marker is used to distinguish different processing sources such as model completion, rule repair, or evidence backfilling; the source marker is used to record the data table, feature source, or model version involved in the completion; and the write-back status and review status are used to distinguish records that have been written back, have not been written back, and are awaiting manual confirmation. Subsequently, according to the sample segment... The primary key writes the completion result back to the target field position of the corresponding record in the sampling test table, forming the completed sampling test dataset. During the write-back process, the original state information is not directly deleted or overwritten; instead, the original missing identifier, completed value, processing method, source marker, and review status are retained. Records with insufficient model confidence, severe input feature loss, inability to uniquely match the primary key of the sample segment, predicted values ​​significantly exceeding a reasonable range, or unsuitable for automatic modification are not directly written back but are retained as records awaiting review. By simultaneously retaining the original field state, completion result, completion method marker, source marker, and write-back status, the write-back step not only achieves the estimation and completion of missing values ​​in the target field but also enables the completion result to directly serve borehole database repair and subsequent 3D modeling, ore body interpretation, and geological analysis.

[0075] To verify the feasibility and effectiveness of this embodiment, randomly occluded portions of the real observation values ​​in the complete sample segment can be used as simulated missing values. After model training, the difference between the predicted values ​​and the original real values ​​can be compared to evaluate the accuracy and stability of the completion process. Furthermore, mean imputation, median imputation, multiple linear regression imputation, K-nearest neighbor imputation, random forest regression, or other models can be set at the same occlusion location as comparison methods, and RMSE, MAE, R², and other indicators can be used for evaluation. Comparisons can also be carried out under different input conditions such as element-based collaborative features only, element-based collaborative features plus spatial location features, and element-based collaborative features plus spatial location features plus geological text semantic features to illustrate the contribution of multi-source constraint information to the completion results. If necessary, feature contribution interpretation methods can be combined to analyze the relative roles of element-based collaborative features, spatial location features, and geological text semantic features in the target element completion process. It should be noted that the above verification steps are effect verification means in the preferred embodiment.

[0076] In this embodiment of the invention, borehole samples are used as the basic processing unit. Without relying on external heterogeneous training data, a completion model is constructed using existing complete sample records in the target borehole dataset. Elemental coherence relationships, sample spatial locations, and geological textual semantic information are jointly introduced to complete missing fields. Compared with existing missing value completion methods that mainly rely on numerical statistical relationships or general table prediction processes, this method better adapts to the "bore number—sample number—start and end depth" organization of borehole data, and makes fuller use of existing numerical, spatial, and textual information in the mine borehole data. This improves the relevance and rationality of the missing value completion results to the target mine data structure and geological background. Finally, the generated completion results can be directly written back to the sampling and analysis table. Therefore, it can not only complete the estimation of missing values ​​in the element content field but also directly serve borehole database repair and subsequent 3D modeling, ore body interpretation, and geological analysis.

[0077] This invention provides a computer-readable storage medium storing a computer program. When executed by a processor, this program implements the missing value completion method of any of the above embodiments. The computer-readable storage medium includes, but is not limited to, any type of disk (including floppy disk, hard disk, optical disk, CD-ROM, and magneto-optical disk), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards, or optical cards. In other words, the storage device includes any medium that stores or transmits information in a readable form by a device (e.g., a computer, a mobile phone), and can be a read-only memory, a disk, or an optical disk, etc.

[0078] This invention also provides a computer application running on a computer, which is used to execute the missing value completion method of any of the above embodiments.

[0079] also, Figure 3 This is a schematic diagram of the structural composition of the electronic device in an embodiment of the present invention.

[0080] This invention also provides an electronic device, such as... Figure 3 As shown. The electronic device includes a processor 302, a memory 303, an input unit 304, and a display unit 305, among other devices. Those skilled in the art will understand that... Figure 3The structural components of the illustrated electronic device do not constitute a limitation on all devices and may include more or fewer components than illustrated, or combine certain components. Memory 303 can be used to store application program 301 and various functional modules. Processor 302 runs application program 301 stored in memory 303, thereby performing various functional applications and data processing of the device. Memory can be internal memory or external memory, or both. Internal memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, or random access memory. External memory may include hard disks, floppy disks, ZIP disks, USB flash drives, magnetic tapes, etc. The memory disclosed in this invention includes, but is not limited to, these types of memory. The memory disclosed in this invention is only an example and not a limitation.

[0081] Input unit 304 is used to receive signal input and user-input keywords. Input unit 304 may include a touch panel and other input devices. The touch panel can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel) and drive the corresponding connection device according to a pre-set program; other input devices may include, but are not limited to, one or more of physical keyboards, function keys (such as play control buttons, power buttons, etc.), trackballs, mice, joysticks, etc. Display unit 305 can be used to display user-input information or information provided to the user, as well as various menus of the terminal device. Display unit 305 may be in the form of a liquid crystal display, organic light-emitting diode, etc. Processor 302 is the control center of the terminal device, connecting various parts of the entire device through various interfaces and lines, and performing various functions and processing data by running or executing software programs and / or modules stored in memory 303, and calling data stored in memory.

[0082] As one embodiment, the electronic device includes: one or more processors 302, a memory 303, and one or more application programs 301, wherein the one or more application programs 301 are stored in the memory 303 and configured to be executed by the one or more processors 302, and the one or more application programs 301 are configured to execute the missing value completion method corresponding to any of the embodiments described above.

[0083] In this embodiment of the invention, borehole samples are used as the basic processing unit. Without relying on external heterogeneous training data, a completion model is constructed using existing complete sample records in the target borehole dataset. Elemental coherence relationships, sample spatial locations, and geological textual semantic information are jointly introduced to complete missing fields. Compared with existing missing value completion methods that mainly rely on numerical statistical relationships or general table prediction processes, this method better adapts to the "bore number—sample number—start and end depth" organization of borehole data, and makes fuller use of existing numerical, spatial, and textual information in the mine borehole data. This improves the relevance and rationality of the missing value completion results to the target mine data structure and geological background. Finally, the generated completion results can be directly written back to the sampling and analysis table. Therefore, it can not only complete the estimation of missing values ​​in the element content field but also directly serve borehole database repair and subsequent 3D modeling, ore body interpretation, and geological analysis.

[0084] Furthermore, the above provides a detailed description of a method and related apparatus for completing missing values ​​in borehole sampling and testing data provided by the embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for completing missing values ​​in borehole sampling and testing data, characterized in that, The method includes: Extract the sampling test table, borehole basic information table, borehole inclination table and the corresponding layer description information or lithological description information of the sample section from the target borehole database as basic input data, and use the missing values ​​of the content field of at least one continuous target element in the sampling test table as target completion objects. Based on the target completion object, a training sample set is constructed by selecting sample segments with real observation values ​​in the target elements in the target borehole database as candidate samples, and multi-source constraint processing is performed on the candidate samples in the training sample set to form a multi-source constraint training dataset. Based on the target element, a target element missing value completion model is constructed. The multi-source constrained training dataset is used as the input of the target element missing value completion model, and the actual observed values ​​of the target element in the multi-source constrained training dataset are used as the output of the target element missing value completion model for training. The target element missing value completion model is an XGBoost regression model. The basic input data is input into the target element missing value completion model that has been trained and converged. The content field completion value corresponding to the target completion object is input, and the content field completion value is filled into the sampling test table to complete the completion of the missing value of the content field of at least one continuous target element in the sampling test table.

2. The missing value completion method according to claim 1, characterized in that, The sampling and analysis table provides the primary key of the sample segment, the sampling start and end depths, and elemental observation information; the borehole basic information table provides the spatial attributes of the borehole opening plane coordinates and the final borehole depth; the borehole inclination table provides borehole attitude information such as borehole angle and borehole dip angle; the stratification description information or lithological description information corresponding to the sample segment provides the corresponding bottom layer category and geological text semantic information; and the selection of the sample segment is jointly determined by the borehole number, sample number, sampling start point, and sampling end point; the sample number is used to represent a single sampling record; the borehole number and the sampling start point and sampling end point are used to determine the position of the corresponding sampling record in the borehole and depth range.

3. The missing value completion method according to claim 1, characterized in that, The method of constructing a training sample set by selecting sample segments with real observation values ​​from the target elements in the target borehole database based on the target completion object includes: Obtain the target elements in the target completion object, and filter the sample segments with real observation values ​​in the target elements as candidate samples in the target borehole database; Each sample segment in the candidate samples is initially screened according to the following criteria: the primary key information of the sample segment is complete, the target field has real observed values, and it can establish a correspondence with the borehole basic information table, borehole inclination table, stratification description information, or lithology description information. The initial screening of the candidate samples is then obtained. A secondary screening process is performed on the candidate samples after the initial screening. Samples with abnormal data, missing fields, non-standard record formats, non-unique primary keys, misaligned depth ranges, or misalignment with other basic attributes are removed. The candidate sample data after the secondary screening is obtained and used as the training sample set.

4. The missing value completion method according to claim 1, characterized in that, The step of performing multi-source constraint processing on candidate samples in the training sample set to form a multi-source constrained training dataset includes: Based on the sampling test table in each candidate sample in the training sample set, the element collaboration relationship features of the corresponding sample segment are constructed and processed to obtain the element collaboration relationship features of each candidate sample in the training sample set. Based on the borehole basic information table in each candidate sample in the training sample set, the spatial location features of the corresponding sample segment are constructed and processed to obtain the sample segment spatial location features of each candidate sample in the training sample set. Based on the hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set, the geological text semantic features of the corresponding sample segments are constructed and processed to obtain the geological text semantic features of each candidate sample in the training sample set. The element collaboration relationship features, sample spatial location features, and geological text semantic features of each candidate sample in the training sample set are associated according to the primary key to form a multi-source constrained training dataset. Each candidate sample in the multi-source constrained training dataset simultaneously contains the element collaboration relationship features representing the combination relationship of elements, the sample spatial location features representing the spatial background, and the geological text semantic features representing the geological environment.

5. The missing value completion method according to claim 4, characterized in that, The step of constructing the element-coordinated relationship features of corresponding sample segments based on the sampling test table of each candidate sample in the training sample set, to obtain the element-coordinated relationship features of each candidate sample in the training sample set, includes: Based on the sampling test table in each candidate sample in the training sample set, the element observation values ​​corresponding to several elements given in the corresponding sample segment are extracted. Based on the element observation values ​​of several elements given in the sample segment, the element coordination relationship feature is established to form the element coordination relationship feature of each candidate sample in the training sample set. The element coordination relationship feature is used to express the combination relationship and numerical correlation between different elements in the sample segment of each candidate sample.

6. The missing value completion method according to claim 4, characterized in that, The process of constructing spatial location features of corresponding sample segments based on the borehole basic information table in each candidate sample of the training sample set to obtain the sample segment spatial location features of each candidate sample in the training sample set includes: Based on the borehole basic information table and borehole inclination table in each candidate sample in the training sample set, the spatial attributes of the corresponding sample segment's open plane coordinates, final borehole depth, and borehole attitude are extracted. Based on the spatial attributes of the corresponding sample segment's open plane coordinates, final hole depth, and drilling posture, spatial position features are constructed to obtain the sample segment spatial position features of each candidate sample in the training sample set. The sample segment spatial position features are used to characterize the spatial position difference of the sample segment in the hole and in space, as well as its spatial posture information.

7. The missing value completion method according to claim 4, characterized in that, The process of constructing geological text semantic features for each candidate sample segment based on the hierarchical description information or lithological description information corresponding to the sample segment in each candidate sample in the training sample set, to obtain the geological text semantic features for each candidate sample in the training sample set, includes: The lithological codes corresponding to the sample segments within each candidate sample in the training sample set are systematically organized, and the main category codes are extracted and grouped into category groups according to prefix or controlled classification rules; simultaneously, The hierarchical description information or lithological description information corresponding to the sample segments in each candidate sample in the training sample set is segmented according to rules in order to extract relevant terms such as mineral name, lithology, color, tectonic phenomenon, occurrence mode, mineralization intensity and alteration characteristics of the sample segments; The relevant terms related to mineral names, lithology, color, tectonic phenomena, occurrence mode, mineralization intensity and alteration characteristics corresponding to the sample segments are semantically vectorized and clustered. Combined with expert review and terminology dictionary or controlled term set, terms that are too general, unclear or difficult to stably participate in modeling are removed to form the core terms of geological text. One-hot encoding is used to encode the category group and the core terms of the geological text into binary features of 0-1, forming the geological text semantic features of each candidate sample in the training sample set.

8. A device for completing missing values ​​in borehole sampling and testing data, characterized in that, The device includes: The target completion confirmation module is used to extract the sampling test table, borehole basic information table, borehole inclination table and the corresponding layer description information or lithological description information of the sample section from the target borehole database as basic input data, and to use the missing value of the content field of at least one continuous target element in the sampling test table as the target completion object. Training set construction module: used to construct a training sample set by selecting sample segments with real observation values ​​in the target borehole database based on the target completion object, and to perform multi-source constraint processing on the candidate samples in the training sample set to form a multi-source constraint training dataset; Model training module: used to construct a target element missing value completion model based on the target element, using the multi-source constrained training dataset as the input of the target element missing value completion model, and using the real observed values ​​of the target element in the multi-source constrained training dataset as the output of the target element missing value completion model for training processing, to form a target element missing value completion model that has been trained and converged, and the target element missing value completion model is an XGBoost regression model; The completion module is used to input the basic input data into the target element missing value completion model that has been trained and converged, input the content field completion value corresponding to the target completion object, and fill the content field completion value into the sampling test table to complete the completion of the missing value of the content field of at least one continuous target element in the sampling test table.

9. An electronic device comprising a processor and a memory, characterized in that, The processor runs a computer program or code stored in the memory to implement the missing value completion method as described in any one of claims 1 to 7.

10. A computer-readable storage medium for storing computer programs or code, characterized in that, When the computer program or code is executed by a processor, the missing value completion method as described in any one of claims 1 to 7 is implemented.