A knowledge-driven based geochemical data analysis processing and report generation method
By mapping, cleaning, density outlier screening, and generating coexistence deviation tables for geochemical data, and combining this with a geological knowledge base-trained model, the problem of misjudging anomalies in geochemical data analysis was solved, enabling reliable anomaly identification and traceable report generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU PROVINCIAL GEOLOGICAL BUREAU BIG DATA CENTER
- Filing Date
- 2026-05-15
- Publication Date
- 2026-06-26
AI Technical Summary
Existing geochemical data analysis methods are prone to misjudging mineral-induced anomalies as noise, and the generated reports lack integration with geological patterns, resulting in unreliable and untraceable anomaly interpretations.
By receiving raw geochemical tables and mapping, cleaning, and standardizing them, density outlier screening and coexistence benchmarks are constructed, a coexistence deviation table is generated, a reconstruction model is trained, and a semantic interpretation package is generated by combining a geological knowledge base to form a geochemical report.
This improves the reliability and traceability of geochemical analysis results, ensuring that anomaly identification and interpretation are based on reliable samples and geological properties, and that anomaly interpretations in reports have clear data sources and spatial basis.
Smart Images

Figure CN122287642A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of geochemical exploration technology, and in particular to a knowledge-driven method for geochemical data analysis, processing and report generation. Background Technology
[0002] Geochemical data reflects the regional geochemical background, elemental enrichment and migration characteristics, and potential mineralization information, serving as a crucial basis for mineral exploration and anomaly delineation. Current technologies often employ statistical analysis, clustering algorithms, machine learning, knowledge graphs, or large language models to identify anomalies, predict target areas, or generate reports from geochemical data. While these methods improve automation, they still exhibit easily overlooked technical limitations.
[0003] First, existing methods often directly use numerical or spatial outliers as the basis for anomaly judgment, ignoring whether the outliers violate the elemental symbiosis patterns of the study area. In actual geochemical data, high-value points may originate from mineralization, sample contamination, testing errors, or data entry mistakes. Relying solely on DBSCAN, threshold discrimination, or general machine learning models can easily misclassify genuine mineral-induced anomalies as noise, or retain distorted data with no geological significance as mineral exploration clues, affecting the reliability of subsequent anomaly interpretation and data reconstruction.
[0004] Second, existing report generation methods mostly remain at the level of summarizing statistical results or generating general text, lacking a mechanism to bind anomalies with faults, rock masses, strata, elemental assemblages, regional metallogenic models, and evidence sources. Even when using general RAGs or knowledge graphs, they tend to generate generalized geological descriptions, making it difficult to explain the spatial rationality, geochemical consistency, and uncertainty of each anomaly conclusion. Therefore, there is an urgent need for a geochemical data analysis and processing method that can identify and reconstruct anomalous data in conjunction with geological patterns, and transform the analysis results into traceable and verifiable professional interpretation reports. Summary of the Invention
[0005] To address the aforementioned problems, embodiments of the present invention provide a knowledge-driven method for geochemical data analysis, processing, and report generation, the method comprising:
[0006] Receive the original geochemical table and the first geological knowledge base, map, clean and standardize the sample field, spatial field, element field and geological attribute field, and generate the first standard table and the first element statistical table;
[0007] A first sample set is constructed based on the first standard table. Density outlier screening is performed on the first sample set to form a second sample set with outlier labels and a third sample set with background labels.
[0008] A first co-occurrence benchmark is generated based on the third sample set, and a second co-occurrence representation is generated based on the second sample set. The first co-occurrence benchmark and the second co-occurrence representation are compared to form a first co-occurrence deviation table. Based on the first co-occurrence deviation table, the second sample set is divided into a fourth sample set and a fifth sample set. The fifth sample set is merged into the third sample set to form a sixth sample set.
[0009] Based on the element field, an element-derived field is formed. Based on the sixth sample set, spatial field, geological attribute field and element-derived field, a first reconstruction training table is formed and a first reconstruction model is trained. The element to be reconstructed is determined in the fourth sample set, and the first reconstruction model is called to generate the first reconstruction data table.
[0010] Based on the fourth sample set, the first reconstructed data table, and the first geological knowledge base, a first spatial association table, a first combination matching table, and a first evidence table are generated.
[0011] The first semantic interpretation package is formed by combining the first element statistics table, the first symbiotic deviation table, the first reconstructed data table, the first spatial correlation table, the first combination matching table, and the first evidence table, and a geochemical exploration report is generated.
[0012] Furthermore, the mapping, cleaning, and standardization include: establishing a first field mapping table; writing the sample field into a first sample index; writing the spatial field into a first spatial index; writing the element field into a first element index; and writing the geological attribute field into a first geological attribute index; performing field normalization, missing data marking, detection limit marking, duplicate sample merging, and dimensional unification on the original geochemical table according to the first field mapping table to form the first standard table; and generating the first element statistical table based on the first standard table.
[0013] Further, the density outlier screening includes: extracting the element field and the spatial field from the first standard table to generate a first sample vector table; generating a first adjacency table based on the first sample vector table; writing density reachability identifiers into the first sample set based on the first adjacency table; writing samples with outlier identifiers into the second sample set, and writing samples with background identifiers into the third sample set.
[0014] Further, the first co-occurrence benchmark includes a first element pair table, a first sorting table, a first symbol table, and a first combined cover table; the second co-occurrence representation includes a second element pair table, a second sorting table, a second symbol table, and a second combined cover table; the first co-occurrence deviation table is generated by comparing the first co-occurrence benchmark and the second co-occurrence representation according to the field order of the element fields.
[0015] Furthermore, the first co-occurrence deviation table includes a first symbol difference field, a first sorting difference field, a first coverage difference field, and a first sample attribution field; the second sample set is split into the fourth sample set and the fifth sample set according to the first sample attribution field; the fifth sample set and the third sample set are merged by sample indexing to form the sixth sample set.
[0016] Further, the method for forming the first reconstruction training table includes: generating a first input field table using the spatial field, the geological attribute field, the element field, and the element-derived field in the sixth sample set; generating a first label field table using the original element values corresponding to the element to be reconstructed; associating the first input field table and the first label field table to form the first reconstruction training table; and training the first reconstruction model based on the first reconstruction training table.
[0017] Further, the method for generating the first reconstructed data table includes: locating the element to be reconstructed in the fourth sample set, marking the original element value of the element to be reconstructed as the first estimated field; calling the first reconstruction model to generate the first reconstruction field; and writing the sample field, the first estimated field, the first reconstruction field, and the corresponding record in the first co-occurrence deviation table into the first reconstructed data table.
[0018] Furthermore, the first geological knowledge base includes a first geological entity index, a first element combination index, and a first data source index; the first geological entity index includes fault records, rock mass records, stratigraphic records, and lithological unit records; the method for generating the first spatial association table includes: retrieving the first geological entity index based on the spatial field of the fourth sample set to obtain the geological entity record corresponding to the sample location; and writing the sample field, the spatial field, the geological entity record, and the first spatial relationship field into the first spatial association table.
[0019] Further, the first combined matching table is generated by matching the element fields of the fourth sample set, the first reconstructed data table, and the first element combination index by covering the field names and element combination members; the first evidence table is generated by binding the first spatial association table, the first combined matching table, and the first data source index to the source; the first semantic interpretation includes a first abnormal fact field, a first spatial evidence field, a first element evidence field, a first source field, and a first verification field; the report chapter table is filled according to the first semantic interpretation package to generate the geochemical report.
[0020] A knowledge-driven geochemical data analysis, processing, and report generation method further includes: generating a first consistency verification table based on a first spatial association table, a first combination matching table, and a first evidence table, and writing the first consistency verification table into the first verification field of a first semantic interpretation package.
[0021] The technical effects and advantages of the knowledge-driven geochemical data analysis, processing, and report generation method provided by this invention are as follows:
[0022] This invention improves the reliability and traceability of geochemical analysis results through continuous processing of anomaly identification, data reconstruction, and evidence report generation. First, it uses density outlier screening to form a second sample set with outlier identifiers and a third sample set with background identifiers. Then, it uses a first symbiotic benchmark and a second symbiotic characterization to generate a first symbiotic deviation table, further subdividing outlier samples and avoiding direct anomaly identification based solely on numerical or spatial outliers. A fifth sample set is merged into the third sample set to form a sixth sample set, and a first reconstruction model is trained based on this sixth sample set. This ensures that the first reconstructed data table is built on reliable samples and geological attribute constraints, reducing the impact of distorted data on subsequent analysis. Finally, through a first spatial correlation table, a first combination matching table, and a first evidence table, it binds anomaly samples, geological entities, element combinations, and data sources, forming a first semantic interpretation package. This ensures that the anomaly interpretation in the geochemical report has clear data sources, spatial basis, and elemental evidence. Attached Figure Description
[0023] Figure 1 This is a flowchart of a knowledge-driven geochemical data analysis, processing, and report generation method in Example 1.
[0024] Figure 2 This is a schematic diagram of the data flow for anomaly data identification and reconstruction in Example 1;
[0025] Figure 3 This is a flowchart of the consistency verification and first verification field generation method in Example 2. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1:
[0027] Please see Figure 1 As shown, embodiments of the present invention provide a knowledge-driven method for geochemical data analysis, processing, and report generation, the method comprising:
[0028] S1: Receive the original geochemical table and the first geological knowledge base, map, clean and standardize the sample field, spatial field, element field and geological attribute field, and generate the first standard table and the first element statistical table;
[0029] S2: Construct a first sample set based on the first standard table, perform density outlier screening on the first sample set, and form a second sample set with outlier labels and a third sample set with background labels;
[0030] S3: Generate a first co-occurrence benchmark based on the third sample set, generate a second co-occurrence representation based on the second sample set, compare the first co-occurrence benchmark with the second co-occurrence representation to form a first co-occurrence deviation table, divide the second sample set into a fourth sample set and a fifth sample set based on the first co-occurrence deviation table, and merge the fifth sample set into the third sample set to form a sixth sample set;
[0031] S4: Based on the element field, form the element-derived field; based on the sixth sample set, spatial field, geological attribute field and element-derived field, form the first reconstruction training table and train the first reconstruction model; determine the element to be reconstructed in the fourth sample set; call the first reconstruction model to generate the first reconstruction data table.
[0032] S5: Generate the first spatial association table, the first combination matching table, and the first evidence table based on the fourth sample set, the first reconstructed data table, and the first geological knowledge base;
[0033] S6: The first element statistics table, the first coexistence deviation table, the first reconstruction data table, the first spatial correlation table, the first combination matching table, and the first evidence table are combined to form the first semantic interpretation package, and a geochemical exploration report is generated.
[0034] In this embodiment, mapping, cleaning, and standardization are used to convert raw geochemical tables from inconsistent sources into a first standard table that can be directly called upon for subsequent density outlier screening, coexistence benchmark construction, and reconstruction training. The raw geochemical table can be a data table formed by summarizing field sampling records, laboratory test records, and geological logging records, which includes at least sample fields, spatial fields, element fields, and geological attribute fields. The sample field is used to characterize the sample number, sampling batch, and sample type; the spatial field is used to characterize the location coordinates of the sample point; the element field is used to characterize the test values of geochemical elements; and the geological attribute field is used to characterize the stratigraphy, lithology, structural location, and geological unit attributes corresponding to the sample point.
[0035] Specifically, a first field mapping table is first established. This table records the correspondence between the original field names in the original geochemical table and the field names within this method. For example, original fields such as "sample number" and "sampling number" are uniformly mapped to sample fields, original fields such as "X coordinate", "longitude", and "ordinate" are uniformly mapped to spatial fields, element names and their test columns are uniformly mapped to element fields, and original fields such as "lithology", "stratigraphic code", and "tectonic unit" are uniformly mapped to geological attribute fields. After the mapping is completed, the sample fields are written to the first sample index, the spatial fields are written to the first spatial index, the element fields are written to the first element index, and the geological attribute fields are written to the first geological attribute index. These indexes are used to store the field source, field affiliation, and record location, so that the subsequently formed second sample set, third sample set, and first coexistence deviation table can be traced back to the original sample record.
[0036] During the cleaning process, based on the first field mapping table, the original geochemical tables are subjected to field normalization, missing data marking, detection limit identification, duplicate sample merging, and dimensional unification. Field normalization means writing fields with the same meaning but different names or formats into a unified field structure; missing data marking means writing records with missing element test values, spatial fields, and geological attribute fields into a missing status, rather than directly deleting the records; detection limit identification means splitting the status of element test values with symbols indicating below, above, or detection limit, and saving the test value field and the detection limit status field separately; duplicate sample merging means performing consistency checks on duplicate records under the same sample field based on the first sample index, and writing the retained records, merged records, and abnormal records into the cleaning record; dimensional unification means converting test values expressed in different units under the same element field into the same dimensional expression, while retaining the original dimensional source.
[0037] For example: If a certain element field in the original geochemical table is recorded as "<0.01", it is not directly written as a normal value during cleaning. Instead, the test value, detection limit status and original expression corresponding to the element field are recorded separately, so that the actual test value and the detection limit processed value can be distinguished when the first element statistical table is formed later. For another example, if the same sample field corresponds to the same spatial field in two records but there are duplicate test values in the element field, the duplicate records are first identified according to the first sample index, and then the merged sample record is generated according to the cleaning rules, and the source of duplication is retained in the cleaning record.
[0038] After the above processing, a first standard table is generated. The first standard table includes at least sample field, spatial field, element field, geological attribute field, detection limit status field, missing status field, and cleaning record field. Subsequently, a first element statistical table is generated based on the first standard table. The first element statistical table summarizes the element fields and records the effective record range, missing status distribution, detection limit status distribution, and dimensional source information of the element fields. The first element statistical table serves as the data source for the subsequent construction of the first sample set and the generation of the first co-occurrence benchmark, and maintains a corresponding relationship with the sample field in the first standard table.
[0039] like Figure 2 As shown, in this embodiment, density outlier screening, co-occurrence deviation comparison, sample attribution splitting, and reconstructed data generation are performed according to... Figure 2 The data stream shown is executed. Figure 2 The first standard table provides the basic data after mapping, cleaning, and standardization. The first sample set is constructed from the first standard table. The second and third sample sets are formed by density outlier screening. The first co-occurrence deviation table is generated by comparing the first co-occurrence benchmark with the second co-occurrence characterization. The fourth, fifth, and sixth sample sets are formed from the attribution results of the first co-occurrence deviation table. The first reconstruction training table, the first reconstruction model, and the first reconstruction data table are used to complete the reconstruction processing of the elements to be reconstructed. Density outlier screening is not directly used as the final anomaly determination result, but is used to form the second sample set to be further verified and the third sample set representing the regional geochemical background from the first sample set. After obtaining the first standard table, the system determines the order of element fields participating in the calculation according to the first element index, and extracts the element fields and spatial fields corresponding to each sample from the first standard table using the first sample index as the row identifier, forming the first sample vector table. One record in the first sample vector table corresponds to one sample, and the record content includes sample fields, spatial fields, and element field values arranged in a unified field order. Among them, element fields with missing status, detection limit status, and cleaning records are not directly used as ordinary raw values in the calculation, but are uniformly converted according to the status identifier in the first standard table, so that the element values and spatial positions of the same sample can enter the same calculation structure.
[0040] After generating the first sample vector table, a first adjacency table is formed based on the proximity relationships between sample vectors. The first adjacency table is used to record the proximity relationships between samples in terms of element content characteristics and spatial location characteristics. It is established as follows: first, element feature sub-vectors are formed based on the element field, and then spatial feature sub-vectors are formed based on the spatial field. The two are written into the vector record of the same sample. Then, the distance relationship between sample records is calculated, and the sample relationship that satisfies the proximity rule is written into the first adjacency table. Since there are usually symbiotic, associated, and differentiated relationships between geochemical elements, the distance calculation not only compares the level of a single element, but also combines the synergistic change relationship between element fields to avoid the local high value of a certain element directly dominating the outlier judgment.
[0041] When generating the first adjacency table, the first sample vector table can be represented as a multi-element content matrix, where each row corresponds to a sample field and each column corresponds to an element field. For the samples in the first sample set... and samples Their distance relationship can be determined as follows:
[0042] ;
[0043] In the formula, and Samples and The corresponding element field vector, Given the covariance matrix of the element field of the first sample set, the system generates a first adjacency table based on the distance relationship and writes density reachability identifiers into the first sample set accordingly.
[0044] After obtaining the first adjacency table, the system writes density reachability identifiers to the first sample set based on the adjacency relationships. The density reachability identifier is used to indicate whether a sample can be included in the main sample group through adjacency relationships. For samples that can be connected to the main sample group through adjacency relationships, a background identifier is written; for samples that cannot be included in the main sample group through adjacency relationships, an outlier identifier is written. Samples with outlier identifiers are written to the second sample set, and samples with background identifiers are written to the third sample set. The second sample set only represents spatial-numerical outliers initially screened out by density relationships, while the third sample set serves as the data basis for subsequently generating the first co-occurrence benchmark.
[0045] For example, when processing samples with high Au content, even if a sample has a significantly high Au field, if its spatial field and other elemental fields can still maintain a density reachable relationship with the main sample group through the first adjacency table, then the sample will not directly enter the second sample set simply because of a single high value. Conversely, if the sample is difficult to integrate into the main sample group in terms of elemental field combination and spatial field relationship, it will first be written into the second sample set and then compared with the subsequent first co-existence benchmark and second co-existence characterization. Thus, density outlier screening only plays the role of initial screening and sample stratification, providing a continuous data entry point for the subsequent formation of the first co-existence deviation table, the fourth sample set, the fifth sample set, and the sixth sample set.
[0046] In this embodiment, the first symbiotic benchmark is used to characterize the normal symbiotic relationship of elements in the geochemical background of the region corresponding to the third sample set, and the second symbiotic characterization is used to characterize the element combination state of the outlier samples in the second sample set. Since the second sample set is only obtained by density outlier screening, its outlier may come from real geological anomalies, or from sampling contamination, testing errors and input deviations. Therefore, the second sample set is not directly used as the final outlier result, but is compared with the first symbiotic benchmark formed by the third sample set.
[0047] Specifically, the system pairs the element fields in the third sample set according to the order of the element fields in the first standard table to form a first element pair table. The first element pair table is used to record the field source and pairing relationship of each element pair, so that subsequent comparisons are not affected by the differences in field arrangement in the original geochemical table. Based on the first element pair table, the system calculates the relative change relationship of each element pair in the third sample set and generates a first sorting table, a first symbol table, and a first combination coverage table. The first sorting table is used to record the relative high and low order of element fields in the third sample set, the first symbol table is used to record the directional relationship of same-direction and opposite-direction changes between element pairs, and the first combination coverage table is used to record the range of element combinations that appear stably in the third sample set.
[0048] For the second sample set, the system uses the same element field order and element pairing rules as the third sample set to generate a second element pair table, a second sorting table, a second symbol table, and a second combination coverage table. The second element pair table uses the sample fields in the second sample set as indexes and records the element pair relationships for each sample. The second sorting table records the relative order of element fields in outlier samples. The second symbol table records the element pair directional relationships of outlier samples. The second combination coverage table records the element combinations actually covered by outlier samples. Thus, the second co-occurrence characterization is consistent with the first co-occurrence benchmark in terms of field source, element pair order, and combination caliber.
[0049] Subsequently, the system aligns the first element pair table with the second element pair table, the first sorting table with the second sorting table, the first symbol table with the second symbol table, and the first combined coverage table with the second combined coverage table according to the field order of the element fields, and generates difference records item by item. Each difference record includes at least the source of the element pair, sorting changes, direction changes, and combined coverage changes. When generating difference records, the element pair relationships under the same sample field can be summarized to form the co-occurrence deviation of that sample relative to the first co-occurrence benchmark. For samples in the second sample set... Its symbiotic deviation It can be represented as:
[0050] ;
[0051] In the formula, Represents the elements in the third sample set and Symbiotic relationship in the first symbiotic benchmark Indicates sample Corresponding element relationships in the second symbiotic representation The number of element pairs participating in the comparison.
[0052] The system summarizes the above difference records by sample field and generates a first coexistence deviation table. The first coexistence deviation table does not represent the level of element content alone, but rather the deviation of outlier samples in the second sample set from the background coexistence relationship of the third sample set.
[0053] For example: If a sample is included in the second sample set in the density outlier screening due to a high Au field, but the ordering, direction, and combination coverage relationships of its accompanying elements are still consistent with those of the third sample set, then it will show a weak co-occurrence deviation in the first co-occurrence deviation table. If the sample not only has a high Au field, but also does not have a consistent combination relationship with the elements that appear stably in the third sample set, then a corresponding ordering difference, sign difference, and coverage difference will be formed in the first co-occurrence deviation table. Through this processing, the first co-occurrence deviation table becomes the data basis for subsequently dividing the second sample set into the fourth and fifth sample sets.
[0054] In this embodiment, the first co-occurrence deviation table is used to take over the comparison results between the first co-occurrence benchmark and the second co-occurrence characterization, and further divides the samples in the second sample set into the fourth sample set and the fifth sample set. The first co-occurrence deviation table does not only record whether the element content is higher than the background, but records the deviation of the outlier sample from the regional geochemical background represented by the third sample set in terms of elemental orientation relationship, elemental relative order and elemental combination coverage relationship.
[0055] Specifically, the first symbolic difference field is used to record the difference between the directional relationship of element pairs in the second sample set and the first co-occurrence benchmark. If a certain element pair in the third sample set shows a stable unidirectional change, while the corresponding sample in the second sample set shows an opposite change, then the corresponding difference identifier is written in the first symbolic difference field. The first sorting difference field is used to record the change in the relative order of the element fields, that is, to determine whether the arrangement relationship of the main elements in the outlier samples still conforms to the background sorting in the third sample set. The first coverage difference field is used to record whether the actual element combinations in the second sample set cover the background combination relationship in the first combination coverage table, and whether there are element combinations that are inconsistent with the background combination. The first sample attribution field is used to save the sample diversion results after the above differences are summarized, and to maintain a correspondence with the first sample index.
[0056] When generating the first sample attribution field, the system reads the first symbol difference field, the first sorting difference field, and the first coverage difference field on a sample-by-sample basis, and aggregates the three types of difference records within the same sample. If an outlier sample exhibits at least one of the following preset attribution conditions: mismatch of element pair orientation relationship, mismatch of element sorting relationship, and mismatch of combination coverage relationship, an abnormal attribution identifier is written into the first sample attribution field, and the sample is written into the fourth sample set. If an outlier sample is marked as an outlier by density outlier screening, but its element pair orientation relationship, element sorting relationship, and combination coverage relationship are still consistent with the first co-occurrence benchmark, a merge attribution identifier is written into the first sample attribution field, and the sample is written into the fifth sample set.
[0057] The fourth sample set is used to store sample records that still need to be reconstructed after density outlier screening and co-occurrence deviation verification. Its sample fields, element fields, spatial fields and corresponding first co-occurrence deviation table records are kept related. The fifth sample set is used to store sample records that belong to the second sample set but do not show obvious co-occurrence deviation. Subsequently, based on the first sample index, the fifth sample set and the third sample set are merged by sample index, duplicate sample records are removed and the original field sources are retained to form the sixth sample set. The sixth sample set is composed of the third sample set and the fifth sample set and is used for the formation of the first reconstruction training table.
[0058] For example: If a high-Au sample is selected as an outlier by density outlier screening and enters the second sample set, but its directional and ordinal relationships with elements such as Ag, Cu, and Pb still conform to the background co-occurrence relationships in the third sample set, then this sample can be written into the fifth sample set and merged into the sixth sample set along with the third sample set. If another high-Au sample is not only outlier in spatial-numerical relationships, but its element combination coverage relationship is also significantly inconsistent with the first co-occurrence benchmark, then this sample is written into the fourth sample set as the object for subsequent element location to be reconstructed and the generation of the first reconstruction data table. In this way, the second sample set is not directly equated with outlier samples, but is diverted through the attribution field of the first co-occurrence deviation table.
[0059] In this embodiment, the first reconstruction training table is used to receive the sixth sample set and provide a trainable data structure for the first reconstruction model. The sixth sample set is formed by merging the third sample set and the fifth sample set through sample indexing. The third sample set is a set of samples with background labels after density outlier screening. The fifth sample set is a set of samples that, although they entered the second sample set, still conform to the background co-occurrence relationship as determined by the first co-occurrence deviation table. Therefore, the sixth sample set does not directly use all the original samples, but uses the samples after outlier screening and co-occurrence deviation verification as the reconstruction training source.
[0060] Specifically, the system uses the sample field in the sixth sample set as the primary key of the record, extracts spatial fields, geological attribute fields, element fields, and element-derived fields from the first standard table, and generates the first input field table. The spatial field is used to record the positional relationship of the sample points, the geological attribute field is used to record the lithology, strata, structural location, and geological unit attributes to which the sample points belong, the element field is used to record the element content after cleaning and standardization, and the element-derived fields are generated by the element fields according to the element combination relationship, including the element ratio field, the element combination summary field, and the element standardization transformation field. The generation of the element-derived fields is based on the field order in the first element index, so that the same element combination has a fixed field position in the first input field table.
[0061] When forming the first input field table, the system first determines the elements to be reconstructed in the fourth sample set, and then reads the original element values corresponding to the elements to be reconstructed from the sixth sample set to generate the first label field table. The first label field table is indexed by the sample field, and its label field is the original element value of the element to be reconstructed in the sixth sample set. In order to avoid the label field being repeatedly entered into the training as a normal input field, the original element values directly corresponding to the elements to be reconstructed in the first input field table are not written into the model input as independent prediction basis, but are composed of the other element fields, element-derived fields, spatial fields and geological attribute fields.
[0062] Subsequently, the system associates the first input field table with the first label field table based on the sample fields to generate the first reconstruction training table. The association is not a simple concatenation, but a consistency check is performed on the sample fields, the completeness of the spatial fields, the attribution of the geological attribute fields, and the status of the label fields. Records with sample fields that cannot be matched, missing label fields that have not been resolved, or geological attribute fields that cannot be assigned are not written into the valid training records of the first reconstruction training table. Thus, each record in the first reconstruction training table includes an input field that can be used for prediction and its corresponding label field.
[0063] The first reconstruction model is trained based on the first reconstruction training table. During training, the first reconstruction model learns the correspondence between spatial fields, geological attribute fields, element fields, and element-derived fields and the elements to be reconstructed. This enables it to generate corresponding reconstruction values based on the geological background and element combination state of the same sample after the elements to be reconstructed in the fourth sample set are marked as to be estimated. For example, when the element to be reconstructed is Au, the first label field table records the original element value of Au in the sixth sample set, and the first input field table records the spatial fields, lithological attributes, structural attributes, and related element fields and element-derived fields other than the direct original value of Au of the corresponding sample. The first reconstruction model formed by training is then used to reconstruct the Au elements to be reconstructed in the fourth sample set.
[0064] In this embodiment, the first reconstruction data table is used to store the original state of the elements to be reconstructed in the fourth sample set, the reconstruction results, and the sources of their co-occurrence deviations. The fourth sample set is a sample set formed after the second sample set is split by the first co-occurrence deviation table. Its samples are not simply numerical outliers, but samples that simultaneously have outlier identifiers and co-occurrence deviation records. The elements to be reconstructed refer to the element fields in the fourth sample set that need to be reconstructed. The determination criteria include the first sign difference field, the first sorting difference field, and the first coverage difference field in the first co-occurrence deviation table.
[0065] Specifically, the system first reads the fourth sample set using the sample field as an index, and locates the original element value corresponding to the element to be reconstructed in the fourth sample set. For the located original element value, it is not directly deleted, nor is it directly overwritten with the model output value. Instead, it is written into the first estimated field in the first reconstruction data table. The first estimated field is used to save the original record, field source and estimated status of the element to be reconstructed, so that the original observation value and the model reconstruction value can be distinguished when the subsequent report is generated.
[0066] Subsequently, the system calls the first reconstruction model. During the call, the spatial field, geological attribute field, element field, and element-derived field of the corresponding sample in the fourth sample set are organized into input records according to the field structure of the first reconstruction training table. The original element value of the element to be reconstructed is not used as a normal input field for reconstruction, but is retained as the first estimated field. The first reconstruction model generates the reconstruction value corresponding to the element to be reconstructed based on the input record and writes the reconstruction value into the first reconstruction field. The first reconstruction field is used to represent the element value state that the element to be reconstructed should correspond to under the reliable background relationship formed by the sixth sample set.
[0067] When generating the first reconstruction data table, the system associates and writes the sample field, the first field to be estimated, the first reconstruction field, and the corresponding records in the first co-occurrence deviation table according to the sample field. The corresponding records include at least the sign difference, sorting difference, coverage difference, and sample attribution information of the sample in the first co-occurrence deviation table. Thus, the first reconstruction data table not only saves the field values before and after reconstruction, but also saves the co-occurrence deviation basis for the sample to be included in the fourth sample set, avoiding the reconstruction result from deviating from the previous anomaly identification process.
[0068] For example: If a sample is entered into the fourth sample set due to an abnormality in the Au field, the system first writes the original Au element value of the sample into the first field to be estimated, and then calls the first reconstruction model to generate the first reconstruction field of Au based on the spatial field, geological attribute field and Au-related element-derived field of the sample; at the same time, the sorting difference and coverage difference records of the sample in the first coexistence deviation table are written into the first reconstruction data table. In this way, the first reconstruction data table can provide source-constrained reconstruction results for the subsequent first spatial association table, first combination matching table and first evidence table.
[0069] In this embodiment, the first geological knowledge base is used to store geological entities, element combinations and data sources related to the study area, and the first geological entity index is used to uniformly register spatial objects in regional geological data, including fault records, rock mass records, stratigraphic records and lithological unit records. Each type of record contains entity name, entity type, spatial range, source data and entity number, so that the samples in the fourth sample set can retrieve the corresponding geological background based on the spatial fields, rather than relying solely on text similarity to generate interpretations.
[0070] Specifically, the system uses the sample field in the fourth sample set as an index to read the spatial field corresponding to the sample. The spatial field is the location expression of the sample point, which can be derived from the coordinate field in the first standard table and corresponds to the first spatial index. After the system converts the spatial field into a sample point spatial object, it sequentially searches the fault record, rock mass record, stratigraphic record and lithological unit record in the first geological entity index to determine the spatial correspondence between the sample point and each geological entity. The spatial correspondence can include relationship types such as the sample point falling within the lithological unit range, the sample point being located within the stratigraphic range, the sample point being adjacent to the fault record, and the sample point being adjacent to the rock mass boundary.
[0071] After obtaining the geological entity records corresponding to the sample location, the system does not directly generate report text. Instead, it first forms a first spatial association table. The first spatial association table uses the sample field as the primary key and writes the sample field, spatial field, geological entity records, and first spatial relationship field into the same record. The first spatial relationship field is used to record the relationship category, relationship source, and matching basis between the sample point and the geological entity. The geological entity records are used to store the names of the matched faults, rock masses, strata, and lithological units. The spatial field is used to retain the source of the sample point location, which is convenient for subsequent association with the first combination matching table and the first evidence table.
[0072] For example: the spatial field of a certain Au anomaly sample in the fourth sample set corresponds to a certain lithological unit and is adjacent to a fault record. After the system obtains the lithological unit record and the fault record from the first geological entity index, it writes the sample field, the original spatial field, the lithological unit record, the fault record, and the corresponding first spatial relationship field into the first spatial association table. Thus, the first spatial association table provides a clear spatial evidence entry point for the subsequent first combination matching table and first evidence table, enabling the anomaly interpretation in the geochemical report to be traced back to the specific sample location and geological entity record.
[0073] In this embodiment, the first combination matching table is used to record the correspondence between abnormal samples in the fourth sample set and the first element combination index. The first element combination index comes from the first geological knowledge base. It registers the element field name, element combination name, data source identifier, and combination applicable conditions according to mineralization type, geological background, and element combination relationship. When generating the first combination matching table, the system uses the sample field in the fourth sample set as the primary key, reads the element field corresponding to the sample, and calls the corresponding records of the first estimated field, the first reconstructed field, and the first co-occurrence deviation table in the first reconstruction data table. Subsequently, the above element fields are matched with the element combination fields in the first element combination index. Field matching means that the actual element fields of the sample, the reconstructed element field status, and the element status are aligned with the element combination fields registered in the knowledge base according to the element field name, element combination members, and element status, and cover records, missing records, and inconsistent records are formed. Therefore, the first combination matching table does not simply record whether a certain element has a high value, but records whether the element combination of the abnormal sample can correspond to the combination relationship in the first element combination index.
[0074] The first evidence table is generated by binding the first spatial association table, the first combination matching table, and the first data source index. Specifically, the system first reads the sample field, spatial field, geological entity record, and first spatial relationship field from the first spatial association table, then reads the element combination matching results under the same sample field from the first combination matching table, and obtains the corresponding data name, data type, record location, and source identifier according to the first data source index. When binding the source, the system uses the sample field as the primary key and writes the spatial evidence, element evidence, and data source into the same evidence record, so that the spatial location, geological entity, element combination, and data source of the same abnormal sample form a corresponding relationship.
[0075] The first semantic interpretation package is used to receive the above structured results. It includes a first anomaly fact field, a first spatial evidence field, a first element evidence field, a first source field, and a first verification field. The first anomaly fact field is written into the sample field, the element to be reconstructed, the first estimated field, and the first reconstructed field in the fourth sample set. The first spatial evidence field is written into the geological entity record and the first spatial relationship field in the first spatial association table. The first element evidence field is written into the field coverage matching result in the first combination matching table. The first source field is written into the source identifier in the first data source index. The first verification field is written into the corresponding record in the first coexistence deviation table, the field status before and after reconstruction, and the evidence integrity status.
[0076] When generating geochemical exploration reports, the system does not directly generate the main text based on the natural language model. Instead, it first fills the report chapter table with the first semantic interpretation package. The report chapter table sets the field entries according to the chapter order of data processing, geochemical characteristics, anomaly delineation and interpretation, and conclusion recommendations. Among them, the data processing chapter calls the first element statistics table and the first reconstructed data table, the anomaly delineation and interpretation chapter calls the first anomaly fact field, the first spatial evidence field, and the first element evidence field, and the conclusion recommendations chapter calls the first source field and the first verification field. Through the above processing, the anomaly interpretation in the geochemical exploration report can be traced back item by item to the fourth sample set, the first reconstructed data table, the first spatial correlation table, the first combination matching table, and the first evidence table.
[0077] For example: After a certain Au anomaly sample enters the fourth sample set, the first combination matching table performs field coverage matching between its Au-related element field and the first reconstruction field and the Au-related combination in the first element combination index; the first evidence table further binds the sample's adjacent fault records, its lithological unit records, and corresponding data sources; the first semantic interpretation package writes the anomaly facts, spatial evidence, elemental evidence, and source information into the report chapter table, thereby generating a geochemical report paragraph with data source and evidence correspondence. Example 2:
[0078] like Figure 3 As shown, this embodiment further improves the design based on Embodiment 1. The difference is that in the actual operation of Embodiment 1, it was found that when the abnormal samples in the fourth sample set simultaneously hit the first spatial association table and the first combination matching table, although the first semantic interpretation package has formed abnormal facts, spatial evidence, elemental evidence and source information, there are situations where the evidence granularity is inconsistent, the source binding is incomplete, and the element status before and after reconstruction is not synchronously checked between the first coexistence deviation table, the first reconstruction data table and the first evidence table corresponding to some samples. This causes the geochemical report to easily output the abnormal interpretation with sufficient evidence and the abnormal interpretation with insufficient evidence with the same argumentation strength when filling the report chapter table, and fails to form a verification chain jointly limited by abnormal identification, element reconstruction, spatial association, combination matching and data source for each abnormal sample. Based on this, a knowledge-driven geochemical data analysis and processing and report generation method further includes: before forming the first semantic interpretation package, generating a first consistency verification table based on the first coexistence deviation table, the first reconstruction data table, the first spatial association table, the first combination matching table and the first evidence table, and updating the first verification field based on the first consistency verification table.
[0079] Specifically, the system uses the sample field as the primary key, reads the first symbol difference field, the first sorting difference field, and the first coverage difference field from the first coexistence deviation table, reads the first field to be estimated and the first reconstruction field from the first reconstruction data table, reads the geological entity record and the first spatial relationship field from the first spatial association table, reads the field coverage matching result from the first combination matching table, and reads the data source binding result from the first evidence table. The above records are not directly entered into the report text, but are first aggregated according to the same sample field to form the first consistency check table.
[0080] The first consistency check table includes the first anomaly source check field, the first reconstruction consistency check field, the first spatial evidence check field, the first elemental evidence check field, and the first source completeness check field. The first anomaly source check field is used to record whether the samples in the fourth sample set have corresponding records in the first coexistence deviation table. The first reconstruction consistency check field is used to record the correspondence between the first field to be estimated and the first reconstruction field, and to mark whether the reconstructed value is consistent with the element to be reconstructed and the sample field. The first spatial evidence check field is used to record whether the sample field has obtained the corresponding geological entity record and the first spatial relationship field. The first elemental evidence check field is used to record whether the element field of the sample and the first reconstruction field have completed field coverage matching with the first element combination index. The first source completeness check field is used to record whether both spatial evidence and elemental evidence can obtain the source identifier in the first data source index.
[0081] After generating the first consistency check table, the system updates the first check field in the first semantic interpretation package based on the first consistency check table. For samples that can be matched with anomaly source records, reconstruction records, spatial evidence records, element evidence records, and data source records, the first check field is written with a complete check identifier. For samples that lack geological entity records, element combination matching records, or data source identifiers, the first check field is written with a missing item check identifier, and the corresponding missing item position is retained. For samples that cannot be matched with the first estimated field, the first reconstruction field, and the first coexistence deviation table records according to the sample field, the first check field is written with a conflict check identifier.
[0082] When generating a geochemical exploration report, the report chapter table is no longer filled solely based on the first anomaly fact field, the first spatial evidence field, and the first element evidence field. Instead, it simultaneously reads the first verification field. For anomaly samples corresponding to complete verification identifiers, a complete anomaly explanation paragraph is written into the report chapter table; for anomaly samples corresponding to missing verification identifiers, a supplementary evidence paragraph is written into the report chapter table; and for anomaly samples corresponding to conflict verification identifiers, a review prompt paragraph is written into the report chapter table, distinguishing it from the ordinary anomaly explanation paragraph. Thus, the anomaly conclusions in the report are no longer merely expressed as natural language descriptions, but are generated by the joint constraints of the first symbiotic deviation table, the first reconstructed data table, the first spatial correlation table, the first combination matching table, and the first evidence table.
[0083] For example: After an Au anomaly sample enters the fourth sample set, if there is a corresponding sorting difference record in the first coexistence deviation table, a corresponding first estimated field and a first reconstructed field in the first reconstructed data table, a fracture record and a lithological unit record in the first spatial association table, an Au-related element combination coverage record in the first combination matching table, and the first evidence table can be bound to the data source index, then the first verification field is written with a complete verification identifier; if the sample only has an Au-related element combination coverage record, but lacks a corresponding geological entity record or data source identifier, then the first verification field is written with a missing item verification identifier, and the report chapter table is simultaneously written with a missing item explanation when generating the anomaly interpretation. In this way, Example 2 is not a simple supplement to Example 1, but rather establishes a sample-level verification chain before the semantic interpretation package is generated, so that the report conclusions correspond one-to-one with the anomaly identification, reconstruction results, spatial evidence, elemental evidence, and data source.
[0084] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0085] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present application, based on the technical solution and concept of the present application, should be covered within the scope of protection of the present application.
Claims
1. A knowledge-driven method for geochemical data analysis, processing, and report generation, characterized in that the method... include: Receive the original geochemical table and the first geological knowledge base, map, clean and standardize the sample field, spatial field, element field and geological attribute field, and generate the first standard table and the first element statistical table; A first sample set is constructed based on the first standard table. Density outlier screening is performed on the first sample set to form a second sample set with outlier labels and a third sample set with background labels. A first co-occurrence benchmark is generated based on the third sample set, and a second co-occurrence representation is generated based on the second sample set. The first co-occurrence benchmark and the second co-occurrence representation are compared to form a first co-occurrence deviation table. Based on the first co-occurrence deviation table, the second sample set is divided into a fourth sample set and a fifth sample set. The fifth sample set is merged into the third sample set to form a sixth sample set. Based on the element field, an element-derived field is formed. Based on the sixth sample set, spatial field, geological attribute field and element-derived field, a first reconstruction training table is formed and a first reconstruction model is trained. The element to be reconstructed is determined in the fourth sample set, and the first reconstruction model is called to generate the first reconstruction data table. Based on the fourth sample set, the first reconstructed data table, and the first geological knowledge base, a first spatial association table, a first combination matching table, and a first evidence table are generated. The first semantic interpretation package is formed by combining the first element statistics table, the first symbiotic deviation table, the first reconstructed data table, the first spatial correlation table, the first combination matching table, and the first evidence table, and a geochemical exploration report is generated.
2. The knowledge-driven geochemical data analysis, processing, and report generation method according to claim 1, characterized in that, The mapping, cleaning, and standardization include: establishing a first field mapping table; writing the sample field into a first sample index; writing the spatial field into a first spatial index; writing the element field into a first element index; and writing the geological attribute field into a first geological attribute index; performing field normalization, missing data marking, detection limit marking, duplicate sample merging, and dimensional unification on the original geochemical table based on the first field mapping table to form the first standard table; and generating the first element statistical table based on the first standard table.
3. The knowledge-driven geochemical data analysis, processing, and report generation method according to claim 1, characterized in that, The density outlier screening includes: extracting the element field and the spatial field from the first standard table to generate a first sample vector table; generating a first adjacency table based on the first sample vector table; writing density reachability identifiers into the first sample set based on the first adjacency table; writing samples with outlier identifiers into the second sample set, and writing samples with background identifiers into the third sample set.
4. The knowledge-driven geochemical data analysis, processing, and report generation method according to claim 1, characterized in that, The first co-occurrence benchmark includes a first element pair table, a first sorting table, a first symbol table, and a first combined cover table; the second co-occurrence representation includes a second element pair table, a second sorting table, a second symbol table, and a second combined cover table. The first co-existence benchmark and the second co-existence representation are compared according to the field order of the element fields to generate the first co-existence deviation table.
5. The knowledge-driven geochemical data analysis, processing, and report generation method according to claim 4, characterized in that, The first coexistence deviation table includes a first symbol difference field, a first sorting difference field, a first coverage difference field, and a first sample attribution field; The second sample set is divided into the fourth sample set and the fifth sample set based on the first sample attribution field; the fifth sample set and the third sample set are merged by sample indexing to form the sixth sample set.
6. The knowledge-driven geochemical data analysis, processing, and report generation method according to claim 1, characterized in that, The method for forming the first reconstruction training table includes: generating a first input field table using the spatial field, the geological attribute field, the element field, and the element-derived field in the sixth sample set; generating a first label field table using the original element values corresponding to the element to be reconstructed; associating the first input field table and the first label field table to form the first reconstruction training table; and training the first reconstruction model based on the first reconstruction training table.
7. The knowledge-driven geochemical data analysis, processing, and report generation method according to claim 6, characterized in that, The method for generating the first reconstructed data table includes: locating the element to be reconstructed in the fourth sample set, marking the original element value of the element to be reconstructed as the first estimated field; calling the first reconstruction model to generate the first reconstruction field; and writing the sample field, the first estimated field, the first reconstruction field, and the corresponding record in the first co-occurrence deviation table into the first reconstructed data table.
8. The knowledge-driven geochemical data analysis, processing, and report generation method according to claim 1, characterized in that, The first geological knowledge base includes a first geological entity index, a first element combination index, and a first data source index; the first geological entity index includes fault records, rock mass records, stratigraphic records, and lithological unit records; The method for generating the first spatial association table includes: retrieving the first geological entity index based on the spatial field of the fourth sample set, and obtaining the geological entity record corresponding to the sample location; Write the sample field, the spatial field, the geological entity record, and the first spatial relationship field into the first spatial association table.
9. The knowledge-driven geochemical data analysis, processing, and report generation method according to claim 8, characterized in that, The first combined matching table is generated by matching the element fields of the fourth sample set, the first reconstructed data table, and the first element combined index by covering the field names and element combined members; the first evidence table is generated by binding the first spatial association table, the first combined matching table, and the first data source index to the source; the first semantic interpretation includes a first abnormal fact field, a first spatial evidence field, a first element evidence field, a first source field, and a first verification field; the report chapter table is filled according to the first semantic interpretation package to generate the geochemical report.
10. The knowledge-driven geochemical data analysis, processing, and report generation method according to claim 9, characterized in that, Also includes: A first consistency verification table is generated based on the first spatial association table, the first combination matching table, and the first evidence table, and the first consistency verification table is written into the first verification field of the first semantic interpretation package.