Heterogeneous data analysis and knowledge graph construction method in gold industry

By performing semantic normalization and cross-modal consistency verification on multimodal heterogeneous data of gold mines, and combining causal inference and spatiotemporal evolution patterns, a high-confidence set of gold knowledge units is generated, solving the problem of multimodal geological data fusion and improving the intelligent prediction capability of deep concealed ore bodies.

CN122047418APending Publication Date: 2026-05-15TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-01-29
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing exploration knowledge modeling methods based on single data sources or simple statistical correlations are difficult to achieve semantic fusion and spatiotemporal coupling analysis of multimodal heterogeneous geological data. They lack the ability to deeply characterize the complex causal mechanisms and dynamic evolution laws of metallogenic systems, resulting in static, shallow, and weakly interpretable knowledge graphs that cannot effectively support high-confidence intelligent inference of the spatial location and genetic mechanisms of deep concealed ore bodies.

Method used

By acquiring a multimodal heterogeneous original dataset of gold mines, semantic normalization parsing based on domain ontology and cross-modal consistency verification are performed to generate a multimodal gold geological knowledge information set. Deep correlation mining of causal inference and spatiotemporal evolution patterns is carried out. Credibility backtracking verification and causal correlation strength screening are performed using known typical gold deposits to generate a verification gold knowledge unit set. Finally, knowledge graph modeling and iterative construction are carried out.

Benefits of technology

It greatly enhances the depth of knowledge utilization and the interpretability of reasoning in deep mineral exploration, providing a reliable new knowledge-driven paradigm for breaking through the bottleneck of intelligent prediction of concealed ore bodies, and forming a machine-understandable knowledge graph that combines causal logic and spatiotemporal constraints to support intelligent prediction of deep and concealed gold deposits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122047418A_ABST
    Figure CN122047418A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data analysis, in particular to a gold industry heterogeneous data analysis and knowledge graph construction method. The method comprises the steps that a gold mine multi-modal heterogeneous original data set is acquired, semantic normalization analysis and cross-modal consistency verification based on domain ontology are carried out on the gold mine multi-modal heterogeneous original data set, and a multi-modal gold geological knowledge information set is generated; based on this, performing deep association mining of fusion causal inference and a spatio-temporal evolution mode, and generating a candidate golden knowledge unit set with causal logic and spatio-temporal constraints; on the basis, taking a known typical gold deposit as a verification anchor point, carrying out credibility backtracking verification and causal association strength screening on knowledge units, and generating a verification gold knowledge unit set; and based on this, modeling and iterative construction of the knowledge graph are carried out, and the gold exploration knowledge graph serving deep and hidden gold mine intelligent prediction is generated. In the heterogeneous data analysis process, the intelligent prediction accuracy of the ore body is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis, and in particular to methods for heterogeneous data parsing and knowledge graph construction in the gold industry. Background Technology

[0002] In the field of gold geological exploration, knowledge-driven intelligent prediction is the core technology direction for addressing the global shortage of gold resources and breaking through the bottleneck of deep and concealed mineral exploration. It is also a key enabling tool to support the national mineral resource security strategy and a new round of mineral exploration breakthroughs.

[0003] However, existing exploration knowledge modeling methods based on a single data source or simple statistical correlation are difficult to achieve semantic fusion and spatiotemporal coupling analysis of multimodal heterogeneous geological data. They lack the ability to deeply characterize the complex causal mechanisms and dynamic evolution laws of metallogenic systems, resulting in static, shallow, and weakly interpretable knowledge graphs that cannot effectively support high-confidence intelligent inference of the spatial location and genetic mechanism of deep concealed ore bodies. Summary of the Invention

[0004] This application provides a method for heterogeneous data parsing and knowledge graph construction in the gold industry to solve the above-mentioned technical problems. The method includes: Obtain a multimodal heterogeneous original dataset of gold mines, perform semantic normalization parsing based on domain ontology and cross-modal consistency verification on the original dataset, and generate a multimodal gold geological knowledge information set. Based on the multimodal gold geological knowledge information set, a deep correlation mining process is performed that integrates causal inference and spatiotemporal evolution patterns to generate a candidate gold knowledge unit set with causal logic and spatiotemporal constraints. Based on the candidate gold knowledge unit set, and using known typical gold deposits as verification anchors, the credibility of the knowledge units is verified by backtracking and the causal relationship strength is screened to generate a verified gold knowledge unit set. Based on the verified gold knowledge unit set, a knowledge graph is modeled and iteratively constructed to generate a gold exploration knowledge graph that serves intelligent prediction of deep and concealed gold deposits.

[0005] Through the above technical solutions, based on systematic knowledge extraction, verification and graph construction, multi-source heterogeneous exploration data is transformed into machine-understandable knowledge that combines causal logic and spatiotemporal constraints. This greatly improves the depth of knowledge utilization and the interpretability of reasoning in the process of deep mineral exploration, and provides a reliable knowledge-driven new paradigm for breaking through the bottleneck of intelligent prediction of concealed ore bodies.

[0006] Optionally, the generation process of the multimodal gold geological knowledge information set includes: the multimodal heterogeneous original dataset of gold mines includes unstructured text, exploration geological images, and geophysical and chemical data; based on the multimodal heterogeneous original dataset, entity recognition, relation extraction, and semantic annotation based on the gold geological ontology are performed, and natural language processing technology is used to realize the semantic normalization expression of the same geological concepts in data from different sources, thereby constructing semantically normalized information; based on the multimodal heterogeneous original dataset, cross-modal correlation analysis is performed, and multimodal fusion technology is applied to identify the phenomenon of identical expressions and logic in different modal data when describing the same geological object or event, thereby constructing cross-modal consistency information; based on the semantic normalization information and the constructed cross-modal consistency information, the multimodal gold geological knowledge information set is generated.

[0007] Optionally, the process of constructing the semantically normalized information includes: based on the unstructured text, using a gold-related large model pre-trained on a gold domain corpus, performing deep analysis and entity relation extraction to generate a geological text semantic relation representation; based on the exploration geological images, using computer vision technology to identify geological structures and lithological units, and parsing them into vectorized spatial objects with geological semantic annotations; based on the geophysical and chemical data, using geological domain expertise to map the abnormal features in the geophysical and chemical data into response pattern descriptions representing geological, geographical, and geochemical information; and uniformly mapping the geological text semantic relation representation, the vectorized spatial objects, and the response pattern descriptions to standard concept nodes of the gold geological domain ontology to eliminate conceptual ambiguities within various types of data, thereby constructing the semantically normalized information.

[0008] Optionally, the process of constructing the cross-modal consistency information includes: establishing a cross-validation bridge connecting the unstructured text, the exploration geological images, and the geophysical and chemical data, using spatial coordinates or geological strata as a unified association benchmark; comparing and analyzing the descriptions of the same geological object or process from different modal data at the logical and attribute levels through the cross-validation bridge to identify potential expression consistency phenomena; based on the identified expression consistency phenomena, mutually verifying and enhancing the descriptions from different modal sources to generate association corroborating information that is supported by multi-source evidence and is logically self-consistent; and integrating all the association corroborating information generated through mutual verification to construct the cross-modal consistency information.

[0009] Optionally, the generation process of the candidate gold knowledge unit set includes: based on the multimodal gold geological knowledge information set, analyzing the temporal sequence and spatial association of geological events, mining and constructing event causal chains characterizing the mineralization process; simultaneously, analyzing the spatial distribution patterns and temporal evolution characteristics of geological elements including lithological units, tectonic traces, alteration zones, and mineralization bodies in the region, and extracting spatiotemporal evolution pattern information; and assigning the logical relationships of the event causal chains and the constraint attributes of the spatiotemporal evolution pattern information to the relevant knowledge units to form the candidate gold knowledge unit set.

[0010] Optionally, the process of constructing the event causal chain includes: extracting a set of geological event instances with clear spatiotemporal attributes from the multimodal gold geological knowledge information set, wherein the set of geological event instances includes rock mass intrusion events, tectonic activity events, hydrothermal alteration events, and mineral precipitation events; mining causal dependencies based on temporal constraints and spatial correlations in the set of geological event instances using a causal discovery algorithm to generate a preliminary causal network topology; performing logical constraints and path verification on the preliminary causal network topology based on a pre-set gold mineralization theory and geodynamic prior knowledge base to select candidate causal paths that conform to the laws of geological evolution; and analyzing the confidence score of each causal edge in the candidate causal path, wherein the confidence score comprehensively considers temporal tightness, spatial proximity, and the corroboration strength in the cross-modal consistency information, thereby constructing the event causal chain with multi-dimensional confidence assessment.

[0011] Optionally, the extraction process of the spatiotemporal evolution pattern information includes: extracting spatial geometric distribution data and geological time series data of the geological elements from the multimodal gold geological knowledge information set; based on the spatial geometric distribution data, using spatial topology analysis and density field simulation, quantitatively analyzing the spatial aggregation, orientation, and symbiotic combination patterns of the geological elements in the region, and generating a spatial distribution pattern; based on the geological time series data, constructing geological age sequences and superposition and cutting relationships of different geological elements, inferring their formation sequence and periodicity, and generating a time series evolution sequence; coupling the spatial distribution pattern with the time series evolution sequence to generate the spatiotemporal evolution pattern information that characterizes how the structure of the mineralization system dynamically evolves with geological time.

[0012] Optionally, the process of generating the verification gold knowledge unit set includes: selecting multiple known typical gold deposits with clear genetic types and distinct geological characteristics as verification benchmarks, and extracting their complete geological knowledge maps as verification templates; matching and mapping the event causal chain and spatiotemporal evolution pattern information one by one with the verified patterns in the verification templates; analyzing the matching degree and coverage of each candidate knowledge unit in multiple verification templates based on a machine learning classifier, marking successfully matched units as trustworthy, and quantifying their causal association strength based on the breadth and consistency of their matching; and finally selecting trustworthy knowledge units that pass backtracking verification and whose association strength is higher than a preset threshold to form the verification gold knowledge unit set.

[0013] Optionally, the generation process of the gold exploration knowledge graph includes: constructing an initial graph network using the verified gold knowledge unit set as nodes and relationships, and using the event causal chain and the spatiotemporal evolution pattern information; using a dynamic graph engine to perform knowledge fusion, reasoning completion, and logical conflict resolution on the initial graph network to generate the gold exploration knowledge graph that satisfies consistency constraints and has multi-dimensional reasoning capabilities; inputting new exploration evidence into the gold exploration knowledge graph to trigger its adaptive iterative update and verification, so as to serve the intelligent prediction of deep and concealed gold deposits.

[0014] Optionally, the multi-dimensional reasoning capability includes: based on the verified gold knowledge unit set, using a graph neural network to simulate the spatial probability diffusion and resource aggregation process of mineralization, generating a probability target area of ​​concealed ore bodies in three-dimensional geological space; transforming the probability target area of ​​concealed ore bodies into a multi-dimensional exploration strategy scheme that includes drilling locations, survey line deployment, and sampling priorities; based on the multi-dimensional exploration strategy scheme, using newly acquired exploration data as feedback information, automatically reconstructing and enhancing the corresponding causal paths and spatiotemporal patterns in the gold exploration knowledge graph, forming a self-evolving intelligent exploration closed loop. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram illustrating an application scenario provided in one embodiment of this application; Figure 2 A flowchart illustrating a method for heterogeneous data parsing and knowledge graph construction in the gold industry, provided as an embodiment of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0018] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0019] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.

[0020] Existing exploration knowledge modeling methods based on a single data source or simple statistical correlations are difficult to achieve semantic fusion and spatiotemporal coupling analysis of multimodal heterogeneous geological data. They lack the ability to deeply characterize the complex causal mechanisms and dynamic evolution laws of metallogenic systems, resulting in static, shallow, and weakly interpretable knowledge graphs that cannot effectively support high-confidence intelligent inference of the spatial location and genetic mechanisms of deep concealed ore bodies.

[0021] Based on this, this application provides a method for heterogeneous data parsing and knowledge graph construction in the gold industry. First, a multimodal heterogeneous raw dataset of gold deposits, consisting of geological reports, exploration images, and geophysical and geochemical data, is acquired. Through semantic parsing and cross-modal verification based on domain ontology, this dataset is transformed into a semantically unified and logically consistent multimodal gold geological knowledge information set. Then, this information set undergoes deep correlation mining by fusing causal inference and spatiotemporal evolution patterns to generate candidate knowledge units containing mineralization mechanisms. Subsequently, using known typical gold deposits as anchors, the candidate units are subjected to credibility backtracking verification and correlation strength screening to form a high-confidence verification knowledge unit set. Finally, based on the verification unit set, a knowledge graph is modeled and iteratively constructed to form a gold exploration knowledge graph serving intelligent prediction of deep and concealed deposits, which is then output to industry researchers. This solution, through systematic knowledge extraction, verification, and graph construction, elevates multi-source heterogeneous exploration data into machine-understandable knowledge that combines causal logic and spatiotemporal constraints. This greatly enhances the depth of knowledge utilization and the interpretability of reasoning in deep mineral exploration, providing a reliable knowledge-driven new paradigm for overcoming the bottleneck of intelligent prediction of concealed ore bodies.

[0022] Figure 1This is a schematic diagram illustrating an application scenario provided by this application. In the process of heterogeneous data parsing, the method provided in this application greatly enhances the depth of knowledge utilization and the interpretability of reasoning in deep mineral exploration, providing a reliable knowledge-driven new paradigm for overcoming the bottleneck of intelligent prediction of concealed ore bodies.

[0023] Specifically, the method of this application is applied to any server that communicates with a large gold data model. Through this server, the method obtains a multimodal heterogeneous raw dataset of gold deposits provided by the large gold data model. First, it acquires a multimodal heterogeneous raw dataset of gold deposits composed of geological reports, exploration images, and geophysical and geochemical data. Then, through semantic parsing and cross-modal verification based on domain ontology, it transforms this dataset into a semantically unified and logically consistent multimodal gold geological knowledge information set. Next, it performs deep correlation mining on this information set, fusing causal inference and spatiotemporal evolution patterns, to generate candidate knowledge units containing mineralization mechanisms. Subsequently, using known typical gold deposits as anchors, it performs credibility backtracking verification and correlation strength screening on the candidate units, forming a high-confidence verification knowledge unit set. Finally, based on the verification unit set, it models and iteratively constructs a knowledge graph to form a gold exploration knowledge graph serving intelligent prediction of deep and concealed deposits, and outputs it to industry researchers.

[0024] For specific implementation details, please refer to the following examples.

[0025] Figure 2 This is a flowchart illustrating a method for heterogeneous data parsing and knowledge graph construction in the gold industry, provided in one embodiment of this application. The method of this embodiment can be applied to servers in the above scenarios. Figure 2 As shown, the method includes: S201. Obtain the original multimodal heterogeneous dataset of gold mines, perform semantic normalization parsing based on domain ontology and cross-modal consistency verification on the original multimodal heterogeneous dataset of gold mines, and generate a multimodal gold geological knowledge information set.

[0026] A multimodal heterogeneous raw dataset for gold mines can be a collection of raw data of varying forms and structures generated during gold mine exploration and research. Specifically, it includes unstructured text, exploration geological images, and geophysical and chemical data, sourced from a large-scale gold data model. Semantic normalization parsing involves mapping and labeling different terms, symbols, or representations describing the same geological entity or phenomenon from heterogeneous raw data to corresponding standard concepts and relationships in an ontology. This aims to eliminate terminological ambiguity and achieve semantic interoperability among multi-source data. Cross-modal consistency verification, based on semantic normalization, establishes association rules between different modalities of data (such as text descriptions, image features, and numerical anomalies) using a unified spatial coordinate system, geological strata, or time series as a benchmark, and then performs logical comparison and mutual verification. A multimodal gold geological knowledge information set can be a structured information set with unified semantic identifiers and key information cross-verified by multi-source evidence, generated after semantic normalization parsing and cross-modal consistency verification.

[0027] Specifically, in the field of gold geological exploration, data exhibits typical characteristics of being multi-source, multi-modal, and multi-scale. Traditional data processing methods often treat various types of data in isolation, such as extracting keywords from text or performing statistical interpolation only on geophysical and geochemical data. This approach severs the inherent connections between data that describe the same geological entity or process, leading to the formation of "data silos" and "semantic gaps." For example, the "fault structure" described in a text report may spatially match the "linear anomaly zone" revealed by geophysical data, and both indicate "hydrothermal migration channels." However, traditional methods struggle to automatically establish such cross-modal connections with deep geological semantics. This step aims to fundamentally address this problem by introducing a gold geological ontology—a standardized conceptual system that formally defines various entities (such as "granite bodies" and "gold-bearing quartz veins"), attributes (such as "occurrence" and "grade"), and relationships (such as "intruding into" and "controlling") within the gold mineralization system—as the cornerstone of semantic understanding. Semantic normalization parsing utilizes technologies such as natural language processing and computer vision to map geological descriptions in unstructured text, geological structures in images, and anomaly patterns in geophysical and geochemical data onto the standard concepts defined in the ontology, eliminating terminological ambiguity and achieving "synonymization." Cross-modal consistency verification further uses spatial location and geological stratigraphy as links to actively examine whether descriptions of different geological objects from text, images, and data are logically consistent or contradictory, thereby filtering out information with consistent evidence chains and high credibility. The necessity of this step lies in its foundational role in transforming raw "data" into usable "knowledge," providing standardized information input that is semantically clear, reliably related, and of controllable quality for subsequent deep knowledge mining.

[0028] S202. Based on the multimodal gold geological knowledge information set, conduct in-depth correlation mining of causal inference and spatiotemporal evolution patterns to generate a candidate gold knowledge unit set with causal logic and spatiotemporal constraints.

[0029] Fusion causal inference, in knowledge mining, not only identifies statistical correlations between geological elements but also focuses on using causal discovery algorithms (under temporal and spatial proximity constraints) to analyze the potential driving and response relationships between geological events, thereby revealing the "cause-effect" chain in mineralization. Spatiotemporal evolution patterns refer to the extraction and formal characterization of the spatial distribution patterns (such as zonation and symbiotic combinations) and temporal evolution sequences (such as mineral formation sequence and tectonic activity periods) of geological entities (such as rock masses and structures) or attributes (such as elemental content). Candidate gold knowledge unit sets can be collections of basic knowledge blocks with "causal logic" and "spatiotemporal constraints" attributes, initially mined from multimodal knowledge information sets through fusion causal inference and spatiotemporal evolution pattern analysis.

[0030] Specifically, traditional geological correlation analysis often relies on expert empirical induction or simple statistical correlation, making it difficult to reveal the complex causal mechanisms and dynamic spatiotemporal evolution laws controlling mineralization. For example, it is known that "alteration" and "mineralization" are spatially associated and correlated, but clarifying the process of "tectonic activity providing a channel → hydrothermal migration triggering alteration → changes in physicochemical conditions leading to mineral precipitation" is causal. The core purpose of this step is to go beyond shallow correlations and uncover the deep-seated dynamics and evolutionary paths driving mineralization. Fusion causal inference aims to construct an event causal chain characterizing the mineralization process from information with timestamps or relative time sequences (such as age data of different rock masses and the chronological order of tectonic activities), using causal discovery algorithms, and combining the geological a priori knowledge of causal laws (such as "intrusive events precede the alteration of the surrounding rocks they cause"),. Meanwhile, spatiotemporal evolution pattern analysis focuses on the distribution patterns of geological elements (such as rock masses, faults, and ore bodies) in three-dimensional space and their changes over time. Through spatial statistics and sequence analysis, patterns such as "ore bodies are distributed at equal intervals along tectonic zones" and "alteration zoning changes regularly with depth" are extracted. By attributing causal logic and spatiotemporal constraints to relevant geological entities, "candidate gold knowledge units" are formed. This step is crucial, as it represents a leap from static factual description to dynamic process understanding. The generated knowledge units are no longer isolated points but rather three-dimensional knowledge fragments embedded with "why" and "how," serving as the core raw material for constructing knowledge graphs with reasoning capabilities.

[0031] S203. Based on the candidate gold knowledge unit set, and using known typical gold deposits as verification anchors, the credibility of the knowledge units is backtracked for verification and the causal relationship strength is screened to generate a verification gold knowledge unit set.

[0032] Known typical gold deposits refer to gold deposit examples that have undergone long-term exploration and research, have clear geological characteristics, relatively complete genetic models, and enjoy consensus and benchmark significance within the industry, such as the Zhenyuan Gold Mine. Credibility backtesting verification refers to the process of matching and comparing the established geological laws of "candidate gold knowledge units" with those of "known typical gold deposits" using patterns and examples, assigning a quantitative credibility score to the candidate knowledge units based on their degree of conformity with typical deposit knowledge. Causal correlation strength screening refers to, based on credibility backtesting verification, further calculating and evaluating the robustness of causal relationships within candidate knowledge units based on indicators such as causal relationship matching coverage, logical consistency, and statistical significance, and then selecting knowledge units with high-strength correlations based on preset thresholds. The verified gold knowledge unit set refers to the set of gold knowledge units with high credibility and strong causal correlations retained after "credibility backtesting verification" and "causal correlation strength screening."

[0033] Specifically, causal chains and patterns mined from data may contain spurious correlations or overly localized specificities, and directly using them for prediction would pose significant risks. Geology is essentially a historical science based on observation and verification. Therefore, this step introduces "known typical gold deposits" as objective verification anchors, necessary to provide geological constraints and quality benchmarks for the results of data-driven mining. By matching candidate knowledge units (such as a causal hypothesis of "shear zone activation → gold precipitation") with verified knowledge templates from multiple typical deposits, the extent to which the knowledge unit reproduces the patterns of real metallogenic systems can be assessed. For example, if a causal chain can be found in multiple typical deposits of different types, its universality and credibility are high; conversely, if it only appears in a few cases, it may be specific. The causal correlation strength screening further uses indicators such as the breadth and consistency of the matching to quantify and score the correlation, filtering out low-confidence, contradictory, or redundant knowledge fragments. This process essentially involves placing the "black box" output of big data mining into a solid prior knowledge base of geology for "white box" verification and purification, ensuring that the final knowledge model not only has the novelty of data insights but also does not deviate from basic geological principles.

[0034] S204. Based on the verified gold knowledge unit set, perform knowledge graph modeling and iterative construction to generate a gold exploration knowledge graph that serves intelligent prediction of deep and concealed gold deposits.

[0035] A gold exploration knowledge graph can refer to a large-scale domain knowledge base that uses "verification of gold knowledge units" as basic elements and adopts a graph structure (composed of nodes, attributes, and edges) for systematic organization and storage.

[0036] Specifically, the validated, high-quality knowledge units remain scattered. The necessity of this step lies in systematically organizing these units into an interconnected, machine-understandable knowledge graph, thereby forming a comprehensive knowledge base usable for complex reasoning. The modeling process defines the graph pattern (node ​​type, edge relationship type) based on the domain ontology, adding validated knowledge units as instances to the graph to form a vast network containing "entity-relationship-attribute". The unique value of this graph lies in its explicit orientation of "serving intelligent prediction of deep and concealed gold deposits". Therefore, its construction is iterative: when new exploration data (such as new borehole information) or new research results are input, it can trigger automatic or semi-automatic updates and expansions of the graph, allowing its knowledge base to continuously grow and evolve. The resulting gold exploration knowledge graph can integrate the causal and spatiotemporal knowledge contained in multi-source heterogeneous data, providing a computable, reasonable, and traceable knowledge platform. It directly supports intelligent prediction tasks such as delineation of mineralized target areas and probability estimation of resource quantities, and models and systematizes the collective wisdom and experience of exploration experts, greatly improving the ability to predict deep concealed ore bodies and the scientific nature of exploration decisions.

[0037] The method provided in this embodiment first acquires a multimodal heterogeneous raw dataset of gold deposits, consisting of geological reports, exploration images, and geophysical and geochemical data. Through semantic parsing and cross-modal verification based on domain ontology, this dataset is transformed into a semantically unified and logically consistent multimodal gold geological knowledge information set. Then, this information set undergoes deep correlation mining by fusing causal inference and spatiotemporal evolution patterns to generate candidate knowledge units containing mineralization mechanisms. Subsequently, using known typical gold deposits as anchors, the candidate units are subjected to credibility backtracking verification and correlation strength screening to form a high-confidence verified knowledge unit set. Finally, based on the verified unit set, a knowledge graph is modeled and iteratively constructed to form a gold exploration knowledge graph serving intelligent prediction of deep and concealed minerals, which is then output to industry researchers. This solution, through systematic knowledge extraction, verification, and graph construction, elevates multi-source heterogeneous exploration data into machine-understandable knowledge with both causal logic and spatiotemporal constraints. This greatly improves the depth of knowledge utilization and the interpretability of reasoning in deep mineral exploration, providing a reliable new knowledge-driven paradigm for overcoming the bottleneck of intelligent prediction of concealed mineral bodies.

[0038] In some embodiments, the multimodal heterogeneous raw dataset for gold mines includes unstructured text, exploration geological images, and geophysical and chemical data. Based on the multimodal heterogeneous raw dataset, entity recognition, relation extraction, and semantic annotation based on the gold geology ontology are performed. Natural language processing technology is used to achieve semantic normalization of the same geological concepts in data from different sources, thereby constructing semantically normalized information. Based on the multimodal heterogeneous raw dataset, cross-modal correlation analysis is performed. Multimodal fusion technology is applied to identify the phenomenon of identical descriptions and logic in different modal data when describing the same geological object or event, thereby constructing cross-modal consistency information. Based on the semantically normalized information and the cross-modal consistency information, a multimodal gold geology knowledge information set is generated.

[0039] Unstructured text can refer to textual materials such as geological survey reports, scientific research papers, and exploration logs, which are described in natural language and lack a fixed format. Exploration geological images can refer to image data containing rich spatial and visual geological information, such as borehole core scans, geological profiles, and field outcrop photographs. Geophysical and chemical data can refer to quantitative datasets that characterize subsurface physical properties and elemental distributions in numerical and curve forms, obtained through geophysical exploration (such as gravity and magnetic methods) and geochemical sampling analysis. Semantic normalization information refers to unambiguous structured information formed by mapping the diverse but identical geological concepts from the aforementioned heterogeneous data to a standard geological terminology system. Cross-modal consistency information refers to correlated information supported by multi-source evidence, generated through correlation and comparison of descriptions of the same geological object from different modal data, after mutual verification and contradiction resolution.

[0040] Specifically, traditional geological data analysis methods operate in isolation, processing textual, image, and geophysical / geochemical data in fragmented ways, leading to severe information fragmentation and semantic gaps. For example, the "NE-trending fault zone" described in the geological report, the "linear tectonic traces" identified in the profile image, and the "gravity gradient zone" revealed by geophysical data are treated as three separate pieces of information under traditional methods. They cannot be automatically linked and identified as the same geological entity, nor can their descriptions be assessed for contradictions (e.g., the text describes a "fault zone width of 50 meters," while the image interpretation only states "30 meters"). This fragmented processing results in a lack of a reliable and consistent knowledge base for subsequent analysis. To address the above issues, this step first utilizes a large-scale model pre-trained on millions of geological documents to deeply analyze unstructured text, extracting semantic relationships such as "granite body intrudes into Early Cretaceous strata." Simultaneously, computer vision technology is employed to identify stratigraphic boundaries and faults in exploration images, resolving them into vectorized spatial objects with coordinates and lithological annotations. For geophysical and geochemical data, geological expertise (e.g., "lower limit of gold anomaly is 0.05 ppm") transforms raw values ​​into response pattern descriptions such as "high Au anomaly zone." Subsequently, the entire information is mapped to a unified conceptual node using the Gold Geological Domain ontology (e.g., normalizing "granite body," "batterwater," and "acidic intrusive rock" to the standard concept "granite"), generating semantically normalized information. Furthermore, using unified spatial coordinates or geological strata as a benchmark, a cross-validation bridge is constructed to automatically compare the logical consistency of information from different modalities (for example, verifying whether the location of the alteration zone described in the text overlaps spatially with the geochemical anomaly area and the fading zone in the image). By verifying and enhancing the credible description, contradictions are marked and addressed, and finally, cross-modal consistent information is integrated to generate a high-quality multimodal gold geological knowledge information set.

[0041] The method provided in this embodiment systematically solves the two major bottlenecks of "semantic ambiguity" and "evidence contradiction" among multi-source heterogeneous geological data, providing semantically clear, evidence-consistent, and high-quality structured information input for subsequent deep knowledge mining, and laying the foundation for the credibility of knowledge throughout the entire process.

[0042] In some embodiments, based on unstructured text, a large gold model pre-trained on a gold domain corpus is used for deep analysis and entity relation extraction to generate a semantic relation representation of geological text. Based on exploration geological images, computer vision technology is used to identify geological structures and lithological units, and these are parsed into vectorized spatial objects with geological semantic annotations. Based on geophysical and chemical data, geological domain expertise is used to map the abnormal features in the geophysical and chemical data into response pattern descriptions representing geological, geographical, and geochemical information. The semantic relation representation of geological text, vectorized spatial objects, and response pattern descriptions are uniformly mapped to standard concept nodes of the gold geological domain ontology to eliminate conceptual ambiguities within various types of data, thereby constructing semantically normalized information.

[0043] The "Golden Big Model" can refer to a specialized large-scale language model pre-trained on massive corpora of literature and reports in the field of gold geology, possessing the ability to deeply understand geological professional texts and identify geological entities and relationships. Geological text semantic relationship representation can refer to geological facts extracted from unstructured text using the aforementioned Golden Big Model and expressed in a structured form (such as triplets: subject-relationship-object). Computer vision technology can refer to a series of algorithms and models (such as U-Net) applied to geological exploration image analysis, whose core function is to automatically identify, segment, and interpret geological elements in images. Vectorized spatial objects can refer to vector graphics with semantic labels that encode geological elements (such as faults and lithological boundaries) identified from exploration geological images using computer vision technology, along with their geometric shape, spatial location, and geological attributes (such as "silicification alteration"). Response pattern description can refer to the interpretation and conversion of raw observation data (such as magnetic anomaly values ​​and elemental contents) into qualitative or semi-quantitative descriptions with geological significance based on geophysical and geochemical expertise (such as "high magnetic anomaly zones may correspond to concealed rock masses" and "gold and arsenic elemental anomalies indicate hydrothermal activity centers"). Standard concept nodes can be the most basic and authoritative knowledge units defined in the ontology of gold geology, used to uniquely and unambiguously represent a geological entity, attribute or relationship.

[0044] Specifically, traditional geological data interpretation relies heavily on expert manual interpretation, which suffers from fundamental flaws such as strong subjectivity, low efficiency, and inconsistent standards. For example, different experts may describe anomalies in the same geophysical data as "gradually changing gradient zones" or "locally high-value areas," leading to confusion in the basic information for subsequent analysis; the interpretation of rock mass boundaries on a geological profile map also varies from person to person. This non-standardized raw information cannot be directly understood and associated by machines, constituting the first barrier to automated knowledge construction. To address these issues, this step first utilizes the Gold Model to perform deep analysis of unstructured text: given a geological report, the model can automatically extract semantic relationships in the geological text, such as "(Jiaojia Fault Zone) - (Control) - (Linglong Gold Mine Field)". Simultaneously, deep learning-based image segmentation and object detection algorithms are used to process the exploration geological images: identifying all stratigraphic boundaries and fault lines in the images and converting them into vectorized spatial objects with accompanying coordinates and attributes (such as "fault, nature: compressional-shear"). For geophysical and chemical data, a built-in professional knowledge rule base is invoked. For example, when the Au content exceeds the background value (e.g., 0.05 ppm) by three consecutive points, a response pattern description of "strong gold anomaly" is automatically generated and associated with its spatial location. Finally, all the above outputs—whether it's the relationships extracted from text, the objects recognized in images, or the patterns of data interpretation—are mapped to a unified standard concept node in the ontology of gold geology. For example, "granite body" in text, "flesh-red batholith" in images, and "high tungsten-molybdenum anomaly zone" in geochemical models are all mapped to the standard node "Yanshanian granite" in the ontology, thereby completely eliminating terminological ambiguity and constructing machine-understandable and computable semantically normalized information.

[0045] The method provided in this embodiment realizes the transformation from relying on human experience to automated and standardized semantic parsing, providing unambiguous and computable structured representations for multimodal data, which is an indispensable underlying data preparation step for building high-quality knowledge graphs.

[0046] In some embodiments, spatial coordinates or geological strata are used as a unified reference to establish a cross-validation bridge connecting unstructured text, exploration geological images, and geophysical and chemical data. Through the cross-validation bridge, descriptions of the same geological object or process from different modal data are compared and analyzed at the logical and attribute levels to identify potential consistency phenomena. Based on the identified consistency phenomena, descriptions from different modal sources are mutually corroborated and enhanced to generate corroborating information with multi-source evidence support and logical self-consistency. All corroborating information generated through mutual corroboration is integrated to construct cross-modal consistency information.

[0047] Cross-validation bridges can refer to information channels established based on a unified spatial coordinate system (such as geodetic coordinates) or a standard geological stratigraphic sequence (such as specific stratigraphic numbers) for associating and comparing descriptions of the same geological entity or event from unstructured text, exploration geological images, and geophysical and chemical data. Statement consistency refers to situations where, after comparison through the aforementioned cross-validation bridge, descriptions of the same geological object or process from different modalities are mutually consistent and supportive in terms of attributes (such as location, scale, and properties) or logic. Corroborating information refers to comprehensive information units generated after integrating and enhancing the identified statement consistency phenomena, containing multi-source evidence chains and their mutual support relationships.

[0048] Specifically, in traditional geological research, data from different professional fields are often analyzed in isolation, lacking a systematic cross-validation mechanism. This leads to conclusions drawn from a single data source that may be biased or even erroneous, and these errors are difficult to detect. For example, geophysical interpretation may infer the existence of a concealed fault, but geological mapping (images) may not reveal any surface signs, and borehole cores (textual descriptions) may not record corresponding structural features. Such contradictions are often overlooked or discovered late in traditional workflows, resulting in subsequent exploration deployments based on unreliable information and causing significant waste. To address these issues, this step first uses a fine spatial coordinate grid or standard geological stratigraphic codes as a unified correlation benchmark to construct a cross-validation bridge. For example, the area with "gold anomaly concentration > 1.0 ppm" in geochemical data (spatial range X), the "silicified alteration zone" interpreted on the geological profile (spatial range Y), and the "altered rocks in the footwall of the Jiaojia fault zone" described in the exploration report (spatial location Z) are correlated through spatial overlay analysis. The system automatically performs logical and attribute comparisons: if the spatial locations of the three elements highly overlap (e.g., overlap area exceeds 70%) and the attribute descriptions are compatible ("silicification alteration" and "altered rock" have the same geological meaning), it is identified as a consistent description. For such phenomena, the system generates a related corroborating information, integrates all evidence sources, and marks their consistency strength. Conversely, if contradictory descriptions are found (e.g., the image shows the fault dips northwest, while the text record dips northeast), an inconsistency alarm is triggered, prompting experts to review the information. Finally, all mutually corroborated, logically consistent related corroborating information is integrated to construct a cross-modal consistent information system with solid evidence and explicit contradictions.

[0049] The method provided in this embodiment establishes an active cross-validation and contradiction detection mechanism for multi-source data, which significantly improves the reliability and consistency of the original knowledge information and lays a solid evidentiary foundation for the subsequent construction of a high-confidence knowledge graph.

[0050] In some embodiments, based on a multimodal gold geological knowledge information set, the temporal sequence and spatial association of geological events are analyzed to mine and construct event causal chains that characterize the mineralization process. Simultaneously, the spatial distribution patterns and temporal evolution characteristics of geological elements, including lithological units, tectonic traces, alteration zones, and mineralized bodies, are analyzed in the region to extract spatiotemporal evolution pattern information. The logical relationships of the event causal chains and the constraint attributes of the spatiotemporal evolution pattern information are jointly assigned to relevant knowledge units to form a candidate gold knowledge unit set.

[0051] A causal chain of events can refer to a logical chain, constructed based on the temporal sequence and spatial association of geological events, used to characterize the causal dependencies between key geological events (such as intrusion, tectonic activity, hydrothermal alteration, and mineral precipitation) during mineralization. Spatiotemporal evolution model information can refer to comprehensive model information extracted by analyzing the spatial distribution patterns and temporal evolution characteristics of geological elements (including lithological units, tectonic traces, alteration zones, and mineralized bodies) in a region, depicting how the structure of the mineralization system dynamically evolves over geological time.

[0052] Specifically, traditional studies of metallogenic regularities often separate the "causal process" from the "spatiotemporal pattern." They either merely list isolated geological events while ignoring their inherent causal driving mechanisms, or they only statically describe mineralization distribution while neglecting its dynamic evolutionary history. This results in knowledge units lacking logical coherence and spatiotemporal constraints, making it difficult to truly reflect the complexity of metallogenic systems. For example, knowing only that a certain area contains "granite bodies" and "gold ore bodies" fails to answer the core scientific question of "how intrusion of the rock mass leads to gold mineralization"; or only delineating the range of "alteration zones" without understanding how they migrate and expand spatially over time. This fragmented and static knowledge cannot support effective prediction of concealed ore bodies. To address the above issues, this step involves conducting two types of in-depth analysis in parallel from a multimodal collection of geological knowledge on gold: First, it systematically sorts out geological event instances with clear spatiotemporal labels (such as the "Early Cretaceous intrusion event" located at coordinate X, age YMa). By analyzing the spatiotemporal proximity of events (such as the "hydrothermal alteration event" occurring slightly later near the contact zone of the intrusion) and their chronological order, it uncovers potential causal relationships and constructs event causal chains such as "regional tectonic stress field changes → triggering fault activity → leading to deep magma upwelling and intrusion → accompanied by hydrothermal activity and alteration → ultimately precipitating and forming minerals in favorable locations." Second, it quantitatively analyzes various geological elements: using spatial statistical methods (such as kernel density estimation) to analyze the spatial aggregation of alteration zones (such as the formation of high-density areas at fault intersections), and combining this with their formation age data, it constructs a temporal evolution sequence (such as early planar potassic alteration, mid-term linear silicification, and late-term vein-type carbonatization), thereby extracting spatiotemporal evolution pattern information of "alteration zoning evolving orderly from the center to the edge and from the deep to the shallow over time." Ultimately, the event causal chain representing the process logic and the spatiotemporal evolution pattern information describing the structural evolution are jointly endowed with the relevant knowledge unit (such as a complete "porphyry gold mineralization system" knowledge unit), so that it simultaneously possesses the causal explanation of "why it was formed" and the spatiotemporal constraints of "how it is distributed", thereby forming a high-quality candidate gold knowledge unit set.

[0053] The method provided in this embodiment achieves the organic integration of "process" and "pattern" research in metallogenesis, enabling the constructed knowledge units to simultaneously contain causal logic and spatiotemporal constraints, laying a core foundation for the subsequent generation of knowledge graphs that have both scientific explanatory power and exploration guidance.

[0054] In some embodiments, a set of geological event instances with clear spatiotemporal attributes is extracted from a multimodal gold geological knowledge information set. The set of geological event instances includes intrusive events, tectonic activity events, hydrothermal alteration events, and mineral precipitation events. Based on a causal discovery algorithm, causal dependency mining based on temporal constraints and spatial correlation is performed on the set of geological event instances to generate a preliminary causal network topology. Based on a pre-set gold mineralization theory and geodynamic prior knowledge base, logical constraints and path verification are applied to the preliminary causal network topology to screen out candidate causal paths that conform to the laws of geological evolution. For each causal edge in the candidate causal path, its confidence score is analyzed. The confidence score integrates temporal tightness, spatial proximity, and the strength of corroboration in cross-modal consistency information to construct an event causal chain with multi-dimensional confidence assessment.

[0055] A collection of geological event instances refers to a set of geological event instances with clear spatiotemporal attributes (such as geological age and coordinate location) extracted from a multimodal gold geology knowledge information set. Intrusive rock mass events can be instances of underground magma upwelling and cooling to form intrusive rock masses. Tectonic activity events can be instances of geological processes where crustal stress causes deformation of rock strata, such as fracturing, folding, or shearing. Hydrothermal alteration events can be instances of geological processes where high-temperature fluids react chemically with surrounding rocks, leading to changes in the composition and structure of the original rock minerals and the formation of new mineral assemblages. These instances have alteration types (such as silicification and sericitization), zoning characteristics, and occurrence timeframes, and are important indicators of ore-forming fluid activity. Mineral precipitation events can be instances of geological processes where ore-forming materials (such as gold) precipitate from fluids and accumulate to form ore or mineralization. These instances have mineralization types (such as vein-like and disseminated), grade distribution, and formation age, and are direct manifestations of the final stage of mineralization. Causal discovery algorithms can be a class of algorithms used to automatically mine potential causal relationships between variables from observational data. In this step, it specifically refers to methods applied to temporal and spatial data of geological events to discover causal dependencies between events. The preliminary causal network topology can be a network structure diagram describing the preliminary causal links between events, generated after analyzing a set of geological event instances using the aforementioned causal discovery algorithm. The pre-built gold mineralization theory and geodynamics prior knowledge base can be a structured knowledge base storing accepted mineralization models, geodynamic process laws, and basic geological constraints within the field, used for logical verification of data-driven results. Candidate causal paths can be sequences of potential causal chains that conform to the basic laws of geological evolution after being filtered by the prior knowledge base. The confidence score can be a comprehensive numerical value that quantitatively evaluates the reliability of each causal edge (i.e., cause-effect pair) in the candidate causal path, integrating multi-dimensional information such as temporal, spatial, and evidentiary support.

[0056] Specifically, traditional causal inferences of geological events based on expert experience, while utilizing domain knowledge, lack systematic data verification and uncertainty quantification, are easily influenced by subjective cognitive biases, and the reasoning process is opaque and unreproducible. For example, professionals might conclude a direct causal relationship between intrusion and mineralization solely based on the statement "intrusion precedes mineralization," ignoring potentially missing key links (such as the activation and migration of hydrothermal systems) or failing to consider the existence of spatial material migration channels. This "gut feeling" approach to causal chain construction makes the subsequent knowledge graph's reasoning foundation fragile and unable to cope with complex geological situations. To address these issues, this step first extracts event instances with precise spatiotemporal tags from the knowledge set, such as: "Linglong Intrusion Event" (age 120±2 Ma, coordinate range X), "Jiaojia Fault Main Activity Period Event" (age 115±5 Ma, spatial trajectory Y), "Sericite Alteration Event" (age 110±3 Ma, range Z), and "Gold Mineralization Event" (age 105±4 Ma, location W). Subsequently, a causal discovery algorithm combining temporal constraints and spatial proximity (such as a constraint-based PC algorithm) is applied to analyze these event instances, automatically inferring a preliminary causal network topology such as "rock mass intrusion event → tectonic activity event → hydrothermal alteration event → mineral precipitation event". Next, the system calls a pre-set knowledge base of gold mineralization theory and geodynamics for logical verification: for example, the knowledge base rules may stipulate that "large-scale mineral precipitation requires pre-existing structures as the ore-guiding channel." If the path generated by the algorithm directly links "mineral precipitation" to "rock mass intrusion" and skips "tectonic activity", the path will be marked as suspicious. Paths that pass the verification become candidate causal paths. Finally, a confidence score is calculated for each causal edge in the path: for example, the edge "tectonic activity event → hydrothermal alteration event" receives a high confidence score (e.g., 0.92) because of its short time interval (e.g., 5 Ma), high spatial overlap between fault zones and alteration zones (e.g., 85%), and corroboration in multi-source data (text, images, geochemical exploration); conversely, if the evidence is weak, the score is lower. Thus, a traceable and questionable causal chain of events with multi-dimensional confidence assessment is constructed.

[0057] The method provided in this embodiment deeply integrates data-driven causal discovery with knowledge-driven logical verification and introduces quantitative confidence assessment, making the constructed event causal chain both objective and reliable and in line with geological laws, significantly improving the scientificity and interpretability of causal reasoning.

[0058] In some embodiments, spatial geometric distribution data and geological time series data of geological elements are extracted from a multimodal gold geological knowledge information set. Based on the spatial geometric distribution data, spatial topology analysis and density field simulation are used to quantitatively analyze the spatial aggregation, orientation and symbiotic combination laws of geological elements in the region, and generate spatial distribution patterns. Based on the geological time series data, geological age sequences and superposition and cutting relationships of different geological elements are constructed to infer their formation sequence and periodicity, and generate time series evolution sequences. The spatial distribution patterns and time series evolution sequences are coupled to generate spatiotemporal evolution pattern information that describes how the structure of the mineralization system dynamically evolves with geological time.

[0059] Spatial geometric distribution data refers to data describing the specific morphology, location, and extent of geological elements (such as rock masses, faults, alteration zones, and mineralization bodies) in geological space. Geological time-series data refers to data obtained through isotopic dating (such as U-Pb method, Ar-Ar method) or other dating methods, characterizing the specific time or time period of formation of geological events or geological elements. Spatial topological analysis can be a process that focuses on analyzing the spatial relationships between geological elements (such as alteration zones and faults, rock masses and surrounding rocks), such as inclusion, adjacency, and intersection. Density field simulation can be a process that uses methods such as kernel density estimation to convert discrete geological element point or surface data into continuous spatial density distribution surfaces to intuitively reveal their degree of aggregation in a region. Spatial distribution patterns can be a comprehensive description of the spatial aggregation (such as whether it is clustered), orientation (such as whether it extends along a specific tectonic line), and symbiotic combination patterns (such as specific lithologies often being associated with specific alterations) of geological elements in a region, quantified through spatial analysis methods. A geological age sequence refers to a sequence formed by arranging multiple geological time series data according to their chronological formation, directly reflecting the sequential order of different geological events or elements on the timeline. Overlay and cutting relationships refer to the physical relationships of geological bodies or structures in space, such as overlapping or cutting. A temporal evolution sequence can be a sequential expression of the chronological order of formation and possible periodicity (such as multiple phases of magmatic-hydrothermal activity) inferred by constructing and comparing geological age data of different geological elements and their overlay and other geological relationships.

[0060] Specifically, traditional geological research often employs a fragmented "space first, time later" approach when dealing with the spatial distribution and temporal evolution of elements, or only provides qualitative descriptions (such as "alteration zones are distributed along faults" or "mineralization has multiple phases"). This lacks a dynamic perspective that quantitatively couples spatial patterns with temporal series, resulting in static and fragmented patterns that fail to truly reflect the dynamic process of the mineralization system's origin and evolution. For example, merely delineating the extent of an alteration zone cannot answer how it expands outward from its center on a million-year scale, or how its internal mineral assemblage changes over time, making it difficult to accurately predict the hidden mineralization potential in its periphery or deeper areas. To address these issues, this step first extracts precise spatial geometric distribution data (such as polygon boundary coordinates) and high-precision geological time series data (such as a sericitization age of 120±2 Ma and a potassic age of 125±3 Ma) of the target geological element (e.g., the "Jiaojia-type alteration zone") from a multimodal gold geology knowledge information set. In the spatial dimension, spatial topological analysis is used to automatically identify the "adjacent" or "included" relationship between alteration zones and ore-controlling faults (e.g., 85% of alteration zone boundaries are less than 50 meters from faults). Spatial contour maps of alteration intensity are generated through density field simulation to quantify their accumulation cores (e.g., in high-density core areas, the alteration area can reach 70%). In the temporal dimension, a time-series evolution sequence of "potassic alteration → sericitization → silicification → carbonatization" is constructed, clarifying the age range of each stage (e.g., potassic alteration 125-120 Ma, sericitization 120-115 Ma) and their superposition relationships (late-stage silicified veins cutting early-stage sericitized zones). Ultimately, the spatial distribution patterns revealing spatial patterns (such as "alteration intensity decreases in a ring-like manner from the intersection of northeast-trending faults outwards") will be deeply coupled with the temporal evolution sequences depicting the time process (such as "hydrothermal activity evolves from high-temperature potassic alteration and sericitization in the early and middle stages to medium-low-temperature silicification and carbonatization in the late stages") to generate truly dynamic spatiotemporal evolution pattern information, such as "the hydrothermal alteration system migrates synchronously with time and space from deep to the hanging wall of northeast-trending faults and from high temperature to low temperature".

[0061] The method provided in this embodiment realizes a fundamental shift from describing static elements to depicting dynamic evolution processes. The extracted spatiotemporal coupling model profoundly reveals the structural evolution law of the ore-forming system and provides key constraints for predicting the spatiotemporal location of ore bodies.

[0062] In some embodiments, multiple known typical gold deposits with clear genetic types and distinct geological characteristics are selected as verification benchmarks, and their complete geological knowledge maps are extracted as verification templates. The event causal chain and spatiotemporal evolution pattern information are matched and mapped to the verified patterns in the verification templates one by one. Based on a machine learning classifier, the matching degree and coverage of each candidate knowledge unit in multiple verification templates are analyzed. The successfully matched units are marked as trustworthy, and their causal association strength is quantified according to the breadth and consistency of their matching. Finally, trustworthy knowledge units that pass backtracking verification and whose association strength is higher than a preset threshold are selected to form a set of verified gold knowledge units.

[0063] The verification template can be a standard knowledge structure constructed based on a complete geological knowledge map of known typical gold deposits, containing verified genetic patterns, event sequences, and spatial distribution patterns, used for retrospective verification of candidate knowledge units. Pattern matching and instance mapping can be the process of comparing and associating the event causal chains and spatiotemporal evolution patterns in candidate knowledge units with the established geological laws in the verification template at the structural and instance levels. Matching degree can refer to the degree of similarity between the candidate unit and a single verification template in key patterns. Coverage can refer to the extent to which the candidate unit has been verified in multiple different types of verification templates. Causal correlation strength can be a quantitative score of the reliability of causal relationships in candidate knowledge units by comprehensively considering matching degree, coverage, and consistency of multi-source evidence. The preset threshold can be the minimum threshold value for causal correlation strength set to screen high-confidence knowledge units.

[0064] Specifically, when constructing an exploration knowledge graph, directly incorporating candidate knowledge units that have not been tested in geological practice can lead to a weak foundation for the graph and unreliable inference results, potentially misleading exploration decisions. Traditional methods lack systematic verification processes or rely solely on subjective judgment based on individual expert experience, resulting in inconsistent verification standards, low efficiency, and an inability to quantify reliability assessments. For example, a "fault-controlled ore deposit" model derived from mining area A might be directly applied to mining area B without verification, even though mining area B is primarily controlled by the contact zone of the rock mass. This would lead to complete failure of predictions based on the flawed model. To address these issues, this step first selects several known deposits with typical genetic types and clear geological characteristics (such as "Jiaojia-type" fractured zone altered rock gold deposits and "Linglong-type" quartz vein gold deposits) as verification benchmarks, extracting their complete knowledge graphs established through years of exploration research as verification templates. Subsequently, the candidate knowledge unit set to be verified (such as the "tectonic-hydrothermal-mineralization" causal chain extracted from the new block) is matched one by one with these verification templates for pattern matching and instance mapping. For example, the "NE-trending fault activity event" of the new block is matched with the "main fault-controlled ore-forming structure" pattern of the Jiaojiashi template, and the "sericite alteration event" is mapped to the "main alteration type" instance of the template. On this basis, a machine learning classifier (such as support vector machine) is used to automatically analyze the matching degree (such as similarity 0.85) of each candidate unit with each template and the coverage in all templates (such as successful matching in 3 out of 5 templates), and the successfully matched units are marked as trustworthy. Finally, the system quantifies the causal association strength based on the breadth and consistency of the matching (for example, a "magmatic hydrothermal mineralization" causal chain that has obtained moderate or higher matching in 4 different types of ore deposit templates and has consistent multi-source evidence can have an association strength of 0.9), and selects units with a strength higher than a preset threshold (such as 0.7), thus forming a high-quality and highly reliable set of verified gold knowledge units.

[0065] The method provided in this embodiment establishes a systematic and quantifiable retrospective verification mechanism based on known mineral deposit templates, which fundamentally ensures the reliability and geological applicability of the core knowledge incorporated into the knowledge graph, laying a solid and credible foundation for intelligent prediction.

[0066] In some embodiments, an initial graph network is constructed using a set of verified gold knowledge units as nodes and relationships, and information on event causal chains and spatiotemporal evolution patterns. A dynamic graph engine is then used to perform knowledge fusion, reasoning completion, and logical conflict resolution on the initial graph network to generate a gold exploration knowledge graph that satisfies consistency constraints and has multi-dimensional reasoning capabilities. New exploration evidence is then input into the gold exploration knowledge graph to trigger its adaptive iterative updates and verification, serving intelligent prediction of deep and concealed gold deposits.

[0067] The initial knowledge graph network can be a graph with a basic topological structure, initially constructed using knowledge units from the verification gold knowledge unit set as nodes and logical and spatiotemporal relationships such as "causality," "symbiosis," and "sequence" contained in the event causal chain and spatiotemporal evolution pattern information as edges. The dynamic knowledge graph engine can be a core software module capable of real-time online processing of the knowledge graph, performing operations such as knowledge fusion, reasoning completion, and conflict resolution. Knowledge fusion refers to the engine comparing, deduplicating, and merging descriptions from different sources targeting the same geological entity or relationship but potentially differing in attributes, forming a unified and authoritative expression. Reasoning completion refers to the engine's ability to automatically infer and add hidden relationships or entities that are not explicitly stated but logically necessary, based on existing causal relationships and spatiotemporal constraints in the graph. Logical conflict resolution refers to the engine's automatic adjudication and correction process based on rules such as evidence strength and source reliability when conclusions derived from different paths in the graph contradict each other (e.g., relationship A indicates "rock mass predates fracture," while relationship B indicates "fracture predates rock mass"). Consistency constraints can be a set of rules that must be followed during knowledge fusion, reasoning, and resolution to ensure that the internal logic of the graph is self-consistent and without contradictions. Multi-dimensional reasoning capability can mean that the graph not only supports "why" reasoning based on causal chains and "where / when" reasoning based on spatiotemporal patterns, but also predictive reasoning based on resource aggregation simulation to determine "how much / how rich".

[0068] Specifically, traditional knowledge graph construction often stops at static knowledge listing and linking, forming a "dead" database. Its fatal flaw lies in the lack of intelligent processing of inconsistent knowledge, reasoning to complete hidden knowledge, and the ability to self-evolve based on new evidence. For example, when descriptions of "metallogenic epoch" extracted from different documents conflict, static graphs cannot automatically resolve the issue; or, the graph records "rock body" and "ore body" but does not clearly define their "metallogenic parent rock" relationship, requiring manual supplementation; more importantly, new drilling mineralization information cannot be automatically integrated into the graph and correct existing knowledge, causing the graph to quickly become outdated and unable to guide subsequent exploration. To address these problems, this step first uses a highly reliable set of verified gold knowledge units as material to construct an initial graph network. Subsequently, the core dynamic graph engine begins operation: In the knowledge fusion phase, it compares in-depth descriptions of the "Jiaojia Fault" from geological reports and scientific research papers, selecting more precise values ​​(such as dip angles of 150°-160° and dip heights of 30°-40°) for merging; in the reasoning completion phase, based on the existing relationship between "hydrothermal alteration surrounding the rock mass" and "the ore body occurs within the alteration zone," it automatically infers and adds the implicit "geological parent rock" relationship that "this rock mass provides hydrothermal and material sources for mineralization"; in the logical conflict resolution phase, when two conflicting data points regarding the age of the same alteration event, 115 Ma and 125 Ma, are found, the engine automatically decides to adopt 115 Ma and updates relevant edges based on the data source (such as high-precision dating reports taking precedence over regional comparison inferences) and cross-modal consistency information (such as the matching degree of the rock mass age corresponding to that age). This generates a living gold exploration knowledge graph that satisfies consistency constraints and possesses multi-dimensional reasoning capabilities. Finally, the system establishes a closed loop: when new exploration evidence (such as the discovery of new mylonite-type mineralization at the depth of borehole ZK001 in the new target area) is input, the engine automatically matches and verifies it with the "tectonic control of mineralization" pattern in the map. If they match, the confidence of the corresponding causal path is enhanced (such as increasing the correlation strength between "ductile-brittle shear zone activity" and "gold precipitation" from 0.7 to 0.85). If they do not match, a new round of conflict detection and map structure adjustment is triggered to achieve adaptive iterative updates of the map.

[0069] The method provided in this embodiment constructs a dynamic knowledge graph with self-fusion, reasoning, purification and evolution capabilities, which fundamentally overcomes the rigidity of traditional static knowledge bases, enabling it to continuously absorb new knowledge and maintain its optimal state, providing a powerful, reliable and ever-evolving core reasoning engine for intelligent prediction of deep and hidden minerals.

[0070] In some embodiments, based on a verified gold knowledge unit set, a graph neural network is used to simulate the spatial probability diffusion and resource aggregation process of mineralization, generating a probability target area of ​​concealed ore bodies in three-dimensional geological space; the probability target area of ​​concealed ore bodies is transformed into a multi-dimensional exploration strategy scheme that includes drilling locations, survey line deployment, and sampling priorities; based on the multi-dimensional exploration strategy scheme, the newly acquired exploration data is used as feedback information to automatically reconstruct and enhance the corresponding causal paths and spatiotemporal patterns in the gold exploration knowledge graph, forming a self-evolving intelligent exploration closed loop.

[0071] Graph neural networks can be deep learning models specifically designed for processing graph-structured data. They are used to simulate the process of information transmission and influence diffusion between geological element nodes (such as faults, rock masses, and alteration zones) through knowledge graph relationships (such as control, heat source provision, and symbiosis). Three-dimensional geological space refers to a three-dimensional digital environment constructed based on geographic coordinates and elevation / depth, used to accurately characterize the spatial distribution of underground geological bodies. A concealed orebody probability target area refers to a three-dimensional region within the three-dimensional geological space, calculated through the above simulations, indicating a probability (e.g., 0.1 to 0.9) of the possible existence of a concealed orebody; a higher probability indicates a greater prospective potential. A multi-dimensional exploration strategy scheme refers to comprehensive engineering deployment suggestions generated based on probability target areas, which can directly guide field construction. Its dimensions include spatial location (drilling locations, survey line deployment), construction sequence (sampling priority), and expected objectives. A self-evolving intelligent exploration closed loop refers to an automated and continuously optimized complete loop system of "knowledge graph generation predicting target areas → guiding exploration engineering to obtain new data → new data feedback driving knowledge graph self-updating → generating more accurate new target areas."

[0072] Specifically, even when traditional geological prediction methods construct knowledge graphs, they often remain at the level of theoretical connections, producing conclusions that are mostly qualitative "favorable mineralization areas." These conclusions cannot be directly translated into actionable, quantitative engineering instructions, leading to exploration decisions still relying on expert experience and resulting in high risks and long cycles. For example, even if a prediction indicates that "a certain fault zone is a key area for mineral exploration," the specific location, depth, and drilling sequence still require manual judgment. Furthermore, if the first batch of boreholes fails to encounter minerals, the entire prediction model cannot be adjusted due to a lack of real-time feedback, resulting in significant waste. To address the above issues, this step first drives the operation of the existing gold exploration knowledge graph with multi-dimensional reasoning capabilities: the graph neural network treats nodes such as "rock mass", "fracture", and "alteration" in the graph as information sources, and performs multiple rounds of message passing and probability diffusion simulation along the relationship edges such as "hydrothermal migration" and "tectonic control of ore". (For example, it simulates hydrothermal fluids starting from the rock mass node, spreading outwards along the high-conductivity fault edge, and accumulating at the geophysical and geochemical anomaly node.) Finally, it calculates the mineralization probability of each unit in the three-dimensional geological space grid, generating a three-dimensional result such as "a hidden ore body probability target area with a probability greater than 0.7 exists at -500 meters to -800 meters on the footwall of the F2 fault". Next, the system transforms this probabilistic model into an actionable multi-dimensional exploration strategy: automatically optimizing and generating three optimal drilling locations (e.g., ZK01: X=123, Y=456, design depth 800 meters) to verify the high-probability target center; planning two wide-area electromagnetic survey lines traversing the target area to explore the extension of the ore-controlling structure; and setting sampling priorities based on probability values, burial depth, and surface conditions (e.g., ZK01 is prioritized for P1 level drilling). Finally, new exploration data obtained after these projects are implemented (e.g., ZK01 shows 5-meter-thick gold mineralization at -650 meters) are used as feedback input to the system, triggering the knowledge graph to automatically reconstruct and enhance the confidence of causal paths such as "F2 fault controlling ore," thereby producing more accurate target areas in the next round of prediction, forming a self-evolving intelligent exploration closed loop that continuously learns from practice and becomes increasingly intelligent with use.

[0073] The method provided in this embodiment bridges the "last mile" from geological theoretical knowledge to field exploration operations, directly transforming intelligent predictions into executable engineering solutions. By relying on closed-loop feedback to achieve continuous self-optimization of the model, the uncertainty and risk of exploration decisions are greatly reduced, and the efficiency and success rate of deep mineral exploration are significantly improved.

Claims

1. A method for heterogeneous data analysis and knowledge graph construction in the gold industry, characterized by: include: Obtain a multimodal heterogeneous original dataset of gold mines, perform semantic normalization parsing based on domain ontology and cross-modal consistency verification on the original dataset, and generate a multimodal gold geological knowledge information set. Based on the multimodal gold geological knowledge information set, a deep correlation mining process is performed that integrates causal inference and spatiotemporal evolution patterns to generate a candidate gold knowledge unit set with causal logic and spatiotemporal constraints. Based on the candidate gold knowledge unit set, and using known typical gold deposits as verification anchors, the credibility of the knowledge units is verified by backtracking and the causal relationship strength is screened to generate a verified gold knowledge unit set. Based on the verified gold knowledge unit set, a knowledge graph is modeled and iteratively constructed to generate a gold exploration knowledge graph that serves intelligent prediction of deep and concealed gold deposits.

2. The method according to claim 1, characterized in that, The generation process of the multimodal gold geological knowledge information set includes: The gold mine multimodal heterogeneous raw dataset includes unstructured text, exploration geological images, and geophysical and chemical data; Based on the aforementioned multimodal heterogeneous original dataset, entity recognition, relation extraction, and semantic annotation based on the ontology of the gold geology field are performed. Natural language processing technology is used to realize the semantic normalization expression of the same geological concepts in data from different sources, thereby constructing semantically normalized information. Based on the aforementioned multimodal heterogeneous original dataset, cross-modal correlation analysis is performed, and multimodal fusion technology is applied to identify the phenomenon that different modal data have the same description and logic when describing the same geological object or event, thereby constructing cross-modal consistency information; Based on the semantic normalization information and the cross-modal consistency information, the multimodal gold geological knowledge information set is generated.

3. The method according to claim 2, characterized in that, The process of constructing the semantically normalized information includes: Based on the unstructured text, a large gold model pre-trained on a gold domain corpus is used to perform deep analysis and entity relation extraction to generate a semantic relation representation of the geological text. Based on the geological exploration images, computer vision technology is used to identify geological structures and lithological units, and they are parsed into vectorized spatial objects with geological semantic annotations. Based on the aforementioned geophysical and chemical data, the anomalous features in the geophysical and chemical data are mapped into response pattern descriptions that characterize geological, geographical, and geochemical information using professional knowledge in the geological field. The geological text semantic relationship description, the vectorized spatial object, and the response pattern description are uniformly mapped to the standard concept nodes of the gold geology ontology to eliminate conceptual ambiguity within various types of data, thereby constructing the semantically normalized information.

4. The method according to claim 2, characterized in that, The process of constructing the cross-modal consistency information includes: Using spatial coordinates or geological strata as a unified reference, a cross-validation bridge is established to connect the unstructured text, the exploration geological images, and the geophysical and chemical data. The cross-validation bridge is used to compare and analyze descriptions of the same geological object or process from different modal data at the logical and attribute levels to identify potential consistency phenomena in description. Based on the identified consistency of the statements, descriptions from different modal sources are mutually corroborated and enhanced to generate corroborating information that is supported by multi-source evidence and is logically self-consistent. All the corroborating information generated through mutual verification is integrated to construct the cross-modal consistency information.

5. The method according to claim 2, characterized in that, The process of generating the candidate gold knowledge unit set includes: Based on the aforementioned multimodal gold geological knowledge information set, the temporal sequence and spatial association of geological events are analyzed to mine and construct event causal chains that characterize the mineralization process. Simultaneously, the spatial distribution patterns and temporal evolution characteristics of geological elements, including lithological units, tectonic traces, alteration zones, and mineralization bodies, are analyzed in the region to extract spatiotemporal evolution pattern information. The logical relationship of the event causal chain and the constraint attributes of the spatiotemporal evolution pattern information are jointly assigned to relevant knowledge units to form the candidate golden knowledge unit set.

6. The method according to claim 5, characterized in that, The process of constructing the event causal chain includes: From the multimodal gold geological knowledge information set, a set of geological event instances with clear spatiotemporal attributes is extracted. The set of geological event instances includes rock mass intrusion events, tectonic activity events, hydrothermal alteration events, and mineral precipitation events. Based on the causal discovery algorithm, the causal dependency relationship mining based on temporal constraints and spatial correlation is performed on the set of geological event instances to generate a preliminary causal network topology. Based on a pre-established gold mineralization theory and a prior knowledge base of geodynamics, logical constraints and path verification are applied to the preliminary causal network topology to select candidate causal paths that conform to the laws of geological evolution. For each causal edge in the candidate causal path, its confidence score is analyzed. The confidence score integrates temporal tightness, spatial proximity, and the corroboration strength in the cross-modal consistency information, thereby constructing the event causal chain with multi-dimensional confidence assessment.

7. The method according to claim 5, characterized in that, The process of extracting the spatiotemporal evolution pattern information includes: From the multimodal gold geological knowledge information set, extract the spatial geometric distribution data and geological time series data of the geological elements; Based on the aforementioned spatial geometric distribution data, spatial topology analysis and density field simulation are used to quantitatively analyze the spatial aggregation, orientation, and symbiotic combination patterns of the geological elements in the region, and generate spatial distribution patterns. Based on the geological time series data, geological age sequences and superposition and cutting relationships of different geological elements are constructed to infer their formation order and periodicity, and a time series evolution sequence is generated. The spatial distribution pattern is coupled with the temporal evolution sequence to generate the spatiotemporal evolution pattern information that describes how the structure of the mineralization system dynamically evolves over geological time.

8. The method according to claim 5, characterized in that, The process of generating the verification gold knowledge unit set includes: Several known typical gold deposits with clear genetic types and distinct geological characteristics were selected as verification benchmarks, and their complete geological knowledge maps were extracted as verification templates. The event causal chain and spatiotemporal evolution pattern information are matched one by one with the verified patterns in the verification template for pattern matching and instance mapping. The matching degree and coverage of each candidate knowledge unit in multiple verification templates are analyzed based on machine learning classifiers. Units that match successfully are marked as trustworthy, and their causal relationship strength is quantified based on the breadth and consistency of their matching. Finally, the credibility knowledge units that pass the backtracking verification and have a correlation strength higher than the preset threshold are selected to form the verification gold knowledge unit set.

9. The method according to claim 8, characterized in that, The generation process of the gold exploration knowledge graph includes: Using the set of verified gold knowledge units as nodes and relationships, and the event causal chain and the spatiotemporal evolution pattern information, an initial graph network is constructed; Using a dynamic graph engine, knowledge fusion, reasoning completion, and logical conflict resolution are performed on the initial graph network to generate the gold exploration knowledge graph that satisfies consistency constraints and has multi-dimensional reasoning capabilities. New exploration evidence is input into the gold exploration knowledge graph, triggering its adaptive iterative updates and verifications to serve intelligent prediction of deep and concealed gold deposits.

10. The method according to claim 9, characterized in that, The multi-dimensional reasoning ability includes: Based on the verified gold knowledge unit set, a graph neural network is used to simulate the spatial probability diffusion and resource accumulation process of mineralization, and generate a probability target area of ​​concealed ore bodies in three-dimensional geological space. The probability target area of ​​the concealed ore body is transformed into a multi-dimensional exploration strategy scheme that includes drilling locations, survey line deployment and sampling priorities. Based on the aforementioned multidimensional exploration strategy, the newly acquired exploration data is used as feedback information to automatically reconstruct and enhance the corresponding causal paths and spatiotemporal patterns in the gold exploration knowledge graph, forming a self-evolving intelligent exploration closed loop.