Mineral-forming anomaly extraction method based on geological'large model '

By constructing a geological "big model", using ERNIE-UIE and convolutional neural networks, the problem of low geological data in ore prospecting prediction in unfamiliar workspaces is solved, and efficient mineralization anomaly extraction and ore prospecting prediction under very few samples are achieved.

CN120336807AActive Publication Date: 2025-07-18CHINA UNIV OF GEOSCIENCES (BEIJING)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510400620.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

In the case of unfamiliar work areas or few geological data and few discovered ore sites, existing deep learning methods are difficult to effectively predict mineralization, resulting in a low level of mineral exploration.

Method used

A method of mineralization anomaly extraction based on geological "big model" is constructed. By constructing text and spatial databases, using the ERNIE-UIE model to extract information, establish a knowledge graph, and combining convolutional neural networks to perform feature space modeling to realize dual-driven ore prospecting prediction of knowledge and data.

Benefits of technology

Under the condition of very few samples, the exploration prospecting prospecting area is realized with dual-driven knowledge and data, which improves the accuracy and efficiency of mineral exploration prediction in unfamiliar work areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336807A_ABST
    Figure CN120336807A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of solid mineral exploration, and particularly relates to a metallogenic anomaly extraction method based on a geological'large model '. Compared with a traditional mineralization prediction method based on deep learning, the method can solve the actual problems under the conditions of low working degree, few geological data, few found ore occurrences and the like in prospecting prediction of an unfamiliar working area. According to the method, a semantic space and a feature space are communicated by establishing space mapping, and a conversion bridge between a'large model 'and'big data' is constructed, so that knowledge and data dual-driven prospecting prospective area delineation is realized under the condition of extremely few samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of solid mineral exploration, and particularly relates to a method for extracting metallogenic anomalies based on a geological "big model". Background Art

[0002] Geological data records the specific and final results of geological work, while geological big data is more comprehensive and in-depth geological information formed on the basis of these data through technical means such as data integration, mining, and analysis. Mineralization prediction, as the core link of mineral exploration, is an important means and method to achieve scientific ore prospecting and increase resource reserves. This research involves multidisciplinary knowledge such as geology, geophysics, geochemistry, and remote sensing, and uses deposit models and metallogenic system prediction models to evaluate mineral resources for the purpose of ore prospecting prediction. Big data and artificial intelligence are significantly changing the human's cognitive and research paradigms, especially in the data-intensive field of mineral exploration, which not only brings huge challenges of transformation but also provides unprecedented opportunities for its innovation and development. Academician Zhao Pengda, the founder of mathematical geology in China, once pointed out that geological big data analysis is a new interdisciplinary subject. It is based on geological science and information technology, describes geological bodies, geological processes, and geological work methods by establishing and applying various mathematical models, and uses data science methods to intelligently process geological big data, analyze and mine valuable core information and key data from it, form concentrated digital knowledge, and reveal information that may be ignored by traditional logics or methods, thereby significantly improving exploration efficiency, reducing uncertainty, lowering exploration costs and risks, and promoting the development of mineral exploration towards a more quantitative, refined, and visual direction. For example, some scholars have applied machine learning methods such as artificial neural networks, support vector machines, and random forests to the geoscience field and achieved good application results. Deep learning methods process and analyze information by constructing neural networks to imitate the human brain's analysis and learning process. Typical models include convolutional neural networks, recurrent neural networks, stacked autoencoders, etc. Applying machine learning methods, especially deep learning methods, to mineralization prediction is an inevitable trend in the development of quantitative geoscience research. With the assistance of artificial intelligence technologies such as machine learning and deep learning, valuable information in geological big data can be further mined, thus promoting the development of geological science and contributing to social and economic development.

[0003] In the field of mineral resource exploration, the application of geological big data is particularly prominent. By integrating data from multiple aspects such as geology, geophysics, geochemistry, and remote sensing, it significantly improves the prediction accuracy of the distribution and reserves of mineral resources, thus providing a solid scientific foundation for ensuring the rational development and utilization of mineral resources. Incorporating intelligent analysis techniques of geological big data into mineral resource exploration and assessment work is an indispensable step in the development of mathematical geology and also the initial stage towards intelligent ore prospecting. Mineralization prediction is based on the extraction and interpretation of multi-disciplinary and multi-source geological information. Although, based on the idea of big data, all data can be handed over to machine learning for processing at one time, simply relying on data to study the characteristics of the data itself or the relationship between the data and known ore deposits. The essence of mineralization prediction based on deep learning is the mathematical statistical analysis of multi-source geological information, and its advantage lies in the ability to extract abstract feature representations at different levels. Since different types of ore deposits may have the same geological origin, the same type of ore deposit in the same study area can also show different ore deposit types due to differences in geological processes. Through deep learning methods such as convolutional neural networks, optimal representation learning can be carried out, and supervised learning can also be used to obtain supervised learning representations related to the ore deposits in the study area. In theory, the same type of ore deposit in the same study area should learn the same representation, but due to the multi-solution nature of geological information and the problem of data dimension (the number of prediction elements), the consistency of learning features is the core issue in deep learning mineralization prediction. In deep learning tasks, we face the problem of insufficient numbers of positive and negative samples, so it is necessary to study deep learning mineralization prediction under few-sample or even zero-sample conditions. Summary of the Invention

[0004] Based on the above problems, the purpose of this application is to provide an innovative deep learning mineral resource prediction method that can solve the practical problems under conditions such as low work intensity, few geological data, or few discovered ore points in prospecting prediction in unfamiliar work areas.

[0005] To achieve the above purpose, the technical solution of this application provides a method for extracting ore-forming anomalies based on a geological "big model", including the following steps:

[0006] S1. Construct a text and spatial database: Determine the retrieval theme of the target mining area for ore-forming anomaly extraction, conduct text big data discovery and spatial big data discovery through data collection tools to obtain geological corpora and geological maps respectively, and obtain a text database and a vector layer library after preprocessing; among them, the text database is used as the corpus, and the vector layer library is used for spatial connection;

[0007] S2. Construct a knowledge graph: Determine the ontology of the knowledge graph to be constructed and the entity types in the ontology, use an information extraction pre-trained model to extract information from the text database, obtain "entity-relationship-entity" triples, construct the nodes and edges of the knowledge graph, and form a knowledge graph; where the nodes are entities and the edges are relationships;

[0008] S3. Predictive model training: Use knowledge graph embedding technology to model the feature space of the obtained knowledge graph, use a convolutional neural network to train the model, establish a mapping from data to the knowledge space, and obtain a trained predictive model;

[0009] S4. Use the trained predictive model for ore-forming prediction and ore-forming anomaly extraction.

[0010] Further, the specific construction method of the information extraction pre-trained model in S2 is as follows:

[0011] (1) Geological corpus fine-tuning: Use Label Studio to randomly extract some texts from the text database for data annotation to obtain a micro-called corpus;

[0012] (2) Allocate the micro-called corpus into a training set, a validation set, and a test set, and obtain fine-tuned geological corpus after training;

[0013] (3) Use the fine-tuned geological corpus to fine-tune the pre-trained model to obtain a fine-tuned information extraction pre-trained model.

[0014] Further, the pre-trained model is the ERNIE-UIE open-source model.

[0015] Further, the specific steps of using knowledge graph embedding technology to model the feature space of the obtained knowledge graph in S3 are as follows: Extract all entities and relationships in the knowledge graph, use a knowledge graph embedding tool to train the vector representations of entities and relationships, and map the embedding vectors of entities and relationships into the semantic space to achieve feature space modeling.

[0016] Further, the knowledge graph embedding tool is the KG2E model.

[0017] Further, the specific method of using a convolutional neural network to train the model in S3 is as follows:

[0018] (1) Pair the data features of entities in the model after achieving feature space modeling with the corresponding semantic labels, construct a set of sample pairs composed of paired semantics and data as the training set, and use a convolutional neural network encoder CNN for training;

[0019] (2) During each training, select one sample pair with different semantics from the training set. Obtain the feature vectors of the data part through a convolutional neural network, and calculate the cosine similarity between the semantics and the data features to form a similarity matrix;

[0020] (3) Update the CNN encoder through the gradient descent method. Through optimization, make the values on the diagonal of the similarity matrix the largest in both the horizontal and vertical directions, that is, the similarity between samples belonging to the same sample pair is the largest, and the similarity between samples belonging to different sample pairs is the smallest.

[0021] (4) Repeat step (3) until the model converges, that is, the value of the loss function no longer decreases significantly, and obtain the trained prediction model; among them, during the training process, save the model parameters with the best performance.

[0022] Beneficial effects: Compared with the traditional deep learning-based mineralization prediction method, it can solve the practical problems under the conditions of low working degree, few geological data, or few discovered ore points in unfamiliar working areas. This method builds a spatial mapping to connect the semantic space and the feature space, constructs a conversion bridge between the "big model" and the "big data", and thus realizes the delineation of prospecting areas driven by both knowledge and data with extremely few samples. Brief Description of the Drawings

[0023] Figure 1 is the flowchart of the knowledge graph construction of this application;

[0024] Figure 2 is the flowchart of the knowledge graph embedding based on the KG2E model;

[0025] Figure 3 is the technical flowchart of the model training process;

[0026] Figure 4 is the prediction algorithm for calling the CNNEncoder;

[0027] Figure 5 is the visualization result of the similarity;

[0028] Figure 6 is the superposition of the prediction result and the geochemical interpolation map, the ore-bearing stratum distribution map, the structure and buffer zone map;

[0029] Figure 7 is the superposition map of the similarity prediction result and the geochemical comprehensive anomaly;

[0030] Figure 8 is the superposition map of the similarity prediction result and the geochemical single element and the first principal component anomaly. Detailed Implementation Manner

[0031] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0032] Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present application. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the authorization specification. In all the examples shown and discussed here, any specific value should be construed as merely exemplary, rather than as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.

[0033] Unless otherwise specified, the meanings of the scientific and technical terms in this specification are the same as those generally understood by those skilled in the art. However, in case of conflict, the definitions in this specification shall prevail.

[0034] In this application, the term "low degree of work" refers to a relatively low degree of systematic geological research on the target area, and the basic or key exploration work has not been completed. The research area is in the preliminary exploration / general survey stage (not the detailed exploration or exploration stage), and only the 1:200,000 or 1:50,000 regional geological survey has been completed, and geophysical exploration (such as gravity, magnetic method), geochemical prospecting (such as stream sediment measurement), or drilling verification has not been carried out.

[0035] The term "less geological data" means that the multi-source geological data (structure, rock, geochemistry, etc.) coverage is incomplete or the accuracy is insufficient, making it difficult to support refined modeling.

[0036] The term "few discovered ore spots" means that the number of known ore deposits (spots) in the research area is low, or no industrial-grade ore deposits have been discovered in the exploration history.

[0037] As Figure 1 shown, the embodiments of the present application disclose a method for extracting ore-forming anomalies based on a geological "big model", including the following steps:

[0038] S1. Construct a text and spatial database: Determine the retrieval theme of the target mining area for ore-forming anomaly extraction, perform text big data discovery and spatial big data discovery through data collection tools, obtain geological corpus and geological maps respectively, and obtain a text database and a vector layer library after preprocessing; among them, the text database serves as the corpus, and the vector layer library is used for spatial connection;

[0039] S2. Construct a knowledge graph: Determine the ontology of the knowledge graph to be constructed and the entity types in the ontology, use the pre-trained model for information extraction to extract information from the text database, obtain "entity-relationship-entity" triples, and construct the nodes and edges of the knowledge graph to form the knowledge graph; where, the nodes are entities and the edges are relationships.

[0040] S3. Predictive model training: Use knowledge graph embedding technology to model the feature space of the obtained knowledge graph, use a convolutional neural network to train the model, establish a mapping from data to the knowledge space, and obtain the trained predictive model.

[0041] S4. Use the trained predictive model for ore-forming prediction and ore-forming anomaly extraction.

[0042] In this embodiment, the data collection tool in S1 is a commonly used data retrieval tool in the prior art. The preferred method for text big data discovery and spatial big data discovery through the data collection tool is as follows:

[0043] (1) After determining the retrieval topic, perform word segmentation using a geological dictionary to determine the retrieval topic words.

[0044] (2) Read in the geological thesaurus, use the knowledge tree based on the geological thesaurus to calculate and analyze the input keywords, and obtain more related keywords.

[0045] (3) Use the obtained related keywords as input for further retrieval until the retrieval is complete or other termination conditions are met. In the retrieval link iteration process, use the keywords given in the previous step as the topic input, generate the initial URL seeds for each keyword using the mainstream search engine API, and perform link analysis and data crawling.

[0046] (4) Use the network resources and algorithms of the search engine to obtain a large number of URLs.

[0047] (5) Through the theme relevance analysis of the URL page data (including title, abstract, source, and related URLs) returned by the search engine, store the URLs related to the geological theme in the crawled URL queue library, and crawl the research data according to the relevance size.

[0048] (6) Remove invalid, incorrect, and duplicate data; merge data from different data sources to solve contradictions and conflicts between data.

[0049] In this embodiment, in the construction of the knowledge graph in S2, the information extraction pre-trained model is a conventional pre-trained model in the prior art, such as open-source models like DeepDive and ERNIE-UIE. The geological text is mainly composed of common modern Chinese vocabulary and geological and mineral professional terms. The characteristics of geological text are mainly reflected in its writing standardization and word accuracy. Terms in the geological field and other disciplines such as statistics and computer science are widely used, and even show the characteristics of mixed Chinese and English compilation. Since geological research usually targets specific regions, the text contains a large number of place names. In addition, there is a nesting phenomenon among professional terms, which may be due to the need for refined professional descriptions. These characteristics reflect the characteristics of geological literature as an exact science, and at the same time reflect the complexity of professional descriptions. These characteristics of geological text make it suitable for a general information extraction framework. There are many common open-source tools for information extraction, such as DeepDive. In this embodiment, the open-source large model ERNIE-UIE is preferably used. The reasons are as follows: First, it is an open-source pre-trained model and is easy to obtain; second, due to the use of prompt learning, it has excellent few-shot fine-tuning capabilities, greatly reducing the workload of corpus annotation. Due to the professionalism presented by geological corpus and the fact that the pre-trained language model uses a small amount of professional corpus, in order to ensure the accuracy of knowledge graph extraction, it is necessary to fine-tune the pre-trained model obtained from open source. In this embodiment, Label Studio is preferably used to perform data annotation on randomly selected geological text, convert the data exported by Label Studio into the form required when inputting the model through a script, and then select a suitable structural pattern for annotation according to requirements to achieve the fine-tuning of geological corpus.

[0050] In a further embodiment, a preferred method for fine-tuning the ERNIE-UIE pre-trained model using Label Studio is disclosed:

[0051] (1) For the constructed text database, randomly select a certain number of geological texts for annotation to form a geological corpus;

[0052] (1.1) Random sampling: Randomly select a certain number of geological texts from the text database, and preprocess the selected texts, including removing format information, unifying fonts, and splitting sentences.

[0053] (1.2) Annotation specification: Define the entity types in the geological field, define the relationships between entities, and formulate annotation specifications, including the use of annotation tools, annotation principles, and ambiguity handling.

[0054] (1.3) Corpus annotation: Use the annotation tool to perform entity and relationship annotation.

[0055] (2) For the task of knowledge extraction, construct a language model and use the model to extract prospecting-related knowledge from text data; store the extracted knowledge in the form of a graph database.

[0056] (2.1) Data preparation: Divide the corpus into a training set, a validation set, and a test set.

[0057] (2.2) Model selection and training: Select pre-trained language models suitable for geological texts, such as BERT, GPT, etc. Based on the pre-trained models, use the geological corpus for fine-tuning training to adapt to the characteristics of the geological field. Optimize the model performance by adjusting parameters such as the learning rate and batch size.

[0058] (2.3) Knowledge extraction: Use the trained model to identify geological entities from the text; identify the relationships between entities and classify them into predefined relationship types.

[0059] (2.4) Knowledge storage: Convert the extracted knowledge into a format supported by the graph database; use the import tools or APIs provided by the graph database to import the knowledge data into the graph database.

[0060] In this embodiment, the specific steps of using the knowledge graph embedding technology to perform feature space modeling on the obtained knowledge graph in S3 are as follows: Extract all entities and relationships in the knowledge graph, use the knowledge graph embedding tool to train the vector representations of the entities and relationships, and map the embedding vectors of the entities and relationships into the semantic space to achieve feature space modeling.

[0061] In a further embodiment, the specific method of using the convolutional neural network to train the model in S3 is as follows:

[0062] (1) Pair the data features of the entities in the model after achieving feature space modeling with the corresponding semantic labels, construct a set of sample pairs composed of paired semantics and data as the training set, and use the convolutional neural network encoder CNN for training;

[0063] (2) Each time during training, select one sample pair with different semantics from the training set, obtain the feature vector of the data part through the convolutional neural network, and calculate the cosine similarity between the semantics and the data features to form a similarity matrix;

[0064] (3) Update the CNN encoder through the gradient descent method, and optimize to make the values on the diagonal of the similarity matrix the largest in both the horizontal and vertical directions, that is, the similarity between samples belonging to the same sample pair is the largest, and the similarity between samples belonging to different sample pairs is the smallest.

[0065] (4) Repeat step (3) until the model converges, that is, the value of the loss function no longer decreases significantly, and obtain the trained prediction model; wherein, during the training process, the model parameters with the best performance are saved.

[0066] The Xiarihamu area will be selected as the study area below, and the technical means and technical effects of the mineralization anomaly extraction method of the geological "big model" of this application will be explained in detail.

[0067] Although the degree of geological work in the Xiarihamu study area is relatively high, with partial coverage of 1:200,000 stream sediment geochemical data, only one known mineral site has been obtained, thus meeting the criterion of “few discovered mineral sites” in this application.

[0068] S1. Constructing text and spatial databases: Aiming at the specific problems of the mining area, we used "Xiarihamu nickel-cobalt mine" as the search theme, collected a variety of data on Xiarihamu nickel-cobalt mine, including geological reports, academic papers, and geological maps, through LAN distributed search and WAN topic crawler and double-iteration website discovery, and made the national mineral deposit database available at the National Geological Information Center. Through LAN data discovery, we discovered and obtained 14 geological reports, documents, monographs, etc. in the local LAN, and a 1:250,000 structural structure map drawn by the Qinghai Provincial Institute of Geological Survey. Figure 2 The research team collected 100 pieces of geochemical data (Bruntai piece and Dazaohuo piece), 1 statistical table of geochemical data of stream sediments in the study area (containing 30 elements and 1779 sampling points in total), 30 single-element geochemical maps drawn based on the national regional geochemical database; and crawled 197 related papers in batches from China National Knowledge Infrastructure through wide area network big data discovery; and obtained open access to 1 national mineral deposit database at the National Geological Information Center.

[0069] S2. Constructing knowledge graph: Knowledge graph is a technical method to describe knowledge with graph model. Its early concept comes from semantic network. This paper mainly uses the method of constructing domain knowledge graph to condense the knowledge contained in the geological literature of the study area. The technical process is as follows: Figure 1As shown below: After determining the research objectives, through the discovery and preprocessing of text and spatial big data, a text database and a vector layer library are obtained. The text database among them serves as the corpus for constructing the knowledge graph, while the vector layer library will be spatially connected with the knowledge graph as attribute information, laying a foundation for the next prediction work. By sorting out the core requirements of the research, the ontology of the knowledge graph is determined; then, the types of entities in the ontology are determined, that is, the collection of entities with the same characteristics or attributes. After that, we enter the stage of entity extraction (named entity recognition) and relationship extraction, that is, constructing the nodes and edges in the knowledge graph respectively. This stage can also be collectively referred to as information extraction. Relying on knowledge extraction technology, we will obtain the smallest unit hidden in the text - triples. Then, we will obtain a complete knowledge graph through knowledge fusion and knowledge representation. Geological texts are mainly composed of commonly used modern Chinese vocabulary and geological and mineral professional terms. The characteristics of geological texts are mainly reflected in their writing standardization and word - using accuracy. Terms in the geological field and other disciplines such as statistics and computer science are widely used, and even show the characteristics of mixed Chinese - English compilation. Since geological research usually targets specific regions, a large number of place names are included in the texts. In addition, there is a nesting phenomenon among professional terms, which may be due to the need for refined professional descriptions. These characteristics reflect the characteristics of geological literature as an exact science, and at the same time reflect the complexity of professional descriptions. These characteristics of geological texts make them suitable for general information extraction frameworks. There are many common open - source tools for information extraction, such as DeepDive, etc. In the research example, the reason we choose the open - source large - model ERNIE - UIE is as follows: First, it is an open - source pre - trained model and is easy to obtain; second, due to the use of prompt learning, it has excellent few - shot fine - tuning ability, greatly reducing the workload of corpus annotation.

[0070] Due to the professionalism presented by geological corpora and the fact that the pre - trained language model uses a small amount of professional corpora, for the accuracy of knowledge graph extraction, we need to fine - tune the pre - trained model obtained through open source. We use Label Studio to annotate randomly selected geological texts, and convert the data exported from Label Studio into the form required when inputting into the model through scripts. We need to select a suitable structural pattern for annotation according to the requirements. For example, in the research example of Xiamarihamu, the structural pattern we selected is "[spot] entity category [asso] relationship category [text]".

[0071] 1. Spotting: Locate the target information segments. The information that needs to be located in the research example is mainly some entities, such as deposit names, geographical locations, metallogenic geological backgrounds, metallogenic epochs, genetic types, mineralization types, geology, geophysics, geochemistry, remote sensing, etc.

[0072] 2. Associating: Identifying the relationships between target information segments. In the research example, to represent the prospecting-related knowledge of the Xiamerhamu deposit, we only need to know its geographical location, metallogenic geological background, metallogenic era, genetic type, and mineralization type, as well as which geological, geophysical, geochemical, and remote sensing information is favorable for its exploration. Therefore, in the research example, we chose to identify three associations, namely "located in", "belongs to", and "favorable for".

[0073] In this extraction task, on the label-studio platform, the entity type tags we constructed include: deposit name, genetic type, mineralization type, metallogenic era, metallogenic geological background, geology, geophysics, geochemistry, remote sensing, and geographical location; the relationship type tags we constructed include: belongs to, located in, and favorable for. The annotation operation can be mainly described as: annotating the subject and object in the triple with entity type tags; adding a relationship connection line with the arrow direction pointing from the subject to the object; and selecting a relationship type tag for the relationship. According to the defined entity and relationship type tags, we annotated the text data, with a total of 1883 sentences annotated, generating 16424 pieces of corpus. The sentences were randomly assigned to the training set (1507 sentences, 12913 pieces), validation set (188 sentences, 1877 pieces), and test set (188 sentences, 1634 pieces) according to the ratio of 8:1:1.

[0074] After obtaining the correctly formatted geological corpus, we fine-tuned the ERNIE-UIE pre-trained model on the downstream task of geological information extraction. Then, using the fine-tuned model, we completed the information extraction in the collected text database. The extracted information was output in the form of structured extraction language and could be organized into a large number of triples. In this study, we used UIE-base (with 12 layers, 768 hidden units, and 12 attention heads) as the initial pre-trained model and fine-tuned it. During the evaluation process, we used a single-stage evaluation method and evaluated each positive example category separately. The validation / test set would use all the tags at the same level to generate corresponding negative examples. Table 1 shows the performance of the model on geological texts before and after fine-tuning with geological corpus, and it can be found that the performance of the fine-tuned model in geological text information extraction has been significantly improved.

[0075] Table 1 Performance of the model on geological texts before and after fine-tuning with geological corpus

[0076]

[0077] As can be seen from Table 1, before fine-tuning, UIE-base had a certain degree of recognition ability for ore deposit names, metallogenic epochs, geographical locations, etc., but the overall recognition ability was poor. After fine-tuning with geological corpus, its overall extraction ability has been significantly improved. Finally, we stored the sorted triples in the graph database. The research example uses Neo4j as the knowledge storage tool. It has high scalability, efficient representation and storage methods, and has low requirements for hardware. It is one of the current popular graph databases.

[0078] S3. Prediction model training;

[0079] First, use the neural system to solve the reasoning problem of the symbolic system, that is, use knowledge graph embedding technology to realize the semantic space modeling of ore prospecting knowledge in the study area; then use the symbolic system to improve the neural system, that is, use the principle of jointly expressing the same entity in the semantic space and the feature space, and realize the feature space modeling of the comprehensive information of ore prospecting in the study area by establishing a mapping.

[0080] In the knowledge graph representation stage, the KG2E model is used in the example. According to the principle of the KG2E model, the pseudo-code for realizing the knowledge graph embedding based on the KG2E model is shown in Table 2, and the technical flow chart is as Figure 2 shown.

[0081] Table 2 Learning algorithm of KG2E model

[0082]

[0083] In the prior art, generally, according to the knowledge extracted from the knowledge graph, a prospecting conceptual model is constructed for the study area to further guide the prospecting prediction work in the study area. However, in this application, we hope to obtain the semantic embedding of the knowledge in the national mineral deposit information database as the training set. Among them, the semantic embedding of the national mineral deposit information database is the semantic training set of the model used for the two-dimensional prospective area delineation driven by both knowledge and data, while the semantic embedding of the prospecting knowledge graph in the study area is the data used for prediction. The knowledge in the national mineral deposit information database is structured, so it is easy to be converted into triples. By connecting the two knowledge graphs, a total of 8,767 entities, 3 relationships, a training set size of 36,651, and test set and validation set sizes of 4,072 are obtained. Then, the learning algorithm of the KG2E model we developed is used for embedding to obtain the Gaussian distributions of all entities and relationships as the final semantic embedding results. We use the MR value (Mean Rank) as the parameter to evaluate the model performance, which can be understood as calculating the average rank of the true values of all triples in the prediction value list. Therefore, the smaller this value is, the better. The evaluation process can be understood as follows: by constructing incorrect triple entities and calculating the similarity between the head entity and the tail entity, and by comparing the similarity values and ranks of the correct and incorrect triples, the quality of the knowledge graph representation vectors can be evaluated. Ideally, the score of the correct triple should be less than that of the incorrect triple and the rank should be more forward. Our evaluation results show that after 90 rounds of training, the average rank of the model is reduced from 860 to about 650.

[0084] On this basis, it is necessary to fuse knowledge and data and establish a mapping from data to knowledge through training. Through the above experiments, we have successfully obtained the Gaussian embeddings of all entities and relationships in the triples and exported the semantic vector set composed of all entities. Considering the subsequent model training, we will make the sample pairs required for training 4 mappings, which will be used as positive sample pairs in the subsequent training process.

[0085] Taking the genetic type as an example, briefly describe the production process of these sample pairs: obtain the number of unique values of the genetic type, that is, count how many different genetic types there are in the semantic vectors; count the number of entity types of "ore deposit name" belonging to each genetic type; remove the genetic types with too few entities, and retain the genetic types with sufficient entities for training; query all the ore deposit name entities connected by the retained genetic type entities; query the comprehensive information attributes obtained by spatial connection of these ore deposit name entities; use the semantic vectors of the genetic type entities connected by each ore deposit name entity and the affiliated comprehensive information attributes to form the set of semantic and data sample pairs for training the genetic type CNN encoder.

[0086] The Convolutional Neural Network Encoder (CNNEncoder) is a key component in deep learning architectures, and its main responsibility is to extract and encode the features of the input data in the convolutional neural network. In a convolutional neural network, the encoder usually refers to the first few layers of the network, and these layers reduce the spatial dimension of the data through convolution and pooling operations while retaining key feature information. To simplify the training process, we designed a CNN encoder structure consisting of 5 convolutional layers and activation functions, 1 max-pooling layer, and 1 fully connected layer, as shown in Table 3, and its technical flow chart is as shown in Figure 3 shown. The convolutional layers, activation functions, and max-pooling layer of this structure are responsible for the transformation from comprehensive information to the feature space, while the fully connected layer is responsible for the mapping from data to the knowledge space. When training the CNN encoder, we adopt the following steps: (1) Construct a set of sample pairs composed of paired semantics and data as the training set; (2) In each training, select one sample pair with different semantics from the training set, obtain the feature vector by passing the data part through the CNN encoder, and calculate the cosine similarity between the semantics and the data features to form a similarity matrix; (3) Update the CNN encoder through the gradient descent method to make the values on the diagonal of the similarity matrix the largest in both the horizontal and vertical directions; (4) Repeat step 3 until convergence, and save the best model. After the training is completed, we will use the best model to predict the comprehensive information of the study area.

[0087] Table 3 Learning Algorithm of CNN Encoder

[0088]

[0089] Based on in-depth research on the metallogenic age, metallogenic geological background, genetic type, and mineralization type, we used the prepared training set and validation set to write Python scripts to train four different convolutional neural network (CNN) encoders. These four encoders respectively focus on the above different geological feature classification tasks. Using the test set data for classification testing, the average values of three evaluation indicators, namely the test accuracy, recall rate, and F1 score of the metallogenic age, metallogenic geological background, genetic type, and mineralization type, are obtained as shown in Table 4. By analyzing the performance of the model in different geological feature classifications, we found that the classification performance of the metallogenic geological background is the best among these three indicators. Generally speaking, the classification result of the metallogenic geological background is the most reliable among these four categories, and its performance in training is also the best, which may be related to the fact that it uses the most training samples and has the fewest classifications. The classification performance of the metallogenic age ranks second only to the metallogenic geological background in terms of accuracy and recall rate, and the F1 score also shows relatively high performance. This indicates that the model performs well in the classification of the metallogenic age and can accurately predict and capture positive class samples. This is consistent with its performance during training. We found that it uses the second most training samples. When the number of classifications is the same as that of the metallogenic geological background, the obvious decline in its performance may be related to the sharp reduction of training samples. The classification performance of the genetic type is similar to that of the metallogenic age in terms of accuracy, but slightly lower than that of the metallogenic geological background. The F1 score also shows a similar trend, indicating that the model also performs well in the classification of the genetic type, but may miss some positive class samples. The genetic type encoder uses fewer samples during training and needs to distinguish more categories, but its performance in terms of accuracy is comparable to that of the metallogenic age, and it rarely predicts negative samples as positive samples. This may indicate that geochemical data can better describe the characteristics of the genetic type than the description of the metallogenic age. The classification performance of the mineralization type is the worst among these three indicators, indicating that the classification result of the mineralization type is the least reliable among these four categories. The number of categories gradually increases from 14 for the metallogenic geological background to 19 for the mineralization type, which may mean that as the number of categories increases, the classification difficulty of the model also increases, which may be one of the reasons for the poor classification performance of the mineralization type.

[0090] Based on the above analysis, we can see that there are certain differences in the performance of the model in different geological feature classifications. Among them, the classification performance of the metallogenic geological background is the best, followed by the metallogenic age and genetic type, and the performance of the mineralization type is the worst. At the same time, as the number of categories increases and the number of samples decreases, the classification performance of the model will also be affected.

[0091] Table 4 Comparison list of model evaluation indicators

[0092]

[0093] S4. Use the trained prediction model for metallogenic prediction and extraction of metallogenic anomalies

[0094] We will train and predict based on four attributes: metallogenic geological background, metallogenic era, genetic type, and mineralization type. Through training and prediction, we will obtain the metallogenic probabilities predicted by four encoders for each grid unit in the study area (the pseudo-code of the prediction algorithm for calling the CNN encoder is shown in Table 5, and the flow chart of the prediction algorithm for calling the CNNEncoder is as Figure 4 shown). The average value of the four metallogenic probabilities is used as the final prediction result. These models do not need to be retrained and can also be used to predict other types of minerals in the study area. When predicting other types of minerals, we only need to change the semantic vector input during prediction.

[0095] Table 5 Pseudo-code of the prediction algorithm for calling the CNN encoder

[0096]

[0097] Take the data of the study area in grid units through four trained CNN encoders, and take the average value of the four obtained similarities as the similarity of the grid unit to the "Xiarihamu" standard point. The visualization result is as Figure 5 shown. This visualization result can be regarded as the result of extracting geochemical anomalies of stream sediments in the study area (extremely strong similarity: top 5%, strong similarity: top 15%, medium similarity: top 25%, weak similarity: similarity greater than 0 but not in the top 25%, dissimilar area: similarity less than 0).

[0098] After analysis, we can observe that within the study area, the known Xiarihamu copper-nickel-cobalt deposit is located in the area with medium similarity, that is, within the abnormal range of the top 15%-25%. We observed four significant abnormal areas, three of which are distributed in the northern part of the study area, and the other is located in the southwestern corner of the study area. In the southern part of the study area, the strong abnormal areas show a relatively continuous distribution pattern, while in the central part of the study area, the strong anomalies are mainly scattered in a punctate distribution. The central part of the study area is mainly composed of large areas of medium similarity areas. In addition, the similarity in the eastern part of the study area is generally low.

[0099] To deeply understand the specific element anomalies revealed by these anomaly extraction results, we conducted an overlay analysis on the extracted anomaly results and the results of interpolating the geochemical data of each element in the study area by the inverse distance weighting method. According to existing academic research, the magmatic liquation-type copper-nickel sulfide deposit in Xiarihamu is closely related to the anomalies in stream sediment measurements of nickel (Ni), cobalt (Co), copper (Cu), and chromium (Cr) elements. As Figure 6: The superposition results of the prediction results with the geochemical interpolation map, the ore-bearing strata distribution map, the structure and buffer zone map, Figure 7 : As shown in the superposition map of the similarity prediction results and the comprehensive geochemical anomaly, through superposition analysis, we found that the predicted high-value areas match extremely well with the high-value areas of Ni, Co, and Cu elements, and also show good agreement with the high-value areas of elements such as manganese (Mn), titanium (Ti), vanadium (V), and iron oxide (Fe2O3). In addition, the median area of the prediction results (between the top 15%-25%) shows a relatively consistent corresponding relationship with the spatial distribution of the ore-bearing strata Jinshuikou rock group in the area, the structural characteristics, and its 1500-meter buffer zone. This discovery further enhances our understanding of the distribution law of ore points in the study area.

[0100] As Figure 8 Shown is the superposition map of the similarity prediction results with the single geochemical element and the first principal component anomaly: The similarity prediction results are respectively superposed with the single element anomalies of Cu, Ni, and Co, which are closely related to mineralization, and the Fac1 anomaly that may reflect the magmatic activity influence in the study area. It is found that the similarity prediction results identify the high-value anomalies of the three single elements in the western part of the study area, but do not identify the high-value area in the southeastern part of the study area as a favorable ore-forming anomaly. The similarity prediction results also have a certain degree of superposition with the Fac1 anomaly. The anomaly distribution in the middle of the study area mentioned above has a good superposition with the Fac1 anomaly in the middle. The above phenomena may indicate that the similarity prediction results are not simply a stacking of high-value areas of geochemical anomalies, but have learned the dimensionality reduction representation of high-dimensional geochemical anomalies of a certain type of ore deposit.

[0101] As mentioned above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for extracting metallogenic anomalies based on a geological "big model", characterized in that It includes the following steps: S1. Construct a text and spatial database: Determine the retrieval theme of the target mining area for ore-forming anomaly extraction. Through data collection tools, conduct text big data discovery and spatial big data discovery to obtain geological corpora and geological maps respectively. After preprocessing, obtain a text database and a vector layer library. Among them, the text database serves as the corpus, and the vector layer library is used for spatial connection; S2. Construct a knowledge graph: Determine the ontology of the knowledge graph to be constructed and the entity types in the ontology. Use an information extraction pre-trained model to extract information from the text database to obtain "entity-relationship-entity" triples, and construct the nodes and edges of the knowledge graph to form a knowledge graph. Among them, the nodes are entities, and the edges are relationships; S3. Predictive model training: Use knowledge graph embedding technology to perform feature space modeling on the obtained knowledge graph, and use a convolutional neural network to train the model to establish a mapping from data to the knowledge space to obtain a trained predictive model; S4. Use the trained predictive model for ore-forming prediction and ore-forming anomaly extraction.

2. The extraction method according to claim 1, characterized in that, The specific construction method of the information extraction pre-trained model described in S2 is as follows: (1) Geological corpus fine-tuning: Use Label Studio to randomly extract some texts from the text database for data annotation to obtain a micro-called corpus; (2) Allocate the micro-called corpus into a training set, a validation set, and a test set, and obtain fine-tuned geological corpus after training; (3) Use the fine-tuned geological corpus to fine-tune the pre-trained model to obtain a fine-tuned information extraction pre-trained model.

3. The extraction method according to claim 2, wherein The pre-trained model is the ERNIE-UIE open-source model.

4. The extraction method according to claim 1, wherein The specific steps of using knowledge graph embedding technology to perform feature space modeling on the obtained knowledge graph described in S3 are as follows: Extract all entities and relationships in the knowledge graph, use a knowledge graph embedding tool to train the vector representations of entities and relationships, and map the embedding vectors of entities and relationships into the semantic space to achieve feature space modeling.

5. The extraction method according to claim 4, wherein The knowledge graph embedding tool is the KG2E model.

6. The extraction method according to claim 1, wherein The specific method of using a convolutional neural network to train the model described in S3 is as follows: (1) Pair the data features of entities in the model after achieving feature space modeling with the corresponding semantic labels, construct a set of sample pairs composed of paired semantics and data as the training set, and use a convolutional neural network encoder CNN for training; (2) Each time during training, select one sample pair with different semantics from the training set, obtain a feature vector through the convolutional neural network for the data part, and calculate the cosine similarity between the semantic and data features to form a similarity matrix; (3) Update the CNN encoder through the gradient descent method, and optimize to make the values on the diagonal of the similarity matrix the largest in both the horizontal and vertical directions, that is, the similarity between samples belonging to the same sample pair is the largest, and the similarity between samples belonging to different sample pairs is the smallest; (4) Repeat step (3) until the model converges, that is, the value of the loss function no longer decreases significantly, to obtain a trained predictive model. Among them, during the training process, save the model parameters with the best performance.

Citation Information

Patent Citations

  • Phosphorite metallogenic rule text data mining method and system

    CN116089629A

  • Mineral resource prediction method based on knowledge graph driving and storage medium

    CN116307123A

  • Bridge health maintenance knowledge graph construction method based on deep learning

    CN117131200A

  • Safety propaganda and education training knowledge graph and data management method and system based on AI

    CN118585658A

  • Method for metallogenic prediction by using multi-source heterogeneous information

    US20240310554A1