Natural resource data intelligent acquisition method

CN121658911BActive Publication Date: 2026-08-07WUDA GEOINFORMATICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUDA GEOINFORMATICS CO LTD
Filing Date
2026-02-06
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]对关键词依赖过高:一旦公告表述方式发生变化或关键词缺失,就容易导致识别错误或遗漏

Benefits of technology

[0018] This invention provides an intelligent data acquisition method for natural resources. The method involves: acquiring raw input data from multiple data sources; extracting structured semantic features from the raw input data; determining whether the raw input data belongs to the natural resources industry based on the structured semantic features and calculating a relevance score; triggering an evolutionary acquisition strategy stage, generating candidate acquisition action sequences based on the structured semantic features of the raw input data using a genetic algorithm mechanism; determining the optimal acquisition action sequence from the candidate sequences based on a natural resources industry knowledge graph using a deep reinforcement learning optimization module (DQN); and acquiring data based on the optimal acquisition action sequence. This invention combines deep reinforcement learning, evolutionary optimization, and industry knowledge enhancement mechanisms to construct an end-to-end intelligent acquisition framework for the natural resources industry. This framework enables efficient and accurate data acquisition in the natural resources industry and addresses issues such as multi-source heterogeneity, semantic complexity, and information gaps in the acquisition and processing of bidding/tendering data in the natural resources industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658911B_ABST
    Figure CN121658911B_ABST
Patent Text Reader

Abstract

The application provides a natural resource data intelligent collection method, which comprises the following steps: obtaining original input data from multiple data sources; extracting the structured semantic features of the original input data; judging whether the original input data belongs to the natural resource industry according to the structured semantic features and calculating the correlation score; triggering the collection strategy evolution stage, generating a candidate collection action sequence based on the genetic algorithm mechanism according to the structured semantic features of the original input data; deciding the optimal collection action sequence from the candidate collection action sequence based on the deep reinforcement learning optimization module DQN according to the natural resource industry knowledge graph; and collecting data based on the optimal collection action sequence. The application combines deep reinforcement learning, evolutionary optimization and industry knowledge enhancement mechanism to construct an end-to-end intelligent collection framework for the natural resource industry, and can realize efficient and accurate collection of natural resource industry data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data acquisition in the natural resources industry, and more specifically, to an intelligent method for acquiring natural resource data. Background Technology

[0002] Traditional data collection, feature extraction, and data classification methods mainly rely on keyword matching, regular expressions, rule bases, and template-based field parsing. In the collection phase, web crawlers are typically used to obtain web announcements or PDF documents, which are then initially filtered using keywords or specific tags. In the feature extraction phase, regular expressions or predefined templates are used to extract key fields such as amount, time, location, and project name from the announcement text.

[0003] However, these traditional methods have significant problems:

[0004] Over-reliance on keywords: If the way the announcement is worded changes or keywords are missing, it can easily lead to identification errors or omissions.

[0005] Poor adaptability: The announcement formats vary greatly across different regions and platforms, making it difficult for the rule base to cover all scenarios and resulting in high maintenance costs.

[0006] Difficulty in handling unstructured data: Many announcements exist in the form of long texts, scanned copies, or complex PDFs, and traditional methods are insufficient for parsing complex documents.

[0007] Weak semantic understanding: It is unable to understand the meaning of the context by relying solely on keywords and rules, and it is difficult to distinguish similar terms or vague descriptions across industries.

[0008] Limited scalability: As the scope of industries expands and the scale of data grows, traditional methods struggle to support refined classification across multiple industries and contexts.

[0009] With the development of artificial intelligence technology, text processing capabilities based on large language models and classification algorithms based on deep learning can solve some data extraction and classification problems, but there is currently a lack of a method for extracting and classifying data features specifically for the natural resources industry. Summary of the Invention

[0010] This invention addresses the technical problems existing in the prior art by providing an intelligent method for collecting natural resource data.

[0011] This invention provides an intelligent method for collecting natural resource data, comprising:

[0012] Obtain raw input data from multiple data sources;

[0013] The structured semantic features of the original input data are extracted based on the Structured Self-Attention and Cognitive Encoding (SSAM) module.

[0014] Based on the structured semantic features, the industry relevance enhancement mechanism module IREM determines whether the original input data belongs to the natural resources industry, calculates the relevance score, and filters the original input data based on the relevance score.

[0015] The acquisition strategy evolution phase is triggered. Based on the structured semantic features of the filtered original input data, the Evolutionary Action Generation Module (EAG) generates candidate acquisition action sequences through a genetic algorithm mechanism.

[0016] Based on the knowledge graph of the natural resources industry, the optimal collection action sequence is determined from the candidate collection action sequences using the deep reinforcement learning optimization module DQN.

[0017] Data is collected based on the optimal collection action sequence to obtain collected data and generate structured project data.

[0018] This invention provides an intelligent data acquisition method for natural resources. The method involves: acquiring raw input data from multiple data sources; extracting structured semantic features from the raw input data; determining whether the raw input data belongs to the natural resources industry based on the structured semantic features and calculating a relevance score; triggering an evolutionary acquisition strategy stage, generating candidate acquisition action sequences based on the structured semantic features of the raw input data using a genetic algorithm mechanism; determining the optimal acquisition action sequence from the candidate sequences based on a natural resources industry knowledge graph using a deep reinforcement learning optimization module (DQN); and acquiring data based on the optimal acquisition action sequence. This invention combines deep reinforcement learning, evolutionary optimization, and industry knowledge enhancement mechanisms to construct an end-to-end intelligent acquisition framework for the natural resources industry. This framework enables efficient and accurate data acquisition in the natural resources industry and addresses issues such as multi-source heterogeneity, semantic complexity, and information gaps in the acquisition and processing of bidding / tendering data in the natural resources industry. Attached Figure Description

[0019] Figure 1 A flowchart of a method for intelligent acquisition of natural resource data provided in one embodiment of the present invention;

[0020] Figure 2 This is a schematic diagram showing the connection between four modules in one embodiment of the present invention;

[0021] Figure 3 A flowchart of the workflow for the Industry Relevance Enhancement Mechanism (IREM) module;

[0022] Figure 4 Workflow diagram for the Evolution Action Generation Module (EAG). Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined with each other to form feasible technical solutions. Such combinations are not constrained by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0024] This invention aims to address the problems of multi-source heterogeneity, semantic complexity, and information gaps in the collection and processing of bidding / tendering data in the natural resources industry. It proposes an intelligent data collection method for natural resources: a Cognitive Evolution & Industry-Relevance Enhanced Deep Q-Network (CE-IREM-DQN). By combining deep reinforcement learning, evolutionary optimization, and industry knowledge enhancement mechanisms, an end-to-end intelligent collection framework for the natural resources industry is constructed, enabling efficient extraction and accurate classification of data from web announcements, PDF files, and structured project data.

[0025] Figure 1 The following is a flowchart illustrating an intelligent data acquisition method for natural resources according to an embodiment of the present invention. Figure 1 As shown, the method includes the following steps:

[0026] Step 1: Obtain raw input data from multiple data sources.

[0027] It is understood that this invention includes four core modules: the Structural Self-Attention Module (SSAM), the Industry-Relevance Enhancement Mechanism Module (IREM), the Evolutionary Action Generator Module (EAG), and the Deep Q-Network Decision Optimizer Module (DQN). The connections between these four modules can be found in [reference needed]. Figure 2 .

[0028] This invention obtains information to be processed from multiple data sources, including government transaction platform announcements, industry websites, PDF files, scanned copies, etc., as raw input data.

[0029] Step 2: Extract the structured semantic features of the original input data based on the Structured Self-Attention and Cognitive Encoding (SSAM) module.

[0030] In one embodiment of the present invention, the Structural Self-Attention and Cognitive Encoding Module (SSAM) includes a structural parsing submodule and a cognitive structure graph convolutional encoding module. The extraction of structured semantic features from the original input data based on the SSAM module includes:

[0031] The original input data is transformed into a graph structure based on the structure parsing submodule.

[0032] The cognitive structure graph convolutional coding module represents the semantic feature vector of each node in the graph structure.

[0033] Understandably, the Structural Self-Attention and Cognitive Encoding Module (SSAM) is used to extract structured semantic features from the original input data. The SSAM module consists of two parts: a structure parsing submodule and a cognitive structure graph convolutional encoding module.

[0034] The Structure Parsing submodule is responsible for "building the skeleton of the graph". Its core goal is to convert the original web page / PDF / OCR document into a structured graph representation.

[0035] The Cognitive Structure Graph Convolutional Encoding module is responsible for "enabling this graph to learn and understand industry semantics." It is the semantic understanding and feature enhancement stage of SSAM. Based on the structural analysis results, it uses graph neural networks and cognitive attention mechanisms to achieve integrated representation learning of structure, semantics, and cognition.

[0036] (1) Structure parsing submodule.

[0037] The structure parsing submodule serves as the input and graph construction phase of SSAM, with the core objective of converting raw web pages / PDFs / OCR documents into structured graph representations.

[0038] 1. First, convert the input data (web page announcements, PDF documents, or scanned document recognition results) into a structured representation.

[0039] 2. For web page data, parse its DOM tree structure and extract node levels, tag types, text content, and adjacency relationships;

[0040] 3. Construct a text paragraph hierarchy tree for PDF documents or OCR recognition results, recording features such as paragraph number, title position, font size, and indentation level.

[0041] 4. The parsing results are represented in the form of a graph structure, where nodes correspond to text units and edges represent structural or semantic connections.

[0042] (2) Cognitive Structural Graph Convolution Encoding.

[0043] To enhance the understanding of the complex hierarchical structure and semantic dependencies in natural resource industry announcements, this invention proposes a Cognitive Structural Graph Convolution (CSGC) mechanism based on traditional Graph Convolutional Networks (GCN) and Graph Attention Networks (GAT) to achieve multi-dimensional semantic modeling of webpage / PDF document structures. The main implementation details are as follows:

[0044] ① Node Feature Construction (Multi-Source Feature Fusion):

[0045] The initial feature vector of each node is no longer composed solely of general text word embeddings and structural attributes, but rather integrates three types of information: word semantic embedding, structural feature embedding, and domain semantic embedding.

[0046]

[0047] in:

[0048] : Contextual semantic embeddings generated by the BERT pre-trained language model;

[0049] Based on the node's DOM hierarchy, font style, paragraph number, and tag category (e.g., ...). <title>、< / title> Structural features of encoding (etc.);

[0050] Domain concept embeddings generated through the industry knowledge graph (provided by the IREM module) are used to reflect the relevance of nodes to natural resource domain terms such as “mining area scope”, “land use planning”, and “forest land occupation”.

[0051] This triple feature fusion endows nodes with both semantic understanding capabilities and industry contextual awareness. This design significantly enhances the model's comprehensive perception of industry terminology, contextual semantics, and structural context, providing a more semantically discriminative representational foundation for subsequent graph neural network structure encoding.

[0052] ② Cognitive Attention Weighting.

[0053] Traditional graph attention mechanisms rely solely on the feature similarity of adjacent nodes to calculate attention weights. This invention introduces a cognitive relevance function (CRF) to improve attention calculation:

[0054]

[0055] in:

[0056] Node characteristics; In bidding documents in the natural resources industry (such as mining, water conservancy, land development, etc.), a node can represent a document element (title, paragraph, table item, indicator field, etc.). Node characteristics These are the semantic vectors of these elements. By capturing semantic content or node attributes, the attention mechanism can understand whether two nodes are related in terms of content.

[0057] For example, the contents of node i and node j are shown in Table 1 below.

[0058] Table 1. Contents of node i and node j

[0059] node i Project Name: Announcement of Mining Rights Transfer for a Gold Mine <![CDATA[(h i ): A text vector encoded by BERT, reflecting the semantics of "mining rights transfer". Expressing project type and regional characteristics Node j Resource reserves: 28 tons of gold. <![CDATA[(h j ): Semantic vectors extracted through structured extraction, representing specific resource metrics. Expressing information on resource quantity and economic value

[0060] This enables the attention mechanism to understand whether two nodes are semantically related, for example:

[0061] (1) The terms “mining rights transfer announcement” and “resource reserves” are highly related semantically;

[0062] (2) "Contact information of bidding entities" is not directly related to "resource reserves". Therefore, in the attention weighting, The former will be higher.

[0063] Structural relationship: Describes the structural or logical relationship between nodes i and j in the bidding document, for example:

[0064] (1) Whether they belong to the same chapter;

[0065] (2) Does a title-content subordination exist?

[0066] (3) Whether they are in the same row / column in the table;

[0067] (4) Whether they are adjacent to each other.

[0068] : The cognitive context representation of a node (generated by an external model or cognitive memory module).

[0069] Cognitive Relevance Function.

[0070] This improvement enables attention weights to be dynamically adjusted based on industry knowledge, thereby strengthening semantically close but structurally dispersed node connections, such as establishing a cross-layer association between "Land Use Size" and "Total Area" in the table below it.

[0071] The implementation approach of CRF in bidding projects in the natural resources industry:

[0072] Cognitive Relevance Function Used to evaluate nodes and The correlation weights between them at the semantic, structural, and cognitive levels guide the graph attention mechanism in allocating attention intensity.

[0073] In bidding documents for projects in the natural resources industry (such as mining rights transfer announcements, water conservancy project bidding documents, land transfer results announcements, etc.), CRF can significantly improve the model's ability to understand "information logical relationships", especially in complex hierarchical structures and industry terminology.

[0074] The structure of a CRF is as follows:

[0075] in, This is a semantic relevance subfunction, used to measure the semantic relevance between two text nodes. Industry explanation: In natural resource bidding projects, different text segments often express different dimensions of the project (such as "project overview", "geological conditions", "investment scale", "bidding qualifications", etc.).

[0076] CRF uses this sub-function to identify semantically relevant parts, for example:

[0077] “Location of mining area” ↔ “Latitude and longitude coordinates”;

[0078] "Resource reserves" ↔ "Mineral type and grade";

[0079] "Transfer Method" ↔ "Bidding Method, Deposit".

[0080] in, Using a learnable weight matrix Model semantic similarity.

[0081] in, These are structure-related sub-functions that encode structural relationships within a document or webpage, such as hierarchy, chapters, tables, or logical order. Industry Note: Tender announcements or winning bid notices typically have a fixed logical structure.

[0082] Title (Project Name) → Paragraph (Project Overview) → Table (Resource Indicators, Investment Amount);

[0083] Parent-child relationships between chapters, fields in the same column in tables, and the location of appendix information are all structural information that helps the model understand which fields are related to each other.

[0084] in, ,in, This is the structurally relevant weight matrix. Includes structural features (such as hierarchy difference, relative position, chapter ID, table row and column numbers).

[0085] in, This is a cognitive-related subfunction, whose role is to introduce external cognitive information, domain knowledge, or task context.

[0086] Industry Explanation: In the natural resources industry, cognitive-level information can come from:

[0087] (1) Industry knowledge graph (e.g., “mineral rights transfer” associated with “Ministry of Natural Resources announcement”);

[0088] (2) Task intent (e.g., the current task is "extracting winning bid information" or "identifying land parcel attributes");

[0089] (3) Historical learning experience (patterns identified by the model in past data collection tasks); this module enables the attention mechanism to "understand industry semantics", for example:

[0090] (1) Knowing that "transfer method" and "payment method" belong to the same logical domain;

[0091] (2) Knowing that the information on “geological exploration blocks” has a higher priority in the “mining rights announcement”.

[0092] in, .

[0093] In the formula, This is the Cognitive Relation Weight Matrix, used to weight cognitive vectors. and The interaction relationships between them are parametrically modeled, thereby explicitly injecting industry knowledge, task context and historical experience. It carries high-level cognitive information, such as:

[0094] Industry concepts (mining rights type, transfer method, geological block); task context (extraction of mining rights announcement / bid results); historical experience (successful field combinations in previous data collection tasks).

[0095] Industry example: "Transfer method" and "Payment method" may appear far apart in text, but they belong to the same business logic domain at the cognitive level. It will amplify the interaction weight of the corresponding dimensions of the two.

[0096] The comprehensive expression of CRF is as follows:

[0097] ;

[0098] After softmax normalization:

[0099] Step 3: Based on the structured semantic features, determine whether the original input data belongs to the natural resources industry based on the industry relevance enhancement mechanism module IREM, calculate the relevance score, and filter the original input data based on the relevance score.

[0100] It is understood that the embodiments of the present invention mainly collect relevant data in the natural resources industry. Therefore, the original input data collected in step 1 is filtered to select relevant data in the natural resources industry.

[0101] Specifically, the structured semantic features of the original input data proposed in step 2 are input into the Industry Relevance Enhancement Mechanism (IREM) module to determine whether the original input data belongs to the natural resources industry.

[0102] In one embodiment of the present invention, based on the structured semantic features, the industry relevance enhancement mechanism module IREM determines whether the original input data belongs to the natural resources industry and calculates a relevance score, including:

[0103] Vectorize the entities, terms, and announcement texts in the domain knowledge graph of the natural resources industry to obtain knowledge vectors, and form a knowledge vector index library.

[0104] The structured semantic features are matched against the vector feature index library to calculate a relevance score;

[0105] If the correlation score is greater than the preset threshold score, it means that the original input data belongs to the natural resources industry and is retained; otherwise, the original input data does not belong to the natural resources industry and is discarded.

[0106] The goal of the Industry Relevance Enhancement Mechanism (IREM) module is to inject domain knowledge (knowledge graph + terminology ontology) of the natural resources industry into document encoding, calculate the multidimensional relevance score between the document and industry entities / concepts, and use this relevance score as a reward / guiding signal for reinforcement learning (DQN) or evolutionary module (EAG), thereby prioritizing the collection of data that is highly relevant to natural resources and improving the accuracy of data collection and business availability.

[0107] See Figure 3 The Industry Relevance Enhancement Mechanism (IREM) module includes the following sub-modules:

[0108] Knowledge layer: Industry knowledge graph (KG) and terminology ontology (including synonyms, units, regular expression templates, etc.).

[0109] Index layer: Vector Database and Graph Database.

[0110] Semantic alignment layer (core): entity recognition and linking, semantic similarity retrieval, path reasoning, score fusion and normalization.

[0111] Interface layer: External API, used to provide SSAM, EAG, and DQN with correlation, interpretation path, and confidence information, and to receive feedback for online updates.

[0112] The following is an introduction to each level.

[0113] 1. Knowledge Layer.

[0114] Purpose: To build the industry knowledge base for the natural resources sector and provide structured information support for subsequent semantic alignment.

[0115] Provides entity and relationship information for the index layer, used for vectorization and graph database construction.

[0116] It provides a reference for entity recognition and semantic matching for the semantic alignment layer.

[0117] The knowledge layer is composed of the following:

[0118] (1) Industry Knowledge Graph (KG):

[0119] Describe the entities involved in the bidding announcement, such as enterprises, projects, regions, resource types (e.g., minerals, forestry, water conservancy), and regulatory agencies, and their relationships.

[0120] For example, "Mine A belongs to Company B", "Project C is located in Region D".

[0121] (2) Terminology:

[0122] This includes professional terminology, synonyms, unit conversions, regular expression templates, etc.

[0123] For example, "iron ore mining" and "iron ore mining" are considered synonymous, and "km²" and "square kilometer" can be treated uniformly.

[0124] Specific implementation:

[0125] When constructing the ontology, standards in the field of natural resources (such as announcements from the Ministry of Land and Resources and mineral industry standards) and existing publicly available bidding data are combined.

[0126] The knowledge graph uses an RDF / OWL structure for storage, where nodes represent entities and edges represent relationships.

[0127] The thesaurus, unit conversion rules, and regular expression templates are stored in a dictionary or configuration file and can be dynamically updated.

[0128] 2. Index Layer.

[0129] Function: To efficiently index entities and relationships in the knowledge layer to support fast retrieval and similarity calculation.

[0130] The index layer is composed of the following:

[0131] (1) Vector Database:

[0132] Vectorize entities, terms, announcement texts, etc. in the knowledge graph (using language models such as Word2Vec, BERT, and SimCSE).

[0133] It supports semantic similarity calculation, such as vector cosine similarity.

[0134] (2) Graph Database:

[0135] Store knowledge graph nodes and relationships for path reasoning and relationship querying (such as Neo4j and ArangoDB).

[0136] Specific implementation:

[0137] Entity extraction and vectorization are performed on the announcement text, company information, and project description.

[0138] Embed knowledge graph nodes into a vector space to form a vector index library (Faiss or Milvus can be used).

[0139] In a graph database, nodes store entity types and attribute information, while edges store relationship types and weights.

[0140] 3. Semantic Alignment Layer (core module).

[0141] Function: To achieve accurate matching between announcement text and industry knowledge graph, calculate relevance and provide explanation paths, it is the core of IREM.

[0142] Submodules and functions of the semantic alignment layer:

[0143] (1) Entity recognition and linking (NER & EL).

[0144] The system identifies entities such as companies, projects, regions, and resource types in the announcement text and links them to knowledge graph nodes.

[0145] Example: Identify "Hunan Mining Co., Ltd." as a corporate entity and link it to the corresponding node in the knowledge graph.

[0146] (2) Semantic similarity search.

[0147] Calculate the similarity between announcement text and entities or terms in a knowledge graph using a vector library.

[0148] It can be combined with word vectors, sentence vectors, and contextual relationships.

[0149] (3) Path Reasoning.

[0150] Calculate the relationship path from text entities to target entities based on graph databases.

[0151] Example: Announcement Project → Related Companies → Industry Category → Project Type.

[0152] (4) Score Fusion & Normalization.

[0153] The matching scores (vector similarity, path distance, entity matching confidence) from different sources are weighted and fused.

[0154] Output a uniform relevance score [0,1] to facilitate subsequent decision-making.

[0155] Specific implementation:

[0156] NER / EL can use deep learning models (such as BERT+CRF) combined with regular templates to improve the recognition rate of technical terms.

[0157] Vector retrieval can be accelerated using the ANN (Approximate Nearest Neighbor) algorithm.

[0158] Path reasoning can be performed using graph search algorithms (BFS, Dijkstra, Path Ranking Algorithm) to compute and interpret paths.

[0159] Fusion can be achieved by using weighted averaging or machine learning models (such as LightGBM) to combine multiple scores.

[0160] 4. Interface Layer (API Layer).

[0161] Function: Provides IREM services to external users, connects to SSAM, EAG, and DQN modules, and enables data sharing and online updates.

[0162] Function:

[0163] (1) Provide relevance and explanation path.

[0164] For each announcement, return: related entity, relevance score, path explanation, and confidence level.

[0165] (2) Receiving feedback.

[0166] Feedback collected from SSAM, EAG, and DQN (such as action execution results and collection accuracy).

[0167] The knowledge graph, vector library, and score fusion weights are updated online.

[0168] (3) Supports batch and real-time queries.

[0169] It can score and analyze the paths of single or batch announcements.

[0170] Specific implementation:

[0171] RESTful API or gRPC service;

[0172] JSON or Protobuf can be used as the data exchange format.

[0173] Supports asynchronous task queues (such as Celery) for handling large-scale announcements.

[0174] Step 4: Trigger the acquisition strategy evolution stage. Based on the structured semantic features of the filtered original input data, candidate acquisition action sequences are generated using the genetic algorithm mechanism based on the Evolutionary Action Generation Module (EAG).

[0175] In one embodiment of the present invention, based on the structured semantic features of the filtered original input data, a candidate acquisition action sequence is generated using a genetic algorithm mechanism based on the Evolutionary Action Generation Module (EAG), including:

[0176] Based on the structured semantic features of the filtered raw input data, the minimum executable atomic action is defined, and the atomic actions are combined in sequence to generate the initial generation of executable acquisition action sequences;

[0177] Calculate the fitness of each executable acquisition action sequence in the initial generation, select multiple executable acquisition action sequences in the initial generation based on the fitness, and perform evolutionary operations to generate the next generation of executable acquisition action sequences.

[0178] The iterative evolution operation continues until the iteration converges or the maximum number of iterations is reached, resulting in a candidate collection action sequence.

[0179] Specifically, the Evolutionary Action Generation (EAG) module is mainly used in the process of collecting bidding announcements for natural resources to optimize action sequences through evolutionary algorithms, thereby improving the adaptability, completeness, and relevance of the collection strategy. EAG treats "collection actions" as individuals of a genetic algorithm, generating high-quality collection actions through mechanisms such as population evolution, fitness evaluation, and crossover mutation, thereby guiding the data collection system to achieve efficient, accurate, and domain-relevant content capture.

[0180] The core objectives of the Evolutionary Action Generation Module (EAG) include:

[0181] 1. Discover highly adaptive action combinations in complex action spaces;

[0182] 2. Maximize the industry relevance, completeness, and spatial consistency of the collected data;

[0183] 3. Support the integration with reinforcement learning (such as DQN) to achieve a closed loop of action execution and policy optimization.

[0184] The Evolution Action Generation Module (EAG) mainly consists of the following sub-modules:

[0185] 1. The Atomic Action Generator has the following functions:

[0186] Define the minimum executable action, such as field extraction, category recognition, content filtering, etc. Each atomic action includes attributes such as action type, execution conditions, and output format.

[0187] 2. Action Sequence Encoder, its function is:

[0188] Atomic actions are combined sequentially to generate executable action sequences. Vectorized representation of action sequences is supported, facilitating fitness evaluation and evolutionary operations.

[0189] 3. Fitness Evaluator, its functions are:

[0190] A comprehensive evaluation of the execution effect of the action sequence is conducted.

[0191] The fitness function formula is as follows:

[0192] in, Action sequence; Industry relevance score; Completeness score of data collection; Spatial consistency score; : Calculation and execution costs; Novelty of action sequences; Adjustable weight parameters.

[0193] 4. Evolutionary Operator:

[0194] The previous generation of action sequences is processed through operations including selection, crossover, and mutation. In each iteration, a new generation of action sequences is generated and compared with and replaced with the best historical sequence.

[0195] 5. Action Optimization Loop, its function is as follows:

[0196] The results of the action sequence execution are fed back to the fitness evaluator, which supports the reinforcement learning strategy to further optimize the action execution strategy and achieve adaptive evolution.

[0197] For a detailed flowchart of EAG's execution process, please refer to [link / reference]. Figure 4 This includes the following steps:

[0198] 1. Atomic Action Initialization: Generate an initial set of atomic actions based on a predefined action library.

[0199] 2. Action sequence generation: Atomic actions are combined in sequence to form the initial population. .

[0200] 3. Fitness assessment: For each action sequence Perform a data acquisition simulation and calculate... .

[0201] 4. Evolutionary Operation:

[0202] Selection: High-quality sequences are retained based on fitness;

[0203] Crossover: Randomly combines parent sequences to generate offspring;

[0204] Mutation: Randomly modifying atomic actions in a sequence to increase diversity.

[0205] 5. Iterative updates:

[0206] Repeat steps 3-4 until the maximum number of iterations or the convergence condition is reached, and output the candidate action sequence. .

[0207] 6. Action sequence execution and feedback:

[0208] Candidate action sequence The data collection task is actually performed, and the results are fed back to the reinforcement learning module to optimize the strategy.

[0209] Step 5: Based on the natural resources industry knowledge graph, the optimal collection action is determined from the candidate collection action sequence using the deep reinforcement learning optimization module DQN.

[0210] Understandably, the Reinforcement Decision and Optimization (DQN) module aims to intelligently optimize action sequences in natural resource data collection. Its core objective is to automatically select the optimal action from a given "collection action space" (candidate collection action sequences) using reinforcement learning methods, thereby improving the relevance, completeness, and spatial consistency of the data collection.

[0211] The enhanced decision-making and optimization module DQN can work in conjunction with IREM (Industry Relevance Enhancement Mechanism) to optimize action strategies in real time by utilizing industry knowledge graphs, terminology ontology, and semantic matching results.

[0212] In one embodiment of the present invention, the optimal acquisition action is determined from the candidate acquisition action sequence based on the natural resources industry knowledge graph and a deep reinforcement learning optimization module (DQN), including:

[0213] Encode each action in the candidate acquisition action sequence into an action vector a;

[0214] Construct the environmental state s for each action vector a, wherein the environmental state s includes a correlation score based on the action acquisition data. Integrity score Spatial consistency score Current resource consumption and novelty index The environmental state s is represented as ;

[0215] Select the current action vector Calculate the current action vector environmental conditions and reward function value ;

[0216] Select the next action vector based on the reward function value. ;

[0217] The selection process continues until all actions in the candidate acquisition action sequence have been selected. The optimal acquisition action sequence is then determined, and the candidate acquisition action sequence is updated based on the optimal acquisition action sequence.

[0218] Among them, the current action vector is calculated. environmental conditions and reward function value ,include:

[0219] In the formula, The weights can be dynamically adjusted according to business needs.

[0220] Through reinforcement learning via the reinforcement decision-making and optimization module DQN in step 5, the optimal action sequence is finally output. It interacts with the candidate action sequence generated by the EAG module, that is, it uses the output optimal action sequence to update the candidate action sequence generated by the EAG module for subsequent iterative optimization, and can be directly fed back into the state vector with the industry relevance score calculated by the IREM module to form a closed-loop optimization.

[0221] Step 6: Collect data based on the optimal collection action to obtain the collected data and generate structured project data.

[0222] Understandably, after obtaining the optimal acquisition action sequence in step 5, the optimal acquisition action sequence is executed to acquire data, and the acquired data is output to the system database to generate structured project data.

[0223] The output includes key fields such as project name, type, geographical location, plot number, area, and planned use;

[0224] The system records the status, actions, and reward information during the data collection process, forming a data collection log;

[0225] By continuously training and using an experience replay mechanism, the model strategy is continuously optimized to achieve adaptive learning.

[0226] This invention provides an intelligent data acquisition method for natural resources. By combining deep reinforcement learning, evolutionary optimization, and industry knowledge enhancement mechanisms, an end-to-end intelligent acquisition framework for the natural resources industry is constructed. This framework enables efficient and accurate data acquisition in the natural resources industry and can solve problems such as multi-source heterogeneity, semantic complexity, and information loss in the acquisition and processing of bidding / tendering data in the natural resources industry.

[0227] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0228] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0229] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0230] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0231] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0232] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0233] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for intelligent collection of natural resource data, characterized in that, include: Obtain raw input data from multiple data sources; The structured semantic features of the original input data are extracted based on the Structured Self-Attention and Cognitive Encoding (SSAM) module. Based on the structured semantic features, the industry relevance enhancement mechanism module IREM determines whether the original input data belongs to the natural resources industry, calculates the relevance score, and filters the original input data based on the relevance score. Trigger the evolution phase of the acquisition strategy. Based on the structured semantic features of the filtered original input data, the Evolutionary Action Generation Module (EAG) generates candidate acquisition action sequences through a genetic algorithm mechanism. Based on the knowledge graph of the natural resources industry, the optimal collection action sequence is determined from the candidate collection action sequences using the deep reinforcement learning optimization module DQN. Data is collected based on the optimal collection action sequence to obtain collected data and generate structured project data; The Structural Self-Attention and Cognitive Encoding (SSAM) module includes a structural parsing submodule and a cognitive structure graph convolutional encoding module. The extraction of structured semantic features from the original input data based on the SSAM module includes: The original input data is transformed into a graph structure based on the structure parsing submodule. The cognitive structure graph convolutional coding module represents the semantic feature vector of each node in the graph structure. The structure parsing submodule transforms the original input data into a graph structure, including: The original input data is parsed in a structured manner. For web page data, its DOM tree structure is parsed to extract node levels, tag types, text content, and adjacency relationships. For PDF documents or OCR recognition results, a text paragraph hierarchy tree is constructed, which is used to record paragraph numbers, heading positions, font sizes, and indentation levels; The parsing results are represented in the form of a graph structure, where nodes correspond to text units and edges represent structural or semantic connections between nodes.

2. The intelligent data acquisition method for natural resources according to claim 1, characterized in that, Based on the cognitive structure graph convolutional coding module, each node in the graph structure is represented by a semantic feature vector, including: Construct the feature vector of each node in the graph structure and the cognitive attention weights of its neighboring nodes; The feature vector of each node in the graph structure is represented as follows: ; in: : The contextual semantic embedding of the i-th node generated by the BERT pre-trained language model; The structural features of the i-th node, including its DOM hierarchy, font style, paragraph number, and tag category encoding. The domain concept embedding of the i-th node generated by the natural resources industry knowledge graph is used to reflect the degree of association between the node and natural resources domain terms. The cognitive attention weights of the adjacent nodes are represented as follows: ; in: : The eigenvectors of node i and node j in the graph structure; : The structural relationship between node i and node j, used to describe the structural or logical relationship between node i and j in the document; Cognitive context representations of nodes i and j; Cognitive relevance function.

3. The intelligent data acquisition method for natural resources according to claim 2, characterized in that, Represented as: ; in, express and The semantic correlation sub-function is used to measure the semantic correlation between node i and node j; express The structure-related sub-functions are used to measure the structural relationship between node i and node j in an encoded document or webpage; express and Cognitive-related sub-functions are used to introduce external cognitive information, domain knowledge, or task context; in, ; , This is the semantic relevance weight matrix; In the formula, This is the structurally relevant weight matrix; ; In the formula, This is a cognitive-related weight matrix used to... and The interaction relationships between them are modeled parametrically.

4. The intelligent data acquisition method for natural resources according to claim 1, characterized in that, Based on the structured semantic features, the industry relevance enhancement mechanism module IREM determines whether the original input data belongs to the natural resources industry and calculates a relevance score, including: Vectorize the entities, terms, and announcement texts in the domain knowledge graph of the natural resources industry to obtain knowledge vectors, and form a vector feature index library. The structured semantic features are matched against the vector feature index library to calculate a relevance score; If the correlation score is greater than the preset threshold score, the original input data belongs to the natural resources industry and is retained; otherwise, the original input data does not belong to the natural resources industry and is discarded.

5. The intelligent data acquisition method for natural resources according to claim 1, characterized in that, The process of generating candidate acquisition action sequences based on the structured semantic features of the filtered original input data and using a genetic algorithm mechanism based on the Evolutionary Action Generation (EAG) module includes: Based on the structured semantic features of the filtered raw input data, the minimum executable atomic action is defined, and the atomic actions are combined in sequence to generate the initial generation of executable acquisition action sequences; Calculate the fitness of each executable acquisition action sequence in the initial generation, select multiple executable acquisition action sequences in the initial generation based on the fitness, and perform evolutionary operations to generate the next generation of executable acquisition action sequences. The iterative evolution operation continues until the iteration converges or the maximum number of iterations is reached, resulting in a candidate collection action sequence.

6. The intelligent data acquisition method for natural resources according to claim 5, characterized in that, The calculation of the fitness of each executable acquisition action sequence in the initial generation includes: ;in: : Executable data acquisition action sequence; Executable data acquisition action sequence Industry relevance score; Executable data acquisition action sequence The score for the integrity of the data collection; Executable data acquisition action sequence Spatial consistency score; Executable data acquisition action sequence The computation and execution costs; Executable data acquisition action sequence Novelty; Adjustable weight parameters.

7. The intelligent data acquisition method for natural resources according to claim 1, characterized in that, The step of determining the optimal collection action sequence from the candidate collection action sequences based on the natural resources industry knowledge graph and the deep reinforcement learning optimization module DQN includes: Encode each action in the candidate acquisition action sequence into an action vector. ; Construct each action vector environmental conditions The environmental state Including correlation scoring based on motion capture data Integrity score Spatial consistency score Current resource consumption and novelty index The environmental state Represented as ; Select the current action vector Calculate the current action vector environmental conditions and reward function value ; Select the next action vector based on the reward function value. ; The selection process continues until all actions in the candidate acquisition action sequence have been selected. The optimal acquisition action sequence is then determined, and the candidate acquisition action sequence is updated based on the optimal acquisition action sequence.

8. The intelligent data acquisition method for natural resources according to claim 7, characterized in that, The calculation of the current action vector environmental conditions and reward function value ,include: In the formula, The weights can be dynamically adjusted according to business needs.

Citation Information

Patent Citations

  • Dynamic data pipeline construction method based on artificial intelligence and multi-modal data processing

    CN119830200A

  • Content recommendation method and system based on industry knowledge graph and reinforcement learning

    CN120994903A