Exploration and development data query method based on knowledge graph and large language model

By building an exploration and development data query system based on knowledge graphs and large language models, and analyzing and storing exploration and development report documents, non-professionals can quickly and conveniently obtain core information of exploration and development data, solving the problem of low information query efficiency, and improving user experience.

CN120256600APending Publication Date: 2025-07-04CHINA NATIONAL OFFSHORE OIL (CHINA) CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510372275.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the increase in the number of professional reporting documents makes it difficult for non-professionals to quickly obtain core information of exploration and development data, and the information is mixed and the query efficiency is low.

Method used

Using a method based on knowledge graphs and large language models, we analyze exploration and development report documents, build vector databases and graph databases, and provide fast query and visual display through keyword matching and synthesis.

Benefits of technology

It has realized that non-professionals can quickly and conveniently query exploration and development data, lowered the threshold for data query, and improved query efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256600A_ABST
    Figure CN120256600A_ABST
Patent Text Reader

Abstract

The invention relates to an exploration and development data query method based on a knowledge graph and a large language model, comprising: receiving an exploration and development report document uploaded by a client, the exploration and development report document carrying unstructured data under titles of all levels; analyzing and splitting the exploration and development report document, vectorizing the split titles and unstructured data under the titles, and storing the vectorized titles and the unstructured data in a vector database; constructing a knowledge graph according to each level of title of the exploration and development report document, and storing the knowledge graph in a graph database, the knowledge graph database comprising a plurality of nodes; receiving a query statement of the client based on the large language model, and extracting a keyword in the query statement; according to the keyword, performing likelihood calculation of the vector, determining a node which is most matched with the graph database, and extracting content under the node; and synthesizing the extracted contents under the nodes by using a large oracle model to obtain text information and / or chart information displayed to the client by the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of oil and gas exploration and development, and particularly to a method for querying exploration and development data based on a knowledge graph and a large language model. Background Art

[0002] All aspects of oil and gas exploration and development are becoming increasingly efficient and convenient.

[0003] However, the development of technology has also brought new challenges: the number of professional report documents has increased sharply. These professional reports have also erected a knowledge barrier for non-professionals, making it difficult for them to quickly grasp the core content. At the same time, the massive reports have led to information chaos, and it is also difficult for professionals to quickly find the key information they need.

[0004] Therefore, how to achieve efficient storage and quick query of report content has become the key to improving the overall work efficiency. If this problem can be solved, it can not only provide great convenience for relevant practitioners, but also help non-professionals quickly extract target information from complex exploration and development data reports, and further lower the threshold of information acquisition through visual display, providing a more intuitive and convenient experience for non-professionals. Summary of the Invention

[0005] The present invention provides a method for querying exploration and development data based on a knowledge graph and a large language model, which can automatically parse, split and store massive unstructured exploration and development data, and at the same time support users to perform quick, convenient and efficient query operations through natural language, without the need for professional knowledge of data processing, significantly reducing the data query threshold for users, shortening the user query time, and greatly improving the user experience and work efficiency.

[0006] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present application provides a method for querying exploration and development data based on a knowledge graph and a large language model, including: S1. Receive an exploration and development report document uploaded by a client, where the exploration and development report document carries unstructured data under each level of heading; S2. Parse and split the exploration and development report document, perform vectorization processing on the split headings at each level and the unstructured data thereunder, and store them in a vector database; S3. Construct a knowledge graph based on the headings at each level of the exploration and development report document, and store it in a graph database, where the knowledge graph includes a number of nodes; S4. Receive a query statement from the client based on the large language model, and extract keywords in the query statement; S5. Calculate the similarity of phase vectors based on the keywords, determine the most matching node in the graph database, and extract the content under the node; S6. Use the large prediction model to synthesize the content under the extracted node to obtain the text information and / or chart information presented to the client by the user.

[0007] In one implementation, in S2, it includes: Segment the exploration and development data report according to the headings at all levels therein; Divide the text content under each heading into text blocks according to a set threshold; Perform vectorization processing on each text block and store it in the vector database.

[0008] In one implementation, in S3, it specifically includes: Build a hierarchical relationship for the headings at all levels in the document and store them in the MarkDown text; Import the content of the text paragraphs under the headings at all levels into the corresponding hierarchical relationship; Parse the headings and hierarchical relationships of the MarkDown document and generate Cypher statements; Import the headings and hierarchical relationships into the knowledge graph through Cypher statements.

[0009] In one implementation, in S5, the specific steps include: Match the keywords with the nodes in the knowledge graph. If a matching node can be obtained, further search for the text vectors under the matching node; if no matching node can be obtained, search and match all the text vectors in the vector database.

[0010] In one implementation, before S6, it also includes identifying and classifying data items in the content under the extracted node to feedback the validity of the searched content.

[0011] In one implementation, if the searched content is valid, further submit it to the large language model for synthesis. If the searched content is not valid, return a search failure message.

[0012] Due to the above technical solutions adopted by the present invention, the following advantages are achieved: In the technical solution provided by the embodiments of the present invention, in terms of the storage of exploration and development unstructured data, the exploration and development report document input by the user transmitted from the client is received. By parsing and splitting the input document, the split fragments are vectorized and stored in the vector database. A knowledge graph is constructed based on the headings at all levels of the document and stored in the graph database, realizing the efficient storage of exploration and development unstructured data; in terms of the query of exploration and development unstructured data, the user query statement transmitted from the client is received, and the keywords in the query statement are extracted. The nodes in the knowledge graph are searched according to the keywords, and the content contained in the nodes is extracted. According to the extracted content, it is summarized and processed by the large language model, and the retrieved original text blocks and chart information are displayed to the user, reducing the technical requirements of the user for the database and improving the query efficiency of the data and the query experience of the user. Description of the Drawings

[0013] Figure 1 It is a flowchart of a method for storing and querying exploration and development unstructured data based on a knowledge graph and a large model according to an embodiment of the present invention; Figure 2 It is a flowchart of parsing, splitting and vectorizing an exploration and development document according to an embodiment of the present invention; Figure 3 It is a flowchart of a method for constructing a knowledge graph based on an exploration and development document according to an embodiment of the present invention; Figure 4 It is a flowchart of a method for querying exploration and development knowledge according to an embodiment of the present invention; Figure 5 It is a flowchart of a knowledge retrieval feedback according to an embodiment of the present invention. Detailed Embodiments

[0014] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the scope of protection of the present invention.

[0015] In view of the defects and problems of the prior art, the present application provides a method for querying exploration and development data based on a knowledge graph and a large language model, which is characterized by including: S1. Receive the exploration and development report document uploaded by the client, and the exploration and development report document carries unstructured data under each level of heading; S2. Parse and split the exploration and development report document, vectorize the split headings at all levels and the unstructured data thereunder, and store them in the vector database; S3. Construct a knowledge graph based on the headings at all levels of the exploration and development report document, and store it in a graph database. The knowledge graph includes several nodes. S4. Receive a query statement from the client based on a large language model, and extract the keywords in the query statement. S5. Calculate the similarity of the phase vectors according to the keywords, determine the most matching node in the graph database, and extract the content under the node. S6. Use the large prediction model to synthesize the content under the extracted node to obtain the text information and / or chart information displayed to the user by the client.

[0016] The above method will be described in a more detailed embodiment with reference to more attached drawings. After a detailed embodiment of the present application, an example will be described to illustrate the practical effect of the method.

[0017] Detailed Embodiment As Figure 1 shown, the present invention provides a method for storing and querying unstructured data for exploration and development based on a knowledge graph and a large language model, which specifically includes the following steps: S110. Receive the exploration and development report document transmitted from the client; S120. Parse and split the document, perform vectorization processing on the split segments, and store them in a vector database; S130. Construct a knowledge graph according to the headings at all levels of the document, and store it in a graph database; S140. Receive the user query statement transmitted from the client, and extract the keywords in the query statement; S150. Search for nodes in the knowledge graph according to the keywords, and extract the content included in the nodes; S160. Summarize and process the extracted content using a large language model, and display the retrieved original text blocks and chart information to the user.

[0018] In the first example, the exploration and development report document transmitted from the client, where the client can be understood as the user side, and the user can upload the exploration and development report document through devices such as mobile phones and personal computers.

[0019] In the first example, the method of parsing and splitting the document, performing vectorization processing on the split segments, and storing them in a vector database, as Figure 2 shown, the specific steps include: S210. Split the exploration and development data report in a preset manner.

[0020] It should be noted that the preset splitting method is mainly set according to the document title, and it is necessary to split through headings at different levels.

[0021] S220. Subdivide the text content under each title according to the set threshold.

[0022] It should be noted that the subdivision operation refers to further subdividing the corresponding text paragraphs under each level of title into multiple small text blocks. By identifying each piece of text content and intercepting the text content according to the threshold, smaller text blocks are obtained for subsequent further detailed operations and processing. The advantage of this is that on the one hand, when constructing a knowledge graph, it can prevent the content stored in a leaf node from being too large, reducing the matching speed and knowledge search speed. On the other hand, during the process of text vectorization, smaller text blocks can also be quickly vectorized and stored in the vector database, reducing the overall processing time and improving the speed of the model's initial parsing and processing of the document.

[0023] It should be known that the threshold refers to the size of how many characters are to be divided into a text block. The setting of the threshold is subjectively operated by people, and the setting of the threshold is greatly affected by the experience of developers and the text structure. If the threshold is set too large, each subdivided text block will be larger, affecting the model's subsequent knowledge recall from the vector library. If the threshold is set too small, it will also cause the text blocks to be cut too small, resulting in loss of context and semantics.

[0024] S230. Vectorize each text block and store it in the vector database.

[0025] It should be noted that in the vectorization process, the BERT encoding model is selected to map the text data to numerical vectors in a high-dimensional space. Subsequently, a Milvus vector database is created to store these numerical vectors. In the vector database, the vector of each text block is associated with a unique identifier for easy management and retrieval. At the same time, to improve the retrieval efficiency, an index is established for the vectors in the database, and an inverted index or a tree-based index structure can be selected. Finally, using the interface provided by the vector database, the converted vector data is inserted into the database, and the synchronization update of the index is ensured to achieve efficient vector retrieval, ensuring the quick and accurate retrieval of relevant information from a large amount of text data.

[0026] It should be known that the vector database is essentially a data storage structure that stores entities and their relationships in a vectorized manner, thus supporting efficient and fast similarity calculation and retrieval operations.

[0027] In the first example, the method of constructing a knowledge graph according to the titles at all levels of the document and storing it in the graph database is as Figure 3 shown, and the specific steps include: S310. Build the hierarchical relationship for headings at all levels in the document and store it in a MarkDown text.

[0028] It should be noted that the hierarchical structure within the document refers to the organizational relationship between various headings and subheadings in the document. For example, structures like "Chapter 1", "2.1", "2.1.1" are usually used to represent different levels and logical relationships of the content. Such a hierarchical structure can clearly show the organization method of the document content. On the one hand, within the article, readers can quickly understand the overall framework of the document and the interrelationships of each part. On the other hand, it also has a good organizational ability for the document content. Especially when dealing with exploration and development documents, it can significantly improve the efficiency and accuracy of information retrieval.

[0029] It should be explained that building the hierarchical relationship requires analyzing the segmented heading blocks, identifying the relationships between different headings based on the numbers before the headings, and constructing the hierarchical structure. By identifying the hierarchy, order, and inclusion or subordination relationships between headings, a hierarchical tree structure can be formed. Each heading will be assigned a corresponding hierarchical number according to the heading level in front of it, such as chapter headings, first-level subheadings, second-level subheadings, third-level subheadings, etc. In this way, the document can be combined from different initial document formats into a unified organizational structure, while ensuring that the context information of each part can be accurately identified and processed during subsequent processing.

[0030] It should be noted that the main reason for choosing a Markdown document to store the hierarchical structure lies in the simplicity of its format and wide compatibility. Markdown files are not only easy to read and edit, but also can be easily processed and parsed by various text processing tools, version control systems, and document generation tools. On the other hand, the structured characteristics of Markdown documents make them an ideal choice for storing hierarchical relationships. In Markdown, the hierarchical relationship can use Markdown's heading syntax, such as "# Chapter 1", "## Chapter 2", " Section 2.1", etc., to intuitively display headings at different levels and keep the document content structured. It is convenient for subsequent automated processing and parsing, improving the operability and processing efficiency of the document.

[0031] S320. Import the content of the text paragraphs under each level of headings into the corresponding hierarchical relationship.

[0032] It should be noted that, according to the hierarchical relationship constructed in S310, each text paragraph should be automatically classified into the corresponding hierarchical node according to the level to which its title belongs. For example, all the content under "Chapter 1" will be classified into the hierarchical node of this chapter. In this way, the content of the document can be stored in a standardized manner according to different levels such as large chapters and subsections at all levels, ensuring that the original document content is not lost after document parsing and document reconstruction. At the same time, maintaining the logical order and hierarchical relationship of the document content ensures that each part can be accurately located and quickly accessed in subsequent operations.

[0033] S330, parse the headings and hierarchical relationships of each level in the Markdown document and generate Cypher statements.

[0034] It should be noted that the hierarchical relationship of the headings in the Markdown file needs to be accurately matched through regular expressions and associated with the corresponding text paragraph content. For example, directly identify the first-level headings (starting with #) as the highest level of the document, and the second-level headings, third-level headings, etc. as its sub-levels, and so on. In this way, a tree structure or nested dictionary of the document can be constructed to clearly represent the hierarchical relationship of the document. In addition, other elements in the Markdown document, such as tables and pictures, can be properly processed to ensure the integrity of the parsing results.

[0035] After specifically parsing the Markdown document, a data structure in the form of a document tree or nested dictionary in a tree structure is output, which contains the hierarchical structures and corresponding contents of the parsed Markdown document. Subsequently, traverse this data structure, with each heading as (:Title {name: "heading name"}) and the text content as (:Content{text: "text content"}), and the hierarchical relationship is connected by MERGE (parent)-[:CONTAINS]->(child). Finally, CREATE or MERGE statements are generated, so that a complete set of insertion statements for the knowledge graph structure based on Cypher statements can be generated.

[0036] S340, import the headings and hierarchical relationships into the knowledge graph through Cypher statements.

[0037] It should be known that the Neo4j graph database is mainly used to store the knowledge graph. As a high-performance native graph database, Neo4j is specifically optimized for complex entity-relationship networks. It mainly uses the property graph model for data management, and its core consists of entities, relationships, and attributes, enabling it to flexibly represent multi-level association information in the real world.

[0038] It should be noted that the knowledge graph uses the titles recognized during the document parsing process as nodes, and the relationships are the subordinate relationships between various title nodes and the subordinate relationships between paragraph contents and titles. Through Cypher insertion statements, entities (title nodes and paragraph content nodes) and relationships (the subordinate structures between nodes) are inserted, and it is checked whether they conform to the schema constraints of the knowledge graph. During the parsing process, existing data is matched to determine whether to perform a create or merge operation to avoid redundant storage. After the data insertion is completed, a consistency check is performed on the knowledge graph, including entity integrity, relationship validity, and data redundancy detection, to ensure that all newly added data conforms to the expected storage specifications and does not create isolated nodes or redundant data. For abnormal data, the system automatically generates logs and provides adjustment suggestions, or automatically corrects them according to preset rules to ensure the integrity and rationality of the knowledge graph.

[0039] In the first example, the method of receiving the user query statement transmitted by the client and extracting the keywords in the query statement needs to use a large language model to extract the keywords from the query statement. Exemplarily, "In this document, what is the permeability of Well XXX". Then when the large language model searches and parses it, it will first extract the keywords as "this document", "Well XXX", "permeability", and "what". Subsequently, these keywords are used to search and match the content of the knowledge graph. After obtaining the keywords in the query statement, it is also necessary to perform a vectorization operation on the keywords for subsequent answer matching in the knowledge graph and semantic similarity search in the vector database.

[0040] In the first example, the method of searching for nodes in the knowledge graph according to the keywords and extracting the content contained in the nodes is as Figure 4 shown, and the specific steps include: S410, match the keywords with the nodes in the knowledge graph. If a match is found, go to S420; if no match is found, go to S430.

[0041] It should be noted that the nodes in the knowledge graph represent the title information used when constructing the knowledge graph from the document. The matching process calculates the semantic similarity between the keywords and the graph nodes, and selects the top-level node with the highest similarity. Next, enter the title node for detailed content matching and knowledge search.

[0042] It should be noted that the matching method based on vector semantic similarity used in the keyword matching process mainly uses semantic similarity to measure whether similar nodes are found. The semantic similarity matching method is an index to measure the similarity degree of two words or concepts at the semantic level. Different from the traditional similarity method based on, semantic similarity evaluates whether two words are semantically close by considering the semantic relationship factors of the words. The method for calculating semantic similarity provided in this implementation example includes the evaluation of the structural similarity of graph nodes. By calculating the relationship between the input keywords and the nodes in the knowledge graph, the cosine similarity between them is calculated to quantify their semantic similarity. Specifically, given a vector representation of a keyword and a vector representation of a node , their semantic similarity is calculated by this formula:

[0043] represents the dot product of vectors, is the L2 norm, representing the length of the vector. Through this similarity calculation, the proximity of the keyword and the node in the semantic space can be evaluated. When the similarity exceeds a certain threshold, it can be determined that relevant content is queried, and then the nodes with inclusion relationships under this node are entered to perform a detailed match on the text content of the passage.

[0044] S420, The knowledge graph node matching is successful, and search for the text vectors included under this node in the vector database.

[0045] It should be noted that if the keyword successfully matches the node, then a detailed search will be performed on all the document contents included under this title node, and the most similar document information will be returned in a dictionary structure, where the search and matching methods are the same as S410.

[0046] S430, The knowledge graph node matching fails, and comprehensively match the text vectors in the vector database.

[0047] It should be noted that if the input keyword fails to match the nodes in the knowledge graph, the data in the vector database will be comprehensively matched, and the closest matching item can be found through the semantic level not covered by the knowledge graph. For example, assuming the keyword input by the user is "permeability", first convert this keyword into a high-dimensional vector representation, then calculate the similarity between this vector and all stored vectors through the vector database, and finally return the text or concept most relevant to "permeability". In this embodiment, the relevant texts that may be returned include statements such as "The permeability of well XXX is XXX" and "The calculation method of the permeability of well XXX is as follows", which have a high semantic similarity to "permeability". In this way, the system can efficiently and accurately achieve semantic matching, thereby providing strong support for subsequent knowledge extraction and reasoning.

[0048] After successfully matching the keyword with the most relevant text in the vector database, according to the predefined rules, determine which parts of the text best meet the user's needs, and the text that meets the requirements will be encapsulated into a dictionary data structure together with the keyword. When no content related to the input keyword can be successfully matched, an empty query dictionary will be returned, indicating that no matching item meeting the conditions for this keyword is found.

[0049] In the first instance of this example, the method of using the large language model to summarize and process the extracted content and display the retrieved original text blocks and chart information to the user, as Figure 5 shown, the specific steps include: S510, Parse the feedback content. If the feedback content is valid, go to S520; if the feedback content is invalid, go to S530.

[0050] It should be noted that after receiving the dictionary structure data, it is necessary to identify and classify each data item therein to determine whether valid content is searched.

[0051] S520, The feedback content is valid. Submit it to the large model and organize the answer output.

[0052] It should be noted that after searching for valid content, extract the core content of the query according to the keyword field, and further analyze the keyword information in the matching result and the text knowledge searched. Input the matched content into the large language model, and the large language model generates an answer according to the text information content that is matched. For example, if the user queries whether a certain well is in a certain operation stage within a certain time period, relevant production time, time range, and operation company name and other information will be extracted from the dictionary and input into the large model as search knowledge at the same time. The large model generates an answer and outputs it, and at the same time displays the retrieved original text blocks and chart information to the user.

[0053] S530. If the feedback content is invalid, return a search failure message.

[0054] It should be noted that when processing the feedback content, if the knowledge documents related to a certain keyword in the dictionary structure are empty, that is, no content related to the query is retrieved. In this case, directly return a query failure message.

[0055] In several embodiments provided by the present invention, it should be understood that the disclosed method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above unit division is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0056] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for querying exploration and development data based on a knowledge graph and a large language model, characterized in that, Including: S1. Receive the exploration and development report document uploaded by the client, where the exploration and development report document carries unstructured data under each level of heading; S2. Parse and split the exploration and development report document, perform vectorization processing on the split headings at each level and the unstructured data thereunder, and store them in a vector database; S3. Construct a knowledge graph based on the headings at each level of the exploration and development report document, and store it in a graph database, where the knowledge graph database includes several nodes; S4. Receive a query statement from the client based on a large language model, and extract the keywords in the query statement; S5. According to the keywords, calculate the similarity of the phase vectors, determine the most matching node in the graph database, and extract the content under the node; S6. Use the large prediction model to synthesize the content under the extracted node to obtain the text information and / or chart information presented to the client by the user.

2. The exploration and development data query method based on a knowledge graph and a large language model according to claim 1, wherein In S2, it includes: Slice the exploration and development data report according to the headings at each level therein; Divide the text content under each heading into text blocks according to a set threshold; Perform vectorization processing on each text block and store it in the vector database.

3. The exploration and development data query method based on a knowledge graph and a large language model according to claim 1, wherein In S3, specifically include: Build a hierarchical relationship for the headings at each level in the document and store them in a MarkDown text; Import the content of the text paragraphs under each heading into the corresponding hierarchical relationship; Parse the headings and hierarchical relationships at each level of the Markdown document and generate Cypher statements; Import the headings and hierarchical relationships into the knowledge graph through Cypher statements.

4. The exploration and development data query method based on a knowledge graph and a large language model according to claim 1, wherein In S5, the specific steps include: Match the keywords with the nodes in the knowledge graph. If a matching node can be obtained, further search for the text vectors under the matching node; if no matching node can be obtained, search and match all the text vectors in the vector database.

5. The exploration and development data query method based on a knowledge graph and a large language model according to claim 3, wherein Before S6, it also includes identifying and classifying the data items in the content under the extracted node to feedback the validity of the searched content.

6. The exploration and development data query method based on a knowledge graph and a large language model according to claim 5, wherein If the searched content is valid, further submit it to the large language model for synthesis. If the searched content is not valid, return a search failure message.

Citation Information

Cited By

  • Intelligent retrieval method and device, electronic equipment and computer storage medium

    CN120873149A