Standard query method and device, equipment, medium and product

By constructing a multi-scale standard knowledge graph and a large language model to generate standard query results, the problems of insufficient mining of correlation information between standards, low degree of structuring and insufficient intelligent application are solved, and efficient and intelligent standard query and application are achieved.

CN120705331APending Publication Date: 2025-09-26SHANGHAI AIRCRAFT DESIGN & RES INST COMML AIRCRAFT OF CHINA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510968318.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies are difficult to fully meet the efficient and intelligent needs for standardization in industrial design and production processes. There is insufficient guidance on the association between standards, insufficient degree of structuring, low retrieval efficiency, and insufficient intelligent application.

Method used

By constructing multiple standard knowledge graphs at different scales (document level, paragraph level, word level), and using a large language model to generate standard query results, we can mine the implicit structure and associated semantics between standards, provide rich contextual information and prompt words, and enhance semantic guidance capabilities.

Benefits of technology

It improves the coverage and correlation accuracy of standard queries, enhances query efficiency, enables users to obtain accurate answers without having to be proficient in standard details, and improves the consistency, reliability and intelligent response level of standard applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705331A_ABST
    Figure CN120705331A_ABST
Patent Text Reader

Abstract

The invention discloses a standard query method and device, equipment, a medium and a product. The method comprises the steps that a standard query request is acquired, and query elements are extracted from the query request; obtaining knowledge retrieval results from a plurality of standard knowledge maps with different scales according to the query elements, and fusing the plurality of knowledge retrieval results to obtain context information; and according to the query element and the context information, constructing a cue word, and based on the cue word, generating a standard query result. According to the method provided by the invention, the core problems of insufficient mining of associated information of a standard system, weak term structure modeling, low semantic retrieval capability, poor standard service response intelligence and the like are effectively solved, and the standardization efficiency, the standard organization capability, the standard service capability and the intelligent application level are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of standard digitization technology, and in particular to a standard query method, device, equipment, medium and product. Background Art

[0002] A standard is a document developed and approved by a recognized organization that specifies rules, guidelines or characteristic values ​​for an activity or its results for common and repeated use by relevant parties to achieve the optimal order and effectiveness of the expected results in a specific field.

[0003] Currently, the vast majority of standards and specifications are still constructed and stored in document form. Their dissemination has gradually evolved from traditional paper documents to electronic formats such as PDF and Word. This shift has, to a certain extent, improved the efficiency of standards transmission, retrieval, access, and utilization. However, in the era of information and knowledge explosion, this documented standards management approach still cannot fully meet the demand for efficient and intelligent standardization in industrial design and production processes. Summary of the Invention

[0004] This application provides a standard query method, device, equipment, medium and product, aiming to effectively solve the technical problem that the existing technology is difficult to fully meet the efficient and intelligent requirements of standardization in industrial design and production processes.

[0005] According to a first aspect of the present application, the present application provides a standard query method, the method comprising:

[0006] Obtain a standard query request and extract query elements from the query request;

[0007] According to the query elements, knowledge retrieval results are obtained from multiple standard knowledge graphs of different scales, and multiple knowledge retrieval results are fused to obtain contextual information;

[0008] Construct prompt words based on query elements and context information, and generate standard query results based on the prompt words.

[0009] In some embodiments of the present application, multiple standard knowledge graphs of different scales include a document-level knowledge graph, a paragraph-level knowledge graph, and a word-level knowledge graph.

[0010] In some embodiments of the present application, knowledge retrieval results are obtained from multiple standard knowledge graphs of different scales according to query elements, and multiple knowledge retrieval results are fused to obtain context information, including:

[0011] Obtaining a first criterion corresponding to the query element and a second criterion associated with the first criterion from the document-level knowledge graph, wherein the first criterion and the second criterion are included in a criterion cluster;

[0012] Obtain corresponding paragraph-level knowledge nodes and first phrase-level knowledge nodes from the paragraph-level knowledge graph and the phrase-level knowledge graph according to the standard cluster, and obtain corresponding second phrase-level knowledge nodes based on the paragraph-level knowledge nodes;

[0013] The first word-sentence level knowledge node, the second word-sentence level knowledge node, the paragraph level knowledge node and the entities corresponding to the standard cluster are fused to obtain context information.

[0014] In some embodiments of the present application, obtaining a first criterion corresponding to a query element and a second criterion associated with the first criterion from a document-level knowledge graph includes:

[0015] Convert query features into query vectors, and convert document-level knowledge graphs into document-level vectors;

[0016] According to the similarity between the query vector and the document-level vector, a target document-level vector is determined, and a first criterion corresponding to the target document-level vector and a second criterion associated with the first criterion are obtained.

[0017] In some embodiments of the present application, obtaining corresponding paragraph-level knowledge nodes and first phrase-level knowledge nodes from the paragraph-level knowledge graph and the phrase-level knowledge graph respectively according to the standard cluster, and obtaining corresponding second phrase-level knowledge nodes based on the paragraph-level knowledge nodes, includes:

[0018] Locate the target paragraph-level graph and the first target phrase-level graph from multiple paragraph-level knowledge graphs and phrase-level knowledge graphs according to the standard clusters;

[0019] Based on the standard cluster, queries are performed in the target paragraph-level graph and the first target word-sentence-level graph respectively, and paragraph-level knowledge nodes and first word-sentence-level knowledge nodes are obtained accordingly;

[0020] Locating a second target phrase-level graph from multiple phrase-level knowledge graphs according to the paragraph-level knowledge nodes;

[0021] Based on the paragraph-level knowledge nodes, a query is performed in the second target word-sentence level graph to obtain the second word-sentence level knowledge nodes.

[0022] In some embodiments of the present application, extracting query elements from a query request includes:

[0023] A large language model is used to understand the query request and generate query elements.

[0024] In some embodiments of the present application, generating standard query results based on prompt words includes:

[0025] Use a large language model to understand prompt words and query requests and generate standard query results.

[0026] In some embodiments of the present application, the document-level knowledge graph is obtained by:

[0027] Extract the standard name in the standard specification document as an entity node, and use the keywords in the standard name as attribute information of the entity node;

[0028] Mining reference relationships of standard specification documents to obtain reference relationships between entity nodes;

[0029] Perform semantic logic analysis on standard specification documents to obtain the hierarchical relationship between entity nodes;

[0030] Perform semantic similarity analysis on standard specification documents to obtain the relevant relationships between entity nodes;

[0031] Obtain document-level knowledge graph based on entity nodes, attribute information, reference relationships, hierarchical relationships, and related relationships.

[0032] In some embodiments of the present application, the paragraph-level knowledge graph is obtained by:

[0033] Extract title information at each level from the standard specification document, and construct the title information into a structure tree based on the hierarchical relationship. The smallest hierarchical unit in the structure tree is a paragraph.

[0034] Obtain paragraph-level knowledge graph based on the structure tree.

[0035] In some embodiments of the present application, the word-level knowledge graph is obtained by:

[0036] Build industry standard architecture based on standard specification documents;

[0037] Based on the industry standard architecture, triples are extracted from paragraphs in standard specification documents, and a word-level knowledge graph is obtained based on the extracted triples.

[0038] Second invention, the present application also provides a standard query device, the device comprising:

[0039] An acquisition module, used for acquiring a standard query request and extracting query elements from the query request;

[0040] The fusion module is used to obtain knowledge retrieval results from multiple standard knowledge graphs of different scales according to the query elements, and fuse the multiple knowledge retrieval results to obtain contextual information;

[0041] The feedback module is used to construct prompt words according to query elements and context information, and generate standard query results based on the prompt words.

[0042] In a third aspect, the present application also provides an electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned standard query method through the computer program.

[0043] In a fourth aspect, the present application further provides a storage medium storing a plurality of instructions suitable for being loaded by a processor to execute the above-mentioned standard query method.

[0044] In a fifth aspect, the present application also provides a computer program product, including a computer program, which implements the above-mentioned standard query method when executed by a processor.

[0045] Through one or more of the above embodiments in this application, at least the following technical effects can be achieved: in the technical solution disclosed in this application, knowledge retrieval results are obtained from multiple standard knowledge graphs of different scales according to the query elements, and multiple knowledge retrieval results are fused to obtain context information, which can mine the implicit structure and associated semantics between standards, provide rich, upstream and downstream related knowledge retrieval results, improve the semantic guidance ability between standards, enhance the query coverage and association accuracy, solve the problem of isolated standard information and lack of guidance in traditional retrieval, and improve retrieval efficiency. In addition, prompt words are constructed according to the query elements and context information, and standard query results are generated based on the prompt words, which realizes query target reconstruction and result organization guidance, so that users can obtain accurate answers without being proficient in the details of the standards, improves the consistency, reliability and response intelligence level of standard application, and is suitable for various scenarios such as industrial design and compliance audit. In summary, the standard query method provided by the embodiment of the present application effectively solves the core problems of insufficient mining of associated information of the standard system, weak clause structure modeling, low semantic retrieval ability and poor intelligence of standard service response, significantly improving the efficiency of standardization, as well as the organizational ability, service ability and intelligent application level of the standard. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The following detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings will make the technical solutions and other beneficial effects of the present application apparent.

[0047] Figure 1 One of the flow charts of a standard query method provided in an embodiment of the present application;

[0048] Figure 2 Schematic diagram of multiple standard knowledge graphs of different scales provided in an embodiment of the present application;

[0049] Figure 3 This is the second flow chart of the standard query method provided in the embodiment of the present application;

[0050] Figure 4A schematic diagram of document-level knowledge retrieval provided in an embodiment of the present application;

[0051] Figure 5 Schematic diagram of paragraph-level knowledge retrieval provided in an embodiment of the present application;

[0052] Figure 6 A schematic diagram of word-level knowledge retrieval provided in an embodiment of the present application;

[0053] Figure 7 A schematic diagram of the fusion of knowledge retrieval results provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] Existing standard digitalization and intelligent application technologies have the following shortcomings:

[0055] 1) Insufficient guidance on inter-standard linkages: Standards systems are systematic and interconnected, often with complex relationships between standards, including complementarity, references, and overall hierarchical structures. However, this potential linkage has not been effectively explored and utilized, resulting in low coverage and accuracy during standard search and retrieval by businesses, thus hindering the effectiveness of digital and intelligent application of standards.

[0056] 2) Insufficient Standard Structure: Standard documents are typically compiled based on long-term practical experience and are highly standardized and repeatable. Their content should possess a clear framework and logical structure. However, existing technologies have not yet fully explored the structure of standard texts and expressed them in a document-level structured manner. This results in a loss of logical relationships between items, hindering the organization, transmission, and efficient use of standard information.

[0057] 3) Inefficient Standard Retrieval: Traditional document retrieval is too coarse-grained to meet the demand for efficient and accurate information acquisition. Although some methods attempt to construct standard documents into vector knowledge bases or triple-based knowledge graphs, the former struggles to leverage the implicit structural information within the standards, while the latter suffers from a loss of semantic expression. Consequently, neither approach has significantly improved knowledge retrieval efficiency and quality.

[0058] 4) Insufficient intelligent application of standards: Standards play a key role in ensuring consistency, reliability, and product quality in industrial design. However, the acquisition and application of standards currently rely primarily on manual retrieval, reading, and understanding by designers. Faced with a vast and constantly updated list of standard entries, designers struggle to stay current, making it difficult to ensure consistency and standards compliance during the design process. More intelligent tools are urgently needed to provide designers with fast, accurate, and efficient standards support services.

[0059] In summary, there is still much room for improvement in the digital and intelligent application of existing standards, and there is an urgent need for systematic breakthroughs in standard structure expression, semantic association mining, knowledge retrieval optimization, and intelligent application services.

[0060] In view of this, the standard query method, device, equipment, medium and product provided in the embodiments of the present application, wherein the standard query method provided in the embodiments of the present application obtains knowledge retrieval results from multiple standard knowledge graphs of different scales according to the query elements, and fuses the multiple knowledge retrieval results to obtain contextual information. It can mine the implicit structure and associated semantics between standards, provide rich, upstream and downstream related knowledge retrieval results, improve the semantic guidance ability between standards, enhance the query coverage and association accuracy, solve the problem of isolated standard information and lack of guidance in traditional retrieval, and improve retrieval efficiency. In addition, prompt words are constructed according to the query elements and contextual information, and standard query results are generated based on the prompt words. It realizes the reconstruction of query targets and the organization and guidance of results, so that users can obtain accurate answers without being proficient in the details of the standards, improve the consistency, reliability and intelligent response level of standard applications, and is suitable for various scenarios such as industrial design and compliance audits. In summary, the standard query method provided in the embodiment of the present application effectively solves core problems such as insufficient mining of related information in the standard system, weak clause structure modeling, poor semantic retrieval capabilities, and poor intelligence of standard service responses, significantly improves the efficiency of standardization, as well as the organizational capabilities, service capabilities and intelligent application levels of standards, and at least solves some of the above problems.

[0061] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making any creative work are within the scope of protection of this application.

[0062] In the description of this application, it should be noted that, unless otherwise specified or limited, the term "and / or" herein is merely a description of an association relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the character " / " herein, unless otherwise specified, generally indicates that the associated objects are in an "or" relationship.

[0063] The standard query method, device, equipment, medium and product provided by this application are introduced below with reference to the accompanying drawings.

[0064] Figure 1 The following is a flowchart of the steps of the standard query method provided in the embodiment of the present application. Figure 1 As shown, the standard query method includes the following steps:

[0065] S101: Obtain a standard query request and extract query elements from the query request.

[0066] Among them, the query request is a question or task instruction raised by the user in natural language, which is used to obtain information related to a certain technical content, standard clause, applicable conditions or version change, such as "What standard stipulates the fire protection requirements for energy storage batteries?", "In the aircraft design process, what test standards does the XX component need to meet?", etc.

[0067] In this step, the query request can be processed by a natural language processing model using at least one of a variety of processing methods such as semantic analysis, intent recognition, entity extraction, and condition extraction to obtain query elements, which include at least one of a standard name, technical keywords, query intent, limiting conditions, target objects, etc.

[0068] S102: Obtain knowledge retrieval results from multiple standard knowledge graphs of different scales according to the query elements, and fuse the multiple knowledge retrieval results to obtain context information.

[0069] Among them, standard knowledge graphs of different scales refer to semantic modeling of standard documents according to the knowledge granularity level, and constructing knowledge graphs that reflect information at different structural levels. The knowledge granularity level can be standard system level, document level, chapter level, paragraph level, sentence level, word level, etc., and the corresponding standard knowledge graphs can be standard system level knowledge graphs, document level knowledge graphs, chapter level knowledge graphs, paragraph level knowledge graphs, sentence level knowledge graphs, word level knowledge graphs, etc. The multiple standard knowledge graphs in the embodiments of the present application can be at least two of the standard system level knowledge graphs, document level knowledge graphs, chapter level knowledge graphs, paragraph level knowledge graphs, sentence level knowledge graphs, word level knowledge graphs, etc.

[0070] In this step, knowledge retrieval results corresponding to the query elements are obtained from standard knowledge graphs of different scales, and the knowledge retrieval results of each knowledge granularity are fused to obtain the context information of the query elements.

[0071] It is understandable that multiple standard knowledge graphs of different scales are constructed in advance based on standard specification documents.

[0072] S103: construct prompt words according to the query elements and context information, and generate standard query results based on the prompt words.

[0073] In this step, query elements and context information are combined using structured templates, language concatenation, or rule-based construction to form input prompts suitable for the Large Language Model (LLM). The prompts are then fed into the Large Language Model (LLM). LLM leverages its language understanding and generation capabilities to generate standard query results based on the prompts and context information.

[0074] Prompt words can include restatements of user questions, contextual guidance such as standard numbers, names, and clause numbers, keyword extraction content, and structured task instructions. Standard query results can include standard documents related to the query request and their meta-information, the original text, numbers, and sources of matching clauses, natural language summaries of standards or clauses, differences and applicability descriptions between multiple standards, and structured output (such as tables, triples, and tag sets).

[0075] The standard query method provided in the embodiment of the present application obtains knowledge retrieval results from a plurality of standard knowledge graphs of different scales according to the query elements, and fuses the plurality of knowledge retrieval results to obtain contextual information. It can mine the implicit structure and associated semantics between standards, provide rich, upstream and downstream related knowledge retrieval results, improve the semantic guidance ability between standards, enhance the query coverage and association accuracy, solve the problem of isolated standard information and lack of guidance in traditional retrieval, and improve retrieval efficiency. In addition, prompt words are constructed according to the query elements and contextual information, and standard query results are generated based on the prompt words. It realizes the reconstruction of query targets and the organization and guidance of results, so that users can obtain accurate answers without being proficient in the details of the standards, improves the consistency, reliability and intelligent response level of standard applications, and is suitable for various scenarios such as industrial design and compliance audits. In summary, the standard query method provided in the embodiment of the present application effectively solves the core problems such as insufficient mining of associated information of the standard system, weak clause structure modeling, low semantic retrieval ability and poor intelligent response of standard services, and significantly improves the efficiency of standardization, as well as the organizational ability, service ability and intelligent application level of the standards.

[0076] In some embodiments of the present application, Figure 2 As shown, multiple standard knowledge graphs of different scales include document-level knowledge graphs, paragraph-level knowledge graphs, and word-level knowledge graphs.

[0077] Among them, the document-level knowledge graph reflects the macro-relationships between standards in standard specification documents, such as reference relationships, version levels, and applicable fields; the paragraph-level knowledge graph reflects the structural hierarchy and thematic organization within the standard, such as the table of contents, clause numbers, and structure tree; the phrase-level knowledge graph reflects the entity-relationship-entity semantic details in the standard clauses, such as parameter restrictions, behavioral constraints, and logical conditions. The knowledge retrieval results corresponding to the query elements obtained by searching the document-level, paragraph-level, and phrase-level standard knowledge graphs not only cover the full-level information structure from inter-standard references to clause details, but also provide retrieval responses at multiple knowledge granularities, enhance contextual understanding, and improve structured output capabilities and interpretability, significantly improving the accuracy, intelligence, and practicality of the standard semantic service system.

[0078] In some embodiments of the present application, Figure 3 As shown in the figure, knowledge retrieval results are obtained from multiple standard knowledge graphs of different scales according to the query elements, and multiple knowledge retrieval results are fused to obtain contextual information, including:

[0079] The first standard corresponding to the query element and the second standard associated with the first standard are obtained from the document-level knowledge graph, and the first standard and the second standard are included in the standard cluster. That is, based on the extracted query element, a semantic search is performed in the document-level knowledge graph to obtain the first standard that is semantically related to the query element. Furthermore, based on the reference relationship, hierarchical relationship or semantic similarity relationship in the document-level graph, the second standard that has a reference or semantic association with the first standard is obtained, such as Figure 4 As shown, the first criterion and the second criterion are both document-level knowledge retrieval results, and the criterion cluster includes the first criterion and the second criterion.

[0080] According to the standard cluster, the corresponding paragraph-level knowledge node and the first sentence-level knowledge node are obtained from the paragraph-level knowledge graph and the sentence-level knowledge graph respectively, and the corresponding second sentence-level knowledge node is obtained based on the paragraph-level knowledge node. That is, for each standard included in the standard cluster, the corresponding paragraph-level knowledge node is searched in the paragraph-level knowledge graph respectively, and the node may include a clause structure node, a title level node or a subject word node, etc. At the same time, in the sentence-level knowledge graph, semantic matching is performed based on the query elements to obtain the first sentence-level knowledge node associated with its semantics. The node is usually an entity-relationship-entity triple extracted from the clause. Furthermore, based on the above paragraph-level knowledge node, the second sentence-level knowledge node existing in its internal or subordinate structure can be obtained to supplement the structural integrity and semantic coherence in the context. Among them, the paragraph-level knowledge node is the paragraph-level knowledge retrieval result (such as Figure 5 As shown), the first sentence-level knowledge node and the second sentence-level knowledge node are both sentence-level knowledge retrieval results (as shown Figure 6shown).

[0081] like Figure 7 As shown in the figure, the first phrase-level knowledge node, the second phrase-level knowledge node, the paragraph-level knowledge node, and the entities corresponding to the standard cluster are fused to obtain context information. In other words, the first phrase-level knowledge node, the second phrase-level knowledge node, the paragraph-level knowledge node, and the entity nodes corresponding to the standard cluster are fused to form a context information set with characteristics of upper and lower semantics, structural logic, and entity association.

[0082] The query method provided in the embodiment of the present application obtains a first standard by matching query elements in the document-level knowledge graph, and then automatically expands along the reference relationship, hierarchical relationship and semantic similarity relationship to obtain a second standard that is supplementary or reference-related to the first standard. The two together constitute a standard cluster, so that the document-level retrieval results are expanded from a single document to a collection of related documents, which significantly improves the query coverage and association accuracy. Secondly, within the scope of the standard cluster, the corresponding paragraph-level knowledge nodes such as chapter nodes, clause nodes or keyword nodes are extracted from the paragraph-level knowledge graph, and the first phrase-level knowledge node within the clause is extracted from the phrase-level knowledge graph; at the same time, the subordinate phrase nodes are further supplemented by the paragraph-level knowledge nodes to realize the hierarchical mapping of chapter-paragraph-triplets. In this way, the traditional plain text clauses are converted into a computable multi-layer structure, the logical dependencies and upper and lower semantics between the clauses are retained, and the deep structured expression capability of the standard content is significantly improved. Finally, knowledge retrieval results at different levels—document-level entities, paragraph-level nodes, and word-level triples—are fused into a contextual information set. Based on this set, prompt words are constructed to guide the large language model in generating multi-modal standard query results, including the original text of clauses, summaries, difference comparisons, or structured tables. Users can obtain accurate, explainable, and directly referenced answers without having to read each clause individually, significantly improving the consistency, reliability, and intelligence of standards in design, verification, and compliance review scenarios.

[0083] In some embodiments of the present application, obtaining a first criterion corresponding to a query element and a second criterion associated with the first criterion from a document-level knowledge graph includes:

[0084] Convert query features into query vectors, and convert document-level knowledge graphs into document-level vectors.

[0085] According to the similarity between the query vector and the document-level vector, a target document-level vector is determined, and a first criterion corresponding to the target document-level vector and a second criterion associated with the first criterion are obtained.

[0086] In some examples, the query element q is converted into a query vector v(q) through a text embedding algorithm (such as BERT, Word2Vec, Sentence-BERT, etc.). At the same time, the name, subject keywords, and summary text of each standard entity node in the document-level knowledge graph are converted into a document-level vector set v(F) through the aforementioned text embedding algorithm. Then, the similarity between the query vector v(q) and the document-level vector v(F) is calculated, and the target document-level vector corresponding to the query vector is determined based on the similarity. Then, the entity f corresponding to the target document-level vector is obtained. 1,2,3,...,n , and entity f 1,2,3,...,n The corresponding first standard F 1,2,3,...,n , and with F 1,2,3,...,n The associated second criterion F′ 1,2,3,...,n , the first standard and the second standard together form a standard cluster {F 1,2,3,...,n , F′ 1,2,3,...,n}.

[0087] Among them, the similarity between the query vector v(q) and the document-level vector v(F) can be calculated using cosine similarity or Euclidean distance. The target document-level vector can be selected from the document-level vector v(F) according to a preset similarity threshold. In addition, according to the first criterion F 1,2,3,...,n In the document-level knowledge graph, we traverse the graph according to the predefined reference relationship edges, hierarchical relationship edges, semantic similarity relationship edges, etc., and filter out the edges that match the first criterion F. 1,2,3,...,n There are criteria of direct reference, subordination, complementation or high semantic relevance as secondary criteria.

[0088] The query method provided in the embodiment of the present application determines the target document-level vector based on the similarity between the query vector and the document-level vector, and obtains a first criterion corresponding to the target document-level vector and a second criterion associated with the first criterion. It not only achieves precise matching of the query vector and the document vector, but also realizes automatic expansion of document-level retrieval results and realizes in-depth mining of standard associations.

[0089] In some embodiments of the present application, obtaining corresponding paragraph-level knowledge nodes and first phrase-level knowledge nodes from the paragraph-level knowledge graph and the phrase-level knowledge graph respectively according to the standard cluster, and obtaining corresponding second phrase-level knowledge nodes based on the paragraph-level knowledge nodes, includes:

[0090] According to the standard cluster, the target paragraph-level graph and the first target sentence-level graph are located from multiple paragraph-level knowledge graphs and sentence-level knowledge graphs respectively.

[0091] Based on the standard cluster, queries are performed in the target paragraph-level graph and the first target word-sentence-level graph respectively, and paragraph-level knowledge nodes and first word-sentence-level knowledge nodes are obtained accordingly.

[0092] A second target phrase-level graph is located from multiple phrase-level knowledge graphs according to paragraph-level knowledge nodes.

[0093] Based on the paragraph-level knowledge nodes, a query is performed in the second target word-sentence level graph to obtain the second word-sentence level knowledge nodes.

[0094] In some examples, by querying the graph database with standard clusters {F 1,2,3,...,n , F′ 1,2,3,...,n} corresponding to the target paragraph-level graph P and the first target phrase-level graph W. And in the target paragraph-level graph P and the first target phrase-level graph W, the cluster {F 1,2,3,...,n , F′ 1,2,3,...,n} corresponding paragraph-level knowledge node p 1,2,3,...,n and the first word-level knowledge node w 1,2,3,...,n Then, locate the second target phrase-level graph in the phrase-level knowledge graph based on the paragraph-level knowledge node. Then, the second word-level knowledge node w′ is obtained through fuzzy retrieval 1,2,3,...,n Among them, fuzzy search can be regular expression, containment matching, prefix / suffix matching, etc.

[0095] The first word-level knowledge node w 1,2,3,...,n , the second sentence level knowledge node w′ 1,2,3,...,n , paragraph-level knowledge node p 1,2,3,...,n And the standard cluster {F 1,2,3,...,n , F′ 1,2,3,...,n} corresponding entity {f 1,2,3,...,n , f′ 1,2,3,...,n} to fuse and obtain context information K = {w 1,2,3,...,n ,w′ 1,2,3,...,n ,p 1,2,3,...,n ,f 1,2,3,...,n , f′ 1,2,3,...,n}.

[0096] The query method provided in the embodiment of the present application obtains paragraph-level knowledge nodes and first sentence-level knowledge nodes from paragraph-level knowledge graphs and sentence-level knowledge graphs according to standard clusters, and obtains second sentence-level knowledge nodes based on the paragraph-level knowledge nodes, thereby realizing complete knowledge retrieval from macro-standard documents to meso-paragraph structures and then to micro-semantic details, significantly enhancing the contextual semantic integrity of subsequent prompt word construction and the accuracy of large language model responses.

[0097] In some embodiments of the present application, extracting query elements from a query request includes:

[0098] The large language model is used to understand the query request to generate query elements. In other words, the large language model is used to understand the query Q (including the question or task raised by the user) and generate query elements q related to Q (including the subject or entity object).

[0099] In some embodiments of the present application, generating standard query results based on prompt words includes:

[0100] Use a large language model to understand prompt words and query requests and generate standard query results.

[0101] In some examples, a large language model is used to perform semantic understanding and information generation processing on the constructed prompt words and the original query request, and output standard query results in structured or natural language form.

[0102] Specifically, the previously fused context information set (including document-level standard entities, paragraph-level structural nodes, word-level triples, etc.) is jointly constructed with the semantic elements in the query request to generate a prompt word template (Prompt) that is suitable for the language model input. Subsequently, the prompt word is used as input, together with the original natural language query request, to input the locally deployed or called LLM engine, triggering its language understanding and generation mechanism under context constraints. During the reasoning process, the language model, based on its extensive semantic reasoning ability and context modeling ability obtained through training, comprehensively analyzes the standard content, clause logic, semantic relations and user question intentions contained in the prompt word, performs task-driven language generation operations, and outputs standard query results that meet the query request requirements. The standard query results can be clause matching results, natural language interpretations, standard comparison results, information summaries presented in formats such as tables, lists, tag sets, triples or JSON, reference relationship paths, etc.

[0103] The query method provided in the embodiment of the present application is driven by the prompt words and the original query together, realizing a large-model semantic enhancement response mechanism supported by rich context. It can not only provide standard answers with semantic coherence, content focus, and clear structure, but also has strong reasoning and interpretation capabilities and context adaptability, effectively improving the intelligence level and practical usability of standard queries.

[0104] In some embodiments of the present application, the document-level knowledge graph is obtained by:

[0105] The standard names in standard specification documents are extracted as entity nodes, and the keywords in the standard names are used as attribute information for the entity nodes. For example, taking the structural design standards in aircraft design as an example, thousands of standard documents specify strict design processes and requirements for aircraft design. We obtain standard specification PDF documents for parsing, use a parsing tool for character content, and use the Python language pymupdf tool to read the document content and save it as plain text markup language Markdown for storage.

[0106] Mining the reference relationships of standard specification documents to obtain the reference relationships between entity nodes.

[0107] Perform semantic logical analysis on standard specification documents to obtain the hierarchical relationship between entity nodes.

[0108] Perform semantic similarity analysis on standard specification documents to obtain the relevant relationships between entity nodes.

[0109] Obtain document-level knowledge graph based on entity nodes, attribute information, reference relationships, hierarchical relationships, and related relationships.

[0110] In some examples, the standard name is used as the entity node, and the name keywords are extracted as its attributes; reference relationship mining is used to obtain the reference relationship between standards; intelligent semantic logic analysis is used to obtain the hierarchical relationship between standards; semantic similarity analysis is used to obtain the correlation relationship between standards, and a knowledge graph at the standard document level is constructed based on the aforementioned entity nodes, attributes, reference relationships, hierarchical relationships, correlation relationships, etc.

[0111] In some embodiments of the present application, the paragraph-level knowledge graph is obtained by:

[0112] The title information of each level is extracted from the standard specification document, and the title information is constructed into a structure tree based on the hierarchical relationship. The smallest hierarchical unit in the structure tree is a paragraph.

[0113] Obtain paragraph-level knowledge graph based on the structure tree.

[0114] In some examples, through the title regular matching method, the standard text content obtained through parsing is used to obtain standard first-level, second-level, third-level and other titles, and establish a standard structure tree; the smallest level unit is the paragraph, and through the keyword extraction model, the paragraph keyword is obtained as the end node of the paragraph-level graph.

[0115] In some embodiments of the present application, the word-level knowledge graph is obtained by:

[0116] Build industry-standard architectures based on standard specification documents.

[0117] Based on the industry standard architecture, we extract triples from paragraphs in standard specification documents and generate a word-level knowledge graph based on the extracted triples. In other words, we build an industry standard schema and use the knowledge extraction model to extract the entity-relationship triples of each paragraph to construct a word-level graph.

[0118] The present application also provides a standard query device, comprising:

[0119] An acquisition module, used for acquiring a standard query request and extracting query elements from the query request;

[0120] The fusion module is used to obtain knowledge retrieval results from multiple standard knowledge graphs of different scales according to the query elements, and fuse the multiple knowledge retrieval results to obtain contextual information;

[0121] The feedback module is used to construct prompt words according to query elements and context information, and generate standard query results based on the prompt words.

[0122] The standard query device provided in this embodiment corresponds to the standard query method provided in any of the above embodiments, and will not be described in detail here.

[0123] Based on any of the above embodiments, another embodiment of the present application further provides an electronic device, which may include: a processor, a communications interface, a memory, and a communication bus, wherein the processor, the communications interface, and the memory communicate with each other via the communication bus. The processor may call logic instructions in the memory to execute the above standard query method.

[0124] In addition, the logical instructions in the above-mentioned memory can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0125] In addition, the logical instructions in the above-mentioned memory can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0126] On the other hand, an embodiment of the present application further provides a storage medium on which a plurality of instructions are stored, and the instructions are suitable for being loaded by a processor to execute the standard query method provided in the above embodiments to execute the above standard query method.

[0127] On the other hand, an embodiment of the present application further provides a computer program product, including a computer program, which implements the above-mentioned standard query method when executed by a processor.

[0128] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. That is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or certain parts of the embodiment.

[0130] In summary, although the present application has been disclosed as above with preferred embodiments, the above preferred embodiments are not intended to limit the present application. Ordinary technicians in this field can make various changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims.

Claims

1. A standard query method, characterized in that: The method comprises: Obtaining a standard query request and extracting query elements from the query request; Obtaining knowledge retrieval results from a plurality of standard knowledge graphs of different scales according to the query elements, and fusing the plurality of knowledge retrieval results to obtain context information; A prompt word is constructed according to the query elements and the context information, and a standard query result is generated based on the prompt word.

2. The method according to claim 1, wherein The multiple standard knowledge graphs of different scales include a document-level knowledge graph, a paragraph-level knowledge graph, and a word-level knowledge graph.

3. The method according to claim 1 or 2, wherein: The step of obtaining knowledge retrieval results from a plurality of standard knowledge graphs of different scales according to the query elements and fusing the plurality of knowledge retrieval results to obtain context information includes: Acquire a first criterion corresponding to the query element and a second criterion associated with the first criterion from a document-level knowledge graph, wherein the first criterion and the second criterion are included in a criterion cluster; Obtaining corresponding paragraph-level knowledge nodes and first phrase-level knowledge nodes from the paragraph-level knowledge graph and the phrase-level knowledge graph according to the standard cluster, and obtaining corresponding second phrase-level knowledge nodes based on the paragraph-level knowledge nodes; The first phrase-level knowledge node, the second phrase-level knowledge node, the paragraph-level knowledge node, and entities corresponding to the standard cluster are fused to obtain the context information.

4. The method according to claim 3, wherein The acquiring, from the document-level knowledge graph, a first criterion corresponding to the query element and a second criterion associated with the first criterion includes: Converting the query elements into query vectors, and converting the document-level knowledge graph into document-level vectors; A target document-level vector is determined according to the similarity between the query vector and the document-level vector, and a first criterion corresponding to the target document-level vector and a second criterion associated with the first criterion are obtained.

5. The method according to claim 3, wherein The step of obtaining corresponding paragraph-level knowledge nodes and first phrase-level knowledge nodes from the paragraph-level knowledge graph and the phrase-level knowledge graph according to the standard cluster, and obtaining corresponding second phrase-level knowledge nodes based on the paragraph-level knowledge nodes, includes: Locating a target paragraph-level graph and a first target phrase-level graph from the plurality of paragraph-level knowledge graphs and the phrase-level knowledge graphs according to the standard cluster; Based on the standard cluster, query the target paragraph-level graph and the first target word-sentence-level graph respectively to obtain paragraph-level knowledge nodes and first word-sentence-level knowledge nodes respectively; Locating a second target phrase-level graph from the plurality of phrase-level knowledge graphs according to the paragraph-level knowledge node; A query is performed in the second target phrase-level graph based on the paragraph-level knowledge node to obtain a second phrase-level knowledge node.

6. The method according to claim 1, wherein The extracting query elements from the query request includes: The query request is understood using a large language model to generate the query elements.

7. The method according to claim 1, wherein Generating a standard query result based on the prompt word includes: The large language model is used to understand the prompt word and the query request to generate the standard query result.

8. The method according to claim 2, wherein The document-level knowledge graph is obtained in the following way: Extracting the standard name in the standard specification document as an entity node, and using the keywords in the standard name as attribute information of the entity node; Mining the reference relationship of the standard specification document to obtain the reference relationship between the entity nodes; Performing semantic logic analysis on the standard specification document to obtain the hierarchical relationship between the entity nodes; Performing semantic similarity analysis on the standard specification document to obtain the correlation between the entity nodes; The document-level knowledge graph is obtained based on the entity nodes, the attribute information, the reference relationships, the hierarchical relationships, and the correlation relationships.

9. The method according to claim 2, wherein The paragraph-level knowledge graph is obtained in the following way: Extracting title information of each level from the standard specification document, and constructing the title information into a structure tree based on the hierarchical relationship, wherein the smallest hierarchical unit in the structure tree is a paragraph; The paragraph-level knowledge graph is obtained based on the structure tree.

10. The method according to claim 2, wherein The word-level knowledge graph is obtained in the following way: Build industry standard architecture based on standard specification documents; Based on the industry standard architecture, triples are extracted from the paragraphs in the standard specification document, and a word-level knowledge graph is obtained based on the extracted triples.

11. A standard query device, characterized in that: The device comprises: An acquisition module, configured to acquire a standard query request and extract query elements from the query request; A fusion module is used to obtain knowledge retrieval results from multiple standard knowledge graphs of different scales according to the query elements, and to fuse the multiple knowledge retrieval results to obtain context information; The feedback module is configured to construct prompt words according to the query elements and the context information, and generate standard query results based on the prompt words.

12. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to execute the standard query method according to any one of claims 1 to 10 through the computer program.

13. A storage medium, characterized in that: The storage medium stores a plurality of instructions, which are suitable for being loaded by a processor to execute the standard query method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the standard query method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Standard knowledge graph construction method and device, and standard query method and device

    CN112732945A

  • Large model knowledge base retrieval method based on multi-granularity retrieval

    CN119646201A

  • Large model knowledge base construction method based on multi-granularity retrieval

    CN119647581A

  • Method and system for extracting contextual information from a knowledge base

    US20210182709A1

  • Method and apparatus for querying knowledge map of legal cases, device and storage medium

    WO2021164226A1