Standard analysis atlas construction method and system using large-scale language model
By using large language models to build standard analysis maps, the complexity of standard documents and multi-layer hierarchical structure processing problems in the existing technology are solved, and efficient and accurate standard analysis and management are achieved.
Patent Information
- Application Number
- CN202510114691.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
AI Technical Summary
It is difficult for the prior art to effectively build standard analysis maps, especially when dealing with the complexity and multi-layered hierarchical structure of standard documents, traditional methods have labor-intensive and error-prone problems.
The standard analysis map is constructed using large language model (LLM), and the standard attribute structure units are extracted through a predefined standard document framework, entities are extracted and knowledge maps of a single standard are constructed, and the standard analysis map is then generated through multi-level alignment.
It significantly improves the accessibility and analysis efficiency of standard files, reduces the possibility of human error, and provides a scalable, accurate and efficient standard management system.
Smart Images

Figure CN120012891A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power systems, and more specifically, to a method and system for constructing a standard analysis graph using a large language model. Background Art
[0002] In an era characterized by rapid technological growth and digital transformation, the need for efficient and structured knowledge representation has become even more critical. Standards are a fundamental cornerstone that provides guidelines, specifications, and frameworks to ensure the quality and interoperability of products, services, and systems. Despite this, the complexity and extensiveness of standards documents pose significant challenges in terms of extraction, alignment, and organization. Traditional manual processing methods often prove to be labor-intensive and error-prone, hindering the capture of the intricate relationships and hierarchies in these standards.
[0003] Knowledge graphs (KGs) provide a powerful way to organize information, facilitating improved advanced search, reasoning, and analytical capabilities. Despite their potential, building KGs from standard documents remains challenging due to unstructured text, domain-specific terminology, and the complexity of accurate information extraction. Recent advances in natural language processing (NLP), especially the emergence of large language models (LLMs) like GPT-3, have opened new avenues for automatically extracting and structuring information from unstructured content. These models have demonstrated remarkable proficiency in understanding and generating human-like text, positioning them as viable solutions to the complexities associated with standard documents.
[0004] Knowledge graphs have emerged as a powerful tool for representing and organizing structured information in a semantic manner. These graph-based representations facilitate efficient data retrieval, analysis, and reasoning in various fields, ranging from search engines and recommender systems to scientific research and healthcare. By encapsulating rich relationships and dependencies between entities, knowledge graphs enable a deeper understanding and exploitation of complex information landscapes. In recent years, the construction and utilization of knowledge graphs have received tremendous attention due to their inherent capabilities. One of the main advantages of knowledge graphs is their ability to represent complex interconnections between entities, thereby providing a comprehensive and intuitive view of the data. This capability is particularly useful in scenarios involving complex and multifaceted datasets, which are not met by traditional data representation methods.
[0005] Among the various applications of knowledge graphs, the field of standardization is crucial. Standardization is an important aspect of multiple industries, including technology, manufacturing, healthcare, and finance. It involves the development, dissemination, and implementation of standards that serve as guidelines, benchmarks, or specifications for products, services, processes, and systems. Structured and systematic organization of this vast and diverse information is essential to maintain quality, ensure interoperability, and promote innovation. Knowledge graphs are uniquely positioned to address the challenges inherent in standardization work. By representing standards, regulations, and related entities as nodes and their interrelationships as edges in a graph structure, complex dependencies and hierarchical arrangements can be effectively captured and navigated. This not only facilitates efficient retrieval and analysis of standards-related information, but also supports the decision-making process by providing a holistic view of the standards environment.
[0006] Standards knowledge graphs, in particular, present unique characteristics and challenges that require specialized applications. Standards themselves contain detailed technical descriptions, legal terminology, and structured metadata such as claims, classifications, and citations. These layers of information require advanced techniques to accurately extract and represent in knowledge graphs. Furthermore, standards involve a multi-layered hierarchical structure where higher-level concepts and broader classifications give way to specific embodiments and detailed claims. Effectively capturing these hierarchical dependencies in a knowledge graph is critical for meaningful analysis and application. Tailored methods for the specific nature of standards are essential to building knowledge graphs that accurately reflect the complex relationships in standards data.
[0007] The advent of Large Language Models (LLMs) has revolutionized the field of Natural Language Processing (NLP), profoundly impacting the way text data is processed and understood. Trained on a wide corpus of text data, LLMs have demonstrated superior capabilities in tasks such as information extraction, entity recognition, and semantic interpretation. These capabilities are essential for converting unstructured text into a structured data format suitable for knowledge graph construction. Leveraging the power of LLMs to extract standards-related information and subsequently build knowledge graphs has great potential in improving the efficiency and accuracy of the standards analysis process. LLMs can handle the subtle and complex language commonly found in standards documents, facilitating the precise extraction of relevant entities and their interrelationships. This automation not only simplifies the creation of knowledge graphs, but also ensures consistency and reduces the possibility of human error.
[0008] The ability of knowledge graphs to represent intricate relationships and hierarchies is perfectly aligned with the requirements of standards analysis. By organizing standards as nodes and their interconnections as edges in a graph, knowledge graphs can succinctly capture the multidimensional nature of standards-related information. This process transforms a large number of independent standards documents into a coherent, navigable structure. However, current approaches to building and leveraging knowledge graphs in the context of standards analysis often lack a unified framework. This missing systematic approach hinders the ability to fully exploit the potential of knowledge graphs in the standardization process.
[0009] Recent studies have examined the intersection of knowledge graphs and NLP, revealing promising avenues for enhancing the information extraction process. Research efforts have explored the use of LLMs for automatically extracting entities and relations from text data, and subsequently integrating these elements into knowledge graphs. Nevertheless, the specific requirements of standard analysis require tailored approaches that must take into account the unique nature of standards, which often involve hierarchical structures and detailed procedural descriptions.
[0010] To address the above issues, there is an urgent need for a standard analysis graph construction method and system using a large language model. Summary of the invention
[0011] To address the deficiencies in the prior art, this paper provides an automated framework called StandardKG Builder, which uses LLM to build knowledge graphs tailored for standards analysis from multiple perspectives for complex information extraction. Our evaluation on a comprehensive dataset of standards documents highlights the effectiveness and scalability of the framework. By combining complex knowledge representation with advanced NLP techniques, this work significantly improves the accessibility and analysis of standards documents, paving the way for more efficient and intelligent standards management systems.
[0012] The present invention adopts the following technical solution.
[0013] The first aspect of the present invention relates to a method for constructing a standard analysis graph using a large language model, characterized in that the method includes the following steps: extracting standard attribute structure units from a standard document using a predefined standard document framework, extracting entities using the standard attribute structure units, and constructing a single standard knowledge graph; for the constructed multiple single standard knowledge graphs, obtaining the association of the multiple single standard knowledge graphs through multi-level alignment, and generating a standard analysis graph.
[0014] Preferably, a predefined standard document framework is used to extract standard attribute structure units from a standard document, including: a predefined standard document framework, wherein the standard document framework includes at least keywords, the relationship between keyword positions and standard attribute structure units; keyword matching is used to extract a keyword outline from a corresponding position of the standard document, and corresponding structural contents under each standard attribute structure unit are extracted under the keyword outline; the standard attribute structure units include the standard Chinese name, the standard English name, the standard number, the standard level, the introduction, the scope, the normative reference documents, the terms and definitions, and the core clauses.
[0015] Preferably, entities are extracted using standard attribute structure units, and a single standard knowledge graph is constructed, including: taking the standard number of the standard document as the root node, extracting multi-level ontologies from the standard attribute structure units and the corresponding structural content; constructing a single standard knowledge graph through the association between multi-level ontologies and the association between multi-level ontologies and root nodes.
[0016] Preferably, the standard number of the standard document is used as the root node, and a multi-level ontology is extracted from the standard attribute structure unit and the corresponding structure content, including: taking the standard clauses as the structure content, extracting multiple clause names from one or more structure contents corresponding to each standard attribute structure unit; taking the clause names as the ontology, and constructing a multi-level ontology with the hierarchies between the clause names as the hierarchy.
[0017] Preferably, for the constructed knowledge graphs of multiple single standards, the association of the knowledge graphs of multiple single standards is obtained through multi-level alignment, including: collecting normative reference files in multiple standard documents, and extracting keywords from the normative reference files by keyword matching, and corresponding the keywords one by one to the root nodes in the knowledge graphs of multiple single standards generated in step 1; extracting the hierarchical relationship between all standard documents from the normative references, and associating the root nodes of the knowledge graphs of multiple single standards by using the hierarchical relationship, and using the association as a connection method between knowledge ontologies in the standard analysis diagram.
[0018] Preferably, for the constructed knowledge graphs of multiple single standards, the association of the knowledge graphs of multiple single standards is obtained through multi-level alignment, including: extracting reference information located inside each standard attribute structure unit from each standard document, wherein the reference information includes at least information position and reference position; constructing the association relationship between different standards and between different standard attribute structure units of the same standard according to the information position and the reference position, and using the association as a connection method between knowledge ontologies in the standard analysis diagram.
[0019] Preferably, for the constructed multiple single-standard knowledge graphs, the association of the multiple single-standard knowledge graphs is obtained through multi-level alignment, including: for the standard attribute structure units of the term and definition type, all terms and definitions are extracted from such structure units; for each term or definition, a BERT model is used to perform semantic analysis to extract words with high similarity to the term or definition from multiple standard documents, and extract the position of the words; the words and word positions are used as attribute information of the current term or definition, and the word position at least includes the term name and the standard attribute structure unit corresponding to the word.
[0020] Preferably, for the constructed multiple single-standard knowledge graphs, the association of the multiple single-standard knowledge graphs is obtained through multi-level alignment, including: for the lowest level ontology in the multi-level ontology, using a large language model to extract the clause content corresponding to the lowest level ontology, extracting the keywords of the lowest level clauses, and generating a core vocabulary of the lowest level ontology; combining the keyword summaries of one or more lowest level clauses corresponding to the upper level clauses to obtain the core vocabulary of the upper level ontology; for any two ontologies in the multi-level ontologies of multiple standard documents, using the similarity calculation method between the core vocabularies of the ontology to obtain alignment of any two ontologies whose similarity of all core vocabularies exceeds a preset threshold.
[0021] Preferably, for the constructed multiple single-standard knowledge graphs, the association of the multiple single-standard knowledge graphs is obtained through multi-level alignment, including: extracting high-level keywords in the introduction, overview, and other document contents from each standard document, and combining them into a topic description for each standard document; using Word2Vec to extract the vocabulary distribution from the topic description, using the LDA topic model to analyze the vocabulary distribution in the topic descriptions of multiple standard documents, and generating the probability of each standard document under different topics based on the vocabulary distribution; taking each topic in different topics as a knowledge ontology, adding the topic knowledge ontology to the standard analysis graph, and associating the root nodes of multiple standard documents to different topic knowledge nodes respectively.
[0022] Preferably, for the constructed multiple single-standard knowledge graphs, the association of the multiple single-standard knowledge graphs is obtained through multi-level alignment, and a standard analysis graph is generated, including: using depth-first search and breadth-first search methods to traverse entities and relationships to generate a standard analysis graph.
[0023] The second aspect of the present invention relates to a system for constructing a standard analysis graph using a large language model, wherein the system is implemented using a method for constructing a standard analysis graph using a large language model as described in the first aspect of the present invention; wherein the system includes a construction module and a generation module; wherein the utilization module is used to extract standard attribute structure units from a standard document using a predefined standard document framework, extract entities using the standard attribute structure units, and construct a single standard knowledge graph; and the generation module is used to obtain the association of multiple single standard knowledge graphs constructed through multi-level alignment, and generate a standard analysis graph.
[0024] A third aspect of the present invention relates to a terminal, comprising a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method described in the first aspect of the present invention.
[0025] A fourth aspect of the present invention relates to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect of the present invention.
[0026] The beneficial effect of the present invention is that, compared with the prior art, a new automated framework in the present invention is specifically used to build a knowledge graph specifically for standard analysis. The proposed framework seamlessly integrates the functions of LLM to handle the complex language of standard documents, thereby facilitating the extraction of relevant entities and their relationships. The method proposes a hierarchical framework and topology for standard knowledge representation and organization. A detailed processing method based on a large language model is proposed for complex standard information extraction, alignment and knowledge graph construction. The method evaluates the proposed framework in a representative task, multi-hop standard text search, using more than 1,000 standard documents, demonstrating the effectiveness of the proposed method.
[0027] The beneficial effects of the present invention also include:
[0028] The method proposes an innovative automated framework for building knowledge graphs tailored for standards analysis, addressing the unique challenges posed by standards text and hierarchical standards structure. This invention advances the field by combining the power of LLM with sophisticated graph construction methods to provide a scalable, accurate, and efficient solution for managing and analyzing standards. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A schematic diagram of a large language model in the present invention;
[0030] Figure 2 A schematic diagram of hierarchical knowledge extraction in the present invention;
[0031] Figure 3 A schematic diagram of the association relationship between ontologies in the present invention;
[0032] Figure 4 This is a diagram of the experimental results of the K-hop search of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the present invention clearer and more accurate, the technical solution of the present invention is described in detail below through multiple specific implementation methods. The embodiments adopted by the present invention are only used to explain the present invention and are not used to limit the content of the present invention.
[0034] In the area of standards-related text analysis methods, a large amount of research has been devoted to developing effective techniques for extracting and understanding standards documents. Traditional methods often involve manual inspection or keyword-based searches, which can be time-consuming and error-prone. Recent advances have seen the integration of natural language processing and machine learning algorithms to simplify the analysis of standards texts. Techniques such as topic modeling, document clustering, and sentiment analysis have been applied to reveal the complexity of standards documents, helping to better understand and interpret standards in various fields.
[0035] The emergence of large language models has greatly changed the landscape of knowledge graph construction. Leveraging the power of LLMs, researchers have explored innovative methods for automatically creating knowledge graphs from large datasets, including standard documents. By leveraging the contextual understanding and representation learning capabilities of LLMs, specific entities, relationships, and hierarchies in standards can be systematically identified and constructed. This approach not only speeds up the process of knowledge graph development, but also improves the accuracy and depth of standard knowledge representation.
[0036] The integration of standard text analysis methods and LLM-based knowledge graph construction technology marks a paradigm shift in the field of standardization research. By combining the analytical power of text processing algorithms with the cognitive power of LLM, a more comprehensive and automated standard knowledge representation framework is expected to revolutionize the way standards are analyzed, classified, and utilized in different fields. The synergy between these two fields is expected to advance the development of the standardization frontier and promote more efficient navigation and application of standard knowledge across industries and sectors.
[0037] Figure 1 It is a schematic diagram of a large language model in the present invention. The first aspect of the present invention relates to a method for constructing a standard analysis graph using a large language model, and the method includes steps 1 and 2. Step 1, using a predefined standard document framework to extract standard attribute structure units from a standard document, using the standard attribute structure units to extract entities, and constructing a single standard knowledge graph; Step 2, for the constructed multiple single standard knowledge graphs, obtaining the association of multiple single standard knowledge graphs through multi-level alignment, and generating a standard analysis graph.
[0038] First, the above process is implemented using an automated framework. This automated framework for building a knowledge graph for standard analysis consists of four main processes, each of which is enriched by the functions of LLM to improve effectiveness and efficiency. By integrating these four key processes with the advanced functions of large language models, the framework significantly improves the accessibility, organization, and analytical depth of standard documents, paving the way for a smarter and more efficient standard management system.
[0039] The basic information distillation process involves extracting basic data from standard documents. Leveraging LLM, the framework identifies and extracts key information, such as definitions, objectives, and related specifications. By efficiently parsing the text, LLM ensures that relevant information is accurately captured, laying a solid foundation for subsequent steps. Basic information distillation is the cornerstone of an automated framework for building a knowledge graph for standard analysis. This process aims to extract basic data from standard documents, enabling the identification and distillation of key information, such as definitions, objectives, and related specifications. By leveraging large language models, the distillation process ensures that the text is parsed efficiently and accurately, thereby extracting relevant details and laying a solid foundation for further analysis. This initial step is critical because it captures the essential components of the document, laying the foundation for subsequent processes built on this refined information.
[0040] The technical implementation of this process relies on natural language processing techniques facilitated by LLM. The process starts with preprocessing standard documents to effectively clean and format the text. Next, LLM leverages named entity recognition (NER) and contextual interpretation to identify key terms and relationships in the text, so that meaningful components such as definitions and objectives can be extracted. These models can summarize lengthy sections and provide a concise representation of key information. In addition, validation mechanisms are employed to ensure the accuracy and relevance of the extracted data, typically involving cross-validation against structured templates or standards. Finally, the refined and organized information is stored in a structured format, making it easily accessible and available for the next stage of hierarchical knowledge extraction and knowledge graph construction. In order to effectively guide LLM to perform basic information distillation.
[0041] In one embodiment, the following prompt may be used: "Analyze the following standards documents and extract essential data, including definitions, objectives, and related specifications. Organize the extracted information under clear headings, ensuring that each category is concise yet comprehensive."
[0042] The program builds on the refined information and organizes the extracted data into a hierarchical structure that reflects the multiple layers of organization inherent in the standard. The LLM helps to categorize the information into broader areas and sub-areas, down to specific standards and their specific chapters and clauses. The inclusion of a semantic layer further enhances interpretability, allowing for efficient semantic search and advanced maintenance of relationships between various elements.
[0043] Hierarchical knowledge extraction is a key process in the automated framework for building knowledge graphs, which aims to systematically organize the information extracted from standard documents into a multi-layered structure. This step ensures that the extracted data reflects the inherent hierarchy present in the standard, classifying the information into broader domains, subdomains, specific standards and their respective parts and clauses. By using large language models, this process enhances the relevance and accessibility of information, allowing users to navigate complex standards more intuitively. Integrating a semantic layer in this extraction process further improves interpretability and facilitates advanced semantic search and the maintenance of relationships between various elements, which is critical to fully understand the interconnectedness within the standard.
[0044] Initially, the extracted data is categorized and structured using LLM-driven mechanisms. The extracted information is analyzed for key topics, terms, and concepts and then organized into a hierarchy that reflects the original structure of the document. LLMs use their contextual understanding to identify terms and classify them into relevant categories, ensuring accurate representation of the broader domain. In addition, the process incorporates semantic tagging to enrich the data with contextual meaning to enhance search capabilities. By creating interconnections between related elements, the framework enables users to effectively track relationships between different levels of the hierarchy. This structured organization not only facilitates access to information, but also enables a deeper understanding of the relationships between standards, supporting improved standards analysis and management.
[0045] Through the specified prompt words, the big model extracts N keywords and key concept entities and concepts from the given standard document title and document content. This is based on the natural language processing and reasoning capabilities of the big model itself. Based on the keywords extracted from the document, the domain and subdomain of the document are determined, and then the document is identified as a hierarchical structure as shown in the following figure to form hierarchical and refined document information. After all documents are organized according to the refined hierarchical information of domain-subdomain-chapter-subchapter-clause, the big model's reasoning ability is used through prompt words to identify and judge the correlation between different domains and subdomains of each standard at different levels. For those with correlation, an associated information edge is established, which horizontally associates multiple standards at different levels. Compared with traditional standard analysis, the big model constructs hierarchical standard information, expands the correlation between standards, and improves the depth of cross-domain and other types of analysis.
[0046] The following prompts can be used: "Using information extracted from a given standards document, organize the data into a hierarchical structure. Identify broader areas and sub-areas and categorize the information into specific standards, parts, and clauses. Clearly label each category and maintain the relationships between different elements to reflect the multi-level organization inherent in the standards. Provide a structured representation to enable easy navigation and semantic search capabilities."
[0047] Preferably, a predefined standard document framework is used to extract standard attribute structure units from a standard document, including: a predefined standard document framework, wherein the standard document framework includes at least keywords, the relationship between keyword positions and standard attribute structure units; keyword matching is used to extract a keyword outline from a corresponding position of the standard document, and corresponding structural contents under each standard attribute structure unit are extracted under the keyword outline; the standard attribute structure units include the standard Chinese name, the standard English name, the standard number, the standard level, the introduction, the scope, the normative reference documents, the terms and definitions, and the core clauses.
[0048] Guided by an organized hierarchy, this process involves creating a structured knowledge graph that interconnects the various components extracted in the previous stages. The LLM plays a vital role in ensuring accurate representation of the relationships and dependencies between standards, including cross-references and nested clauses, thereby facilitating a comprehensive understanding of the structure of the standards document.
[0049] This process takes hierarchically organized data and transforms it into a knowledge graph to vividly illustrate the relationships between different elements in the standard. By using a large language model, this construction phase ensures accurate representation of the intricate relationships and dependencies in the standard, such as cross-references, nested clauses, and hierarchical connections. The resulting knowledge graph not only supports intuitive navigation and understanding of complex standards documents, but also facilitates advanced query capabilities, allowing users to gain deeper insights from interconnected information. This process utilizes LLM to extract and model entities and their relationships based on the hierarchical structure established in the previous process. Initially, key entities such as terms, sections, and related clauses are identified, which are used as nodes in the graph. Next, LLM analyzes the semantic and syntactic relationships of the extracted information to create edges that accurately represent the connections between these entities. This includes determining the nature of each relationship (such as "is a", "defines", or "references") to ensure the clarity and relevance of the graph. In addition, validating relationships and integrating ontological frameworks can improve the accuracy and functionality of the graph. The final output is a structured interactive knowledge graph that helps users effectively retrieve and analyze information in the context of the standards document.
[0050] The following prompts can be used: "Using organized hierarchical data extracted from standard documents, extract all relevant entities and relations. Identify key terms, parts, and clauses as entities (nodes) and define their interconnections (edges). Explicitly categorize relations such as 'defines', 'references', and 'is a' to accurately reflect the relationship between these entities. Build a structured knowledge graph to easily explore and understand the connected information.
[0051] For standard documents of different levels, knowledge is extracted through predefined templates. For example, for high-level standard documents, the knowledge contained therein may be more abstract and broad, and the knowledge graph is constructed by extracting themes, keywords, definitions, etc. First, an extraction template is defined, which mainly contains fields such as themes, keywords, and definitions. Then, we use the big model technology to use the above-extracted templates, themes, etc. as part of the structured prompt words, etc., to extract the relevant information from the standard document and fill it into the template.
[0052] For low-level standard documents, the knowledge they may contain is more specific and detailed. The knowledge graph is constructed by extracting the rules, steps, methods, etc. An extraction template is also defined. This template mainly contains fields such as rules, steps, and methods. Then, the relevant information in the standard document is extracted through the large model and structured prompt words, and filled into the template.
[0053] For standard documents in different fields, the knowledge is extracted through predefined domain dictionaries. For example, in the medical field, the knowledge graph is constructed by extracting diseases, symptoms, drugs, etc. For the engineering field, the knowledge graph is constructed by extracting engineering projects, engineering methods, engineering materials, etc. First, a domain dictionary needs to be built. This dictionary contains professional terms and concepts in the field, which can be combined by expert annotation and large model production. Then, we use this dictionary as a structured prompt and use the large model to identify domain-specific terms in the standard document, which will be used as entities in the knowledge graph. After identifying the entities, we need to extract the relationship between the entities. This can be achieved through NLP technologies such as dependency syntactic analysis and relation extraction. For example, we can use dependency syntactic analysis to identify the grammatical relationship between entities and use relation extraction technology to extract the semantic relationship between entities.
[0054] For knowledge graphs of different levels and fields, there may be entities or relationships with the same name but different meanings, which need to be aligned. This is done through multi-dimensional information such as context information, entity attributes, and relationship types. First, collect the context information of the entity or relationship in the knowledge graph, including its position in the graph, adjacent entities or relationships, and its attributes. This information is obtained through graph traversal algorithms such as depth-first search and breadth-first search. Secondly, for different graphs across levels, first perform direct text matching between entities based on their direct hierarchical relationships, and then perform pairwise similarity analysis of adjacent entities through entity similarity analysis to align the relationship between adjacent layers of standards with reference and inheritance relationships. For knowledge graphs in different fields, structured summaries of standards are performed through large models, and then topic models such as LDA are used to cluster different standards. Within each topic cluster, contextual similarity analysis is performed between standards and entities.
[0055] Among them, these contextual information are used to calculate the semantic similarity of entities or relationships in knowledge graphs in different fields. This is achieved through word embedding models such as Word2Vec, BERT and large models. First, the word embedding model converts the text into a high-dimensional vector, and then calculates the cosine similarity between the vectors to obtain the semantic similarity of the entity or relationship. Finally, disambiguation is performed based on the semantic similarity. A threshold interval is set to determine whether it is the same entity concept based on its height. If it is higher than the upper threshold, the confirmation results of basic language models such as BERT are directly used to consider the entities related. If it is within the threshold interval, the relationship is re-judged and confirmed through the large model and structured prompt words. Below the minimum threshold, it is considered that there is no direct connection between the entities.
[0056] For different standard documents, the main document structure units are identified. The extraction is mainly carried out through keyword matching combined with the basic standard document framework. The extracted structure mainly includes standard attribute structure units, including the standard name, standard English name, standard number, standard level, introduction, and standard content structure units, including scope, normative references, terms and definitions, core clauses and related content.
[0057] Entities are extracted using standard attribute structure units, and a single standard knowledge graph is constructed, including: taking the standard number of the standard document as the root node, extracting multi-level ontologies from the standard attribute structure units and the corresponding structural content; and constructing a single standard knowledge graph through the association between multi-level ontologies and the association between multi-level ontologies and root nodes.
[0058] Taking the standard number of the standard document as the root node, a multi-level ontology is extracted from the standard attribute structure unit and the corresponding structure content, including: taking the standard clauses as the structure content, extracting multiple clause names from one or more structure contents corresponding to each standard attribute structure unit; taking the clause names as the ontology, and constructing a multi-level ontology with the hierarchies between the clause names as the hierarchy.
[0059] Any individual standard document is constructed separately to form a single standard knowledge graph. This knowledge graph is based on the hierarchical structure of the standard, with the standard number as the root node, forming a knowledge graph with the standard attribute structure unit as the main attribute node and the standard content structure unit as the main content node. For example, taking the GB / T 36643-2018 standard as an example, the "GB / T 36643-2018" is the initial node of the ontology node. Based on this, the ontology node also includes the aforementioned attribute structure unit and content structure unit. Among them, the core clause content of the content structure unit is mainly divided into N ontology nodes such as the first level, second level, third level, and fourth level, and the name attribute of each node is the clause name.
[0060] Different standards are associated based on the two dimensions of level and field of standards to form a fused knowledge graph.
[0061] Figure 2 The schematic diagram of hierarchical knowledge extraction in the present invention. For the knowledge graphs of multiple single standards constructed, the association of the knowledge graphs of multiple single standards is obtained through multi-level alignment, including: collecting normative reference files in multiple standard documents, and extracting keywords from the normative reference files by keyword matching, and corresponding the keywords one by one to the root nodes in the knowledge graphs of multiple single standards generated in step 1; extracting the hierarchical relationship between all standard documents from the normative references, and using the hierarchical relationship to associate the root nodes of the knowledge graphs of multiple single standards, and using the association as a connection method between knowledge ontologies in the standard analysis diagram.
[0062] The normative reference documents of the standards are extracted and identified through the large model, and the standard numbers and relationships are formed to form a mandatory alignment between multi-level standards. Among them, the order of alignment is carried out according to the general standard levels of international standards, national standards, industry standards, and group standards, that is, the industry standard level refers to the national level, etc., which is a directional relationship.
[0063] The basic citation alignment is based on normative citations and extracts the specific citation relationships between the content structure units of specific standards based on the big model, forming an alignment between the content structure units of standards at different levels. Note that this alignment realizes the alignment of the clause level nodes between cross-level standards and is also directional.
[0064] Preferably, for the constructed knowledge graphs of multiple single standards, the association of the knowledge graphs of multiple single standards is obtained through multi-level alignment, including: extracting reference information located inside each standard attribute structure unit from each standard document, wherein the reference information includes at least information position and reference position; constructing the association relationship between different standards and between different standard attribute structure units of the same standard according to the information position and the reference position, and using the association as a connection method between knowledge ontologies in the standard analysis diagram.
[0065] For the constructed multiple single-standard knowledge graphs, the association of the multiple single-standard knowledge graphs is obtained through multi-level alignment, including: for the standard attribute structure units of term and definition types, all terms and definitions are extracted from such structure units; for each term or definition, the BERT model is used to perform semantic analysis to extract words with high similarity to the term or definition from multiple standard documents, and extract the position of the words; the words and word positions are used as attribute information of the current term or definition, and the word position at least includes the term name and standard attribute structure unit corresponding to the word.
[0066] Direct matching of term keywords and semantic similarity matching are used. For example, the BERT model is used to evaluate the similarity between two words, and word-level alignment in term and definition units between cross-level standards is constructed to form alignment of belonging and definition.
[0067] Figure 3 The schematic diagram of the association relationship between ontologies in the present invention. For the constructed multiple single-standard knowledge graphs, the association of the multiple single-standard knowledge graphs is obtained through multi-level alignment, including: for the bottom-level ontology in the multi-level ontology, a large language model is used to extract the clause content corresponding to the bottom-level ontology, extract the keywords of the bottom-level clauses, and generate the core vocabulary of the bottom-level ontology; the keyword abstracts of one or more bottom-level clauses corresponding to the upper-level clauses are combined to obtain the core vocabulary of the upper-level ontology; for any two ontologies in the multi-level ontology of multiple standard documents, the similarity calculation method between the core vocabulary of the ontology is used to obtain any two ontologies whose core vocabulary similarity exceeds the preset threshold for alignment.
[0068] For each lowest-level clause of the content structure unit, a large model is used to perform M keyword summaries, and the upper-level titles are combined in sequence. For example, T1-T2-T3-C, where T1 represents the title of the first-level clause, and so on, C represents the content. M keywords are extracted from C, and a core vocabulary of the T3 level is formed in T3. The core vocabulary of each layer is then calculated upwards layer by layer. Finally, the summary keyword table of all levels is obtained. Finally, the summary keyword tables of different levels between all standard documents are reversely matched layer by layer. That is, the similarity of the keyword tables at the bottom level is matched pairwise first. If the similarity meets the threshold X, the association is judged and aligned. Once aligned, the alignment of this pair of matching keyword tables of the upper level is no longer considered, reducing the matching overhead.
[0069] For the constructed multiple single-standard knowledge graphs, the association of multiple single-standard knowledge graphs is obtained through multi-level alignment, including: extracting high-level keywords in the introduction, overview, and other document contents from each standard document, and combining them into a topic description for each standard document; using Word2Vec to extract vocabulary distribution from the topic description, using the LDA topic model to analyze the vocabulary distribution in the topic description of multiple standard documents, and generating the probability of each standard document under different topics based on the vocabulary distribution; taking each topic in different topics as a knowledge ontology, adding the topic knowledge ontology to the standard analysis graph, and associating the root nodes of multiple standard documents to different topic knowledge nodes respectively.
[0070] Based on the introduction, overview and all the most advanced keyword lists, a list of topic descriptions for each standard document is combined. For the topic descriptions of all standard documents, LDA is used for topic analysis. For all standards under the same topic, a separate topic node is constructed and associated with each. The name of the topic node is generated by combining all the topic descriptions under the topic through the large model.
[0071] Based on the above core steps, different levels of alignment of cross-domain and cross-level standards are achieved.
[0072] Preferably, for the constructed multiple single-standard knowledge graphs, the association of the multiple single-standard knowledge graphs is obtained through multi-level alignment, and a standard analysis graph is generated, including: using depth-first search and breadth-first search methods to traverse entities and relationships to generate a standard analysis graph.
[0073] For knowledge graphs of different levels and fields, there may be entities or relationships with the same name but different meanings, which need to be aligned. This is done through multi-dimensional information such as context information, entity attributes, and relationship types. First, collect the context information of the entity or relationship in the knowledge graph, including its position in the graph, adjacent entities or relationships, and its attributes. This information is obtained through graph traversal algorithms such as depth-first search and breadth-first search. Secondly, for different graphs across levels, first perform direct text matching between entities based on their direct hierarchical relationships, and then perform pairwise similarity analysis of adjacent entities through entity similarity analysis to align the relationship between adjacent layers of standards with reference and inheritance relationships. For knowledge graphs in different fields, structured summaries of standards are performed through large models, and then topic models such as LDA are used to cluster different standards. Within each topic cluster, contextual similarity analysis is performed between standards and entities.
[0074] Among them, these contextual information are used to calculate the semantic similarity of entities or relationships in knowledge graphs in different fields. This is achieved through word embedding models such as Word2Vec, BERT and large models. First, the word embedding model converts the text into a high-dimensional vector, and then calculates the cosine similarity between the vectors to obtain the semantic similarity of the entity or relationship. Finally, disambiguation is performed based on the semantic similarity. A threshold interval is set to determine whether it is the same entity concept based on its height. If it is higher than the upper threshold, the confirmation results of basic language models such as BERT are directly used to consider the entities related. If it is within the threshold interval, the relationship is re-judged and confirmed through the large model and structured prompt words. Below the minimum threshold, it is considered that there is no direct connection between the entities.
[0075] The last part focuses on achieving efficient knowledge retrieval and analysis capabilities. Leveraging the power of LLM, the framework facilitates advanced search capabilities, including multi-hop queries and similarity analysis. This not only enhances user interaction with the knowledge graph, but also facilitates deeper insights into standard documents, supporting informed decision-making and strategic planning in standards management.
[0076] This process enables users to perform advanced searches and analyses, such as multi-hop queries that traverse different nodes in the graph, gaining a deeper understanding of the relationships and context embedded in standards documents. By leveraging the power of LLM, the system enhances user interaction with the knowledge graph, enabling fast and intuitive retrieval of specific information. This capability supports tasks such as identifying similarities between standards, detecting information gaps, and revealing patterns between documents, ultimately facilitating informed decision-making and strategic planning in standards management.
[0077] The knowledge retrieval and analysis process leverages the power of LLM to process queries in a context-aware manner, enabling them to efficiently navigate the knowledge graph. When a query is posed, the system first interprets the user's intent and identifies relevant nodes and relationships based on semantic understanding. The framework then employs a multi-hop query algorithm that enables it to quickly retrieve interrelated information, even if they are several layers apart in the graph; this involves not only searching for direct matches, but also understanding the context and relationships between entities to enrich the output with relevant data.
[0078] During the retrieval process, two types of operations are mainly performed. First, the user's natural language search query is converted into an accurate graph query statement based on the basic ontology information of the knowledge graph for precise matching. Secondly, the domain user's natural language is converted into a vector and compared with the standard text using basic embedding models such as TFIDF for fuzzy matching. Through precise matching and fuzzy matching, the standards related to the user's natural language query are recalled in multiple standard information associated with the graph. Identifying user intent refers to converting natural language query statements into query statements of the graph database through LLM, and the correlation between multi-standard documents has been stored in the graph through the previous steps. Similarity analysis, based on TFIDF or other embedding models such as BERT, uses Euclidean distance to judge similarity.
[0079] In addition, similarity analysis enables the system to make comparisons between different standards or clauses, helping users to identify best practices or relevant examples. Through these advanced analytical capabilities, the search process enhances accessibility and usability, providing valuable insights into the complex network of information contained in standards documents.
[0080] During this process, the following prompts can be used: "Leverage the structured knowledge graph to conduct comprehensive queries based on the query context provided to extract relevant information. Identify key entities, their relationships, and any similarities with other standards or clauses. Ensure that the analysis captures multi-hop connections and provides contextual insights to support strategic decisions in standards management.
[0081] Ontology is a collection of entity types and entity type relationships. It is essentially an organizational paradigm, that is, what entity types, relationship types, and attribute types are there. Ontology is critical to the effectiveness and efficiency of the entire framework. It serves as the infrastructure for organizing concepts, relationships, and rules, managing knowledge domains, and ensuring coherent and consistent representation of information. By providing a clear taxonomy, ontology promotes meaningful connections between various elements in the knowledge graph, thereby enabling advanced semantic search and retrieval capabilities. In addition, it enhances the interpretability of data, enabling users and LLMs to better understand the nuances and meanings of standards. This structured approach optimizes query processing, reduces ambiguity, and ensures fast and accurate retrieval of relevant information, ultimately facilitating informed decision-making and strategic planning in the context of standards management. Through a carefully designed ontology, the framework achieves a higher level of precision and relevance in knowledge retrieval, significantly improving the user experience and overall practicality of the knowledge graph.
[0082] The design of the ontology of this method reflects the particularity of standard analysis, that is, on the one hand, standards have the aforementioned hierarchical structure, and on the other hand, different standards have associations. This is the advantage of this method in standard analysis. The ontology successfully and fully integrates the basic pattern structure of the knowledge edge graph, while ensuring the consistency and interrelationship of entities and relationships across multiple levels and domains. The ontology consists of a comprehensive set of classes (nodes) that classify information into high-level domains, such as "security" and "IT", and further subdivisions include sub-domains and various levels of standards (e.g., international, national). Each of these nodes is interconnected through defined relationships, forming a robust framework that links standards to their respective parts, clauses, and other related components. The attributes in the ontology not only illustrate the hierarchical relationships, but also key metadata, including publication date, version, keywords, and source organization of the standard. This multi-layered approach enhances the ability of the knowledge graph to effectively manage and analyze standard documents, and promotes coherent understanding and application in different research fields.
[0083] This experiment will utilize a curated standard text dataset consisting of various standard documents, totaling 1,030 entries. Each document in this dataset includes at least the following components, such as section, definition, objective, etc. In addition, to enhance the evaluation of the construction algorithm, we provide an additional 80 annotations reflecting the relationships between standards, which are achieved through manual labeling. These relationship labels will help evaluate the quality and effectiveness of the knowledge graph framework and its ability to interpret and connect different standard information.
[0084] The main goal of this experiment is to validate the accuracy of multi-hop queries using the proposed knowledge graph framework on a given standard text dataset. We will compare the effectiveness of the framework with traditional text embedding methods. A set of multi-hop queries will be constructed, which will require the framework to traverse multiple relations in the knowledge graph. These queries will be designed to test the framework's ability to accurately retrieve interconnected information from standard documents.
[0085] In the present invention, StandardKG Builder will be compared with a text embedding method (Naive Embedding), which is a basic KG built from native references of the original standard (NaiveKG). Native Embedding converts standard documents into dense vector representations to perform similarity searches. The embedding method will be compared with multi-hop queries using cosine similarity. GPT-3.5 is used as the base LLM model for all programs in StandardKG Builder. Before LLM processing, standardized text formatting, segmentation, and chunking processes are applied to all standard documents and texts.
[0086] Method Evaluation The evaluation metric for multi-hop related criteria retrieval is based on the accuracy of the retrieval process. This metric is based on manually annotated data that establishes the relevance of criteria and their relations in the knowledge graph. For example, if criterion A links to criterion B, and B links to criterion C, a multi-hop query starting from A aims to retrieve relevant information about C to B.
[0087] The accuracy of multi-hop related standard retrieval is calculated using the standard machine learning accuracy formula, which measures the proportion of correct predictions made by the retrieval framework. We use F1-score as the final metric for evaluation. The F1 score is the harmonic mean of precision and recall, providing a balance between the two. The calculation formula is:
[0088] F1-score = 2 × precision × recall / precision + recall
[0089] In multi-hop queries, the F1-Score can be evaluated at different levels of hops (1 hop, 2 hops, ..., N hops). This allows analyzing the performance of the retrieval framework as the query depth increases.
[0090] Figure 4Figure 1 is the experimental result of the present invention. The F1 scores of the three methods evaluated in the multi-hop search task, specifically showing their performance in 1 to 5 hops. The results show that there is a clear difference in the effectiveness of each method as the number of hops increases. In general, StandardKG Builder consistently outperforms the other two methods at all hop levels, starting with an F1 score of 0.95 for 1 hop and then gradually decreasing to 0.68 for 5 hops. As the number of hops increases, the effectiveness of NaiveEmbedded decreases significantly, eventually obtaining a low score of 0.20 at 5 hops. Although NaiveKG shows moderate performance, starting from 0.83 hops for 1 hop and decreasing to 0.42 hops for 5 hops, indicating better resilience than Naive Embedding, it still lags behind the proposed method.
[0091] The results confirm that StandardKG Builder demonstrates strong capabilities in multi-hop search scenarios, with significant advantages over the comparison methods. On the one hand, the most notable benefit of StandardKG Builder is its robust performance in high-hop scenarios. Although its F1 score decreases with increasing hops, it maintains significantly higher overall accuracy compared to other options. This is particularly important in applications that require complex reasoning across multiple association criteria, where maintaining high accuracy is critical for reliability. In addition, the high initial F1 score of 0.95 shows that StandardKGBuilder effectively captures relationships and information in a single-hop context. This effective knowledge representation allows for better retrieval of relevant information.
[0092] Table 1 Time consumption of the comparison methods on multi-hop search
[0093]
[0094] Table 1 records the average retrieval time of the compared methods. This shows that the introduction of LLM computation in StandardKG Builder leads to increased computational cost and longer processing time. In contrast, other methods (Naive Embedding and NaiveKG) maintain more stable and controllable search times in the same multi-hop scenario, ranging from 0.43 to 1.97 seconds. This difference highlights that while StandardKG Builder can provide advanced capabilities through LLM integration, it comes at the expense of efficiency compared to the more stable performance of other methods.
[0095] In summary, this paper introduced StandardKG Builder, an automated framework that leverages the power of large language models to build knowledge graphs designed for standards analysis from different perspectives, thereby enhancing complex information extraction. Our evaluation on a comprehensive dataset of standards documents demonstrated the effectiveness and scalability of the framework. By combining knowledge representation with advanced natural language processing techniques, we significantly improved the accessibility and analysis capabilities of standards documents, laying the foundation for a more efficient and intelligent standards management system.
[0096] StandardKG Builder excels at interpreting subtle relationships in data and synthesizing this information to produce accurate, contextually relevant results. Its fine-grained knowledge edge graph topology helps fully represent interconnections, enabling the model to navigate multi-hop queries with greater accuracy and reliability. This innovative approach not only advances the field of standard analytics, but also opens up new avenues for future research in intelligent information retrieval and management systems.
[0097] The second aspect of the present invention relates to a system for constructing a standard analysis graph using a large language model, wherein the system is implemented using a method for constructing a standard analysis graph using a large language model in the first aspect of the present invention; wherein the system includes a construction module and a generation module; wherein the utilization module is used to extract standard attribute structure units from a standard document using a predefined standard document framework, extract entities using the standard attribute structure units, and construct a single standard knowledge graph; and the generation module is used to obtain the association of multiple single standard knowledge graphs constructed through multi-level alignment, and generate a standard analysis graph.
[0098] A third aspect of the present invention relates to a terminal, comprising a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method described in the first aspect of the present invention.
[0099] A fourth aspect of the present invention relates to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect of the present invention.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention is described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the technical solutions of the present invention still include contents that can modify or replace the specific embodiments of the present invention. Any modification or replacement that does not depart from the spirit and scope of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for constructing a standard analysis graph using a large language model, characterized in that: The method comprises the following steps: Use the predefined standard document framework to extract standard attribute structure units from standard documents, use the standard attribute structure units to extract entities, and build a single standard knowledge graph; For the multiple single-standard knowledge graphs constructed, the association of the multiple single-standard knowledge graphs is obtained through multi-level alignment, and a standard analysis graph is generated.
2. The method for constructing a standard analysis graph using a large language model according to claim 1, characterized in that: The method of extracting the standard attribute structure unit from the standard document by using the predefined standard document framework includes: Predefine a standard document framework, wherein the standard document framework at least includes a relationship between keywords, keyword positions and standard attribute structure units; By using keyword matching, a keyword outline is extracted from the corresponding position of the standard document, and the corresponding structural content under each standard attribute structure unit is extracted under the keyword outline; The standard attribute structure units include the standard Chinese name, standard English name, standard number, standard level, introduction, scope, normative reference documents, terms and definitions, and core clauses.
3. The method for constructing a standard analysis graph using a large language model according to claim 2, characterized in that: The method of extracting entities using standard attribute structure units and constructing a single standard knowledge graph includes: Taking the standard number of the standard document as the root node, a multi-level ontology is extracted from the standard attribute structure unit and the corresponding structure content; A single standard knowledge graph is constructed through the association between multi-level ontologies and the association between multi-level ontologies and root nodes.
4. The method for constructing a standard analysis graph using a large language model according to claim 3, characterized in that: The method uses the standard number of the standard document as the root node and extracts a multi-level ontology from the standard attribute structure unit and the corresponding structure content, including: Taking the standard clauses as structural contents, extracting multiple clause names from one or more structural contents corresponding to each standard attribute structural unit; A multi-level ontology is constructed with the clause names as the ontology and the levels between the clause names as the levels.
5. The method for constructing a standard analysis graph using a large language model according to claim 4, characterized in that: The method of obtaining associations of the plurality of knowledge graphs of the single standard by multi-level alignment for the plurality of knowledge graphs of the single standard comprises: Collect normative reference documents from multiple standard documents, extract keywords from the normative reference documents using keyword matching, and map the keywords one by one to the root nodes in the knowledge graphs of multiple single standards generated in step 1; The hierarchical relationship between all standard documents is extracted from normative references, and the root nodes of the knowledge graphs of multiple single standards are associated using the hierarchical relationship, and the association is used as a connection method between knowledge ontologies in the standard analysis graph.
6. The method for constructing a standard analysis graph using a large language model according to claim 5, characterized in that: The method of obtaining associations of the plurality of knowledge graphs of the single standard by multi-level alignment for the plurality of knowledge graphs of the single standard comprises: Extracting reference information located inside each standard attribute structure unit from each standard document, wherein the reference information at least includes information position and reference position; According to the information position and the reference position, the association relationship between different standards and between different standard attribute structure units of the same standard is constructed, and this association is used as a connection method between knowledge ontologies in the standard analysis diagram.
7. The method for constructing a standard analysis graph using a large language model according to claim 6, characterized in that: The method of obtaining associations of the plurality of knowledge graphs of the single standard by multi-level alignment for the plurality of knowledge graphs of the single standard comprises: For standard attribute structural units of term and definition types, all terms and definitions are extracted from such structural units; For each term or definition, a BERT model is used to perform semantic analysis to extract words with high similarity to the term or definition from multiple standard documents and extract the positions of the words; The word and the word position are used as the attribute information of the current term or definition, and the word position at least includes the clause name and the standard attribute structure unit corresponding to the word.
8. The method for constructing a standard analysis graph using a large language model according to claim 7, characterized in that: The method of obtaining associations of the plurality of knowledge graphs of the single standard by multi-level alignment for the plurality of knowledge graphs of the single standard comprises: For the bottom-level ontology in the multi-level ontology, a large language model is used to extract the clause content corresponding to the bottom-level ontology, extract the keywords of the bottom-level clauses, and generate a core vocabulary of the bottom-level ontology; Combine the keyword summaries of one or more bottom-level clauses corresponding to the upper-level clauses to obtain the core vocabulary of the upper-level ontology; For any two ontologies in the multi-level ontologies of multiple standard documents, the similarity calculation method between the core vocabularies of the ontology is adopted to obtain any two ontologies whose core vocabularies have similarities exceeding a preset threshold for alignment.
9. The method for constructing a standard analysis graph using a large language model according to claim 8, characterized in that: The method of obtaining associations of the plurality of knowledge graphs of the single standard by multi-level alignment for the plurality of knowledge graphs of the single standard comprises: Extract high-level keywords from the introduction, overview, and other document contents of each standard document and combine them into a topic description for each standard document; Word2Vec is used to extract vocabulary distribution from topic descriptions, and the LDA topic model is used to analyze the vocabulary distribution in the topic descriptions of multiple standard documents. The probability of each standard document under different topics is generated based on the vocabulary distribution; Each of the different topics is taken as a knowledge ontology, the subject knowledge ontology is added to the standard analysis map, and the root nodes of multiple standard documents are respectively associated with different subject knowledge nodes.
10. The method for constructing a standard analysis graph using a large language model according to claim 9, characterized in that: The method of obtaining the association of the knowledge graphs of the multiple single standards through multi-level alignment and generating a standard analysis graph for the constructed knowledge graphs of the multiple single standards includes: Depth-first search and breadth-first search methods are used to traverse entities and relationships to generate standard analysis graphs.
11. A standard analysis graph construction system using a large language model, characterized by: The system is implemented using a method for constructing a standard analysis graph using a large language model as described in any one of claims 1 to 10; wherein, The system comprises utilizing a building module and a generating module; wherein, The utilization module is used to extract standard attribute structure units from standard documents using a predefined standard document framework, extract entities using the standard attribute structure units, and construct a single standard knowledge graph; The generation module is used to obtain the association of multiple single-standard knowledge graphs through multi-level alignment for the constructed multiple single-standard knowledge graphs, and generate a standard analysis graph.
12. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Cited By
Method for automatically identifying difference between domestic and overseas standard files
CN120197607A
Large model-based standard document automatic generation and multi-dimensional auditing method and system
CN120597846A
Method and system for automatic generation and multi-dimensional review of standard documents based on large models
CN120597846B