Industrial knowledge generation method and device fusing large model and knowledge graph
By constructing a knowledge graph in industrial knowledge documents and combining large-scale models and semantic transformation models, the problem of low accuracy of large language models in industrial knowledge query is solved, achieving efficient association and semantic reasoning of industrial knowledge documents and improving the accuracy of target content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, the accuracy of querying industrial knowledge documents through large language models is low, making it difficult to meet the complex needs of semantic understanding, cross-document fusion, and knowledge reasoning.
By combining large models and knowledge graphs, a knowledge graph of industrial knowledge documents is constructed. Through semantic transformation models and similarity matching, chapter content related to keywords is determined, and deduplication and relevance analysis are performed to generate target content.
It improves the accuracy of target content in industrial knowledge documents, enables effective association and semantic reasoning of chapter content, and can accurately locate logically coherent knowledge content related to keywords.
Smart Images

Figure CN121301504B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and apparatus for generating industrial knowledge that integrates large models and knowledge graphs. Background Technology
[0002] To meet the need for intelligent industrial knowledge systems, staff need to quickly retrieve target content related to specific tasks from massive amounts of industrial documents.
[0003] In related technologies, after receiving demand information, the server inputs it into a large language model. The large language model performs semantic analysis on the demand information to determine at least one primary keyword corresponding to the demand information, and then determines the target content in the industrial document based on the primary keyword. However, the accuracy of determining the target content using the above method is low. Summary of the Invention
[0004] This application provides an industrial knowledge generation method and apparatus that integrates large models and knowledge graphs to improve the accuracy of generating target content.
[0005] In a first aspect, embodiments of this application provide an industrial knowledge generation method that integrates large models and knowledge graphs, including:
[0006] Receive requirement information sent by the client, which includes at least one keyword. The requirement information is used to instruct the generation of document content related to at least one keyword based on an industrial knowledge document. The industrial knowledge document includes multiple paragraphs.
[0007] Among multiple paragraphs, identify at least one target paragraph that is associated with the required information;
[0008] At least one primary keyword is matched with the knowledge graph corresponding to the industrial knowledge document to determine at least one chapter content in the industrial knowledge document; wherein, the industrial knowledge document includes multiple chapter titles and the chapter content corresponding to each of the multiple chapter titles, and the knowledge graph includes multiple knowledge nodes and the relationships between the multiple knowledge nodes;
[0009] Based on the requirements information, at least one target paragraph, and at least one chapter, generate target content; wherein, the target content is content in the industrial knowledge document related to at least one keyword.
[0010] In one possible implementation, identifying at least one target paragraph content associated with the requirement information among multiple paragraph contents includes:
[0011] The semantic transformation model is used to process the demand information to obtain the semantic vector corresponding to the demand information.
[0012] The semantic transformation model is used to process the content of multiple paragraphs separately, resulting in semantic vectors corresponding to each paragraph.
[0013] For each paragraph in multiple paragraphs, the semantic similarity between the requirement information and the paragraph content is determined based on the semantic vector corresponding to the requirement information and the semantic vector corresponding to the paragraph content.
[0014] Based on the semantic similarity between the demand information and the content of multiple paragraphs, at least one target paragraph is identified among the multiple paragraphs.
[0015] In one possible implementation, at least one first keyword is matched with the knowledge graph corresponding to the industrial knowledge document, and at least one chapter of content is determined in the industrial knowledge document, including:
[0016] Based on at least one primary keyword, multiple knowledge nodes, and the relationships between the multiple knowledge nodes, determine at least one candidate node that is related to at least one primary keyword among the multiple knowledge nodes;
[0017] Based on at least one candidate node and at least one first keyword, determine at least one chapter of content in the industrial knowledge document.
[0018] In one possible implementation, based on at least one candidate node and at least one first keyword, at least one chapter of content is determined in the industrial knowledge document, including:
[0019] For each primary keyword in at least one primary keyword, perform the following steps:
[0020] For each candidate node in at least one candidate node, determine the similarity between the candidate node and the first keyword;
[0021] The chapter content corresponding to the first keyword is determined based on the similarity between each candidate node and the first keyword.
[0022] Among them, at least one chapter contains the chapter content corresponding to at least one primary keyword.
[0023] In one possible implementation, the method further includes:
[0024] Determine the prompt text, which includes questions and prompt words for identifying the second keyword of an industrial knowledge document to a pre-trained large model;
[0025] The prompt text, multiple chapter titles, and the chapter content corresponding to each chapter title are input into a pre-trained large model to obtain multiple secondary keywords for industrial knowledge documents;
[0026] Based on multiple secondary keywords, multiple chapter titles, and the chapter content corresponding to each chapter title, determine multiple knowledge nodes and the relationships between these knowledge nodes.
[0027] In one possible implementation, target content is generated based on demand information, at least one target paragraph content, and at least one chapter content, including:
[0028] Deduplication is performed on at least one paragraph and at least one chapter to obtain at least one candidate content;
[0029] Generate target content based on the demand information and at least one candidate content.
[0030] In one possible implementation, the target content is generated based on demand information and at least one candidate content, including:
[0031] For each candidate content in at least one candidate content, the candidate content and demand information are input into the relevance analysis model to obtain the semantic relevance information between the candidate content and the demand information;
[0032] Based on the semantic relevance information between each of the at least one candidate content and the demand information, at least one target candidate content is determined from the at least one candidate content.
[0033] Determine the generated prompt text, which includes posing questions and prompts to the pre-trained large model to determine the target content to be generated;
[0034] The generated prompt text and at least one target candidate content are input into a pre-trained large model to obtain the target content.
[0035] Secondly, embodiments of this application provide an industrial knowledge generation device that integrates large models and knowledge graphs, comprising:
[0036] The receiving module is used to receive the requirement information sent by the client. The requirement information includes at least one keyword and is used to instruct the generation of document content related to the at least one keyword based on the industrial knowledge document. The industrial knowledge document includes multiple paragraphs.
[0037] The first determining module is used to determine at least one target paragraph content associated with the requirement information from multiple paragraph contents;
[0038] The processing module is used to match at least one first keyword with the knowledge graph corresponding to the industrial knowledge document, and to determine at least one chapter content in the industrial knowledge document; wherein, the industrial knowledge document includes multiple chapter titles and the chapter content corresponding to each of the multiple chapter titles, and the knowledge graph includes multiple knowledge nodes and the relationships between the multiple knowledge nodes;
[0039] The generation module is used to generate target content based on the requirements information, at least one target paragraph content, and at least one chapter content; wherein, the target content is content in the industrial knowledge document related to at least one keyword.
[0040] In one possible implementation, the first determining module is specifically used for:
[0041] The semantic transformation model is used to process the demand information to obtain the semantic vector corresponding to the demand information.
[0042] The semantic transformation model is used to process the content of multiple paragraphs separately, resulting in semantic vectors corresponding to each paragraph.
[0043] For each paragraph in multiple paragraphs, the semantic similarity between the requirement information and the paragraph content is determined based on the semantic vector corresponding to the requirement information and the semantic vector corresponding to the paragraph content.
[0044] Based on the semantic similarity between the demand information and the content of multiple paragraphs, at least one target paragraph is identified among the multiple paragraphs.
[0045] In one possible implementation, the processing module is specifically used for:
[0046] Based on at least one primary keyword, multiple knowledge nodes, and the relationships between the multiple knowledge nodes, determine at least one candidate node that is related to at least one primary keyword among the multiple knowledge nodes;
[0047] Based on at least one candidate node and at least one first keyword, determine at least one chapter of content in the industrial knowledge document.
[0048] In one possible implementation, the processing module is specifically used for:
[0049] For each primary keyword in at least one primary keyword, perform the following steps:
[0050] For each candidate node in at least one candidate node, determine the similarity between the candidate node and the first keyword;
[0051] The chapter content corresponding to the first keyword is determined based on the similarity between each candidate node and the first keyword.
[0052] Among them, at least one chapter contains the chapter content corresponding to at least one primary keyword.
[0053] In one possible implementation, the industrial knowledge generation apparatus integrating large models and knowledge graphs further includes a second determining module, wherein the second determining module is used for:
[0054] Determine the prompt text, which includes questions and prompt words for identifying the second keyword of an industrial knowledge document to a pre-trained large model;
[0055] The prompt text, multiple chapter titles, and the chapter content corresponding to each chapter title are input into a pre-trained large model to obtain multiple secondary keywords for industrial knowledge documents;
[0056] Based on multiple secondary keywords, multiple chapter titles, and the chapter content corresponding to each chapter title, determine multiple knowledge nodes and the relationships between these knowledge nodes.
[0057] In one possible implementation, the generation module is specifically used for:
[0058] Deduplication is performed on at least one paragraph and at least one chapter to obtain at least one candidate content;
[0059] Generate target content based on the demand information and at least one candidate content.
[0060] In one possible implementation, the generation module is specifically used for:
[0061] For each candidate content in at least one candidate content, the candidate content and demand information are input into the relevance analysis model to obtain the semantic relevance information between the candidate content and the demand information;
[0062] Based on the semantic relevance information between each of the at least one candidate content and the demand information, at least one target candidate content is determined from the at least one candidate content.
[0063] Determine the generated prompt text, which includes posing questions and prompts to the pre-trained large model to determine the target content to be generated;
[0064] The generated prompt text and at least one target candidate content are input into a pre-trained large model to obtain the target content.
[0065] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0066] The memory stores the instructions that the computer executes;
[0067] The processor executes computer execution instructions stored in memory, causing the processor to perform the methods involved in the first aspect and / or any possible implementation of the first aspect.
[0068] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the methods involved in the first aspect and / or any possible implementation of the first aspect.
[0069] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0070] The memory stores the instructions that the computer executes;
[0071] The processor executes computer execution instructions stored in memory, causing the processor to perform the methods involved in the first aspect and / or any possible implementation of the first aspect.
[0072] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods involved in the first aspect and / or any possible implementation of the first aspect.
[0073] The industrial knowledge generation method and apparatus that integrates large models and knowledge graphs provided in this application, after receiving demand information, combines the demand information and the knowledge graph to obtain at least one chapter content, including chapter content that is related to at least one keyword. For industrial knowledge documents that design multiple standards, process documents, and design schemes, it realizes effective association, integration, and semantic reasoning of chapter content, and can accurately locate chapter content that is related to at least one keyword and has coherent knowledge logic, thereby improving the accuracy of generating target content. Attached Figure Description
[0074] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0075] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application;
[0076] Figure 2 A flowchart illustrating an industrial knowledge generation method that integrates large models and knowledge graphs, provided for an embodiment of this application;
[0077] Figure 3 A flowchart illustrating the process of determining the content of at least one chapter, as provided in an embodiment of this application;
[0078] Figure 4 A schematic diagram illustrating the determination of a knowledge graph corresponding to an industrial knowledge document, provided as an embodiment of this application;
[0079] Figure 5 A flowchart illustrating the process of determining target content provided in an embodiment of this application;
[0080] Figure 6 A schematic diagram illustrating an industrial knowledge generation method that integrates large models and knowledge graphs, provided as an embodiment of this application;
[0081] Figure 7 A schematic diagram illustrating the determination of target content provided in an embodiment of this application;
[0082] Figure 8 A schematic diagram of an industrial knowledge generation device that integrates large models and knowledge graphs is provided for embodiments of this application;
[0083] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0084] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0085] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0086] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0087] With the continuous upgrading of manufacturing industries such as high-end equipment manufacturing, energy equipment, rail transportation, and aerospace, industrial production processes are becoming increasingly standardized, complex, and refined. Knowledge of process design, assembly operation procedures, and parameter control specifications has become a key component of the industrial knowledge system. This type of knowledge is typically recorded in unstructured documents such as industry standards, technical specifications, operating procedures, and quality manuals. It is characterized by high-density professional terminology, numerous logical dependencies, and cross-task information overlap, making it a typical example of complex industrial knowledge.
[0088] In some embodiments, staff need to retrieve task-related knowledge from a large number of industrial knowledge documents. For example, assuming the task involves part design changes, assembly path adjustments, or exception handling, staff would need to retrieve corresponding process segments, reference standards, or historical processing solutions from a large number of industrial knowledge documents.
[0089] In related technologies, workers can search for target content in industrial knowledge documents using traditional information retrieval methods, such as first-keyword matching and full-text document retrieval. However, traditional information retrieval methods lack deep semantic understanding and structural association recognition capabilities, making it difficult to meet complex needs such as semantic understanding, cross-document fusion, knowledge reasoning, and standard generation. Therefore, natural language processing technologies such as semantic retrieval, knowledge question answering, and text generation can be introduced into industrial knowledge document query scenarios.
[0090] In some embodiments, after receiving the demand information, the server performs semantic vectorization processing on the industrial knowledge document and the demand information based on the embedding model, obtaining at least one semantic vector corresponding to the industrial knowledge document and a semantic vector corresponding to the demand information. Then, the server uses a re-ranker model to determine the semantic similarity between the semantic vector corresponding to the demand information and at least one semantic vector corresponding to the industrial knowledge document, and determines the target content based on the semantic similarity between the semantic vector corresponding to the demand information and at least one semantic vector corresponding to the industrial knowledge document. The embedding model is used to convert unstructured text into semantic vectors.
[0091] In some embodiments, with the development of Bidirectional Encoder Representations from Transformers (BERT), Chat Generative Pre-trained Transformer (ChatGPT), and General Language Model (GLM), the server, upon receiving the requirement information, inputs the requirement information and prompt text into the large language model. The large language model performs semantic analysis on the requirement information, determines at least one primary keyword corresponding to the requirement information, and identifies the target content in the industrial knowledge document based on the at least one primary keyword.
[0092] However, since large language models are usually trained on general corpora, they have difficulty understanding professional concepts such as numbering systems, process status, procedure templates, and mixed text and graphics structures that are unique to the industrial knowledge domain. Furthermore, large language models lack guidance and structural constraints from industrial knowledge, resulting in low accuracy of the target content obtained through queries using large language models.
[0093] Based on this, this application provides an industrial knowledge generation method that integrates large models and knowledge graphs. It determines the knowledge graph corresponding to the industrial knowledge document. After receiving the requirement information, it combines the requirement information and the knowledge graph to obtain at least one chapter content, including chapter content that is related to at least one keyword. For industrial knowledge documents that design multiple standards, process documents, and design schemes, it realizes effective association, integration, and semantic reasoning of chapter content, and can accurately locate chapter content that is related to at least one keyword and has coherent knowledge logic, thereby improving the accuracy of determining the target content.
[0094] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0095] Figure 1 Please refer to the schematic diagram of an application scenario provided in this application embodiment. Figure 1 It includes a client 11 and a server 12. The client 11 can be a device such as a mobile phone or a computer, and the server 12 stores industrial knowledge documents.
[0096] In practical applications, data can be transmitted between the client 11 and the server 12. For example, when it is necessary to query document content related to at least one first keyword from an industrial knowledge document, the client 11 can send a request to the server 12. The request is used to request the query of document content related to at least one first keyword in the industrial knowledge document. After receiving the request, the server 12 determines the target content in the industrial knowledge document based on the request and sends the target content to the client 11.
[0097] It should be noted that, Figure 1 This is merely an example to illustrate one application scenario, and is not intended to limit the application scenario.
[0098] The technical solutions of this application and how they solve the aforementioned technical problems are described in detail below with specific embodiments. These specific embodiments may exist independently or in combination with each other. Identical or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0099] Figure 2 This is a flowchart illustrating an industrial knowledge generation method that integrates large models and knowledge graphs, provided as an embodiment of this application. Please refer to... Figure 2 The method may include the following steps:
[0100] S21. Receive the requirement information sent by the client. The requirement information includes at least one keyword. The requirement information is used to instruct the generation of document content related to at least one keyword based on the industrial knowledge document. The industrial knowledge document includes multiple paragraphs.
[0101] The client can be a terminal device such as a mobile phone or computer. In some embodiments, when staff need to query relevant document content for a specific task in the industrial knowledge document, they can send the request information to the server through the client to request the determination of content related to at least one first keyword based on the industrial knowledge document.
[0102] It should be noted that the multiple paragraphs in an industrial knowledge document can be divided based on content paragraphs or content chapters.
[0103] In some embodiments, at least one keyword is obtained based on requirement information. Specifically, after receiving the requirement information, the server can perform semantic analysis on the requirement information to determine at least one keyword in the requirement information, and the at least one keyword is used to represent the task requirements of a specific task.
[0104] For example, assuming the requirement information is "What are the welding parameters of an aero-engine blade", semantic analysis is performed on the requirement information to determine that at least one keyword includes "aero-engine blade" and "welding parameters".
[0105] S22. Among multiple paragraph contents, identify at least one target paragraph content that is associated with the requirement information.
[0106] Among them, at least one target paragraph associated with the demand information can be, for example, at least one target paragraph that has a semantic relationship with the demand information.
[0107] In some embodiments, after receiving the requirement information, the server performs semantic analysis on the requirement information and multiple paragraph contents to determine the degree of semantic association between each paragraph and the requirement information, and then determines the paragraph contents among the multiple paragraph contents whose degree of semantic association is greater than the degree threshold as the target paragraph.
[0108] For example, suppose that multiple paragraph contents include paragraph content 1, paragraph content 2 and paragraph content 3, wherein the semantic correlation between paragraph content 1 and the requirement information is 50%, the semantic correlation between paragraph content 2 and the requirement information is 0%, and the semantic correlation between paragraph content 3 and the requirement information is 60%, and the degree threshold is 40%, then at least one target paragraph is determined to include paragraph content 1 and paragraph content 3.
[0109] S23. Match at least one primary keyword with the knowledge graph corresponding to the industrial knowledge document to determine at least one chapter content in the industrial knowledge document; wherein, the industrial knowledge document includes multiple chapter titles and the chapter content corresponding to each of the multiple chapter titles, and the knowledge graph includes multiple knowledge nodes and the relationships between the multiple knowledge nodes.
[0110] The knowledge graph corresponding to the industrial knowledge document is a knowledge representation based on the structured relationships built upon the industrial knowledge document. It is used to decompose the chapter structure and core content of the industrial knowledge document into multiple knowledge nodes, as well as the relationships between these knowledge nodes. In some embodiments, multiple knowledge nodes may include chapter titles, core terms, process parameters, etc., from the industrial knowledge document, and the relationships between multiple knowledge nodes may include, for example, relationships such as inclusion, belonging, referencing, dependency, and comparison.
[0111] Multiple chapter titles are hierarchical text identifiers used in industrial knowledge documents to distinguish content from different topics. In some embodiments, multiple chapter titles can be determined based on the hierarchical structure of the industrial knowledge document. Then, for each chapter title, the specific knowledge text under that chapter title is determined as the chapter content corresponding to that chapter title. In some embodiments, after determining the chapter titles and their corresponding chapter content, the semantic mapping relationship between "chapter title - chapter content" can be determined based on the hierarchical dependency between chapter titles and chapter content.
[0112] At least one chapter is the chapter content among multiple chapter contents that is related to the requirement information. In some embodiments, after receiving the requirement information, the server, for each of the at least one first keyword, performs matching processing on each first keyword with multiple knowledge nodes in the knowledge graph to determine at least one knowledge node that matches the first keyword. Then, combining the relationships between multiple knowledge nodes and the at least one knowledge node, the knowledge node that is related to the first keyword is determined. Finally, the chapter content corresponding to the knowledge node in the industrial knowledge document that is related to at least one first keyword is determined as at least one chapter content.
[0113] S24. Generate target content based on the requirement information, at least one target paragraph content, and at least one chapter content; wherein, the target content is content in the industrial knowledge document related to at least one keyword.
[0114] In some embodiments, since at least one paragraph and at least one chapter are initially selected as content related to at least one keyword, to more accurately determine the content related to at least one keyword, the content of at least one chapter and at least one paragraph can be further filtered based on the requirements information to determine content that is related to at least one keyword and has a high degree of relevance to at least one related term. Then, target content related to the requirements information is generated based on the content that is related to at least one keyword and has a high degree of relevance to at least one related term.
[0115] For example, semantic analysis can be performed on the demand information, at least one paragraph, and at least one chapter to determine the relevance between each paragraph and the demand information, and the relevance between each chapter and the demand information. Chapters and / or paragraphs with relevance greater than a preset threshold are identified as content that is associated with at least one keyword and has a high degree of relevance to at least one related word. Then, the server inputs the content associated with at least one keyword and having a high degree of relevance to at least one related word into a pre-trained large model. The pre-trained large model processes this content to generate the target content.
[0116] exist Figure 2 In the illustrated embodiment, this application utilizes a knowledge graph corresponding to the industrial knowledge document to effectively associate and integrate complex knowledge content within the document. Then, upon receiving the requirement information, the application combines the requirement information and the knowledge graph to obtain at least one chapter, including chapter content associated with at least one keyword. For industrial knowledge documents containing multiple standards, process documents, and design schemes, this approach achieves effective association, integration, and semantic reasoning of chapter content, accurately locating chapter content related to at least one keyword and with coherent knowledge logic, thereby improving the accuracy of generating the target content.
[0117] exist Figure 2 Based on the illustrated embodiment, the following, in conjunction with Figure 3 The process of identifying at least one chapter in an industrial knowledge document is described in detail.
[0118] Figure 3 This is a flowchart illustrating a process for determining the content of at least one chapter, as provided in an embodiment of this application. Figure 3 As shown, the process may include the following steps:
[0119] S31. Process the industrial knowledge documents and determine the knowledge graph corresponding to the industrial knowledge documents.
[0120] In some embodiments, the knowledge graph corresponding to an industrial knowledge document can be determined as follows: a prompt text is determined, which includes a question and prompt words posing a second keyword to the pre-trained large model to determine the industrial knowledge document; the prompt text, multiple chapter titles, and the chapter content corresponding to each chapter title are input into the pre-trained large model to obtain multiple second keywords of the industrial knowledge document; based on the multiple second keywords, multiple chapter titles, and the chapter content corresponding to each chapter title, multiple knowledge nodes and the relationships between the multiple knowledge nodes are determined.
[0121] Among these, a pre-trained large model can be, for example, a pre-trained large language model. It's important to note that this pre-trained large model is trained based on industry knowledge, enabling it to learn the language habits, knowledge systems, and task logic of the industry domain. Therefore, a pre-trained large model possesses the knowledge understanding and professional task processing capabilities of the industry domain.
[0122] The prompt text is used to instruct the pre-trained large model to process multiple chapter titles and their corresponding chapter contents to determine the second keyword of the industrial knowledge document. Furthermore, the prompt text also constrains and guides the pre-trained large model's process of determining the second keyword, ensuring that the second keyword meets the construction requirements of the knowledge graph. In some embodiments, the prompt text can be determined through prompt word engineering.
[0123] For example, the prompt text could be: "Based on multiple chapter titles and their corresponding chapter contents in the industrial knowledge document, please determine the second keywords that reflect the core knowledge of each chapter in the industrial knowledge document. Extract at least 2-3 second keywords from each chapter, and the second keywords should focus on industrial professional terms (such as process parameters, standard numbers, equipment names, etc.)."
[0124] Multiple secondary keywords are data or concepts extracted from industrial knowledge documents that represent the core meaning of each chapter's content. For example, assuming the chapter content of "high-temperature alloy welding process for turbine blades" in an industrial knowledge document, multiple secondary keywords could include "welding groove form," "heat input control," "cooling rate," "grain coarsening," and "defect detection," etc.
[0125] In some embodiments, prompt text, multiple chapter titles, and the corresponding chapter content for each chapter title can be input into a pre-trained large model. The pre-trained large model processes the multiple chapter titles and their corresponding chapter content based on the prompt text to determine multiple candidate second keywords for the industrial knowledge document. Then, the multiple candidate second keywords are filtered and merged to obtain the multiple second keywords for the industrial knowledge document.
[0126] Then, using rule parsing and semantic extraction, multiple secondary keywords, multiple chapter titles, and the chapter content corresponding to each chapter title are processed to determine multiple knowledge nodes and the relationships between them. In some embodiments, a knowledge graph of industrial knowledge documents can be constructed based on multiple knowledge nodes and the relationships between them, using "knowledge node-relationship-knowledge node" triples. For example, the resulting triple could be: "High-temperature alloy welding process" - includes - "Heat input control" - "Heat input control should maintain... (chapter content corresponding to heat input control)"
[0127] In some embodiments, a graph structure can be built based on NetworkX, and attribute fields (such as the document to which it belongs, level, keyword tags, etc.) can be added to each knowledge node. Each knowledge node generates hierarchical connection paths based on its title level, and a cross-standard reference graph is constructed (such as the mapping relationship between CPS1000 and CPS0100).
[0128] In some embodiments, when there are multiple industrial knowledge documents, the pre-trained large model will automatically perform knowledge graph fusion processing. The fusion process includes node merging (based on standardized labels), path completion (preserving the integrity of the structure tree), and conflict node handling (preserving the semantically strongest branch). The fused knowledge graph is imported into Neo4j as a graph database or output to PyVis for visualization, and serves as the knowledge foundation for subsequent retrieval and generation modules.
[0129] In some embodiments, tag fusion, node merging, and hierarchical path identification can be combined to further improve the completeness and reusability of the graph. Specifically, triples can be batch-organized into a JSON structure, supporting multi-source document aggregation. Each triple has a unique number field for subsequent knowledge node location and dependency analysis. Furthermore, if the triple is generated through inference rather than direct extraction, a device information file field will be recorded to identify its inference source.
[0130] Furthermore, considering the diversity of terminology and inconsistencies in expression within industrial documents, the pre-trained large model incorporates a vocabulary mapping table and a semantic clustering model to perform synonym merging, format standardization, and domain mapping on extracted keywords. For example, "heat input control" and "welding heat energy control" are grouped into a unified entity; "CPS1000-A chapter," "Section A," and "Standard number A" are mapped to unified nodes. These standardized entities are used to construct a consistent graph structure, improving cross-document matching capabilities.
[0131] Specifically, it can be combined with Figure 4 To understand, Figure 4This is a schematic diagram illustrating how to determine the knowledge graph corresponding to an industrial knowledge document, as provided in an embodiment of this application. Figure 4 As shown, the server determines the titles at different levels in the industrial knowledge document according to the process specifications, as well as the content corresponding to each title. Then, it adds tags and performs other operations to at least one chapter title and the chapter content corresponding to each chapter title in the industrial knowledge document.
[0132] Then, the prompt text, multiple chapter titles, and their corresponding chapter contents are input into a pre-trained large model. Based on the prompt text, the pre-trained large model processes the chapter titles and their corresponding chapter contents to determine multiple candidate secondary keywords for the industrial knowledge document. These candidate secondary keywords are then filtered and merged to obtain the final secondary keywords for the industrial knowledge document.
[0133] The algorithm processes multiple secondary keywords, multiple chapter titles, and the corresponding chapter content of each chapter title to determine multiple knowledge nodes and the relationships between them. This results in a "knowledge node-relationship-knowledge node" triple. In some embodiments, the obtained triples, along with at least one chapter title and its corresponding chapter content, can be used to determine the chapter content corresponding to each of the multiple knowledge nodes. Then, based on the triples and the corresponding chapter content of each knowledge node, a knowledge graph is determined. The resulting knowledge graph includes multiple knowledge nodes, the relationships between them, and the corresponding chapter content of each knowledge node.
[0134] It should be noted that knowledge graphs can be saved in standard JSON format, supporting visualization and graph structure navigation.
[0135] S32. Based on at least one primary keyword, multiple knowledge nodes, and the relationships between the multiple knowledge nodes, determine at least one candidate node that is related to at least one primary keyword among the multiple knowledge nodes.
[0136] In some embodiments, for each keyword in the at least one keyword, at least one knowledge node similar to the keyword can be determined from multiple knowledge nodes based on methods such as similarity retrieval. For each knowledge node among the at least one knowledge node similar to the keyword, at least one candidate node related to the keyword is determined by combining the association relationship between the knowledge node and other knowledge nodes in the knowledge graph. That is, the at least one candidate node related to the keyword includes at least one knowledge node similar to the keyword, as well as knowledge nodes that have an association relationship with at least one knowledge node.
[0137] S33. Based on at least one candidate node and at least one first keyword, determine at least one chapter in the industrial knowledge document.
[0138] In some embodiments, the method for determining at least one chapter content may be as follows: for each of the at least one first keywords, perform the following steps: for each of the at least one candidate nodes, determine the similarity between the candidate node and the first keyword; determine the chapter content corresponding to the first keyword based on the similarity between each of the at least one candidate node and the first keyword; wherein, the at least one chapter content includes the chapter content corresponding to each of the at least one first keyword.
[0139] In some embodiments, for each candidate node among at least one candidate node, if the similarity between the candidate node and the first keyword is greater than or equal to a similarity threshold, then the chapter content corresponding to the candidate order is determined as the chapter content corresponding to the first keyword.
[0140] For example, suppose at least one candidate node includes candidate node 1, candidate node 2, and candidate node 3, wherein the similarity between candidate node 1 and the first keyword 1 is 30%, the similarity between candidate node 2 and the first keyword 1 is 70%, the similarity between candidate node 3 and the first keyword 1 is 10%, the similarity between candidate node 1 and the first keyword 2 is 60%, the similarity between candidate node 2 and the first keyword 2 is 32%, and the similarity between candidate node 3 and the first keyword 2 is 23%, and the similarity threshold is 40%. Then, the chapter content corresponding to the first keyword 1 is determined to be the chapter content corresponding to candidate node 2; the chapter content corresponding to the first keyword 2 is determined to be the chapter content corresponding to candidate node 1. Therefore, at least one chapter content is determined to include the chapter content corresponding to candidate node 1 and the chapter content corresponding to candidate node 2.
[0141] exist Figure 3 In the illustrated embodiment, by automatically constructing an industrial knowledge graph, structural information such as assembly sequence, parameter dependencies, and material constraints in complex industrial documents is explicitly expressed. This overcomes the limitation of traditional embedding models, which only support plain text similarity retrieval, enabling the system to possess stronger semantic organization and logical modeling capabilities. Training a large model based on industrial knowledge allows the pre-trained model to understand and apply professional semantics such as numbering rules, process levels, and standard terminology, thus compensating for the shortcomings of general-purpose large language models in terms of industry adaptability.
[0142] Based on the above embodiments, the following, in conjunction with Figure 5 This application will further explain the process of determining the target content.
[0143] Figure 5 This is a schematic diagram illustrating a process for determining target content, provided as an embodiment of this application. Figure 5 As shown, the process may include the following steps:
[0144] S51. Among multiple paragraph contents, identify at least one target paragraph content that is associated with the requirement information.
[0145] In some embodiments, the method for determining at least one target paragraph content can be as follows: processing the requirement information according to a semantic transformation model to obtain a semantic vector corresponding to the requirement information; processing multiple paragraph contents according to the semantic transformation model to obtain semantic vectors corresponding to each of the multiple paragraph contents; determining the semantic similarity between the requirement information and the paragraph content based on the semantic vectors corresponding to the requirement information and the paragraph content; and determining at least one target paragraph content among the multiple paragraph contents based on the semantic similarity between the requirement information and each of the multiple paragraph contents.
[0146] Semantic transformation models are used to convert unstructured text into semantic vectors. For example, an embedding model can be used. It's important to note that the semantic vectors obtained after semantic transformation of unstructured text do not contain random numbers. Instead, they are vectors representing the semantic meaning of the text, learned by the semantic transformation model.
[0147] In some embodiments, the semantic transformation model is trained based on first sample data and a first label. The first sample data may be, for example, unstructured text, and the first label may be, for example, the semantic vector label corresponding to the unstructured text.
[0148] It should be noted that the first sample data can include one or more unstructured texts. Correspondingly, the first label can also contain one or more semantic vector labels corresponding to the unstructured texts, and the number of unstructured texts is the same as the number of semantic vector labels corresponding to the unstructured texts.
[0149] During the training of the semantic translation model, the server can input the first sample data into the semantic translation model, process the first sample data, and obtain the semantic vector of the first sample data output by the semantic translation model. Then, the loss value is calculated based on the semantic vector and semantic vector label of the first sample data, and the model parameters of the semantic translation model are adjusted according to the calculated loss value, thus completing one round of training.
[0150] The server can train the semantic translation model in one or more rounds until the training termination condition is met, at which point the training process stops, and the trained semantic translation model is obtained. The training termination condition can be set according to actual needs; for example, it can be set that the loss value is less than or equal to a preset loss value, or that the number of training iterations reaches a preset number, etc. The trained semantic translation model can process multiple paragraphs of content in an industrial knowledge document, obtaining semantic vectors corresponding to each paragraph.
[0151] In some embodiments, after receiving demand information, the server can input the demand information into a semantic transformation model. The semantic transformation model processes the demand information to obtain a semantic vector corresponding to the demand information. Furthermore, the server can also input multiple paragraphs from an industrial knowledge document into the semantic transformation model. The semantic transformation model processes the multiple paragraphs to obtain semantic vectors corresponding to each paragraph.
[0152] Then, for each paragraph within the multiple paragraph contents, a preliminary similarity matching is performed between the corresponding semantic vector and the semantic vector corresponding to that paragraph content, based on the paragraph vector index library, to obtain the semantic similarity between the requirement information and the paragraph content. Here, semantic similarity is used to quantify the degree of association between the requirement information and the paragraph content at the "semantic level." In some embodiments, semantic similarity can be expressed as a percentage, with a higher semantic similarity indicating a higher degree of association between the requirement information and the paragraph content at the "semantic level," and a lower semantic similarity indicating a lower degree of association between the requirement information and the paragraph content at the "semantic level."
[0153] Therefore, paragraphs with a semantic similarity to the required information that is greater than or equal to a preset similarity can be identified as target paragraphs.
[0154] For example, suppose that multiple paragraph contents include paragraph content 1, paragraph content 2 and paragraph content 3, wherein the semantic similarity between paragraph content 1 and the requirement information is 60%, the semantic similarity between paragraph content 2 and the requirement information is 75%, the semantic similarity between paragraph content 3 and the requirement information is 40%, and the preset similarity is 50%, then it is determined that at least one target paragraph content includes paragraph content 1 and paragraph content 2.
[0155] S52. Match at least one primary keyword with the knowledge graph corresponding to the industrial knowledge document, and determine at least one chapter in the industrial knowledge document.
[0156] For a detailed introduction to identifying at least one chapter in the industrial knowledge document, please refer to [link / reference needed]. Figure 3 The embodiments shown will not be described in detail here.
[0157] S53. Perform deduplication on at least one paragraph and at least one chapter to obtain at least one candidate content.
[0158] In some embodiments, at least one paragraph and at least one chapter may contain duplicate or highly similar content. Therefore, deduplication can be performed on at least one paragraph and at least one chapter to remove redundant or highly similar content.
[0159] It should be noted that redundant content that is repeated or highly similar in at least one paragraph and at least one chapter does not mean that the content in at least one paragraph and at least one chapter is "completely identical". Rather, it is judged as redundant from two dimensions, "text surface" and "semantic depth", taking into account the professionalism of industrial documents.
[0160] S54. Generate target content based on the demand information and at least one candidate content.
[0161] In some embodiments, the method for determining the target content among at least one candidate content can be as follows: For each candidate content among at least one candidate content, the candidate content and demand information are input into a relevance analysis model to obtain semantic relevance information between the candidate content and the demand information; based on the semantic relevance information between each of the at least one candidate content and the demand information, at least one target candidate content is determined among the at least one candidate content; a prompt text is generated, which includes proposing a question and prompt words to a pre-trained large model to determine the target content; the generated prompt text and at least one target candidate content are input into the pre-trained large model to obtain the target content.
[0162] In some embodiments, the relevance analysis model is used to analyze the semantic relevance between candidate content and demand information, and the relevance analysis model can be, for example, a fusion reordering model, which can include, for example, BERT with a dual-tower or interactive semantic encoding structure, Robustly Optimized BERT Approach (RoBERTa), or a cross encoder optimized for industrial text.
[0163] In some embodiments, the relevance analysis model is trained based on second sample data and second labels. The second sample data may include, for example, first sample content and multiple second sample contents, and the second labels may include, for example, semantic relevance labels between each of the multiple second sample contents and the first sample content. It should be noted that the first sample content includes a "question-document fragment-matching level" dataset, covering multiple tasks such as process question answering, standard comparison, and parameter interpretation, and possesses strong domain generalization ability.
[0164] During the training of the correlation analysis model, the server can input second sample data into the model. The model processes the first sample content and multiple second sample contents within the second sample data to determine the semantic relevance information between each of the second sample contents and the first sample content. Then, based on the semantic relevance labels between each of the second sample contents and the first sample content, and the semantic relevance information between each of the second sample contents and the first sample content, a loss value is determined. The model parameters of the correlation analysis model are then adjusted based on the calculated loss value, thus completing one round of training.
[0165] The server can train the relevance analysis model through one or more rounds until the training termination condition is met, at which point the training process stops, resulting in a completed relevance analysis model. The training termination condition can be set according to actual needs; for example, it can be set that the loss value is less than or equal to a preset loss value, or that the number of training iterations reaches a preset number, etc. The trained relevance analysis model can process candidate content and demand information to obtain semantic relevance information between them.
[0166] In some embodiments, after determining at least one candidate content, the server inputs the candidate content and demand information into a relevance analysis model for each candidate content. The relevance analysis model processes the candidate content and demand information through a deep semantic interaction mechanism to obtain semantically related information between the candidate content and demand information. The semantic relevance information is used to indicate the degree of semantic association between the candidate content and demand information, and the semantic relevance information can be represented, for example, in the form of semantic relevance degree, semantic relevance confidence, etc.
[0167] It should be noted that, to improve the accuracy of determining semantic relevance information, the semantic relevance information can also be determined by combining the title level of the candidate content in the industrial knowledge document, the frequency of its occurrence in the industrial knowledge document, and whether the candidate content is a key knowledge node in the knowledge graph. For example, a weighted calculation can be performed on the title level of the candidate content in the industrial knowledge document, the frequency of its occurrence in the industrial knowledge document, and whether the candidate content is a key knowledge node in the knowledge graph, and then the semantic relevance information can be determined by combining the weighted sum obtained from the calculation.
[0168] In some embodiments, a similarity re-ranking strategy can be used to determine the degree of matching between at least one candidate content and the requirement information. The similarity re-ranking strategy calculates the degree of matching between each candidate content and the requirement information from multiple dimensions, such as document level, structural path position, and vocabulary matching.
[0169] The matching degree is used to represent the similarity between candidate content and requirement information across multiple dimensions, including document hierarchy, structural path location, and lexical matching. The matching degree can be expressed as a percentage. A higher matching degree indicates a greater similarity between the candidate content and requirement information in terms of textual surface features and / or structural features; a lower matching degree indicates a lower similarity between the candidate content and requirement information in terms of textual surface features and / or structural features.
[0170] In some embodiments, the semantic path covers issues such as open expressions, ambiguous terms, and implicit intentions, while the structural path strengthens the modeling of logical relationships, terminology dependencies, and contextual constraints in process standards. The two complement each other, overcoming the limitation of traditional retrieval methods in recalling key information when faced with complex expressions in industrial contexts. This ensures that the obtained target content not only includes paragraph text but also metadata such as source documents, heading levels, and associated graph nodes, providing structural contextual support for the output of the target content.
[0171] In some embodiments, after determining the semantic relevance information and matching degree between each of the at least one candidate content and the demand information, the server determines the candidate content whose semantic relevance degree indicated by the semantic relevance information is greater than or equal to a preset degree and whose matching degree is greater than or equal to a preset matching degree as at least one target candidate content.
[0172] Among these, a pre-trained large model can be, for example, a pre-trained large language model. It's important to note that this pre-trained large model is trained based on industry knowledge, enabling it to learn the language habits, knowledge systems, and task logic of the industry domain. Therefore, the pre-trained large model possesses the ability to understand industry-specific knowledge, handle specialized tasks, and generate professional knowledge.
[0173] It should be noted that the pre-trained large model for generating the target content and the pre-trained large model for constructing the knowledge graph can be the same large model or different large models.
[0174] The generated prompt text is used to instruct the pre-trained large model to process at least one target candidate content, generating a question containing target content relevant to the demand information. Furthermore, the generated prompt text also serves as a constraint and guide for the pre-trained large model in generating target content, ensuring that the generated target content conforms to professional knowledge. In some embodiments, the generated prompt text can be determined through prompt word engineering.
[0175] In some embodiments, the server may input the generated prompt text and at least one target candidate content into a pre-trained large model, and the pre-trained large model processes the at least one target candidate content according to the generated prompt text to generate the target content.
[0176] It should be noted that the number of target content items can be one or multiple. When there are multiple target content items, they are categorized according to their structural paths, ensuring that the output content of the pre-trained large model is not only locally relevant but also possesses cross-paragraph semantic continuity. To further control information density, the pre-trained large model also incorporates a "redundancy compression mechanism," merging and pruning paragraphs with semantic repetition or highly overlapping content to ensure that the output target content meets information coverage requirements while remaining within the contextual constraints of the pre-trained large model.
[0177] exist Figure 5 In the illustrated embodiment, the paragraph or chapter content that is related to the required information is determined by combining semantic vectors and knowledge graphs. This effectively supports the handling of complex semantic problems across chapters, specifications, and documents, and improves the system's knowledge coverage and reasoning capabilities in heterogeneous industrial documents.
[0178] Based on the above embodiments, the following is combined with Figure 6 The industrial knowledge generation method that integrates large-scale models and knowledge graphs provided in this application is further explained.
[0179] Figure 6 This is a schematic diagram illustrating an industrial knowledge generation method that integrates large models and knowledge graphs, as provided in an embodiment of this application. Figure 6As shown, the server can input industrial knowledge documents and prompt text into a pre-trained large model. After processing the industrial knowledge documents and prompt text, the pre-trained large model obtains the knowledge graph corresponding to the industrial knowledge documents. Then, when it is necessary to query the knowledge content of a specific task from the industrial knowledge documents, the server sends the request information to the server. The server inputs the request information and the knowledge graph into the pre-trained large model, which processes the request information and the knowledge graph to obtain the target content.
[0180] Specifically, the process by which the pre-trained large model processes demand information and knowledge graphs to obtain the target content can be found in [reference needed]. Figure 7 . Figure 7 This is a schematic diagram illustrating the determination of target content as provided in an embodiment of this application. Figure 7 As shown, after receiving the demand information and knowledge graph, the pre-trained large model processes the demand information according to the semantic transformation model to obtain the semantic vector corresponding to the demand information. It then processes multiple paragraph contents separately according to the semantic transformation model to obtain the semantic vector corresponding to each paragraph content. Finally, for each paragraph content, based on the semantic vector corresponding to the demand information and the semantic vector corresponding to the paragraph content, it determines the semantic similarity between the demand information and the paragraph content. Then, at least one paragraph content with a semantic similarity greater than the preset similarity is identified as at least one paragraph content.
[0181] In addition, the pre-trained large model also processes the demand information, determines at least one keyword of the demand information, and performs matching processing with the knowledge graph based on the at least one keyword to determine at least one chapter content.
[0182] In some embodiments, after determining at least one chapter and at least one paragraph, a pre-trained large model performs semantic analysis, knowledge reorganization, and content generation on the at least one chapter and at least one paragraph based on knowledge graphs, original information tags, and selected paragraphs. Specifically, the pre-trained large model determines the target content based on demand information, at least one chapter, at least one paragraph, and the relationships between multiple knowledge nodes. Simultaneously, it combines prompt text to guide the pre-trained large model to output the target content in accordance with industry-standard language.
[0183] In some embodiments, to improve response accuracy and content consistency, this application designs a structure guidance mechanism that automatically constructs a "content framework sketch" before outputting the target content. This sketch includes elements such as the process structure, control parameters, and numbering references that the target content should contain, and uses this as the logical anchor point for generating the target content. Simultaneously, a terminology standardization component is introduced to automatically proofread and standardize the professional terms in the generated target content, avoiding terminology misuse and format inconsistencies. To reduce the probability of illusions, this application also introduces a knowledge consistency verification module, which automatically compares the entity coverage and terminology consistency between the target content and the original graph nodes and recalled paragraphs, providing feedback and correction for obvious conflicts or inconsistencies. Finally, the generated target content is output in paragraph or document form, and the output target content possesses characteristics such as semantic completeness, standardized structure, accurate terminology, and clear standard references. It can be directly used in practical application scenarios such as process planning assistance, document generation, and work instruction writing, achieving a reliable and automated response to complex industrial knowledge.
[0184] Then, deduplication, semantic relevance analysis, and matching degree analysis are performed on at least one chapter and at least one paragraph to determine the target content.
[0185] exist Figure 6 In the illustrated embodiments, by introducing structure guidance, prompting engineering, and meta-information control mechanisms, this invention can effectively constrain the generation behavior of large models, significantly reducing problems such as content duplication, logical jumps, and terminology errors, and improving the standardization and engineering usability of the generated results. The industrial knowledge generation method integrating large models and knowledge graphs provided in this application embodiment achieves a closed-loop process from industrial knowledge documents to structured knowledge graphs and then to standardized process content generation. This not only improves the query response speed and accuracy of industrial knowledge but also enhances the automation level and intelligence capabilities of the knowledge service system.
[0186] Figure 8 This is a schematic diagram of an industrial knowledge generation device that integrates a large model and a knowledge graph, provided as an embodiment of this application. Figure 8 As shown, the device includes a receiving module 81, a first determining module 82, a processing module 83, and a generating module 84, wherein:
[0187] The receiving module 81 is used to receive the demand information sent by the client. The demand information includes at least one keyword and is used to instruct the generation of document content related to the at least one keyword based on the industrial knowledge document. The industrial knowledge document includes multiple paragraphs.
[0188] The first determining module 82 is used to determine at least one target paragraph content associated with the requirement information from multiple paragraph contents;
[0189] The processing module 83 is used to match at least one first keyword with the knowledge graph corresponding to the industrial knowledge document, and to determine at least one chapter content in the industrial knowledge document; wherein, the industrial knowledge document includes multiple chapter titles and the chapter content corresponding to each of the multiple chapter titles, and the knowledge graph includes multiple knowledge nodes and the relationship between the multiple knowledge nodes;
[0190] The generation module 84 is used to generate target content based on the requirement information, at least one target paragraph content, and at least one chapter content; wherein, the target content is content in the industrial knowledge document related to at least one keyword.
[0191] In one possible implementation, the first determining module 82 is specifically used for:
[0192] The semantic transformation model is used to process the demand information to obtain the semantic vector corresponding to the demand information.
[0193] The semantic transformation model is used to process the content of multiple paragraphs separately, resulting in semantic vectors corresponding to each paragraph.
[0194] For each paragraph in multiple paragraphs, the semantic similarity between the requirement information and the paragraph content is determined based on the semantic vector corresponding to the requirement information and the semantic vector corresponding to the paragraph content.
[0195] Based on the semantic similarity between the demand information and the content of multiple paragraphs, at least one target paragraph is identified among the multiple paragraphs.
[0196] In one possible implementation, the processing module 83 is specifically used for:
[0197] Based on at least one primary keyword, multiple knowledge nodes, and the relationships between the multiple knowledge nodes, determine at least one candidate node that is related to at least one primary keyword among the multiple knowledge nodes;
[0198] Based on at least one candidate node and at least one first keyword, determine at least one chapter of content in the industrial knowledge document.
[0199] In one possible implementation, the processing module 83 is specifically used for:
[0200] For each primary keyword in at least one primary keyword, perform the following steps:
[0201] For each candidate node in at least one candidate node, determine the similarity between the candidate node and the first keyword;
[0202] The chapter content corresponding to the first keyword is determined based on the similarity between each candidate node and the first keyword.
[0203] Among them, at least one chapter contains the chapter content corresponding to at least one primary keyword.
[0204] In one possible implementation, the industrial knowledge generation apparatus integrating large models and knowledge graphs further includes a second determining module, wherein the second determining module is used for:
[0205] Determine the prompt text, which includes questions and prompt words for identifying the second keyword of an industrial knowledge document to a pre-trained large model;
[0206] The prompt text, multiple chapter titles, and the chapter content corresponding to each chapter title are input into a pre-trained large model to obtain multiple secondary keywords for industrial knowledge documents;
[0207] Based on multiple secondary keywords, multiple chapter titles, and the chapter content corresponding to each chapter title, determine multiple knowledge nodes and the relationships between these knowledge nodes.
[0208] In one possible implementation, the generation module 84 is specifically used for:
[0209] Deduplication is performed on at least one paragraph and at least one chapter to obtain at least one candidate content;
[0210] Generate target content based on the demand information and at least one candidate content.
[0211] In one possible implementation, the generation module 84 is specifically used for:
[0212] For each candidate content in at least one candidate content, the candidate content and demand information are input into the relevance analysis model to obtain the semantic relevance information between the candidate content and the demand information;
[0213] Based on the semantic relevance information between each of the at least one candidate content and the demand information, at least one target candidate content is determined from the at least one candidate content.
[0214] Determine the generated prompt text, which includes posing questions and prompts to the pre-trained large model to determine the target content to be generated;
[0215] The generated prompt text and at least one target candidate content are input into a pre-trained large model to obtain the target content.
[0216] The industrial knowledge generation device 80 that integrates large models and knowledge graphs provided in this application embodiment can execute the industrial knowledge generation method that integrates large models and knowledge graphs provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described again here.
[0217] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 9 As shown, the electronic device 90 provided in this application embodiment includes: a memory 91 and a processor 92;
[0218] Memory 91 stores instructions executed by the computer;
[0219] The processor 92 executes the computer execution instructions stored in the memory 91, causing the processor 92 to execute the industrial knowledge generation method that integrates large models and knowledge graphs provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0220] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0221] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0222] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0223] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described industrial knowledge generation method that integrates large models and knowledge graphs.
[0224] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the aforementioned industrial knowledge generation method that integrates large models and knowledge graphs.
[0225] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0226] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0227] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0228] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0229] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0230] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0231] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0232] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. An industrial knowledge generation method fusing a large model and a knowledge graph, characterized in that, The method comprises the following steps: receiving demand information sent by a client, the demand information including at least one keyword, and the demand information being used to indicate that document content related to the at least one keyword is generated according to an industrial knowledge document; the industrial knowledge document including a plurality of paragraph contents; determining at least one target paragraph content associated with the demand information from the plurality of paragraph contents; matching the at least one first keyword with a knowledge graph corresponding to the industrial knowledge document to determine at least one chapter content in the industrial knowledge document; the industrial knowledge document including a plurality of chapter titles and chapter contents corresponding to the plurality of chapter titles respectively, and the knowledge graph including a plurality of knowledge nodes and association relationships between the plurality of knowledge nodes; performing deduplication processing on the at least one target paragraph content and the at least one chapter content to obtain at least one candidate content; for each candidate content in the at least one candidate content, inputting the candidate content and the demand information into a relevance analysis model to obtain semantic relevance information between the candidate content and the demand information; determining at least one target candidate content from the at least one candidate content according to the semantic relevance information between each of the at least one candidate content and the demand information; determining a generation prompt text, the generation prompt text including a question and a prompt word for a pre-trained large model to determine a target content; inputting the generation prompt text and the at least one target candidate content into the pre-trained large model to obtain the target content; wherein the target content is content related to the at least one keyword in the industrial knowledge document; The method further comprises: determining a prompt text, the prompt text including a question and a prompt word for a pre-trained large model to determine a second keyword of the industrial knowledge document; inputting the prompt text, the plurality of chapter titles and chapter contents corresponding to the plurality of chapter titles into the pre-trained large model to obtain a plurality of second keywords of the industrial knowledge document; determining the plurality of knowledge nodes and association relationships between the plurality of knowledge nodes according to the plurality of second keywords, the plurality of chapter titles and chapter contents corresponding to the plurality of chapter titles respectively.
2. The method of claim 1, wherein, The method further comprises: processing the demand information according to a semantic conversion model to obtain a semantic vector corresponding to the demand information; processing the plurality of paragraph contents according to the semantic conversion model respectively to obtain semantic vectors corresponding to the plurality of paragraph contents respectively; for each paragraph content in the plurality of paragraph contents, determining a semantic similarity between the demand information and the paragraph content based on the semantic vector corresponding to the demand information and the semantic vector corresponding to the paragraph content; determining the at least one target paragraph content from the plurality of paragraph contents according to the semantic similarity between the demand information and each of the plurality of paragraph contents respectively.
3. The method according to claim 1 or 2, characterized in that, The matching processing of the at least one first keyword and the knowledge graph corresponding to the industrial knowledge document comprises: According to the at least one first keyword, the plurality of knowledge nodes, and the association relationship between the plurality of knowledge nodes, at least one candidate node related to the at least one first keyword is determined in the plurality of knowledge nodes; According to the at least one candidate node and the at least one first keyword, the at least one chapter content is determined in the industrial knowledge document.
4. The method of claim 3, wherein, According to the at least one candidate node and the at least one first keyword, the at least one chapter content is determined in the industrial knowledge document. For each first keyword in the at least one first keyword, the following steps are performed: For each candidate node in the at least one candidate node, the similarity between the candidate node and the first keyword is determined; According to the similarity between each of the at least one candidate node and the first keyword, the chapter content corresponding to the first keyword is determined; The at least one chapter content comprises the chapter content corresponding to each of the at least one first keyword.
5. An industrial knowledge generation device that fuses a large model and a knowledge graph, characterized by, Comprise: The receiving module is used for receiving the demand information sent by the client, and the demand information comprises at least one keyword, and the demand information is used to indicate that the document content related to the at least one keyword is generated according to the industrial knowledge document; the industrial knowledge document comprises a plurality of paragraph contents; The first determination module is used for determining at least one target paragraph content associated with the demand information in the plurality of paragraph contents; The processing module is used for matching processing of the at least one first keyword and the knowledge graph corresponding to the industrial knowledge document, and determining at least one chapter content in the industrial knowledge document; wherein the industrial knowledge document comprises a plurality of chapter titles and the chapter content corresponding to each of the plurality of chapter titles, and the knowledge graph comprises a plurality of knowledge nodes and the association relationship between the plurality of knowledge nodes; The generation module is used for generating target content according to the demand information, the at least one target paragraph content and the at least one chapter content; wherein the target content is the content related to the at least one keyword in the industrial knowledge document; The generation module is specifically used for performing deduplication processing on the at least one target paragraph content and the at least one chapter content to obtain at least one candidate content; For each candidate content in the at least one candidate content, the candidate content and the demand information are input into a relevance analysis model to obtain semantic relevance information between the candidate content and the demand information; According to the semantic relevance information between each of the at least one candidate content and the demand information, at least one target candidate content is determined in the at least one candidate content; The generation prompt text is determined, and the generation prompt text comprises a question and a prompt word for a pre-trained large model to determine the generation of target content; inputting the generated prompt text and the at least one target candidate content into the pre-trained large model to obtain the target content; a second determination module, configured to determine a prompt text, the prompt text including a question of asking the pre-trained large model to determine a second keyword of the industrial knowledge document and a prompt word; inputting the prompt text, the plurality of chapter titles and the chapter content corresponding to each of the plurality of chapter titles into the pre-trained large model to obtain a plurality of second keywords of the industrial knowledge document; determining the plurality of knowledge nodes and the association relationship between the plurality of knowledge nodes according to the plurality of second keywords, the plurality of chapter titles and the chapter content corresponding to each of the plurality of chapter titles.
6. An electronic device, comprising: comprising: a memory, a processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory, so that the processor executes the method in any one of claims 1-4.
7. A computer readable storage medium characterized in that, the computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the method in any one of claims 1-4.
Citation Information
Patent Citations
Retrieval generation method and device based on large language model and knowledge graph
CN119848168A
Retrieval enhancement generation method based on document knowledge base and knowledge graph
CN120780849A