An industrial formula and process knowledge answer generation method and device

By cleaning industrial formula data, identifying named entities and extracting relationships, forming a domain knowledge graph, and using a large language model to generate answers containing self-certified information, the problems of high construction and maintenance costs, lack of intelligence and low retrieval efficiency in the existing technology are solved, and efficient and intelligent industrial formula and process knowledge management and application are achieved.

CN119494402BActive Publication Date: 2025-05-30SHENYANG INST OF AUTOMATION GUANGZHOU CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411493804.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-05-30
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

The existing technology has problems such as high construction and maintenance costs, lack of intelligence and low retrieval efficiency in the management and application of industrial formula and process knowledge.

Method used

By cleaning and preprocessing industrial formula data and process knowledge data, naming entity recognition, relationship extraction and data fusion, forming a domain knowledge graph, and establishing an inverted index to improve query efficiency. Combining user input questions and optimal subgraphs, a large language model is used to generate answers containing self-certified information.

Benefits of technology

It reduces the construction and maintenance costs of the knowledge graph, improves the system's retrieval efficiency, provides more intelligent and efficient services, and the generated answers include decision suggestions and self-certified information, which improves the credibility and explanatory nature of the answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119494402B_ABST
    Figure CN119494402B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for generating answers to industrial formula and process knowledge. The method includes: forming a domain knowledge graph for industrial formula and process knowledge according to an entity set and a relationship set; establishing an inverted index based on the node attributes and edge attributes of the domain knowledge graph; encoding and retrieving the nodes and edges of the domain knowledge graph according to the inverted index and a query embedding vector generated from a user input question to obtain an optimal subgraph; and inputting the user input question after text encoding and the optimal subgraph after graph encoding into a preset large language model to generate an answer with self-certifying information. By adopting the present invention, automated data processing and knowledge graph construction reduce the need for manual intervention and lower the construction and maintenance costs; the introduction of an inverted index and an efficient retrieval algorithm improves the query speed and ensures that the system can quickly respond to user needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial data processing, and particularly to a method and device for generating answers to industrial formulas and process knowledge. Background Art

[0002] Existing technologies have their own advantages and disadvantages in the management and application of private formulas and process knowledge in various industries. Traditional database systems are stable but lack intelligence; expert systems have certain reasoning capabilities but are inflexible; knowledge graphs can handle complex relationships but have high construction and maintenance costs. Although the above three main existing technologies have to some extent solved the problems of managing and applying private industrial formulas and process knowledge in the textile industry, their flexibility and intelligence are very limited, and there are still many deficiencies:

[0003] (1) High construction and maintenance costs: Traditional databases and expert systems rely on manual maintenance, and the construction and update of knowledge graphs also require a large amount of human input. At the same time, as time goes by, the amount of data increases exponentially. Especially in the face of more and more unstructured and semi-structured data, maintaining and updating databases and knowledge graphs will become more and more complex.

[0004] (2) Lack of intelligence: The system can only operate according to predefined structures and query statements, lacking flexible knowledge reasoning and automated processing capabilities. Once the rules are defined, it is difficult to make dynamic adjustments. It can be seen that these technologies have limited capabilities in knowledge reasoning and automated processing and are difficult to meet dynamic and complex process requirements.

[0005] (3) Low retrieval efficiency: In the face of complex queries and large-scale graphs, the processing capabilities of existing technologies are limited, and the retrieval and reasoning efficiency are relatively low, making it difficult to respond to user needs in real time. Summary of the Invention

[0006] Embodiments of the present invention provide a method and device for generating answers to industrial formulas and process knowledge, which reduce the construction and maintenance costs of industry knowledge graphs, improve the retrieval efficiency of the system, and provide more intelligent and efficient services for users.

[0007] To achieve the above object, the first aspect of the embodiments of the present application provides a method for generating answers to industrial formulas and process knowledge, including:

[0008] Performing data cleaning and preprocessing on industrial formula data and process knowledge data to obtain standardized data;

[0009] Performing named entity recognition, relationship extraction, and data fusion on the standardized data to obtain an entity set and a relationship set;

[0010] Form a domain knowledge graph for industrial formula and process knowledge based on the set of entities and the set of relationships;

[0011] Establish an inverted index according to the node attributes and edge attributes of the domain knowledge graph;

[0012] Encode and retrieve the nodes and edges of the domain knowledge graph according to the inverted index and the query embedding vector generated from the user input question to obtain the optimal subgraph;

[0013] Input the user input question after text encoding and the optimal subgraph after graph encoding into a preset large language model to generate an answer containing self-certifying information.

[0014] In a possible implementation manner of the first aspect, data cleaning and preprocessing are performed on the industrial formula data and process knowledge data to obtain standardized data, specifically including:

[0015] Based on the unique identifier, delete the duplicate items in the industrial formula data and process knowledge data;

[0016] Fill in the missing values in the industrial formula data and process knowledge data through linear interpolation;

[0017] Standardize the industrial formula data and process knowledge data according to the mean and standard deviation of the industrial formula data and process knowledge data.

[0018] In a possible implementation manner of the first aspect, named entity recognition, relationship extraction, and data fusion are performed on the standardized data to obtain a set of entities and a set of relationships, specifically including:

[0019] Use the RoBERTa model to identify the entities in the standardized data;

[0020] Use a pre-trained language model combined with a bidirectional long short-term memory network to identify the relationships between the entities;

[0021] Fuse two entities with a similarity greater than the matching threshold to obtain a set of entities and a set of relationships.

[0022] In a possible implementation manner of the first aspect, the formation of a domain knowledge graph for industrial formula and process knowledge according to the set of entities and the set of relationships specifically includes:

[0023] Fill the set of entities and the set of relationships into preset different entity relationships, use the entities in the different entity relationships as nodes, and the relationships in the different entity relationships as edges to import into a graph database to form a domain knowledge graph for industrial formula and process knowledge.

[0024] In a possible implementation of the first aspect, establishing an inverted index according to the node attributes and edge attributes of the domain knowledge graph specifically includes:

[0025] Establishing a number of node attribute fields according to the node attributes of the domain knowledge graph;

[0026] Establishing a number of edge attribute fields according to the edge attributes of the domain knowledge graph;

[0027] Recording the node numbers of all the node attribute fields and the edge numbers of all the edge attribute fields, and establishing an inverted index.

[0028] In a possible implementation of the first aspect, before encoding and retrieving the nodes and edges of the domain knowledge graph according to the inverted index and the query embedding vector generated from the user input question to obtain the optimal subgraph, it further includes:

[0029] Inputting the user input question into an application language model, and applying a preset encoding strategy to generate a query embedding vector.

[0030] In a possible implementation of the first aspect, encoding and retrieving the nodes and edges of the domain knowledge graph according to the inverted index and the query embedding vector generated from the user input question to obtain the optimal subgraph specifically includes:

[0031] Applying the preset encoding strategy to encode the nodes and edges of the domain knowledge graph respectively;

[0032] Using the k-nearest neighbor retrieval method to identify the nodes of the domain knowledge graph with the greatest similarity to the query embedding vector as the most relevant nodes, and identifying the edges of the domain knowledge graph with the greatest similarity to the query embedding vector as the most relevant edges.

[0033] In a possible implementation of the first aspect, inputting the user input question after text encoding and the optimal subgraph after graph encoding into a preset large language model to generate an answer containing self-certifying information specifically includes:

[0034] Performing text encoding on the user input question through a text embedder to obtain a query input;

[0035] Aligning the optimal subgraph with the vector space of the preset large language model through a graph encoder containing a multi-layer perceptron to obtain a graph input;

[0036] Inputting the query input and the graph input into the preset large language model together to generate an answer containing decision suggestions and self-certifying information; the self-certifying information is the explanatory information of the decision suggestions.

[0037] In a possible implementation manner of the first aspect, the preset large language model is an open source large language model or a commercial large language model that introduces an enhanced attention mechanism.

[0038] A second aspect of an embodiment of the present application provides an industrial formula and process knowledge answer generation device, comprising:

[0039] The standardization module is used to clean and preprocess the industrial formula data and process knowledge data to obtain standardized data;

[0040] A fusion module, used for performing named entity recognition, relationship extraction and data fusion on the standardized data to obtain an entity set and a relationship set;

[0041] A graph module, used to form a domain knowledge graph for industrial formula and process knowledge based on the entity set and the relationship set;

[0042] Establishing a module for establishing an inverted index according to the node attributes and edge attributes of the domain knowledge graph;

[0043] A subgraph module, used to encode and retrieve nodes and edges of the domain knowledge graph according to the inverted index and the query embedding vector generated by the user input question to obtain an optimal subgraph;

[0044] A generation module is used to input the user input question after text encoding and the optimal subgraph after graph encoding into a preset large language model to generate an answer containing self-verification information.

[0045] Compared with the prior art, the method and device for generating answers to industrial formula and process knowledge provided by the embodiment of the present invention ensure the high quality of data through data cleaning and preprocessing; through named entity recognition and relationship extraction, the system can accurately identify and extract key entities in the data and the relationships between them, providing structured information for building a knowledge graph, and then converting the entities and relationships into nodes and edges in the graph database, forming a structured knowledge graph, which is convenient for the management and query of complex relationships. Then, by establishing an inverted index, the system can quickly locate relevant nodes and edges in large-scale data, significantly improving the query speed. Finally, combined with the user input question and the optimal subgraph, the answer generated by the large language model not only contains decision suggestions, but also comes with self-certification information, which improves the credibility and interpretability of the answer. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a flowchart of a method for generating industrial formula and process knowledge answers according to an embodiment of the present invention;

[0047] Figure 2It is a schematic diagram of the acquisition process of the optimal subgraph provided by an embodiment of the present invention;

[0048] Figure 3 It is a schematic diagram of the generation process of an answer containing self-certifying information provided by an embodiment of the present invention. Specific embodiments

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] To solve the above problems, please refer to Figure 1 , an embodiment of the present invention provides an industrial formula and process knowledge answer generation method, including:

[0051] S10. Clean and preprocess the industrial formula data and process knowledge data to obtain standardized data.

[0052] S11. Perform named entity recognition, relationship extraction, and data fusion on the standardized data to obtain an entity set and a relationship set.

[0053] S12. Form a domain knowledge graph for industrial formula and process knowledge according to the entity set and the relationship set.

[0054] S13. Establish an inverted index according to the node attributes and edge attributes of the domain knowledge graph.

[0055] S14. Encode and retrieve the nodes and edges of the domain knowledge graph according to the inverted index and the query embedding vector generated from the user input question to obtain the optimal subgraph.

[0056] S15. Input the user input question after text encoding and the optimal subgraph after graph encoding into a preset large language model to generate an answer containing self-certifying information.

[0057] The above method is essentially a method for managing and utilizing private industrial formula and process knowledge based on a large language model and a domain knowledge graph. Through the RAG (Retrieval Augmented Generation) process, it realizes the comprehensive cleaning of data, the effective extraction and fusion of knowledge, the construction and indexing of the graph, as well as the efficient generation and explanation of answers:

[0058] The standardization process in S10 ensures the integrity and consistency of the data, providing a high-quality data foundation for subsequent processing. In S11, through named entity recognition and relationship extraction, the system can accurately identify and extract key entities in the data and the relationships between them, providing structured information for constructing the knowledge graph. Data fusion reduces redundant entities and improves the degree of data integration, making the knowledge graph more concise and efficient.

[0059] S12 converts entities and relationships into nodes and edges in the graph database, forming a structured knowledge graph, which is convenient for the management and query of complex relationships. Due to the visualization characteristics of the knowledge graph, the data relationships are more intuitive, which is convenient for management and analysis. In S13, by establishing an inverted index, the system can quickly locate relevant nodes and edges in large-scale data, significantly improving the query speed. The inverted index optimizes the retrieval path, enabling complex queries to be completed efficiently and enhancing the system's response speed. S14 generates query embedding vectors and performs similarity calculations, enabling the system to accurately match the questions input by users, improving the accuracy and relevance of retrieval.

[0060] S15 combines the user's input question and the optimal subgraph. The answers generated by the large language model not only contain decision-making suggestions but also come with self-certifying information, improving the credibility and interpretability of the answers. The generated answers are more accurate and comprehensive, allowing users to obtain the required information faster and enhancing the overall user experience of the system.

[0061] The above method can not only efficiently manage and apply industrial formula and process knowledge but also significantly improve the performance of the management system in the following aspects: First, through data cleaning and preprocessing, the high quality of the data is ensured; second, the structuring of the knowledge graph forms a structured and visual knowledge graph, which is convenient for management and query; third, the query efficiency is improved, and the query speed is significantly increased through the inverted index and efficient retrieval algorithms; fourth, the level of intelligence is improved, and intelligent answers are generated using the large language model, enhancing the intelligence of the system and the user experience. These improvements are not only applicable to the textile industry but can also be extended to other industries that need to manage and apply complex knowledge, helping industrial enterprises enhance their competitiveness and innovation capabilities.

[0062] Exemplarily, the industrial formula data and process knowledge data are subjected to data cleaning and preprocessing to obtain standardized data, which specifically includes:

[0063] Based on the unique identifier, duplicate items in the industrial formula data and process knowledge data are deleted;

[0064] Missing values in the industrial formula data and process knowledge data are filled by linear interpolation;

[0065] Standardize the industrial formula data and process knowledge data according to the means and standard deviations of the industrial formula data and process knowledge data.

[0066] Data cleaning and preprocessing is a crucial step in data analysis. Its purpose is to remove noise and inconsistencies in the data, which is completed through three steps: deleting redundant information, filling in missing data, and data standardization, in order to improve data quality and model accuracy. After manually deleting duplicates based on a unique identifier (such as ID), it is also necessary to perform missing value processing measurement and data standardization on the data. Missing value processing refers to using appropriate methods to fill in missing values when they exist in the dataset. In this solution, missing values will be filled in by linear interpolation, and the linear interpolation method is implemented as follows:

[0067]

[0068] where x i is the position of the missing value, x i-1 is the data value at the position before the missing value, and x i+1 is the data value at the position after the missing value. Data standardization is to convert the data to the same scale to eliminate the influence of the dimension between different features, and the data is standardized by the following formula:

[0069]

[0070] where x is the original data value, μ is the mean of the dataset, σ is the standard deviation of the dataset, and x′ is the standardized data value.

[0071] Exemplarily, performing named entity recognition, relation extraction, and data fusion on the standardized data to obtain an entity set and a relation set specifically includes:

[0072] Using the RoBERTa model to identify entities in the standardized data;

[0073] Using a pre-trained language model combined with a bidirectional long short-term memory network to identify the relationships between the entities;

[0074] Fusing two entities with a similarity greater than the matching threshold to obtain an entity set and a relation set.

[0075] The above steps are used for knowledge extraction and fusion, aiming to extract key information from unstructured text data and fuse this information with other data sources to form a unified knowledge representation. It mainly includes three steps: named entity recognition, relation extraction, and data fusion.

[0076] Named entities can be recognized using the RoBERTa model to identify entities in text, such as chemical components, process steps, equipment names, etc. RoBERTa is an improved version of BERT, aiming to improve the performance of the model by optimizing the pre-training process. Its main improvements include: longer pre-training, larger batch sizes, removing the Next Sentence Prediction (NSP) task, and better hyperparameter tuning. The self-attention mechanism of RoBERTa is as follows:

[0077]

[0078] where Q, K, and V are linear transformations of the input, and d k is the dimension of the key vector. For the input text T, the RoBERTa model outputs the entity set E:

[0079] E = {(e i , t i ) | i = 1, 2,..., n} (4)

[0080] where e i is the i-th entity recognized, t i is the type of the entity, and n is the total number of entities.

[0081] Relation extraction can use a pre-trained language model (such as BERT) combined with a bidirectional long short-term memory network (BiLSTM) to identify the relationships between entities. For example, "the combination of component A and component B" or "step X is the pre-step of step Y". Among them, the attention mechanism of BERT is the same as formula (c), and BiLSTM is a bidirectional long short-term memory network that improves the ability to model sequence data by considering both forward and backward context information. The main features of BiLSTM include: bidirectional information flow and long short-term memory. To capture the context information of the sequence. The BiLSTM cell is as follows:

[0082] f t = σ(W f · [h t-1 , x t + b f ) (5)

[0083] i t = σ(W i · [h t-1 , x t + b i ) (6)

[0084]

[0085] o t = σ(W o · [h t-1, x t + b o ) (9)

[0086] h t = o t *tanh(C t ) (10)

[0087] Among them, f t is the forget gate, i t is the input gate, is the candidate memory, C t is the memory cell, o t is the output gate, h t is the hidden state, x t is the input, W and b are weights and biases, σ is the sigmoid function, and tanh is the hyperbolic tangent function. After processing the input text T, BERT combines with the BiLSTM model to output the relationship set R:

[0088] R = {(e i , e j , r ij ) | i, j = 1, 2,..., n} (11)

[0089] Among them, e i and e j are the identified entities, and r ij is the relationship between the entities e i and e j . Data fusion is to fuse the knowledge extracted from the text with other data sources (such as databases, experimental data) to form a unified knowledge representation. According to the following matching rules, the same or similar entities in different data sources are merged.

[0090]

[0091] Among them, (e i , e j ) is the similarity between the entities e i and e j , and threshold is the matching threshold. During the matching process, two entities with a similarity greater than the matching threshold are regarded as the same and merged.

[0092] Through the above steps, key entities can be identified from the standardized data, relationships between entities can be extracted, and data fusion can be performed through similarity calculation, finally generating an entity set and a relationship set. These sets provide important basic data for the subsequent construction of the domain knowledge graph.

[0093] Exemplarily, according to the entity set and the relationship set, a domain knowledge graph for industrial formula and process knowledge is formed, specifically including:

[0094] Fill the entity set and the relationship set into preset different entity relationships, use the entities in the different entity relationships as nodes, and use the relationships in the different entity relationships as edges to import into a graph database, so as to form a domain knowledge graph for industrial formula and process knowledge.

[0095] The goal of constructing a domain knowledge graph is to organize the extracted entities and relationships into a coherent graph structure for storing and retrieving domain-specific knowledge. Determine entity types according to private formula and process knowledge, such as chemical components, process steps, equipment names, etc., and further define relationship types between entities, such as "combination", "preceding step", etc. Subsequently, fill the entity set and relationship set after data fusion into various entities and relationships to form a domain knowledge graph for private formula and process knowledge. Finally, import the node (entity) and edge (relationship) data into the graph database Neo4j to form a complete knowledge graph for efficient retrieval and query.

[0096] When importing, each entity in the entity set is imported as a node into the graph database. Each node contains attributes of the entity, such as name, type, description, etc.; each relationship in the relationship set is imported as an edge into the graph database. Each edge connects two nodes and contains a relationship type and other attributes. A complete knowledge graph is constructed in the graph database to ensure that all nodes and edges are correctly connected.

[0097] By the above steps, the entity set and the relationship set are imported into the graph database, forming a domain knowledge graph for industrial formula and process knowledge. This graph structurally represents the key entities and their relationships in industrial formula and process knowledge, providing strong support for subsequent queries, analyses, and applications.

[0098] Exemplarily, according to the node attributes and edge attributes of the domain knowledge graph, an inverted index is established, specifically including:

[0099] According to the node attributes of the domain knowledge graph, a number of node attribute fields are established;

[0100] According to the edge attributes of the domain knowledge graph, a number of edge attribute fields are established;

[0101] Record the node numbers of all the node attribute fields and the edge numbers of all the edge attribute fields to establish an inverted index.

[0102] The purpose of the graph index is to improve the query and retrieval efficiency of the knowledge graph. By establishing an index, specific nodes or relationships can be quickly located, supporting complex query operations. This step creates an inverted index for the attributes of nodes (entities) and edges (relationships) in the domain knowledge graph to support fast retrieval. First, select the node and edge attribute fields for which to establish the index (such as node names, types, edge types, etc.); second, create an inverted index for each index field, recording the node or edge serial numbers corresponding to each attribute value. Specifically, the inverted index structure is as follows:

[0103] I = {(term, {id 1 , id 2 , …, id n}) ∣ term ∈ attributes, id i ∈ nodes or edges}(13)

[0104] where I is the set of inverted indexes, term is the attribute value of the index field, id i is the corresponding node or edge serial number, attributes is the set of all attributes, and nodes or edges is the set of all entities and relationships.

[0105] If you want to find all nodes of type "raw material", you can quickly find the ID list of these nodes through the inverted index. If you want to find all edges of type "use", you can also quickly find the ID list of these edges through the inverted index. By establishing an inverted index, the search efficiency for node and edge attributes in the domain knowledge graph can be greatly improved. This method is particularly suitable for scenarios with diverse attribute values and frequent queries, which can significantly reduce query time and improve the user experience.

[0106] See Figure 2 , for example, before encoding and retrieving the nodes and edges of the domain knowledge graph according to the inverted index and the query embedding vector generated from the user input question to obtain the optimal subgraph, it further includes:

[0107] Input the user input question into the application language model, and apply a preset encoding strategy to generate a query embedding vector.

[0108] For the text attribute Z n of node n n , apply the language model LM to obtain the representation Z n = LM(x d ) ∈ R q . For a random formulation problem q given in this field by a technician in the textile industry, apply the same encoding strategy to generate the query embedding vector Z q = LM(x d。

[0109] Exemplarily, encoding and retrieving the nodes and edges of the domain knowledge graph according to the inverted index and the query embedding vector generated from the user input question to obtain an optimal subgraph specifically includes:

[0110] Applying the preset encoding strategy to encode the nodes and edges of the domain knowledge graph respectively;

[0111] Using the k-nearest neighbor retrieval method, identifying the nodes of the domain knowledge graph with the highest similarity to the query embedding vector as the most relevant nodes, and identifying the edges of the domain knowledge graph with the highest similarity to the query embedding vector as the most relevant edges.

[0112] The above-mentioned encoding and retrieval of the nodes and edges of the domain knowledge graph based on the cross-domain inverted index established above aims to accurately locate relevant subgraph and node information in specific decision-making problems and scenarios

[0113] The k-nearest neighbor retrieval method can be used to identify the most relevant nodes and edges based on the similarity between the query and each node or edge. Subsequently, it is planned to study graph refinement and embedding, aiming to construct a subgraph containing as many relevant nodes and edges as possible while keeping the size of the graph within a manageable range. Finally, a reward algorithm can be planned to assign higher reward values to the nodes (V k ) and edges (E k ) most relevant to the query, and identify an optimal-sized and relevant subgraph S * . By optimizing the total reward value of the nodes and edges minus the cost related to the size of the subgraph, an optimal subgraph is obtained, aiming to contextualize the retrieved recipe process scenarios.

[0114] In practical applications, according to the keywords (such as "cotton fiber", "equipment") in the query embedding vector Zq, the inverted index can be used to initially screen out relevant nodes and edges. Calculate the similarity between the embedding vectors Zn and Ze of the initially screened nodes and edges and the query embedding vector Zq. Cosine similarity or other distance measurement methods can be used.

[0115] Then select the top k nodes and edges with the highest similarity as the most relevant nodes and edges. Select the most relevant nodes and edges: for example, the most relevant node: node 2 (ring spinning frame); the most relevant edge: edge 1 (cotton fiber is used by the ring spinning frame). Connect the most relevant nodes and edges and their adjacent nodes and edges to form an optimal subgraph.

[0116] Through the above steps, query embedding vectors can be generated based on the questions input by users, and the inverted index and k-nearest neighbor retrieval method can be used to efficiently locate the most relevant nodes and edges in the domain knowledge graph, and finally construct the optimal subgraph. This method not only improves the retrieval efficiency but also ensures the relevance and accuracy of the retrieval results.

[0117] See Figure 3 , exemplarily, inputting the user input question after text encoding and the optimal subgraph after graph encoding into a preset large language model to generate an answer containing self-certifying information specifically includes:

[0118] Performing text encoding on the user input question through a text embedder to obtain a query input;

[0119] Aligning the optimal subgraph with the vector space of the preset large language model through a graph encoder containing a multi-layer perceptron to obtain a graph input;

[0120] Inputting the query input and the graph input into the preset large language model together to generate an answer containing decision-making suggestions and self-certifying information; the self-certifying information is the explanatory information of the decision-making suggestions.

[0121] First, at the prompt input end, after the knowledge alignment process described above, the optimal subgraph is obtained; at the enhanced retrieval input end, the retrieved optimal subgraph is encoded using a graph encoder, and the graph tokens are aligned with the vector space of the large language model through the graph encoder of the multi-layer perceptron. Then, the encoded graph tokens (graph input) and the query (query input) processed by the text embedder are input into the pre-trained and frozen large language model together to generate an answer.

[0122] Suppose the optimal subgraph contains node 1 (cotton fiber), node 2 (ring spinning frame), and edge 1 (cotton fiber -> ring spinning frame). Encoding the text attributes of the nodes and edges using a pre-trained language model to generate node embedding vectors and edge embedding vectors. Then using a graph encoder containing a multi-layer perceptron to align the embedding vectors of the nodes and edges into the vector space of the large language model to generate a graph input.

[0123] Through the above steps, the user input question and the optimal subgraph can be encoded respectively, and these encodings are input into the preset large language model to generate an answer containing decision-making suggestions and self-certifying information. This method not only improves the accuracy and relevance of the answer but also provides detailed explanatory information to help users better understand and apply the generated answer.

[0124] Exemplarily, the preset large language model is an open-source large language model or a commercial large language model introducing an enhanced attention mechanism.

[0125] To further enhance the credibility and interpretability of the generated answers, an enhanced attention mechanism is introduced here to examine the relevance of different sub-graph components to the query, ensuring that the generated self-certifying information is both accurate and rooted in the domain-specific knowledge structure. In addition, the optimized training process of the graph encoder and text embedder combined with domain expert knowledge will be studied to further improve the model's understanding and answering quality of complex questions.

[0126] Compared with the prior art, a method for generating answers to industrial formula and process knowledge provided by an embodiment of the present invention ensures high-quality data through data cleaning and preprocessing; through named entity recognition and relationship extraction, the system can accurately identify and extract key entities in the data and the relationships between them, providing structured information for constructing a knowledge graph, and then converting the entities and relationships into nodes and edges in a graph database to form a structured knowledge graph, which is convenient for the management and query of complex relationships. Then, by establishing an inverted index, the system can quickly locate relevant nodes and edges in large-scale data, significantly improving the query speed. Finally, combined with the user input question and the optimal sub-graph, the answers generated by the large language model not only contain decision-making suggestions but also come with self-certifying information, improving the credibility and interpretability of the answers.

[0127] An apparatus for generating answers to industrial formula and process knowledge provided by an embodiment of the present application includes: a standardization module, a fusion module, a graph module, an establishment module, a sub-graph module, and a generation module.

[0128] The standardization module is used to perform data cleaning and preprocessing on industrial formula data and process knowledge data to obtain standardized data.

[0129] The fusion module is used to perform named entity recognition, relationship extraction, and data fusion on the standardized data to obtain an entity set and a relationship set.

[0130] The graph module is used to form a domain knowledge graph for industrial formula and process knowledge according to the entity set and the relationship set.

[0131] The establishment module is used to establish an inverted index according to the node attributes and edge attributes of the domain knowledge graph.

[0132] The sub-graph module is used to encode and retrieve the nodes and edges of the domain knowledge graph according to the inverted index and the query embedding vector generated from the user input question to obtain the optimal sub-graph.

[0133] The generation module is used to input the user input question after text encoding and the optimal sub-graph after graph encoding into a preset large language model to generate an answer with self-certifying information.

[0134] Compared with the prior art, an industrial formula and process knowledge answer generation device provided by an embodiment of the present invention ensures high-quality data through data cleaning and preprocessing; through named entity recognition and relationship extraction, the system can accurately identify and extract key entities in the data and the relationships between them, providing structured information for constructing a knowledge graph, and then converting the entities and relationships into nodes and edges in a graph database to form a structured knowledge graph, which is convenient for the management and query of complex relationships. Then, by establishing an inverted index, the system can quickly locate relevant nodes and edges in large-scale data, significantly improving the query speed. Finally, in combination with the user input question and the optimal subgraph, the answers generated by the large language model not only include decision-making suggestions but also come with self-certifying information, improving the credibility and interpretability of the answers.

[0135] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.

[0136] An embodiment of the present application provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned industrial formula and process knowledge answer generation method.

[0137] The computer device may be a computing device such as a smart phone, a tablet computer, a desktop computer, and a cloud server. The computer device may include but is not limited to a processor and a memory. Those skilled in the art can understand that the figure is only an example of the computer device and does not constitute a limitation on the computer device. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0138] The so-called processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0139] In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit and the external storage device of the computer device. The memory is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory may also be used to temporarily store data that has been output or is to be output.

[0140] Embodiments of the present application provide a computer program product. When the computer program product runs on a computer device, it enables the computer device to execute the steps in the above-mentioned method embodiments.

[0141] In several embodiments provided in the present application, it can be understood that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved.

[0142] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0143] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A method for generating answers to industrial formula and process knowledge, characterized in that: include: Clean and preprocess industrial formula data and process knowledge data to obtain standardized data; Performing named entity recognition, relationship extraction and data fusion on the standardized data to obtain an entity set and a relationship set; According to the entity set and the relationship set, a domain knowledge graph for industrial formula and process knowledge is formed; Establish an inverted index according to the node attributes and edge attributes of the domain knowledge graph; According to the inverted index and the query embedding vector generated by the user input question, the nodes and edges of the domain knowledge graph are encoded and retrieved to obtain the optimal subgraph, specifically including: applying a preset encoding strategy to encode the nodes and edges of the domain knowledge graph respectively; using a k-nearest neighbor retrieval method, identifying the node of the domain knowledge graph with the greatest similarity to the query embedding vector as the most relevant node, and identifying the edge of the domain knowledge graph with the greatest similarity to the query embedding vector as the most relevant edge; The text-encoded user input question and the graph-encoded optimal subgraph are input into a preset large language model to generate an answer containing self-verification information.

2. A method for generating answers to industrial formula and process knowledge as claimed in claim 1, characterized in that: The industrial formula data and process knowledge data are cleaned and preprocessed to obtain standardized data, specifically including: Based on the unique identifier, deleting duplicates in the industrial formula data and process knowledge data; Filling missing values ​​in the industrial formula data and process knowledge data by linear interpolation; The industrial formula data and the process knowledge data are standardized according to the mean and standard deviation of the industrial formula data and the process knowledge data.

3. A method for generating answers to industrial formula and process knowledge as claimed in claim 1, characterized in that: The performing named entity recognition, relationship extraction and data fusion on the standardized data to obtain an entity set and a relationship set specifically includes: Using the RoBERTa model to identify entities in the standardized data; Using a pre-trained language model combined with a bidirectional long short-term memory network to identify the relationship between the entities; Two entities whose similarity is greater than the matching threshold are merged to obtain an entity set and a relationship set.

4. A method for generating answers to industrial formula and process knowledge as claimed in claim 1, characterized in that: The forming of a domain knowledge graph for industrial formula and process knowledge according to the entity set and the relationship set specifically includes: The entity set and the relationship set are filled into different preset entity relationships, the entities in the different entity relationships are used as nodes, and the relationships in the different entity relationships are imported into a graph database as edges to form a domain knowledge graph for industrial formula and process knowledge.

5. The method for generating industrial formula and process knowledge answers as claimed in claim 1, characterized in that: The step of establishing an inverted index according to the node attributes and edge attributes of the domain knowledge graph specifically includes: According to the node attributes of the domain knowledge graph, several node attribute fields are established; According to the edge attributes of the domain knowledge graph, a plurality of edge attribute fields are established; The node serial numbers of all the node attribute fields and the edge serial numbers of all the edge attribute fields are recorded to establish an inverted index.

6. A method for generating answers to industrial formula and process knowledge as claimed in claim 1, characterized in that: Before encoding and retrieving the nodes and edges of the domain knowledge graph according to the inverted index and the query embedding vector generated by the user input question to obtain the optimal subgraph, the method further includes: The user input question is fed into the application language model, and a preset encoding strategy is applied to generate a query embedding vector.

7. A method for generating answers to industrial formula and process knowledge as claimed in claim 1, characterized in that: The step of inputting the text-encoded user input question and the graph-encoded optimal subgraph into a preset large language model to generate an answer containing self-verification information specifically includes: Text encoding is performed on the user input question by a text embedder to obtain a query input; Aligning the optimal subgraph with the vector space of a preset large language model through a graph encoder including a multi-layer perceptron to obtain a graph input; The query input and the graph input are input together into the preset large language model to generate an answer containing decision suggestions and self-verification information; the self-verification information is the explanation information of the decision suggestions.

8. A method for generating answers to industrial formula and process knowledge as claimed in claim 6, characterized in that: The preset large language model is an open source large language model or a commercial large language model that introduces an enhanced attention mechanism.

9. An industrial formula and process knowledge answer generation device, characterized in that: include: The standardization module is used to clean and preprocess the industrial formula data and process knowledge data to obtain standardized data; A fusion module, used for performing named entity recognition, relationship extraction and data fusion on the standardized data to obtain an entity set and a relationship set; A graph module, used to form a domain knowledge graph for industrial formula and process knowledge based on the entity set and the relationship set; Establishing a module for establishing an inverted index according to the node attributes and edge attributes of the domain knowledge graph; The subgraph module is used to encode and retrieve the nodes and edges of the domain knowledge graph according to the inverted index and the query embedding vector generated by the user input question to obtain the optimal subgraph, which specifically includes: applying a preset encoding strategy to encode the nodes and edges of the domain knowledge graph respectively; using a k-nearest neighbor retrieval method to identify the node of the domain knowledge graph with the greatest similarity to the query embedding vector as the most relevant node, and to identify the edge of the domain knowledge graph with the greatest similarity to the query embedding vector as the most relevant edge; A generation module is used to input the user input question after text encoding and the optimal subgraph after graph encoding into a preset large language model to generate an answer containing self-verification information.

Citation Information

Patent Citations

  • Method for constructing knowledge graph based on large language model and vector library

    CN119129722A

  • Method and system for generating enhanced knowledge questions and answers for mixed retrieval of heterogeneous database

    CN119311831A