Intelligent question and answer method and system for chemical safety
By constructing a chemical professional knowledge graph and entity relationship extraction model, combined with a large language model, the accuracy and efficiency of the intelligent question-and-answer system in the chemical field is solved, and accurate answers to chemical safety issues are achieved.
Patent Information
- Application Number
- CN202410110253.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-29
AI Technical Summary
The existing intelligent question-and-answer system is difficult to accurately answer professional questions in the chemical field, and there are problems with confusion in facts and incorrect answers.
Construct a chemical professional knowledge graph, combine pre-trained entity extraction and relationship extraction models, conduct entity recognition and semantic relationship extraction, use the chemical knowledge graph for answer search, and combine large language models for supplementary Q&A when necessary.
It improves the accuracy and efficiency of answering questions in the field of chemical safety, reduces the burden of manual judgment, and gives full play to the understanding and reasoning ability of large language models.
Smart Images

Figure CN120386836A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic recognition, and in particular to a chemical safety intelligent question-answering method and a chemical safety intelligent question-answering system. Background Art
[0002] In chemical production and operation, safety issues have always been the focus of attention. Due to its extremely strong professionalism, it is very difficult for ordinary user groups to distinguish and obtain correct chemical safety knowledge. Even scholars in the field of chemical safety need to spend a lot of time and energy to screen and obtain knowledge. Most traditional intelligent question-answering systems focus on the general field, and there is less research on question-answering in the chemical field. Recently, research on question-answering using knowledge graphs mostly has problems such as inconvenient information acquisition and difficult semantic understanding. A knowledge graph is a data representation evolved from a semantic network, and its characteristic is to describe things, concepts, and relationships in a strongly structured way, and various knowledge can be organized in the form of a graph, so as to effectively organize and manage knowledge. It answers factual questions well, but due to limited data sources, there is a problem of low knowledge coverage, and it cannot answer questions flexibly that are not included in the knowledge graph and cannot effectively handle questions with strong subjectivity. The emergence of large language models provides new ideas and methods for solving these problems. A large language model (LLM) refers to a type of language model based on neural networks with a large number of parameters (usually billions or more). Compared with models with small-scale parameters, the large language model has a qualitative leap in the ability of natural language understanding and reasoning, and this performance is called "emergence of ability". Although large language models perform very well in various natural language tasks in the general field, they have limitations such as chaotic answers to facts and incorrect generated answers in specific professional fields. To solve this problem, a new intelligent question-answering system applied to the chemical field needs to be proposed. Summary of the Invention
[0003] The purpose of the embodiments of the present invention is to provide a chemical safety intelligent question-answering method and system to at least solve the problem that existing intelligent question-answering solutions are prone to chaotic answers to facts and incorrect generated answers in the chemical field.
[0004] To achieve the above object, a first aspect of the present invention provides a chemical safety intelligent question-answering method, the method comprising: collecting question information of a user, and processing the question information into text information; performing matching with a preset question-answering template based on the text information, and directly feeding back answer information based on the preset question-answering template after successful matching; otherwise, performing entity recognition in the text information based on a pre-trained entity extraction model to obtain entity mentions; performing semantic relationship extraction in the text information based on a pre-trained relationship extraction model; performing answer retrieval in a pre-constructed chemical knowledge graph based on the entity mentions and the extracted semantic relationships, and assembling and feeding back the answer to the user side.
[0005] Optionally, the question information of the user is voice question information or text question information; if the question information of the user is voice question information, the voice question information is processed into corresponding text information based on a voice recognition algorithm.
[0006] Optionally, after processing the question information into text information, the method further comprises: performing first-level normalization processing on pinyin, special characters, and punctuation marks in the text information; performing stop word list and synonym list normalization processing on the text information after the first-level normalization processing to obtain processed text information; performing vector embedding processing on the processed text information, and performing text vectorization processing on the text information after the vector embedding processing.
[0007] Optionally, the performing matching with a preset question-answering template based on the text information, and directly feeding back answer information based on the preset question-answering template after successful matching, comprises: calculating a similarity between a question vector after text vectorization processing and a question in a preset question-answering template; if the similarity is greater than a preset threshold, directly feeding back answer information based on the question in the corresponding preset question-answering template.
[0008] Optionally, the entity extraction model comprises: an ALBERT model layer for initializing word vectors of text information; a BiLSTM network layer for learning context feature information based on the word vectors, performing entity recognition, and predicting the probability that each character belongs to different labels; a CRF processing layer for calculating the optimal entity mentions.
[0009] Optionally, the loss function of the entity extraction model is:
[0010]
[0011] where entityLoss is the loss function of the entity extraction model; p is the true distribution of BIO annotations of questions in training samples; q is the distribution of BIO of questions predicted by the entity extraction model; and n is the dimension of question representations.
[0012] Optionally, the relation extraction model includes: an Embedding layer for mapping each word in the text information to a low-dimensional space; a bidirectional LSTM layer for feature extraction based on the word vectors mapped by the Embedding layer; an Attention layer for generating a weight vector and merging the lexical-level features in each iteration into sentence-level features based on the weight vector; and an output layer for classifying the sentence-level feature vectors to obtain semantic relations.
[0013] Optionally, based on the entity mentions and the extracted semantic relations, answer retrieval is performed in a pre-constructed chemical knowledge graph, and the answers are assembled and fed back to the user side, including: in the pre-constructed chemical knowledge graph, using the entity mentions as the central nodes, retrieving all subgraphs within one-hop distance containing the central nodes as query graphs; pruning the edges of the query graphs and selecting the top k edges as candidate relations; calculating the similarity between the semantic relations and the candidate relations respectively; if the similarity is greater than a preset similarity threshold, taking the retrieved answers as feedback answers, assembling the answers and feeding them back to the user side; otherwise, performing semantic understanding and question answering based on a large language model.
[0014] Optionally, the method further includes: constructing a chemical knowledge graph, including: collecting structured, semi-structured, and unstructured data related to the field of chemical safety; preprocessing the obtained data, and extracting the preprocessed data into structured triples for storage in a Neo4j database as the chemical knowledge graph.
[0015] Optionally, performing semantic understanding and question answering based on a large language model includes: storing the questions, normalized questions, entity mentions, topic entities, candidate relations, relations, and answers proposed by users during the question parsing process in a PostgreSQL or MySQL database; constructing a training data set based on the stored data, and using the constructed chemical knowledge graph as domain knowledge and the training data set as training samples together; performing large language model training based on the training samples, and performing semantic understanding and question answering based on the training results.
[0016] In a second aspect of the present invention, a chemical safety intelligent question-answering system is provided. The system includes: a collection unit for collecting question information of a user and processing the question information into text information; a matching unit for performing preset question-answering template matching based on the text information and directly feeding back answer information based on the preset question-answering template after successful matching; an entity recognition unit for, conversely, performing entity recognition in the text information based on a pre-trained entity extraction model to obtain entity mentions; a relationship extraction unit for performing semantic relationship extraction in the text information based on a pre-trained relationship extraction model; and an answer generation unit for performing answer retrieval in a pre-constructed chemical knowledge graph based on the entity mentions and the extracted semantic relationships, and assembling and feeding back the answers to the user side.
[0017] In a third aspect of the present invention, a computer-readable storage medium is provided. Instructions are stored on the computer-readable storage medium, and when running on a computer, the computer is caused to execute the above-mentioned chemical safety intelligent question-answering method.
[0018] In a fourth aspect of the present invention, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned chemical safety intelligent question-answering method is implemented.
[0019] Through the above technical solutions, the solution of the present invention constructs a chemical engineering professional knowledge graph to summarize and store knowledge in the field of chemical safety. At the same time, a chemical safety intelligent question-answering system is constructed based on the knowledge graph, which can accurately select the most matching answer for the user's input question, effectively reduce the workload of manual judgment, and improve work efficiency. Using the constructed chemical engineering professional knowledge graph to provide the background knowledge required for reasoning by the large language model reduces factual errors in the reasoning of the large language model; providing background knowledge information rather than direct answers to the large language model can give full play to the powerful understanding and reasoning abilities of the large language model.
[0020] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementation, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0022] Figure 1 is a flowchart of the steps of a chemical safety intelligent question-answering method provided by an embodiment of the present invention;
[0023] Figure 2It is the system structure diagram of the chemical safety intelligent question-answering system provided by an embodiment of the present invention. Specific embodiments
[0024] The following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.
[0025] In chemical production and operation, safety issues have always been the focus of attention. Due to its extremely strong professionalism, it is very difficult for ordinary user groups to distinguish and obtain correct chemical safety knowledge. Even scholars in the field of chemical safety need to spend a lot of time and energy to screen and obtain knowledge. Most traditional intelligent question-answering systems focus on the general field, and there is less research on question-answering in the chemical field. Recently, research on question-answering using knowledge graphs mostly has problems such as inconvenient information acquisition and difficult semantic understanding. A knowledge graph is a data representation evolved from a semantic network, and its characteristic is to describe things, concepts, and relationships in a strongly structured way, and various knowledge can be organized in the form of a graph, so as to effectively organize and manage knowledge. It answers factual questions well, but due to limited data sources, there is a problem of low knowledge coverage, and it cannot answer questions flexibly for questions not included in the knowledge graph, and cannot effectively handle questions with strong subjectivity. The emergence of large language models provides new ideas and methods for solving these problems. A large language model (LLM) refers to a type of language model based on neural networks with a large number of parameters (usually billions or more). Compared with models with small-scale parameters, large language models have a qualitative leap in the ability of natural language understanding and reasoning, and this performance is called "emergence of ability". Although large language models perform very well in various natural language tasks in the general field, they have limitations such as chaotic answers and incorrect generated answers in specific professional fields.
[0026] In response to the above problems, the solution of the present invention constructs a chemical engineering professional knowledge graph to summarize and store the knowledge in the field of chemical safety. At the same time, a chemical safety intelligent question-answering system is constructed based on the knowledge graph, which can accurately select the most matching answer for the user's input question, effectively reduce the workload of manual judgment, and improve work efficiency. By constructing a chemical engineering professional knowledge graph, background knowledge required for reasoning is provided for the large language model, reducing factual errors in the reasoning of the large language model; providing background knowledge information rather than direct answers for the large language model can give full play to the powerful understanding and reasoning ability of the large language model.
[0027] Figure 1 It is the method flow chart of the method for mining the influencing factors of the safe production status of a refining device provided by an embodiment of the present invention. As Figure 1As shown in the figure, an embodiment of the present invention provides a method for mining influencing factors of the safe production status of a refining and chemical device. The method includes:
[0028] Step S10: Collect the problem information of the user and process the problem information into text information.
[0029] Specifically, the problem information of the user is voice problem information or text problem information; if the problem information of the user is voice problem information, then based on the speech recognition algorithm, the voice problem information is processed into the corresponding text information.
[0030] After processing the problem information into text information, the method further includes: performing first-level normalization processing on the pinyin, special characters, and punctuation marks in the text information; performing stop word list and synonym list normalization processing on the text information after the first-level normalization processing to obtain the processed text information; performing vector embedding processing on the processed text information, and performing text vectorization processing on the text information after the vector embedding processing.
[0031] Example 1:
[0032] Obtain the problem input by the user. For the voice input of the user, text conversion should be performed using speech processing technology. For the obtained problem text, first perform normalization processing, including processing the pinyin, special characters, and punctuation marks in the problem text. At the same time, normalize the processed text using a common stop word list and synonym list. For example, replace some colloquial words of the user (such as "what is", "why", etc.) with (such as "definition", "concept", "reason", etc.).
[0033] Step S20: Based on the text information, perform matching with a preset question-answer template, and directly feedback answer information based on the preset question-answer template after successful matching.
[0034] Specifically, calculate the similarity between the question vector after text vectorization processing and the question in the preset question-answer template; if the similarity is greater than the preset threshold, directly feedback the answer information in the corresponding preset question-answer template.
[0035] In the embodiment of the present invention, first perform vector embedding processing on the normalized question sentence, use pre-trained models such as Bert and Glove for text vectorization, compare the similarity between the obtained question vector and the question in the common question-answer pairs. If the similarity is greater than the set threshold, directly feedback the answer in the common question-answer pairs to the user. If the similarity is less than the set threshold, use an entity extraction model to perform entity recognition on the question to obtain the key entity of the question sentence.
[0036] Step S30: Conversely, based on the pre-trained entity extraction model, perform entity recognition in the text information to obtain entity mentions.
[0037] Specifically, the entity extraction model includes: an ALBERT model layer for initializing word vectors of text information; a BiLSTM network layer for learning context feature information based on the word vectors, performing entity recognition, and predicting the probability that each word belongs to different tags; and a CRF processing layer for calculating the optimal entity mentions.
[0038] Embodiment 2:
[0039] The entity extraction process is as follows: First, perform matching using a dictionary. By constructing a dictionary of common chemical safety knowledge entities, tokenize the question text and perform dictionary matching. Entities are mined from data such as common chemical equipment, device operation instructions, and process cards, and entities are exported from the constructed knowledge graph to build a custom dictionary. The custom dictionary is imported into the custom dictionary of jieba tokenizer, and the part-of-speech of all words in the custom dictionary is set to "Entity". The question sentence is tokenized and part-of-speech tagged using jieba. Select the words with the part-of-speech "Entity" as the topic entities of the question sentence. For example, if the question sentence is "When was xx oil depot built", the jieba part-of-speech tagging result is "xx oil depot / Entity's / uj built / v time / n is / v what / r when / n", then the entity (also called entity mention) mentioned in this question sentence is selected as "xx oil depot".
[0040] Furthermore, if the entity matching is successful, relation extraction is performed. If the entity matching fails, the entity extraction model is used for entity recognition. The entity extraction model uses ALBert-BiLSTM-CRF to train an initial entity extraction model of ALBerT+BiLSTM+CRF. For the entity extraction model, its first layer uses a pre-trained ALBERT model (a pre-trained language representation model) to initialize the word vectors of text information, and the obtained word vectors can extract the main features of the text through the relationships between words; the second layer is a bidirectional long short-term memory neural network BiLSTM (composed of a forward LSTM and a backward LSTM, which can pay attention to the front and back information of the text and avoid loss of important information). The word vectors obtained in the first layer are used as the input of each time step of the bidirectional long short-term memory neural network BiLSTM. Through the BiLSTM model, context feature information is learned, entity recognition is performed, and the probability that each word belongs to different tags is predicted; in the third layer, since the tag probabilities predicted by BiLSTM do not consider the actual relationships between tags, the output sequence of BiLSTM must be processed by CRF, and combined with the relationships between tags, that is, the state transition matrix (for example, the I tag cannot appear before the B tag), to calculate the optimal entity mentions. The loss function of the entity extraction model is:
[0041]
[0042] Among them, entityLoss is the loss function of the entity extraction model; p is the true distribution of the BIO annotation of the questions in the training samples; q is the distribution of the BIO of the questions predicted by the entity extraction model; and n is the dimension of the question representation.
[0043] Example 3:
[0044] The question is "When was xx Oil Depot built", and the prediction result of the entity extraction model is "x / Bx / I oil / I depot / I of / O built / O into / O time / O is / O what / O time / O". Among them, B indicates the start of the entity mention. Extract all the words marked with I after the start of B to obtain the entity mention "xx Oil Depot".
[0045] Step S40: Perform semantic relation extraction in the text information based on the pre-trained relation extraction model.
[0046] Specifically, the relation extraction model includes: an Embedding layer for mapping each word in the text information to a low-dimensional space; a bidirectional LSTM layer for feature extraction based on the word vectors mapped by the Embedding layer; an Attention layer for generating a weight vector and merging the word-level features in each iteration into sentence-level features based on this weight vector; and an output layer for performing relation classification on the sentence-level feature vectors to obtain semantic relations.
[0047] Example 4:
[0048] After obtaining the entity mentions by using dictionary matching and the entity extraction model, a relation extraction model is used to obtain the relations of the questions. The relation extraction model adopts a relation extraction model combining the attention mechanism and bidirectional LSTM. The model is divided into: Embedding layer: Map each word to a low-dimensional space. LSTM layer: Use bidirectional LSTM to obtain high-level features from the Embedding layer; in the embedding layer, the extraction model can obtain the embedding of each word and use it as a sequence I=(i1,i2,...i n )). There are two LSTM models in the Bi-LSTM, one extracting forward hidden features The other extracts backward hidden features The Bi-LSTM layer can be expressed as follows:
[0049]
[0050]
[0051] Merging the forward and backward hidden layer features can obtain the output Extract with an activation function and then map it to a d l -dimensional space. d l is the number of entity types to be recognized.
[0052] Attention layer: Generate a weight vector. By multiplying with this weight vector, the word-level features in each iteration are combined into sentence-level features; Represent the set of vectors input to the LSTM layer as H: [h1, h2, …, hT]. The weight matrix obtained by its Attention layer is obtained in the following way:
[0053] M = tanh(H)
[0054] a = softmax(w T M)
[0055] r = Hα T
[0056] where d w is the dimension of the word vector, and w T is the transpose of a parameter vector obtained through training and learning. The sentence finally used for classification will be represented as follows:
[0057] h* = tanh(r)
[0058] Output layer: Use the sentence-level feature vector for relation classification.
[0059] Step S50: Based on the entity mentions and the extracted semantic relations, perform answer retrieval in the pre-constructed chemical knowledge graph, and assemble the answers and feedback them to the user side.
[0060] Specifically, in the pre-constructed chemical knowledge graph, use the entity mentions as the central nodes, retrieve all subgraphs within one-hop distance containing the central nodes as the query graph; prune the edges of the query graph, select the top k edges as candidate relations; calculate the similarity between the semantic relations and the candidate relations respectively; if the similarity is greater than the preset similarity threshold, use the retrieved answer as the feedback answer, and assemble the answers and feedback them to the user side; otherwise, perform semantic understanding and question answering based on the large language model.
[0061] In an embodiment of the present invention, after obtaining the key entities and semantic relationships in the question sentence, a retrieval is performed in the knowledge graph. In the knowledge graph, the key entities are used as the central nodes, and all subgraphs within one-hop distance containing the central nodes are retrieved as the query graph. Pruning is performed on the edges of the query graph, and the top k edges are selected as candidate relationships. Since the knowledge graph may not contain the semantic relationships obtained from the query, the top k edges obtained from the query are used as candidate relationships, and the similarity between each candidate path and the semantic relationship is calculated. The specific model consists of a BERT as the basic layer, an average pooling layer, a contrastive learning loss layer, and a softmax layer. The BERT layer encodes the sentence into word vectors, the average pooling layer normalizes the word vectors, the contrastive learning loss layer enhances the connection between the identified semantic relationship and the candidate relationship through contrastive learning, and finally the softmax layer outputs the probability of the candidate relationship. Finally, the optimal candidate path is selected as the actual relationship. If the score of the candidate path exceeds the set threshold, it is used as the final relationship. If the score is lower than the threshold, a large language model is used for in-depth semantic understanding and question answering.
[0062] After obtaining the key entities and the finally determined relationships, a Cypher query statement is constructed to retrieve the answer in the graph, and the answer is assembled and fed back to the user.
[0063] Furthermore, if the score of the candidate path retrieved in the knowledge graph is less than the set threshold, it indicates that there is no suitable semantic relationship in the knowledge graph and the question cannot be answered and fed back. Therefore, a large language model is used for supplementary question answering. Models used in the industry such as ChatGPT, Wenxin Yiyan, and Tongyi Qianwen can be adopted as the initial large language model. Information such as the question raised by the user during the question parsing process, the normalized question sentence, entity mentions, topic entities, candidate relationships, relationships, and answers is stored in a PostgreSQL or MySQL database to construct the training dataset required by the large language model. At the same time, the structured entities, relationships, etc. data included in the constructed chemical engineering professional knowledge graph are implicitly added to the training process of the large language model, integrating the entity representation in the knowledge graph into the text representation, and using the background knowledge of the knowledge graph as the background information for the large language model to reason, enabling the large language model to better exert its semantic understanding and reasoning capabilities.
[0064] Preferably, the method further includes: constructing a chemical engineering knowledge graph, including: collecting structured, semi-structured, and unstructured data related to the field of chemical engineering safety; preprocessing the obtained data, and extracting the preprocessed data into structured triples for storage in a Neo4j database as the chemical engineering knowledge graph.
[0065] Example Five:
[0066] 1) Collect structured, semi-structured, and unstructured data related to the field of chemical safety, such as structured databases like the National Work Safety Risk Monitoring and Early Warning System and the Sinopec Dual-Prevention Digital Intelligence Control Platform; semi-structured and unstructured data such as technical specifications, operation documents, standards and specifications, enterprise annual reports, emergency plans, and accident cases.
[0067] 2) Preprocess the obtained data, and then extract it into structured triples for storage in a Neo4j database. For structured data, perform operations such as data cleaning and outlier removal; for unstructured data, use deep learning techniques to extract entities, attributes, and relationships; finally, import the obtained triple data into the Neo4j graph database for storage and display.
[0068] Figure 2 It is the system structure diagram of the chemical safety intelligent question-answering system provided by an embodiment of the present invention. As Figure 2 shown, an embodiment of the present invention provides a chemical safety intelligent question-answering system, and the system includes:
[0069] An acquisition unit, configured to acquire the question information of the user and process the question information into text information.
[0070] Specifically, the question information of the user is voice question information or text question information; if the question information of the user is voice question information, then based on the speech recognition algorithm, the voice question information is processed into the corresponding text information.
[0071] After processing the question information into text information, the method further includes: performing primary normalization processing on pinyin, special characters, and punctuation marks in the text information; performing stop word list and synonym list normalization processing on the text information after primary normalization processing to obtain the processed text information; performing vector embedding processing on the processed text information, and performing text vectorization processing on the text information after vector embedding processing.
[0072] Obtain the question input by the user. For the voice input of the user, use speech processing technology to perform text conversion. For the obtained question text, first perform normalization processing, including processing pinyin, special characters, and punctuation marks in the question text, and at the same time perform normalization processing on the processed text using a common stop word list and synonym list, such as replacing some spoken words of the user (such as what, why, etc.) with (such as definition, concept, reason, etc.).
[0073] A matching unit, configured to perform preset question-answer template matching based on the text information, and directly feedback answer information based on the preset question-answer template after successful matching.
[0074] Specifically, a similarity calculation is performed between the question vector after text vectorization processing and the questions in the preset Q&A template; if the similarity is greater than the preset threshold, answer information is directly fed back based on the questions in the corresponding preset Q&A template.
[0075] In the embodiment of the present invention, first, the normalized question sentence is subjected to vector embedding processing, and pre-trained models such as Bert and Glove are used for text vectorization. A similarity comparison is made between the obtained question vector and the questions in the common Q&A pairs. If the similarity is greater than the set threshold, the answer in the common Q&A pairs is directly fed back to the user. If the similarity is less than the set threshold, an entity extraction model is used to perform entity recognition on the question to obtain the key entities of the question sentence.
[0076] The entity recognition unit, on the contrary, is used to perform entity recognition in the text information based on the pre-trained entity extraction model to obtain entity mentions.
[0077] Specifically, the entity extraction model includes: an ALBERT model layer for initializing the word vectors of the text information; a BiLSTM network layer for learning context feature information based on the word vectors, performing entity recognition, and predicting the probability that each word belongs to different labels; a CRF processing layer for calculating the optimal entity mentions.
[0078] The entity extraction process is as follows: First, dictionary matching is performed using a dictionary. By constructing a common chemical safety knowledge entity dictionary, the question text is segmented, and dictionary matching is carried out. Entities are mined from data such as common chemical equipment, device operation instructions, and process cards, and entities are derived from the constructed knowledge graph to build a custom dictionary. The dictionary is imported into the custom dictionary of jieba segmentation, and the part-of-speech of all words in the custom dictionary is set to "Entity". The question sentence is segmented and part-of-speech tagged using jieba segmentation. The words with the part-of-speech of "Entity" are selected as the topic entities of the question sentence. For example, if the question sentence is "When was xx oil depot built", the jieba part-of-speech tagging result is "xx oil depot / Entity of / uj built / v time / n is / v what / r when / n", then the entity (also called entity mention) mentioned in this question sentence is selected as "xx oil depot".
[0079] Further, if the entity matching is successful, relation extraction is performed. If the entity matching is unsuccessful, an entity extraction model is used for entity recognition. The entity extraction model uses ALBert-BiLSTM-CRF, and an initial entity extraction model of ALBerT+BiLSTM+CRF is trained. For the entity extraction model, its first layer uses a pre-trained ALBERT model (a pre-trained language representation model) to initialize the word vectors of the text information, and the obtained word vectors can extract the main features of the text through the relationships between words; the second layer is a bidirectional long short-term memory neural network BiLSTM (composed of a forward LSTM and a backward LSTM, which can pay attention to the front and back information of the text and avoid the loss of important information). The word vectors obtained in the first layer are used as the input of the bidirectional long short-term memory neural network BiLSTM at each time step. Through the BiLSTM model, the context feature information is learned for entity recognition, and the probability that each word belongs to different labels is predicted; in the third layer, since the label probabilities predicted by BiLSTM do not consider the actual relationships between labels, the output sequence of BiLSTM must be processed by CRF, and combined with the relationship between labels, that is, the state transition matrix (such as the I label cannot appear before the B label), to calculate the optimal entity mention. The loss function of the entity extraction model is:
[0080]
[0081] where entityLoss is the loss function of the entity extraction model; p is the true distribution of the BIO annotation of the questions in the training samples; q is the distribution of the BIO of the questions predicted by the entity extraction model; and n is the dimension of the question representation.
[0082] For example, for the question "When was xx oil depot built", the prediction result of the entity extraction model is "x / Bx / I oil / I depot / I of / O built / O into / O time / O is / O what / O time / O". Among them, B indicates the start of the entity mention, and all words marked with I after the start of B are extracted to obtain the entity mention "xx oil depot".
[0083] The relation extraction unit is used to perform semantic relation extraction in the text information based on a pre-trained relation extraction model.
[0084] Specifically, the relation extraction model includes: an Embedding layer for mapping each word in the text information to a low-dimensional space; a bidirectional LSTM layer for feature extraction based on the word vectors mapped by the Embedding layer; an Attention layer for generating a weight vector and merging the word-level features in each iteration into sentence-level features based on the weight vector; and an output layer for performing relation classification on the sentence-level feature vectors to obtain semantic relations.
[0085] After obtaining entity mentions using dictionary matching and entity extraction models, a relation extraction model is used to obtain the relations in the question. The relation extraction model is a relation extraction model that combines an attention mechanism and bidirectional LSTM. The model consists of: Embedding layer: Maps each word to a low-dimensional space. LSTM layer: Uses bidirectional LSTM to obtain high-level features from the Embedding layer; in the embedding layer, the extraction model can obtain the embedding of each word and use it as a sequence I = (i1, i2,... i n ) In the Bi-LSTM, there are two LSTM models, one extracting forward hidden features and the other extracting backward hidden features The Bi-LSTM layer can be represented as follows:
[0086]
[0087]
[0088] Merging the forward and backward hidden layer features can obtain the output Use an activation function to extract and then map it to a d l -dimensional space. d l is the number of entity types to be recognized.
[0089] Attention layer: Generates a weight vector. By multiplying with this weight vector, the word-level features in each iteration are combined into sentence-level features; represent the set of vectors input to the LSTM layer as H: [h1, h2,..., hT]. The weight matrix obtained by its Attention layer is obtained in the following way:
[0090] M = tanh(H)
[0091] a = softmax(w T M)
[0092] r = Hα T
[0093] where d w is the dimension of the word vector, and w T is the transpose of a parameter vector obtained through training. The sentence finally used for classification will be represented as follows:
[0094] h* = tanh(r)
[0095] Output layer: Uses the sentence-level feature vector for relation classification.
[0096] An answer generation unit, configured to retrieve answers in a pre-constructed chemical knowledge graph based on the entity mention and the extracted semantic relationship, and feedback the assembled answers to the user side.
[0097] Specifically, in the pre-constructed chemical knowledge graph, using the entity mention as the central node, retrieve all subgraphs within one-hop distance containing the central node as the query graph; prune the edges of the query graph, and select the top k edges as candidate relationships; calculate the similarity between the semantic relationship and the candidate relationships respectively; if the similarity is greater than the preset similarity threshold, use the retrieved answer as the feedback answer, and feedback the assembled answer to the user side; otherwise, perform semantic understanding and question answering based on the large language model.
[0098] In the embodiment of the present invention, after obtaining the key entity and semantic relationship in the question sentence, retrieve in the knowledge graph. In the knowledge graph, use the key entity as the central node, retrieve all subgraphs within one-hop distance containing the central node as the query graph, prune the edges of the query graph, and select the top k edges as candidate relationships. Since the knowledge graph may not contain the retrieved semantic relationship, the top k retrieved edges are used as candidate relationships, and calculate the similarity between each candidate path and the semantic relationship. The specific model consists of a BERT as the basic layer, an average pooling layer, a contrastive learning loss layer, and a softmax layer. The BERT layer encodes the sentence into word vectors, the average pooling layer normalizes the word vectors, the contrastive learning loss layer enhances the connection between the recognized semantic relationship and the candidate relationships through contrastive learning, and finally the softmax layer outputs the probability of the candidate relationships. Finally, select the optimal candidate path as the actual relationship. If the score of the candidate path exceeds the set threshold, use it as the final relationship. If the score is lower than the threshold, use the large language model for in-depth semantic understanding and question answering.
[0099] After obtaining the key entity and the finally determined relationship, construct a cypher query statement to retrieve the answer in the graph, assemble the answer, and feedback it to the user.
[0100] Furthermore, if the score of the candidate path retrieved from the knowledge graph is less than the set threshold, it indicates that there is no suitable semantic relationship in the knowledge graph and the question cannot be answered and feedback cannot be provided. Therefore, supplementary question answering needs to be performed through a large language model. Models commonly used in the industry such as ChatGPT, ERNIE Bot, and Tongyi Qianwen can be used as the initial large language model. Information such as the questions raised by users during the question parsing process, the normalized questions, entity mentions, topic entities, candidate relationships, relationships, and answers are stored in a PostgreSQL or MySQL database to construct the training dataset required by the large language model. At the same time, the structured entities, relationships, and other data included in the constructed chemical engineering professional knowledge graph are implicitly incorporated into the training process of the large language model as domain knowledge, integrating the entity representations in the knowledge graph into the text representation, and using the background knowledge of the knowledge graph as the background information for the large language model's reasoning, enabling the large language model to better utilize its semantic understanding and reasoning capabilities.
[0101] Preferably, the method further includes: constructing a chemical engineering knowledge graph, including: collecting structured, semi-structured, and unstructured data related to the chemical engineering safety field; preprocessing the obtained data, and extracting the preprocessed data into structured triples for storage in a Neo4j database as the chemical engineering knowledge graph.
[0102] 1) Collect structured, semi-structured, and unstructured data related to the chemical engineering safety field, such as structured databases like the National Work Safety Risk Monitoring and Early Warning System and the Sinopec Dual-Prevention Digital Control Platform; semi-structured and unstructured data such as technical specifications, operation manuals, standards, enterprise annual reports, emergency plans, and accident cases.
[0103] 2) Preprocess the obtained data and then extract it into structured triples for storage in a Neo4j database. For structured data, perform operations such as data cleaning and outlier removal; for unstructured data, use deep learning techniques to extract entities, attributes, and relationships; finally, import the obtained triple data into the Neo4j graph database for storage and display.
[0104] The embodiment of the present invention also provides a computer-readable storage medium, on which instructions are stored, and when they run on a computer, the computer is enabled to execute the above-mentioned chemical engineering safety intelligent question-answering method.
[0105] Those skilled in the art can understand that all or part of the steps in the methods for implementing the above embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a single-chip microcomputer, a chip, or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0106] The optional embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the technical concept scope of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention. Additionally, it should be noted that in the above specific embodiments, the various specific technical features described can be combined in any appropriate manner without conflict. To avoid unnecessary repetition, the embodiments of the present invention will not separately describe various possible combination methods.
[0107] Furthermore, any combination can be made among various different embodiments of the present invention as long as it does not violate the idea of the embodiments of the present invention, and it should also be regarded as the content disclosed in the embodiments of the present invention.
Claims
1. A chemical safety intelligent Q&A method, characterized in that The method includes: Collect the question information of the user and process the question information into text information; Based on the text information, perform matching with a preset Q&A template, and directly feedback answer information based on the preset Q&A template after successful matching; Conversely, perform entity recognition in the text information based on a pre-trained entity extraction model to obtain entity mentions; Perform semantic relationship extraction in the text information based on a pre-trained relationship extraction model; Based on the entity mentions and the extracted semantic relationships, perform answer retrieval in a pre-constructed chemical knowledge graph, and assemble and feedback the answers to the user side.
2. The method according to claim 1, wherein The question information of the user is voice question information or text question information; If the question information of the user is voice question information, then process the voice question information into corresponding text information based on a speech recognition algorithm.
3. The method according to claim 1, characterized in that, After processing the question information into text information, the method further includes: Perform first-level normalization processing on the pinyin, special characters, and punctuation marks in the text information; Perform stop word list and synonym list normalization processing on the text information after the first-level normalization processing to obtain the processed text information; Perform vector embedding processing on the processed text information, and perform text vectorization processing on the text information after vector embedding processing.
4. The method according to claim 3, characterized in that, The performing matching with a preset Q&A template based on the text information and directly feedbacking answer information based on the preset Q&A template after successful matching includes: Calculate the similarity between the question vector after text vectorization processing and the questions in the preset Q&A template; If the similarity is greater than a preset threshold, directly feedback answer information based on the questions in the corresponding preset Q&A template.
5. The method according to claim 1, wherein The entity extraction model includes: The ALBERT model layer, which is used to initialize the word vectors of the text information; The BiLSTM network layer, which is used to learn context feature information based on the word vectors, perform entity recognition, and predict the probability that each character belongs to different labels; The CRF processing layer, which is used to calculate the optimal entity mentions.
6. The method according to claim 1, wherein The loss function of the entity extraction model is: Where entityLoss is the loss function of the entity extraction model; p is the true distribution of the BIO annotation of the question sentence in the training sample; q is the distribution of the BIO of the question sentence predicted by the entity extraction model; n is the dimension of the question sentence representation.
7. The method according to claim 1, characterized in that The relationship extraction model includes: The Embedding layer, which is used to map each word in the text information to a low-dimensional space; The bidirectional LSTM layer, which is used to perform feature extraction based on the word vectors mapped by the Embedding layer; The Attention layer, which is used to generate a weight vector, and based on this weight vector, merge the lexical-level features in each iteration into sentence-level features; The output layer, which is used to perform relationship classification on the sentence-level feature vectors to obtain semantic relationships.
8. The method according to claim 1, wherein The performing answer retrieval in a pre-constructed chemical knowledge graph based on the entity mentions and the extracted semantic relationships, and assembling and feedbacking the answers to the user side includes: In the pre-constructed chemical knowledge graph, use the entity mentions as the central nodes, and retrieve all subgraphs within one-hop distance containing the central nodes as the query graph; Prune the edges of the query graph and select the top k edges as candidate relationships; Calculate the similarity between the semantic relationship and the candidate relationships respectively; If the similarity is greater than the preset similarity threshold, use the retrieved answer as the feedback answer and assemble the answer to be fed back to the user side; Otherwise, perform semantic understanding and question answering based on the large language model.
9. The method according to claim 1, wherein The method further includes: Construct a chemical industry knowledge graph, including: Collect structured, semi-structured and unstructured data related to the chemical safety field; Preprocess the obtained data, and extract the preprocessed data into structured triples for storage in a Neo4j database as the chemical industry knowledge graph.
10. The method according to claim 8, characterized in that The performing semantic understanding and question answering based on the large language model includes: Store the questions raised by the user, the normalized questions, entity mentions, topic entities, candidate relationships, relationships, and answers during the question parsing process in a PostgreSQL or MySQL database; Construct a training data set based on the stored data, and use the constructed chemical industry knowledge graph as domain knowledge and the training data set as training samples together; Perform large language model training based on the training samples, and perform semantic understanding and question answering based on the training results.
11. A chemical safety intelligent Q&A system, characterized in that, The system includes: A collection unit for collecting the question information of the user and processing the question information into text information; A matching unit for performing preset question-answer template matching based on the text information and directly feeding back answer information based on the preset question-answer template after successful matching; An entity recognition unit for performing entity recognition in the text information based on a pre-trained entity extraction model to obtain entity mentions after the matching fails; A relationship extraction unit for performing semantic relationship extraction in the text information based on a pre-trained relationship extraction model; An answer generation unit for retrieving answers in the pre-constructed chemical industry knowledge graph based on the entity mentions and the extracted semantic relationships, and assembling the answers to be fed back to the user side.
12. A computer-readable storage medium, characterized in that, Instructions are stored on the computer-readable storage medium, and when running on a computer, the computer executes the chemical safety intelligent question-answering method described in any one of claims 1-10.
13. An electronic device, the electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the chemical safety intelligent question-answering method described in any one of claims 1-10.