Question answering method and device based on knowledge graph and electronic equipment
By constructing a knowledge graph-based question-and-answer method, using a sequence model of attention mechanism to supplement entities and relationships, the problem of low efficiency in querying rules for dangerous cargo transportation in relational databases in aviation, realizing multi-hop question-and-answer and intelligent interaction, improving query efficiency and accuracy.
Patent Information
- Application Number
- CN202510385148.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, when using relational databases to store and manage air hazardous cargo transportation rules, there is a problem that inter-entity relationships lead to low query efficiency. Especially when dealing with complex multi-factor associations, SQL statements are lengthy and difficult to maintain, and the knowledge question and answer system is inefficient and has a high error rate when facing unstructured or semi-structured rule texts.
Using a question-and-answer method based on knowledge graphs, we use a sequence model of the initial knowledge graph and use an attention mechanism to supplement it, establish a target knowledge graph, realize efficient supplementation and query of entities and relationships, and support the ability of multi-hop questions and answers.
It improves the query efficiency and accuracy of aviation dangerous cargo transportation rules, supports multi-hop Q&A, provides intelligent Q&A interactive experience, promotes information sharing and collaboration, and improves operational efficiency.
Smart Images

Figure CN120338066A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of air transportation, and in particular, to a question answering method, device and electronic device based on a knowledge graph. Background Art
[0002] Currently, in the scenario of air transportation of dangerous goods, most of the business systems in the related technologies rely on relational databases to store and manage dangerous goods transportation rules. However, since relational databases represent data in the form of two-dimensional tables, each table represents an entity or concept, such as "dangerous goods catalog table", "dangerous goods category table", "packaging grade table", "transportation condition table", etc., and use foreign keys to associate the relationships between different tables. However, there are many limitations in dealing with complex multi-factor associations, specifically as follows:
[0003] (1) When retrieving data using SQL (Structured Query Language), when it comes to complex multi-table joins and nested subqueries, the SQL statements become long and difficult to maintain, and the efficiency of executing the corresponding SQL statements to return results is relatively low;
[0004] (2) For knowledge question answering systems, they mainly rely on keyword matching and simple pattern recognition algorithms. When facing a large amount of unstructured or semi-structured rule texts, it is difficult to query, and during the process of querying rules, the relationships between dangerous goods transportation rules cannot be understood and require manual parsing, with low efficiency and high error rate.
[0005] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0006] Embodiments of the present invention provide a question answering method, device and electronic device based on a knowledge graph, so as to at least solve the technical problem that in the related technologies, when using a relational database to store and manage air dangerous goods transportation rules, due to the complex relationships between entities, the query efficiency is affected during the process of determining the answer corresponding to the question.
[0007] According to one aspect of the embodiments of the present invention, a question answering method based on a knowledge graph is provided, including: obtaining a target question, where the target question includes: a question related to dangerous goods transported by air; querying a target answer that matches the target question in a target knowledge graph, where the target knowledge graph includes: a knowledge graph obtained by supplementing entities and relationships of an initial knowledge graph, the initial knowledge graph includes: a knowledge graph constructed based on transportation rules associated with the dangerous goods and a target language model, and the model type of the target language model includes: a sequence model based on an attention mechanism.
[0008] Further, the target knowledge graph is obtained in the following manner: acquiring transportation rules related to the dangerous goods to obtain a target rule set, and constructing the initial knowledge graph based on the target rule set; acquiring a question data set, where the question data set includes: S questions related to the dangerous goods, where S is a positive integer; determining a question-answer pair data set based on the target language model and the question data set, and supplementing entities and relationships in the initial knowledge graph based on the question-answer pair data set to obtain the target knowledge graph, where the question-answer pair data set includes: the S questions in the question data set and the reply answers corresponding to each question in the question data set.
[0009] Further, supplementing entities and relationships in the initial knowledge graph based on the question-answer pair data set to obtain the target knowledge graph includes: acquiring entities and relationships related to the dangerous goods in the question-answer pair data set, and adding the acquired entities and relationships to the initial knowledge graph to obtain an updated initial knowledge graph; in the updated initial knowledge graph, extracting entities and relationships related to each question in the question-answer pair data set to obtain S question entity-relationship sets; supplementing entities and relationships in the initial knowledge graph based on the S question entity-relationship sets to obtain the target knowledge graph.
[0010] Further, in the updated initial knowledge graph, extracting entities and relationships related to each question in the question-answer pair data set to obtain S question entity-relationship sets includes: converting the S questions into word vectors respectively to obtain S question vectors; converting each relationship in the updated initial knowledge graph into a word vector to obtain a relationship vector set, and stacking the relationship vector set into a matrix to obtain a relationship matrix; for the S questions, based on the relationship matrix and the question vector corresponding to each question, extracting entities and relationships related to each question in the updated initial knowledge graph to obtain the S question entity-relationship sets.
[0011] Further, for the S questions, based on the relationship matrix and the question vector corresponding to each question, extracting entities and relationships related to each question in the updated initial knowledge graph to obtain the S question entity-relationship sets includes: for the S questions, determining the similarity between the question and each relationship in the updated initial knowledge graph based on the question vector corresponding to each question and the relationship matrix to obtain a similarity set for each question; based on the similarity set for each question, extracting entities and relationships related to each question in the updated initial knowledge graph to obtain the S question entity-relationship sets.
[0012] Further, supplementing the entities and relationships in the initial knowledge graph based on the S problem entity relationship sets to obtain the target knowledge graph includes: in each problem entity relationship set, deleting the relationships existing in the initial knowledge graph to obtain S candidate entity relationship sets; scoring all the relationships in each candidate entity relationship set based on the triples associated with all the relationships in each candidate entity relationship set to obtain a scoring result, where the scoring result includes: the score value of each relationship in each candidate entity relationship set, and each triple associated with a relationship includes: this relationship and the two entities associated with this relationship; supplementing the entities and relationships in the initial knowledge graph based on the scoring result to obtain the target knowledge graph.
[0013] Further, the data types of the data in the target rule set include at least one of the following: structured data, semi-structured data, and unstructured data; constructing the initial knowledge graph based on the target rule set includes: preprocessing the data in the target rule set to obtain a processed target rule set, where the preprocessing methods include at least one of the following: data cleaning, data standardization; extracting the entities and relationships in the processed target rule set, and combining each extracted relationship with the entities associated with this relationship into a triple to obtain T triples, where T is a positive integer; constructing the initial knowledge graph based on the T triples; after constructing the initial knowledge graph based on the target rule set, it further includes: storing the initial knowledge graph in a graph database.
[0014] Further, querying for an answer matching the target question in the target knowledge graph to obtain the target answer includes: determining the similarity between the target question and each triple associated with the target knowledge graph to obtain a target similarity set; based on the target similarity set, determining the triple most similar to the target question in the target knowledge graph to obtain a target triple; determining the target answer based on the target triple.
[0015] According to another aspect of the embodiments of the present invention, there is also provided a question-answering device based on a knowledge graph, including: a first acquisition unit for acquiring a target question, where the target question includes: a question related to dangerous goods in air transportation; a query unit for querying for an answer matching the target question in the target knowledge graph to obtain a target answer, where the target knowledge graph includes: a knowledge graph obtained by supplementing the entities and relationships of the initial knowledge graph, and the initial knowledge graph includes: a knowledge graph constructed based on the transportation rules associated with the dangerous goods and a target language model, and the model type of the target language model includes: a sequence model based on an attention mechanism.
[0016] Further, the target knowledge graph is obtained through the following units: a first processing unit, configured to obtain the transportation rules related to the dangerous goods, obtain a target rule set, and construct the initial knowledge graph based on the target rule set; a second obtaining unit, configured to obtain a question data set, where the question data set includes: S questions related to the dangerous goods, where S is a positive integer; a supplement unit, configured to determine a question-answer pair data set based on the target language model and the question data set, and supplement the entities and relationships in the initial knowledge graph based on the question-answer pair data set to obtain the target knowledge graph, where the question-answer pair data set includes: the S questions in the question data set and the reply answers corresponding to each question in the question data set.
[0017] Further, the supplement unit includes: an obtaining subunit, configured to obtain the entities and relationships related to the dangerous goods in the question-answer pair data set, and add the obtained entities and relationships to the initial knowledge graph to obtain an updated initial knowledge graph; an extracting subunit, configured to extract the entities and relationships related to each question in the question-answer pair data set from the updated initial knowledge graph to obtain S question entity-relationship sets; a supplementing subunit, configured to supplement the entities and relationships in the initial knowledge graph based on the S question entity-relationship sets to obtain the target knowledge graph.
[0018] Further, the extracting subunit includes: a conversion module, configured to convert the S questions into word vectors respectively to obtain S question vectors; a processing module, configured to convert each relationship in the updated initial knowledge graph into a word vector to obtain a relationship vector set, and stack the relationship vector set into a matrix to obtain a relationship matrix; an extracting module, configured to, for the S questions, based on the relationship matrix and the question vector corresponding to each question, extract the entities and relationships related to each question from the updated initial knowledge graph to obtain the S question entity-relationship sets.
[0019] Further, the extracting module includes: a determining sub-module, configured to, for the S questions, determine the similarity between each question and each relationship in the updated initial knowledge graph based on the question vector corresponding to each question and the relationship matrix to obtain a similarity set for each question; an extracting sub-module, configured to extract the entities and relationships related to each question from the updated initial knowledge graph based on the similarity set for each question to obtain the S question entity-relationship sets.
[0020] Further, the supplementary subunit includes: a deletion module, configured to delete the relationships existing in the initial knowledge graph in each question entity relationship set to obtain S candidate entity relationship sets; a scoring module, configured to score all the relationships in each candidate entity relationship set based on the triples associated with all the relationships in each candidate entity relationship set, to obtain a scoring result, where the scoring result includes: the score value of each relationship in each candidate entity relationship set, and each triple associated with a relationship includes: this relationship and the two entities associated with this relationship; a supplementary module, configured to supplement the entities and relationships in the initial knowledge graph based on the scoring result to obtain the target knowledge graph.
[0021] Further, the data types of the data in the target rule set include at least one of the following: structured data, semi-structured data, and unstructured data; the first processing unit includes: a preprocessing subunit, configured to preprocess the data in the target rule set to obtain a preprocessed target rule set, where the preprocessing methods include at least one of the following: data cleaning, data standardization; an extraction subunit, configured to extract the entities and relationships in the preprocessed target rule set, and combine each extracted relationship and the entities associated with this relationship into a triple to obtain T triples, where T is a positive integer; a construction subunit, configured to construct the initial knowledge graph based on the T triples; after constructing the initial knowledge graph based on the target rule set, it further includes: storing the initial knowledge graph in a graph database.
[0022] Further, the query unit includes: a first determination subunit, configured to determine the similarity of each triple associated with the target question and the target knowledge graph to obtain a target similarity set; a second determination subunit, configured to determine the triple most similar to the target question in the target knowledge graph based on the target similarity set to obtain a target triple; a third determination subunit, configured to determine the target answer based on the target triple.
[0023] On the other hand, according to an embodiment of the present invention, there is also provided an electronic device, including: a processor; and a memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the question-answering method based on a knowledge graph according to any one of the above via executing the executable instructions.
[0024] On the other hand, according to an embodiment of the present invention, there is also provided a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the question-answering method based on a knowledge graph according to any one of the above.
[0025] In the present invention, a target problem is obtained, where the target problem includes: problems related to dangerous goods in air transportation; an answer matching the target problem is queried in a target knowledge graph to obtain a target answer, where the target knowledge graph includes: a knowledge graph obtained by supplementing entities and relationships of an initial knowledge graph, and the initial knowledge graph includes: a knowledge graph constructed based on transportation rules associated with dangerous goods and a target language model, and the model type of the target language model includes: a sequence model based on an attention mechanism. Furthermore, it solves the technical problem in the related art that when using a relational database to store and manage air dangerous goods transportation rules, due to the complex relationships between entities, the query efficiency is affected. In the present invention, based on the target knowledge graph (a knowledge graph established using transportation rules related to air transportation of dangerous goods and supplemented using a target language model), question and answer for air transportation of dangerous goods are realized, avoiding the situation in the related art where storing transportation rules through a relational database results in low efficiency in querying answers corresponding to questions, thereby achieving the technical effect of improving the answer efficiency for questions regarding shipping dangerous goods. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the illustrative embodiments and descriptions of the present invention are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:
[0027] Figure 1 is a flowchart of an optional question and answer method based on a knowledge graph according to an embodiment of the present invention;
[0028] Figure 2 is a schematic diagram of an entity relationship set of a question taking "lithium battery metal" as an example according to an embodiment of the present invention;
[0029] Figure 3 is a schematic diagram of an optional candidate relationship set according to an embodiment of the present invention;
[0030] Figure 4 is a flowchart of an optional construction of an initial knowledge graph according to an embodiment of the present invention;
[0031] Figure 5 is a schematic diagram of an optional knowledge graph according to an embodiment of the present invention;
[0032] Figure 6 is a schematic diagram of an optional question and answer device based on a knowledge graph according to an embodiment of the present invention;
[0033] Figure 7 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0036] For the convenience of description, some terms or nouns related to the present invention will be explained below.
[0037] A Knowledge Graph is a structured semantic topology network used to describe entities, concepts and their interrelationships in the real world. It is based on a graph structure, where nodes represent entities or concepts and edges represent the associations or attributes between entities. By integrating multi-source heterogeneous data, the Knowledge Graph supports intelligent applications such as semantic search, reasoning and question answering, and is widely used in search engines, recommendation systems and the field of artificial intelligence.
[0038] A Graph Database is a tool specifically used to store, manage and query graph-structured data. It is based on graph theory and uses nodes, edges and properties to represent and store data, and is suitable for processing highly associated complex data relationships. By storing the Knowledge Graph in a Graph Database, fast retrieval and analysis of knowledge can be achieved, supporting application scenarios such as intelligent search, recommendation systems and intelligent question answering.
[0039] Entity Relation Completion is a key technology in Knowledge Graph construction and optimization, aiming to complete the incomplete information in the Knowledge Graph by predicting or inferring the missing relationships between entities, thereby improving the integrity and practicality of the Knowledge Graph.
[0040] Single-hop question answering requires finding a triple in the knowledge graph to determine the answer entity. That is to say, there is a relationship directly between the question and the answer entity. For example, for the question "In which type of aircraft can lithium batteries be transported?", the correct answer, i.e., "cargo aircraft transportation", can be directly returned through this triple.
[0041] Multi-hop question answering (Multi-hop QA) is a complex question answering task, referring to a multi-hop question answering system based on a knowledge graph. In the present invention, multiple knowledge graph entity relationships are inferred and integrated to answer questions that require multi-step reasoning. Multi-hop question answering can "jump" in multiple steps to connect relevant information and finally deduce the correct answer.
[0042] Transformer is a sequence model based on the attention mechanism. Different from traditional recurrent neural networks (RNNs) and convolutional neural networks (CNNs), Transformer only uses the self-attention mechanism to process the input sequence and output sequence, so it can perform parallel computing, greatly improving the computing efficiency.
[0043] Large Language Model (LLM) is a large-scale natural language processing model based on deep learning technology, capable of understanding and generating human language. It learns the statistical laws and semantic representations of language by training on a vast amount of text data, thus possessing powerful language understanding and question answering capabilities.
[0044] It should be noted that the user information involved in this application (including but not limited to user device information, user personal information, etc.), the collected information and data (including but not limited to data for analysis, stored data, displayed data, question data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with the relevant laws, regulations, and standards of the relevant regions, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.
[0045] Embodiment 1
[0046] According to an embodiment of the present invention, a method embodiment of an optional question answering method based on a knowledge graph is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0047] Figure 1is a flowchart of an optional question - answering method based on a knowledge graph according to an embodiment of the present invention. As Figure 1 shown, the method includes the following steps:
[0048] Step S101, obtain a target question, where the target question includes: questions related to dangerous goods transported by air.
[0049] The above - mentioned target question may include: questions related to dangerous goods transported by air.
[0050] Step S102, query for an answer matching the target question in the target knowledge graph to obtain a target answer, where the target knowledge graph includes: a knowledge graph obtained by supplementing entities and relationships of an initial knowledge graph. The initial knowledge graph includes: a knowledge graph constructed based on transportation rules associated with dangerous goods and a target language model. The model type of the target language model includes: a sequence model based on an attention mechanism.
[0051] The above - mentioned target knowledge graph may be composed of an entity set E and a relationship set R, and is a labeled directed attribute graph. Entities and relationships can be combined into a triple form as the basic unit of the knowledge graph. In the form of (e s , r, e t ), where e s and e t respectively represent the head entity and the tail entity, and r represents the relationship between the head entity and the tail entity. Entities in the target knowledge graph may include: dangerous - goods - related entities (for example, specific transport name, packaging description, special rules, packaging class, aircraft type, label, etc.). In this embodiment, entities such as dangerous - goods name, packaging class, and special rules can be used as nodes in the target knowledge graph, and the association relationships between them are used as edges. Compared with the two - dimensional table in the related art, it can more intuitively display the connections between entities, facilitating understanding and operation. In addition, in the target knowledge graph, in addition to the basic nodes and edges, attribute values such as dangerous - goods category numbers and applicable aircraft types can be added to each element (node) to further enrich the content of the knowledge graph.
[0052] In this embodiment, during the question - answering process using the knowledge graph, since the connection relationship between entities is abstracted as an edge connecting two entities. The information of entities and edges can be used to search for the answer (response) corresponding to the question. According to the relationship distance between the question entity and the answer entity, or, according to the similarity between the target question and the triples in the target knowledge graph, search for the target answer corresponding to the target question.
[0053] To avoid the need for multi-hop question answering that requires multiple steps of reasoning and path search to find the correlation path between the question and the answer. For example, for the question "When transporting lithium batteries, what packaging differences does China Eastern Airlines need to consider?", the main entity in the question is first identified as "lithium battery", which is connected to the one-hop entity "packaging instructions" through the relationship "cargo aircraft transportation rules", and then through the relationship of the two-hop entity "operator differences" to the question answer "China Eastern Airlines packaging instruction differences". This intermediate path involves multiple entity nodes and relationships. Currently, through existing knowledge graph tools, most are limited to the single-hop question answering mode and are difficult to match complex questions. In this embodiment, based on a target language model (e.g., a large language model, a sequence model based on the attention mechanism), entities and relationships in the initial knowledge graph can be supplemented so that the knowledge graph constructed in this embodiment has the ability of multi-hop question answering.
[0054] Specifically, based on the knowledge graph of air transportation rules for dangerous goods constructed in the early stage (corresponding to the initial knowledge graph), to further improve the performance of intelligent question answering, the knowledge graph is combined with a large language model. By integrating external data (e.g., question-answer pair datasets), the initial knowledge graph is supplemented based on the large language model to make the intelligent question answering for dangerous goods transportation rules more flexible and accurate. By integrating the large language model, a deeper context understanding and learning mechanism is given to the question answering. So as to be able to understand the questions raised by users more accurately. At the same time, the knowledge graph uses its own structured information to make up for the hallucination problem of the large language model. It helps to process complex questions more effectively and enables users to obtain knowledge related to dangerous goods transportation rules more easily. Introducing the large language model to supplement entities and relationships in the knowledge graph can solve the limitations of single-hop question answering, be applicable to multi-hop question answering scenarios, communicate with users in a more natural way, and improve the question answering efficiency.
[0055] In this embodiment, through the above steps, based on the target knowledge graph (a knowledge graph established using transportation rules related to air transportation of dangerous goods and supplemented using a target language model), question answering for dangerous goods transported by air is realized, avoiding the low efficiency of querying the answer corresponding to the question in the related technology by storing transportation rules in a relational database. Thus, the technical effect of improving the answering efficiency of questions about shipping dangerous goods is achieved. Furthermore, the technical problem in the related technology of using a relational database to store and manage air transportation rules for dangerous goods and the complex entity relationships affecting the query efficiency during the process of determining the answer corresponding to the question is solved.
[0056] Optionally, the target knowledge graph is obtained as follows: acquiring transportation rules related to dangerous goods to obtain a target rule set, and constructing an initial knowledge graph based on the target rule set; acquiring a question data set, where the question data set includes: S questions related to dangerous goods, where S is a positive integer; determining a question-and-answer pair data set based on the target language model and the question data set, and supplementing entities and relationships in the initial knowledge graph based on the question-and-answer pair data set to obtain the target knowledge graph, where the question-and-answer pair data set includes: the S questions in the question data set and the corresponding reply answers for each question in the question data set.
[0057] The above-mentioned transportation rules related to dangerous goods may include: transportation rules contained in e-books, laws and regulations, cargo station operation specifications, etc. related to dangerous goods transportation rules.
[0058] In this embodiment, first, starting from e-books, laws and regulations, and cargo station operation specifications related to dangerous goods transportation rules, text mining techniques can be used, and the experience of industry experts in the civil aviation field can also be combined to extract keywords extracted from the text, so as to collect, integrate, and clean the prior knowledge scattered in various books and rule texts. To further improve the accuracy, it can also be guided and reviewed by industry experts, and natural language processing tools can be used to perform entity recognition and relationship extraction on the text information in the field of dangerous goods (transportation rules related to dangerous goods) to form structured triple data of aviation dangerous goods transportation rules. Through knowledge graph technology, the extracted entities and relationships of dangerous goods and rules can be organized and connected to form an aviation dangerous goods transportation rule knowledge graph (corresponding to the initial knowledge graph), which visualizes the complex relationships between entities related to dangerous goods (such as special transportation names, packaging instructions, special rules, packaging grades, aircraft types, labels, etc.) in the form of a graph, providing a real, comprehensive, and systematic knowledge base of dangerous goods transportation rules for researchers and practitioners.
[0059] The aviation dangerous goods transportation rule knowledge graph can be a labeled directed attribute graph, which can display entities such as special transportation names, hazard classifications, labels, and packaging grades in the knowledge base (corresponding to the target rule set) in the form of nodes, and connect the relationships between related entities with directed edges. The entities and entity association relationships extracted from the knowledge base are presented in a topological network, and analysis and application capabilities such as cargo item name data retrieval, knowledge calculation, and graph visualization are provided for dangerous goods transportation operations. The management and visualization of the aviation dangerous goods transportation rule knowledge graph use a graph database to fuse the attributes of the knowledge base and realize the storage of the cargo knowledge graph data model. Different from ordinary graph processing or in-memory databases, graph data provides complete database features, including support for ACID (the four major characteristics of database transactions, namely atomicity, consistency, isolation, and durability) transactions, cluster support, backup and failover, etc.
[0060] To enable the knowledge graph to achieve the effect of multi-hop question answering, in this embodiment, the initial knowledge graph can also be combined with a large language model (corresponding to the target language model), and external data (corresponding to the question-answer pair dataset) can be integrated into the initial knowledge graph to make the intelligent question answering of dangerous goods transportation rules more flexible and accurate. Solve the limitations of single-hop question answering and be applicable to the scenario of multi-hop question answering.
[0061] In this embodiment, based on the question-answer pair dataset, the TransE model (a knowledge graph completion model) is used to supplement the entities and relationships in the initial knowledge graph to obtain the target knowledge graph.
[0062] In this embodiment, in the waybill information, appraisal report information, and cargo security inspection and storage information in the business system, a semi-automatic mode and expert experience are used to construct a dangerous goods transportation rule question dataset (corresponding to the question dataset), and then the large language model can be called to generate the question-answer pair dataset. Specifically,
[0063] The pseudo-code of the question-answer pair dataset generation process is shown in Table 1, where F represents all the triple data of the constructed dangerous goods transportation rule knowledge graph, di represents one of the triples, and m represents a generated piece of data. Through the large language model, answers are generated for each question in the question set one by one. D QA represents the question-answer pair data extracted and summarized from the generated responses.
[0064] Table 1
[0065]
[0066] In this embodiment, D QA can also be stored in the format of <question, answer> key-value pairs, which is convenient for subsequent processing using scripts and stored as a txt format file. Then, it can be screened and proofread in combination with a preset script. Some question-answer pairs irrelevant to domain questions, question-answer pairs with unclear answers, question-answer pairs that do not answer any practical questions, and repetitive question-answer pairs are deleted. During the deletion process, domain experts can also be involved in the review and modification.
[0067] Finally, the remaining question-answer pair data stored in txt format can be processed using a preset script and stored as CSV (Comma-Separated Values) type data. The CSV type data can contain two columns, namely "question" and "answer".
[0068] On the basis of initially establishing the knowledge graph and Q&A dataset for air dangerous goods transportation rules, entities and relationships in the <question, answer> dataset can also be identified (which can be automatically extracted or combined with the experience analysis of industry experts), and systematically transformed into triples to supplement the original knowledge graph (initial knowledge graph), generating a complete knowledge graph entity relationship set to be evaluated. Since the knowledge graph is constructed by extracting and matching attributes from rule texts and can also be combined with the experience analysis of industry experts, due to the limitations of manual experience, there are invalid triples in the graph database, so there are problems of missing links and incompleteness. The process of knowledge graph completion can infer missing content using existing knowledge to obtain the target knowledge graph, making the knowledge graph more complete. The dangerous goods rule triples consist of entities, attribute relationships, and attribute value entities. The knowledge graph completion task aims to learn the triple information within the knowledge graph to infer the missing part and improve the data completeness of the knowledge graph.
[0069] Optionally, based on the Q&A pair dataset, entities and relationships in the initial knowledge graph are supplemented to obtain the target knowledge graph, including: obtaining entities and relationships related to dangerous goods in the Q&A pair dataset and adding the obtained entities and relationships to the initial knowledge graph to obtain an updated initial knowledge graph; in the updated initial knowledge graph, extracting entities and relationships related to each question in the Q&A pair dataset to obtain S question entity relationship sets; supplementing entities and relationships in the initial knowledge graph based on the S question entity relationship sets to obtain the target knowledge graph.
[0070] In this embodiment, entities and relationships related to dangerous goods in the Q&A pair data can also be extracted and added to the initial knowledge graph to obtain an updated initial knowledge graph. For example, entities and relationships in the <question, answer> dataset can be identified and transformed into triples to supplement the initial knowledge graph, generating a complete knowledge graph entity relationship set to be evaluated (corresponding to the updated initial knowledge graph).
[0071] In this embodiment, the original question node e in the Q&A pair dataset s , and among all the relationships in the knowledge graph, find the relationship "semantically directly related to the original question e" s and use it as the question entity relationship set. Since the question-answer entity relationship set, then the entities and relationships in the initial knowledge graph can be supplemented according to the question entity relationship set to obtain the target knowledge graph, so as to improve the data completeness in the target knowledge graph.
[0072] Optionally, in the updated initial knowledge graph, entities and relationships related to each question in the question-answer pair dataset are extracted to obtain S question entity-relationship sets, including: converting the S questions into word vectors respectively to obtain S question vectors; converting each relationship in the updated initial knowledge graph into a word vector to obtain a relationship vector set, and stacking the relationship vector set into a matrix to obtain a relationship matrix; for the S questions, based on the relationship matrix and the question vector corresponding to each question, in the updated initial knowledge graph, entities and relationships related to each question are extracted to obtain S question entity-relationship sets.
[0073] Taking one of the S questions, question e s as an example, the process of obtaining the question entity-relationship set corresponding to this question is illustrated as follows, specifically including:
[0074] (1) Convert each word [e1, e2, …, e s in the original question e n into a word vector, and use the TransE model (a knowledge graph completion model) to map the entities in the dangerous goods transportation rules knowledge graph (the updated initial knowledge graph) of this question into a low-dimensional vector space, and use a simple linear transformation to represent the relationships between entities. Input it into the Transformer encoder to obtain the encoded representation (corresponding to the question vector) of question e s as:
[0075] e = [e1, e2, …, e n ;
[0076] (2) For each relationship r in the updated knowledge graph, use the TransE model to convert its composition [r1, r2, …, r m into a word vector, and input it into the same Transformer encoder to obtain the encoded representation of this relationship as:
[0077] r = [r1, r2, …, r m
[0078] After that, stack the encoded representations of all relationships (corresponding to the relationship vector set) into a "relationship encoded representation matrix" R ∈ Rn×m (corresponding to the relationship matrix), where n represents the number of relationships in the knowledge graph, and m represents the dimension of the relationship embedding. When dealing with multi-hop questions, this parameter can be adjusted to achieve the completion of multi-layer entities and relationships, so as to achieve multi-hop question answering.
[0079] Finally, based on the relationship matrix and the question vector corresponding to each question, in the updated initial knowledge graph, entities and relationships related to each question are extracted to obtain S question entity-relationship sets.
[0080] Optionally, for S questions, based on the relationship matrix and the question vectors corresponding to each question, in the updated initial knowledge graph, extract the entities and relationships related to each question to obtain S question entity-relationship sets, including: for S questions, based on the question vectors and the relationship matrix corresponding to each question, determine the similarity between the question and each relationship in the updated initial knowledge graph to obtain the similarity set of each question; based on the similarity set of each question, in the updated initial knowledge graph, extract the entities and relationships related to each question to obtain S question entity-relationship sets.
[0081] Taking one of the S questions, question e s as an example, the process of obtaining the question entity-relationship set corresponding to this question is illustrated as follows. Specifically, multiply e by the relationship matrix R to obtain the semantic similarity between the question node and all relationships, and then select R max relationships and the entities corresponding to these relationships in descending order of semantic similarity as the question entity-relationship set of question e s in the whole entity and relationship complement process.
[0082] Figure 2 FIG. is a schematic diagram of an optional question entity-relationship set taking "lithium metal battery" as an example according to an embodiment of the present invention. As Figure 2 shown, "lithium battery metal" is the central entity. The black solid arrows and black circles respectively represent the existing relationships and entities in the knowledge graph. For example, the entity "miscellaneous dangerous goods" and the relationship "lithium battery metal-1: main hazard -> miscellaneous dangerous goods", and the number represents the relationship level from the central entity. The gray solid lines and gray balls respectively represent the supplementary relationships and predicted tail entities. For example, "970-3: type restriction -> UN3091".
[0083] Optionally, based on the S question entity-relationship sets, supplement the entities and relationships in the initial knowledge graph to obtain the target knowledge graph, including: in each question entity-relationship set, delete the relationships existing in the initial knowledge graph to obtain S candidate entity-relationship sets; based on the triples associated with all relationships in each candidate entity-relationship set, score all relationships in the candidate entity-relationship set to obtain a scoring result, where the scoring result includes: the score value of each relationship in each candidate entity-relationship set, and each triple associated with a relationship includes: the relationship and the two entities associated with the relationship; based on the scoring result, supplement the entities and relationships in the initial knowledge graph to obtain the target knowledge graph.
[0084] During the process of updating the initial knowledge graph, the entities and relationships updated to the initial knowledge graph may contain invalid triples. Therefore, the relationships in the problem entity-relationship set can be scored to select reliable triples to be added to the initial knowledge graph. However, since the problem entity-relationship set may include relationships that already exist in the initial knowledge graph, these relationships do not need to be scored and added to the initial knowledge graph. To improve processing efficiency, in each problem entity-relationship set, the relationships that already exist in the initial knowledge graph can be removed to obtain the candidate entity-relationship set in this problem entity-relationship set.
[0085] Taking the following problem e s and the corresponding problem entity-relationship set R max as an example for illustration:
[0086] In this embodiment, the relationships that already exist for the central entity (relationships that already exist in the initial knowledge graph) can be removed from the problem entity-relationship set R max , and the remaining relationships are used as the candidate relationship set Rp. An example is shown Figure 3 as follows.
[0087] After that, in the candidate relationship set, all relationships can be scored, and the closest relationships can be selected to complete the initial knowledge graph network. In this embodiment, the core idea of the TransE model can be applied. For example, if (e s , r, et) is a valid triple, then the vector of the head entity plus the vector of the relationship should be close to the vector of the tail entity, that is:
[0088] e s + r ≈ e t
[0089] For each relationship rp in Rp and the central entity e s , for triples in the form of (es, rp,?) (? represents the undetermined tail entity), a scoring function is defined using the TransE model. The TransE model evaluates based on the distance between vectors:
[0090] f r (e s , e t ) = ||e s + r - e t ||2
[0091] where ||·||2 represents the L2 norm (Euclidean distance). By setting a preset parameter α as the distance measurement standard, the tail entity with a predicted distance close to 0 for each query triple can be selected. If f r (e s , e t) If ≤α, then add the tail entity and relationship that meet the requirements to the initial knowledge graph, where the prediction range of the tail entity is all entities in the candidate relationship set.
[0092] Next, taking the multi-hop question as an example, the process of the entity relationship prediction and completion algorithm will be described, as shown in Table 2 specifically:
[0093] Table 2
[0094]
[0095]
[0096] After that, the principle of the Decoder in the Transformer framework can be used to decode the entity set and relationship set encoded as word vectors in the completed knowledge graph (corresponding to the target knowledge graph) into new triples to improve the data completeness of the knowledge graph.
[0097] Optionally, the data types of the data in the target rule set include at least one of the following: structured data, semi-structured data, and unstructured data; constructing the initial knowledge graph based on the target rule set includes: preprocessing the data in the target rule set to obtain the processed target rule set, where the preprocessing methods include at least one of the following: data cleaning, data standardization; extracting entities and relationships from the processed target rule set, and combining each extracted relationship with the entities associated with the relationship into triples to obtain T triples, where T is a positive integer; constructing the initial knowledge graph based on the T triples; after constructing the initial knowledge graph based on the target rule set, it further includes: storing the initial knowledge graph in the graph database.
[0098] In this embodiment, the process of constructing the key target knowledge graph may include multiple steps such as data collection, processing, modeling, and optimization. Figure 4 is a flowchart of an optional method for constructing the initial knowledge graph according to an embodiment of the present invention. As Figure 4 shown, it mainly includes: data extraction, knowledge extraction, and triple construction. Data extraction includes: extracting data from structured data, semi-structured data, and unstructured data, and the extraction strategies may include but are not limited to: attribute extraction, message parsing, and image information extraction. Knowledge extraction may include: entity extraction and relationship extraction. The construction process of the air dangerous goods transportation rules knowledge graph (corresponding to the initial knowledge graph) is as follows:
[0099] 1. Data collection and preprocessing:
[0100] Collect and prepare structured, semi-structured, and unstructured data. For example, the target rule set can be obtained through means such as OCR text recognition tools for e-books, some other data search tools, experience analysis of industry experts, and collection and collation by the project team. The specific data can include:
[0101] Structured data: It includes tabular data in relational databases such as the Directory of Dangerous Goods in Air Transport, Civil Aviation Special Data Dictionary, Cargo Packaging Grade Dictionary, and Packaging Instruction Dictionary.
[0102] Semi-structured data: Extract business data messages such as waybill declarations, manifest declarations, and cargo security inspections and releases from the business systems of major freight stations. It mainly includes message data in formats such as XML and JSON.
[0103] Unstructured data: Collect image data on civil aviation dangerous goods transportation, including dangerous goods label images, security inspection X-ray images, etc.
[0104] Initial cleaning and standardization:
[0105] Data cleaning: Remove noise, correct error values, fill in missing values, etc.
[0106] Data standardization: Unify the data formats from different sources to ensure consistency.
[0107] 2. Entity recognition and relationship extraction:
[0108] A knowledge graph consists of an entity set E and a relationship set R. In this embodiment, entities and relationships can be combined into a triple form, which is used as the basic unit of the knowledge graph. It is represented in the form of (es, r, e t )), where es and et represent the head entity and the tail entity respectively, and r represents the relationship. By mining the text in the transportation rules and integrating the experience of industry experts, nodes (entities) and relationships (part of the data is shown in Table 3) are obtained:
[0109] Table 3
[0110]
[0111]
[0112] 3. Graph database modeling:
[0113] After obtaining the entity-relationship triples related to dangerous goods transportation rules (corresponding to T triples), they can be saved in CSV format. To more effectively display the mutual relationships of various entities in the civil aviation dangerous goods transportation rules knowledge graph, a graph database can be selected for data storage. The graph database can organize data using nodes and relationships. The instances of the dangerous goods transportation rules knowledge graph stored in the graph database, where various relationships are used as edges of different types to express, and the edges of different colors represent the mutual connections of each entity. As shown in Figure XX, an example of the knowledge graph is given. With the storage and presentation functions of the graph database, the complex connection relationships between entities can be obtained more directly and clearly.
[0114] In this embodiment, the entities and relationships are mapped to the graph database, creating nodes to represent entities, edges to represent relationships, and attributes to store additional information. Figure 5 is a schematic diagram of an optional knowledge graph according to an embodiment of the present invention, taking Figure 5 as an example, (lithium battery metal, main hazard, miscellaneous dangerous goods) is a triple, representing the instance that "the main hazard of lithium batteries is miscellaneous dangerous goods". "Lithium battery" is an entity of the type of special name for transportation. "Lithium battery" has its own UN number, Chinese and English names, Chinese and English descriptions, etc. These information can be called the attributes of the entity. Similarly, "miscellaneous dangerous goods" is an entity of the location type. There is a "main hazard" relationship between "lithium battery" and "miscellaneous dangerous goods". Through entities, relationships, and attributes, understandable knowledge can be effectively organized and visually displayed.
[0115] 4. Maintenance and update:
[0116] In this embodiment, the initial knowledge graph and the target knowledge graph can also be updated. For example, the graph can be maintained regularly. Specifically, the changes in the graph data can be monitored, and the entity and relationship information can be updated in a timely manner. In this embodiment, an incremental update mechanism can be adopted to track the dynamics of the field, regularly introduce new knowledge into the knowledge graph, and eliminate outdated information.
[0117] By constructing a high-quality knowledge graph, various intelligent applications and services are supported to ensure that the finally generated knowledge graph is both accurate and has good expressive ability.
[0118] Optionally, in the target knowledge graph, query for answers that match the target question to obtain the target answer, including: determining the similarity of each triple associated with the target question and the target knowledge graph to obtain the target similarity set; based on the target similarity set, in the target knowledge graph, determine the triple most similar to the target question to obtain the target triple; and determine the target answer based on the target triple.
[0119] In this embodiment, first, the target problem can be analyzed and transformed into a representation form that can be compared with the triples in the knowledge graph. Specifically, the keywords and entities in the problem can be identified, and the type of relationship being asked in the problem can be determined. For example, the question "What is the maximum net quantity of lithium battery metals on a cargo plane?" can be parsed into the entity "lithium battery metals", the relationship type "cargo plane transportation", and the target entity attribute "maximum net quantity".
[0120] Secondly, for each triple (e, r, e) in the target knowledge graph, the entities and relationships can be encoded into vector representations to facilitate similarity calculation. For each triple in the knowledge graph, its similarity with the target problem representation can be calculated. Specifically, cosine similarity, Euclidean distance, or other similarity measurement strategies can be adopted. For example, if the problem is represented as vector P and the triple is represented as vector T, the similarity S can be obtained by calculating the cosine similarity between P and T. After calculating the similarities between all triples in the knowledge graph and the target problem, these similarity values can be collected into a set. This set contains the similarity scores of each triple in the knowledge graph related to the problem. The form of the target similarity set is set S, where each element in set S represents the similarity between a triple and the target problem. Based on the target similarity set, in the target knowledge graph, the triple most similar to the target problem is determined to obtain the target triple. For example, from the target similarity set, those triples with the highest similarity scores can be selected as the target triples. Based on the selected target triples, the information that can directly or indirectly answer the target problem is determined. For example, in the above question, if the target triple is (cargo plane transportation, lithium battery metals, maximum net quantity = 35kg), then the target answer can be "35kg".
[0121] In this embodiment, by constructing a knowledge graph of aviation dangerous goods transportation rules (corresponding to the initial knowledge graph) and combining a multi-hop question-answering system based on the Transformer entity-relationship completion framework and a large language model, the efficiency, accuracy, and user experience of querying dangerous goods transportation rules have been significantly improved. The question-answering method based on the knowledge graph provided in this embodiment can achieve the following effects:
[0122] (1) Enhanced knowledge representation ability: Since most of the civil aviation dangerous goods transportation rules in the related technologies exist in text form, the information is scattered and difficult to integrate. This embodiment adopts the triple data form of "entity-relationship-entity" and can construct a knowledge graph containing 9 categories of dangerous goods, 3,475 kinds of dangerous goods, 6,160 nodes, and 33,251 multi-factor association relationships. It can not only intuitively display the complex associations between information such as dangerous goods names, packaging levels, and special rules, but also support deeper data mining and analysis.
[0123] (2) Improve query accuracy and speed: The relational database retrieval method in related technologies is incapable of handling complex relational networks, especially when there are multiple complex relationships between entities, and multiple two-dimensional tables need to be searched simultaneously. The question-answering method based on the knowledge graph uses the connection relationship between entities to accurately and quickly return concise node information, thereby improving the accuracy and speed of the query. In addition, by supplementing the initial knowledge graph to obtain the target knowledge graph, the knowledge graph can be endowed with the ability of a multi-hop question-answering mechanism. Users can obtain complete answers from a huge rule system across multiple texts or fragments, avoiding the limitations of traditional single-hop question-answering.
[0124] (3) Intelligent question-answering interactive experience: The introduction of a large language model and reinforcement learning framework enables a more intelligent question-answering interactive experience. It can not only understand the user's natural language query intent, but also adopt multiple answer strategies based on the multi-layer entity relationships of the knowledge graph to provide more accurate answers. For example, in the actual civil aviation cargo transportation outbound business, the user only needs to enter a simple description or question to automatically infer the relevant dangerous goods transportation rules and provide detailed explanations and guidance, which greatly facilitates the operation of front-line staff.
[0125] (4) Promote multi-party collaboration and information sharing: Since dangerous goods transportation may involve multiple roles such as shipping agents, airlines, ground agents, and cargo security checkpoints, each link needs to strictly abide by the corresponding rules. Through the visual presentation of the knowledge graph, personnel from all parties can intuitively understand and query the relevant knowledge of dangerous goods transportation, ensuring the consistency and transparency of information. This not only helps to reduce misunderstandings and errors, but also promotes efficient collaboration between departments and improves overall operational efficiency.
[0126] Embodiment 2
[0127] Embodiment 2 of the present invention provides an optional question-answering device based on a knowledge graph, and each implementation unit in the question-answering device corresponds to each implementation step in Embodiment 1.
[0128] Figure 6 is a schematic diagram of an optional question-answering device based on a knowledge graph according to an embodiment of the present invention, such as Figure 6 As shown, it includes: a first acquisition unit 61 and a query unit 62.
[0129] The first acquisition unit 61 is used to acquire a target problem, wherein the target problem includes: a problem related to dangerous goods transported by air;
[0130] A query unit 62, configured to query an answer matching the target question in a target knowledge graph to obtain a target answer, where the target knowledge graph includes: a knowledge graph obtained by supplementing entities and relationships of an initial knowledge graph, and the initial knowledge graph includes: a knowledge graph constructed based on transportation rules associated with dangerous goods and a target language model, and the model type of the target language model includes: a sequence model based on an attention mechanism.
[0131] In the question-answering device based on a knowledge graph provided in the second embodiment of the present invention, a target question can be obtained through a first acquisition unit 61, where the target question includes: a question related to dangerous goods transported by air. The query unit 62 queries an answer matching the target question in the target knowledge graph to obtain a target answer, where the target knowledge graph includes: a knowledge graph obtained by supplementing entities and relationships of an initial knowledge graph, and the initial knowledge graph includes: a knowledge graph constructed based on transportation rules associated with dangerous goods and a target language model, and the model type of the target language model includes: a sequence model based on an attention mechanism. Furthermore, it solves the technical problem in the related art that when using a relational database to store and manage air dangerous goods transportation rules, due to the complex relationships between entities, the query efficiency is affected during the process of determining the answer corresponding to the question. In this embodiment, based on the target knowledge graph (a knowledge graph established using transportation rules related to air-transported dangerous goods and supplemented using a target language model), question answering for air-transported dangerous goods is realized, avoiding the situation in the related art where storing transportation rules through a relational database results in low efficiency in querying the answer corresponding to the question, thereby achieving the technical effect of improving the answer efficiency for questions about shipping dangerous goods.
[0132] Optionally, in the question-answering device based on a knowledge graph provided in the second embodiment of the present invention, the target knowledge graph is obtained through the following units: a first processing unit, configured to obtain transportation rules related to dangerous goods to obtain a target rule set, and construct an initial knowledge graph based on the target rule set; a second acquisition unit, configured to obtain a question data set, where the question data set includes: S questions related to dangerous goods, where S is a positive integer; a supplementing unit, configured to determine a question-answer pair data set based on the target language model and the question data set, and supplement entities and relationships in the initial knowledge graph based on the question-answer pair data set to obtain a target knowledge graph, where the question-answer pair data set includes: the S questions in the question data set and the reply answers corresponding to each question in the question data set.
[0133] Optionally, in the question answering device based on a knowledge graph provided in the second embodiment of the present invention, the supplement unit includes: an acquisition subunit, configured to acquire entities and relationships related to dangerous goods in the question-answer pair dataset, and add the acquired entities and relationships to the initial knowledge graph to obtain an updated initial knowledge graph; an extraction subunit, configured to extract entities and relationships related to each question in the question-answer pair dataset from the updated initial knowledge graph to obtain S question entity-relationship sets; a supplement subunit, configured to supplement the entities and relationships in the initial knowledge graph based on the S question entity-relationship sets to obtain a target knowledge graph.
[0134] Optionally, in the question answering device based on a knowledge graph provided in the second embodiment of the present invention, the extraction subunit includes: a conversion module, configured to convert the S questions into word vectors respectively to obtain S question vectors; a processing module, configured to convert each relationship in the updated initial knowledge graph into a word vector to obtain a relationship vector set, and stack the relationship vector set into a matrix to obtain a relationship matrix; an extraction module, configured to, for the S questions, based on the relationship matrix and the question vector corresponding to each question, extract entities and relationships related to each question from the updated initial knowledge graph to obtain S question entity-relationship sets.
[0135] Optionally, in the question answering device based on a knowledge graph provided in the second embodiment of the present invention, the extraction module includes: a determination sub-module, configured to, for the S questions, determine the similarity between each question and each relationship in the updated initial knowledge graph based on the question vector corresponding to each question and the relationship matrix to obtain a similarity set for each question; an extraction sub-module, configured to extract entities and relationships related to each question from the updated initial knowledge graph based on the similarity set for each question to obtain S question entity-relationship sets.
[0136] Optionally, in the question answering device based on a knowledge graph provided in the second embodiment of the present invention, the supplement subunit includes: a deletion module, configured to delete the relationships existing in the initial knowledge graph from each question entity-relationship set to obtain S candidate entity-relationship sets; a scoring module, configured to score all the relationships in each candidate entity-relationship set based on the triples associated with all the relationships in each candidate entity-relationship set to obtain a scoring result, where the scoring result includes: the score value of each relationship in each candidate entity-relationship set, and each triple associated with a relationship includes: the relationship and the two entities associated with the relationship; a supplement module, configured to supplement the entities and relationships in the initial knowledge graph based on the scoring result to obtain a target knowledge graph.
[0137] Optionally, in the question-answering device based on a knowledge graph provided in the second embodiment of the present invention, the data types of the data in the target rule set include at least one of the following: structured data, semi-structured data, and unstructured data; the first processing unit includes: a preprocessing subunit, configured to preprocess the data in the target rule set to obtain a processed target rule set, where the preprocessing methods include at least one of the following: data cleaning and data standardization; an extraction subunit, configured to extract entities and relationships from the processed target rule set, and combine each extracted relationship with the entities associated with the relationship into a triple to obtain T triples, where T is a positive integer; a construction subunit, configured to construct an initial knowledge graph based on the T triples; after constructing the initial knowledge graph based on the target rule set, it further includes: storing the initial knowledge graph in a graph database.
[0138] Optionally, in the question-answering device based on a knowledge graph provided in the second embodiment of the present invention, the query unit includes: a first determination subunit, configured to determine the similarity between the target question and each triple associated with the target knowledge graph to obtain a target similarity set; a second determination subunit, configured to determine the triple most similar to the target question in the target knowledge graph based on the target similarity set to obtain a target triple; a third determination subunit, configured to determine a target answer based on the target triple.
[0139] The above-mentioned question-answering device based on a knowledge graph may further include a processor and a memory. The above-mentioned first acquisition unit 61 and query unit 62, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions.
[0140] The above-mentioned processor includes a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and by adjusting the kernel parameters, question-answering for dangerous goods in air transportation can be realized based on the target knowledge graph (a knowledge graph established using transportation rules related to dangerous goods in air transportation and supplemented using a target language model), avoiding the situation in the related art where storing transportation rules in a relational database results in low efficiency in querying answers to questions, thereby achieving the technical effect of improving the efficiency of answering questions about dangerous goods in shipping.
[0141] The above-mentioned memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0142] According to another aspect of the embodiments of the present invention, an electronic device is further provided, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the knowledge graph-based question answering method according to any one of the above via executing the executable instructions.
[0143] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the knowledge graph-based question answering method according to any one of the above.
[0144] Figure 7 is a schematic diagram of an electronic device according to an embodiment of the present invention, as Figure 7 shown, an embodiment of the present invention provides an electronic device 70. The electronic device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the knowledge graph-based question answering method according to any one of the above.
[0145] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0146] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0147] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0148] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0149] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0150] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0151] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A question answering method based on a knowledge graph, characterized in that Including: Obtain a target question, where the target question includes: questions related to dangerous goods transported by air; Query for an answer matching the target question in the target knowledge graph to obtain a target answer, where the target knowledge graph includes: a knowledge graph obtained by supplementing the entities and relationships of an initial knowledge graph, and the initial knowledge graph includes: a knowledge graph constructed based on the transportation rules associated with the dangerous goods and a target language model, and the model type of the target language model includes: a sequence model based on an attention mechanism.
2. The question-and-answer method according to claim 1, wherein The target knowledge graph is obtained through the following method: Obtain the transportation rules related to the dangerous goods to obtain a target rule set, and construct the initial knowledge graph based on the target rule set; Obtain a question data set, where the question data set includes: S questions related to the dangerous goods, where S is a positive integer; Based on the target language model and the question data set, determine a question-and-answer pair data set, and based on the question-and-answer pair data set, supplement the entities and relationships in the initial knowledge graph to obtain the target knowledge graph, where the question-and-answer pair data set includes: the S questions in the question data set and the reply answers corresponding to each question in the question data set.
3. The Q&A method according to claim 2, wherein Based on the question-and-answer pair data set, supplement the entities and relationships in the initial knowledge graph to obtain the target knowledge graph, including: Obtain the entities and relationships related to the dangerous goods in the question-and-answer pair data set, and add the obtained entities and relationships to the initial knowledge graph to obtain an updated initial knowledge graph; In the updated initial knowledge graph, extract the entities and relationships related to each question in the question-and-answer pair data set to obtain S question entity-relationship sets; Based on the S question entity-relationship sets, supplement the entities and relationships in the initial knowledge graph to obtain the target knowledge graph.
4. The Q&A method according to claim 3, characterized in that In the updated initial knowledge graph, extract the entities and relationships related to each question in the question-and-answer pair data set to obtain S question entity-relationship sets, including: Convert the S questions into word vectors respectively to obtain S question vectors; Convert each relationship in the updated initial knowledge graph into a word vector to obtain a relationship vector set, and stack the relationship vector set into a matrix to obtain a relationship matrix; For the S questions, based on the relationship matrix and the question vector corresponding to each question, extract the entities and relationships related to each question in the updated initial knowledge graph to obtain the S question entity-relationship sets.
5. The Q&A method according to claim 4, characterized in that, For the S questions, based on the relationship matrix and the question vector corresponding to each question, extract the entities and relationships related to each question in the updated initial knowledge graph to obtain the S question entity-relationship sets, including: For the S questions, based on the question vector corresponding to each question and the relationship matrix, determine the similarity between the question and each relationship in the updated initial knowledge graph to obtain a similarity set for each question; Based on the similarity set for each question, in the updated initial knowledge graph, extract the entities and relationships related to each question to obtain S sets of question-entity relationships.
6. The Q&A method according to claim 3, wherein Supplement the entities and relationships in the initial knowledge graph based on the S sets of question-entity relationships to obtain the target knowledge graph, including: In each set of question-entity relationships, delete the relationships existing in the initial knowledge graph to obtain S sets of candidate entity relationships; Based on the triples associated with all the relationships in each set of candidate entity relationships, score all the relationships in that set of candidate entity relationships to obtain a scoring result, where the scoring result includes: the score value of each relationship in each set of candidate entity relationships, and each triple associated with a relationship includes: the relationship and the two entities associated with the relationship; Based on the scoring result, supplement the entities and relationships in the initial knowledge graph to obtain the target knowledge graph.
7. The Q&A method according to claim 2, wherein The data types of the data in the target rule set include at least one of the following: structured data, semi-structured data, and unstructured data; Construct the initial knowledge graph based on the target rule set, including: Preprocess the data in the target rule set to obtain a preprocessed target rule set, where the preprocessing methods include at least one of the following: data cleaning, data standardization; Extract the entities and relationships in the preprocessed target rule set, and combine each extracted relationship with the entities associated with the relationship into a triple to obtain T triples, where T is a positive integer; Construct the initial knowledge graph based on the T triples; After constructing the initial knowledge graph based on the target rule set, it further includes: storing the initial knowledge graph in a graph database.
8. The Q&A method according to claim 1, wherein Query for an answer matching the target question in the target knowledge graph to obtain the target answer, including: Determine the similarity of the target question to each triple associated with the target knowledge graph to obtain a target similarity set; Based on the target similarity set, determine the triple most similar to the target question in the target knowledge graph to obtain the target triple; Determine the target answer based on the target triple.
9. A question answering device based on a knowledge graph, characterized in that, Including: A first acquisition unit for acquiring a target question, where the target question includes: a question related to dangerous goods in air transportation; A query unit for querying for an answer matching the target question in the target knowledge graph to obtain the target answer, where the target knowledge graph includes: a knowledge graph obtained by supplementing the entities and relationships of the initial knowledge graph, and the initial knowledge graph includes: a knowledge graph constructed based on the transportation rules associated with the dangerous goods and the target language model, and the model type of the target language model includes: a sequence model based on an attention mechanism.
10. An electronic device, characterized in that, Comprising one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the knowledge graph-based question-answering method according to any one of claims 1 to 8.