An interactive question-answering method for chemical knowledge and an interactive question-answering device thereof
By constructing a chemical knowledge graph database and using an interactive question-and-answer method based on vector models, the problem of chemical engineers struggling to fully grasp chemical knowledge has been solved, enabling efficient information acquisition and safety accident prevention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- EAST CHINA UNIV OF SCI & TECH
- Filing Date
- 2023-05-17
- Publication Date
- 2026-05-29
AI Technical Summary
Technical personnel in the chemical industry often struggle to quickly and comprehensively grasp knowledge from multiple branches of chemicals. Existing technologies are unable to efficiently address the wide range, diverse types, and large quantities of knowledge in the chemical field, leading to obstacles in efficiency and comprehensiveness when handling incidents related to hazardous chemicals.
Based on a chemical knowledge graph database, multiple field information is obtained through a data source connection module, ontology definitions and constraints are determined using a knowledge graph ontology module, and data linking and fusion are performed using a mapping configuration module to construct a chemical knowledge graph database. Semantic analysis is then conducted using a vector model to achieve interactive question answering.
It provides a highly efficient entry point for technical personnel in the chemical industry to access chemical information, answer questions and provide solutions in real time, reduce the incidence of safety accidents, and protect personal and economic property safety.
Smart Images

Figure CN116561282B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of information integration, data fusion, and knowledge graphs, and particularly to an interactive question-and-answer method for chemical knowledge, an interactive question-and-answer device for chemical knowledge, and a computer-readable storage medium. Background Technology
[0002] As a vital pillar industry driving national economic development and meeting the needs of people's production and daily life, the chemical industry places particular emphasis on balancing development and safety. Because the entire chemical industry process involves a large number of hazardous chemicals, which are diverse in type and properties, every stage of the production, transportation, storage, and use of hazardous chemicals presents safety hazards that could threaten life and property.
[0003] Currently, knowledge in the chemical industry is characterized by its wide range of sources, diverse types, and large quantity. This makes it difficult for technicians in the field to quickly and comprehensively master knowledge of various branches of chemicals. While existing technologies offer some solutions for intelligent question answering based on retrieval or deep learning matching techniques, these technologies are problematic in two ways: firstly, they do not involve knowledge of the chemical industry, making direct application difficult; secondly, they cannot efficiently address the wide range of sources, diverse types, and large quantities of knowledge in the chemical industry, often failing to understand the user's true needs and truly solve their chemical industry problems. This poses a significant obstacle to the efficiency and comprehensiveness of technicians handling emergencies, especially those involving hazardous chemicals.
[0004] To overcome the aforementioned deficiencies in existing technologies, there is an urgent need in this field for an interactive question-and-answer technology for chemical knowledge. This technology should provide chemical engineers with an efficient entry point for accessing chemical information from multiple sources, and provide real-time explanations and solutions to the problems faced by chemical engineers, thereby reducing the incidence of safety accidents and better protecting the personal safety of the people and the economic and property security of enterprises and the nation. Summary of the Invention
[0005] The following provides a brief overview of one or more aspects to offer a basic understanding of them. This overview is not an exhaustive summary of all conceived aspects, nor is it intended to identify key or decisive elements of all aspects, nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed descriptions that follow.
[0006] To overcome the aforementioned deficiencies in the existing technology, this invention provides an interactive question-and-answer method for chemical knowledge, a question-and-answer device for chemical knowledge, and a computer-readable storage medium. These methods provide chemical industry technicians with an efficient access point for obtaining chemical information from multiple sources, and offer real-time explanations and solutions to problems faced by chemical industry technicians, thereby reducing the incidence of safety accidents and better protecting the personal safety of the people and the economic and property security of enterprises and the nation.
[0007] Specifically, the interactive question-and-answer method for chemical knowledge provided by the first aspect of the present invention includes the following steps: obtaining a question involving chemical knowledge; identifying keywords in the question to determine at least one keyword, wherein the keyword includes entity keywords and / or ontology relation attribute keywords; and performing semantic analysis based on the keyword via a vector model, and outputting the result of the semantic analysis as the answer to the question.
[0008] Furthermore, the interactive question-and-answer device for chemical knowledge provided according to the second aspect of the present invention includes a memory and a processor. The memory stores computer instructions. The processor is connected to the memory and configured to execute the computer instructions stored in the memory to implement the interactive question-and-answer method for chemical knowledge provided in the first aspect of the present invention.
[0009] Furthermore, according to a third aspect of the present invention, a computer-readable storage medium is provided thereon storing computer instructions. When executed by a processor, these computer instructions implement the interactive question-and-answer method for chemical knowledge provided in the first aspect of the present invention. Attached Figure Description
[0010] The above-described features and advantages of the present invention will be better understood after reading the following detailed description of embodiments of the present disclosure in conjunction with the accompanying drawings. In the drawings, components are not necessarily drawn to scale, and components having similar related characteristics or features may have the same or similar reference numerals.
[0011] Figure 1 A flowchart illustrating a method for constructing a chemical knowledge graph database according to some embodiments of the present invention is shown;
[0012] Figure 2 A block diagram illustrating the construction of a chemical knowledge graph database according to other embodiments of the present invention is shown;
[0013] Figure 3 A flowchart illustrating the mapping relationship provided according to some embodiments of the present invention is shown;
[0014] Figure 4A schematic diagram of a data link provided according to some embodiments of the present invention is shown;
[0015] Figure 5 A flowchart illustrating an interactive question-and-answer method for chemical knowledge provided according to some embodiments of the present invention is shown.
[0016] Figure 6 This diagram illustrates a semantic analysis process for keywords according to some embodiments of the present invention.
[0017] Figure 7 This diagram illustrates a flowchart of obtaining a vector solution to a problem based on a vector model, according to some embodiments of the present invention; and
[0018] Figure 8 A structural block diagram of an interactive question-and-answer device for chemical knowledge provided according to some embodiments of the present invention is shown.
[0019] Figure Labels
[0020] 10. Data source connection module;
[0021] 11. Knowledge Graph Ontology Module;
[0022] 12 Mapping Configuration Modules;
[0023] 13. Data Linking and Fusion Module;
[0024] An interactive question-and-answer device for knowledge about 800 chemicals;
[0025] 810 memory; and
[0026] 820 processor. Detailed Implementation
[0027] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention is presented in conjunction with preferred embodiments, this does not mean that the features of the invention are limited to these embodiments. On the contrary, the purpose of describing the invention in conjunction with embodiments is to cover other options or modifications that may be derived based on the claims of the present invention. To provide a thorough understanding of the invention, many specific details will be included in the following description. The invention may also be implemented without using these details. Furthermore, to avoid confusion or obscuring the focus of the invention, some specific details will be omitted in the description.
[0028] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0029] Furthermore, the terms "upper," "lower," "left," "right," "top," "bottom," "horizontal," and "vertical" used in the following description should be understood as the orientations shown in the relevant paragraphs and accompanying drawings. These relative terms are for illustrative purposes only and do not imply that the described apparatus must be manufactured or operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0030] It is understood that although terms such as "first," "second," and "third" may be used herein to describe various components, regions, layers, and / or parts, these components, regions, layers, and / or parts should not be limited by these terms, and these terms are only used to distinguish different components, regions, layers, and / or parts. Therefore, the first components, regions, layers, and / or parts discussed below may be referred to as second components, regions, layers, and / or parts without departing from some embodiments of the present invention.
[0031] As mentioned above, knowledge in the chemical industry is characterized by its wide range of sources, diverse types, and large quantity, making it difficult for those skilled in the art to quickly and comprehensively master relevant chemical knowledge across multiple branches. While existing technologies offer some solutions for intelligent question answering based on retrieval techniques or deep learning matching techniques, these technologies are problematic in two ways: firstly, they do not involve knowledge of the chemical industry, making direct application difficult; secondly, they cannot efficiently address the wide range of sources, diverse types, and large quantities of knowledge in the chemical industry, often failing to understand the user's true needs and truly solve the user's chemical industry problems. This significantly hinders the efficiency and comprehensiveness of technical personnel in handling emergencies, especially those involving hazardous chemicals.
[0032] In order to overcome the above-mentioned deficiencies of the existing technology, there is an urgent need in the field for an interactive question-and-answer method for chemical knowledge, an interactive question-and-answer device for chemical knowledge, and a computer-readable storage medium, which can provide chemical technicians with an efficient access point to chemical information from multiple sources, and provide relevant explanations and solutions to the problems faced by chemical technicians in real time, thereby reducing the incidence of safety accidents and better protecting the personal safety of people and the economic and property security of enterprises and the country.
[0033] In some non-limiting embodiments, the interactive question-and-answer method for chemical knowledge provided in the first aspect of the present invention can be implemented based on the interactive question-and-answer device for chemical knowledge provided in the second aspect of the present invention.
[0034] First, the interactive question-and-answer method for chemical knowledge provided in the first aspect of this invention is based on a chemical knowledge graph database. In some embodiments of this invention, the method for constructing the chemical knowledge graph database mainly includes the following steps:
[0035] S110: Retrieves multiple fields of information related to chemical knowledge from the data source.
[0036] Specifically, you can refer to Figure 2 , Figure 2 A block diagram illustrating the construction of a chemical knowledge graph database according to other embodiments of the present invention is shown.
[0037] like Figure 2 As shown, in some embodiments of the present invention, the data source connection module 10 can connect to multiple external data sources to facilitate users in quickly obtaining various types of file data from different departments and fields within an enterprise. The file data from the connectable data sources can include various unstructured or semi-structured data, as well as direct data files, such as JSON (JavaScript Object Notation) files, comma-separated value (CSV) files, and various structured database files. The various drivers in the data source connection module 10 can directly connect to, load, and read the aforementioned multiple data sources and access data information. For example, for structured databases, operations such as selecting the database, selecting the data table, and selecting fields can be performed. For direct data files, such as Excel and CSV files, data file upload, column selection, and saving operations are provided, thereby obtaining multiple field information related to hazardous chemicals from the aforementioned multiple data sources. Optionally, the drivers may include MySQL, PostgreSQL Drivers for data sources such as Microsoft Access, SQLite, DB2, DM, Oracle, SQL Server, Excel, and CSV files.
[0038] Furthermore, in some preferred embodiments, if the present invention is subsequently applied to intelligent interactive question-and-answer of chemical knowledge, the data source connection module 10 can also collect a large number of user questions and determine the classification of these question types. For example, user questions may include questions about hazardous chemicals raised by various departments and production teams of an enterprise, and knowledge extraction methods are used to extract knowledge from different data sources by referring to the hazardous chemical knowledge graph ontology. Using the database data of the chemical enterprise or regulatory unit as the data basis for interactive question-and-answer, multiple field information of chemical knowledge involved in these sub-databases related to user questions is obtained, where the field information may include equipment information, product information, etc. Optionally, the obtained field information can be transmitted to the mapping configuration module 12 for subsequent mapping configuration operations. The data connection and fusion module 13 can then extract knowledge from the data and answer data involved in the questions according to the mapping configuration table and persistently store it in the graph database, forming the knowledge and data basis for interactive question-and-answer.
[0039] S120: Based on the ontology definition and ontology constraints in the knowledge graph, obtain the mapping configuration record represented by semantic triples corresponding to the information of each field.
[0040] like Figure 2 As shown, in some embodiments of the present invention, the ontology definition and ontology constraints of the subsequently acquired fields can be determined first through the knowledge graph ontology module 11.
[0041] Specifically, the knowledge graph ontology module 11 can be a hazardous chemical knowledge graph ontology that conforms to the Web Ontology Language (OWL) specification and the Resource Description Framework (RDFS) specification. The knowledge graph ontology module 11 mainly includes ontology definitions and ontology constraints. Here, ontology definitions can include category attributes, relation attributes, and / or data attributes of entity classes in the ontology, and ontology constraints can be used to define the mutual constraint relationships among the above three ontology definitions, such as mutual exclusion relationships.
[0042] For example, in some optional embodiments of the present invention, the class definition of the hazardous chemicals ontology may or may not include specific types of entity classes such as "construction workers," "personnel," or "human" and their constraints. The knowledge graph ontology module 11 can be defined as an ontology in the hazardous chemicals domain, but does not strictly constrain the specific ontology form. For example, ontology 1 can define ontology classes including "construction workers," "production workers," "transport workers," "gaseous chemicals," "liquid chemicals," and "solid chemicals," "limited company," "group company," "four-wheeled truck," and "eight-wheeled truck," etc., and these similar classes have mutual exclusion constraints. Ontology 2 can define ontology classes such as "personnel," "chemicals," "chemical enterprises," and "vehicles," and there are no ontology constraints between them. Therefore, as long as the ontology can summarize the abstract concepts involved in the above multiple data sources and / or the problems of sorting and classifying, and can complete the subsequent mapping configuration process, the forms of ontology 1 and ontology 2 can both meet the requirements.
[0043] Furthermore, when the ontology constraint includes the mutual exclusion constraint relationship between the three ontology definitions mentioned above, that is, there is a mutual exclusion constraint between the ontology class corresponding to ontology 1 and the ontology class corresponding to ontology 2, then an entity that has been defined as ontology 1 cannot be defined as ontology 2.
[0044] The knowledge graph ontology module 11 can transmit the ontology definitions and ontology constraints in the generated hazardous chemicals domain knowledge graph to the mapping configuration module 12. The mapping configuration module 12 combines the multiple field information related to chemical knowledge obtained from the data source connection module 10, organizes and classifies the multiple data sources, and / or the questions that users may or frequently encounter, and maps the ontology content to the data source / database fields one by one, and persists the mapping configuration.
[0045] Specifically, such as Figure 2 As shown, select the desired data source from multiple connected data sources and / or sub-databases related to the user's issue, and retrieve the required data columns or data fields from the selected data source. Then, map the retrieved data field information to the ontology content. See details below. Figure 3 , Figure 3 A flowchart illustrating the mapping relationships provided according to some embodiments of the present invention is shown. For example... Figure 3 As shown, step S120 above may further include steps S121 to S123.
[0046] S121: Based on the category attribute, relation attribute, and / or data attribute of the ontology, determine at least one mapping rule corresponding to the field information.
[0047] In some optional embodiments, the mapping rule can be a semantic triple in the form of subject-verb-object. Specifically, based on the information of the first field, which includes a category attribute as the subject, the corresponding first mapping rule can be determined as <first field name, rdf:type, category attribute>, where the first field name is universal, rdf:type is a type definition descriptor, and the category attribute represents the category attribute corresponding to the first field name. For example, in this invention, if a specific field in the data source is specified as the entity type of the knowledge graph ontology in the field of hazardous chemicals (i.e., the ontology class corresponding to the objective data entity), and the first field data corresponds to chemicals in the objective world, then the first field can be defined as the "chemicals" class in the ontology content, resulting in a mapping configuration record of <first field name, rdf:type, unique identifier of the chemical class in the ontology>.
[0048] Furthermore, in some optional embodiments, the second field information is connected to the relation attribute in the ontology as a predicate (i.e., the second item in the semantic triple), which is then connected to the object in the semantic triple. This determines the corresponding second mapping rule as <second field name, relation attribute, third field name>, where the second and third field names are subject to the constraints of the aforementioned first mapping rule. For example, if all data in the second field is "chemicals" and has been specified as the "chemicals" class of the ontology through the first mapping rule, and all data in the third field is the name of a transportation company and has been specified as the "chemical company" class through the first mapping rule, then the relation attribute "has a carrier" defined in the ontology can be used as the predicate of this mapping configuration, resulting in the mapping configuration record <second field name, has a carrier, third field name>.
[0049] Furthermore, in some optional embodiments, the third field information is connected to the data attribute in the ontology as a predicate (i.e., the second item in the semantic triple), which is then connected to the object in the semantic triple. This determines the corresponding third mapping rule as <fourth field name, data attribute, fifth field name>, where the fourth field name is subject to the constraints of the first mapping rule, and the fifth field name is universal and contains numerical information. For example, if all data in the fourth field represents a chemical enterprise and has already been specified as the "chemical enterprise" class of the ontology through the first mapping rule, then the data attribute "address located at" defined in the ontology can be used as the predicate of the mapping configuration, resulting in the mapping configuration record as <fourth field name, address located at, fifth field name>.
[0050] Preferably, in some embodiments, the corresponding fourth mapping rule can be determined as <sixth field name, rdfs:label, seventh field name> based on the fourth field information, including the field name as an alias. Here, the sixth field name is subject to the constraints of the first mapping rule, rdfs:label is the specified alias label, and the seventh field name is the alias name corresponding to the sixth field name. For example, when all data in the sixth field is a company tax number and has been defined as the "chemical enterprise" ontology class by the first mapping rule, a user-friendly company name can be added to the sixth field data. This name appears in the seventh field of the data table, resulting in a mapping configuration record of <sixth field name, rdfs:label, seventh field name>. The alias label can be a human-friendly label for easy user identification.
[0051] S122: Determine the uniqueness of at least one mapping rule based on the mutual constraints between the category attributes, relation attributes, and / or data attributes.
[0052] Specifically, in response to any of the aforementioned field information being defined as a first category, first relation, or first data attribute, it will no longer be defined as a second category, second relation, or second data attribute in the same dimension as the previously defined first category, first relation, or first data attribute type, thereby determining the uniqueness of at least one of the aforementioned mapping rules. For example, when the ontology content contains attributes of the "Chemicals" class and the "Apparatus" class in the same dimension and with mutually exclusive constraints, then when a field is defined as the "Chemicals" class, that field cannot be defined as the "Apparatus" class.
[0053] S123: Combine multiple mapping configuration records corresponding to at least one mapping rule that satisfies uniqueness into a mapping relationship table.
[0054] In some optional embodiments, the mapping configuration module 12 obtains multiple mapping configuration records, and can integrate these records that conform to the above mapping rules and uniqueness into a mapping relationship table. Optionally, the mapping configuration module 12 can also classify and organize the multiple mapping configuration records to obtain multiple mapping relationship tables of different types related to chemical knowledge, and persistently store these mapping relationship tables.
[0055] S130: Link each mapping configuration record to the data source to convert each mapping configuration record into a data triplet connected to the data source.
[0056] Specifically, you can refer to Figure 4 , Figure 4 A schematic diagram of a data link provided according to some embodiments of the present invention is shown. For example... Figure 4As shown, step S130 above may further include steps S131 to S133.
[0057] S131: Retrieve the contents of the data source table in the data source.
[0058] In some alternative embodiments, data from multiple data sources can be organized into a tabular format.
[0059] S132: Connect each mapping configuration record to the corresponding data source table content to obtain multiple data triples connected to the data source.
[0060] For example, please refer to Table 1 below. As shown in Table 1, the mapping relationship table in the mapping configuration module 12 contains information such as "field name," i.e., "company name," mapping configuration records that correspond one-to-one with the ontology content, and data source connection information, i.e., "company tax number." It is important to note that the field name in the mapping relationship table is merely a name and does not contain actual table data. Furthermore, the data source connection information is used to assist the backend program in obtaining the actual data for that field; it does not include the actual data for that field itself, but rather serves to aid the data linking process. The data linking process ultimately links the actual data in the data source table with the ontology content into a data triplet. At this point, one mapping configuration record can link to many actual data triples, the specific number of which depends on the amount of data in the data source table fields.
[0061] Company tax number Main business Company Name Tax code 1 Chemical production Company A Tax code 2 Chemical logistics carrier Company B …… …… …… HS code 10000 Chemical equipment sales Company N
[0062] Table 1
[0063] S133: Based on the mapping rules, classify each mapping configuration record and divide its corresponding data triples into type definition triples, relation attribute triples, data attribute triples, and label triples.
[0064] Specifically, in some embodiments, the records in the constructed mapping table are queried and read sequentially, including the recorded hazardous chemical data source fields and / or hazardous chemical ontology information. The ontology information records the ontology class, relation attributes, and data attributes. The data source fields can be fields involved in the user's question and its answer. After completing the linking of data and semantic triples, four types of semantic triples can be ultimately linked: type definition triple <entity, rdf:type, entity type>, label triple <entity, rdfs:label, attribute value>, ontology relation attribute triple <entity, relation attribute, entity>, and ontology data attribute triple <entity, data attribute, attribute value>. These four types of triples correspond to at least one of the first to fourth mapping rules mentioned above. It is worth noting that at this point, the data used in the data linking triples is no longer the field name, but all the data contained in that field. Therefore, a single mapping configuration record in the mapping configuration table can be linked to obtain tens of thousands of triple data pairs.
[0065] For example, in the embodiment shown in Table 1 above, if the field named "Tax ID" contains 10,000 tax ID data, then the type definition triples can be obtained from <Tax ID No. 1, rdf:type, Chemical Enterprise>, <Tax ID No. 2, rdf:type, Chemical Enterprise>, ... all the way up to <Tax ID No. 10000, rdf:type, Chemical Enterprise>, a total of 10,000 pairs of type definition triples. Correspondingly, the label triples range from <Tax ID No. 1, rdfs:label, Company A> to <Tax ID No. 10000, rdfs:label, Company N>, a total of 10,000 pairs of triples. Correspondingly, the ontology relation attribute triples will have the corresponding number of fields: <a certain chemical product, has a carrier, a certain chemical enterprise>. And the ontology data attribute triples will have the corresponding number of fields: <a certain enterprise, located at, a certain address>.
[0066] S140: Construct a chemical knowledge graph database based on data triples that connect to data sources.
[0067] After completing the above data linking, the data that the chemical knowledge graph database needs to store will be obtained. For example... Figure 2 As shown, in this embodiment, a graph database, which is different from a traditional structured database, is used to perform data fusion and persistent storage on the standardized data after data linking.
[0068] Specifically, since the data sources may include chemical enterprises, government agencies, and other sectors, and this data may be stored in different types of structured databases, such as Excel and CSV files, as well as semi-structured databases, such as JSON files, data fusion can be performed on the data from these different sources. This involves standardizing the aggregation of data from different storage types, different sources, and different sectors into a single graph database. Furthermore, in this embodiment, different types of graph databases can be selected, such as Neo4j, OrientDB, ArangoDB, JanusGraph, and Dgraph, to persistently store this data as a data stream to construct a chemical knowledge graph database.
[0069] According to the knowledge extraction and graph knowledge base construction method provided by the present invention, the concept of mapping configuration is used to associate knowledge graph ontology with data from different sources, and the data and ontology are linked according to semantic triples. Finally, a knowledge graph knowledge base conforming to the top-level structure of the ontology is merged and persisted to graph data, laying the foundation for subsequent interactive question-and-answer applications.
[0070] In some embodiments of the present invention, machine learning methods can be used to train a model on the data in the aforementioned chemical knowledge graph database. Then, keywords in the user's question are identified and question-and-answer logic is constructed. The machine learning model is then used to calculate the results based on the question-and-answer logic, and finally, the interactive question-and-answer results are fed back to the user in a standardized manner. This interactive question-and-answer method, based on machine learning technology, trains a model on the knowledge base from which knowledge is extracted, transforming the links of semantic triples into vector operations, facilitating the computer program to analyze the user's intent and perform calculations to retrieve results.
[0071] Specifically, please see Figure 5 , Figure 5 A flowchart illustrating an interactive question-and-answer method for chemical knowledge provided according to some embodiments of the present invention is shown. Figure 5 As shown, in some embodiments of the present invention, the interactive question-and-answer method for chemical knowledge mainly includes the following steps:
[0072] S510: Questions involving knowledge of chemicals.
[0073] In some embodiments of this invention, a front-end platform for interactive question answering can be implemented using mainstream front-end frameworks such as Vue / React. The front-end interface can obtain the user's desired questions through various interaction methods, such as hardware human-computer interaction interfaces, network communication interfaces, or persistent hardware reading interfaces, including microphone, keyboard human-computer interaction interfaces, brain-computer interfaces, and external storage interfaces. The front-end calls the back-end interface and passes the obtained user questions about hazardous chemicals as parameters to the back-end. The back-end performs subsequent processing, achieving a separation of front-end and back-end programs. The back-end program can convert the calculation result format and then return the standardized answer to the interactive interface.
[0074] S520: Perform keyword identification on the question to determine at least one keyword.
[0075] In some embodiments of the present invention, keywords may include entity keywords and / or ontology relational attribute keywords, wherein entity keywords and ontology relational attribute keywords may be represented in the form of a triple of <head entity, relational attribute, tail entity>. For example, assuming the obtained question is "What is the main substance of storage tank No. 5", the identified entity keyword is "storage tank No. 5" and the relational attribute keyword is "main substance".
[0076] S530: Perform keyword-based semantic analysis via a vector model and output the results of the semantic analysis as the answer to the question.
[0077] In some embodiments of the present invention, the backend can use keyword matching technology to perform keyword matching on the above user questions. The keyword matching may include entity keyword matching and ontology relation attribute keyword matching, so as to correspond to the <subject, predicate, object> in the semantic triple respectively.
[0078] Specifically, please see Figure 6 , Figure 6 A schematic diagram of a semantic analysis process for keywords according to some embodiments of the present invention is shown. The above-described semantic analysis based on keywords via a vector model may further include steps S531 to S533.
[0079] S531: In response to the fact that the keywords in the question contain only one entity keyword, the first semantic analysis rule is adopted, taking the entity corresponding to the entity keyword as the head entity, linking all the ontology relation attributes corresponding to the entities in sequence, and matching the object tail entity according to the algorithm calculated by the vector model.
[0080] For example, when a user enters only "Company Name 1", the backend program only recognizes one entity keyword, "Company Name 1", which is the head entity. After that, it will link relational attributes such as "production" and "ownership", and calculate and match all tail entities in triples such as <Company Name 1, production, tail entity 1> and <Company Name 1, ownership, tail entity 2> according to the vector model.
[0081] S532: In response to a question containing two entity keywords, the second semantic analysis rule is adopted to determine the position of the two entity keywords in the question. The entity corresponding to the first entity keyword, which is in the first position, is taken as the head entity subject, and the entity corresponding to the second entity keyword, which is in the second position, is taken as the tail entity object. The predicate is matched according to the algorithm calculated by the vector model, where the predicate is the ontology relation attribute.
[0082] For example, in the question "What is toluene in storage tank A?", the entity keywords "storage tank A" and "toluene" appear simultaneously, and there are no other entity or relation attribute keywords. Then, according to the vector model, all relation attributes in the triple <storage tank A, relation attributes, toluene> are calculated and matched. That is, all predicates of the triple are queried. The relation attribute may be "main substance", "storage", etc., which mainly depends on the ontology definition and knowledge extraction data when the chemical knowledge graph database was built previously.
[0083] Furthermore, in some optional embodiments, in response to the question containing both an entity keyword and a relational attribute keyword, the relative positions of the entity keyword and the ontology relational attribute keyword in the question are analyzed, and semantic analysis is performed based on the relative positions.
[0084] Specifically, if the entity keyword precedes the ontology relation attribute keyword, the third semantic analysis rule can be adopted, using the entity corresponding to the entity keyword as the head entity, linking the ontology relation attribute keyword as the predicate, and matching the object tail entity according to the algorithm calculated by the vector model.
[0085] For example, in the question "What is the main substance of storage tank A?", the entity keyword "storage tank A" appears before the relation attribute keyword "main substance". According to the vector model, all tail entities in the triple <storage tank A, main substance, tail entity> are calculated and matched, which means all objects of the query triple.
[0086] Conversely, if the entity keyword is located after the ontology relation attribute keyword, the fourth semantic analysis rule can be adopted, using the entity corresponding to the entity keyword as the tail entity, linking the ontology relation attribute keyword as the predicate, and matching the subject head entity according to the algorithm calculated by the vector model.
[0087] For example, in the question "What is the main substance of toluene?", the entity keyword "toluene" appears after the relation attribute keyword "main substance". According to the vector model, all head entities in the triple <head entity, main substance, toluene> are calculated and matched, that is, all subjects of the query triple are queried.
[0088] Furthermore, in some alternative embodiments, if the question contains more than three entity keywords, the first semantic analysis rule can be applied to the first entity keyword that appears. For example, if the user inputs "storage tank A, toluene, ethylene oxide", then the first semantic analysis rule is only applied to "storage tank A".
[0089] If a question contains two or more entity keywords and one or more ontology relational attribute keywords, then the third or fourth semantic analysis rule mentioned above can be applied to the first entity keyword and the first ontology relational attribute. For example, for the question "What is the main substance of storage tank A, and what about storage tank B?", the entity keywords are "storage tank A" and "storage tank B", and the relational attribute keyword is "main substance". According to the vector model, all tail entities in the triple <storage tank A, main substance, tail entity> are calculated and matched.
[0090] If a question contains more than one entity keyword and more than two ontology relation attribute keywords, the first entity keyword and the first ontology relation attribute keyword will be subject to the third or fourth semantic analysis rule mentioned above. For example, for the question "What is the main substance in storage tank A, and where is toluene stored?", the entity keywords are "storage tank A" and "toluene", and the relation attribute keywords are "main substance" and "storage". Therefore, according to the vector model, all tail entities in the triple <storage tank A, main substance, tail entity> will be calculated and matched.
[0091] S533: In response to the absence of entity keywords or ontology relation attribute keywords in the question, prompt the user to re-enter the information.
[0092] In some embodiments of the present invention, machine learning models are trained primarily on the knowledge base data extracted from the aforementioned hazardous chemicals knowledge. The trained model provides the algorithmic basis for question answering using a vector computation method, i.e., given two vectors, a third vector is calculated. The third vector is calculated sequentially during practical application. With each entity vector in the machine learning model L2 norm of the difference Thresholds are determined based on experience. If the corresponding L2 norm Then the entity vector is considered to be Calculate the answer to the user's question.
[0093] Specifically, the triples <head entity, relation attribute, tail entity> obtained above are projected onto the vector space of the vector model, so that each entity and relation attribute has a corresponding vector. These vectors are called the vector model. During the training of the vector model, the target vectors corresponding to entities or relation attributes in the vector model can be adjusted to satisfy the corresponding vector operations of the link semantics of the triples <head entity, relation attribute, tail entity>. For details, please refer to [link to documentation]. Figure 7 , Figure 7 A schematic diagram of a process for obtaining a vector solution to a problem based on a vector model, according to some embodiments of the present invention, is shown.
[0094] like Figure 7 As shown, S710: Based on the sample of the link semantics of the above triple <head entity, relation attribute, tail entity>, the vector model is trained by machine learning, and the vector model trained by machine learning is used to simulate the head entity, tail entity and relation attribute.
[0095] Specifically, first, the entity vector and relation attribute vector Perform initial normalization, where the entity vector , A set of all head entities and tail entities, a relation attribute vector. , For a set of ontology relation attributes, <head entity, ontology relation attribute, tail entity> corresponds to a vector triple < , , >. Then, perform a number of operations in units of triples <head entity, relation attribute, tail entity>. The batch sampling, and the corresponding set of batch sampling triple vectors are denoted as . Subsequently, the batch sampling set was processed according to a preset probability. = Perform the operation of replacing the head and tail entity vectors to obtain the set of counterexamples. Then, based on the aforementioned batch sampling set and counterexample set, the set is determined. ; and through For each element of the above triplet , , Perform iterative updates to obtain the mathematical formula that satisfies the semantic transformation of triples into vectors. ,in, This represents the vector distance between h+r and t. Representing 0 and The value is taken as the maximum.
[0096] In the above embodiments, the distance of negative examples is introduced to minimize the difference between the distance of positive examples and the distance of negative examples. That is, the model mimics the semantic linking of triples through vector mathematical operations. This is continuously adjusted. , , Three vectors, ultimately satisfying the mathematical formula for transforming the semantics of a triplet into a vector. That is, it satisfies the link semantics of <head entity, ontology relation attribute, tail entity>.
[0097] S720: Determine the first vector corresponding to the entity keywords and the second vector corresponding to the relation attribute keywords in the problem.
[0098] S730: Perform vector operations on the first and second vectors to obtain the numerical solution for the third vector. vector.
[0099] According to the above formula By performing vector operations on two known vectors, we can find the numerical solution for the other unknown vector using these operations. vector.
[0100] S740: Logarithmic solution The vectors are validated to obtain the vector solution to the problem.
[0101] In some preferred embodiments, a deviation threshold can also be set to determine the obtained numerical solution. The vector is subtracted from all entity vectors in the above vector model. If the numerical solution... If the magnitude of the difference between the vector and the entity vector is less than the deviation threshold, then the entity keyword corresponding to the entity vector is the answer to the previous question that meets the requirements.
[0102] The present invention also provides a specific embodiment to further and more fully understand the interactive question-and-answer method for chemical knowledge based on a chemical knowledge graph database.
[0103] In this embodiment, the question is "What is the main substance of storage tank No. 5?". First, keyword identification is performed on the entities in the question, then keyword identification is performed on the relational attributes, and finally, the question intent is analyzed to determine the quantity and relative position of entity keywords and relational attribute keywords. In this embodiment, the identified entity keyword is "storage tank No. 5", and the relational attribute keyword is "main substance". The entity keyword precedes the relational attribute keyword, according to the semantics of a subject-verb-object triple. That is, given the subject "storage tank No. 5" and the predicate "main substance", the tail entity in the triple <storage tank No. 5, main substance, tail entity> is solved. Semantically, the tail entity is also the object, indicating that the question requires querying the object. Then, the intent is transformed into a vector model for calculation through intent-query conversion. The vector corresponding to the entity "storage tank No. 5" is used as the head entity, and the vector corresponding to the relational attribute "main substance" is used as the predicate vector. After vector operations, a third vector for reference in determining the outcome is obtained. Preferably, the results can be formatted and adjusted by subtracting the third vector from each entity vector in the entity vector model, and then taking the norm of the difference. A threshold can be set empirically, and all differences less than the threshold are considered valid query results. For example, with a threshold of 0.05, vector entities with differences of 0.06 and 0.071 are considered invalid query results, while those with differences of 0.04 and 0.001 are considered valid query results. Finally, the standardized results can be displayed and fed back to the user through an interactive interface in the form of text, voice, images, or video.
[0104] Although the methods described above are illustrated and depicted as a series of actions for the sake of simplicity, it should be understood and appreciated that these methods are not limited by the order of the actions, as some actions may occur in a different order and / or concurrently with other actions from the illustrations and descriptions herein or not illustrated and described herein but which may be understood by those skilled in the art, according to one or more embodiments.
[0105] This concludes the introduction of an interactive question-and-answer method for chemical knowledge provided by the first aspect of the present invention. The second aspect of the present invention also provides an interactive question-and-answer device for chemical knowledge. Please refer to... Figure 8 , Figure 8 A structural block diagram of an interactive question-and-answer device for chemical knowledge provided according to some embodiments of the present invention is shown.
[0106] like Figure 8The interactive question-and-answer device 800 for chemical knowledge shown includes a memory 810 and a processor 820. The memory 810 includes, but is not limited to, the computer-readable storage medium described in the third aspect of the present invention, on which computer instructions are stored. The processor 820 is connected to the memory 810 and configured to execute the computer instructions stored in the memory to implement the interactive question-and-answer method for chemical knowledge provided in the first aspect of the present invention.
[0107] Furthermore, those skilled in the art will understand that the embodiments of the interactive question-and-answer methods for chemical knowledge described in the first aspect above are merely some non-limiting implementations provided by the present invention, intended to clearly demonstrate the main concepts of the present invention and provide some specific solutions that are convenient for the public to implement, rather than being used to limit all functions or all working methods of the interactive question-and-answer device 800 for chemical knowledge. Similarly, the interactive question-and-answer device 800 for chemical knowledge is also merely some non-limiting implementations provided by the present invention, and does not constitute a limitation on the executing subject or execution order of each step in these interactive question-and-answer methods for chemical knowledge. Any solutions implemented by adding, subtracting, and / or replacing steps in the prior art according to the principles of the present invention should be included within the protection scope of the present invention.
[0108] In summary, this invention provides an interactive question-and-answer method for chemical knowledge, a question-and-answer device for chemical knowledge, and a computer-readable storage medium. This invention's interactive question-and-answer construction method, based on knowledge extraction, employs keyword identification of hazardous chemical entities and relational attributes in user questions; it analyzes user semantics by analyzing the number and relative position of keywords. By establishing and training a vector model of entities and ontology relational attributes, it obtains feedback answers to user questions through vector computation, and also includes standardizing the type of backend return values. This invention can provide chemical industry technicians with an efficient entry point for obtaining chemical information from multiple sources, and provides real-time explanations and solutions to problems faced by chemical industry technicians, thereby reducing the incidence of safety accidents and better protecting the personal safety of the people and the economic and property security of enterprises and the nation.
[0109] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An interactive question-and-answer method for chemical knowledge, characterized in that, Includes the following steps: Questions involving knowledge of chemicals; The problem is subjected to keyword identification to determine at least one keyword, wherein the keyword includes entity keywords and / or ontology relation attribute keywords, wherein the triples <head entity, relation attribute, tail entity> represented by the entity keywords and the ontology relation attribute keywords correspond to semantic triples <subject, predicate, object> respectively; In response to the fact that the keywords in the question contain only one entity keyword, the first semantic analysis rule is adopted, taking the entity corresponding to the entity keyword as the head entity, sequentially linking all the ontology relation attributes corresponding to the entity, and matching the object tail entity according to the algorithm calculated by the vector model; In response to the presence of two entity keywords in the question, a second semantic analysis rule is applied to determine the positions of the two entity keywords in the question. The entity corresponding to the first entity keyword, which appears earlier, is designated as the head entity subject, and the entity corresponding to the second entity keyword, which appears later, is designated as the tail entity object. Predicate matching is then performed according to the algorithm calculated by the vector model, where the predicate is an ontology relation attribute; and / or If the entity keyword or the ontology relation attribute keyword does not appear in the question, the user is prompted to re-enter the information. Project the triple <head entity, relation attribute, tail entity> onto the vector space of the vector model, so that each entity and relation attribute has a one-to-one vector correspondence; Based on the sample of the link semantics of the triple <head entity, relation attribute, tail entity>, the vector model is trained by machine learning, and the vector model trained by machine learning is used to simulate the head entity, the tail entity and the relation attribute. Determine the first vector corresponding to the entity keyword in the problem and the second vector corresponding to the relation attribute keyword; Perform vector operations on the first and second vectors to obtain the numerical solution for the third vector. Vectors; and For the numerical solution The vectors are validated to obtain the vector solution to the problem; and The result of the semantic analysis is output as the answer to the question.
2. The interactive question-and-answer method as described in claim 1, characterized in that, The step of identifying keywords in the problem to determine at least one keyword further includes: In response to the fact that the question contains both an entity keyword and a relational attribute keyword, the relative positions of the entity keyword and the ontology relational attribute keyword in the question are analyzed, and semantic analysis is performed based on the relative positions.
3. The interactive question-and-answer method as described in claim 2, characterized in that, The step of performing semantic analysis based on the relative position includes: In response to the entity keyword preceding the ontology relation attribute keyword, a third semantic analysis rule is adopted, using the entity corresponding to the entity keyword as the head entity, linking the ontology relation attribute keyword as the predicate, and matching the object tail entity according to the algorithm calculated by the vector model; and In response to the fact that the entity keyword is located after the ontology relation attribute keyword, the fourth semantic analysis rule is adopted, taking the entity corresponding to the entity keyword as the tail entity, linking the ontology relation attribute keyword as the predicate, and matching the subject head entity according to the algorithm calculated by the vector model.
4. The interactive question-and-answer method as described in claim 3, characterized in that, After the step of identifying keywords in the question to determine at least one keyword, the method further includes: In response to the fact that the question contains more than three entity keywords, the first semantic analysis rule is applied to the first entity keyword that appears. In response to the question containing two or more entity keywords and one or more ontology relation attribute keywords, the third semantic analysis rule or the fourth semantic analysis rule is applied to the first occurrence of the entity keyword and the first occurrence of the ontology relation attribute; and / or In response to the fact that the problem contains one or more entity keywords and two or more ontology relation attribute keywords, the third semantic analysis rule or the fourth semantic analysis rule is applied to the first entity keyword and the first ontology relation attribute keyword that appear.
5. The interactive question-and-answer method as described in claim 1, characterized in that, The step of using the vector model trained by the machine learning method to simulate the head entity, the tail entity, and the relational attributes includes: For entity vectors and relation attribute vector Perform initial normalization, wherein the entity vector , The relation attribute vector is the set of all head entities and tail entities. , For a set of ontology relation attributes, the <head entity, ontology relation attribute, tail entity> corresponds to the vector triple < , , >; The quantity is calculated using the triple <head entity, relation attribute, tail entity> as the unit. The batch sampling, and the corresponding set of batch sampling triple vectors are denoted as . ; Batch sampling set with preset probability = Perform the operation of replacing the head and tail entity vectors to obtain the set of counterexamples. ; Based on the batch sampling set and the counterexample set, determine the set. ;as well as pass For each element of the triplet , , Perform iterative updates to obtain the mathematical formula that satisfies the semantic transformation of triples into vectors. ,in This represents the vector distance between h+r and t. Representing 0 and The value is taken as the maximum.
6. The interactive question-and-answer method as described in claim 5, characterized in that, The vector operations are performed on the first vector and the second vector to obtain the numerical solution of the third vector. The steps for vectorization include: According to the formula To find the numerical solution for another unknown vector, we perform vector operations on two known vectors. vector.
7. The interactive question-and-answer method as described in claim 1, characterized in that, The numerical solution The steps for verifying vectors to obtain the vector solution to the problem include: Set a deviation threshold; The numerical solution The vector is subtracted from all entity vectors in the vector model; and In response to the numerical solution If the magnitude of the difference between the vector and the entity vector is less than the deviation threshold, it is determined that the entity keyword corresponding to the entity vector is a satisfactory answer to the question.
8. An interactive question-and-answer device for chemical knowledge, characterized in that, include: Memory; as well as A processor, connected to the memory, and configured to implement an interactive question-and-answer method for chemical knowledge as described in any one of claims 1 to 7.
9. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instructions are executed by the processor, the interactive question-and-answer method for chemical knowledge as described in any one of claims 1 to 7 is implemented.