A first aid knowledge question-answering method and system based on multimodal knowledge graph

By building a multimodal first aid knowledge graph and deep learning model, the problems of inconvenience and poor results of offline first aid education and training are solved, online first aid knowledge learning and multilingual translation are realized, and first aid efficiency and citizens' first aid ability are improved.

CN114064931BActive Publication Date: 2025-05-16XINJIANG YEMA ELECTRONIC BUSINESS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111434019.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-05-16
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

The offline first aid education and training in the existing technology is inconvenient and poor, resulting in low first aid knowledge popularization rate and insufficient first aid ability for citizens.

Method used

The first aid knowledge question-and-answer method and system based on multimodal knowledge graph is adopted. By constructing a multimodal first aid knowledge graph, the entities and relationships in the question are extracted, the similarity is calculated using a deep learning model, the answers to the question are determined, and the translation is performed according to the target language selected by the user.

Benefits of technology

It realizes online learning of first aid knowledge, improves learning convenience and first aid effect, can perform multilingual translation, and improves first aid efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114064931B_ABST
    Figure CN114064931B_ABST
Patent Text Reader

Abstract

The present invention relates to a first aid knowledge question-answering method and system based on a multimodal knowledge graph, the method comprising: obtaining first aid related knowledge based on the Internet, and constructing a multimodal first aid knowledge graph based on the first aid related knowledge; obtaining a question sentence input by a user, and extracting entities and relationships in the question sentence using an entity-relationship joint extraction model; locating entities in the multimodal first aid knowledge graph according to the entities in the question sentence, and determining matching entities; calculating the similarity of all relationships between the entities in the question sentence and the matching entities using a deep learning model; determining the answer to the question sentence according to the similarity; inputting the answer into a machine translation model according to the target language selected by the user, and outputting the translated answer. The first aid knowledge question-answering method and system based on a multimodal knowledge graph of the present invention can be used to learn first aid knowledge online, improve the convenience of learning for social citizens and the first aid effect, and can perform multilingual translation, thereby improving the efficiency of first aid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of first aid knowledge question answering, and in particular to a first aid knowledge question answering method and system based on a multimodal knowledge graph. Background Art

[0002] Out-of-hospital emergency care is a basic link in the "three-ring theory" of the trauma treatment system (out-of-hospital emergency medical ring, in-hospital emergency ring, and critical illness ICU ring), and it is also a very important link. However, at present, out-of-hospital emergency care is a relatively weak link in China's medical rescue. Taking out-of-hospital cardiac arrest as an example, data shows that 500,000 people die from cardiac arrest in China every year, and the proportion of those who can be saved is less than 3%. When a patient has cardiac arrest, 90% of the patients can be saved if they receive effective cardiopulmonary resuscitation within 1 minute; 50% of the patients can be saved if they receive effective cardiopulmonary resuscitation within 4 minutes; only 10% of the patients can be saved if they receive effective cardiopulmonary resuscitation within 6 minutes; and patients with cardiac arrest can hardly be treated if cardiopulmonary resuscitation is started after 10 minutes. About 70% to 80% of cardiac arrests occur at home, on the streets, and in public places, and it is difficult for 120 ambulances to arrive within 4 minutes. For accidents such as sudden death, drowning, foreign objects stuck in the throat, car accidents, poisoning, and other types of disasters, timely and effective rescue such as social first aid and public self-help and mutual rescue can minimize casualties before the arrival of professional first aid forces. Therefore, the popularization of public first aid knowledge and skills is directly related to the life safety of the masses and social development. Once an emergency occurs, the first witness at the first aid scene must be able to bear the burden of quickly implementing first aid with the help of simple first aid equipment. However, because first aid incidents can occur at any time and in any occasion, the first witnesses at the first aid scene are often ordinary citizens, rather than first aid professionals.

[0003] However, Chinese citizens have a weak awareness of first aid, and their first aid knowledge and ability are insufficient. According to survey data, the first aid knowledge popularization rate in developed countries in the world is more than 10%. Among them, the proportion of the population in the United States who have received on-site first aid training is more than 33%, the penetration rate in Germany is 80%, and the penetration rate in France and Australia is 40%. However, the average penetration rate of first aid knowledge and skills in China is less than 1%, and it is urgent to increase emergency first aid skills training and pay attention to the popularization of first aid knowledge. After investigating the level of first aid knowledge mastery of college students in Zhejiang Province, it is concluded that the first aid knowledge level of college students in Zhejiang Province still needs to be improved. All colleges and universities should strengthen the popularization training and publicity and education of first aid knowledge to effectively improve the first aid knowledge level of college students. When investigating the mastery of first aid knowledge and skills and the existing problems of college students in medical schools, it is concluded that the first aid knowledge and skills level of medical school students is generally low, and the demand for first aid knowledge and skills is high. Certain improvements need to be made in first aid safety education to achieve the development of first aid knowledge and skills theory and practice. College students are the backbone of the future society, but in fact only 29.90% of them have participated in first aid training, and only 16.60% of them know the correct position of chest compression. College students are extremely lacking in first aid knowledge on the spot, and 95.40% of college students said they need first aid skills training. It can be seen that the popularization rate of first aid knowledge among the Chinese people is seriously low, including college students with higher education.

[0004] The level of out-of-hospital emergency care has become one of the specific social indicators that reflects the degree of modern civilization, economic development level, and comprehensive strength of a country or region, and it is also related to the life safety of every citizen. However, at present, the development of pre-hospital emergency care in China is relatively weak, citizens are not willing to provide emergency care, their awareness of emergency care is weak, and their knowledge and ability of emergency care are seriously insufficient. In view of the current situation in the field of pre-hospital emergency care, it is necessary to strengthen the popularization of emergency care knowledge, improve the emergency care ability of social citizens, and ensure that they can reasonably apply emergency care measures for self-rescue before the arrival of emergency personnel, so as to gain more rescue time for subsequent pre-hospital emergency care.

[0005] In order to popularize pre-hospital emergency knowledge and improve citizens' awareness of emergency safety, all countries have conducted research. Outside of China, the popularization of emergency knowledge has a relatively complete system and relevant legal support. Among them, 46 states in the United States have certification examination standards for first-aid witnesses. A certified first-aid witness refers to a first-aid witness who has undergone 40 to 60 hours of first-aid training and is able to respond to pre-hospital medical emergency situations. In addition to institutions that specialize in cardiopulmonary resuscitation training, all major primary and secondary schools and community service agencies have professional instructors to provide out-of-hospital first aid training. At the same time, in the United States, the law stipulates that personnel in special positions such as firefighters and drivers must have a certain level of first aid ability certification. This series of measures has made the American public have a strong awareness of first aid, and the popularization rate of basic first aid techniques has reached 89.95%, making the lifesaving rate as high as 99.87%. The French National Institute of Medical Research has proposed many measures to address the lack of public first aid knowledge, including: in driving school training, college and primary and secondary school examinations, fulfilling group obligations, and high-risk operation practices, more first aid training opportunities must be provided to train public first aid personnel, and funding should be provided to these institutions to ensure daily first aid training in schools, the military and workplaces. The units that carry out pre-hospital emergency knowledge popularization in China are mainly Chinese emergency hospitals, the Chinese Red Cross Society and the American Heart Association (AHA). The existing traditional methods of popularizing emergency knowledge are mostly through offline training, inviting professionals to conduct skills training on site. Although the effect is good and the participants are impressed, with the development of society, its defects are also becoming more prominent. The current shortcomings include the lack of freedom in learning time and place, which is very inconvenient for people who go to work every day and students. For example, many corporate staff find it difficult to take time to participate in emergency knowledge learning during work, resulting in very few people around them who know emergency knowledge, and the training time is not continuous. Even if they have participated in emergency knowledge training, they can not remember much completely. On the other hand, a training requires more manpower and material resources, and the number of trainees is also limited. In addition, the professional theoretical knowledge involved in the training is difficult to understand, especially for ethnic minority compatriots who do not understand the language fluently. The popularization of emergency knowledge for college students is often carried out through the way of conducting emergency knowledge lectures to intermittently cultivate college students' emergency quality. China has long begun to attach importance to CPR training for the public, such as the Red Cross and the emergency center provide CPR training for police, security guards, firefighters, school teachers and students; medical schools and hospitals provide CPR training for medical workers; and other institutions promote first aid. However, the reality is that after the CPR training, only a few first aid workers can still perform effective CPR. In order to solve the above problems in the field of first aid education, the present invention intends to solve them by combining first aid knowledge and knowledge graphs, natural language processing and mobile Internet technology.

[0006] Knowledge graph is the foundation and bridge for intelligent semantic retrieval, laying a solid foundation for knowledge interconnection on the World Wide Web. Compared with the traditional web page network, the nodes in the knowledge graph have changed from web pages to different types of entities, and the edges in the graph have changed from hyperlinks connecting web pages to rich semantic relationships between entities. Knowledge graphs can be divided into general domain knowledge graphs and vertical domain knowledge graphs according to their coverage. Medicine is one of the most widely used vertical fields for knowledge graphs, and it is also a hot spot in the field of artificial intelligence research in various countries. It has great application value in smart medical fields such as disease risk assessment, intelligent auxiliary diagnosis and treatment, medical quality control and medical care, knowledge question and answer, etc. Existing medical knowledge graphs include IBM's Watson Health, Ali Health's "Yizhilu" medical think tank, Sogou's AI medical knowledge graph APGC, etc. The application of medical knowledge graphs has also begun to come into people's sight in the past two years. In the medical field, typical medical knowledge graphs include SNOMED-CT, IBM's Watson Health, and China's traditional Chinese medicine knowledge graph such as Shanghai Shuguang Hospital. These traditional knowledge graphs only contain textual knowledge, and do not cover knowledge in other modalities, such as medical images, and audio and video resources for first aid publicity. Traditional knowledge graphs mainly focus on the study of entities and relationships in texts and databases, while multimodal knowledge graphs build entities in multiple modalities (such as visual modalities) and multimodal semantic relationships between entities in multiple modalities based on traditional knowledge graphs. For example, in Richpedia, the latest multimodal encyclopedia graph, first constructs a multimodal semantic relationship (rpo:imageof) between the image modality London Eye image and the text modality knowledge graph entity (DBpedia entity: London eye), and then constructs a multimodal semantic relationship (rpo:nextTo) between the image modality entity London Eye and the image modality entity Big Ben. The latest research on multimodal knowledge graphs in the medical field is medical question-answering based on multimodal knowledge graphs. The method of constructing multimodal knowledge graphs is to add visual information to the original Chinese symptom database. The specific method is to collect multiple images of entities in the Chinese symptom database from Google Images, and obtain the visual representation of the entity based on different image noise values. Finally, the visual information of the entity is well integrated into the Chinese symptom database. First aid education resources are rich, and the promotional materials for a first aid skill include text, language, video and pictures.

[0007] Knowledge graphs for emergency treatment are rarely heard of, but traditional knowledge graphs in the medical field have been widely studied. The construction process of traditional knowledge graphs in the medical field can be summarized into three steps: medical knowledge extraction, medical knowledge fusion, and medical knowledge computing. Medical knowledge extraction extracts the constituent elements of knowledge graphs such as entities, relationships, and attributes from a large amount of structured, semi-structured or unstructured medical data, and selects a reasonable and efficient way to store the elements in the knowledge base. Medical knowledge fusion integrates, disambiguates, and processes the content of the medical knowledge base, enhances the logic and expression ability within the knowledge base, and updates old knowledge or supplements new knowledge for the medical knowledge graph. Medical knowledge computing uses knowledge reasoning to infer missing facts and automatically complete disease diagnosis and treatment.

[0008] Compared with traditional named entity recognition tasks, first aid information entity recognition is quite different, specifically in the following aspects: 1) It mainly recognizes information such as sudden illnesses in first aid, descriptions of patients' sudden symptoms, first aid equipment, first aid drugs, first aid skills, first aid operation steps, and symptoms of patients during first aid; 2) The patient's sudden illness and description of sudden symptoms are also different from traditional named entity recognition, such as the use of a large number of professional terms and abbreviations, which makes the medical entity recognition task more dependent on prior knowledge; 4) Medical named entities are generally nested. However, most current named entity recognition systems can only process a single entity, ignore the entities nested inside, and cannot capture fine-grained semantic information in the text. To solve the above problems, Meizhi Ju proposed a new type of dynamically stacked (flatNER layers) neural network model, using internal entities to promote external entity detection; JuntaoY proposed a multi-head annotation strategy, which constructed a Biaffine mechanism for the encoding layer through two feedforward networks, and used SoftMax encoding to reconstruct entity recognition into a structured prediction task, and the model's accuracy on the data set was improved by 2.2%. However, multi-head annotation has the problem of sparsity of the representation matrix. Therefore, the present invention introduces a BERT (Bidirectional Encoder Representation from Transformers) pre-trained language model as the basis, searches all spans in the sentence through fragment arrangement and annotation, flexibly handles complex extraction problems in recognition tasks, and decouples sequence length to achieve joint extraction of entity relationships.

[0009] Relation extraction is not only the main task of information extraction, but also a key link in building and supplementing knowledge graphs. In recent years, joint learning models have made great progress in the research of entity relationship extraction. Compared with previous pipeline methods, joint learning can make full use of the connection between the two subtasks of entity and relationship extraction, and avoid the problem of error propagation. Existing joint models can be divided into two categories: structured prediction and multi-task learning. The structured prediction method integrates the two tasks into a unified framework and uses a decoding module to output the extracted information. Meishan Zhang, Jue Wang adopted the table filling method proposed in Makoto Miwa; Arzoo Katiyar and Suncong Zheng adopted a sequence labeling-based method; Changzhi Sun and Tsu-Jui Fu proposed a graph-based method to jointly predict entity and relationship types; Xiaoya Li converted the task into a multi-round question-answering problem. All of the above methods need to solve the global optimization problem and use beam search or reinforcement learning to perform joint decoding during reasoning. The multi-task learning method basically constructs two separate entity recognition and relationship extraction models and optimizes them together through parameter sharing. Makoto Miwa proposed using a sequence labeling model for entity prediction and a tree-based LSTM model for relation extraction. The two models share an LSTM layer for contextualized word representations, and they found that sharing parameters improved the performance of both models. Giannis Bekoulis's approach is similar, except that they modeled relation classification as a multi-label head selection problem. But these methods still perform pipeline decoding: first extract entities, and apply a relation model to the predicted entities.

[0010] Knowledge representation learning is based on the idea of ​​distributed representation. It maps the semantic information of entities (or relationships) into a low-dimensional dense real-valued vector space, so that the distance between two semantically similar objects is also close. It can achieve effective representation and calculation of knowledge graphs and is effectively applied in relational reasoning, link prediction, and entity clustering. Compared with symbolic representation, distributed representation has the advantages of improving computational efficiency, alleviating data sparsity, and realizing heterogeneous information fusion. Therefore, designing a reasonable knowledge representation learning model lays the foundation for knowledge computing and further realizing natural language understanding. Representation learning technology originated from the word2vec language learning model proposed by Mikolov et al. in 2013. Mikolov et al. found that the word vector expression trained by the language model has translation invariance. Inspired by this, Bordes et al. designed the TransE model, which adopts the h+r=t modeling assumption and uses addition as a semantic composition calculation scheme to realize relational reasoning between entities. This modeling is simple, feasible, and efficient, and has become the mainstream technology for scholars from various countries to study representation learning. However, the relationship types in knowledge graphs are usually complex, and TransE cannot accurately represent relationship types such as "one-to-many" and "many-to-many". To solve this problem, a series of research results have been produced: Wang et al. proposed the TransH model, which maps entities to the hyperplane of the corresponding relationship r, and then uses the addition method to achieve semantic synthesis mapping, effectively solving the one-to-many and many-to-many problems. However, the premise requires that the entities and relationships are in the same space, so it cannot better represent the multi-semantic characteristics of relationships in medical clinical data. Based on the design ideas of TransE and TransH, Lin et al. proposed the TransR model by changing the vector expression of entities and relationships without changing the semantic synthesis operation mode. This model maps entities and relationships to different vector spaces respectively, and then performs semantic synthesis, which is more suitable for the multi-semantic representation of relationships in medical clinical data.

[0011] Knowledge reasoning is to further explore implicit knowledge on the basis of the existing medical knowledge base, so as to enrich and expand the knowledge base. Traditional knowledge reasoning methods include descriptive logic-based reasoning, rule-based reasoning, and case-based reasoning. Although traditional knowledge reasoning methods have promoted the development of medical knowledge graphs to a certain extent, they also have defects such as low accuracy, low data utilization, and insufficient learning ability, and have not met the requirements of practical applications. With the rapid growth of the scale of medical big data, traditional knowledge reasoning methods will have problems such as information omission and prolonged diagnosis time. Artificial intelligence technology has a natural advantage in mining useful information from massive medical data, which can improve the efficiency and accuracy of knowledge reasoning. Commonly used models include artificial neural network models, genetic algorithms, and back propagation network models.

[0012] Entity linking maps entity references in natural language text to corresponding entities in the knowledge graph. For example, "high fever is one of the manifestations of heat stroke", and the two entity references "high fever" and "heat stroke" are mapped to the corresponding entities in the knowledge graph. With the development of the Internet, data is growing at an exponential level, and it is challenging to quickly obtain effective information in text. However, entity linking helps users to accurately obtain effective information in a short time. In the medical field, entity linking correctly links entity references in electronic medical records to corresponding entities in the medical knowledge graph, and can also solve the diversity and ambiguity of medical entities. Current research on knowledge graphs mainly focuses on static knowledge graphs where facts do not change over time. In addition, entity linking relies on the perfection of knowledge graphs.

[0013] There are two key steps in the entity linking task: reference identification and entity disambiguation. Entity disambiguation includes candidate entity generation and candidate entity ranking. The reference identification method here is consistent with the medical entity relationship identification method.

[0014] Candidate entity generation takes the specifically identified entity reference as the target, searches in the knowledge graph, and finds the corresponding entity as the candidate entity. This process is candidate entity generation. There are often more than one candidate entity, so the candidate entities need to be sorted. Candidate entity generation requires a high recall rate to improve the accuracy of entity linking. Generating candidate entities by building a name dictionary is the most common method, but this method has a relatively low recall rate for candidate entities. We adopt the method of building a name dictionary and construct an empirical probability entity graph through the knowledge graph, that is, the entity popularity corresponding to the entity reference is calculated through the knowledge graph to assist in candidate entity generation.

[0015] Sorting of candidate entities There are mainly traditional feature methods, binary classification methods, graph-based methods, and deep learning-based methods for sorting candidate entities. Lin Zefei et al. developed a disambiguation method that integrates multiple features, combining and weighting entity popularity, question similarity features, and similar entity reference features, sorting the results, and obtaining the entity corresponding to the entity reference, but this traditional feature method is difficult to capture fine-grained structural information and semantic information; Pilz et al. constructed an entity reference and entity vector representation method for topic information, by calculating the similarity topic distance between the entity reference context and the candidate entity context, and then inputting the topic distance into the SVM classifier for classification; Zhou Jin et al. proposed a graph-based joint feature entity linking method, using a restarted random walk algorithm to fuse multiple features into the initial edge weight calculation. This graph-based method has good interpretability for global disambiguation, but it is difficult to combine with local methods to optimize disambiguation, and it is difficult to distinguish entities with similar semantics in texts with insufficient context information; the deep learning-based method does not require manual feature annotation, and Wu Xiaochong et al. used the C-DSSM model to implement short text entity linking. The present invention proposes to input entity reference-entity similarity and empirical probability logarithm into a feed-forward neural network (FFNN), calculate local scores, add global voting scores, and achieve ranking of candidate entities.

[0016] At present, there are two ways to construct multimodal knowledge graphs. One is the traditional approach of extracting different modalities separately and fusing them to form the final multimodal graph. Li Manling proposed the first comprehensive open source multimedia knowledge extraction system, which takes a large amount of unstructured and heterogeneous multimedia data from various sources and languages ​​as input, and creates a coherent and structured knowledge base, indexing entities, relationships and events in accordance with rich fine-grained ontologies. However, the traditional method has the following problems: the dependency and correspondence between different modal features are not considered at the source, so that the final fusion result cannot well characterize the various associations contained in the multimodal data itself. Therefore, further, the graph itself has multimodal characteristics from the beginning, and the constructed multimodal graph can help understand multimodal data and complete tasks such as visual relationship recognition and cross-modal entity linking. Therefore, the second construction method is to add additional modal information on the basis of traditional knowledge graphs through network links or image searches to construct a multimodal knowledge graph. Based on the Chinese symptom database, Zhang Yingying et al. constructed a multimodal knowledge graph in the medical field by integrating image information of entities in the knowledge base. In the multimodal encyclopedia graph Richpedia, the multimodal semantic relationship (rpo:imageof) between the image modality London Eye image and the text modality knowledge graph entity (DBpedia entity: London eye) was first constructed, and then the multimodal semantic relationship (rpo:nextTo) between the image modality entity London Eye and the image modality entity Big Ben was also constructed. How to extract information from other modalities and structure unstructured visual information is the key to building a multimodal knowledge graph.

[0017] The common method of structural processing of voice information is to first define a global identifier for the voice information and convert it into text form through voice processing tools such as iFlytek, then identify the entities in the text and link them with the entities in the knowledge graph, and finally connect the voice information and the corresponding entities in the knowledge graph in a relational manner. In constructing the multimodal course knowledge graph, Qi Xiaohui first recognized the teacher's lecture voice into text through voice recognition technology, and then realized the matching link between voice and knowledge point entities through text matching, and defined the relationship between the two as an association, thus completing the multimodal entity linking work.

[0018] In order to integrate traditional knowledge graphs with image information, it is first necessary to obtain the attributes of objects in the image and the relationships between different objects. The main existing research methods include manual annotation of image information, image recognition and image description. The manual annotation method is to manually generate structured text data for the various attributes of objects in the image and the relationships between objects, so as to expand the entities and the relationships between entities in the knowledge graph; the image recognition method can extract objects in the image and their attributes using image recognition technology. However, the image recognition method cannot identify the relationship between different objects in the picture. Image description is based on image recognition, which can not only identify objects in the image but also identify the relationship between objects. Chen Shizhe and others proposed a more fine-grained control signal in their research on image description, called Abstract Scene Graph (ASG), which makes it easy to control the objects, attributes and relationships that users want to express through ASG. However, both image recognition and image description require a large number of manually annotated images. Most of the images in existing first aid publicity materials are not annotated.

[0019] The purpose of integrating video information into traditional knowledge graphs is to mine more connections between entities in the knowledge base. There are many first aid training videos in existing first aid publicity resources, such as first aid equipment teaching videos, first aid skills teaching videos, and common trauma treatment videos. Almost all the objects contained in such videos are in text-based first aid training resources, so there is little value in adding new entities to the knowledge graph, but the relationship between objects in the video is ubiquitous. By mining the relationship between objects in the video, the relationship in the knowledge graph can be further enriched. The latest multimodal knowledge graph Richpedia contains the relationship between image entities. This relationship is discovered through the description information of the image containing the two image entities, and this description information comes from the image file name in Wikipedia. For example, the file name of an airplane picture in Wikipedia is "an airplane parked on the runway waiting to fly", which includes two entities "runway" and "aircraft", and the relationship between them "stayon".

[0020] Knowledge graphs provide a more effective way to express, organize, manage and utilize massive, heterogeneous and dynamic data, making the system more intelligent and closer to human cognitive thinking.

[0021] As the most common application of knowledge graphs, intelligent question answering has also become a hot topic in current research. Intelligent question answering based on knowledge graphs mainly includes the following methods:

[0022] The earliest intelligent question answering based on knowledge graph was a template-based question answering method, which formed a query expression by constructing a set of template parameters to match the question text. The whole process does not involve question analysis, and the relevant entity relationship mapping is replaced by a preset query template. Question templates need to be manually written by domain experts, which is time-consuming and labor-intensive, and difficult to maintain. To solve this problem, research on the automatic generation of question templates has been paid attention to. Cui et al. proposed an optimization solution for large-scale automatic template generation for simple factual question answering. Abujabal et al. proposed the QUINT model, which automatically learns templates through corpus and converts natural language questions into knowledge base queries with the help of generated templates. Cocco et al. proposed an object-oriented question answering system, which uses the LinkedSpeding dataset in RDF form to learn SPARQL templates through machine learning methods on the existing training set (paired question and answer pairs).

[0023] The advantages of this method are: it can obtain relatively accurate answers and the response speed is fast. The disadvantages are: it requires a lot of manpower to proofread the template and maintain the template library. However, for the complex multi-hop problems in the question-answering field, the latest template method can also provide solutions. The current research focus of this method is more on automatic template generation to overcome the time-consuming and labor-intensive problems.

[0024] The core idea of ​​the question-answering method based on semantic parsing is to convert the query into a logical expression by parsing the components of natural language questions, and then use the semantic information of the knowledge graph to convert the logical expression into a knowledge graph query, and finally obtain the corresponding result. The logical expression is used for structured queries facing the knowledge graph to find entities and knowledge related to the entity in the knowledge base. The implementation of the semantic parsing question-answering system based on the knowledge graph requires two key steps: 1) using a semantic parser to convert the question into a semantic representation that the machine can understand and run; 2) using the semantics to generate a structured query language, query the knowledge graph, and find answers from the returned entity set. There are three types of semantic parsing methods: semantic parsing based on dictionary-grammar, semantic parsing based on semantic graph construction, and semantic parsing based on neural network. Semantic representations based on symbolic logic usually lack flexibility. In the process of analyzing the semantics of questions, they are also easily affected by the semantic gap between symbols. At the same time, it takes many steps to obtain structured semantic representations from natural language questions, and the error transmission between multiple steps affects the accuracy of question answering. In addition, a large amount of corpus is required to train neural networks, so the question-answering method based on semantic parsing is not effective.

[0025] The answer ranking method based on deep learning requires projecting the rich semantic information (characters, words, contextual relationships, entities, relationships, and attributes in the knowledge graph) contained in the question and knowledge graph into a high-dimensional vector space to obtain character vectors or word vectors, and then calculate the similarity of the vectors through a deep learning model, and then obtain the candidate ranking through the corresponding scoring mechanism to obtain the final question and answer results. The latest research is that Zhou et al. combined rules with neural networks and won the first place in the 2019 CCKS evaluation. In the method of answer ranking based on deep learning, calculating the correlation between the input question and the candidate answer entity is the core task. The current question and answer model that uses direct training of question and answer pairs has achieved good results.

[0026] Machine translation, in the translation task, hopes to obtain a translation from the source language to the target language. The main research methods of machine translation at present are statistical machine translation and neural machine translation. The statistical machine translation system performs a mathematical modeling of machine translation. It can be trained on the basis of big data. Its cost is very low because this method is language-independent. Once this model is established, it can be applied to all languages. Common statistical machine translation models include word-based machine translation modeling, distortion and reproduction rate-based models, phrase-based models, and syntax-based models. Statistical machine translation is a corpus-based method, so if the amount of data is relatively small, it will face a data sparsity problem. At the same time, it also faces another problem. Its translation knowledge comes from the automatic training of big data, and how to add expert knowledge is not yet mature. Neural network translation has risen rapidly in recent years. Compared with statistical machine translation, neural network translation is relatively simple in terms of model. It mainly consists of two parts: encoder and decoder. The encoder represents the source language as a high-dimensional vector after a series of neural network transformations. The decoder is responsible for re-decoding (translating) this high-dimensional vector into the target language. With the development of deep learning technology, neural network translation systems have surpassed statistical methods in most languages. The biggest difference between neural machine translation and statistical machine translation lies in the representation method of language text strings. Statistical machine translation is a representation model based on discrete space. All word strings are essentially composed of smaller word strings (phrases, rules). Neural machine translation is a representation model based on continuous space. The continuous space representation model can capture more hidden information. All word strings correspond to a point in a continuous space (for example, a point in a multidimensional real space). In this way, the model can be better optimized and has better generalization ability for unseen samples. Existing neural network translation models mainly include models based on recurrent neural networks, models based on convolutional neural networks, and models based on self-attention.

[0027] To sum up, the existing first aid education is generally offline education, which is inconvenient for citizens to learn and has poor learning effects. Therefore, the popularization of first aid knowledge is low, resulting in the problem of low first aid capabilities among citizens. Summary of the invention

[0028] The purpose of the present invention is to provide a first aid knowledge question and answer method and system based on a multimodal knowledge graph to solve the problems of inconvenience and poor effect of offline first aid education and training in the prior art.

[0029] To achieve the above object, the present invention provides the following solutions:

[0030] A first aid knowledge question-answering method based on a multimodal knowledge graph, comprising:

[0031] Acquire first aid related knowledge based on the Internet, and construct a multimodal first aid knowledge graph based on the first aid related knowledge; the first aid related knowledge includes emergency diseases, first aid medicines, first aid methods, first aid equipment and how to use first aid equipment; the first aid related knowledge includes text, voice, picture and video;

[0032] Obtaining a question sentence input by a user, and extracting entities and relations in the question sentence using an entity-relationship joint extraction model;

[0033] Locating entities in the multimodal first aid knowledge graph according to the entities in the question sentence, and determining matching entities; the matching entities are entities in the multimodal first aid knowledge graph that match the entities in the question sentence;

[0034] Calculating the similarity between the relationship in the question and the relationship of all the matching entities using a deep learning model;

[0035] Determining the answer to the question according to the similarity;

[0036] According to the target language selected by the user, the answer is input into a machine translation model and the translated answer is output.

[0037] Optionally, the acquiring of first aid related knowledge based on the Internet and constructing a multimodal first aid knowledge graph according to the first aid related knowledge specifically include:

[0038] Acquire first aid knowledge in text form based on the Internet and construct a traditional first aid knowledge graph;

[0039] Acquire first aid knowledge in the form of voice, image, and video based on the Internet;

[0040] The first aid knowledge in voice form, the first aid knowledge in image form and the first aid knowledge in video form are integrated into the traditional first aid knowledge graph to obtain a multimodal first aid knowledge graph.

[0041] Optionally, the step of integrating the first aid knowledge in the form of voice, the first aid knowledge in the form of images, and the first aid knowledge in the form of videos into the traditional first aid knowledge graph to obtain a multimodal first aid knowledge graph specifically includes:

[0042] Converting the voice information in the voice-based first aid knowledge into text information;

[0043] Jointly extracting the text information to obtain the entity and text of the first aid knowledge in voice form;

[0044] Integrate the entity and text of the first aid knowledge in voice form into the traditional first aid knowledge graph to obtain a first aid knowledge graph containing voice first aid knowledge;

[0045] Annotating the image information in the image-based first aid knowledge to obtain an image annotation result; the image annotation result includes object attributes in the image and the relationship between the objects;

[0046] Integrating the image annotation result into the first aid knowledge graph including voice first aid knowledge to obtain a first aid knowledge graph including voice first aid knowledge and image first aid knowledge;

[0047] Annotating the image information in the video-form first aid knowledge to obtain a video annotation result; the video annotation result includes the name of the first aid skill;

[0048] The video annotation results are integrated into the first aid knowledge graph including voice first aid knowledge and image first aid knowledge to obtain the multimodal first aid knowledge graph.

[0049] Optionally, determining the answer to the question according to the similarity specifically includes:

[0050] Selecting the relationship with the highest similarity among the similarities;

[0051] The matching entity corresponding to the relationship with the highest similarity is used as the answer to the question.

[0052] A first aid knowledge question-answering system based on a multimodal knowledge graph, comprising:

[0053] A multimodal first aid knowledge graph construction module is used to obtain first aid related knowledge based on the Internet and construct a multimodal first aid knowledge graph based on the first aid related knowledge; the first aid related knowledge includes emergency diseases, first aid medicines, first aid methods, first aid equipment and methods of using first aid equipment; the first aid related knowledge includes text, voice, pictures and videos;

[0054] An entity relationship extraction module is used to obtain a question sentence input by a user and extract entities and relationships in the question sentence using an entity relationship joint extraction model;

[0055] A matching module, used to locate entities in the multimodal first aid knowledge graph according to the entities in the question sentence, and determine matching entities; the matching entities are entities in the multimodal first aid knowledge graph that match the entities in the question sentence;

[0056] A similarity calculation module, used to calculate the similarity between the relationship in the question and the relationship of all the matching entities using a deep learning model;

[0057] An answer determination module, used to determine the answer to the question according to the similarity;

[0058] The translation module is used to input the answer into a machine translation model according to the target language selected by the user and output the translated answer.

[0059] Optionally, the multimodal first aid knowledge graph construction module specifically includes:

[0060] A traditional first aid knowledge graph construction unit is used to obtain first aid knowledge in text form based on the Internet and construct a traditional first aid knowledge graph;

[0061] A multimodal first aid knowledge acquisition unit, used to acquire first aid knowledge in the form of voice, image, and video based on the Internet;

[0062] The integration unit is used to integrate the first aid knowledge in voice form, the first aid knowledge in image form and the first aid knowledge in video form into the traditional first aid knowledge graph to obtain a multimodal first aid knowledge graph.

[0063] Optionally, the integration unit specifically includes:

[0064] A conversion subunit, used for converting the voice information in the voice-form first aid knowledge into text information;

[0065] A joint extraction subunit, used for jointly extracting the text information to obtain the entity and text of the first aid knowledge in voice form;

[0066] A voice integration subunit is used to integrate the entity and text of the first aid knowledge in voice form into the traditional first aid knowledge graph to obtain a first aid knowledge graph containing voice first aid knowledge;

[0067] An image annotation subunit is used to annotate the image information in the image-based first aid knowledge to obtain an image annotation result; the image annotation result includes the attributes of objects in the image and the relationship between the objects;

[0068] An image integration subunit is used to integrate the image annotation result into the first aid knowledge graph containing voice first aid knowledge, so as to obtain a first aid knowledge graph containing voice first aid knowledge and image first aid knowledge;

[0069] The video annotation subunit is used to annotate the image information in the video form of first aid knowledge to obtain a video annotation result; the video annotation result includes the name of the first aid skill;

[0070] The video integration subunit is used to integrate the video annotation results into the first aid knowledge graph containing voice first aid knowledge and image first aid knowledge to obtain the multimodal first aid knowledge graph.

[0071] Optionally, the answer determination module specifically includes:

[0072] A selection unit, used for selecting the relationship with the highest similarity among the similarities;

[0073] The answer determination unit is used to use the matching entity corresponding to the relationship with the highest similarity as the answer to the question.

[0074] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0075] The present invention discloses a first aid knowledge question-answering method and system based on a multimodal knowledge graph. By constructing a multimodal first aid knowledge graph, entities and relationships in a question are extracted. According to the entities in the question, entities in the multimodal first aid knowledge graph are located. The similarity of all relationships between the entities in the question and the entities in the multimodal first aid knowledge graph is calculated using a deep learning model. According to the similarity, the answer to the question is determined. According to the target language selected by the user, the answer is translated and the translated answer is output. The present invention combines multimodal knowledge in the first aid field and constructs a multimodal first aid knowledge graph on the basis of a traditional first aid knowledge graph. And a neural network machine translation model is used to translate the answer text obtained by querying in the multimodal first aid knowledge graph according to the target language selected by the user. By using the first aid knowledge question-answering method and system based on a multimodal knowledge graph of the present invention, first aid knowledge can be learned online, which improves the convenience of learning for social citizens and the first aid effect, and can perform multilingual translation, thereby improving the efficiency of first aid. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0077] Figure 1 A flowchart of a first aid knowledge question-answering method based on a multimodal knowledge graph provided by the present invention;

[0078] Figure 2 A structural diagram of a first aid knowledge question-answering system based on a multimodal knowledge graph provided by the present invention;

[0079] Figure 3 A schematic diagram of the entity labeling method provided by the present invention;

[0080] Figure 4 A schematic diagram of the structure of the joint extraction model of entities and relationships provided by the present invention;

[0081] Figure 5 A structural diagram of an intelligent question-answering system according to a specific embodiment of the present invention;

[0082] Figure 6 A smart question-answering flow chart of a specific implementation mode provided by the present invention. DETAILED DESCRIPTION

[0083] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0084] The purpose of the present invention is to provide a first aid knowledge question and answer method and system based on a multimodal knowledge graph to solve the problems of inconvenience and poor effect of offline first aid education and training in the prior art.

[0085] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0086] Figure 1 A flowchart of a first aid knowledge question-answering method based on a multimodal knowledge graph provided by the present invention is shown in FIG. Figure 1 As shown, a first aid knowledge question answering method based on a multimodal knowledge graph includes:

[0087] Step 101: Obtain first aid related knowledge based on the Internet, and construct a multimodal first aid knowledge graph based on the first aid related knowledge. The first aid related knowledge includes emergency diseases, first aid medicines, first aid methods, first aid equipment, and how to use first aid equipment; the first aid related knowledge includes text, voice, pictures, and videos.

[0088] In a specific implementation, the step 101 specifically includes:

[0089] Acquire first aid knowledge in text form based on the Internet and construct a traditional first aid knowledge graph.

[0090] Acquire first aid knowledge in voice form, image form and video form based on the Internet.

[0091] The first aid knowledge in voice form, the first aid knowledge in image form and the first aid knowledge in video form are integrated into the traditional first aid knowledge graph to obtain a multimodal first aid knowledge graph.

[0092] The step of integrating the first aid knowledge in the form of voice, the first aid knowledge in the form of images, and the first aid knowledge in the form of videos into the traditional first aid knowledge graph to obtain a multimodal first aid knowledge graph specifically includes:

[0093] The voice information in the voice-form first aid knowledge is converted into text information.

[0094] The text information is jointly extracted to obtain the entity and text of the first aid knowledge in voice form.

[0095] The entities and texts of the first aid knowledge in voice form are integrated into the traditional first aid knowledge graph to obtain a first aid knowledge graph containing voice first aid knowledge.

[0096] The image information in the image-based first aid knowledge is annotated to obtain an image annotation result; the image annotation result includes object attributes in the image and the relationship between the objects.

[0097] The image annotation result is integrated into the first aid knowledge graph including voice first aid knowledge to obtain a first aid knowledge graph including voice first aid knowledge and image first aid knowledge.

[0098] The image information in the video-format first aid knowledge is annotated to obtain a video annotation result; the video annotation result includes the name of the first aid skill.

[0099] The video annotation results are integrated into the first aid knowledge graph including voice first aid knowledge and image first aid knowledge to obtain the multimodal first aid knowledge graph.

[0100] Step 102: Obtain a question sentence input by a user, and extract entities and relations in the question sentence using an entity-relationship joint extraction model.

[0101] Step 103: Locate entities in the multimodal first aid knowledge graph according to the entities in the question and determine matching entities. The matching entities are entities in the multimodal first aid knowledge graph that match the entities in the question.

[0102] Step 104: Calculate the similarity between the relationship in the question and the relationship of all the matching entities using a deep learning model.

[0103] Step 105: Determine the answer to the question according to the similarity.

[0104] In a specific implementation, the step 105 specifically includes:

[0105] The relationship with the highest similarity among the similarities is selected.

[0106] The matching entity corresponding to the relationship with the highest similarity is used as the answer to the question.

[0107] In actual application, after determining the answer to the question, in addition to translating the text, you can also view the corresponding voice first aid knowledge, image first aid knowledge, and video first aid knowledge.

[0108] Step 106: According to the target language selected by the user, the answer is input into a machine translation model, and the translated answer is output.

[0109] Figure 2 A structural diagram of a first aid knowledge question-answering system based on a multimodal knowledge graph provided by the present invention, such as Figure 2 As shown, a first aid knowledge question-answering system based on a multimodal knowledge graph includes:

[0110] The multimodal first aid knowledge graph construction module 201 is used to obtain first aid related knowledge based on the Internet and construct a multimodal first aid knowledge graph based on the first aid related knowledge. The first aid related knowledge includes emergency diseases, first aid medicines, first aid methods, first aid equipment and how to use first aid equipment; the first aid related knowledge includes text, voice, pictures and videos.

[0111] The entity relationship extraction module 202 is used to obtain a question sentence input by a user and extract entities and relationships in the question sentence using an entity relationship joint extraction model.

[0112] The matching module 203 is used to locate the entity in the multimodal first aid knowledge graph according to the entity in the question sentence and determine the matching entity; the matching entity is the entity in the multimodal first aid knowledge graph that matches the entity in the question sentence.

[0113] The similarity calculation module 204 is used to calculate the similarity between the relationship in the question and the relationship of all the matching entities by using a deep learning model.

[0114] The answer determination module 205 is used to determine the answer to the question according to the similarity.

[0115] The translation module 206 is used to input the answer into a machine translation model according to the target language selected by the user and output the translated answer.

[0116] In a specific implementation, the multimodal first aid knowledge graph construction module 201 specifically includes:

[0117] The traditional first aid knowledge graph construction unit is used to obtain first aid knowledge in text form based on the Internet and construct a traditional first aid knowledge graph.

[0118] The multimodal first aid knowledge acquisition unit is used to acquire first aid knowledge in voice form, image form and video form based on the Internet.

[0119] The integration unit is used to integrate the first aid knowledge in voice form, the first aid knowledge in image form and the first aid knowledge in video form into the traditional first aid knowledge graph to obtain a multimodal first aid knowledge graph.

[0120] In a specific embodiment, the integration unit specifically includes:

[0121] The conversion subunit is used to convert the voice information in the voice-form first aid knowledge into text information.

[0122] The joint extraction subunit is used to jointly extract the text information to obtain the entity and text of the first aid knowledge in voice form.

[0123] The voice integration subunit is used to integrate the entity and text of the first aid knowledge in voice form into the traditional first aid knowledge graph to obtain a first aid knowledge graph containing voice first aid knowledge.

[0124] The image annotation subunit is used to annotate the image information in the image-based first aid knowledge to obtain an image annotation result, which includes the attributes of objects in the image and the relationship between objects.

[0125] The image integration subunit is used to integrate the image annotation result into the first aid knowledge graph containing voice first aid knowledge, so as to obtain the first aid knowledge graph containing voice first aid knowledge and image first aid knowledge.

[0126] The video annotation subunit is used to annotate the image information in the video form of first aid knowledge to obtain a video annotation result; the video annotation result includes the name of the first aid skill.

[0127] The video integration subunit is used to integrate the video annotation results into the first aid knowledge graph containing voice first aid knowledge and image first aid knowledge to obtain the multimodal first aid knowledge graph.

[0128] In a specific implementation, the answer determination module specifically includes:

[0129] A selection unit is used to select the relationship with the highest similarity among the similarities.

[0130] The answer determination unit is used to use the matching entity corresponding to the relationship with the highest similarity as the answer to the question.

[0131] In practical applications, since most of the knowledge in the field of pre-hospital emergency care is unstructured data, the present invention mainly studies information extraction methods for unstructured information. Information extraction methods are used in the process of constructing a multimodal emergency knowledge graph and extracting entities and relationships in the question using an entity-relationship joint extraction model. Information extraction includes two tasks: entity recognition and relationship extraction. The main methods include pipeline information extraction and joint extraction. The pipeline information extraction method is to first perform entity extraction and then perform relationship extraction. Two models need to be trained. Errors in entity extraction will affect the effect of relationship extraction. There is entity redundancy, and the inherent connection and dependency between the two tasks are ignored. Joint extraction can alleviate the problem of entity extraction errors propagating to relationship extraction. Therefore, the present invention defines a unified entity-relationship label space (the union of entity types and relationship types). The input of the joint extraction model of entities and relationships is a two-dimensional n*n table (n is the length of the input text). The joint extraction model of entities and relationships assigns labels to each unit from a unified label space. As Figure 3 As shown, the entities are the squares on the diagonal and the relations are the rectangles on either side of the diagonal.

[0132] The joint extraction method retains the complete expression of entity relationships in real extraction scenarios (overlapping relationships, directed relationships, and undirected relationships). Based on the table form, the joint extraction model of entities and relationships performs two operations: filling and decoding. First, filling the table is to predict the label of each word pair, and a double attention mechanism is used to learn the interaction between word pairs, while also adding regularized structural constraints to the table. Next, an approximate joint decoding algorithm is designed to output the final extracted entities and relationships. Experiments have shown that it can effectively identify nested entities and overlapping relationships. For the nested entity problem, after identifying the span, it is iteratively identified whether there is a nested entity. Figure 3 For example, after identifying David Perkins, the next step is to determine whether David and Perkins are entities. The joint extraction model structure of entities and relations is as follows: Figure 4 As shown, it is mainly divided into three parts, namely the encoding layer, the constraint layer and the decoding layer.

[0133] ①Encoding

[0134] First, use the pre-trained language model to obtain the vector representation of the text (h1, h2, ..., h n ), in order to solve the long-distance dependency problem of sentences, the context of the sentences is concatenated, that is, the sentences are extended to a fixed length (the default setting is 200). At the same time, a deep biaffine attention mechanism is used to better encode the directional information of the words in the table: 1) Two dimensionality-reduced multi-layer perceptrons are used to obtain information in different directions of the sentences; 2) The biaffine model is used to calculate the score vector of each word pair; 3) The predicted label is output through the Softmax function.

[0135] ②Add constraints

[0136] In fact, the predicted labels obtained in the previous step are independent of each other, and the results obtained by the joint extraction model of entities and relations are relatively poor. Intuitively, entities and relations correspond to squares and rectangles in a two-dimensional matrix, respectively, but the corresponding constraints were not explicitly defined in the previous step. In this regard, two independent constraints are added: 1) entities and undirected relations are symmetric about the diagonal; 2) if a relationship exists, the corresponding entity pair must exist, that is, the probability of the relationship label is not higher than the probability of the two entities.

[0137] ③Decoding

[0138] The decoding algorithm is mainly divided into three steps: span decoding (entity and span between entities), entity type decoding, and entity pair relationship decoding. The two-dimensional matrix is ​​predicted in a unified label space, and those exceeding the threshold are marked as spans. Finally, the entities and relationships with the highest scores are predicted.

[0139] The main function of entity linking is to use the entities in the knowledge base to disambiguate the entity references obtained from the text, and identify the mapping entity corresponding to each entity reference in the medical knowledge base. The entity reference here refers to a text representation of the entity. An entity may have many different expressions, such as full name, alias, abbreviation, etc. For example, artificial respiration is also called cardiopulmonary resuscitation in different texts, and cardiac defibrillator is also called AED in other articles. In first aid knowledge, the same objects with these different names contain the same attributes. Therefore, the present invention uses an entity linking method based on entity attributes to determine whether the entities are the same by calculating the similarity of the character strings in the name attributes of the entities. The similarity between entity names and attributes is mainly calculated by Consine distance, jaccard correlation coefficient, etc.:

[0140]

[0141]

[0142] Among them, for the medical entities given by e1 and e2, A(e) represents the attribute string of the medical entity e.

[0143] In practical applications, multimodal first aid knowledge is integrated into the traditional first aid knowledge graph. The specific process is as follows:

[0144] In first aid training resources, voice information is mainly a voice explanation of a certain first aid skill. This type of voice information is mostly contributed by first aid experts, so the voice quality is very high. The present invention first manually recognizes the voice information, identifies the first aid skills currently explained by the voice, and uses the corresponding first aid skill name to define a resource identifier for the voice file. First, use a voice recognition tool to convert the voice information into text information, then jointly extract the text information, and then add the entities and relationships of the voice information to the knowledge graph. Finally, locate the corresponding entity in the knowledge graph according to the name of the voice file and connect it through the "audioof" relationship.

[0145] In first aid training resources, the content of image information mainly includes people and first aid equipment. The objects contained are relatively simple, and the relationship between objects is relatively clear. The present invention adopts a manual annotation method for image information, including annotating the attributes of objects in the image and the relationship between objects. The present invention uses the results of manual annotation as the resource identifier of the image, extracts information from the results of manual annotation, and completes the knowledge graph. For images containing only a single object, the entities in the knowledge graph are located according to the annotation results and connected through the "imageof" relationship.

[0146] In first aid training resources, the content of most video information is first aid skills. The present invention adopts a manual annotation method for image information, including annotating the name of the first aid skill displayed in the video, locating the entity in the knowledge graph according to the annotation result, and connecting through the "videoof" relationship.

[0147] In a specific embodiment, the present invention uses a deep learning-based answer ranking method to construct a question-answering system based on a multimodal first aid knowledge graph. The specific system framework is as follows: Figure 5 shown.

[0148] The answer ranking method based on deep learning needs to project the rich semantic information (characters, words, contextual relationships, entities, relationships and attributes in the knowledge graph) contained in the question and knowledge graph into a high-dimensional vector space to obtain character vectors or word vectors, calculate the similarity of the vectors through the deep learning model, and then obtain the candidate ranking through the corresponding scoring mechanism to obtain the final question and answer result. This system uses the BERT model to compare the similarity of the relationship, so as to find the best answer in the knowledge graph. Figure 6After extracting the entities and relations of the user's question, the triple structure of the multimodal first aid knowledge graph is used. Given two elements of the triple, the information of the remaining element is retrieved in the multimodal first aid knowledge graph as the answer.

[0149] Although the triple knowledge representation form has been widely used and recognized, it has problems such as low computational efficiency when applied to the medical field. In recent years, with the significant progress of representation learning technology such as artificial intelligence, machine learning and deep learning, the semantic information in medical entities can be represented as dense low-dimensional real-valued vectors, thereby calculating the complex semantic associations between entities and relationships in low-dimensional space. The commonly used method for knowledge representation is the distance translation model, in which the distance translation model uses a distance-based scoring function to judge the rationality of facts. Representatives of the distance translation model include the translation model (TransE) and its extended complex relationship models (TransH, TransR, TransD, TransG, KG2E, etc.). TransE is the most representative distance translation model, which represents entities and relationships as vectors in the same space. The relationship vector in the triple can be regarded as the translation of the head entity vector to the tail entity vector, and satisfies the relationship: The translation model has fewer parameters, low computational complexity, and is suitable for large-scale sparse medical knowledge bases. The performance and scalability are relatively good. Therefore, the present invention uses the TransE model for knowledge representation.

[0150] The first aid health knowledge question-answering method and system based on multimodal knowledge graph of the present invention has the following advantages:

[0151] There is a gap in the existing research on first aid, and the present invention has good innovation and practicality. With the development of information technology, the media for the dissemination of first aid knowledge has gradually changed from traditional printed materials such as first aid manuals, books, and bulletin boards to information technology products such as first aid websites and medical software. In addition to traditional offline volunteer activities, the existing methods of popularizing first aid knowledge also include using the Internet to spread knowledge around first aid health subjects. Online dissemination of first aid knowledge has a wider range of popularization and low popularization costs, while offline volunteer activities can better guarantee the quality of popularization. However, in real life, users prefer to directly understand the details of first aid skills knowledge rather than a complete explanation of first aid skills. And at the scene of emergency first aid, the present invention can immediately give answers to relevant first aid knowledge based on the symptom information provided by the user. Compared with the existing methods of popularizing first aid knowledge, the present invention has a wider range of use scenarios and can more accurately understand user questions and give accurate answers. At the same time, considering the existence of multiple languages ​​in some parts of China, the present invention provides a machine translation model in the field of first aid to present first aid knowledge in different languages. The present invention can solve the problems of single first aid training methods, limited offline training effects, and weak public first aid safety awareness in the current first aid education field.

[0152] The data comes from the Urumqi Emergency Center, and the emergency information is true and reliable.

[0153] Emphasize the learning effect of users. The present invention combines multimodal knowledge in the field of first aid and constructs a multimodal first aid knowledge graph based on the traditional knowledge graph. The present invention integrates the information of first aid related voice, image and video into the traditional knowledge graph, and the answer returned to the user also includes information of other modes, so the answer is more comprehensive and easy to understand.

[0154] The present invention has multilingual characteristics. The present invention uses a neural machine translation method, uses Chinese-Uyghur parallel predictions in the field of first aid to build a neural network machine translation model, and translates the text of the answer obtained by querying in the multimodal first aid knowledge graph according to the target language selected by the user. In this way, the text in the first aid knowledge graph can be kept in Chinese characters, which is convenient for updating the knowledge graph in the future. At the same time, it meets the user's need to understand the answer.

[0155] In practical applications, a first aid health knowledge question-answering system based on a multimodal knowledge graph includes:

[0156] Multimodal first aid knowledge graph: crawl first aid related knowledge from the Internet, including emergency diseases, first aid medicines, first aid methods, first aid equipment and how to use them. The knowledge is in the form of text, voice, pictures and videos. Use first aid related knowledge to build a multimodal first aid knowledge graph.

[0157] Joint extraction model of entities and relations: Build a joint extraction model of entities and relations in natural language questions in the field of first aid, and locate entities in the multimodal first aid knowledge graph based on the extracted entities.

[0158] Relationship similarity calculation model based on deep learning: Based on the jointly extracted entities, link to the entities in the multimodal emergency knowledge graph to achieve the purpose of entity positioning. Use the deep learning model to calculate the similarity of all relationships between the extracted relations in the question and the related entities in the knowledge graph. Get the answer to the question based on the similarity score.

[0159] Machine translation model based on neural network: The neural machine translation model can at least achieve Chinese-Uyghur translation.

[0160] When a question is received from the user, the entity-relationship joint extraction model is used to extract the entities and relations in the question. Based on the entities in the question, the entities in the knowledge graph are located. Then, a deep learning model such as Bert is used to calculate the similarity between the relations in the question and all relations of the matching entities in the multimodal first aid knowledge graph. The relation with the highest similarity is selected, and the corresponding entity is obtained as the answer to the question. Then, based on the target language selected by the user, the answer is used as the input of the translation model. Finally, the result of the translation model is fed back to the user as the final answer.

[0161] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0162] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A first aid knowledge question-answering method based on a multimodal knowledge graph, characterized in that: include: Acquire first aid related knowledge based on the Internet, and construct a multimodal first aid knowledge graph based on the first aid related knowledge; the first aid related knowledge includes emergency diseases, first aid medicines, first aid methods, first aid equipment and how to use first aid equipment; the first aid related knowledge includes text, voice, picture and video; Obtaining a question sentence input by a user, and extracting entities and relationships in the question sentence using an entity-relationship joint extraction model; the input of the entity-relationship joint extraction model is a two-dimensional n*n table; the entity-relationship joint extraction model includes an encoding layer, a constraint layer, and a decoding layer; The encoding layer is used to encode the question input by the user. Specifically: a pre-trained language model is used to obtain a vector representation of the text. To address the long-distance dependency problem of sentences, the context of the sentence is concatenated, and the sentence is extended to a fixed length. At the same time, a deep dual-affine attention mechanism is used to encode the directional information of the words in the table: 1) two dimensionality-reduced multi-layer perceptrons are used to obtain information in different directions of the sentence; 2) a dual-affine model is used to calculate the score vector of each word pair; 3) a predicted label is output through a Softmax function; The constraint layer is used to add constraints; the constraints include a first constraint and a second constraint; the first constraint is that entities and undirected relationships are symmetric about a diagonal line; the second constraint includes that if a relationship exists, then the corresponding entity pair must exist; The decoding layer is used for decoding; the decoding includes: span decoding, entity type decoding and entity pair relationship decoding in sequence; predicting the two-dimensional matrix in a unified label space, marking the span as exceeding the threshold, and finally predicting the entity and relationship with the highest score; Locating entities in the multimodal first aid knowledge graph according to the entities in the question sentence, and determining matching entities; the matching entities are entities in the multimodal first aid knowledge graph that match the entities in the question sentence; Calculating the similarity between the relationship in the question and the relationship of all the matching entities using a deep learning model; Determining the answer to the question according to the similarity; According to the target language selected by the user, the answer is input into a machine translation model and the translated answer is output.

2. The first aid knowledge question-answering method based on multimodal knowledge graph according to claim 1 is characterized in that: The acquiring of first aid related knowledge based on the Internet and constructing a multimodal first aid knowledge graph based on the first aid related knowledge specifically include: Acquire first aid knowledge in text form based on the Internet and construct a traditional first aid knowledge graph; Acquire first aid knowledge in the form of voice, image, and video based on the Internet; The first aid knowledge in voice form, the first aid knowledge in image form and the first aid knowledge in video form are integrated into the traditional first aid knowledge graph to obtain a multimodal first aid knowledge graph.

3. The first aid knowledge question-answering method based on multimodal knowledge graph according to claim 2 is characterized in that: The step of integrating the first aid knowledge in the form of voice, the first aid knowledge in the form of images, and the first aid knowledge in the form of videos into the traditional first aid knowledge graph to obtain a multimodal first aid knowledge graph specifically includes: Converting the voice information in the voice-based first aid knowledge into text information; Jointly extracting the text information to obtain the entity and text of the first aid knowledge in voice form; Integrate the entity and text of the first aid knowledge in voice form into the traditional first aid knowledge graph to obtain a first aid knowledge graph containing voice first aid knowledge; Annotating the image information in the image-based first aid knowledge to obtain an image annotation result; the image annotation result includes object attributes in the image and the relationship between the objects; Integrating the image annotation result into the first aid knowledge graph including voice first aid knowledge to obtain a first aid knowledge graph including voice first aid knowledge and image first aid knowledge; Annotating the image information in the video-form first aid knowledge to obtain a video annotation result; the video annotation result includes the name of the first aid skill; The video annotation results are integrated into the first aid knowledge graph including voice first aid knowledge and image first aid knowledge to obtain the multimodal first aid knowledge graph.

4. The first aid knowledge question-answering method based on multimodal knowledge graph according to claim 1, characterized in that: Determining the answer to the question according to the similarity specifically includes: Selecting the relationship with the highest similarity among the similarities; The matching entity corresponding to the relationship with the highest similarity is used as the answer to the question.

5. A first aid knowledge question-answering system based on a multimodal knowledge graph, characterized in that: include: A multimodal first aid knowledge graph construction module is used to obtain first aid related knowledge based on the Internet and construct a multimodal first aid knowledge graph based on the first aid related knowledge; the first aid related knowledge includes emergency diseases, first aid medicines, first aid methods, first aid equipment and methods of using first aid equipment; the first aid related knowledge includes text, voice, pictures and videos; An entity relationship extraction module, used to obtain a question sentence input by a user, and extract entities and relationships in the question sentence using an entity relationship joint extraction model; the input of the entity relationship joint extraction model is a two-dimensional n*n table; the entity relationship joint extraction model includes an encoding layer, a constraint layer and a decoding layer; The encoding layer is used to encode the question input by the user. Specifically: a pre-trained language model is used to obtain a vector representation of the text. To address the long-distance dependency problem of sentences, the context of the sentence is concatenated, and the sentence is extended to a fixed length. At the same time, a deep dual-affine attention mechanism is used to encode the directional information of the words in the table: 1) two dimensionality-reduced multi-layer perceptrons are used to obtain information in different directions of the sentence; 2) a dual-affine model is used to calculate the score vector of each word pair; 3) a predicted label is output through a Softmax function; The constraint layer is used to add constraints; the constraints include a first constraint and a second constraint; the first constraint is that entities and undirected relationships are symmetric about a diagonal line; the second constraint includes that if a relationship exists, then the corresponding entity pair must exist; The decoding layer is used for decoding; the decoding includes: span decoding, entity type decoding and entity pair relationship decoding in sequence; predicting the two-dimensional matrix in a unified label space, marking the span as exceeding the threshold, and finally predicting the entity and relationship with the highest score; A matching module, used to locate entities in the multimodal first aid knowledge graph according to the entities in the question sentence, and determine matching entities; the matching entities are entities in the multimodal first aid knowledge graph that match the entities in the question sentence; A similarity calculation module, used to calculate the similarity between the relationship in the question and the relationship of all the matching entities using a deep learning model; An answer determination module, used to determine the answer to the question according to the similarity; The translation module is used to input the answer into a machine translation model according to the target language selected by the user and output the translated answer.

6. The first aid knowledge question-answering system based on multimodal knowledge graph according to claim 5 is characterized in that: The multimodal first aid knowledge graph construction module specifically includes: A traditional first aid knowledge graph construction unit is used to obtain first aid knowledge in text form based on the Internet and construct a traditional first aid knowledge graph; A multimodal first aid knowledge acquisition unit, used to acquire first aid knowledge in the form of voice, image, and video based on the Internet; The integration unit is used to integrate the first aid knowledge in voice form, the first aid knowledge in image form and the first aid knowledge in video form into the traditional first aid knowledge graph to obtain a multimodal first aid knowledge graph.

7. The first aid knowledge question-answering system based on multimodal knowledge graph according to claim 6 is characterized in that: The integration unit specifically includes: A conversion subunit, used for converting the voice information in the voice-form first aid knowledge into text information; A joint extraction subunit, used for jointly extracting the text information to obtain the entity and text of the first aid knowledge in voice form; A voice integration subunit is used to integrate the entity and text of the first aid knowledge in voice form into the traditional first aid knowledge graph to obtain a first aid knowledge graph containing voice first aid knowledge; An image annotation subunit is used to annotate the image information in the image-based first aid knowledge to obtain an image annotation result; the image annotation result includes the attributes of objects in the image and the relationship between the objects; An image integration subunit is used to integrate the image annotation result into the first aid knowledge graph containing voice first aid knowledge, so as to obtain a first aid knowledge graph containing voice first aid knowledge and image first aid knowledge; The video annotation subunit is used to annotate the image information in the video form of first aid knowledge to obtain a video annotation result; the video annotation result includes the name of the first aid skill; The video integration subunit is used to integrate the video annotation results into the first aid knowledge graph containing voice first aid knowledge and image first aid knowledge to obtain the multimodal first aid knowledge graph.

8. The first aid knowledge question-answering system based on multimodal knowledge graph according to claim 5, characterized in that: The answer determination module specifically includes: A selection unit, used for selecting the relationship with the highest similarity among the similarities; The answer determination unit is used to use the matching entity corresponding to the relationship with the highest similarity as the answer to the question.

Citation Information

Patent Citations

  • Knowledge searching method and system, question and answer device, electronic device and storage medium

    CN109933724A

  • Knowledge graph question-answering method and device

    CN111639171A

  • AR tumor knowledge graph multi-modal demonstration method based on intelligent search

    CN112131405A

  • Linear cultural heritage knowledge graph construction method and system, computing device and medium

    CN112527915A