Method and apparatus for expanding a question and answer knowledge base

By obtaining user behavior logs and web page data, and using information extraction technology to expand the Q&A knowledge base, it solves the problem that the voice interaction system cannot answer emergencies accurately in real time, and achieves more accurate Q&A capabilities.

CN114218364BActive Publication Date: 2025-07-18HISENSE ELECTRONIC TECH (WUHAN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111397544.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-23
Publication Date
2025-07-18
Estimated Expiration
2041-11-23

AI Technical Summary

Technical Problem

Existing voice interaction systems are difficult to accurately store and provide information query answers to emergencies in real time, resulting in the inability to provide accurate answers.

Method used

By obtaining the behavior logs of multiple online users, determining the target topic, and obtaining unstructured data related to the target topic from the web page, using information extraction technology to extract structured data from the unstructured data, expanding the Q&A knowledge base.

Benefits of technology

It realizes the real-time enrichment of the electronic device Q&A knowledge base, and can answer users' questions more accurately, especially information query for emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218364B_ABST
    Figure CN114218364B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method and device for expanding a question-and-answer knowledge base, which can achieve: obtaining the behavior logs of multiple online users, and determining a target topic according to the behavior logs of the multiple online users; obtaining first data related to the target topic in a web page and storing it, where the first data is unstructured data; extracting second data from the first data and saving the second data to a preset question-and-answer knowledge base, where the second data is structured data. The embodiment of the present application can determine the target topic of interest to the current user according to the behavior logs of online users, actively obtain unstructured data related to the target topic from the web page, then extract structured data from the unstructured data, and use the extracted structured data to expand the question-and-answer knowledge base, thereby enriching the question-and-answer knowledge base of the electronic device, and therefore can answer the questions raised by users more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of voice interaction, and in particular, to a method and device for expanding a question and answer knowledge base. Background Art

[0002] At present, due to the development of voice technology, there are more and more intelligent voice interaction devices, and voice interaction has become a very important way of human-computer interaction. Especially in recent years, with the popularization of voice assistants, from mobile terminals to some intelligent home appliances, services can be obtained through voice interaction.

[0003] In existing voice interaction systems, although the question and answer query function is generally supported, most of the question and answer query functions are limited to common sense knowledge questions and answers, such as "Why is the sky blue?", and the answers to such questions are relatively fixed. Therefore, the content stored in the question and answer knowledge base can accurately answer users; for information queries about sudden events, such as "How much loss did the heavy rain in this city cause yesterday?", the answers to such questions are not fixed and it is difficult to store them in the question and answer knowledge base in real time and accurately, resulting in the voice interaction system being unable to provide accurate answers. Summary of the Invention

[0004] The embodiments of the present application provide a method and device for expanding a question and answer knowledge base, which can improve the accuracy of the voice question and answer function of an electronic device.

[0005] In some embodiments, the above method for expanding a question and answer knowledge base includes:

[0006] Obtain the behavior logs of multiple online users, and determine a target topic according to the behavior logs of the multiple online users;

[0007] Obtain and store first data related to the target topic in a web page, where the first data is unstructured data;

[0008] Extract second data from the first data and save the second data to a preset question and answer knowledge base, where the second data is structured data.

[0009] In a feasible implementation manner, the determining a target topic according to the behavior logs of multiple online users includes:

[0010] Preprocess the behavior logs of the multiple online users and determine each sentence included in the preprocessed behavior logs;

[0011] Input each of the sentences into a Transformer-based Bidirectional Encoder Representations from Transformers (BERT) model to obtain sentence vectors corresponding to the respective sentences;

[0012] Cluster the sentence vectors corresponding to the respective sentences, and determine the target topic based on the clustering results.

[0013] In a feasible implementation, extracting second data from the first data includes:

[0014] Extract first structured data from the first data using a preset distributed information extraction model;

[0015] Extract second structured data from the first data using a preset joint information extraction model;

[0016] Obtain the second data based on the first structured data and the second structured data.

[0017] In a feasible implementation, the distributed information extraction model includes a relation classification model and an entity recognition model; the extracting first structured data from the first data using a preset distributed information extraction model includes:

[0018] Using the relation classification model, input each sentence in the first data into the BERT model to obtain a first sentence vector corresponding to each sentence; determine the mapping between the first sentence vectors corresponding to the respective sentences and a preset relation set; based on the mapping between the first sentence vectors corresponding to the respective sentences and the preset relation set, determine the relation category to which each sentence belongs;

[0019] Using the entity recognition model, add the relation category to which each sentence belongs to each sentence, and input the sentences after addition into the BERT model to obtain a second sentence vector corresponding to each sentence; determine the mapping between each element of the second sentence vectors corresponding to the respective sentences and a preset sequence label set; based on the mapping between each element of the second sentence vectors corresponding to the respective sentences and the preset sequence label set, determine the main entity and the guest entity in each sentence;

[0020] Obtain the first structured data based on the main entity and the guest entity in each sentence, and the relation category to which each sentence belongs.

[0021] In a feasible implementation, the extracting second structured data from the first data using a preset joint information extraction model includes:

[0022] Using the combined information extraction model, input each sentence in the first data into the BERT model to obtain the first sentence vector corresponding to each sentence;

[0023] According to the first sentence vectors corresponding to the respective sentences, determine the position information of the main entity in each sentence, where the position information of the main entity includes the start position and the end position of the main entity;

[0024] Add the position information of the main entity in each sentence to the corresponding first sentence vector to obtain the second sentence vector corresponding to each sentence;

[0025] According to the second sentence vectors corresponding to the respective sentences, determine the position information of the guest entity in each sentence, and the relationship category between the main entity and the guest entity in each sentence, where the position information of the guest entity includes the start position and the end position of the guest entity;

[0026] Obtain the second structured data according to the main entity and the guest entity in each sentence, and the relationship category between the main entity and the guest entity in each sentence.

[0027] In a feasible implementation manner, the obtaining the second data according to the first structured data and the second structured data includes:

[0028] Determine the intersection between the first structured data and the second structured data, and determine the intersection as the second data.

[0029] In a feasible implementation manner, the obtaining the second data according to the first structured data and the second structured data includes:

[0030] Based on prior knowledge, verify the first structured data and the second structured data;

[0031] Determine the intersection between the verified first structured data and the second structured data as the second data.

[0032] In a feasible implementation manner, the obtaining and storing the first data related to the target topic in the web page includes:

[0033] Use a preset crawler program to obtain and store the first data related to the target topic in the web page.

[0034] In a feasible implementation manner, the obtaining and storing the first data related to the target topic in the web page

[0035] After obtaining the first data related to the target topic in a web page, store the first data in a preset distributed file system (Hadoop Distributed File System, abbreviated as HDFS) and a search server (ElasticSearch, abbreviated as ES) respectively.

[0036] In some embodiments, the above-mentioned Q&A knowledge base expansion device includes:

[0037] A first acquisition module, configured to acquire the behavior logs of multiple online users, and determine a target topic according to the behavior logs of the multiple online users;

[0038] A second acquisition module, configured to acquire and store the first data related to the target topic in a web page, where the first data is unstructured data;

[0039] A processing module, configured to extract second data from the first data, and save the second data to a preset Q&A knowledge base, where the second data is structured data.

[0040] The Q&A knowledge base expansion method and device provided by the embodiments of the present application can determine the target topic of interest to the current user by acquiring the behavior logs of multiple online users; and after determining the target topic, actively acquire and store the unstructured data related to the target topic from the web page; then based on information extraction technology, extract structured data from the above unstructured data, and use the extracted structured data to expand the Q&A knowledge base, thereby enriching the Q&A knowledge base of the electronic device in real time, and therefore can answer the questions raised by users more accurately. Description of the Drawings

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 It is a schematic flowchart of the electronic device applied in the voice interaction scenario;

[0043] Figure 2 It is another schematic flowchart of the electronic device applied in the voice interaction scenario;

[0044] Figure 3 It is a schematic flowchart of a Q&A knowledge base expansion method provided by an embodiment of the present application;

[0045] Figure 4It is a schematic diagram of a structured data extraction system framework exemplarily shown in the embodiments of the present application;

[0046] Figure 5 It is a schematic diagram of program modules of a question and answer knowledge base expansion device provided in the embodiments of the present application. Detailed implementation manners

[0047] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application. In addition, although the disclosed content in the present application is introduced according to exemplary one or several examples, it should be understood that each aspect of these disclosed contents can also be separately constituted as a complete implementation manner.

[0048] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the subsequent described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.

[0049] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar or same-kind objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances, for example, they can be implemented in an order other than those given in the illustration or description of the embodiments of the present application.

[0050] In addition, the terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device including a series of components does not necessarily have to be limited to those components clearly listed, but may include other components not clearly listed or inherent to these products or devices.

[0051] The term "module" used in the present application refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code that can perform functions related to the element.

[0052] The question and answer knowledge base expansion method provided by the embodiments of the present application can be applied to electronic devices such as mobile terminals, tablet computers, computers, laptop computers, smart TVs, etc., and there is no limitation in the embodiments of the present application.

[0053] In some embodiments, an electronic device includes a sound collector and a controller. The electronic device can collect voice data in real time through its sound collector. Subsequently, the collected voice data is sent to the controller, and the controller identifies the instructions included in the voice data and performs a question-and-answer operation.

[0054] In some embodiments, referring to Figure 1 , Figure 1 is a schematic flow chart of the application of the electronic device in a voice interaction scenario. In S11, the sound collector in the electronic device collects voice data in the surrounding environment where the electronic device is located in real time.

[0055] In S12, after receiving the voice data, the controller identifies the instructions included in the voice data. For example, if the voice data includes the question "Why is the sky blue?" issued by the user, after the controller identifies the instructions included in the voice data, the controller can execute the identified instructions and search for relevant answers in the question-and-answer knowledge base.

[0056] In some embodiments, the electronic device can be connected to a server through the Internet. Then, when the electronic device collects voice data, it can send the voice data to the server through the Internet. The server identifies the instructions included in the voice data and sends the identified instructions back to the electronic device, so that the electronic device can directly execute the received instructions. Compared with the scenario shown in Figure 1 , this scenario reduces the requirement for the computing power of the electronic device and can set a larger recognition model on the server to further improve the accuracy of identifying instructions in voice data and the accuracy of the voice question-and-answer function.

[0057] In some embodiments, referring to Figure 2 , Figure 2 is another schematic flow chart of the application of the electronic device in a voice interaction scenario. In S21, the sound collector in the electronic device collects voice data in the surrounding environment where the electronic device is located in real time and sends the collected voice data to the controller. In S22, the controller further sends the voice data to the server through the communicator. The server identifies the instructions included in the voice data in S23. Subsequently, the server sends the identified instructions back to the electronic device in S24. Correspondingly, the electronic device receives the instructions through the communicator and sends them to the controller. Finally, the controller can directly execute the received instructions in S25.

[0058] In some embodiments, after receiving the voice input by the user, the electronic device generally performs speech recognition on the voice input by the user first, determines the user's intention through the semantic understanding engine, and then provides relevant services for the user according to the user's intention. Although existing electronic devices generally support the Q&A query function, most of the Q&A query functions are limited to common sense knowledge Q&A, such as "Why is the sky blue?" The answers to such questions are relatively fixed, so the user can be accurately answered according to the content stored in the knowledge base; for information queries about sudden events, such as "How much loss did the heavy rain in this city cause yesterday?", the answers to such questions are not fixed and it is difficult to store them in the knowledge base accurately in real time, resulting in the voice interaction system being unable to provide accurate answers.

[0059] In view of the above technical problems, in the embodiments of the present application, a method for expanding the Q&A knowledge base is provided. By obtaining the behavior logs of multiple online users, the target topics of interest to the current user can be determined; and after determining the target topics, unstructured data related to the target topics is actively obtained from the web pages and stored; then based on information extraction technology, structured data is extracted from the above unstructured data, and the extracted structured data is used to expand the Q&A knowledge base, thus enriching the Q&A knowledge base of the electronic device in real time, so that the questions raised by the user can be answered more accurately. The following uses detailed embodiments for detailed description.

[0060] Referring to Figure 3 , Figure 3 is a schematic flow chart of the method for expanding the Q&A knowledge base provided by the embodiments of the present application. The execution subject of this embodiment can be Figure 1 the electronic device in the embodiment shown in Figure 2 , or it can also be Figure 3 the server in the embodiment shown in

[0061]

[0062]

[0063]

[0064] In a feasible implementation manner, the behavior logs of multiple online users can be collected in real time, and the target topics can be determined according to the collected behavior logs of multiple online users.

[0064] In some embodiments, the above target topic can be the topic with the highest question frequency among all the topics asked by multiple online users, or it can also be multiple topics with relatively high question frequencies.

[0064] Exemplarily, if "heavy rain" frequently appears in the collected behavior logs, it can be determined that the user is very concerned about the current heavy rain weather, and it can be determined that the topic of interest to the user is "heavy rain"; if "Olympics" and "medal table" frequently appear in the collected log data, it can be determined that the user is very concerned about the current Olympics and medal rankings, and it can be determined that the topics of interest to the user are "Olympics" and "medal table".

[0065] S302. Obtain and store first data related to the target topic in the web page, where the first data is unstructured data.

[0066] In some embodiments, after determining the target topic of interest to the user, search for and save first data related to the target topic in the web page.

[0067] Among them, the above-mentioned first data is unstructured data.

[0068] In a feasible implementation manner, web crawler technology can be used to search for content related to the above-mentioned target topic in multiple web pages, and the searched content is used as the first data.

[0069] Among them, web crawler technology is a program or script that automatically grabs information on the World Wide Web according to certain rules. For example, focused crawler technology can automatically download web pages, or selectively access web pages and related links on the World Wide Web according to established crawling targets to obtain the required information.

[0070] Different from general web crawlers, focused crawlers do not pursue large coverage, but target at grabbing web pages related to a specific topic content, preparing data resources for topic-oriented user queries, which can greatly save hardware and network resources; in addition, the saved pages are also updated quickly due to the small number, and can also well meet the needs of some specific groups for information in specific fields.

[0071] Among them, the above-mentioned unstructured data can be understood as data with irregular or incomplete data structures, no predefined data models, and inconvenient to be represented by a two-dimensional logic table in a database. It includes all formats of office documents, texts, pictures, Extensible Markup Language (XML), Hyper Text Markup Language (HTML), various reports, images, and audio / video information, etc.

[0072] In some embodiments, after obtaining the above-mentioned first data, the above-mentioned first data can be stored.

[0073] In a feasible implementation manner, the above-mentioned first data can be respectively stored in a preset HDFS and ES; alternatively, the above-mentioned first data can also be stored only in HDFS or ES.

[0074] It can be understood that ES is an open-source distributed search engine based on RESTful web interfaces and built on Apache Lucene. At the same time, ES is also a distributed document database, where each field can be indexed and the data of each field can be searched, and it can be horizontally extended to hundreds of servers for storing and processing PB-level data. It can store, search, and analyze a large amount of data in a very short time, so it can be used as the core search engine in complex search scenarios.

[0075] HDFS has the characteristics of high fault tolerance, and it provides high throughput to access the data of applications, and it is suitable for applications with ultra-large data sets.

[0076] S303. Extract the second data from the first data and save the second data to a preset question-and-answer knowledge base. The second data is structured data.

[0077] In some embodiments, after storing the above-mentioned first data, a complete information extraction system can be built. Through information extraction, the conversion from unstructured data to structured data can be realized, that is, the second data is extracted from the above-mentioned first data, and the second data is structured data.

[0078] Among them, structured data is also called row data, which is data logically expressed and implemented by a two-dimensional table structure, strictly follows the data format and length specifications, and is mainly stored and managed through relational databases.

[0079] After extracting the above-mentioned second data, save the above-mentioned second data to a preset question-and-answer knowledge base. When the electronic device receives a question related to the above-mentioned target topic, the electronic device can search for relevant data from the question-and-answer knowledge base and give an answer.

[0080] The method for expanding the question-and-answer knowledge base provided by the embodiments of the present application can determine the target topic of interest to the current user by obtaining the behavior logs of multiple online users; and after determining the target topic, actively obtain unstructured data related to the target topic from the web page and store it; then based on information extraction technology, extract structured data from the above-mentioned unstructured data, and use the extracted structured data to expand the question-and-answer knowledge base, thereby enriching the question-and-answer knowledge base of the electronic device in real time, so it can answer the questions raised by users more accurately.

[0081] Based on the content described in the above embodiments, in some embodiments of the present application, after obtaining the behavior logs of multiple online users, the obtained behavior logs can be preprocessed first.

[0082] In a feasible implementation manner, the above preprocessing process may include:

[0083] (1) After filtering out invalid characters, sensitive information, relatively long texts, relatively short texts, etc. in the obtained behavior logs, the remaining text information is sorted into multiple sentences.

[0084] (2) Input the above sentences into the BERT model to obtain the sentence vector (representation) of each sentence.

[0085] In a feasible implementation manner, the above BERT model structure mainly includes an Embedding layer and an output layer. Among them, the input of the BERT model is two sentences, separated by the [SEP] symbol from each other, and a special symbol [CLS] is added at the beginning of the sentence to facilitate downstream classification tasks. The input embedding of BERT consists of three parts: token embedding, segment embedding, and position embedding. Among them, token embedding is the vector representation of each word; segment embedding is mainly used to distinguish two sentences and is used for sentence-level MASK tasks, distinguished by "0" and "1"; position embedding represents position encoding, that is, the position of each word is represented in the form of encoding. The position encoding in BERT is different from the trigonometric function position encoding of the original Transformer and is obtained through training in the model. Then, the Embedding layer of BERT is obtained by summing the above three encoding methods.

[0086] In the output layer, the input Embedding is input into an encoder of a multi-layer stacked bidirectional Transformer for feature extraction, and finally each word in the sentence will output a vector with a length of hidden_size.

[0087] (3) Cluster the sentence vectors corresponding to each sentence, and based on the clustering result, determine the above target topic.

[0088] In a feasible implementation manner, the Kmeans algorithm can be used to cluster multiple sentence vectors, and the clustering result is manually analyzed. The clustering behavior logs in the target field are selected for the second clustering,... The above clustering process is repeated until convergence. After the above clustering process converges, the above target topic is determined according to the clustering result.

[0089] The Q&A knowledge base expansion method provided by the embodiments of the present application can determine the target topics of interest to the current user by obtaining the behavior logs of multiple online users and performing BERT encoding and Kmeans clustering. After determining the target topics, it actively obtains and stores data related to the target topics from the web pages, thereby enriching the Q&A knowledge base of the electronic device in real time, and thus can answer the questions raised by users more accurately.

[0090] Based on the content described in the above embodiments, in some embodiments of the present application, after obtaining the first data related to the above target topics, a distributed information extraction model and a joint information extraction model can be used to extract structured data (i.e., the second data) from the first data.

[0091] In a feasible implementation manner, a distributed information extraction model can be used to extract the first structured data from the first data; a preset joint information extraction model can be used to extract the second structured data from the first data; and then, based on the first structured data and the second structured data, the above second data can be obtained.

[0092] For a better understanding of the embodiments of the present application, refer to Figure 4 , Figure 4 which is a schematic diagram of a structured data extraction system framework exemplarily shown in the embodiments of the present application. In some embodiments, the distributed information extraction model includes a relation classification model and an entity recognition model.

[0093] Among them, the relation classification model can implement:

[0094] BERT encoding: Input each sentence in the first data into the BERT model to obtain the first sentence vector corresponding to each sentence.

[0095] Text classification: Determine the mapping between the first sentence vector corresponding to each sentence and the preset relation set; based on the mapping between the first sentence vector corresponding to each sentence and the preset relation set, determine the relation category to which each sentence belongs.

[0096] The entity recognition model can implement:

[0097] BERT encoding: Add the relation category to which each sentence belongs to each sentence, and input the added sentences into the BERT model to obtain the second sentence vector corresponding to each sentence.

[0098] Sequence label annotation: Determine the mapping between each element of the second sentence vector corresponding to each sentence and the preset sequence label set.

[0099] Text classification: Determine the main entity and the guest entity in each of the above sentences according to the mapping between each element of the second sentence vector corresponding to each of the above sentences and the preset sequence label set.

[0100] Among them, according to the main entity and the guest entity in each of the above sentences, and the relationship category to which each of the above sentences belongs, the first structured data can be obtained.

[0101] In a feasible implementation manner, use a relationship classification model to complete the mapping from the first sentence vector after BERT encoding to the preset relationship set, so as to determine the relationship category to which each sentence belongs. Then, append the determined relationship category to the end of each sentence, and perform a second BERT encoding to obtain a new second sentence vector; use an entity recognition model to complete the mapping from each element of the second sentence vector to the sequence label set, so as to determine the main entity and the guest entity in the sentence.

[0102] Among them, text classification is a task in natural language processing (NLP, Natural Language Processing), which can be divided into traditional machine learning methods and deep learning methods. Traditional machine learning methods use SVM classification models, KNN classification models, random forest models (RF), etc., and deep learning methods use fastText models, TextCNN models, TextRNN models, and pre-trained BERT models, etc.

[0103] Text classification based on BERT can be mainly divided into the following three steps:

[0104] 1). Obtain the representation (sentence vector) of the input sentence based on the pre-trained language model BERT;

[0105] 2). After passing the output sentence vector of BERT through a fully connected layer, activate to obtain the classification category;

[0106] 3). Compare the obtained classification with the standard and optimize the difference between the two until the difference is minimized.

[0107] Among them, named entity recognition (NER) is also a task in NLP. An entity can be considered as an instance of a certain concept. For example, "person name" is a concept, or an entity type, then "Zhang San" is a "person name" entity. "Time" is an entity type, then "Mid-Autumn Festival" is a "time" entity. The so-called entity recognition is the process of picking out the entity types that want to be obtained from a sentence.

[0108] NER can be treated as a sequence labeling problem, such as BIO sequence labeling. BIO sequence labeling is to label each element as "B-X", "I-X", or "O". Among them, "B-X" means that the segment where this element is located belongs to type X and this element is at the beginning of this segment, "I-X" means that the segment where this element is located belongs to type X and this element is in the middle or at the end of this segment, and "O" means not belonging to any entity type. For example, for the sentence "Xiaoming watched a game of the Chinese men's basketball team in Yanyuan of Nanjing University", the labeling result is: [B-PER, I-PER, O, B-ORG, I-ORG, I-ORG, I-ORG, O, B-LOC, I-LOC, O, O, B-ORG, I-ORG, I-ORG, I-ORG, O, O, O, O, O], where PER represents the person type, ORG represents the organization type, and LOC represents the location type.

[0109] In some embodiments, the combined information extraction model can achieve:

[0110] BERT encoding: Input each sentence in the above first data into the BERT model to obtain the first sentence vector corresponding to each sentence.

[0111] Determine the main entity: According to the first sentence vectors corresponding to the above sentences, determine the position information (including the start position and the end position) of the main entity in each sentence.

[0112] Extract the main entity, the guest entity, and the relationship category: Add the position information of the main entity in each of the above sentences to the first sentence vector corresponding to each sentence to obtain the second sentence vector corresponding to each sentence; According to the second sentence vectors corresponding to the above sentences, determine the position information (including the start position and the end position information) of the guest entity in each sentence, and the relationship category between the main entity and the guest entity in each sentence; According to the main entity and the guest entity in each of the above sentences, and the relationship category between the main entity and the guest entity in each sentence, obtain the second structured data.

[0113] Among them, the combined extraction network (Cascade-Binary-Tagging-Framework) is a network structure that uses BERT for underlying encoding and uses the Seq2seq probability graph idea to model the predicate relationship (predicate) between the main entity (subject) and the guest entity (object) as a mapping function from the main entity to the guest entity, so as to achieve the purpose of extracting entities and the corresponding relationships between entities from unstructured text. This network is mainly divided into the following steps:

[0114] 1). The original sequence is input into the BERT encoder to obtain an encoded sequence;

[0115] 2). Input the encoded sequence into the fully-connected network of a binary classifier to predict the subject;

[0116] 3). According to the predicted subject, extract the encoding vectors corresponding to the head and tail of the subject from the encoded sequence;

[0117] 4). Using the encoding vector of the subject as a condition, perform a Layer Normalization operation on the encoded sequence to normalize all neuron nodes of a single sample for each layer;

[0118] 5). Input the sequence (sentence vector) after Layer Normalization into another fully-connected network for binary classification to predict the object and predicate corresponding to the subject.

[0119] To better understand the embodiments of the present application, taking the input sentence: "Wu Jing, the director of Wolf Warrior 2, was born in 1974" as an example, after passing through the joint extraction network, the obtained entity relationship triples are: [(Wolf Warrior 2, director, Wu Jing), (Wu Jing, date of birth, 1974)].

[0120] In some embodiments, after obtaining the above first structured data and second structured data, the intersection between the first structured data and the second structured data can be determined, and this intersection is determined as the above second data.

[0121] In other embodiments, after obtaining the above first structured data and second structured data, it is also possible to verify the first structured data and the second structured data based on prior knowledge; the intersection between the verified first structured data and the second structured data is determined as the above second data, thereby improving the accuracy of the second data.

[0122] In some embodiments of the present application, the structured data obtained by the above two models can also be compared, and methods such as threshold screening are used to further improve the accuracy of the second data.

[0123] The method for expanding the Q&A knowledge base provided by the embodiments of the present application, after determining the target topic, can obtain unstructured data related to the target topic from the web page, then extract structured data from the above unstructured data based on a preset information extraction technique, and use the extracted structured data to expand the Q&A knowledge base, thereby enriching the Q&A knowledge base of the electronic device in real time, and thus being able to answer the questions raised by users more accurately.

[0124] In some embodiments, the present application also provides a device for expanding the Q&A knowledge base. Refer toFigure 5 , Figure 5 This is a schematic diagram of the program module of the Q&A knowledge base expansion device provided by the embodiment of the present application.

[0125] As Figure 5 shown, the device includes:

[0126] The first acquisition module 501 acquires the behavior logs of multiple online users, and determines the target topic according to the behavior logs of the multiple online users.

[0127] The second acquisition module 502 acquires and stores the first data related to the target topic in the web page, and the first data is unstructured data.

[0128] The processing module 503 extracts the second data from the first data and saves the second data to a preset Q&A knowledge base, and the second data is structured data.

[0129] The knowledge base expansion device provided by the embodiment of the present application can determine the target topic of interest to the current user by acquiring the behavior logs of multiple online users; and after determining the target topic, actively acquire and store the unstructured data related to the target topic from the web page; then based on the information extraction technology, extract the structured data from the above unstructured data, and use the extracted structured data to expand the Q&A knowledge base, thereby enriching the Q&A knowledge base of the electronic device in real time, so that it can answer the questions raised by the user more accurately.

[0130] In a feasible implementation manner, the first acquisition module 501 is used for:

[0131] Preprocess the behavior logs of the multiple online users, and determine each sentence contained in the preprocessed behavior logs; input each sentence into the BERT model to obtain the sentence vector corresponding to each sentence; cluster the sentence vectors corresponding to each sentence, and based on the clustering result, determine the above target topic.

[0132] In a feasible implementation manner, the processing module 503 is used for:

[0133] Extract the first structured data from the first data by using a preset distributed information extraction model; extract the second structured data from the first data by using a preset joint information extraction model; obtain the second data according to the first structured data and the second structured data.

[0134] In a feasible implementation manner, the above distributed information extraction model includes a relation classification model and an entity recognition model; specifically, the processing module 503 is used for:

[0135] Using the above relationship classification model, input each sentence in the above first data into the BERT model to obtain the first sentence vector corresponding to each sentence; determine the mapping between the first sentence vector corresponding to each sentence and the preset relationship set; according to the mapping between the first sentence vector corresponding to each sentence and the preset relationship set, determine the relationship category to which each sentence belongs.

[0136] Using the above entity recognition model, add the relationship category to which each sentence belongs to each sentence, and input each sentence after the addition into the BERT model to obtain the second sentence vector corresponding to each sentence; determine the mapping between each element of the second sentence vector corresponding to each sentence and the preset sequence label set; according to the mapping between each element of the second sentence vector corresponding to each sentence and the preset sequence label set, determine the main entity and the guest entity in each sentence.

[0137] According to the main entity and the guest entity in each sentence, and the relationship category to which each sentence belongs, obtain the first structured data.

[0138] In a feasible implementation manner, the processing module 503 is specifically configured to:

[0139] Using the joint information extraction model, input each sentence in the first data into the BERT model to obtain the first sentence vector corresponding to each sentence;

[0140] According to the first sentence vector corresponding to each sentence, determine the position information of the main entity in each sentence, including the start position and the end position;

[0141] Add the position information of the main entity in each sentence to the first sentence vector corresponding to each sentence to obtain the second sentence vector corresponding to each sentence;

[0142] According to the second sentence vector corresponding to each sentence, determine the position information of the guest entity in each sentence (including the start position and the end position), and the relationship category between the main entity and the guest entity in each sentence; according to the main entity and the guest entity in each sentence, and the relationship category between the main entity and the guest entity in each sentence, obtain the second structured data.

[0143] In a feasible implementation manner, the processing module 503 is specifically configured to:

[0144] Determine the intersection between the first structured data and the second structured data, and determine the intersection as the second data.

[0145] In a feasible implementation manner, the above determining the intersection between the first structured data and the second structured data, and determining the intersection as the second data includes:

[0146] Based on prior knowledge, verify the first structured data and the second structured data; determine the intersection between the verified first structured data and the second structured data as the second data.

[0147] In a feasible implementation manner, the second acquisition module 503 is specifically configured to:

[0148] Use a preset crawler program to acquire and store the first data related to the target topic in the web page.

[0149] In a feasible implementation manner, the second acquisition module 503 is specifically configured to:

[0150] After acquiring the first data related to the target topic in the web page, store the first data into a preset HDFS and ES respectively.

[0151] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of devices or modules may be in an electrical, mechanical or other form.

[0152] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0153] In addition, in each embodiment of the present application, the various functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The above-mentioned unit integrated with modules can be implemented in the form of hardware, or in the form of a hardware plus software functional unit.

[0154] The above-mentioned integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above-mentioned software functional modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in each embodiment of the present application.

[0155] It should be understood that the above-mentioned processor may be a Central Processing Unit (CPU), or it may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0156] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0157] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0158] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0159] An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a main control device.

[0160] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for expanding a question-and-answer knowledge base, characterized in that, Including: Obtain the behavior logs of multiple online users, and determine a target topic according to the behavior logs of the multiple online users; Obtain and store first data related to the target topic in a web page, where the first data is unstructured data; Extract second data from the first data and save the second data to a preset question and answer knowledge base, where the second data is structured data; The determining the target topic according to the behavior logs of the multiple online users includes: Preprocess the behavior logs of the multiple online users and determine each sentence contained in the preprocessed behavior logs; Input each sentence into a BERT model with bidirectional encoding based on Transformer to obtain a sentence vector corresponding to each sentence; Cluster the sentence vectors corresponding to each sentence, and determine the target topic based on the clustering result.

2. The method according to claim 1, wherein Extracting the second data from the first data includes: Extract first structured data from the first data by using a preset distributed information extraction model; Extract second structured data from the first data by using a preset joint information extraction model; Obtain the second data according to the first structured data and the second structured data.

3. The method according to claim 2, wherein The distributed information extraction model includes a relation classification model and an entity recognition model; the extracting the first structured data from the first data by using the preset distributed information extraction model includes: Using the relation classification model, input each sentence in the first data into the BERT model to obtain a first sentence vector corresponding to each sentence; determine the mapping between the first sentence vector corresponding to each sentence and a preset relation set; according to the mapping between the first sentence vector corresponding to each sentence and the preset relation set, determine the relation category to which each sentence belongs; Using the entity recognition model, add the relation category to which each sentence belongs to each sentence, and input the added sentences into the BERT model to obtain a second sentence vector corresponding to each sentence; determine the mapping between each element of the second sentence vector corresponding to each sentence and a preset sequence label set; according to the mapping between each element of the second sentence vector corresponding to each sentence and the preset sequence label set, determine the main entity and the guest entity in each sentence; Obtain the first structured data according to the main entity and the guest entity in each sentence, and the relation category to which each sentence belongs.

4. The method according to claim 2, wherein The extracting the second structured data from the first data by using the preset joint information extraction model includes: Using the joint information extraction model, input each sentence in the first data into the BERT model to obtain a first sentence vector corresponding to each sentence; According to the first sentence vector corresponding to each sentence, determine the position information of the main entity in each sentence, where the position information of the main entity includes the start position and the end position of the main entity; Add the position information of the main entity in each sentence to the corresponding first sentence vector to obtain a second sentence vector corresponding to each sentence; Based on the second sentence vectors corresponding to the respective sentences, determine the position information of the object entities in the respective sentences, and the relationship categories between the subject entities and the object entities in the respective sentences, where the position information of the object entities includes the start position and the end position of the object entities; Based on the subject entities and the object entities in the respective sentences, and the relationship categories between the subject entities and the object entities in the respective sentences, obtain the second structured data.

5. The method according to any one of claims 2 to 4, characterized in that The obtaining of the second data based on the first structured data and the second structured data includes: Determine the intersection between the first structured data and the second structured data, and determine the intersection as the second data.

6. The method according to claim 5, characterized in that, The obtaining of the second data based on the first structured data and the second structured data includes: Based on prior knowledge, verify the first structured data and the second structured data; Determine the intersection between the verified first structured data and the second structured data as the second data.

7. The method according to claim 1, characterized in that The obtaining and storing of the first data related to the target topic in the web page includes: Use a preset crawler program to obtain and store the first data related to the target topic in the web page.

8. The method according to claim 1 or 7, characterized in that, The obtaining and storing of the first data related to the target topic in the web page includes: After obtaining the first data related to the target topic in the web page, store the first data in a preset distributed file system HDFS and a search server ES respectively.

9. A question and answer knowledge base expansion device, characterized in that, The apparatus includes: A first acquisition module, configured to acquire the behavior logs of multiple online users, and determine a target topic according to the behavior logs of the multiple online users; A second acquisition module, configured to obtain and store the first data related to the target topic in the web page, where the first data is unstructured data; A processing module, configured to extract second data from the first data and save the second data to a preset question and answer knowledge base, where the second data is structured data; The first acquisition module is specifically configured to: Preprocess the behavior logs of the multiple online users, and determine the respective sentences included in the preprocessed behavior logs; Input the respective sentences into a BERT model with bidirectional encoding based on Transformer to obtain the sentence vectors corresponding to the respective sentences; Cluster the sentence vectors corresponding to the respective sentences, and determine the target topic based on the clustering result.

Citation Information

Patent Citations

  • Knowledge graph construction method and system for enclosed switchgear

    CN112883197A

  • Relation triple extraction method, device and equipment and medium

    CN112989788A

  • Customer service robot model training method and device, electronic equipment and medium

    CN113342946A