Knowledge question and answer method and device

By querying the knowledge associated with the target problem in the knowledge base, generating the target prompt words and inputting a large language model, the problems that are not answered accurately in the existing question-and-answer system are solved, and more professional and accurate answers are achieved.

CN119988528APending Publication Date: 2025-05-13BEIJING JINGDONG TUOXIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311498168.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing question-and-answer system, users' questions are brief or grammatical, and their generalization ability of knowledge graphs or large language models is insufficient, resulting in inaccurate answers.

Method used

By querying target knowledge associated with target questions in a pre-built knowledge base, update the prompt word expression to generate target prompt words and input them into the large language model to generate answers.

Benefits of technology

It improves the professionalism and accuracy of answers, improves the user experience, and improves the efficiency and accuracy of knowledge acquisition through different types of knowledge bases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988528A_ABST
    Figure CN119988528A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge question and answer method and device, and relates to the technical field of artificial intelligence. A specific embodiment of the knowledge question and answer method comprises the following steps: in response to a received target question, querying target knowledge associated with the target question in a pre-constructed knowledge base; according to the target question and the target knowledge, updating a preset cue word expression to obtain a corresponding target cue word; and inputting the target cue word into a preset large language model, and generating an answer of the target question. According to the embodiment, the target cue word associated with the target question is generated based on the pre-constructed knowledge base, and the answer of the target question is generated by using the large language model according to the target cue word, so that the professionality and effectiveness of the answer can be improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for answering knowledge questions. Background Art

[0002] The question-answering system can output corresponding answers based on the user's questions to answer the user's questions. The question-answering system usually answers the questions raised by the user based on the knowledge graph or the large language model. For example, the question-answering system maps the question to the entity in the knowledge graph, and infers the relationship between the entities in the knowledge graph to get the answer to the question. For another example, the question-answering system inputs the question into the large language model, and the large language model generates the answer to the question.

[0003] In the process of implementing the present invention, the inventors found that the prior art has at least the following problems:

[0004] Users' question expressions are often brief or contain grammatical errors, and the generalization capabilities of knowledge graphs or large language models are insufficient, resulting in inaccurate answers output by the knowledge graphs or large language models. Summary of the invention

[0005] In view of this, an embodiment of the present invention provides a method and device for knowledge question answering, which can improve the professionalism and accuracy of answers and enhance the user experience.

[0006] To achieve the above object, according to a first aspect of an embodiment of the present invention, a knowledge question answering method is provided, comprising:

[0007] In response to receiving a target question, querying a pre-built knowledge base for target knowledge associated with the target question;

[0008] According to the target question and the target knowledge, a preset prompt word expression is updated to obtain a corresponding target prompt word;

[0009] The target prompt word is input into a preset large language model to generate an answer to the target question.

[0010] Optionally, the knowledge base includes: a structured knowledge base; querying the target knowledge associated with the target problem in the pre-built knowledge base includes:

[0011] Identify key fields from the target problem;

[0012] According to the key fields, the structured data stored in the structured knowledge base is retrieved to obtain retrieval data, and the retrieval data is used as the target knowledge corresponding to the target problem.

[0013] Optionally, before retrieving the structured data stored in the structured knowledge base, the method further includes:

[0014] In response to receiving the target text, identifying target structured data from the target text;

[0015] The target structured data is cleaned, and the cleaned target structured data is stored in the structured knowledge base.

[0016] Optionally, the knowledge base includes: a question-answer knowledge base; querying the pre-built knowledge base for target knowledge associated with the target question includes:

[0017] Vectorizing the target problem to obtain a target problem vector;

[0018] Querying the question-answering knowledge base for a question vector that meets a preset first semantic similarity condition with the target question vector;

[0019] An answer vector associated with the question vector is determined, answer data corresponding to the answer vector is obtained, and the answer data is used as target knowledge corresponding to the target question.

[0020] Optionally, before searching the question-answering knowledge base for a question vector that meets a preset first semantic similarity condition with the target question vector, the method further includes:

[0021] In response to receiving the target text, identifying question data and answer data from the target text, and establishing an association relationship between the question data and the answer data;

[0022] The question data and the answer data are vectorized to obtain a question vector and an answer vector having an associated relationship, and the question vector and the answer vector having an associated relationship are stored in the question-answer knowledge base.

[0023] Optionally, the knowledge base includes: a plain text knowledge base; querying the target knowledge associated with the target problem in the pre-built knowledge base includes:

[0024] Vectorizing the target problem to obtain a target problem vector;

[0025] Searching the plain text knowledge base for a plain text vector that meets a preset second semantic similarity condition with the target question vector;

[0026] Obtain plain text data corresponding to the plain text vector, and use the plain text data as target knowledge corresponding to the target problem.

[0027] Optionally, before searching the plain text knowledge base for a plain text vector that meets a preset second semantic similarity condition with the target question vector, the method further includes:

[0028] In response to receiving the target text, segmenting the target text to obtain a plurality of segmented plain texts;

[0029] The plurality of segmented plain texts are vectorized to obtain a plurality of plain text vectors, and the plurality of plain text vectors are stored in the plain text knowledge base.

[0030] Optionally, the method further comprises:

[0031] Performing accuracy evaluation on the answer to the target question to obtain an evaluation result of the answer to the target question;

[0032] According to the evaluation result, the degree of association between the target problem and the target knowledge is updated, and the degree of association is used to query the target knowledge associated with the target problem again in a pre-built knowledge base.

[0033] According to a second aspect of an embodiment of the present invention, a knowledge question answering device is provided, comprising:

[0034] A query module, configured to query a pre-built knowledge base for target knowledge associated with the target question in response to receiving the target question;

[0035] A prompt module, used for updating a preset prompt word expression according to the target question and the target knowledge to obtain a corresponding target prompt word;

[0036] The answer module is used to input the target prompt word into a preset large language model to generate an answer to the target question.

[0037] Optionally, the knowledge base includes: a structured knowledge base; querying the target knowledge associated with the target problem in the pre-built knowledge base includes:

[0038] Identify key fields from the target problem;

[0039] According to the key fields, the structured data stored in the structured knowledge base is retrieved to obtain retrieval data, and the retrieval data is used as the target knowledge corresponding to the target problem.

[0040] Optionally, the device further comprises:

[0041] a first recognition module, configured to recognize target structured data from the target text in response to receiving the target text;

[0042] The first storage module is used to clean the target structured data and store the cleaned target structured data in the structured knowledge base.

[0043] Optionally, the knowledge base includes: a question-answer knowledge base; querying the pre-built knowledge base for target knowledge associated with the target question includes:

[0044] Vectorizing the target problem to obtain a target problem vector;

[0045] Querying the question-answering knowledge base for a question vector that meets a preset first semantic similarity condition with the target question vector;

[0046] An answer vector associated with the question vector is determined, answer data corresponding to the answer vector is obtained, and the answer data is used as target knowledge corresponding to the target question.

[0047] Optionally, the device further comprises:

[0048] A second recognition module is used for, in response to receiving the target text, recognizing the question data and the answer data from the target text, and establishing an association relationship between the question data and the answer data;

[0049] The second storage module is used to vectorize the question data and the answer data to obtain a question vector and an answer vector with an associated relationship, and store the question vector and the answer vector with an associated relationship in the question and answer knowledge base.

[0050] Optionally, the knowledge base includes: a plain text knowledge base; querying the target knowledge associated with the target problem in the pre-built knowledge base includes:

[0051] Vectorizing the target problem to obtain a target problem vector;

[0052] Searching the plain text knowledge base for a plain text vector that meets a preset second semantic similarity condition with the target question vector;

[0053] Obtain plain text data corresponding to the plain text vector, and use the plain text data as target knowledge corresponding to the target problem.

[0054] Optionally, the device further comprises:

[0055] A segmentation module, configured to segment the target text in response to receiving the target text, to obtain a plurality of segmented plain texts;

[0056] The third storage module is used to vectorize the multiple segmented plain texts to obtain multiple plain text vectors, and store the multiple plain text vectors in the plain text knowledge base.

[0057] Optionally, the device further comprises:

[0058] An evaluation module, used to evaluate the accuracy of the answer to the target question and obtain an evaluation result of the answer to the target question;

[0059] An updating module is used to update the correlation between the target problem and the target knowledge according to the evaluation result, and the correlation is used to query the target knowledge associated with the target problem again in a pre-built knowledge base.

[0060] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, including:

[0061] one or more processors;

[0062] a storage device for storing one or more programs,

[0063] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of the above embodiments.

[0064] According to a fourth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method described in any one of the above embodiments is implemented.

[0065] One embodiment of the above invention has the following advantages or beneficial effects: generating target prompt words associated with the target question based on a pre-built knowledge base, and using a large language model to generate an answer to the target question based on the target prompt words, which can improve the professionalism and effectiveness of the answer and improve the user experience; building a structured knowledge base based on structured data, obtaining structured data associated with the target question from the structured knowledge base as target knowledge, and improving the efficiency of obtaining target knowledge; building a question-and-answer knowledge base based on question-and-answer data, and obtaining question-and-answer data associated with the target question from the question-and-answer knowledge base, which can retrieve historical questions similar to the target question and their answers, and improve the efficiency of obtaining target knowledge; building a structured knowledge base based on pure text data, and obtaining structured data associated with the target question from the structured knowledge base as target knowledge, and improving the efficiency of obtaining target knowledge. This data constructs a plain text knowledge base, and obtains plain text content related to the target question from the plain text knowledge base, which can improve the efficiency of acquiring target knowledge; storing different types of knowledge in different knowledge bases can improve the scalability of knowledge acquisition and improve the efficiency and accuracy of acquiring different types of knowledge; generating target prompt words based on pre-set prompt word expressions can more flexibly obtain target prompt words to meet the various ways and styles of asking questions to the large language model; evaluating the answers to the target questions, and updating the correlation between the target questions and the target knowledge based on the evaluation results, so that when facing similar or identical questions, more relevant knowledge can be obtained from the knowledge base more efficiently.

[0066] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific implementation examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.

[0068] Figure 1 is a schematic diagram of the main process of the knowledge question answering method according to an embodiment of the present invention;

[0069] Figure 2 is a schematic diagram of the overall process of knowledge question answering according to a reference embodiment of the present invention;

[0070] Figure 3 is a schematic diagram of a main process of updating a structured knowledge base according to a reference embodiment of the present invention;

[0071] Figure 4 is a schematic diagram of the main process of updating the question-answer knowledge base according to a reference embodiment of the present invention;

[0072] Figure 5 is a schematic diagram of the main process of updating a plain text knowledge base according to a reference embodiment of the present invention;

[0073] Figure 6 is a schematic diagram of the main process of a knowledge question answering method according to a reference embodiment of the present invention;

[0074] Figure 7 is a schematic diagram of main modules of a knowledge question answering device according to an embodiment of the present invention;

[0075] Figure 8 is an exemplary system architecture diagram to which embodiments of the present invention may be applied;

[0076] Fig. 9 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION

[0077] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0078] It should be noted that in the technical solution of the present invention, the collection, use, storage, sharing and transfer of user personal information involved are in compliance with the provisions of relevant laws and regulations, and it is necessary to inform the user and obtain the user's consent or authorization. When applicable, the user's personal information is de-identified and / or anonymized and / or encrypted.

[0079] The question-answering system can output corresponding answers based on the user's questions to answer the user's questions. The question-answering system usually answers the questions raised by the user based on the knowledge graph or the large language model. For example, the question-answering system maps the question to the entity in the knowledge graph, and infers the relationship between the entities in the knowledge graph to get the answer to the question. For another example, the question-answering system inputs the question into the large language model, and the large language model generates the answer to the question.

[0080] Users' question expressions are often brief or contain grammatical errors, and the generalization capabilities of knowledge graphs or large language models are insufficient, resulting in inaccurate answers output by the knowledge graphs or large language models.

[0081] In view of this, according to a first aspect of an embodiment of the present invention, a knowledge question answering method is provided.

[0082] Figure 1 FIG. 1 is a schematic diagram of the main process of the knowledge question answering method according to an embodiment of the present invention. Figure 1 As shown, the knowledge question answering method according to the embodiment of the present invention mainly includes the following steps S101 to S103.

[0083] Step S101 , in response to receiving a target question, querying a pre-built knowledge base for target knowledge associated with the target question.

[0084] The execution subject of the embodiment of the present invention receives the target question sent by the user, and the target question includes questions in various fields, such as medical questions, financial questions, meteorological questions, legal questions, etc. The target question can be one or more. For example, the execution subject of the embodiment of the present invention identifies multiple "question marks" from the target question, regards the content before each question mark as a question, and splits the target question into multiple questions.

[0085] A knowledge base is pre-built, and the knowledge base is used to store knowledge in various fields. Specifically, for knowledge in different fields, corresponding knowledge bases are built, including: medical knowledge base, financial knowledge base, meteorological knowledge base, legal knowledge base, etc.; for knowledge of different data types, corresponding knowledge bases are built, including: text knowledge base, image knowledge base, audio knowledge base, video knowledge base, etc. The execution subject of the embodiment of the present invention establishes an association relationship between multiple knowledge stored in the knowledge base. Each knowledge does not exist independently, but is interrelated. For example, the execution subject of the embodiment of the present invention establishes a strong association relationship or a weak association relationship between different knowledge based on statistical data such as the number of times and frequency of different knowledge appearing at the same time. The more times and the higher the frequency of different knowledge appearing at the same time, the stronger the association relationship between these knowledge. The knowledge stored in the knowledge base is dynamically updated. The execution subject of the embodiment of the present invention stores new knowledge in the knowledge base, modifies, deletes, hides the knowledge stored in the knowledge base, or adds, deletes, modifies the association relationship between knowledge, etc. according to the received knowledge modification request. The knowledge base for storing knowledge applicable to the embodiment of the present invention includes: a relational database, a non-relational database, a vector database, an ElasticSearch database, etc.

[0086] After receiving the target question, the execution subject of the embodiment of the present invention searches for target knowledge related to the target question in a pre-built knowledge base. For example, the execution subject of the embodiment of the present invention uses the target question as a retrieval condition, searches for data records containing the target question in a relational database, and obtains text knowledge corresponding to the target question from the queried data records. For another example, the execution subject of the embodiment of the present invention classifies the target question, or labels the target question, and searches for image knowledge (i.e., images including text knowledge), audio knowledge, or video knowledge that meets the classification results or label information of the target question in a non-relational database.

[0087] Exemplarily, the target question received by the execution subject of the embodiment of the present invention is "Can I take cold medicine when I have a fever?", and the data record whose "question" field includes the above target question is searched in a pre-set relational database, and the "answer" field is obtained in the found data record, and the data included in the "answer" field is used as the target knowledge associated with the target question; the labels of the above target question include: medical treatment, medication, fever, cold medicine, precautions, taboos, etc., and the image data, audio data or video data with all the above labels are searched in the non-relational database, and the found data is used as the target knowledge associated with the target question.

[0088] Searching for target knowledge associated with the target question in multiple knowledge bases can obtain as much and as accurate reference knowledge as possible, providing a large amount of knowledge background for the subsequent large language model, making it easier to obtain more accurate and professional answers.

[0089] According to a reference embodiment of the present invention, the knowledge base includes: a structured knowledge base, which is used to store structured data and semi-structured data. The structured data and semi-structured data are data with data structure and follow data format, length, type and other specifications. Specifically, the structured data includes: product information (for example, product name, product price, production date, manufacturer, product specifications or product ingredients, etc.), logistics information (for example, starting point, destination, transportation time, transportation method, etc.), and the semi-structured data includes: XML (Extensible Markup Language) data, JSON (JavaScript Object Notation) data, HTML (HyperTextMarkup Language) data, etc.

[0090] Each structured data and semi-structured data has a certain structure and can be represented as tabular data or triple data. Each data includes multiple fields such as entity, attribute, relationship, etc., where an entity field corresponds to one or more attribute fields, and two entity fields are associated through relationship fields. For example, in a structured knowledge base in the medical field, a data includes multiple entity fields: cold, respiratory disease, cold medicine, acetaminophen, etc., where the relationship field between the entity field "cold" and the entity field "respiratory disease" is "belongs to", that is, "cold belongs to respiratory disease", the relationship field between the entity field "cold" and the entity field "cold medicine" is "take", the attribute field between the entity field "cold medicine" and the entity field "acetaminophen" is "includes", that is, "cold medicine includes acetaminophen", and the attribute fields corresponding to the entity field "cold" include "chills, runny nose", etc.

[0091] When searching for target knowledge associated with a target problem in a pre-built knowledge base, first identify key fields from the target problem. Key fields may be entity fields, attribute fields, or relationship fields. For example, the execution subject of an embodiment of the present invention performs entity recognition on the target problem and identifies entity fields from the target problem. Based on the identified key fields, the structured data stored in the structured knowledge base is retrieved. Specifically, the identified key fields are used as screening conditions to screen out data in the entity fields that include or are equal to the key fields, and the screened data is used as retrieval data, and the retrieval data is used as the target knowledge corresponding to the target problem.

[0092] Exemplarily, the target question received by the execution subject of the embodiment of the present invention is "Can I take metformin for a cold?", and entity recognition is performed on the above question to obtain the disease entity "cold" and the drug entity "metformin". The structured knowledge base is queried based on the above disease entity and drug entity, and the target knowledge obtained includes: cold is a respiratory disease, and the main symptoms are sneezing, runny nose, and coughing. Metformin is a drug for treating type 2 diabetes, and its indications include diabetes, hypoglycemia, etc.

[0093] Acquiring structured data associated with the target problem from a structured knowledge base as target knowledge can improve the efficiency and accuracy of acquiring structured data and provide an efficient and convenient query method for structured data in the knowledge base.

[0094] According to another reference embodiment of the present invention, before searching the structured data stored in the structured knowledge base, the method further includes: in response to receiving the target text, identifying the target structured data from the target text. For example, identifying entities, attributes included in the entities, relationships between different entities, etc. from the target text, using the identified entities as entity fields, the identified attributes as attribute fields, and the identified relationships as relationship fields, and forming the target structured data with the entity fields, attribute fields, and relationship fields. Data cleaning is performed on the target structured data, for example, determining whether the value in the target structured data meets the preset value range, and if the value does not meet the value range, deleting the field corresponding to the value, or deleting the structured data including the field; for another example, determining whether the field length meets the preset length range, and if the field length does not meet the length range, deleting the field whose field length is not within the length range, or deleting the structured data including the field; for another example, determining whether the structured data includes null values, and if the structured data includes null values, replacing the null values ​​with default values, or deleting the fields containing null values.

[0095] The target structured data after data cleaning is stored in a structured knowledge base. The structured knowledge base applicable to the embodiment of the present invention includes: an ElasticSearch database. ElasticSearch provides an efficient and convenient text-based retrieval function. The execution subject of the embodiment of the present invention retrieves the structured data and / or semi-structured data stored in the ElasticSearch database according to the identified fields. Preferably, an index is established for one or more fields in the ElasticSearch database to improve data query efficiency.

[0096] It should be noted that the structured knowledge base is dynamically updated. The execution subject of the embodiment of the present invention stores new structured data in the structured knowledge base, and modifies or deletes outdated or erroneous structured data according to the knowledge base update request.

[0097] Exemplarily, the execution subject of the embodiment of the present invention identifies the target text "Ibuprofen is a non-steroidal anti-inflammatory drug. This product inhibits cyclooxygenase, reduces the synthesis of prostaglandins, produces analgesic and anti-inflammatory effects; and has an antipyretic effect through the hypothalamic temperature regulation center." and obtains multiple fields including: "Ibuprofen", "Ibuprofen", "non-steroidal anti-inflammatory drug", "analgesia", "anti-inflammatory", "inhibit cyclooxygenase", and "reduce the synthesis of prostaglandins"; the execution subject of the embodiment of the present invention combines the above fields into structured data, performs data cleaning on the structured data, determines and identifies the presence of typos, whether English meets the uppercase and lowercase specifications, whether Chinese meets the traditional and simplified specifications, etc., and stores the structured data after data cleaning in the ElasticSearch database.

[0098] Building a structured knowledge base based on structured data can classify knowledge of different data types, improve the efficiency of acquiring target knowledge, improve the storage efficiency of structured data, and improve the scalability of the knowledge base.

[0099] According to another reference embodiment of the present invention, the knowledge base includes: a question-and-answer knowledge base, which is used to store question-and-answer data. For example, the question-and-answer data includes: online questions asked by users, and online answers given by real doctors or pharmacists to users' questions. The question-and-answer data includes both actual questions raised by users, covering various common questions asked by users, and high-quality answers from real doctors or pharmacists. The questions and answers in the question-and-answer data have an associated relationship, each question can be associated with one or more answers, and each answer can be used to answer one or more questions. It should be noted that the question-and-answer knowledge base includes a question vector corresponding to the question and an answer vector corresponding to the answer, that is, the question-and-answer knowledge base includes: real question-and-answer text and answer text, and also includes: a vectorized representation of the question-and-answer text and a vectorized representation of the answer text.

[0100] The question and answer knowledge base applicable to the embodiment of the present invention includes: a vector knowledge base (for example, the open source knowledge base Chroma, Pinecone, etc.) and a relational database, wherein the vector knowledge base is used to store question and answer vectors corresponding to question and answer texts and answer vectors corresponding to answer texts, and the relational database is used to store question and answer texts and answer texts.

[0101] When searching for target knowledge associated with a target question in a pre-built knowledge base, the target question is first vectorized to obtain a target question vector. For example, vector encoding (embedding operation) is performed on the target question, and the vectorization result of the target question is used as the target question vector. Then, the question vector that meets the first pre-set semantic similarity condition with the target question vector is searched in the question and answer knowledge base. For example, the first pre-set semantic similarity condition is that the similarity between the target question vector and the question vector stored in the question and answer knowledge base is greater than or equal to the first pre-set similarity threshold. That is, the execution subject of an embodiment of the present invention calculates the similarity between the target question vector and each question vector stored in the question and answer knowledge base, and determines the question vector that meets the first semantic similarity condition according to the first similarity threshold. Then determine the answer vector associated with the question vector, obtain the answer data corresponding to the answer vector, and use the answer data as the target knowledge corresponding to the target question.

[0102] Exemplarily, the question-answering knowledge base includes question vector A1, question vector A2 and question vector A3. The executing body of the embodiment of the present invention respectively calculates the similarities between the target question vector and the above three question vectors, and obtains that the similarity corresponding to question vector A1 is 0.6, the similarity corresponding to question vector A2 is 0.9, and the similarity corresponding to question vector A3 is 0.5; since the pre-set first similarity threshold is 0.8, question vector A2 is used as a vector associated with and similar to the target question vector, and the answer vector B2 associated with question vector A2 is obtained, and then the answer text corresponding to the answer vector B2 and the question text corresponding to the question vector A2 are obtained, and the obtained question text and answer text are used as the target knowledge corresponding to the target question.

[0103] Preferably, when there are multiple question vectors whose similarities are greater than or equal to a first similarity threshold, the multiple question vectors are sorted in descending order of similarity, and according to a pre-set quantity threshold, several question vectors with the greatest similarity are taken as vectors associated with and similar to the target question vector, thereby obtaining several question texts and answer texts.

[0104] By obtaining question and answer data associated with the target question from the question and answer knowledge base as the target knowledge, similar questions and corresponding answers that users have asked can be quickly retrieved, thereby improving the efficiency and accuracy of obtaining question and answer data and providing an efficient and convenient query method for question and answer data in the knowledge base.

[0105] According to another reference embodiment of the present invention, before searching the question data that meets the preset first semantic similarity condition with the target question vector in the question and answer knowledge base, the method also includes: in response to receiving the target text, identifying the question data and the answer data from the target text, and establishing an association relationship between the question data and the answer data. For example, the execution subject of the embodiment of the present invention receives the online question and answer data of the user and the doctor, the online question and answer data includes the questions raised by the user and the corresponding answers of the doctor, the online question data is the data with question tags and answer tags, and the specific data format is "[question]C1[answer]D1[question]C2[answer]D2", wherein "[question]C1" indicates that the data C1 is the question data raised by the user, and "[answer]D1" indicates that the data D1 is the answer data replied by the doctor. Through the labels in the data, the execution subject of the embodiment of the present invention identifies the target text as question data and answer data.

[0106] Then, the question data and the answer data are vectorized to obtain a question vector corresponding to the question data and an answer vector corresponding to the answer data, and an association relationship is established between the question vector and the answer vector. For example, the question vector has an association relationship with the answer vector obtained before it and the answer vector obtained after it. The question vector and the answer vector with an association relationship are stored in the question and answer knowledge base, that is, the question vector and the answer vector are stored in the vector database. It should be noted that the question and answer knowledge base is dynamically updated. The execution subject of the embodiment of the present invention stores new question and answer data and their vectors in the question and answer knowledge base according to the knowledge base update request, and modifies or deletes outdated or erroneous question and answer data and their vectors.

[0107] Exemplarily, the target text received by the execution subject of the embodiment of the present invention is "[Q]E1[A]F1[A]F2[Q]E2[Q]E3[A]F3", wherein "[Q]E1" indicates that text E1 is a question raised by a user, and "[A]F1" indicates that text F1 is the doctor's answer, and so on for other texts; thereby obtaining question text E1, answer text F1 and answer text F3 associated with question text E1, question text E2 and question text E3, and answer text F3 associated with question text E2 and question text E3; vector encoding is performed on the above-identified question text and answer text to obtain corresponding question vectors and answer vectors, and an association relationship between the question vector and the answer vector is established according to the order between the question text and the answer text, and the question vector and answer vector with the association relationship are stored in a question-answering knowledge base, i.e., a vector database.

[0108] Building a question-and-answer knowledge base based on question-and-answer data can classify knowledge of different data types, improve the efficiency of acquiring target knowledge, improve the storage efficiency of question-and-answer data, and improve the scalability of the knowledge base.

[0109] According to another reference embodiment of the present invention, the knowledge base includes: a plain text knowledge base, which is used to store plain text data, such as disease encyclopedia knowledge, pharmacy encyclopedia knowledge, medical popular science articles, etc. Plain text data has no obvious structure, but the plain text data contains a large amount of knowledge content. It should be noted that the plain text knowledge base includes plain text vectors corresponding to the plain text data, that is, the plain text knowledge base includes: plain text data and vectorized representation of the plain text data.

[0110] When searching for target knowledge associated with a target problem in a pre-built knowledge base, first vectorize the target problem to obtain a target problem vector. Search the plain text knowledge base for plain text vectors that meet a preset second semantic similarity condition with the target problem vector. For example, the preset second semantic similarity condition is that the similarity between the target problem vector and the plain text vector stored in the plain text knowledge base is greater than or equal to a preset second similarity threshold. That is, the execution subject of an embodiment of the present invention calculates the similarity between the target problem vector and each plain text vector stored in the plain text knowledge base, and determines the plain text vector that meets the second semantic similarity condition according to the second similarity threshold. Then, the plain text data corresponding to the screened plain text vector is obtained, and the plain text data is used as the target knowledge corresponding to the target problem.

[0111] Exemplarily, the plain text knowledge base includes a plain text vector G1, a plain text vector G2 and a plain text vector G3. The executing body of the embodiment of the present invention respectively calculates the similarities between the target problem vector and the above three plain text vectors, and obtains that the similarity corresponding to the plain text vector G1 is 0.4, the similarity corresponding to the plain text vector G2 is 0.8, and the similarity corresponding to the plain text vector G3 is 0.7; since the pre-set second similarity threshold is 0.75, the plain text vector G2 is used as a vector associated with and similar to the target problem vector, and then the plain text data corresponding to the plain text vector G2 is obtained, and the acquired plain text data is used as the target knowledge corresponding to the target problem.

[0112] Preferably, when there are multiple plain text vectors whose similarities are greater than or equal to a second similarity threshold, the multiple plain text vectors are sorted in descending order of similarity, and according to a pre-set quantity threshold, several plain text vectors with the greatest similarity are taken as vectors associated with and similar to the target question vector, thereby obtaining several plain text data.

[0113] Obtaining plain text data associated with the target question from the plain text knowledge base as the target knowledge can quickly retrieve popular science information related to the user's question, improve the efficiency and accuracy of obtaining plain text data, provide an efficient and convenient query method for the plain text data in the knowledge base, expand the depth and breadth of the target knowledge, and facilitate the professionalism and accuracy of the answers.

[0114] According to a reference embodiment of the present invention, before searching the plain text knowledge base for a plain text vector that meets a preset second semantic similarity condition with the target question vector, the method further includes: in response to receiving the target text, segmenting the target text to obtain a plurality of segmented plain texts. For example, in the case where the target text includes natural paragraphs, the text is segmented by natural paragraphs, and each natural paragraph is used as a segmented plain text; for another example, according to a preset data length, the target text is segmented to obtain a plurality of segmented plain texts with equal data lengths; for another example, in the case where the target text includes a title, the target text is segmented by titles such as main title, subtitle, subtitle, parent title, etc. to obtain a plurality of segmented plain texts. Then, the plurality of segmented plain texts are vectorized to obtain a plurality of plain text vectors, and the plurality of plain text vectors are stored in the plain text knowledge base, i.e., the vector database.

[0115] It should be noted that the plain text knowledge base is dynamically updated. The execution subject of the embodiment of the present invention stores new plain text data and its vectors in the plain text knowledge base, and modifies or deletes outdated or erroneous plain text data and its vectors according to the knowledge base update request.

[0116] Exemplarily, the target text received by the execution subject of the embodiment of the present invention is "H1[SEP]H2[SEP]H3", wherein the target text includes: natural paragraph H1, natural paragraph H2 and natural paragraph H3, and "[SEP]" represents the separator between natural paragraphs; the above-identified natural paragraphs (i.e., segmented plain texts) are vector encoded to obtain corresponding plain text vectors, and an association relationship between multiple plain text vectors is established according to the front-to-back order between different segmented plain texts, and multiple plain text vectors with an associated relationship are stored in a plain text knowledge base, i.e., a vector database.

[0117] Constructing a plain text knowledge base based on plain text data can classify knowledge of different data types, improve the efficiency of acquiring target knowledge, improve the storage efficiency of plain text data, and improve the scalability of the knowledge base.

[0118] Step S102: updating a preset prompt word expression according to the target question and the target knowledge to obtain a corresponding target prompt word.

[0119] After obtaining the target question and target knowledge, obtain a preset prompt word expression, which is the standard style of the target prompt word. The prompt word expression is used to specify the position of the target question and target knowledge in the target prompt word, bring the target question and target knowledge into the corresponding position of the prompt word expression, and use the target question and target knowledge to update the prompt word expression, that is, use the target question and target knowledge to fill in the prompt word expression, and use the updated prompt word expression as the target prompt word.

[0120] Exemplarily, the pre-set prompt word expression is "Known drug instruction manual knowledge: [target knowledge 1]. Known question: The answer to [target knowledge 2-1] is [target knowledge 2-2]. Known [target knowledge 3]. Please refer to the above content to answer the question: [target question]." Among them, "[target knowledge 1]" is used to fill in the target knowledge obtained from the structured knowledge base, "[target knowledge 2-1]" is used to fill in similar questions obtained from the question and answer knowledge base, "[target knowledge 2-2]" is used to fill in the answers to similar questions obtained from the question and answer knowledge base, "[target knowledge 3]" is used to fill in the target knowledge obtained from the plain text knowledge base, and "[target question]" is used to fill in the target question sent by the user; the prompt word expression after the content is filled in is used as the target prompt word.

[0121] It should be noted that the prompt word expression is configurable. The execution subject of the embodiment of the present invention modifies the prompt word expression according to the received expression modification request, for example, adding a new sentence pattern, modifying or deleting an existing sentence pattern, etc.

[0122] The target prompt word is generated according to the preset prompt word expression, which can obtain the target prompt word more flexibly and meet the various ways and styles of asking questions to the large language model.

[0123] Step S103: input the target prompt word into a preset large language model to generate an answer to the target question.

[0124] After obtaining the target prompt word, the target prompt word is input into a preset large language model, the output of the large language model is received, the output of the large language model is used as the answer to the target question, and the answer to the target question is returned to the questioner of the target question. A large language model (LLM) refers to a deep learning model trained with a large amount of text data that can generate natural language text or understand language text. For example, the GPT-3.5 large language model, the GPT-4 large language model, the Claude large language model, etc., but not limited to this.

[0125] According to a reference embodiment of the present invention, the method further includes: evaluating the accuracy of the answer to the target question to obtain an evaluation result of the answer to the target question. Exemplarily, the execution subject of the embodiment of the present invention transmits the target question and the corresponding answer to the scorer, or sends a prompt message to the scorer, so that the scorer scores the answer to the target question, and determines the evaluation result by manual scoring. Another exemplary method is to calculate the perplexity (PPL) of the answer. The perplexity is an indicator used to evaluate whether the output content of the language model is appropriate. In essence, it is to calculate the probability of the output content. The greater the probability of the content output by the language model, the smaller the perplexity, which means that the language model is better and the output content is more appropriate. The corresponding evaluation result is determined according to the perplexity of the answer.

[0126] After obtaining the evaluation result, the correlation between the target question and the target knowledge is updated according to the evaluation result, wherein the correlation is used to query the target knowledge associated with the target question again in the pre-built knowledge base. Exemplarily, the correlation between the target question and all the knowledge in the knowledge base is 0, and then the correlation between the target question and the screened target knowledge is updated according to the evaluation result. The better the evaluation result of the answer, the greater the correlation of the corresponding target knowledge, indicating that the target question and the target knowledge are more related; when the target question is received again, the first target knowledge greater than or equal to the correlation threshold is first queried according to the correlation, and then the knowledge updated after the last target question is received is queried to obtain the second target knowledge, and the first target knowledge and the second target knowledge are used as the target knowledge of the target question received again, and are used to fill in the prompt word expression and generate the target prompt word.

[0127] The answers to the target questions are evaluated, and the correlation between the target questions and the target knowledge is updated based on the evaluation results. This enables more efficient acquisition of knowledge that is more relevant to the question from the knowledge base when facing similar or identical questions.

[0128] Figure 2 is a schematic diagram of the overall process of knowledge question answering according to a reference embodiment of the present invention. Figure 2As shown, the execution subject of the embodiment of the present invention receives the target question sent by the user, then performs entity recognition on the target question, obtains key fields such as entity fields, attribute fields, and relationship fields in the target question, and performs vector encoding on the target question to obtain a target question vector; then, according to the key fields, the relevant target knowledge is queried in the ElasticSearch database (i.e., the structured knowledge base) to obtain target knowledge 1, the question and answer vector whose similarity with the target question vector is greater than or equal to the first similarity threshold is queried in the vector database 1 (i.e., the question and answer knowledge base), the corresponding question and answer text is obtained, and the obtained question and answer text is used as the target knowledge 2, the plain text vector whose similarity with the target question vector is greater than or equal to the second similarity threshold is queried in the vector database 2 (i.e., the plain text knowledge base), the corresponding plain text data is obtained, and the obtained plain text data is used as the target knowledge 3; in the prompt word construction module, the target prompt word expression is selected, and the prompt word construction is performed according to the target knowledge 1, the target knowledge 2 and the target knowledge 3 to obtain the target prompt word; the target prompt word is input into the large language model, the answer to the target question is output using the large language model, and the answer is returned to the user who asked the target question.

[0129] Figure 3 FIG. 1 is a schematic diagram of the main process of updating a structured knowledge base according to a reference embodiment of the present invention. Figure 3 As shown, the execution subject of the embodiment of the present invention receives the target text, performs entity recognition on the target text to obtain entity fields, performs attribute recognition on the target text to obtain attribute fields, performs relationship recognition on the target text to obtain relationship fields, and stores the entity fields, attribute fields and relationship fields in a structured knowledge base or a semi-structured knowledge base.

[0130] Figure 4 FIG. 1 is a schematic diagram of the main process of updating the question-answer knowledge base according to a reference embodiment of the present invention. Figure 4 As shown, the execution subject of the embodiment of the present invention obtains an online conversation between a doctor and a patient, identifies question data and answer data from the online conversation, vectorizes the question data to obtain a question vector, vectorizes the answer data to obtain an answer vector, establishes an association relationship between the question vector and the answer vector, and stores the question vector and answer vector having an association relationship in a vector database.

[0131] Figure 5 FIG. 1 is a schematic diagram of the main process of updating a plain text knowledge base according to a reference embodiment of the present invention. Figure 5As shown, the execution subject of the embodiment of the present invention receives the target text, segments the target text to obtain multiple segmented plain texts, vectorizes the multiple segmented plain texts to obtain multiple plain text vectors, establishes an association relationship between the multiple plain text vectors, and stores the plain text vectors with the association relationship in a vector database.

[0132] Figure 6 FIG. 1 is a schematic diagram of the main process of a method for answering questions according to a reference embodiment of the present invention. Figure 6 As shown, the knowledge question answering method may include:

[0133] Step S601, in response to receiving a target question, querying a pre-built structured knowledge base for structured knowledge associated with the target question;

[0134] Step S602, querying the question-answering knowledge associated with the target question in the pre-built question-answering knowledge base;

[0135] Step S603, searching a pre-built plain text knowledge base for plain text knowledge associated with the target question;

[0136] Step S604, updating the preset prompt word expression according to the target question and structured knowledge, question-answer knowledge, and plain text knowledge to obtain the target prompt word;

[0137] Step S605: input the target prompt word into a preset large language model to generate an answer to the target question.

[0138] The specific implementation content of the knowledge question answering method of the above-mentioned reference embodiment of the present invention has been described in detail in the knowledge question answering method described above, so the repeated content will not be described here.

[0139] According to a second aspect of an embodiment of the present invention, a knowledge question and answering device is provided.

[0140] Figure 7 is a schematic diagram of the main modules of the knowledge question answering device according to an embodiment of the present invention, such as Figure 7 As shown, the knowledge question answering device 700 mainly includes:

[0141] A query module 701 is used for querying target knowledge associated with the target question in a pre-built knowledge base in response to receiving the target question;

[0142] A prompt module 702 is used to update a preset prompt word expression according to the target question and the target knowledge to obtain a corresponding target prompt word;

[0143] The answer module 703 is used to input the target prompt word into a preset large language model to generate an answer to the target question.

[0144] According to a reference embodiment of the present invention, the knowledge base includes: a structured knowledge base; querying the target knowledge associated with the target problem in the pre-built knowledge base includes:

[0145] Identify key fields from the target problem;

[0146] According to the key fields, the structured data stored in the structured knowledge base is retrieved to obtain retrieval data, and the retrieval data is used as the target knowledge corresponding to the target problem.

[0147] According to another reference embodiment of the present invention, the knowledge question answering device 700 further includes:

[0148] a first recognition module, configured to recognize target structured data from the target text in response to receiving the target text;

[0149] The first storage module is used to clean the target structured data and store the cleaned target structured data in the structured knowledge base.

[0150] According to another reference embodiment of the present invention, the knowledge base includes: a question-answer knowledge base; querying the target knowledge associated with the target question in the pre-built knowledge base includes:

[0151] Vectorizing the target problem to obtain a target problem vector;

[0152] Querying the question-answering knowledge base for a question vector that meets a preset first semantic similarity condition with the target question vector;

[0153] An answer vector associated with the question vector is determined, answer data corresponding to the answer vector is obtained, and the answer data is used as target knowledge corresponding to the target question.

[0154] According to another reference embodiment of the present invention, the knowledge question answering device 700 further includes:

[0155] A second recognition module is used for, in response to receiving the target text, recognizing the question data and the answer data from the target text, and establishing an association relationship between the question data and the answer data;

[0156] The second storage module is used to vectorize the question data and the answer data to obtain a question vector and an answer vector with an associated relationship, and store the question vector and the answer vector with an associated relationship in the question and answer knowledge base.

[0157] According to another reference embodiment of the present invention, the knowledge base includes: a plain text knowledge base; querying the target knowledge associated with the target problem in the pre-built knowledge base includes:

[0158] Vectorizing the target problem to obtain a target problem vector;

[0159] Searching the plain text knowledge base for a plain text vector that meets a preset second semantic similarity condition with the target question vector;

[0160] Obtain plain text data corresponding to the plain text vector, and use the plain text data as target knowledge corresponding to the target problem.

[0161] According to a reference embodiment of the present invention, the knowledge question answering device 700 further includes:

[0162] A segmentation module, configured to segment the target text in response to receiving the target text, to obtain a plurality of segmented plain texts;

[0163] The third storage module is used to vectorize the multiple segmented plain texts to obtain multiple plain text vectors, and store the multiple plain text vectors in the plain text knowledge base.

[0164] According to another reference embodiment of the present invention, the knowledge question answering device 700 further includes:

[0165] An evaluation module, used to evaluate the accuracy of the answer to the target question and obtain an evaluation result of the answer to the target question;

[0166] An updating module is used to update the correlation between the target problem and the target knowledge according to the evaluation result, and the correlation is used to query the target knowledge associated with the target problem again in a pre-built knowledge base.

[0167] It should be noted that the specific implementation content of the knowledge question and answer device described in the embodiment of the present invention has been described in detail in the knowledge question and answer method described above, so the repeated content will not be described here.

[0168] According to the technical solution of the embodiment of the present invention, a target prompt word associated with a target question is generated based on a pre-built knowledge base, and an answer to the target question is generated based on the target prompt word using a large language model, which can improve the professionalism and effectiveness of the answer and improve the user experience; a structured knowledge base is built based on structured data, and structured data associated with the target question is obtained from the structured knowledge base as target knowledge, thereby improving the efficiency of obtaining target knowledge; a question-and-answer knowledge base is built based on question-and-answer data, and question-and-answer data associated with the target question is obtained from the question-and-answer knowledge base, which can retrieve historical questions similar to the target question and their answers, thereby improving the efficiency of obtaining target knowledge; a knowledge base is built based on plain text data Plain text knowledge base, obtaining plain text content related to the target question from the plain text knowledge base can improve the efficiency of acquiring target knowledge; storing different types of knowledge in different knowledge bases can improve the scalability of knowledge acquisition and improve the efficiency and accuracy of acquiring different types of knowledge; generating target prompt words based on pre-set prompt word expressions can more flexibly obtain target prompt words to meet the various ways and styles of asking questions to the large language model; evaluating the answers to the target questions, and updating the correlation between the target questions and the target knowledge based on the evaluation results, so that when facing similar or identical questions, more relevant knowledge can be obtained from the knowledge base more efficiently.

[0169] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the first aspect of the embodiment of the present invention.

[0170] According to a fourth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method provided by the first aspect of the embodiment of the present invention is implemented.

[0171] Figure 8 An exemplary system architecture 800 is shown to which the method for knowledge question answering or the apparatus for knowledge question answering according to an embodiment of the present invention may be applied.

[0172] like Figure 8 As shown, system architecture 800 may include terminal devices 801, 802, 803, a network 804, and a server 805. Network 804 is used to provide a medium for communication links between terminal devices 801, 802, 803 and server 805. Network 804 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.

[0173] The user can use the terminal devices 801, 802, 803 to interact with the server 805 through the network 804 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 801, 802, 803, such as knowledge question and answer applications, data query applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0174] The terminal devices 801 , 802 , and 803 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0175] The server 805 may be a server that provides various services, such as a background management server (only an example) that provides support for knowledge question and answer requests sent by the upstream terminal devices 801, 802, and 803. In response to receiving the target question, the background management server may query the target knowledge associated with the target question in the pre-built knowledge base; update the pre-set prompt word expression according to the target question and the target knowledge to obtain the corresponding target prompt word; input the target prompt word into the pre-set large language model to generate the answer to the target question; and feed back the knowledge question and answer situation (only an example) to the terminal device.

[0176] It should be noted that the knowledge question answering method provided in the embodiment of the present invention is generally executed by the server 805, and accordingly, the knowledge question answering device is generally set in the server 805. The knowledge question answering method provided in the embodiment of the present invention can also be executed by the terminal devices 801, 802, and 803, and accordingly, the knowledge question answering device can be set in the terminal devices 801, 802, and 803.

[0177] It should be understood that Figure 8 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0178] Reference below Fig. 9 , which shows a schematic diagram of the structure of a computer system 900 of a terminal device suitable for implementing an embodiment of the present invention. Fig. 9 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0179] like Fig. 9As shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the system 900 are also stored. The CPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0180] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that a computer program read therefrom is installed into the storage section 908 as needed.

[0181] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above-mentioned functions defined in the system of the embodiment of the present invention are executed.

[0182] It should be noted that the computer-readable medium shown in the embodiment of the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In an embodiment of the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In an embodiment of the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0183] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0184] The modules involved in the embodiments of the present invention may be implemented by software or hardware. The modules described may also be set in a processor, for example, it may be described as: a processor includes a query module, a prompt module, and an answer module, wherein the names of these modules do not constitute a limitation on the modules themselves under certain circumstances, for example, the query module may also be described as a "module for querying target knowledge associated with a target question".

[0185] As another aspect, an embodiment of the present invention further provides a computer-readable medium, which may be included in the device described in the above embodiment; or may exist independently without being assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device implements the following method: in response to receiving a target question, querying a pre-built knowledge base for target knowledge associated with the target question; updating a pre-set prompt word expression according to the target question and the target knowledge to obtain a corresponding target prompt word; inputting the target prompt word into a pre-set large language model to generate an answer to the target question.

[0186] According to the technical solution of the embodiment of the present invention, a target prompt word associated with a target question is generated based on a pre-built knowledge base, and an answer to the target question is generated based on the target prompt word using a large language model, which can improve the professionalism and effectiveness of the answer and improve the user experience; a structured knowledge base is built based on structured data, and structured data associated with the target question is obtained from the structured knowledge base as target knowledge, thereby improving the efficiency of obtaining target knowledge; a question-and-answer knowledge base is built based on question-and-answer data, and question-and-answer data associated with the target question is obtained from the question-and-answer knowledge base, which can retrieve historical questions similar to the target question and their answers, thereby improving the efficiency of obtaining target knowledge; a knowledge base is built based on plain text data Plain text knowledge base, obtaining plain text content related to the target question from the plain text knowledge base can improve the efficiency of acquiring target knowledge; storing different types of knowledge in different knowledge bases can improve the scalability of knowledge acquisition and improve the efficiency and accuracy of acquiring different types of knowledge; generating target prompt words based on pre-set prompt word expressions can more flexibly obtain target prompt words to meet the various ways and styles of asking questions to the large language model; evaluating the answers to the target questions, and updating the correlation between the target questions and the target knowledge based on the evaluation results, so that when facing similar or identical questions, more relevant knowledge can be obtained from the knowledge base more efficiently.

[0187] The above specific implementations do not constitute a limitation on the protection scope of the embodiments of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the embodiments of the present invention should be included in the protection scope of the embodiments of the present invention.

Claims

1. A knowledge question answering method, characterized in that: include: In response to receiving a target question, querying a pre-built knowledge base for target knowledge associated with the target question; According to the target question and the target knowledge, a preset prompt word expression is updated to obtain a corresponding target prompt word; The target prompt word is input into a preset large language model to generate an answer to the target question.

2. The method according to claim 1, characterized in that The knowledge base includes: a structured knowledge base; querying the target knowledge associated with the target problem in the pre-built knowledge base includes: Identify key fields from the target problem; According to the key fields, the structured data stored in the structured knowledge base is retrieved to obtain retrieval data, and the retrieval data is used as the target knowledge corresponding to the target problem.

3. The method according to claim 2, characterized in that Before searching the structured data stored in the structured knowledge base, the method further includes: In response to receiving the target text, identifying target structured data from the target text; The target structured data is cleaned, and the cleaned target structured data is stored in the structured knowledge base.

4. The method according to claim 1, characterized in that The knowledge base includes: a question-answering knowledge base; querying the target knowledge associated with the target question in the pre-built knowledge base includes: Vectorizing the target problem to obtain a target problem vector; Querying the question-answering knowledge base for a question vector that meets a preset first semantic similarity condition with the target question vector; An answer vector associated with the question vector is determined, answer data corresponding to the answer vector is obtained, and the answer data is used as target knowledge corresponding to the target question.

5. The method according to claim 4, characterized in that Before searching the question-answering knowledge base for a question vector that meets a preset first semantic similarity condition with the target question vector, the method further includes: In response to receiving the target text, identifying question data and answer data from the target text, and establishing an association relationship between the question data and the answer data; The question data and the answer data are vectorized to obtain a question vector and an answer vector having an associated relationship, and the question vector and the answer vector having an associated relationship are stored in the question-answer knowledge base.

6. The method according to claim 1, characterized in that The knowledge base includes: a plain text knowledge base; querying the target knowledge associated with the target problem in the pre-built knowledge base includes: Vectorizing the target problem to obtain a target problem vector; Searching the plain text knowledge base for a plain text vector that meets a preset second semantic similarity condition with the target question vector; Obtain plain text data corresponding to the plain text vector, and use the plain text data as target knowledge corresponding to the target problem.

7. The method according to claim 6, characterized in that Before searching the plain text knowledge base for a plain text vector that meets a preset second semantic similarity condition with the target question vector, the method further includes: In response to receiving the target text, segmenting the target text to obtain a plurality of segmented plain texts; The plurality of segmented plain texts are vectorized to obtain a plurality of plain text vectors, and the plurality of plain text vectors are stored in the plain text knowledge base.

8. The method according to claim 1, characterized in that The method further comprises: Performing accuracy evaluation on the answer to the target question to obtain an evaluation result of the answer to the target question; According to the evaluation result, the association degree between the target problem and the target knowledge is updated, and the association degree is used to query the target knowledge associated with the target problem again in the pre-built knowledge base.

9. A knowledge question-answering device, characterized in that: include: A query module, configured to query a pre-built knowledge base for target knowledge associated with the target question in response to receiving the target question; A prompt module, used for updating a preset prompt word expression according to the target question and the target knowledge to obtain a corresponding target prompt word; The answer module is used to input the target prompt word into a preset large language model to generate an answer to the target question.

10. An electronic device, characterized in that: include: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.

11. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.