Intelligent question and answer method and device, equipment and storage medium

By identifying the user's Q&A intention and using historical Q&A records, combined with the user's career information, the intelligent Q&A system can provide more accurate and personalized answers, solving the problem of inaccurate answers in existing smart Q&A technologies.

CN120144718APending Publication Date: 2025-06-13创优数字科技(广东)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510303507.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing smart question-and-answer technologies usually only answer users’ questions without considering the user’s characteristics, intentions or tendencies, resulting in inaccurate answers and affecting the user experience.

Method used

By obtaining the user's historical Q&A record, identifying the user's Q&A intent, and determining the target knowledge base based on the intention and question data, thereby obtaining the optimal answer data. At the same time, filter the most suitable answer data based on the user's career information.

Benefits of technology

It improves the accuracy and personalization of the answers, making the answers obtained by users more in line with their actual needs, and greatly improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144718A_ABST
    Figure CN120144718A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent question and answer method and device, equipment and a storage medium, and the method comprises the steps: obtaining question data and historical question and answer records of a user in response to a question and answer request instruction initiated by the user; identifying a question-answering intention of the user based on the question data and the historical question-answering record; determining a corresponding target knowledge base according to the question and answer intention and the question data; determining each piece of answer data corresponding to the question data based on the target knowledge base; and acquiring occupational information of the user, determining optimal answer data from the answer data according to the occupational information, and feeding back the optimal answer data to the user. The question and answer intention of the user is identified based on the question data and the historical question and answer records, the purpose of the user is known, so that the answer meeting the user can be given, the corresponding target knowledge base is determined according to the question and answer intention and the question data, and meanwhile, the occupation of the user is considered; and the answer which is most matched and most suitable for the user is selected from the answer data, so that the user experience is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent question answering, and specifically relates to an intelligent question answering method, device, equipment and storage medium. Background Art

[0002] In this era of information overload, people's demand for efficiently obtaining accurate information is becoming increasingly urgent. The intelligent question answering technology has emerged. The technology of intelligent question answering aims to understand the natural language questions of users and directly give accurate answers, which greatly improves the efficiency of information retrieval and acquisition and has been widely applied in multiple fields.

[0003] The existing intelligent question answering technology usually only answers the questions raised by users in a single way. For example, when multiple different users ask the same question, the answers they get are exactly the same. It does not distinguish users and does not consider factors such as the characteristics, intentions or inclinations of the users themselves. Therefore, although the answers obtained by users are standard, they are not what the users really need, and from the perspective of individual users, the accuracy is not high, which affects the user experience. Summary of the Invention

[0004] In view of this, this application provides an intelligent question answering method, device, equipment and storage medium, which is used to solve the problem that the existing intelligent question answering technology usually only answers the questions raised by users in a single way, but it is not what the users really need, and from the perspective of individual users, the accuracy is not high, which affects the user experience.

[0005] To achieve the above objectives, the following solutions are proposed:

[0006] In a first aspect, an intelligent question answering method includes:

[0007] Respond to the request instruction for question answering initiated by the user, and obtain the question data and the user's historical question answering records;

[0008] Identify the user's question answering intention based on the question data and historical question answering records;

[0009] Determine the corresponding target knowledge base according to the question answering intention and question data;

[0010] Determine each answer data corresponding to the question data based on the target knowledge base;

[0011] Obtain the user's occupation information, determine the optimal answer data from each answer data according to the occupation information, and feedback the optimal answer data to the user.

[0012] Preferably, the identifying the user's question answering intention based on the question data and historical question answering records includes:

[0013] Classify the historical Q&A records according to each preset Q&A type to obtain groups of Q&A data under each Q&A type;

[0014] Extract the data summary of each group of the Q&A data;

[0015] Complete the problem data to obtain complete problem information;

[0016] Identify the Q&A intention of the user based on the complete problem information and the data summary of each group of the Q&A data.

[0017] Preferably, the determining the corresponding target knowledge base according to the Q&A intention and problem data includes:

[0018] Vectorize the problem data to obtain a problem vector;

[0019] Perform a similarity search on the problem vector in a pre-established vector database to screen out the corresponding various tools;

[0020] Determine the description information of each tool;

[0021] Combine the Q&A intention and the problem data, and respectively match the combination with the description information of each tool to calculate the matching degree;

[0022] Screen the target tool according to the matching degree, and use the database corresponding to the target tool as the target knowledge base.

[0023] Preferably, the determining the respective answer data corresponding to the problem data based on the target knowledge base includes:

[0024] Determine the problem vector corresponding to the problem data;

[0025] Obtain a pre-trained fine-ranking model corresponding to the target knowledge base;

[0026] Use the fine-ranking model to process the problem vector corresponding to the problem data to obtain respective answer vectors; the fine-ranking model is trained with a vector sample set as the training sample and the answer vector sample corresponding to each vector sample in the vector sample set as the sample label;

[0027] Convert each answer vector into the corresponding answer data respectively.

[0028] Preferably, the construction process of the vector sample set includes:

[0029] Convert each piece of knowledge data in the target knowledge base into each piece of knowledge vector;

[0030] For each of the said knowledge vectors, calculate the correlation between this knowledge vector and each of the other knowledge vectors respectively to determine the degree of correlation;

[0031] Take each of the other knowledge vectors with a degree of correlation not less than a preset correlation threshold as the corresponding knowledge vectors of this knowledge vector;

[0032] Combine this knowledge vector with each of its corresponding knowledge vectors respectively to obtain respective vector samples;

[0033] Summarize the historical Q&A records of all users to obtain a Q&A set;

[0034] Extract all historical question data and corresponding historical answer data from the Q&A set, convert each piece of the historical question data into a corresponding historical question vector respectively, and at the same time convert each piece of the historical answer data into a corresponding historical answer vector;

[0035] Combine each of the historical question vectors with the corresponding historical answer vector to obtain respective vector samples;

[0036] Summarize all the said vector samples to obtain the vector sample set.

[0037] Preferably, the step of determining the optimal answer data from each of the answer data according to the occupational information and feeding back the optimal answer data to the user includes:

[0038] Extract the occupational characteristics of multiple preset occupations;

[0039] Cluster the occupational characteristics of each of the preset occupations to form respective different types of feature clusters;

[0040] Calculate the feature degree of each of the feature clusters;

[0041] Determine the feature cluster to which the occupational information belongs from each of the feature clusters as the target cluster;

[0042] Compare the feature degree of the target cluster with a preset feature threshold to determine the optimal answer data.

[0043] Preferably, the step of comparing the feature degree of the target cluster with a preset feature threshold to determine the optimal answer data includes:

[0044] Calculate the similarity between each of the answer data and the question data respectively;

[0045] If the feature degree of the target cluster is less than the feature threshold, take the answer data corresponding to the highest similarity as the optimal answer data;

[0046] If the feature degree of the target cluster is not less than the feature threshold, one or more answer data corresponding to a similarity greater than a preset similarity threshold are combined to obtain optimal answer data.

[0047] In a second aspect, an intelligent question and answer device includes:

[0048] A response module, configured to obtain question data and the user's historical question and answer records in response to a request instruction for question and answer initiated by the user;

[0049] A question and answer intention recognition module, configured to recognize the user's question and answer intention based on the question data and the historical question and answer records;

[0050] A target knowledge base determination module, configured to determine a corresponding target knowledge base according to the question and answer intention and the question data;

[0051] An answer data determination module, configured to determine each answer data corresponding to the question data based on the target knowledge base;

[0052] An optimal answer data determination module, configured to obtain the user's occupation information, determine optimal answer data from each of the answer data according to the occupation information, and feed back the optimal answer data to the user.

[0053] In a third aspect, an intelligent question and answer device includes a memory and a processor;

[0054] The memory is used for storing a program;

[0055] The processor is configured to execute the program to implement each step of the intelligent question and answer method according to any one of the first aspects.

[0056] In a fourth aspect, a storage medium stores a computer program, and when the computer program is executed by a processor, each step of the intelligent question and answer method according to any one of the first aspects is implemented.

[0057] As can be seen from the above technical solution, in this application, in response to a request instruction for asking and answering initiated by a user, question data and the user's historical Q&A records are obtained; the Q&A intention of the user is identified based on the question data and the historical Q&A records; a corresponding target knowledge base is determined according to the Q&A intention and the question data; various answer data corresponding to the question data are determined based on the target knowledge base; the user's occupation information is obtained, and the optimal answer data is determined from each of the answer data according to the occupation information, and the optimal answer data is fed back to the user. In this application, first, in response to the request initiated by the user, question data is obtained, and at the same time, the user's historical Q&A records are obtained. The historical Q&A records can indicate information such as the user's characteristics and tendencies. Therefore, based on the question data and the historical Q&A records, the Q&A intention of the user is identified, so that the purpose of the user can be better understood, and thus an answer that suits the user can be given. In order to obtain a more accurate answer, a corresponding target knowledge base is determined according to the Q&A intention and the question data. That is to say, this target knowledge base contains answer data that can be used to answer the question data. At the same time, considering the user's occupation, the answer that is most matched and suitable for the user is selected from each of the answer data. Then, this answer is both standard and what the user really needs, greatly improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0059] Figure 1 An optional flowchart of an intelligent Q&A method provided by an embodiment of this application;

[0060] Figure 2 Another optional flowchart of an intelligent Q&A method provided by an embodiment of this application;

[0061] Figure 3 A schematic structural diagram of an intelligent Q&A device provided by an embodiment of this application;

[0062] Figure 4 A schematic structural diagram of an intelligent Q&A device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0064] An embodiment of the present invention provides an intelligent question-answering method. This method can be applied to various computer terminals or intelligent terminals, and its execution subject can be a processor or server of a computer terminal or an intelligent terminal. The method flow chart of the method is as Figure 1 shown, and specifically includes:

[0065] S1: In response to a request instruction for asking and answering initiated by a user, obtain question data and the user's historical question-and-answer records.

[0066] This application can be applied to any software, APP, platform, system, device, etc. where an intelligent question-answering scenario is located. At the same time, the user can input question data in the interaction interface equipped with the intelligent question-answering software, platform, system, etc., which is equivalent to initiating a request instruction for asking and answering. After the response, the question data input by the user can be obtained, and the user's account information or personal information can also be obtained, so as to extract the user's historical question-and-answer records. The historical question-and-answer records can be the user's historical chat records in this question-answering scenario.

[0067] The user's historical question-and-answer records can reflect the user's interests and preferences, knowledge level, behavior habits, language and expression styles, emotional characteristics, etc. Further, it can make the subsequent question answering more natural, help the user understand and absorb the answers, and improve the fluency of the interaction.

[0068] In the existing intelligent question-answering process, the user's historical question-and-answer records are usually not involved, and the context information is not fully considered, which will result in a relatively simple final answer. The answers obtained by different users asking the same question are all the same, and cannot meet the actual needs of users.

[0069] S2: Identify the user's question-answering intention based on the question data and historical question-and-answer records.

[0070] The question data is the data input by the user and is the most intuitive data reflecting the user's current question-answering intention. However, it is not accurate enough to identify the question-answering intention only based on the question data. It will be more accurate if the question-answering intention is identified based on both the question data and the historical question-and-answer records.

[0071] On the other hand, the problem data only intuitively reflects the user's surface question-and-answer intention, and the historical question-and-answer records only reflect some historical characteristics of the user. Then, combining the two can identify the user's question-and-answer intention more deeply and accurately, so that subsequent question answering can be carried out according to the question-and-answer intention, which can improve the accuracy of the answer and enhance the naturalness of the interaction.

[0072] S3: Determine the corresponding target knowledge base according to the question-and-answer intention and the problem data.

[0073] It can be understood that different knowledge bases (such as databases) contain different knowledge and data. Determining the most suitable and optimal knowledge base as the target knowledge base according to the question-and-answer intention and the problem data can obtain more accurate answers, reduce noise at the same time, avoid mis-extracting some irrelevant knowledge, and thus improve the pertinence of the answer.

[0074] Determining the corresponding target knowledge base can also centrally use knowledge, improve resource utilization rate, and save computing resources.

[0075] S4: Determine each answer data corresponding to the problem data based on the target knowledge base.

[0076] Determining the corresponding target knowledge base to search for answers can improve efficiency, save time, and improve the adaptability of the answer.

[0077] It can be understood that different users will ask different questions, and different questions also exist in different fields. Then, different questions may correspond to the same answer, and the same question may also correspond to multiple different answers. At the same time, to ensure that no correct and appropriate answer is missed, this step is to determine one or more answer data corresponding to the problem data from the target knowledge base.

[0078] S5: Obtain the user's occupation information, determine the optimal answer data from each of the answer data according to the occupation information, and feedback the optimal answer data to the user.

[0079] In order to provide customized answers to users, this step obtains the user's occupation information and determines the optimal answer data from each of the answer data according to the occupation information. This is because users in different occupations and positions have different information needs. Screening out the optimal answer data based on the occupation information will make users feel that the answer is very considerate and make users feel understood, thus greatly enhancing the question-and-answer experience. In addition, some irrelevant or less relevant answer data can also be filtered out according to the occupation information.

[0080] Occupation information can be requested from the user, for example, by asking the user through a dialogue interaction, or by searching the historical Q&A records to see if the user has previously revealed their occupation, or by inferring the user's occupation information based on the historical Q&A records.

[0081] As can be seen from the above technical solution, in this application, in response to a request instruction for Q&A initiated by the user, question data and the user's historical Q&A records are obtained; the Q&A intention of the user is identified based on the question data and the historical Q&A records; a corresponding target knowledge base is determined according to the Q&A intention and the question data; each answer data corresponding to the question data is determined based on the target knowledge base; the occupation information of the user is obtained, the optimal answer data is determined from each of the answer data according to the occupation information, and the optimal answer data is fed back to the user. In this application, first, in response to the request initiated by the user, the question data is obtained, and at the same time, the user's historical Q&A records are obtained. The historical Q&A records can show information such as the user's characteristics and tendencies. Therefore, based on the question data and the historical Q&A records, the Q&A intention of the user is identified, so that the purpose of the user can be better understood, and thus an answer that suits the user can be given. In order to obtain a more accurate answer, a corresponding target knowledge base is determined according to the Q&A intention and the question data. That is to say, the target knowledge base contains answer data that can be used to answer the question data. At the same time, considering the occupation of the user, the answer that is most matched and suitable for the user is selected from each of the answer data. Then, this answer is both standard and what the user really needs, greatly improving the user experience.

[0082] In the method provided by the embodiment of the present invention, the process of identifying the Q&A intention of the user based on the question data and the historical Q&A records is specifically described as follows:

[0083] Classify the historical Q&A records according to each preset Q&A type to obtain each group of Q&A data under each Q&A type;

[0084] Extract the data summary of each group of the Q&A data;

[0085] Perform a completion process on the question data to obtain complete question information;

[0086] Based on the complete question information and the data summary of each group of the Q&A data, identify the Q&A intention of the user.

[0087] In existing solutions, although historical Q&A records are also extracted, they are directly used after extraction without summarizing the historical Q&A records and extracting valuable information, which often delays efficiency and cannot efficiently handle complex, cross-stage problems. Therefore, considering that historical Q&A records contain a lot of data content, in order to improve efficiency, the historical Q&A records are classified according to preset Q&A types. When setting Q&A types, multiple directions can be considered. For example, based on the user's purpose, they can be divided into information type, task type, conversation type, etc.; based on the user's needs, they can be divided into consultation type, guidance type, and entertainment type; based on the openness of the question, they can be divided into closed type and open type.

[0088] When classifying, the type can be determined based on keywords by extracting keywords. For example, if the question data contains words related to "law", then when classifying according to the openness of the question, the question data is of the closed type; for example, if the question data contains interrogative words or sentence patterns, when classifying according to the user's needs, the question data is of the consultation type. Thus, the historical Q&A records are divided into multiple groups of Q&A data. Below, valuable information is extracted in units of groups to generate a concise and clear data summary, avoiding the problem that the data in the historical Q&A records is messy and the summary cannot be quickly extracted.

[0089] In some cases, the user only expresses the question concisely, or the user does not express the question completely, so the user's true intention cannot be recognized. Therefore, the question data can be completed to help the user perfect the question and obtain complete question information, which is helpful for accurately identifying the user's Q&A intention.

[0090] Among them, in the process of identifying the Q&A intention based on the complete question information and the data summary of each group of Q&A data, the text features, semantic features, context features, and behavior features of each group of Q&A data can be extracted, and the Q&A intention is determined by combining these features and the complete question information.

[0091] The following embodiments will detail the process of determining the corresponding target knowledge base according to the Q&A intention and question data in the present application.

[0092] The question data is vectorized to obtain a question vector;

[0093] In a pre-established vector database, similarity retrieval is performed on the question vector to screen out the corresponding various tools;

[0094] Determine the description information of each of the tools;

[0095] The Q&A intention and the question data are combined, and after combination, they are respectively matched with the description information of each of the tools to calculate the matching degree;

[0096] Filter the target tool according to the matching degree, and use the database corresponding to the target tool as the target knowledge base.

[0097] Specifically, LangChain can be introduced in this step. LangChain is a framework for developing applications driven by language models, dedicated to simplifying the development of AI model applications. It provides a set of tools, components, and interfaces that can simplify the process of creating applications supported by large language models (LLMs) and chat models. LangChain can easily manage interactions with language models, link multiple components together, and integrate additional resources such as APIs and databases. It provides a series of tools to enhance the capabilities of the model, including API call tools, data retrieval tools, code execution tools, calculation tools, etc.

[0098] In addition, a vector database is introduced. A vector database is a database system for storing and retrieving vector data. In traditional relational databases, data is stored and queried in tabular form, while vector databases store data in vector form and provide efficient vector indexing and query functions. Vector databases are usually used to process large-scale high-dimensional vector data, such as image features, text representations, and voice features, etc. It can support application scenarios such as complex similarity search and clustering analysis.

[0099] However, in this application, a vector database is established and encapsulated based on the above tools. Various tools and their related descriptions are stored in this vector database. By vectorizing the question data to obtain a question vector, and then performing similarity retrieval on the question vector in this vector database, one or more corresponding tools can be filtered out, and at the same time, the description information of the tools is determined. The description information can be used for matching, combining the Q&A intention and the question data, and then respectively matching the combined result with the description information of each filtered tool. Methods such as cosine similarity, Euclidean distance, and Jaccard similarity coefficient can be used to calculate the matching degree, and filtering is performed according to the matching degree. For example, a value is set, and the tool with a matching degree greater than this value is used as the target tool. Among multiple open-source databases, the database corresponding to the target tool is used as the target knowledge base, and the format of the target tool can be unified to facilitate finding the target knowledge base.

[0100] The vector database established and encapsulated in this application supports multiple output return formats, has a built-in parameter recognition mechanism, newly established a parameter class, so it is more flexible, and at the same time supports the invocation of tools through front-end configuration.

[0101] Optionally, ElasticSearch can also be established. ElasticSearch is an open-source distributed search and analysis engine that has a storage of text vector models. Therefore, the question data can be vectorized to obtain a question vector, and the similarity search of the question vector can be performed in ElasticSearch.

[0102] The process of determining each answer data corresponding to the question data in the present application will be specifically described below.

[0103] Determine the question vector corresponding to the question data;

[0104] Obtain a pre-trained re-ranking model corresponding to the target knowledge base;

[0105] Use the re-ranking model to process the question vector corresponding to the question data to obtain each answer vector; the re-ranking model is trained with a vector sample set as the training sample and the answer vector sample corresponding to each vector sample in the vector sample set as the sample label.

[0106] Convert each answer vector into the corresponding answer data respectively.

[0107] Specifically, vectorizing the question data to obtain a question vector can also facilitate the training of a re-ranking model, which corresponds to the target knowledge base. That is to say, a knowledge base will correspond to a re-ranking model. The re-ranking model uses a vector sample set as the training sample, and the vector sample set contains countless vector samples. It is trained with the answer vector sample corresponding to each vector sample as the sample label. Then, using the trained re-ranking model to process the user's question vector, accurate answer vectors can be obtained, and then converted into non-vector form, such as text-form answer data. The re-ranking model can adopt a Transformer-based model or a text vector model (such as gte-base-zh, that is, the General Text Embedding Base Model), etc. Using the trained re-ranking model to process the question vector can capture richer semantic information in the question vector to generate more natural and coherent answers, and can also support multiple languages, without being restricted by the user's language. And subsequently, the re-ranking model can be continuously optimized using the user's questions to make it more accurate and efficient.

[0108] The refined ranking model includes an answer generation module and a ranking module. Among them, the output end of the answer generation module is connected to the input end of the ranking module. After the question vector is input into the answer generation module, the answer generation module processes the question vector to obtain multiple answer vectors. The answer vectors are input into the ranking module through the connection between the two modules. The ranking module performs refined ranking on the multiple answer vectors and outputs the several answer vectors with higher rankings (such as the top five) in the order of refined ranking.

[0109] Among them, the construction process of the vector sample set may include:

[0110] Convert each piece of knowledge data in the target knowledge base into each piece of knowledge vector;

[0111] For each piece of the knowledge vector, calculate the correlation between this piece of knowledge vector and each other piece of knowledge vector respectively to determine the degree of correlation;

[0112] Take each other piece of knowledge vector with a degree of correlation not less than the preset correlation threshold as each corresponding knowledge vector of this knowledge vector;

[0113] Combine this piece of knowledge vector with each of its corresponding knowledge vectors respectively to obtain each vector sample;

[0114] Summarize the historical Q&A records of all users to obtain a Q&A set;

[0115] Extract all historical question data and corresponding historical answer data from the Q&A set, convert each piece of the historical question data into a corresponding historical question vector respectively, and at the same time convert each piece of the historical answer data into a corresponding historical answer vector;

[0116] Combine each of the historical question vectors with the corresponding historical answer vector to obtain each vector sample;

[0117] Summarize each of the vector samples to obtain the vector sample set.

[0118] Specifically, the sample quality and form used in the training process of the fine-ranking model are crucial. When constructing the vector sample set, considering that one piece of knowledge data is associated with one or more pieces of knowledge data, a combined form is adopted. A knowledge vector is combined with each of the other relevant knowledge vectors respectively to obtain each vector sample, which can enrich the sample data and enable the model to achieve better performance during training. At the same time, considering that one question corresponds to one or more answers, and one answer corresponds to one or more questions, all users' historical Q&A records are summarized, all historical question data and corresponding historical answer data are extracted, and then vectorized. Each historical question vector is combined with the corresponding historical answer vector to obtain each vector sample. Then, training the model with these two forms of vector samples can make the fine-ranking model more accurate and more adaptable.

[0119] During the training process, the fine-ranking model will conduct recall evaluation and screen out appropriate multi-way recalls. Based on the recall results, the model is optimized and debugged until the fine-ranking model reaches the optimal state.

[0120] Optionally, the process of determining the optimal answer data from each of the answer data according to the occupational information and feeding back the optimal answer data to the user may include the following steps:

[0121] Extract the occupational characteristics of multiple preset occupations;

[0122] Cluster the occupational characteristics of each of the preset occupations to form different types of feature clusters;

[0123] Calculate the feature degree of each of the feature clusters;

[0124] Determine the feature cluster to which the occupational information belongs from each of the feature clusters as the target cluster;

[0125] Compare the feature degree of the target cluster with a preset feature threshold to determine the optimal answer data.

[0126] Specifically, multiple occupations can be preset, such as programmers, doctors, teachers, etc. The occupational characteristics include professional knowledge, skills, common terms, work scenarios, etc. Among them, the feature descriptions can be obtained from occupation-related data sources, and then these occupational characteristics can be extracted through natural language processing techniques. Next, these occupational characteristics are clustered to obtain different types of feature clusters. The clustering algorithms can be K-means, hierarchical clustering, etc., which are not limited in this embodiment. In this way, similar occupations can be classified to reduce redundant data processing, and then the feature degree is calculated. The feature degree refers to the concentration of features within the cluster. For example, the average distance within each feature cluster and the sample proportion of each feature cluster can be calculated. For each feature cluster, dividing its sample proportion by the average distance within the cluster can obtain the feature degree of the feature cluster. It is also necessary to determine the feature cluster to which the occupation information belongs as the target feature cluster, and determine the optimal answer data based on the feature degree of the target feature cluster, so as to meet the needs of different occupational groups, optimize resource allocation, and improve the answer efficiency.

[0127] Then, further, the steps of comparing the feature degree of the target cluster with a preset feature threshold to determine the optimal answer data are as follows:

[0128] Calculate the similarity between each of the answer data and the question data respectively;

[0129] If the feature degree of the target cluster is less than the feature threshold, the answer data corresponding to the highest similarity is used as the optimal answer data;

[0130] If the feature degree of the target cluster is not less than the feature threshold, one or more answer data with similarities greater than a preset similarity threshold are combined to obtain the optimal answer data.

[0131] Specifically, considering that the feature degrees of different target clusters are high or low, a high feature degree means that the features of the target cluster are concentrated and representative, while a low feature degree means that the features of the target cluster are dispersed or have more noise. Therefore, in order to obtain more accurate answers, a preset feature threshold is used to distinguish the high and low feature degrees of the target cluster. Then, regardless of whether the feature degree is high or low, the optimal answer data can be obtained.

[0132] Furthermore, as Figure 2As shown, the present application can also set up a feedback mechanism. In steps S1 to S5, if there is a situation where it cannot be successfully executed, such as being unable to recognize the user's question-and-answer intention, or being unable to determine the corresponding target knowledge base, or being unable to determine the answer data corresponding to the question data, or being unable to determine the optimal answer data, then the question data is set as an unknown question, the unknown question is recorded, and the unknown question is used to optimize the fine-ranking model, or the unknown question is used to optimize the knowledge base / target knowledge base, providing more possibilities for subsequent intelligent question answering. In the existing solutions, although the knowledge base is also involved, the involved knowledge base is a static mode, lacking the optimization and update of knowledge, unable to achieve subsequent development and iteration, and unable to meet users in all fields and at all levels. The present application can also cluster all the recorded unknown questions, so that when other unknown questions appear again, the answers can be directly determined.

[0133] Optionally, the present application can be applied to an intelligent question-and-answer platform, built using docker services, with independent operation and maintenance deployment. Docker is an open-source platform for developing, deploying, and running application programs. Based on containerization technology, it can package the application program and its dependencies into an independent container, achieving rapid deployment and portability of the application. Applying the present application to an intelligent question-and-answer platform built using docker services can better implement the intelligent question-and-answer scenario. At the same time, it is necessary to store each question-and-answer record in units of users and perform caching so that historical question-and-answer records can be quickly read subsequently.

[0134] Corresponding to Figure 1 the method described above, an embodiment of the present invention also provides an intelligent question-and-answer device for Figure 1 the specific implementation of the method in. The intelligent question-and-answer device provided by the embodiment of the present invention can be in a computer terminal or various mobile devices. Combining Figure 3 , the intelligent question-and-answer device is introduced. As Figure 3 shown, the device can include:

[0135] A response module 10, configured to respond to a request instruction for question and answer initiated by a user, and obtain question data and the user's historical question-and-answer records;

[0136] A question-and-answer intention recognition module 20, configured to recognize the user's question-and-answer intention based on the question data and historical question-and-answer records;

[0137] A target knowledge base determination module 30, configured to determine a corresponding target knowledge base according to the question-and-answer intention and question data;

[0138] An answer data determination module 40, configured to determine each answer data corresponding to the question data based on the target knowledge base;

[0139] The optimal answer data determination module 50 is configured to obtain the occupation information of the user, determine the optimal answer data from each of the answer data according to the occupation information, and feedback the optimal answer data to the user.

[0140] As can be seen from the above technical solution, in this application, in response to a request instruction for asking and answering initiated by a user, question data and the user's historical question-and-answer records are obtained; the question-and-answer intention of the user is identified based on the question data and the historical question-and-answer records; a corresponding target knowledge base is determined according to the question-and-answer intention and the question data; each answer data corresponding to the question data is determined based on the target knowledge base; the occupation information of the user is obtained, the optimal answer data is determined from each of the answer data according to the occupation information, and the optimal answer data is feedback to the user. In this application, first, in response to the request initiated by the user, the question data is obtained, and at the same time, the user's historical question-and-answer records are obtained. The historical question-and-answer records can indicate information such as the characteristics and tendencies of the user. Therefore, based on the question data and the historical question-and-answer records, the question-and-answer intention of the user is identified, so that the purpose of the user can be better understood, and thus an answer that conforms to the user can be given. In order to obtain a more accurate answer, a corresponding target knowledge base is determined according to the question-and-answer intention and the question data. That is to say, this target knowledge base contains answer data that can be used to answer the question data. At the same time, considering the occupation of the user, the answer that is most matched and suitable for the user is selected from each of the answer data. Then, this answer is both standard and what the user really needs, greatly improving the user experience.

[0141] Furthermore, an embodiment of this application provides an intelligent question-and-answer device. Optionally, Figure 4 shows a hardware structure block diagram of the intelligent question-and-answer device. Referring to Figure 4 , the hardware structure of the intelligent question-and-answer device may include: at least one processor 01, at least one communication interface 02, at least one memory 03, and at least one communication bus 04.

[0142] In the embodiment of this application, the number of the processor 01, the communication interface 02, the memory 03, and the communication bus 04 is at least one, and the processor 01, the communication interface 02, and the memory 03 complete mutual communication through the communication bus 04.

[0143] The processor 01 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.

[0144] The memory 03 may include a high-speed RAM memory and may also include a non-volatile memory, such as at least one disk memory.

[0145] Among them, the memory stores a program, and the processor can call the program stored in the memory. The program is used to execute the following intelligent question-answering method, including:

[0146] In response to a request instruction for question-answering initiated by the user, obtain the question data and the user's historical question-answering records;

[0147] Identify the user's question-answering intention based on the question data and historical question-answering records;

[0148] Determine the corresponding target knowledge base according to the question-answering intention and question data;

[0149] Determine each answer data corresponding to the question data based on the target knowledge base;

[0150] Obtain the user's occupation information, determine the optimal answer data from each of the answer data according to the occupation information, and feedback the optimal answer data to the user.

[0151] Optionally, the refined functions and extended functions of the program can refer to the description of the intelligent question-answering method in the method embodiments.

[0152] The embodiment of the present application also provides a storage medium. The storage medium can store a program suitable for execution by a processor. When the program runs, it controls the device where the storage medium is located to execute the following intelligent question-answering method, including:

[0153] In response to a request instruction for question-answering initiated by the user, obtain the question data and the user's historical question-answering records;

[0154] Identify the user's question-answering intention based on the question data and historical question-answering records;

[0155] Determine the corresponding target knowledge base according to the question-answering intention and question data;

[0156] Determine each answer data corresponding to the question data based on the target knowledge base;

[0157] Obtain the user's occupation information, determine the optimal answer data from each of the answer data according to the occupation information, and feedback the optimal answer data to the user.

[0158] Specifically, the storage medium may be a computer-readable storage medium, and the computer-readable storage medium may be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM.

[0159] Optionally, the refinement function and the extension function of the program may refer to the description of the intelligent question-and-answer method in the method embodiments.

[0160] In addition, in each of the embodiments of the present disclosure, the functional modules may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a live broadcast device, or a network device, etc.) to execute all or part of the steps of the methods in the embodiments of the present disclosure.

[0161] Finally, it should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0162] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments may be referred to each other.

[0163] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An intelligent question-answering method, characterized in that: include: In response to a question-and-answer request from a user, obtain question data and the user's historical question-and-answer records; Identify the user's question-answering intention based on the question data and historical question-answering records; Determine a corresponding target knowledge base according to the question-answering intention and question data; Determine each answer data corresponding to the question data based on the target knowledge base; The occupation information of the user is obtained, the best answer data is determined from each of the answer data according to the occupation information, and the best answer data is fed back to the user.

2. The method according to claim 1, characterized in that The identifying the user's question-answering intention based on the question data and historical question-answering records includes: Classify the historical question and answer records according to the preset question and answer types to obtain each group of question and answer data under each question and answer type; Extracting a data summary of each set of the question-and-answer data; Complete the problem data to obtain complete problem information; Based on the complete question information and the data summary of each set of the question and answer data, the question and answer intention of the user is identified.

3. The method according to claim 1, characterized in that Determining the corresponding target knowledge base according to the question-answering intention and question data includes: Vectorizing the problem data to obtain a problem vector; Perform similarity search on the problem vector in a pre-established vector database to screen out corresponding tools; determining descriptive information for each of said tools; The question-answering intention and the question data are combined, and the combined information is matched with the description information of each tool to calculate the matching degree; The target tool is screened according to the matching degree, and the database corresponding to the target tool is used as the target knowledge base.

4. The method according to claim 1, characterized in that: The determining of each answer data corresponding to the question data based on the target knowledge base includes: Determining a problem vector corresponding to the problem data; Acquire a pre-trained refined ranking model corresponding to the target knowledge base; The refined ranking model is used to process the question vector corresponding to the question data to obtain each answer vector; the refined ranking model is trained using a vector sample set as a training sample and an answer vector sample corresponding to each vector sample in the vector sample set as a sample label; Each of the answer vectors is converted into corresponding answer data respectively.

5. The method according to claim 4, characterized in that The process of constructing the vector sample set includes: Convert each piece of knowledge data in the target knowledge base into each piece of knowledge vector; For each of the knowledge vectors, calculating the correlation between the knowledge vector and other knowledge vectors to determine the correlation; The other knowledge vectors whose relevance is not less than a preset relevance threshold are used as the corresponding knowledge vectors of the knowledge vector; Combining the knowledge vector with each of its corresponding knowledge vectors to obtain vector samples; Summarize all users' historical question and answer records to obtain a question and answer set; Extract all historical question data and corresponding historical answer data from the question and answer set, and convert each piece of the historical question data into a corresponding historical question vector, and convert each piece of the historical answer data into a corresponding historical answer vector; Combining each of the historical question vectors with the corresponding historical answer vectors to obtain vector samples; The vector samples are aggregated to obtain the vector sample set.

6. The method according to any one of claims 1 to 5, characterized in that The determining the best answer data from each of the answer data according to the occupation information, and feeding back the best answer data to the user, comprises: Extracting occupational characteristics of multiple preset occupations; Clustering the occupational characteristics of each of the preset occupations to form characteristic clusters of different types; Calculating the characteristic degree of each of the feature clusters; Determine the feature cluster to which the occupation information belongs from each of the feature clusters as a target cluster; The characteristic degree of the target cluster is compared with a preset characteristic threshold to determine the optimal answer data.

7. The method according to claim 6, characterized in that The step of comparing the characteristic degree of the target cluster with a preset characteristic threshold to determine the optimal answer data includes: Calculate the similarity between each of the answer data and the question data respectively; If the characteristic degree of the target cluster is less than the characteristic threshold, the answer data corresponding to the highest similarity is taken as the optimal answer data; If the characteristic degree of the target cluster is not less than the characteristic threshold, one or more answer data corresponding to similarities greater than a preset similarity threshold are combined to obtain optimal answer data.

8. An intelligent question-answering device, characterized in that: include: A response module, used to respond to a question-and-answer request instruction initiated by a user, and obtain question data and the user's historical question-and-answer records; A question-answering intention recognition module, used to recognize the user's question-answering intention based on the question data and historical question-answering records; A target knowledge base determination module is used to determine the corresponding target knowledge base according to the question-answering intention and question data; An answer data determination module, used to determine each answer data corresponding to the question data based on the target knowledge base; The optimal answer data determination module is used to obtain the user's occupation information, determine the optimal answer data from each of the answer data based on the occupation information, and feed back the optimal answer data to the user.

9. An intelligent question-answering device, characterized in that: including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the intelligent question-answering method as described in any one of claims 1-7.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the intelligent question-answering method as described in any one of claims 1 to 7 is implemented.