Question answering method and device based on natural language processing, equipment and storage medium

By using a question-and-answer method based on natural language processing, the problem of low response efficiency in self-service customer service systems has been solved. This method enables automatic preprocessing of user input questions and the integration of behavioral information, providing efficient and accurate question-and-answer services.

CN112231452BActive Publication Date: 2025-10-24PING AN TRUST CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011085684.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-12
Publication Date
2025-10-24
Estimated Expiration
2040-10-12

AI Technical Summary

Technical Problem

Existing self-service customer service systems are inefficient at responding to questions, especially for users unfamiliar with the business, who find it difficult to accurately obtain the answers they need.

Method used

It adopts a question-answering method based on natural language processing, pre-processes user input questions, matches and re-orders sample questions, and combines user behavior information to provide targeted and accurate answers.

Benefits of technology

It reduces the amount of user operations, improves the pertinence and accuracy of answers, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112231452B_ABST
    Figure CN112231452B_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence, and provides a natural language processing question and answer method, device, equipment and storage medium, input question is acquired when a question and answer instruction is received, and the input question is preprocessed to generate input question characteristics; a corresponding matching sample question set is determined in a preset database according to the input question characteristics; user behavior information of a target user is acquired, each matching sample question is reordered according to the user behavior information, and a target sample question corresponding to the question and answer instruction is determined according to a sorting result; target sample reply information corresponding to the target sample question is acquired, and the target sample reply information is output. In addition, the application can be applied to the field of intelligent medical treatment, and the question and answer of intelligent customer service are carried out. In addition, the application also relates to the blockchain technology, and the preset database and the user behavior information can be stored in the blockchain. The application combines the ideas of artificial intelligence and natural language processing, is favorable for improving the pertinence and accuracy of answers, and improving the use experience of users.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a natural language processing question and answer method, device, equipment and storage medium. BACKGROUND

[0002] The traditional consulting industry is realized through artificial customer service to realize communication with users. This customer service mode needs to invest a large amount of human cost. With the development of science and technology, self-service customer service has appeared on the market. The user is guided to select the problem type layer by layer through the navigation menu, and then the corresponding answer is given according to the user's selection. This method needs the user to operate several times to get the answer he needs. For users who are not familiar with the business, it will lead to their inability to select the correct problem option, and thus they cannot get the answer they want. Therefore, how to solve the low efficiency of the problem reply of the existing self-service customer service system has become a technical problem to be solved at present. SUMMARY

[0003] The main purpose of the present application is to provide a natural language processing-based question and answer method, device, equipment and storage medium, which aims to solve the technical problem of low efficiency of the problem reply of the existing self-service customer service system.

[0004] To achieve the above-mentioned purpose, the embodiment of the present application provides a natural language processing-based question and answer method, which comprises the following steps:

[0005] When receiving a question and answer instruction triggered by a target user operation, an input question in the question and answer instruction is obtained, and the input question is preprocessed to generate an input question feature;

[0006] According to the input question feature, sample question matching is performed in a preset database, the matching sample question corresponding to the input question is determined, and a matching sample question set is generated;

[0007] The user behavior information of the target user is obtained, the matching sample questions in the matching sample question set are reordered according to the user behavior information, and the target sample question corresponding to the question and answer instruction is determined in the matching sample question set according to the sorting result;

[0008] The target sample answer information corresponding to the target sample question is obtained, and the target sample answer information is output.

[0009] In addition, to achieve the above-mentioned purpose, the embodiment of the present application also provides a natural language processing-based question and answer device, which comprises:

[0010] An instruction receiving module is configured to, when receiving a question and answer instruction triggered by a target user operation, acquire an input question in the question and answer instruction, and pre-process the input question to generate input question features.

[0011] A question matching module is configured to perform sample question matching in a preset database according to the input question features, determine a matching sample question corresponding to the input question, and generate a matching sample question set.

[0012] A question rearrangement module is configured to acquire user behavior information of the target user, rearrange the matching sample questions in the matching sample question set according to the user behavior information, and determine a target sample question corresponding to the question and answer instruction in the matching sample question set according to a rearrangement result.

[0013] An information output module is configured to acquire target sample reply information corresponding to the target sample question, and output the target sample reply information.

[0014] In addition, to achieve the above object, the embodiment of the present application further provides a question and answer device based on natural language processing, which comprises a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein the computer program is executed by the processor to implement the steps of the question and answer method based on natural language processing.

[0015] In addition, to achieve the above object, the embodiment of the present application further provides a storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of the question and answer method based on natural language processing.

[0016] The user of the embodiment of the present application can input own question information according to actual conditions, the question and answer device (or terminal, server, etc.) combines the ideas of artificial intelligence and natural language processing, automatically performs pre-processing and question retrieval according to the input of the user, simultaneously obtains corresponding sample questions and sample answers combining the behavior information of the user, and outputs the sample answers to provide question and answer services for the user, which is beneficial to reduce the operation amount of the user, even for the user who is not familiar with the business, the user can also obtain question answers; and since the answers to the questions are determined based on the input questions of the user and the behavior information of the user, the pertinence and accuracy of the answers are improved, and the use experience of the user is improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 The figure is a schematic diagram of the hardware structure of the question and answer device based on natural language processing involved in the embodiment of the present application.

[0018] Figure 2Flowchart of the first embodiment of the application for the question and answer method based on natural language processing;

[0019] Figure 3 Functional module diagram of the first embodiment of the application for the question and answer device based on natural language processing.

[0020] The implementation, functional features and advantages of the application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are merely intended to explain the application and not to limit the application.

[0022] The question and answer method based on sample matching involved in the embodiments of the application is mainly applied to a question and answer device based on natural language processing, which can be a server, a PC, a portable computer, a mobile terminal or the like device having display and processing functions.

[0023] Referring to Figure 1 , Figure 1 The hardware structure diagram of the question and answer device based on natural language processing involved in the embodiments of the application is shown in FIG. 1. In the embodiments of the application, the question and answer device based on natural language processing can include a processor 1001 (such as a CPU), a communication bus 1002, a user interface 1003, a network interface 1004 and a memory 1005. The communication bus 1002 is used to realize the connection and communication among these components; the user interface 1003 can include a display screen (Display) and an input unit such as a keyboard (Keyboard); the network interface 1004 can optionally include a standard wired interface and a wireless interface (such as a WI-FI interface); the memory 1005 can be a high-speed RAM memory or a stable memory (non-volatile memory) such as a disk memory, and the memory 1005 can optionally be a storage device independent of the aforementioned processor 1001.

[0024] Those skilled in the art can understand that Figure 1 The hardware structure shown in FIG. 1 does not constitute a limitation on the question and answer device based on sample matching, and can include more or fewer components than those shown in the figure, or combine certain components, or different component arrangements.

[0025] Continuing to refer to Figure 1 , Figure 1 The memory 1005 as a computer readable storage medium in the embodiments of the application can include an operating system, a network communication module and a computer program.

[0026] In Figure 1In the embodiment, the network communication module is mainly used for connecting the database and communicating data with the database; and the processor 1001 can call the computer program stored in the memory 1005 and execute the question and answer method based on natural language processing provided by the embodiment.

[0027] The embodiment of the application provides a question and answer method based on natural language processing.

[0028] With reference to Figure 2 , Figure 2 FIG. 1 is a flowchart of a first embodiment of the question and answer method based on natural language processing.

[0029] In the embodiment, the question and answer method based on natural language processing comprises the following steps:

[0030] In step S10, when receiving a question and answer instruction triggered by a target user operation, an input question in the question and answer instruction is acquired, and the input question is preprocessed to generate input question features.

[0031] The traditional customer service industry realizes communication with users through manual customer service, which needs to invest a large amount of manpower. With the development of science and technology, self-service customer service begins to appear on the market, which guides users to select problem types layer by layer through a navigation menu, and then gives corresponding answers according to the user's selection. This method needs the user to perform multiple operations to obtain the answer he needs. For users who are not familiar with the business, it will lead to their inability to select the correct problem option, and thus cannot obtain the answer they want. Therefore, how to solve the low problem reply efficiency of the existing self-service customer service system has become a technical problem to be solved at present. In view of this, the embodiment provides a question and answer method based on natural language processing. The user can input his own problem information according to the actual situation, and the question and answer device (or terminal, server, etc.) combines the ideas of artificial intelligence and natural language processing, automatically preprocesses and searches for problems according to the user's input, and at the same time, combines the user's behavior information to obtain corresponding sample questions and sample answers, and outputs the sample answers to provide question and answer services for the user, which is beneficial to reduce the user's operation amount. Even for users who are not familiar with the business, they can also get problem solutions. Moreover, since the answers to the problems are determined based on the user's input questions and behavior information, it is beneficial to improve the pertinence and accuracy of the answers and improve the user's use experience.

[0032] The question answering method based on natural language processing in this embodiment can be implemented by a server. For example, a user sends question information to the server through a user terminal (such as a user's own mobile phone, a dedicated customer service robot, or the like), and the server answers according to the question information. Of course, the method can also be completed independently by the user terminal (or a dedicated customer service robot). For example, the user operates on his own mobile phone, and the mobile phone itself independently implements the present solution. For the sake of convenience, the implementation by the server is taken as an example for description in the following.

[0033] Before answering, the server first acquires question information of the user. The question information can be input by the user after triggering a question instruction on the user terminal, and sent to the server by the user terminal. For example, the user inputs his own question information on his own mobile phone in a manual input or voice input manner, and the mobile phone sends the question information to the server. For the sake of convenience, the question information acquired by the server can be referred to as input question. When receiving the input question, the server pre-processes the input question to obtain standardized and structured input question features, so as to facilitate subsequent question retrieval through the input question features. The input question features can be represented in the form of a feature vector (which can be referred to as an input feature vector), a feature matrix (which can be referred to as an input feature matrix), a vector map (which can be referred to as an input vector map), or the like.

[0034] Further, the preprocessing includes text segmentation, keyword extraction, synonym expansion, sentence vector acquisition, etc. Among them, text segmentation refers to the process of recombining continuous character sequences in the text into word sequences according to certain specifications. Text segmentation can be achieved in various ways, such as forward maximum matching, reverse maximum matching, feature scanning (marker segmentation), and segmentation based on statistical models. Keyword extraction refers to extracting business feature keywords (or problem feature keywords) from the word sequence obtained by segmentation. Keyword extraction can be achieved by string matching, that is, a number of sample keywords are defined in advance according to the business situation, and then the word sequence is compared with the sample keywords to identify and extract the keywords that match the sample keywords. Synonym expansion obtains synonyms of the keywords (words with the same or similar meaning as the keywords). The synonym expansion can be achieved by a corpus database, that is, a corpus database is set in advance, which includes a number of synonym sets, each of which includes a number of words with similar meanings. When performing synonym expansion, find the synonym set where the keyword is located in the database, and the other words in the set are synonyms of the keyword. It is worth noting that the keywords of the word sequence may not have corresponding synonyms. Sentence vector acquisition refers to obtaining the corresponding vector group according to the extracted keywords and expanded synonyms, that is, mapping the keywords and synonyms to obtain corresponding vector elements, and then combining these vector elements to obtain a vector. This vector can be considered as an input feature vector of the input question, which is used to represent the input question features of the input question. Since one keyword may correspond to multiple synonyms, the input question may have multiple input feature vectors, for example, the keywords of the input question are A1 and B1, and the keyword A1 has a synonym A2, then the input feature vector of the input question is (A1, B1) and (A2, B1), and A1, A2 and B1 can also be referred to as elements of the input feature vector.

[0035] In step S20, sample question matching is performed in the preset database according to the input question features, the matching sample question corresponding to the input question is determined, and a matching sample question set is generated.

[0036] In the embodiment, when the preprocessing result (input question feature) is obtained, the sample question retrieval and matching in the database can be performed according to the preprocessing result to determine the matching sample question corresponding to the input question. The database includes a plurality of sample questions, and sample answers and sample feature vectors corresponding to the sample questions. The sample question retrieval process is to calculate the similarity between the input question and the sample question by using the input feature vector of the input question and the sample feature vector of the sample question, and then to filter a plurality of sample questions with greater similarity from the data according to the similarity. The sample questions with greater similarity can be referred to as matching sample questions, and the matching sample questions form a matching sample set. The matching sample question can be a sample question with a feature vector similarity greater than a certain threshold, or a few sample questions with the greatest similarity. When calculating the similarity, the similarity can be calculated by using cosine similarity, BM25, or the like.

[0037] It is worth noting that an input question can have a plurality of input feature vectors, and the input feature vectors can be used to calculate the similarity respectively, and then the greatest similarity is taken as the similarity between the input question and the sample question.

[0038] In step S30, the user behavior information of the target user is obtained, the matching sample questions in the matching sample question set are reordered according to the user behavior information, and the target sample question corresponding to the question and answer instruction is determined from the matching sample question set according to the sorting result.

[0039] In the embodiment, when the matching sample set is obtained, the server will reorder the matching sample questions to determine the target sample question from the plurality of matching sample questions according to the sorting result. In the embodiment, the reordering is performed in combination with the user behavior information, thereby improving the pertinence and accuracy of the answer.

[0040] Specifically, the user behavior information includes historical browsing information and historical transaction information, and the step S30 includes:

[0041] In step S31, the historical browsing information and the historical transaction information of the target user are obtained, and the target interest label of the target user is determined according to the historical browsing information and the historical transaction information.

[0042] In this embodiment, the behavior information of the user can include the historical service record of the user and the historical browsing record, and the interest label of the user can be acquired according to the behavior information, the interest label being used to represent which business scenarios, business nodes, etc. the user is likely to intersect with in the near future, and the interest label is used to predict which business scenario or business node the user is likely to ask questions based on. For example, according to the behavior information of the user, it can be known that the user searches for chronic disease insurance twice in 24 hours, and the interest label of the user can be acquired, including chronic disease and insurance taboo, and the user's current question can be a consultation about the disease for a certain insurance business. Or, the user has input A trust fund multiple times in the historical search, and the interest label of the user can include A trust fund. The user's current question can be a consultation for the trust fund business.

[0043] In step S32, the target interest label is matched with the sample attribute labels of each matched sample question respectively, the number of matching labels corresponding to each matched sample question is determined, and each matched sample question is reordered according to the number of matching labels corresponding to each matched sample question.

[0044] For a sample question, there are corresponding sample attribute labels, and the sample attribute labels are used to represent the problem type involved in the sample question, including the business type, business node, operation flow type / amount type, etc. to which the sample question belongs; for example, the sample attribute label of a certain sample question includes chronic disease insurance, indicating that the sample question is a problem related to disease restrictions for insurance. Or, the sample attribute label includes A trust fund, indicating that the sample question is a problem related to the trust fund. Then, the server can compare the interest label of the user with the sample attribute labels of each matched sample question respectively, determine the number of matching labels of each matched sample question, and sort each matched sample question according to the number of matching labels, the greater the number of matching labels, the higher the sorting, so as to obtain the sorting result.

[0045] In step S33, the target sample question corresponding to the question and answer instruction is determined from the matched sample questions according to the sorting result.

[0046] When the sorting result is obtained, the higher the sorting of the matched sample question, the higher the repetition degree of the user's behavior, and then the target sample question can be determined from the matched sample questions according to the sorting result, for example, the matched sample question with the highest sorting is the target sample question, or the top X matched sample questions are the target sample questions.

[0047] Further, before sorting each matching sample question according to the number of matching labels, the number of matching labels can be mapped to the interval (0, 1) by first normalizing the mapping, for example, by mapping through the softmax function. Through the softmax function, the number of matching labels can be mapped to a value of (0, 1), and the sum of these values is 1 (satisfying the properties of probability). When the mapping value of the number of matching labels is obtained, each matching sample question is sorted according to the size of the mapping value. The larger the mapping value, the higher the ranking. At this time, the mapping value can be considered as the probability that the input question is equivalent to the matching sample question. Then, the maximum mapping value can be compared with a preset threshold (such as 0.8). If the maximum mapping value is greater than the preset threshold, the matching sample question corresponding to the maximum mapping value can be determined as the target sample question. If the maximum mapping value is less than the preset threshold, the top X matching sample questions can be determined as the target sample question, or the top X matching sample questions can be returned to the user end, and the target sample question can be determined according to the selection feedback returned by the user end.

[0048] In step S40, the target sample question corresponding target sample answer information is obtained, and the target sample answer information is output.

[0049] In the embodiment, when the target sample question is determined, the server can obtain the target sample answer corresponding to the target sample question in the database, and output the target sample answer to the user end, so that the user obtains the target sample answer.

[0050] It should be emphasized that, in order to further ensure the privacy and security of the above-mentioned sample questions, sample answers, and user behavior information, the above-mentioned database can be stored in a node of a blockchain, and the user behavior information can also be stored in a node of a blockchain.

[0051] The blockchain referred to in the embodiment is a new application mode of distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, and other computer technologies. Blockchain, in essence, is a decentralized database, which is a series of data blocks associated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-fake) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer.

[0052] In the embodiment, when receiving a question and answer instruction triggered by a target user operation, an input question in the question and answer instruction is acquired, and the input question is preprocessed to generate an input question feature; sample question matching is performed in a preset database according to the input question feature, a matching sample question corresponding to the input question is determined, and a matching sample question set is generated; user behavior information of the target user is acquired, and the matching sample questions in the matching sample question set are reordered according to the user behavior information, and a target sample question corresponding to the question and answer instruction is determined in the matching sample question set according to a sorting result; target sample reply information corresponding to the target sample question is acquired, and the target sample reply information is output. In this way, the user can input own question information according to actual conditions, and a question and answer device (or a terminal, a server, etc.) combines the ideas of artificial intelligence and natural language processing, automatically performs preprocessing and question retrieval according to user input, obtains corresponding sample questions and sample answers in combination with user behavior information, and outputs the sample answers to provide question and answer services for the user, which is beneficial to reducing the operation amount of the user, even for a user unfamiliar with a business, the user can also obtain question answers; and since the answers to the questions are determined based on the questions input by the user and the behavior information of the user, the pertinence and accuracy of the answers are improved, and the use experience of the user is improved.

[0053] Based on the first embodiment of the question and answer method based on natural language processing, the second embodiment of the question and answer method based on natural language processing is provided.

[0054] In the embodiment, the step S20 includes:

[0055] In the step S21, similarity between the input question and each sample question is calculated according to an input feature vector corresponding to the input question feature and a sample feature vector corresponding to each sample question in the database.

[0056] In the embodiment, when the preprocessing result (input question feature) is obtained, sample question retrieval and matching in the database can be performed according to the preprocessing result to determine the matching sample question corresponding to the input question. The database includes a plurality of sample questions, sample answers corresponding to the sample questions, and sample feature vectors. The process of sample question retrieval is to calculate the similarity between the input question and the sample question by using the input feature vector of the input question and the sample feature vector of the sample question.

[0057] In the embodiment, when calculating the similarity, a sample question is sequentially taken from each sample question as a current sample question, and a current sample feature vector of the current sample question is acquired. Then, the input feature vector and the current sample feature vector are substituted into a preset similarity formula to calculate the similarity between the input question and the current sample question, and the preset similarity formula is:

[0058] S =∑(w i *R(q i ))

[0059] S is the similarity between the input question and the current sample question; q i is the i-th element in the input feature vector Q; w i is the weight of q i , the more sample feature vectors containing q i in the database, the smaller w i is; R(q i ) is the relevance score of q i to the current sample feature vector, R(q i ) is determined according to the number of times q i appears in the input feature vector, the number of times q i appears in the current sample feature vector, the number of elements in the current sample feature vector, and the average number of elements of sample feature vectors of all sample questions in the database.

[0060] Further, the w i is calculated according to a preset weight formula, and the preset weight formula is:

[0061]

[0062] N is the number of sample questions in the database; n(q i ) is the number of sample feature vectors containing q i in the database. For the setting of w i , the more sample feature vectors containing q i in the knowledge base, the lower the weight of q i is. That is, when the sample feature vectors of many sample questions all contain q i , the discrimination of q i is not high, and therefore the importance of using q i for query is low.

[0063] Further, regarding R(q i ), R(q i ) is calculated according to the number of times q i appears in the input feature vector, the number of times q i appears in the current sample feature vector, the number of elements in the current sample feature vector, and the average number of elements of sample feature vectors of all sample questions in the database, and a preset relevance formula, and the preset relevance formula is:

[0064]

[0065] wherein, k1, k2, b are preset parameters, and are all greater than zero;

[0066] F1(q i ) is the number of occurrences of q i in the current sample feature vector; F2(q i ) is the number of occurrences of q i in the input feature vector; dl is the number of elements of the current sample feature vector; avgdl is the average number of elements of the sample feature vectors of all sample problems in the database.

[0067] Step S22, determining a matching sample problem in each sample problem according to the similarity between the input problem and each sample problem, wherein the similarity between the matching sample problem and the input problem is greater than a preset threshold.

[0068] After obtaining the similarity between the input problem and each sample problem, a number of matching sample problems are screened from each sample problem according to the similarity, wherein the matching sample problem can be a sample problem with a similarity greater than a certain threshold. It should be noted that one problem information can have multiple feature vectors, and when calculating the similarity, these feature vectors can be used for calculation respectively, and then the maximum similarity is taken as the similarity between the problem information and the sample problem.

[0069] In the above manner, by obtaining the similarity between the input problem and the sample problem, the matching sample problem of the input problem is obtained, and the preliminary retrieval of the input problem is realized, which is convenient for subsequent reordering.

[0070] Based on the first or second embodiment of the above natural language processing-based question and answer method, a third embodiment of the natural language processing-based question and answer method of the present application is proposed.

[0071] In the present embodiment, after the step S40, the method further comprises:

[0072] Step S50, when receiving the answer evaluation information fed back by the target user based on the target sample reply information, determining the answer effect corresponding to the target sample reply information according to the answer evaluation information, so as to distribute artificial customers or adjust the target sample reply information according to the answer effect.

[0073] In this embodiment, after outputting the feedback-related target sample reply information, the server can further acquire the answer evaluation information fed back by the user based on the target sample reply information, and determine the answer effect according to the answer evaluation information, so as to distribute artificial customers or adjust related algorithms according to the answer effect. For example, if the user's evaluation is not satisfied or the problem is not solved, the server can acquire the relevant contact information (such as account information, telephone number, etc.) of the user, and send the contact information of the user to the corresponding artificial customer terminal, so as to provide answers to the user through artificial customer service; in addition, the reply information corresponding to other matching sample problems can also be used as new target sample reply information and output to the user terminal for the user to acquire other reply information. Of course, the related algorithm for retrieval can also be adjusted accordingly.

[0074] In the above manner, the embodiment acquires user feedback after inputting target sample reply information, and further processes according to the feedback, which is beneficial to improve the accuracy of question and answer and improve the user experience.

[0075] In addition, the embodiment of the application also provides a question and answer device based on natural language processing.

[0076] Reference Figure 3 , Figure 3 FIG. 1 is a schematic diagram of functional modules of a first embodiment of the question and answer device based on natural language processing of the application.

[0077] In this embodiment, the question and answer device based on natural language processing comprises:

[0078] The instruction receiving module 10 is configured to acquire an input question in the question and answer instruction triggered by a target user operation when the question and answer instruction is received, and pre-process the input question to generate input question features;

[0079] The question matching module 20 is configured to perform sample question matching in a preset database according to the input question features, determine the matching sample questions corresponding to the input question, and generate a matching sample question set;

[0080] The question rearrangement module 30 is configured to acquire user behavior information of the target user, rearrange the matching sample questions in the matching sample question set according to the user behavior information, and determine the target sample question corresponding to the question and answer instruction in the matching sample question set according to the sorting result;

[0081] The information output module 40 is configured to acquire target sample reply information corresponding to the target sample question, and output the target sample reply information.

[0082] Further, the question matching module 20 comprises:

[0083] a similarity calculation unit, configured to calculate similarity between the input question and each sample question according to the input feature vector corresponding to the input question feature and the sample feature vector corresponding to each sample question in the database;

[0084] a first determination unit, configured to determine a matching sample question from the sample questions according to the similarity between the input question and each sample question, wherein the similarity between the matching sample question and the input question is greater than a preset threshold.

[0085] Further, the similarity calculation unit is specifically configured to sequentially take a sample question from the sample questions as a current sample question, and obtain a current sample feature vector of the current sample question; and substitute the input feature vector and the current sample feature vector into a preset similarity formula to calculate the similarity between the input question and the current sample question, the preset similarity formula being:

[0086] S =∑(w i *R(q i ))

[0087] S is the similarity between the input question and the current sample question; q i is the i-th element in the input feature vector Q; w i is the weight of q i , the database contains more sample feature vectors of q i , w i is smaller; R(q i ) is the relevance score of q i and the current sample feature vector, R(q i ) is determined according to the number of occurrences of q i in the input feature vector, the number of occurrences of q i in the current sample feature vector, the number of elements in the current sample feature vector, and the average number of elements in the sample feature vectors of all sample questions in the database.

[0088] Further, the similarity calculation unit is further configured to calculate w i according to a preset weight formula, the preset weight formula being:

[0089]

[0090] N is the number of sample questions in the database; n(q i ) is the number of sample feature vectors containing q i in the database.

[0091] Further, the similarity calculation unit is further configured to calculate w iThe number of occurrences in the input feature vector, q i The number of occurrences in the current sample feature vector, the number of elements in the current sample feature vector, the average number of elements in the sample feature vectors of all sample questions in the database, and the preset correlation formula are used to calculate R(q i ), the preset correlation formula is:

[0092]

[0093] in, k1, k2, and b are preset parameters, and all are greater than zero; F1(q i ) is q i The number of times it appears in the current sample feature vector; F2(q i ) is q i The number of times it appears in the input feature vector; dl is the number of elements in the current sample feature vector; avgdl is the average number of elements in the sample feature vectors of all sample questions in the database.

[0094] Furthermore, the question rearrangement module 30 includes:

[0095] a tag acquisition unit, configured to acquire the target user's historical browsing information and historical transaction information, and determine the target interest tag of the target user based on the historical browsing information and the historical transaction information;

[0096] a label matching unit, configured to match the target interest label with the sample attribute label of each matching sample question, determine the number of matching labels corresponding to each matching sample question, and reorder each matching sample question according to the number of matching labels corresponding to each matching sample question;

[0097] The second determining unit is used to determine the target sample question corresponding to the question-answering instruction among the matching sample questions according to the sorting result.

[0098] Furthermore, the question-answering device based on natural language processing further includes:

[0099] The effect determination module is used to determine the answer effect corresponding to the target sample reply information according to the answer evaluation information when receiving the answer evaluation information fed back by the target user based on the target sample reply information, so as to allocate artificial customers or adjust the target sample reply information according to the answer effect.

[0100] Among them, each module in the above-mentioned question-answering device based on natural language processing corresponds to each step in the above-mentioned question-answering method embodiment based on natural language processing, and their functions and implementation processes will not be repeated here one by one.

[0101] In addition, the embodiment of the present application further provides a computer readable storage medium.

[0102] The computer readable storage medium of the present application stores a computer program, wherein the computer program is executed by a processor to implement the steps of the question and answer method based on natural language processing.

[0103] The method implemented when the computer program is executed can refer to the embodiments of the question and answer method based on natural language processing of the present application, which will not be described here.

[0104] It should be noted that in this paper, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or system. Without more limitations, the element defined by the sentence "includes a" does not exclude the presence of other identical elements in the process, method, article or system including the element.

[0105] The above-mentioned embodiment numbers of the present application are only for description, not representing the advantages and disadvantages of the embodiments.

[0106] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and necessary general hardware platforms, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions to make a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the methods described in various embodiments of the present application.

[0107] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent flow transformation made by using the content of the present application specification and drawings, or directly or indirectly applied to other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A natural language processing-based question answering method, characterized by, The natural language processing-based question and answer method comprises the following steps: Upon receiving a target user operation triggered question and answer instruction, an input question in the question and answer instruction is acquired, and the input question is preprocessed to generate input question features; Sample question matching is performed in a preset database according to the input question features, a matching sample question corresponding to the input question is determined, and a matching sample question set is generated; User behavior information of the target user is acquired, the user behavior information is stored in a blockchain node, the matching sample questions in the matching sample question set are reordered according to the user behavior information, and a target sample question corresponding to the question and answer instruction is determined in the matching sample question set according to the sorting result, wherein the blockchain node is one node in an application mode integrating distributed data storage, peer-to-peer transmission, consensus mechanism and encryption algorithm; Target sample reply information corresponding to the target sample question is acquired, and the target sample reply information is outputted; The step of performing sample question matching in a preset database according to the input question features, determining a matching sample question corresponding to the input question, and generating a matching sample question set specifically comprises: According to the input feature vector corresponding to the input question features and the sample feature vector corresponding to each sample question in the database, the similarity of the input question and each sample question is calculated through a cosine similarity algorithm and a BM25 algorithm; According to the similarity of the input question and each sample question, a matching sample question is determined in each sample question, and a matching sample question set is generated based on the matching sample question, wherein the similarity of the matching sample question and the input question is greater than a preset threshold; The step of preprocessing the input question to generate input question features specifically comprises: Continuous character sequences in the text of the input question are combined into word sequences according to a preset specification through a forward maximum matching algorithm, a reverse maximum matching algorithm, feature scanning, flag segmentation and a statistical model-based method; Sample keywords are defined according to business conditions; Business feature keywords matching the sample keywords are identified and extracted in the word sequences through comparison between the word sequences and the sample keywords; The same meaning word set in which the business feature keywords are located is searched in a corpus database; Corresponding sentence vectors are acquired according to the business feature keywords and the same meaning word set, and the sentence vectors are taken as input feature vectors of the input question; Input question features are obtained based on the input feature vectors; The user behavior information comprises historical browsing information and historical transaction information, and the step of acquiring user behavior information of the target user, storing the user behavior information in a blockchain node, reordering the matching sample questions in the matching sample question set according to the user behavior information, and determining a target sample question corresponding to the question and answer instruction in the matching sample question set according to the sorting result specifically comprises: Obtaining historical browsing information and historical transaction information of the target user, storing the historical browsing information and the historical transaction information in a blockchain node, and determining a target interest label of the target user according to the historical browsing information and the historical transaction information; Matching the target interest label with sample attribute labels of each matching sample question respectively, determining a matching label number corresponding to each matching sample question, and reordering each matching sample question according to the matching label number corresponding to each matching sample question; Determining a target sample question corresponding to the question and answer instruction from the matching sample questions according to the ordering result; Before the step of reordering each matching sample question according to the matching label number corresponding to each matching sample question, the method further comprises: Mapping the matching label number corresponding to each matching sample question into a mapping value in the range of (0, 1) through a softmax function, and the cumulative sum of the mapping value is 1, the mapping value being a probability that the input question is equivalent to the matching sample question; Correspondingly, the step of reordering each matching sample question according to the matching label number corresponding to each matching sample question comprises: Ordering each matching sample question according to the mapping value; Correspondingly, the step of determining a target sample question corresponding to the question and answer instruction from the matching sample questions according to the ordering result comprises: Comparing the maximum mapping value in the ordering result with a preset threshold value; If the maximum mapping value is greater than the preset threshold value, the matching sample question corresponding to the maximum mapping value is determined as the target sample question corresponding to the question and answer instruction; If the maximum mapping value is less than the preset threshold value, at least one matching sample question at the front of the ordering is determined as the target sample question, or at least one matching sample question at the front of the ordering is returned to a user end, and the target sample question corresponding to the question and answer instruction is determined according to selection feedback returned by the user end.

2. The natural language processing based question answering method of claim 1, wherein, The step of calculating the similarity between the input question and each sample question according to the input feature vector corresponding to the input question feature and the sample feature vector corresponding to each sample question in the database comprises: Taking a sample question from each sample question in turn as a current sample question, and obtaining a current sample feature vector of the current sample question; Substituting the input feature vector and the current sample feature vector into a preset similarity formula to calculate the similarity between the input question and the current sample question, the preset similarity formula being: a similarity of the input question to a current sample question; is the i-th element in the input feature vector Q; For the weight of the database containing the more sample feature vectors, the smaller; For a correlation score with the current sample feature vector, According to the number of occurrences in the input feature vector, the number of occurrences in the current sample feature vector, the number of elements of the current sample feature vector, the average number of elements of the sample feature vectors of all sample problems of the database. 3.The natural language processing based question answering method of claim 2, wherein, Before the step of substituting the input feature vector and the current sample feature vector into the preset similarity formula to calculate the similarity between the input question and the current sample question, the method further comprises: According to a preset weight formula , the preset weight formula is: , N is the number of sample questions in the database; The database contains the number of sample feature vectors. 4.The natural language processing based question answering method of claim 2, wherein, Before the step of substituting the input feature vector and the current sample feature vector into the preset similarity formula to calculate the similarity between the input question and the current sample question, the method further comprises: According to the number of occurrences in the input feature vector, the number of occurrences in the current sample feature vector, the number of elements of the current sample feature vector, the average number of elements of the sample feature vectors of all sample problems in the database, and a preset correlation formula The preset correlation formula is: wherein, k1, k2, b are preset parameters and are all greater than zero; for occurrence in the current sample feature vector; for occurrences in the input feature vector; dl is the number of elements of the current sample feature vector; avgdl is the average number of elements of the sample feature vectors of all sample questions in the database.

5. The natural language processing based question answering method according to any one of claims 1 to 4, characterized in that, The step of acquiring target sample reply information corresponding to the target sample question, outputting and displaying the target sample reply information further comprises: When receiving the answer evaluation information fed back by the target user based on the target sample reply information, determining the answer effect corresponding to the target sample reply information according to the answer evaluation information, and distributing artificial customers or adjusting the target sample reply information according to the answer effect.

6. A natural language processing based question answering apparatus, characterized by, The natural language processing-based question and answer device comprises: An instruction receiving module is configured to acquire an input question in a question and answer instruction triggered by a target user operation and perform preprocessing on the input question to generate input question features when the question and answer instruction is received; A question matching module is configured to perform sample question matching in a preset database according to the input question features, determine a matching sample question corresponding to the input question, and generate a matching sample question set; A question rearrangement module is configured to acquire user behavior information of the target user, store the user behavior information in a blockchain node, and rearrange the matching sample questions in the matching sample question set according to the user behavior information, and determine a target sample question corresponding to the question and answer instruction in the matching sample question set according to a sorting result, wherein the blockchain node is one node in an application mode integrating distributed data storage, point-to-point transmission, consensus mechanism, and encryption algorithm; An information output module is configured to acquire target sample reply information corresponding to the target sample question and output the target sample reply information; The question matching module is further configured to calculate the similarity between the input question and each sample question by using a cosine similarity algorithm and a BM25 algorithm according to an input feature vector corresponding to the input question features and a sample feature vector corresponding to each sample question in the database; determine a matching sample question in each sample question according to the similarity between the input question and each sample question, and generate a matching sample question set based on the matching sample question, wherein the similarity between the matching sample question and the input question is greater than a preset threshold; The instruction receiving module is further configured to combine continuous character sequences in the text of the input question into word sequences according to a preset specification by using a forward maximum matching algorithm, a reverse maximum matching algorithm, feature scanning, flag segmentation, and a statistical model; define sample keywords according to business conditions; identify and extract business feature keywords matching the sample keywords in the word sequences by comparing the word sequences with the sample keywords; find a synonym set in which the business feature keywords are located in a corpus database; acquire a corresponding sentence vector according to the business feature keywords and the synonym set, and use the sentence vector as an input feature vector of the input question; and obtain input question features based on the input feature vector. The user behavior information includes historical browsing information and historical transaction information. The question rearrangement module is further configured to acquire the historical browsing information and the historical transaction information of the target user, store the historical browsing information and the historical transaction information in a blockchain node, and determine a target interest label of the target user according to the historical browsing information and the historical transaction information; match the target interest label with sample attribute labels of each matching sample question respectively, determine a matching label number corresponding to each matching sample question, and rearrange each matching sample question according to the matching label number corresponding to each matching sample question; and determine a target sample question corresponding to the question and answer instruction from the matching sample questions according to a sorting result. The question rearrangement module is further configured to map the matching label number corresponding to each matching sample question into a mapping value in a range of (0, 1) by using a softmax function, and a cumulative sum of the mapping values is 1, where the mapping value is a probability that the input question is equivalent to the matching sample question. The question rearrangement module is further configured to sort each matching sample question according to the mapping value. The question rearrangement module is further configured to compare a maximum mapping value in the sorting result with a preset threshold value; if the maximum mapping value is greater than the preset threshold value, the matching sample question corresponding to the maximum mapping value is determined as the target sample question corresponding to the question and answer instruction; if the maximum mapping value is less than the preset threshold value, at least one matching sample question at a front of the sorting is determined as the target sample question, or at least one matching sample question at the front of the sorting is returned to a user terminal, and a target sample question corresponding to the question and answer instruction is determined according to selection feedback returned by the user terminal.

7. A natural language processing based question answering device, characterized by, The natural language processing-based question and answer device includes a processor, a memory, and a computer program stored on the memory and executable by the processor, where the computer program is executed by the processor to implement steps of the natural language processing-based question and answer method according to any one of claims 1 to 5.

8. A storage medium, characterized by The storage medium stores a computer program, where the computer program is executed by a processor to implement steps of the natural language processing-based question and answer method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and device for recommending question and answer page related questions

    CN104462554A

  • Data processing method and device, terminal device, and computer storage medium

    CN109376298A

  • Short text similarity calculation method based on probability model

    CN109858028A