Text matching method, device, storage medium and computer equipment

The method improves FAQ-based question answering systems by classifying and optimizing similarity matching results based on text class, addressing the issue of non-business queries to enhance accuracy and convenience.

CN113886544BActive Publication Date: 2025-07-15VIPSHOP (GUANGZHOU) SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111152505.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-07-15
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

In the existing question and answer system, when text matching is performed through similarity matching and threshold, invalid text from non-business classes will be matched with text in the FAQ knowledge base, resulting in low convenience and accuracy, and the inability to effectively solve user business problems.

Method used

By obtaining the sentence vector and text category of the target text, combining the sentence vectors of the text to be matched in the FAQ knowledge base, similarity matching is performed, and the similarity matching results are optimized based on the text category to filter out the target matching text.

Benefits of technology

It improves the accuracy and convenience of the Q&A system, can more accurately match business-related texts, and reduces the impact of non-business texts on matching results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113886544B_ABST
    Figure CN113886544B_ABST
Patent Text Reader

Abstract

The text matching method, device, storage medium, and computer device provided by the present invention, when performing similarity matching between a target text and each of the texts to be matched in the text set to be matched, first obtain a first sentence vector and a text category corresponding to the target text, then obtain second sentence vectors corresponding to each of the texts to be matched in the text set to be matched, determine the similarity matching result of each text to be matched according to the first sentence vector and the second sentence vector, and then, for the similarity matching result of each text to be matched, it can be optimized by the text category of the target text. For example, the optimization methods for the similarity matching results of business texts and non-business texts can be different. Therefore, adopting the solution of the present application can support reducing the impact of non-business texts on the final matching result, so as to more accurately and conveniently help users solve business problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a text matching method, apparatus, storage medium, and computer device. Background Art

[0002] Text matching is a common application scenario in the field of natural language processing. A large number of NLP (Neuro-Linguistic Programming) tasks are based on text matching, such as information retrieval, machine translation, question answering systems, etc.

[0003] Most existing question answering systems based on FAQ knowledge bases use text similarity matching methods. By matching the text input by the user with each similar text in the FAQ knowledge base and giving a similarity matching score, and then re-ranking and thresholding the similarity matching scores of each similar text to output the final matching result.

[0004] However, only matching the text input by the user by means of similarity matching scores and thresholding will match some non-business invalid texts, such as "What do you mean?", "What's going on?", "Really?", "Trouble", etc., with the similar texts in the FAQ knowledge base. The answers determined by the text matching results cannot solve the problems that users encounter in business, resulting in low convenience and accuracy of the question answering system. Summary of the Invention

[0005] The purpose of the present invention is to at least solve one of the above technical defects, especially the technical defect of low convenience and accuracy of the question answering system in the prior art.

[0006] The present invention provides a text matching method, and the method includes:

[0007] Obtain a target text and a set of texts to be matched corresponding to the target text;

[0008] Determine a first sentence vector and a text category corresponding to the target text, and second sentence vectors corresponding to each text to be matched in the set of texts to be matched;

[0009] Perform similarity matching between the first sentence vector corresponding to the target text and the second sentence vector corresponding to each text to be matched to obtain a similarity matching result for each text to be matched;

[0010] Optimize the similarity matching results of each text to be matched based on the text category of the target text, and determine a target matching text in the set of texts to be matched based on the optimized similarity matching results.

[0011] Optionally, the step of obtaining the set of texts to be matched corresponding to the target text includes:

[0012] Tokenize the target text to obtain at least one phrase;

[0013] Retrieve the phrase in the FAQ knowledge base to obtain multiple texts to be matched corresponding to the phrase, and form a set of texts to be matched; wherein, an index structure corresponding to multiple texts to be matched is pre-established in the FAQ knowledge base.

[0014] Optionally, the step of determining the first sentence vector and text category corresponding to the target text includes:

[0015] Input the target text into a text classification model to obtain the first sentence vector and text category corresponding to the target text output by the text classification model;

[0016] Wherein, the text classification model is trained with multiple texts to be matched corresponding to different text categories in the FAQ knowledge base as training samples and the text category corresponding to each text to be matched as sample labels.

[0017] Optionally, the step of determining the second sentence vector corresponding to each text to be matched in the set of texts to be matched includes:

[0018] Search in the cache for the second sentence vector corresponding to each text to be matched in the set of texts to be matched;

[0019] Wherein, all texts to be matched in the FAQ knowledge base and the second sentence vector corresponding to each text to be matched obtained through the text classification model are pre-stored in the cache.

[0020] Optionally, the step of optimizing the similarity matching results of each text to be matched based on the text category of the target text includes:

[0021] Determine the corresponding adjustment coefficient according to the text category of the target text;

[0022] Use the adjustment coefficient to optimize the similarity matching results of each text to be matched.

[0023] Optionally, the text category of the target text includes business texts and non-business texts;

[0024] When the target text is a non-business text, the adjustment coefficient of the target text is less than the adjustment coefficient of the business text.

[0025] Optionally, the step of determining the target matching text in the text set to be matched based on the optimized similarity matching result includes:

[0026] Sort the optimized similarity matching results to obtain a sorting result;

[0027] Filter the texts to be matched in the sorting result according to a preset selection number and a preset similarity threshold;

[0028] Use the filtered texts to be matched as the target matching texts of the target text.

[0029] The present invention also provides a text matching device, including:

[0030] A text acquisition module, configured to acquire a target text and a text set to be matched corresponding to the target text;

[0031] A text processing module, configured to determine a first sentence vector and a text category corresponding to the target text, and a second sentence vector corresponding to each text to be matched in the text set to be matched;

[0032] A similarity matching module, configured to perform similarity matching between the first sentence vector corresponding to the target text and the second sentence vector corresponding to each text to be matched, to obtain a similarity matching result for each text to be matched;

[0033] A text matching module, configured to optimize the similarity matching results of each text to be matched based on the text category of the target text, and determine the target matching text in the text set to be matched based on the optimized similarity matching result.

[0034] The present invention also provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the text matching method as described in any one of the above embodiments.

[0035] The present invention also provides a computer device, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the text matching method as described in any one of the above embodiments.

[0036] As can be seen from the above technical solutions, the embodiments of the present invention have the following advantages:

[0037] The text matching method, device, storage medium, and computer device provided by the present invention, when performing similarity matching between a target text and each of the texts to be matched in a text set to be matched, first obtain a first sentence vector and a text category corresponding to the target text, then obtain second sentence vectors corresponding to each of the texts to be matched in the text set to be matched, determine the similarity matching result of each text to be matched according to the first sentence vector and the second sentence vector. Then, for the similarity matching result of each text to be matched, it can be optimized by the text category of the target text, so that the optimized similarity matching result not only considers the similarity between the target text and the text to be matched, but also considers the text category of the target text. Among them, optimizing the similarity matching result, for example, the optimization methods for the similarity matching results of business texts and non-business texts can be different. Therefore, adopting the solution of the present application can support reducing the impact of non-business texts on the final matching result, thereby being able to more accurately and conveniently help users solve business problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 It is a flowchart of a text matching method provided by an embodiment of the present invention;

[0040] Figure 2 It is a structural schematic diagram of the input and output of the BERT model provided by an embodiment of the present invention;

[0041] Figure 3 It is a flowchart of an online prediction process for integrating classification and similarity matching provided by an embodiment of the present invention;

[0042] Figure 4 It is a structural schematic diagram of a text matching device provided by an embodiment of the present invention;

[0043] Figure 5 It is an internal structural schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] Most existing question-and-answer systems based on FAQ knowledge bases use the method of text similarity matching. By matching the text input by the user with each similar text in the FAQ knowledge base and giving a similarity matching score, and then reordering and thresholding the similarity matching scores of each similar text to output the final matching result. Among them, the FAQ knowledge base refers to a pre-edited database for storing pairs composed of business knowledge questions and answers.

[0046] However, only matching the text input by the user through the similarity matching score and thresholding will match some non-business invalid texts, such as "What do you mean?", "What's going on?", "Really?", "trouble", etc., with the similar texts in the FAQ knowledge base. The answers determined by this text matching result cannot solve the problems that users encounter in business, resulting in low convenience and accuracy of the question-and-answer system.

[0047] Therefore, the purpose of the present invention is to solve the technical problems of low convenience and accuracy of the question-and-answer system in the prior art, and the following technical solutions are proposed:

[0048] In one embodiment, as Figure 1 shown, Figure 1 is a schematic flowchart of a text matching method provided by an embodiment of the present invention; the present invention provides a text matching method, which specifically includes the following:

[0049] S110: Obtain a target text and a set of texts to be matched corresponding to the target text.

[0050] In this step, when a user needs to answer some business knowledge questions through the question-and-answer system, the question to be answered can be sent to the interaction interface of the question-and-answer system in the form of a target text. When the question-and-answer system receives the target text input by the user, it can perform text matching based on this target text to obtain an answer corresponding to this target text.

[0051] It can be understood that the question-and-answer system in this application is a question-and-answer system based on an FAQ knowledge base. For example, in the business scenario of an intelligent customer service in the question-and-answer system, this business scenario can significantly reduce the number and cost of human customer services.

[0052] For example, for the online intelligent customer service of 10086, when a user asks a question "How to query the phone bill", the question and answer system can automatically give a corresponding reply "Please send an SMS of 'HF' to the 10086 number to query the current phone bill", and there is no need to spend high-cost manpower to provide answers.

[0053] Specifically, in this application, when the question and answer system obtains the target text input by the user, it can find the corresponding text set to be matched according to the target text. It can be understood that the text set to be matched here refers to the similar questions and standard questions that are pre-edited in the FAQ knowledge base corresponding to the question and answer system and are related to the business of the question and answer system; among them, the standard questions refer to multiple different types of questions pre-edited in the FAQ knowledge base according to the business categories of the question and answer system. One standard question corresponds to multiple similar questions similar to the standard question. The standard questions and similar questions together form the text set in the FAQ knowledge base.

[0054] Therefore, when the question system obtains the target text input by the user, it can find the similar questions and standard questions corresponding to the target text in the FAQ knowledge base, so as to form a text set to be matched.

[0055] Furthermore, when finding the similar questions and standard questions corresponding to the target text in the FAQ knowledge base, the target text can be segmented to obtain the corresponding keywords of the target text, and then the keywords can be used to retrieve in the FAQ knowledge base to obtain the similar questions and the standard questions corresponding to the similar questions that match the keywords, so as to form a text set to be matched.

[0056] S120: Determine the first sentence vector and text category corresponding to the target text, and the second sentence vectors corresponding to each text to be matched in the text set to be matched.

[0057] In this step, after obtaining the target text and the text set to be matched corresponding to the target text through step S110, each text to be matched in the text set to be matched can be further subjected to similarity matching with the target text, so as to determine the target matching text in the text set to be matched according to the similarity matching result.

[0058] When performing similarity matching between each text to be matched in the text set to be matched and the target text, the text features of the target text can be vectorized to obtain the first sentence vector corresponding to the target text, and then the text features of each text to be matched can be vectorized to obtain the second sentence vector corresponding to each text to be matched. Finally, the first sentence vector corresponding to the target text is respectively subjected to similarity matching with the second sentence vectors corresponding to each text to be matched, so as to accurately obtain the similarity matching result of each text to be matched.

[0059] Specifically, when vectorizing the text features of the target text to determine the first sentence vector corresponding to the target text, the target text can be tokenized to obtain the vector representation of each word, and then the vectors of all words are superimposed into a new vector as the sentence vector representation of the target text; or an encoder-decoder model (encoder-decoder model) can be used, and the sentences of the context are predicted through the central sentence, and the vector obtained by the encoder for the sentence is used as the sentence vector representation; an RNN (recurrent neural network), CNN (convolutional neural network), attention mechanism or a more complex model can also be used, and multi-task learning is performed based on the labeled corpus of common tasks in natural language processing (such as named entity recognition, sentence similarity determination, etc.), and the output of the shared layer is used as the sentence vector representation.

[0060] Furthermore, when determining the second sentence vectors corresponding to the respective texts to be matched in the text set to be matched, the determination method of the first sentence vector can be used for determination. Moreover, in order to improve the matching efficiency of the similarity matching, the present application can also pre-vectorize all similar questions in the FAQ knowledge base, and then store each similar question and its corresponding sentence vector. When subsequently determining the sentence vector corresponding to the text to be matched in the text set to be matched, the sentence vector corresponding to the text to be matched is directly searched in the storage area, without repeatedly vectorizing the similar questions in the FAQ knowledge base, so as to improve the matching efficiency of the similarity matching.

[0061] Even further, when determining the first sentence vector corresponding to the target text, the text category of the target text can also be determined. The process of determining the text category of the target text can be carried out simultaneously with the process of determining the first sentence vector, or can be carried out separately. If carried out simultaneously, a relevant classification model can be selected. This classification model has the ability to obtain the sentence vector corresponding to the input text and the text category of the input text during the forward calculation process, so that when the present application uses this classification model, the first sentence vector and the text category corresponding to the target text can be obtained.

[0062] In addition, the text category of the target text in the present application corresponds to the business category of the standard questions in the FAQ knowledge base. For example, the business category of the standard questions can include business-related texts and non-business-related texts, and the business-related texts can further include texts of different types of businesses.

[0063] S130: Perform similarity matching between the first sentence vector corresponding to the target text and the second sentence vector corresponding to each text to be matched, and obtain the similarity matching result of each text to be matched.

[0064] In this step, after determining the first sentence vector and text category corresponding to the target text, as well as the second sentence vectors corresponding to each text to be matched in the text set to be matched through step S120, the first sentence vector corresponding to the target text can be respectively matched with the second sentence vector corresponding to each text to be matched to obtain the similarity matching result of each text to be matched.

[0065] It can be understood that when performing similarity matching on two sentence vectors, it is mainly to calculate the distance between the two sentence vectors. The greater the distance, the greater the similarity. The calculation methods of similarity matching include but are not limited to calculating using the Pearson correlation coefficient, calculating using the Euclidean distance, calculating using the Cosine similarity, calculating using the Manhattan distance, etc., which are not limited here.

[0066] S140: Optimize the similarity matching results of each text to be matched based on the text category of the target text, and determine the target matching text in the text set to be matched based on the optimized similarity matching results.

[0067] In this step, after obtaining the similarity matching results of each text to be matched through step S130, then the similarity matching results of each text to be matched can be optimized according to the text category of the target text, and the target matching text in the text set to be matched can be determined based on the optimized similarity matching results.

[0068] Specifically, when optimizing the similarity matching results of each text to be matched according to the text category of the target text, the pre-set optimization rules can be based on. The optimization rules can be to uniformly lower the similarity matching scores of text categories that are non-business texts, uniformly increase the similarity matching scores of text categories that are business texts, or increase the similarity matching scores of business texts in different categories to different degrees.

[0069] Furthermore, when determining the target matching text in the text set to be matched based on the optimized similarity matching results, the optimized similarity matching results can be sorted, and then the text to be matched corresponding to the similarity matching result with a higher ranking can be selected and output as the target matching text. When selecting the similarity matching result with a higher ranking, it can also be selected according to the similarity threshold to obtain the final target matching text.

[0070] It can be understood that after optimizing the similarity matching results of each text to be matched in this application, when the text category of the target text is a non-business text, the score value of the optimized similarity matching result is lower than that of the similarity matching result before optimization. At this time, if the target matching text is selected according to the traditional selection method, since the score value of the optimized similarity matching result is low, there are fewer target matching texts that exceed the similarity threshold that can be selected, thereby reducing the impact of non-business texts on the final matching result.

[0071] In the above embodiment, when performing similarity matching between the target text and each text to be matched in the text set to be matched, first, obtain the first sentence vector and text category corresponding to the target text, then obtain the second sentence vectors corresponding to each text to be matched in the text set to be matched, determine the similarity matching result of each text to be matched according to the first sentence vector and the second sentence vector. Then, for the similarity matching result of each text to be matched, it can be optimized by the text category of the target text, so that the optimized similarity matching result not only considers the similarity between the target text and the text to be matched, but also considers the text category of the target text; among them, the optimization of the similarity matching result, for example, the optimization methods for the similarity matching results of business texts and non-business texts can be different. Therefore, adopting the solution of this application can support reducing the impact of non-business texts on the final matching result, thereby being able to more accurately and conveniently help users solve business problems.

[0072] The above embodiment describes the text matching method in this application in detail. Next, the process of how to obtain the text set to be matched corresponding to the target text in this application will be described.

[0073] In one embodiment, the step of obtaining the text set to be matched corresponding to the target text in step S110 may include:

[0074] S111: Segment the target text to obtain at least one phrase.

[0075] S112: Retrieve the phrase in the FAQ knowledge base to obtain multiple texts to be matched corresponding to the phrase, and form a text set to be matched; wherein, an index structure corresponding to multiple texts to be matched is pre-established in the FAQ knowledge base.

[0076] In this embodiment, when determining the text set to be matched corresponding to the target text, similar questions and standard questions corresponding to the target text can be found in the FAQ knowledge base, so as to form a text set to be matched.

[0077] Specifically, when searching for similar questions and standard questions corresponding to the target text in the FAQ knowledge base, the target text can be segmented to obtain at least one phrase corresponding to the target text. Then, the segmented phrases can be used to retrieve in the FAQ knowledge base to obtain similar questions that match the phrases and the standard questions corresponding to the similar questions, thereby forming a text set to be matched.

[0078] It should be noted that in order to quickly obtain the text set to be matched corresponding to the target text in this application, an index structure corresponding to similar questions and standard questions can be established in the FAQ knowledge base in advance. This index structure can be constructed through the ElasticSearch retrieval tool. Elastic Search uses the inverted index technology of Lucene to achieve faster filtering than relational databases. Therefore, when using the Elastic Search retrieval tool, the similar questions and standard questions in the FAQ knowledge base can be segmented first, and then the Elastic Search retrieval tool can be used to establish an index for the segmented FAQ knowledge base to obtain the corresponding index structure.

[0079] The above embodiments illustrate the process of how to obtain the text set to be matched corresponding to the target text in this application. Next, the process of how to determine the first sentence vector and text category corresponding to the target text in this application will be described.

[0080] In one embodiment, the steps of determining the first sentence vector and text category corresponding to the target text in step S120 may include:

[0081] S121: Input the target text into the text classification model to obtain the first sentence vector and text category output by the text classification model corresponding to the target text.

[0082] In this embodiment, when determining the first sentence vector and text category corresponding to the target text, the target text can be input into a pre-configured text classification model, so as to predict the text category of the target text through this text classification model, and during the process of predicting the text category of the target text, a sentence vector corresponding to the target text is output.

[0083] Among them, the text classification model of this application can be trained with multiple texts to be matched corresponding to different text categories in the FAQ knowledge base as training samples and the text category corresponding to each text to be matched as sample labels.

[0084] Furthermore, the text classification model in this application can be the BERT model, or the Ernie model, or it can also be the TextCNN model, which is not limited here.

[0085] This application can preferably use the BERT model to predict the target text. The BERT model is one of the popular research areas in the field of natural language processing (NLP) in recent years. The training of the BERT model is mainly divided into two stages. In the pre-trained stage, the model parameters are optimized based on a large amount of data to learn general language representations. In the fine-tuned stage, the model parameters are re-fine-tuned based on specific downstream tasks to improve the accuracy of specific NLP tasks.

[0086] Schematically, as Figure 2 shown, Figure 2 is a schematic structural diagram of the input and output of the BERT model provided by an embodiment of the present invention; when the BERT model is used for pre-training in this application, multiple similar questions corresponding to standard questions of different text categories in the FAQ knowledge base can be used as training samples, and the text category corresponding to each similar question can be used as a sample label for training. When the trained BERT model is used, a user question can be input. After the user question is predicted and classified by the BERT model, the corresponding category and sentence vector can be obtained.

[0087] Specifically, the FAQ knowledge base contains N standard questions, all of which are related to the business. Each standard question contains multiple similar questions similar to this standard question. The present invention defines the N standard questions as N categories, and the samples of each category are all the similar questions under this category. In addition, in addition to the N categories related to the business, the present invention also adds another category, named the "non-business" category. Corresponding corpora such as chatting, skills, and invalid questions can be placed in the "non-business" category to construct a text classification data set of N + 1 categories; then use the pre-trained BERT model to fine-tune the N + 1 category text classification data set to train the text classification model.

[0088] The above embodiments illustrate the process of how to determine the first sentence vector and text category corresponding to the target text in this application. Next, the process of how to determine the second sentence vector corresponding to each matching text in the text set to be matched in this application will be described.

[0089] In one embodiment, the step of determining the second sentence vector corresponding to each matching text in the text set to be matched in step S120 may include:

[0090] S122: Search for the second sentence vector corresponding to each matching text in the text set to be matched in the cache respectively. Among them, all the matching texts in the FAQ knowledge base and the second sentence vector corresponding to each matching text obtained through the text classification model are pre-stored in the cache.

[0091] In this application, to improve the matching efficiency of similarity matching, all similar questions in the FAQ knowledge base can be vectorized in advance, and then each similar question and its corresponding sentence vector are stored. When determining the sentence vector corresponding to the text to be matched in the text set to be matched later, directly search for the sentence vector corresponding to the text to be matched in the storage area, without repeatedly vectorizing the similar questions in the FAQ knowledge base, so as to improve the matching efficiency of similarity matching.

[0092] Specifically, in this application, all similar questions in the FAQ knowledge base can be respectively input into the trained BERT model for forward calculation, and the corresponding sentence vectors are obtained from the output layer of the BERT model, and then can be stored in the cache in the form of key-value, where the key is the similar question and the value is the corresponding sentence vector.

[0093] When respectively searching for the second sentence vectors corresponding to each text to be matched in the text set to be matched in the cache, the text to be matched can be input into the corresponding search bar in the cache, and then search for the second sentence vector corresponding to the text to be matched in the cache.

[0094] The above embodiments describe the process of how to determine the second sentence vector corresponding to each text to be matched in the text set to be matched in this application. Next, the process of optimizing the similarity matching result in this application will be described.

[0095] In one embodiment, the step of optimizing the similarity matching results of each text to be matched based on the text category of the target text in step S140 may include:

[0096] S141: Determine the corresponding adjustment coefficient according to the text category of the target text.

[0097] S142: Optimize the similarity matching results of each text to be matched by using the adjustment coefficient.

[0098] In this embodiment, when optimizing the similarity matching results of each text to be matched according to the text category of the target text, the corresponding adjustment coefficient can be determined according to the text category of the target text, and then the similarity matching results of each text to be matched are optimized by using the adjustment coefficient.

[0099] For example, when the text category of the target text is a non-business text, the corresponding adjustment coefficient can be a coefficient less than 1, and then the similarity matching result corresponding to the text to be matched is multiplied by a coefficient less than 1 to suppress the current similarity matching result and reduce the situation of accidentally touching business knowledge points.

[0100] For business - type texts, the corresponding adjustment coefficient can be set to 1. After multiplying it with the similarity matching result, the obtained similarity matching result remains the original one, thus ensuring that business - type texts can match the exact target matching text.

[0101] In one embodiment, the text categories of the target text may include business - type texts and non - business - type texts; when the target text is a non - business - type text, the adjustment coefficient of the target text is less than that of the business - type text.

[0102] The above two embodiments illustrate the process of optimizing the similarity matching result in this application. Next, the process of determining the target matching text in this application will be described.

[0103] In one embodiment, the step of determining the target matching text in the text set to be matched based on the optimized similarity matching result in step S140 may include:

[0104] A11: Sort the optimized similarity matching result to obtain a sorting result.

[0105] A12: Screen the texts to be matched in the sorting result according to the preset selection number and the preset similarity threshold.

[0106] A13: Use the screened texts to be matched as the target matching text of the target text.

[0107] In this embodiment, when determining the target matching text in the text set to be matched based on the optimized similarity matching result, the optimized similarity matching result can be sorted first to obtain a sorting result, then the texts to be matched in the sorting result are screened according to the preset selection number and the preset similarity threshold, and finally the screened texts to be matched are used as the target matching text of the target text.

[0108] For example, when the similarity matching result is obtained, the similar questions, standard questions, and similarity scores in the FAQ knowledge base corresponding to the target text can be determined. Then, according to the preset selection number and the preset similarity threshold, the similar questions whose similarity scores are greater than the preset similarity threshold and whose rankings are within the preset selection number, as well as the standard questions corresponding to the similar questions, can be selected as the final target matching texts.

[0109] To better explain the text matching method of the present invention, the following will be used Figure 3 to further illustrate. Schematically, as Figure 3 shown, Figure 3 is a schematic diagram of the online prediction process integrating classification and similarity matching provided by the embodiment of the present invention.

[0110] Figure 3 After obtaining the user's question sentence, first segment the user's question sentence, and then quickly retrieve the top 50 pairs of (similar questions, standard questions) with related scores in Elastic Search. This step is also called recall. Then, use the similar questions in the top 50 pairs retrieved as key values and input them into the cache module respectively to obtain the sentence vectors corresponding to 50 similar questions. At the same time, input the user's question sentence into the trained BERT multi-classification model for forward calculation to obtain the corresponding sentence vector and the predicted category of the user's question sentence. Then, judge whether the predicted category is a "non-business" category. If so, set the adjustment coefficient to T = 0.7. If not, set the adjustment coefficient to T = 1.0. Next, calculate the similarity between the sentence vectors of 50 similar questions and the sentence vector of the user's question sentence one by one to obtain the similarity scores. Multiply the similarity scores by the discount coefficient T to obtain new similarity scores. According to the new similarity scores, sort the corresponding 50 pairs of (similar questions, standard questions, similarity scores) and set the threshold Threshold = 0.7. Take the similar questions, standard questions, and similarity scores whose similarity scores are greater than Threshold and rank in the top 5 as the final output of the FAQ text matching.

[0111] The text matching device provided by the embodiments of the present application will be described below. The text matching device described below can be correspondingly referred to the text matching method described above.

[0112] In one embodiment, as Figure 4 shown, Figure 4 is a schematic structural diagram of a text matching device provided by an embodiment of the present invention; the present invention also provides a text matching device, including a text acquisition module 210, a text processing module 220, a similarity matching module 230, and a text matching module 240, which specifically include the following:

[0113] The text acquisition module 210 is used to acquire a target text and a text set to be matched corresponding to the target text.

[0114] The text processing module 220 is used to determine a first sentence vector and a text category corresponding to the target text, and second sentence vectors corresponding to each text to be matched in the text set to be matched.

[0115] The similarity matching module 230 is used to perform similarity matching between the first sentence vector corresponding to the target text and the second sentence vector corresponding to each text to be matched to obtain a similarity matching result for each text to be matched.

[0116] The text matching module 240 is used to optimize the similarity matching results of each text to be matched based on the text category of the target text, and determine the target matching text in the text set to be matched based on the optimized similarity matching results.

[0117] In the above embodiment, when performing similarity matching between the target text and each text to be matched in the text set to be matched, first obtain the first sentence vector and text category corresponding to the target text, then obtain the second sentence vectors corresponding to each text to be matched in the text set to be matched, determine the similarity matching result of each text to be matched according to the first sentence vector and the second sentence vector. Then, for the similarity matching result of each text to be matched, it can be optimized by the text category of the target text, so that the optimized similarity matching result not only considers the similarity between the target text and the text to be matched, but also considers the text category of the target text; among them, the optimization of the similarity matching result, for example, the optimization methods for the similarity matching results of business texts and non-business texts can be different. Therefore, adopting the solution of this application can support reducing the influence of non-business texts on the final matching result, so as to more accurately and conveniently help users solve business problems.

[0118] In one embodiment, the text acquisition module 210 may include:

[0119] The word segmentation module is used to segment the target text to obtain at least one phrase.

[0120] The retrieval module is used to retrieve the phrase in the FAQ knowledge base to obtain multiple texts to be matched corresponding to the phrase, and form a text set to be matched; wherein, an index structure corresponding to multiple texts to be matched is pre-established in the FAQ knowledge base.

[0121] In one embodiment, the text processing module 220 may include:

[0122] The text classification module is used to input the target text into the text classification model to obtain the first sentence vector and text category corresponding to the target text output by the text classification model.

[0123] Among them, the text classification model is trained with multiple texts to be matched corresponding to different text categories in the FAQ knowledge base as training samples and the text category corresponding to each text to be matched as sample labels.

[0124] In one embodiment, the text processing module 220 may include:

[0125] The sentence vector matching module is used to separately find the second sentence vectors corresponding to each text to be matched in the text set to be matched in the cache.

[0126] Among them, all the text to be matched in the FAQ knowledge base, and the second sentence vectors corresponding to each text to be matched obtained by the text classification model are pre-stored in the cache.

[0127] In one embodiment, the text matching module 240 may include:

[0128] A coefficient determination module, configured to determine a corresponding adjustment coefficient according to the text category of the target text.

[0129] An optimization module, configured to optimize the similarity matching results of each text to be matched by using the adjustment coefficient.

[0130] In one embodiment, the text category of the target text may include business text and non-business text; when the target text is non-business text, the adjustment coefficient of the target text is less than the adjustment coefficient of the business text.

[0131] In one embodiment, the text matching module 240 may include:

[0132] A sorting module, configured to sort the optimized similarity matching results to obtain a sorting result;

[0133] A screening module, configured to screen the text to be matched in the sorting result according to a preset selection number and a preset similarity threshold.

[0134] A target confirmation module, configured to use the screened text to be matched as the target matching text of the target text.

[0135] In one embodiment, the present invention further provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the text matching method according to any one of the above embodiments.

[0136] In one embodiment, the present invention further provides a computer device, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the text matching method according to any one of the above embodiments.

[0137] Schematically, as Figure 5 shown, Figure 5 is a schematic internal structure diagram of a computer device provided by an embodiment of the present invention. The computer device 300 may be provided as a server. Referring to Figure 5, the computer device 300 includes a processing component 302, which further includes one or more processors, and memory resources represented by a memory 301 for storing instructions executable by the processing component 302, such as application programs. The application programs stored in the memory 301 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 302 is configured to execute instructions to perform the text matching method of any of the above embodiments.

[0138] The computer device 300 may further include a power component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 may operate based on an operating system stored in the memory 301, such as Windows Server TM, Mac OS XTM, Unix TM, Linux TM, Free BSDTM, or the like.

[0139] Those skilled in the art can understand that Figure 5 the structure shown in

[0140] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.

[0141] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The embodiments can be combined according to needs, and the same or similar parts can be referred to each other.

[0142] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A text matching method, characterized in that, The method includes: Obtaining a target text and a set of texts to be matched corresponding to the target text; Determining a first sentence vector and a text category corresponding to the target text, and second sentence vectors corresponding to each text to be matched in the set of texts to be matched; Performing similarity matching between the first sentence vector corresponding to the target text and the second sentence vector corresponding to each text to be matched to obtain a similarity matching result for each text to be matched; Optimizing the similarity matching results of each text to be matched based on the text category of the target text, and determining a target matching text in the set of texts to be matched based on the optimized similarity matching results; The step of optimizing the similarity matching results of each text to be matched based on the text category of the target text includes: Determining a corresponding adjustment coefficient according to the text category of the target text; wherein, the text category of the target text includes business texts and non-business texts; when the target text is a non-business text, the adjustment coefficient of the target text is less than the adjustment coefficient of the business text; Optimizing the similarity matching results of each text to be matched by using the adjustment coefficient.

2. The text matching method according to claim 1, wherein The step of obtaining the set of texts to be matched corresponding to the target text includes: Performing word segmentation on the target text to obtain at least one phrase; Retrieving the phrase in a FAQ knowledge base to obtain a plurality of texts to be matched corresponding to the phrase, and forming a set of texts to be matched; wherein, an index structure corresponding to a plurality of texts to be matched is pre-established in the FAQ knowledge base.

3. The text matching method according to claim 1, wherein The step of determining the first sentence vector and the text category corresponding to the target text includes: Inputting the target text into a text classification model to obtain a first sentence vector and a text category corresponding to the target text output by the text classification model; Wherein, the text classification model is trained by using a plurality of texts to be matched corresponding to different text categories in the FAQ knowledge base as training samples and the text category corresponding to each text to be matched as sample labels.

4. The text matching method according to claim 3, wherein The step of determining the second sentence vectors corresponding to each text to be matched in the set of texts to be matched includes: Respectively searching in a cache for the second sentence vectors corresponding to each text to be matched in the set of texts to be matched; Wherein, all texts to be matched in the FAQ knowledge base and the second sentence vectors corresponding to each text to be matched obtained through the text classification model are pre-stored in the cache.

5. The text matching method according to claim 1, wherein The step of determining the target matching text in the set of texts to be matched based on the optimized similarity matching results includes: Sorting the optimized similarity matching results to obtain a sorting result; Screening the texts to be matched in the sorting result according to a preset selection number and a preset similarity threshold; Taking the screened texts to be matched as the target matching text of the target text.

6. A text matching device, characterized in that, Includes: A text acquisition module for obtaining a target text and a set of texts to be matched corresponding to the target text; A text processing module, configured to determine a first sentence vector and a text category corresponding to the target text, as well as second sentence vectors corresponding to each of the texts to be matched in the text set to be matched; A similarity matching module, configured to perform similarity matching between the first sentence vector corresponding to the target text and the second sentence vector corresponding to each text to be matched, to obtain a similarity matching result for each text to be matched; A text matching module, configured to optimize the similarity matching results of each text to be matched based on the text category of the target text, and determine a target matching text in the text set to be matched based on the optimized similarity matching results; The step of optimizing the similarity matching results of each text to be matched based on the text category of the target text in the text matching module includes: Determining a corresponding adjustment coefficient according to the text category of the target text; wherein, the text category of the target text includes business text and non-business text; when the target text is non-business text, the adjustment coefficient of the target text is less than the adjustment coefficient of the business text; Optimizing the similarity matching results of each text to be matched by using the adjustment coefficient.

7. A storage medium, characterized in that: The computer-readable instructions are stored in the storage medium, and when the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the text matching method according to any one of claims 1 to 5.

8. A computer device, characterized in that: The computer-readable instructions are stored in the computer device, and when the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the text matching method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and system for carrying out natural language processing NLP on contract text data

    CN111753541A

  • Text matching method and device, computer equipment and readable storage medium

    CN113204629A