A method for intent recognition in multi-turn conversations during question and answer

Through semantic analysis and similarity calculation of user input text, combined with LLM model and clustering algorithm, the problem of user intention changes in multiple rounds of conversations is solved, efficient intention recognition and personalized reply are achieved, and the quality and efficiency of conversations are improved.

CN119441464BActive Publication Date: 2025-07-04ZHEJIANG FULIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510033277.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-07-04
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

During multiple rounds of conversations, user intentions change at any time, resulting in the diversity and long-tail characteristics of conversational search scenarios that increase the difficulty of knowledge retrieval, and it is difficult for the prior art to effectively identify user intentions.

Method used

By obtaining the user input text and performing semantic analysis, calculating similarity and text rewriting, using the preset LLM model for intent recognition and reply, establishing a database to store intent tags, using regular expressions and stop word lists for text filtering, and combining K-means clustering and Euclidean distance calculation to determine the synthetic text.

Benefits of technology

Improve the quality and efficiency of multiple rounds of conversations, ensure the contextual consistency and personalized response of the conversation, reduce the difficulty of knowledge retrieval, and avoid duplicate or irrelevant answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441464B_ABST
    Figure CN119441464B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for intent recognition in multi-turn conversations in question answering, belonging to the technical field of natural language understanding. Specifically, it includes: obtaining the first input text y1 of the current user, determining the intent tags corresponding to each word in the input text, and determining the relevant field corresponding to the input text y1 according to the intent tags; obtaining the second input text of the current user, calculating the similarity between the input text y2 and the input text y1. If the similarity is greater than or equal to a preset similarity threshold, rewrite the input text y2 to determine the synthesized text Y1 and reply to the synthesized text Y1; calculate the similarity between the input text y3 and the synthesized text Y1. If the similarity is greater than or equal to the similarity threshold, rewrite the input text y3 to determine the synthesized text Y2 and reply to the synthesized text Y2. The present invention improves the quality and efficiency of conversations in multi-turn conversations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language understanding, and particularly relates to a method for intent recognition in multi-turn conversations in question and answer. Background Art

[0002] With the popularization of intelligent electronic devices such as mobile phones, tablets, and smart watches, network products and applications attached to such electronic devices have gradually developed to use a conversational human-computer interaction method.

[0003] Currently, a chatbot is usually built into intelligent electronic devices. A chatbot refers to a computer program or software that can use the text, voice, pictures, or touch operations sent by the user as input data, understand the user intent expressed by the input data, and then realize human-computer interaction. It can be understood as an intelligent dialogue system that automatically makes responses according to the content sent by the user. To a certain extent, a chatbot can replace a real person to have a conversation and can be integrated into a dialogue system as an automatic online assistant for scenarios such as intelligent chatting, customer service, and information inquiry. During the human-machine conversation process, whether the chatbot can recognize the user's intent and give an answer that matches the conversation initiated by the user, or execute an operation instruction that matches the user's intent, has a significant impact on improving the conversation quality between humans and machines.

[0004] However, during the multi-turn conversation process, the user's intent may change at any time during different conversation turns. Due to the diversity and long-tail characteristics of the conversational search scenario, this increases the difficulty of knowledge retrieval during the conversation process. Therefore, there is an urgent need to construct a technical solution for intent recognition of users based on the conversation initiated by the user. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for intent recognition in multi-turn conversations in question and answer, and solve the following technical problems:

[0006] During the multi-turn conversation process, the user's intent may change at any time during different conversation turns. Due to the diversity and long-tail characteristics of the conversational search scenario, this increases the difficulty of knowledge retrieval during the conversation process. Therefore, there is an urgent need to construct a technical solution for intent recognition of users based on the conversation initiated by the user.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] A method for intent recognition in multi-turn conversations in question and answer, comprising the following steps:

[0009] S1. Obtain the first input text of the current user and mark it as input text y1. Perform text filtering on the input text y1, conduct semantic analysis on the filtered input text y1, determine the intention tags corresponding to each word in the input text y1 according to the results of semantic analysis, and determine the relevant field corresponding to the input text y1 according to the intention tags.

[0010] S2. Obtain the second input text of the current user and mark it as input text y2. Obtain the relevant field corresponding to the input text y2, calculate the similarity between the input text y2 and the input text y1 according to the relevant field. If the similarity is greater than or equal to the preset similarity threshold, rewrite the input text y2 through the preset LLM model and the input text y1, and determine the synthesized text Y1 according to the result of text rewriting, and reply to the synthesized text Y1.

[0011] S3. Obtain the third input text of the current user and mark it as input text y3. Obtain the relevant field corresponding to the input text y3, and calculate the similarity between the input text y3 and the synthesized text Y1 according to the relevant field.

[0012] If the similarity is greater than or equal to the similarity threshold, rewrite the input text y3 through the preset LLM model and the synthesized text Y1, and determine the synthesized text Y2 according to the result of text rewriting, and reply to the synthesized text Y2.

[0013] As a further solution of the present invention: In the S1, the specific process of text filtering is as follows:

[0014] Remove special characters in the text through regular expressions and replace the special characters with blank strings. The special characters include punctuation marks and special control characters; use a predefined stop word list to remove stop words in the text. The stop words include "de", "shi", and "le".

[0015] As a further solution of the present invention: In the S2, the following steps are further included:

[0016] If the similarity is less than the preset similarity threshold, directly reply to the input text y2.

[0017] As a further solution of the present invention: In the S2, the specific process of determining the synthesized text Y1 is as follows:

[0018] S11. Obtain all the results of text rewriting and mark them as hypothesis texts b1, b2,..., ba, where a is a positive integer. Select any hypothesis text and reply to it to obtain the hypothesis text reply corresponding to the hypothesis text, and obtain all the hypothesis text replies and mark them as h1, h2,..., ha;

[0019] S12. Obtain the non-numerical data of any hypothetical text response, encode the non-numerical data to obtain numerical data, convert the encoded hypothetical text response into a number of feature vectors, and generate a feature set for each hypothetical text response; calculate the Euclidean distance I between all feature sets, and use the Euclidean distance I as the similarity metric for K-means clustering to obtain the target category clusters;

[0020] S13. Taking any feature set in the target category cluster as the center, calculate the aggregation degree of this feature set and mark it as DP, and calibrate the hypothetical text corresponding to the feature set with the smallest DP value as the synthetic text Y1;

[0021] The calculation formula of DP is:

[0022] ;

[0023] where z is the number of feature sets in the target category cluster, z0 is the central feature set point, and v0 is the other feature set points in the target category cluster.

[0024] As a further solution of the present invention: in the said S3, the following steps are further included:

[0025] If the similarity is less than the similarity threshold, execute the following steps:

[0026] S21. Calculate the similarity X1 between the input text y3 and the input text y1, and calculate the similarity X2 between the input text y3 and the input text y2;

[0027] S22. Calculate the difference between the similarity X1 and the similarity X2. If the difference is greater than 0, rewrite the input text y3 through the LLM model and the input text y1, and determine the synthetic text Y3 according to the result of the text rewriting, and reply to the said synthetic text Y3;

[0028] If the difference is less than 0, rewrite the input text y3 through the LLM model and the input text y2, and determine the synthetic text Y4 according to the result of the text rewriting, and reply to the said synthetic text Y4;

[0029] If the difference is equal to 0, rewrite the input text y3 through the LLM model, the input text y1 and the input text y2, and determine the synthetic text Y5 according to the result of the text rewriting, and reply to the said synthetic text Y5.

[0030] As a further solution of the present invention: in the said S1, the specific process of determining the relevant field corresponding to the input text y1 according to the intention label is:

[0031] A database is established, and intention tags of the labeled fields are stored in the database. The intention tags are related to several fields at the same time. The relevance p between the intention tags and different fields is generated. All words in the input text y1 are summarized into corresponding intention tags, and the relevance p1, p2, …, pq of the intention tags involved in any field are statistically calculated and calculated according to the calculation formula The total relevance x of the field is calculated. If the total relevance x is greater than or equal to the preset threshold, the field is designated as the relevant field corresponding to the input text y1; if the total relevance x is less than the preset threshold, the field is not designated as the relevant field corresponding to the input text y1.

[0032] As a further solution of the present invention: in S2, the specific calculation process of the similarity is as follows: the relevant fields corresponding to the input text y2 and the input text y1 are obtained, and then all overlapping relevant fields are obtained and marked as L1, L2, …, LN, and the calculation formula The similarity C is calculated, where M is the number of relevant fields corresponding to the input text y1, and N is the number of overlapping relevant fields.

[0033] Advantages of the present invention:

[0034] Through the text processing of multiple rounds of conversations, by calculating the similarity and text rewriting, the present invention can effectively identify the similarity between the user input and the previous input to identify whether the user's intention has changed, and avoid repeated answers or irrelevant answers. For example, if the user's input is highly similar to the previous content, the user input text is optimized through text rewriting, so as to provide more accurate and relevant content. Through the semantic analysis and intention recognition of each input text, the conversation flow can be adjusted in real time and modified as needed, and more reasonable responses can be made according to the actual situation. This makes the conversation more contextually coherent, personalized and adaptable. The present invention reduces the difficulty of knowledge retrieval in the process of multiple rounds of conversations and improves the quality and efficiency of multiple rounds of conversations. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The present invention will be further described below with reference to the accompanying drawings.

[0036] Figure 1 It is a schematic flow chart of an intention recognition method for multiple rounds of conversations in question and answer of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Please refer to Figure 1 As shown, the present invention is a method for intent recognition in multi-turn conversations in question and answer, including the following steps:

[0039] S1, Obtain the first input text of the current user and label it as input text y1, perform text filtering on the input text y1, perform semantic analysis on the filtered input text y1, determine the intent labels corresponding to each word in the input text y1 according to the results of the semantic analysis, and determine the relevant field corresponding to the input text y1 according to the intent labels;

[0040] S2, Obtain the second input text of the current user and label it as input text y2, obtain the relevant field corresponding to the input text y2, calculate the similarity between the input text y2 and the input text y1 according to the relevant field, if the similarity is greater than or equal to the preset similarity threshold, rewrite the input text y2 through the preset LLM model and the input text y1, and determine the synthesized text Y1 according to the result of the text rewriting, and reply to the synthesized text Y1;

[0041] S3, Obtain the third input text of the current user and label it as input text y3, obtain the relevant field corresponding to the input text y3, and calculate the similarity between the input text y3 and the synthesized text Y1 according to the relevant field;

[0042] If the similarity is greater than or equal to the similarity threshold, rewrite the input text y3 through the preset LLM model and the synthesized text Y1, and determine the synthesized text Y2 according to the result of the text rewriting, and reply to the synthesized text Y2.

[0043] Through the text processing of multi-turn conversations, by calculating similarities and text rewriting, the present invention can effectively identify the similarity between the user input and the previous input to determine whether the user's intent has changed, avoiding repeated or irrelevant answers. For example, if the user's input is highly similar to the previous content, the user input text is optimized through text rewriting to provide more accurate and relevant content. Through semantic analysis and intent recognition of each input text, the conversation flow can be adjusted in real time and modified as needed, and more reasonable responses can be made according to the actual situation. This makes the conversation more contextually coherent, personalized and adaptable. The present invention reduces the knowledge retrieval difficulty in the multi-turn conversation process and improves the quality and efficiency of the multi-turn conversation.

[0044] It should be noted that the present invention only discusses the solutions for the first three conversations of the user. In actual applications, the number of conversations of the user may far exceed three. Therefore, when the user conducts the fourth, fifth, and more conversations, the above-mentioned step S3 can be repeatedly executed, thereby reducing the difficulty of knowledge retrieval in the multi-round conversation process and improving the quality and efficiency of the multi-round conversation.

[0045] In a preferred embodiment of the present invention, in S1, the specific process of text filtering is as follows:

[0046] Remove special characters in the text through regular expressions and replace the special characters with blank strings. The special characters include punctuation marks and special control characters; use a predefined stop word list to remove stop words in the text. The stop words include "de", "shi", and "le".

[0047] In another preferred embodiment of the present invention, in S2, the following steps are further included:

[0048] If the similarity is less than the preset similarity threshold, directly reply to the input text y2.

[0049] In another preferred embodiment of the present invention, in S2, the specific process of determining the synthesized text Y1 is as follows:

[0050] S11, obtain all the results of text rewriting and mark them as hypothetical texts b1, b2,..., ba, where a is a positive integer. Select any one of the hypothetical texts and reply to obtain the hypothetical text reply corresponding to the hypothetical text. Obtain all the hypothetical text replies and mark them as h1, h2,..., ha;

[0051] S12, obtain the non-numerical data of any one of the hypothetical text replies, encode the non-numerical data to obtain numerical data, convert the encoded hypothetical text reply into several feature vectors, and generate a feature set for each hypothetical text reply; calculate the Euclidean distance I between all the feature sets, and use the Euclidean distance I as the similarity metric for K-means clustering to obtain the target category cluster;

[0052] S13, take any one of the feature sets in the target category cluster as the center, calculate the aggregation degree of the feature set and mark it as DP, and calibrate the hypothetical text corresponding to the feature set with the minimum DP value as the synthesized text Y1;

[0053] The calculation formula of DP is:

[0054] ;

[0055] where z is the number of feature sets in the target category cluster, z0 is the central feature set point, and v0 is the other feature set points in the target category cluster.

[0056] The LLM model can analyze the user's query and identify the core intention therein. For example, if the user asks "What's the weather like today?", the LLM will understand that the user wants to know the current weather condition, extract the key information. After understanding the intention, the LLM will extract the key information in the query, such as time, location, topic, etc. In the example of "What's the weather like today?", the LLM will identify "today" and "weather" as the key information. Based on the analysis of the original query intention and key information, the LLM can rewrite the query in different ways. This process is not limited to literal replacement, but may involve changes in different expressions, tones, levels of detail, etc. For example, for "What's the weather like today?", it will generate queries including: "What's the weather forecast for today?", "Will it rain today?", "How's the weather today?", "What's the temperature today?". These different query rewrites, although different in statements, all convey the same core intention - querying about the weather condition today, thus obtaining multiple text rewrite results.

[0057] In another preferred embodiment of the present invention, in S3, the following steps are further included:

[0058] If the similarity is less than the similarity threshold, the following steps are executed:

[0059] S21, calculate the similarity X1 between the input text y3 and the input text y1, and calculate the similarity X2 between the input text y3 and the input text y2;

[0060] S22, calculate the difference between the similarity X1 and the similarity X2. If the difference is greater than 0, rewrite the input text y3 through the LLM model and the input text y1, determine the synthesized text Y3 according to the result of the text rewrite, and reply to the synthesized text Y3;

[0061] If the difference is less than 0, rewrite the input text y3 through the LLM model and the input text y2, determine the synthesized text Y4 according to the result of the text rewrite, and reply to the synthesized text Y4;

[0062] If the difference is equal to 0, rewrite the input text y3 through the LLM model, the input text y1 and the input text y2, determine the synthesized text Y5 according to the result of the text rewrite, and reply to the synthesized text Y5.

[0063] In another preferred embodiment of the present invention, in S1, the specific process of determining the relevant field corresponding to the input text y1 according to the intention label is:

[0064] A database is established, and intention tags of the annotated fields are stored in the database. The intention tags are related to several fields at the same time. The correlation p between the intention tags and different fields is generated. All words in the input text y1 are summarized into corresponding intention tags, and the correlations p1, p2, …, pq of the intention tags involved in any field are counted and calculated according to the calculation formula The total correlation x of this field is calculated. If the total correlation x is greater than or equal to the preset threshold, this field is designated as the relevant field corresponding to the input text y1; if the total correlation x is less than the preset threshold, this field is not designated as the relevant field corresponding to the input text y1.

[0065] In another preferred embodiment of the present invention, in S2, the specific calculation process of the similarity is as follows: the relevant fields corresponding to the input text y2 and the input text y1 are obtained, and then all overlapping relevant fields are obtained and marked as L1, L2, …, LN, and the calculation formula The similarity C is calculated, where M is the number of relevant fields corresponding to the input text y1, and N is the number of overlapping relevant fields.

[0066] The above has described an embodiment of the present invention in detail, but the above content is only a preferred embodiment of the present invention and cannot be considered as limiting the implementation scope of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A method for intent recognition in multi-turn conversations in question and answer, characterized in that, It includes the following steps: S1. Obtain the first input text of the current user and mark it as input text y1. Filter the input text y1, perform semantic analysis on the filtered input text y1, determine the intent tags corresponding to each word in the input text y1 according to the results of the semantic analysis, and determine the relevant field corresponding to the input text y1 according to the intent tags; S2. Obtain the second input text of the current user and mark it as input text y2. Obtain the relevant field corresponding to the input text y2, calculate the similarity between the input text y2 and the input text y1 according to the relevant field. If the similarity is greater than or equal to the preset similarity threshold, rewrite the input text y2 through the preset LLM model and the input text y1, and determine the synthesized text Y1 according to the result of the text rewriting, and reply to the synthesized text Y1; S3. Obtain the third input text of the current user and mark it as input text y3. Obtain the relevant field corresponding to the input text y3, and calculate the similarity between the input text y3 and the synthesized text Y1 according to the relevant field; If the similarity is greater than or equal to the similarity threshold, rewrite the input text y3 through the preset LLM model and the synthesized text Y1, and determine the synthesized text Y2 according to the result of the text rewriting, and reply to the synthesized text Y2; The specific process of determining the synthesized text Y1 is as follows: S11. Obtain all the results of text rewriting and mark them as hypothetical texts b1, b2,..., ba, where a is a positive integer. Select any one of the hypothetical texts and reply, obtain the hypothetical text reply corresponding to the hypothetical text, and obtain all the hypothetical text replies and mark them as h1, h2,..., ha; S12. Obtain the non-numerical data of any one of the hypothetical text replies, encode the non-numerical data to obtain numerical data, convert the encoded hypothetical text reply into several feature vectors, and generate the feature set of each hypothetical text reply; calculate the Euclidean distance I between all the feature sets, use the Euclidean distance I as the similarity measure for K-means clustering, and obtain the target category cluster; S13. Take any one of the feature sets in the target category cluster as the center, calculate the aggregation degree of the feature set and mark it as DP, and calibrate the hypothetical text corresponding to the feature set with the smallest DP value as the synthesized text Y1; The calculation formula of DP is: ; where z is the number of feature sets in the target category cluster, z0 is the central feature set point, and v0 is other feature set points in the target category cluster.

2. The method for intent recognition in multi-turn conversations in question and answer according to claim 1, characterized in that, In the above S1, the specific process of text filtering is as follows: Remove the special characters in the text through regular expressions and replace the special characters with blank strings. The special characters include punctuation marks and special control characters; use a predefined stop word list to remove the stop words in the text. The stop words include "de", "shi", and "le".

3. A method for intention recognition in multi-turn conversations in question and answer, according to claim 1, characterized in that, In the above S2, the following steps are further included: If the similarity is less than the preset similarity threshold, directly reply to the input text y2.

4. The method for intention recognition in multi-round conversations in Q&A according to claim 1, characterized in that, In the above S3, the following steps are further included: If the similarity is less than the similarity threshold, perform the following steps: S21, calculate the similarity X1 between the input text y3 and the input text y1, and calculate the similarity X2 between the input text y3 and the input text y2; S22, calculate the difference between the similarity X1 and the similarity X2. If the difference is greater than 0, rewrite the input text y3 through the LLM model and the input text y1, determine the synthesized text Y3 according to the result of the text rewriting, and reply to the synthesized text Y3; If the difference is less than 0, rewrite the input text y3 through the LLM model and the input text y2, determine the synthesized text Y4 according to the result of the text rewriting, and reply to the synthesized text Y4; If the difference is equal to 0, rewrite the input text y3 through the LLM model, the input text y1 and the input text y2, determine the synthesized text Y5 according to the result of the text rewriting, and reply to the synthesized text Y5.

5. A method for intent recognition in multi-turn conversations in Q&A, as claimed in claim 1, characterized in that In the above S1, the specific process of determining the specific field related to the input text y1 according to the intention label is as follows: Establish a database, in which intention tags of the labeled fields are stored. The intention tags are related to several fields at the same time. Generate the relevance p of the intention tags with different fields. Induce all words in the input text y1 into the corresponding intention tags, and count the relevance p1, p2,..., pq of the intention tags involved in any field. Then, according to the calculation formula calculate the total relevance x of this field. If the total relevance x is greater than or equal to the preset threshold, then label this field as the relevant field corresponding to the input text y1; if the total relevance x is less than the preset threshold, then do not label this field as the relevant field corresponding to the input text y1.

6. The method for intention recognition in multi-turn conversations in question and answer according to claim 1, characterized in that, In S2, the specific calculation process of the similarity is as follows: Obtain the relevant fields corresponding to the input text y2 and the input text y1, and then obtain all the overlapping relevant fields and mark them as L1, L2,..., LN. The calculation formula is used to calculate the similarity C, where M is the number of relevant fields corresponding to the input text y1, and N is the number of overlapping relevant fields.

Citation Information

Patent Citations

  • Label extraction method based on short text clustering technology

    CN111414479A

  • Multi-source and multi-mode fused knowledge reasoning method, system and device and medium

    CN119005340A