Dialogue intention recognition method and system, electronic equipment and storage medium

Through clustering and grading of dialogue texts, combined with multi-layer perceptron model, the pre-trained model is adjusted, which solves the problem of low accuracy in existing intention recognition technology, and realizes efficient and accurate recognition of dialogue text intention recognition.

CN120354860AActive Publication Date: 2025-07-22CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510845763.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The existing intention recognition technology has a low accuracy rate for intent recognition of dialogue texts, especially when the amount of information is insufficient, the recognition effect is poor.

Method used

By clustering and grading intention similarity of multiple historical dialogue texts, extracting keyword information, adjusting it using pre-trained models, building a tuned intention recognition model, and combining with a multi-layer perceptron for intention recognition.

Benefits of technology

The accuracy of dialogue text intention recognition is improved, especially when the amount of information is insufficient, and the accuracy and efficiency of recognition is improved by enhancing the amount of keyword information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354860A_ABST
    Figure CN120354860A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue intention recognition method and system, electronic equipment and a storage medium, and the method comprises the steps: carrying out the intention classification of each type of historical dialogue text according to the intention importance according to a second clustering result, and obtaining an intention classification result; extracting a plurality of first keywords of the dialogue text to be recognized, and judging whether the information amount of the plurality of first keywords is greater than or equal to a preset threshold according to the intention grading result; if the information amount of the plurality of first keywords is greater than or equal to a preset threshold value, acquiring a plurality of target dialogue texts of a category corresponding to the plurality of first keywords from the plurality of historical dialogue texts according to the first clustering result; the multiple target dialogue texts are adopted to adjust a pre-trained intention recognition model, and an optimized intention recognition model is obtained; and performing intention recognition on all keywords corresponding to the dialogue text to be recognized by adopting the optimized intention recognition model to obtain an intention recognition result. The dialogue text intention recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of intent recognition, and particularly to a method, system, electronic device and storage medium for dialogue intent recognition. Background Art

[0002] With the continuous development of deep learning and artificial intelligence technologies, intelligent dialogue systems have been widely applied in multiple fields such as medical treatment, e-commerce, entertainment, and autonomous driving, such as intelligent customer service, digital doctors, intelligent housekeepers, etc. Intelligent dialogue needs to accurately understand the user's intent in the dialogue in order to better solve problems for the user. In order to accurately understand the user's intent in the dialogue, a technology that can understand the meaning of the user's dialogue is required. Therefore, the birth of intent recognition technology is of great significance.

[0003] Intent recognition is an important task in natural language processing (NLP). It aims to determine the intent or purpose expressed in the user's input statement. Simply put, intent recognition is to perform semantic understanding on the user's words in order to better answer the user's questions or provide relevant services. Traditional intent recognition methods are generally based on template matching or artificial feature sets, which are time-consuming and laborious and have poor scalability. If a traditional neural network model such as a convolutional neural network (CNN) is used for intelligent recognition, there is a lack of context information, the accuracy of intent recognition is low, and a large amount of dialogue text information is required to improve the accuracy.

[0004] Therefore, the accuracy of existing related intent recognition technologies for dialogue text intent recognition is relatively low. Summary of the Invention

[0005] The present application aims to propose a method, system, electronic device and storage medium for dialogue intent recognition, which can improve the accuracy of dialogue text intent recognition.

[0006] In a first aspect, an embodiment of the present application provides a method for dialogue intent recognition, the method comprising: Obtaining a plurality of historical dialogue texts and their corresponding first intent labels, a plurality of training dialogue texts and their corresponding second intent labels, and a dialogue text to be recognized; Based on the first intent label, clustering the plurality of historical dialogue texts according to intent similarity to obtain a first clustering result, and, based on the second intent label, clustering the plurality of training dialogue texts according to intent similarity to obtain a second clustering result; According to the second clustering result, grading the historical dialogue texts of each category according to intent importance to obtain an intent grading result; Extract multiple first keywords of the dialogue text to be recognized, and determine whether the information volume of the multiple first keywords is greater than or equal to a preset threshold according to the intention classification result; If the information volume of the multiple first keywords is greater than or equal to the preset threshold, obtain multiple target dialogue texts of the category corresponding to the multiple first keywords from the multiple historical dialogue texts according to the first clustering result; Use the multiple target dialogue texts to adjust a pre-trained intention recognition model to obtain a tuned intention recognition model, where the pre-trained intention recognition model is trained based on the multiple training dialogue texts and their corresponding second intention labels; Use the tuned intention recognition model to perform intention recognition on all keywords corresponding to the dialogue text to be recognized to obtain an intention recognition result.

[0007] Compared with the prior art, the first aspect of this application has the following beneficial effects: This method obtains multiple historical dialogue texts and their corresponding first intention labels, multiple training dialogue texts and their corresponding second intention labels, and the dialogue text to be recognized; based on the first intention labels, cluster the multiple historical dialogue texts according to intention similarity to obtain a first clustering result, and, based on the second intention labels, cluster the multiple training dialogue texts according to intention similarity to obtain a second clustering result; according to the second clustering result, classify the historical dialogue texts of each category according to intention importance to obtain an intention classification result; extract multiple first keywords of the dialogue text to be recognized, and determine whether the information volume of the multiple first keywords is greater than or equal to a preset threshold according to the intention classification result; if the information volume of the multiple first keywords is greater than or equal to the preset threshold, obtain multiple target dialogue texts of the category corresponding to the multiple first keywords from the multiple historical dialogue texts according to the first clustering result; use the multiple target dialogue texts to adjust a pre-trained intention recognition model to obtain a tuned intention recognition model, where the pre-trained intention recognition model is trained based on the multiple training dialogue texts and their corresponding second intention labels; use the tuned intention recognition model to perform intention recognition on all keywords corresponding to the dialogue text to be recognized to obtain an intention recognition result. In this way, by first clustering the multiple historical dialogue texts according to intention similarity, and then classifying the historical dialogue texts of each category according to intention importance, because the intention classification of the same category is better distinguished and the intention classification efficiency can be improved; then determine whether the information volume of the multiple first keywords is greater than or equal to the preset threshold, because sufficient information volume can improve the accuracy of dialogue text intention recognition; and then adjust the pre-trained intention recognition model with the multiple target dialogue texts, which can improve the accuracy of the intention recognition model, so as to further improve the accuracy of dialogue text intention recognition.

[0008] In some embodiments, the intention recognition model includes a word vector layer, normalization, a convolutional layer, a ReLU function, multiple semantic feature extraction modules, and a multi-layer perceptron. Using the optimized intention recognition model to perform intention recognition on all keywords corresponding to the dialogue text to be recognized, the intention recognition result is obtained, including: Input all keywords corresponding to the dialogue text to be recognized into the optimized intention recognition model, and vectorize all keywords through the word vector layer to obtain target keyword vectors; Perform feature vector extraction on the target keyword vectors through normalization, a convolutional layer, and a ReLU function to obtain keyword features; Input the keyword features into the first semantic feature extraction module for semantic feature extraction to obtain a first semantic feature extraction result; Input the first semantic feature extraction result and the keyword features into the second semantic feature extraction module for semantic feature extraction to obtain a second semantic feature extraction result; Input the second semantic feature extraction result and the keyword features into the third semantic feature extraction module for semantic feature extraction to obtain a third semantic feature extraction result; Perform feature fusion on the first semantic feature extraction result, the second semantic feature extraction result, and the third semantic feature extraction result to obtain fused semantic features; Input the fused semantic features into the multi-layer perceptron for intention recognition to obtain an intention recognition result.

[0009] In some embodiments, the pre-trained intention recognition model is obtained in the following manner: Construct an intention recognition model including a word vector layer, normalization, a convolutional layer, a ReLU function, multiple semantic feature extraction modules, and a multi-layer perceptron; Extract the keywords of each training dialogue text in the multiple training dialogue texts to obtain a keyword data set; Screen the keywords in the keyword data set to obtain a screened data set, and obtain a second intention label corresponding to the screened data set; Pre-train the constructed intention recognition model through the screened data set and its corresponding second intention label to obtain a pre-trained intention recognition model.

[0010] In some embodiments, the screening of the keywords in the keyword data set to obtain a screened data set includes: Vectorize the keywords in the keyword data set to obtain a keyword vector data set; Input the keyword vector dataset into the multi-head attention mechanism to obtain the weights corresponding to each keyword; Calculate the importance degree of each keyword according to the weights corresponding to each keyword; Screen the keywords in the keyword dataset according to the importance degree of each keyword to obtain the screened dataset.

[0011] In some embodiments, the step of determining whether the information amount of the multiple first keywords is greater than or equal to a preset threshold according to the intention classification result includes: Determine the category corresponding to the to-be-recognized dialogue text in the second clustering result according to the multiple first keywords; Select the relevant intention classification result according to the category corresponding to the to-be-recognized dialogue text to obtain the target intention classification result; Calculate the information amount value of the multiple first keywords according to the target intention classification result; Compare the information amount value of the multiple first keywords with the preset threshold to determine whether the information amount of the multiple first keywords is greater than or equal to the preset threshold.

[0012] In some embodiments, the step of calculating the information amount value of the multiple first keywords according to the target intention classification result includes: ; wherein, represents the information amount value of the multiple first keywords, represents the correlation between the first keyword i and the intention level j in the target intention classification result. The correlation of the first keyword belonging to the intention level j is 1, and the correlation of the first keyword not belonging to the intention level j is 0. k represents the number of keywords belonging to the same intention level, n represents the number of intention levels corresponding to the dialogue text of the current category, and m represents the number of the multiple first keywords.

[0013] In some embodiments, the dialogue intention recognition method further includes: If the information amount of all the first keywords corresponding to the to-be-recognized dialogue text is less than the preset threshold, input the to-be-recognized dialogue text into the trained multi-turn dialogue model to obtain the multi-turn dialogue text related to the to-be-recognized dialogue text; Extract multiple second keywords from the multi-turn dialogue text, and merge all the second keywords and all the first keywords to obtain the merged keyword set; Calculate the information amount value of all the keywords in the merged keyword set according to the target intention classification result; Compare the information quantity values of all keywords in the merged keyword set with the preset threshold to obtain a comparison result, until the comparison result indicates that the information quantity values of all keywords in the merged keyword set are greater than or equal to the preset threshold, and obtain all keywords corresponding to the to-be-recognized dialogue text.

[0014] In a second aspect, an embodiment of the present application further provides a dialogue intention recognition system, and the system includes: A data acquisition unit, configured to acquire a plurality of historical dialogue texts and their corresponding first intention labels, a plurality of training dialogue texts and their corresponding second intention labels, and a to-be-recognized dialogue text; A data clustering unit, configured to cluster the plurality of historical dialogue texts according to intention similarity based on the first intention label to obtain a first clustering result, and cluster the plurality of training dialogue texts according to intention similarity based on the second intention label to obtain a second clustering result; An intention grading unit, configured to grade the historical dialogue texts of each category according to intention importance based on the second clustering result to obtain an intention grading result; An information quantity judgment unit, configured to extract a plurality of first keywords of the to-be-recognized dialogue text, and judge whether the information quantity of the plurality of first keywords is greater than or equal to a preset threshold according to the intention grading result; A historical data acquisition unit, configured to, if the information quantity of the plurality of first keywords is greater than or equal to the preset threshold, obtain a plurality of target dialogue texts corresponding to the plurality of first keywords from the plurality of historical dialogue texts according to the first clustering result; A model adjustment unit, configured to adjust a pre-trained intention recognition model by using the plurality of target dialogue texts to obtain a tuned intention recognition model, where the pre-trained intention recognition model is trained based on the plurality of training dialogue texts and their corresponding second intention labels; An intention recognition unit, configured to perform intention recognition on all keywords corresponding to the to-be-recognized dialogue text by using the tuned intention recognition model to obtain an intention recognition result.

[0015] In a third aspect, an embodiment of the present application further provides an electronic device, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute a dialogue intention recognition method as described above.

[0016] Fourthly, an embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute a dialogue intention recognition method as described above.

[0017] It can be understood that the beneficial effects of the above second to fourth aspects compared with the related art are the same as those of the above first aspect compared with the related art. For relevant descriptions, reference can be made to the relevant descriptions in the above first aspect and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where: Figure 1 is a schematic flowchart of an embodiment of the dialogue intention recognition method provided by the present application; Figure 2 is a schematic structural diagram of an intention recognition model in the best embodiment of the dialogue intention recognition method provided by the present application; Figure 3 is a schematic structural diagram of an embodiment of the dialogue intention recognition system provided by the present application; Figure 4 is a schematic structural diagram of an embodiment of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.

[0020] In the description of the present application, if the first, second, etc. are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0021] In the description of the present application, it should be understood that the orientation or positional relationship indicated by terms such as up, down, etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present application.

[0022] In the description of this application, it should be noted that unless otherwise clearly defined, terms such as setting, installation, connection, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in this application in combination with the specific content of the technical solution.

[0023] First, analyze several nouns involved in this application: The k-means clustering algorithm, that is, the k-means algorithm, is an iterative clustering analysis algorithm. Its steps are as follows: initially divide the data into K groups, then randomly select K objects as the initial clustering centers, and then calculate the distance between each object and each seed clustering center, and assign each object to the clustering center closest to it. The clustering centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the clustering center of the cluster will be recalculated based on the existing objects in the cluster. This process will be repeated continuously until a certain termination condition is met. The termination condition can be that no (or the minimum number of) objects are reassigned to different clusters, no (or the minimum number of) clustering centers change anymore, or the sum of squared errors is locally minimized.

[0024] The DBSCAN algorithm: is a relatively representative density-based clustering algorithm. Different from partitioning and hierarchical clustering methods, it defines a cluster as the largest set of density-connected points, can divide regions with sufficient high density into clusters, and can discover clusters of any shape in a spatial database with noise.

[0025] The large language model LLM: refers to a deep learning model trained using a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language production on various topics through training on a huge dataset. Its core idea is to learn the patterns and structures of natural language through large-scale unsupervised training, and to simulate the human language cognition and generation process to a certain extent.

[0026] The ChatGPT model: is built based on the GPT system large model. OpenAI uses the training method of "reinforcement learning from human feedback" (RLHF). The essence of ChatGPT is an intelligent tool that improves the human brain's ability to collect, organize, calculate, analyze, etc. various information materials, and is a tool system that provides rich and accurate solutions, schemas, etc. materials or conditions for the "concept construction" of the human brain.

[0027] The multi-layer perceptron MLP: is a feedforward artificial neural network model that maps multiple input datasets to a single output dataset.

[0028] Word2vec: A group of related models used to generate word vectors. These models are shallow and two-layer neural networks, trained to reconstruct linguistic word texts. The network represents words and needs to guess the input words at adjacent positions. Under the bag-of-words model assumption in Word2vec, the order of words is not important. After training, the Word2vec model can be used to map each word to a vector, which can represent the relationship between words.

[0029] With the continuous development of deep learning and artificial intelligence technologies, intelligent dialogue systems have been widely applied in many fields such as healthcare, e-commerce, entertainment, and autonomous driving, such as intelligent customer service, digital doctors, intelligent butlers, etc. Intelligent dialogue needs to accurately understand the user's intention in the dialogue in order to better solve problems for the user. To accurately understand the user's intention in the dialogue, a technology that can understand the meaning of the user's dialogue is required. Therefore, the birth of intent recognition technology is of great significance.

[0030] Intent Recognition is an important task in Natural Language Processing (NLP). It aims to determine the intention or purpose expressed in the user's input statement. Simply put, intent recognition is to perform semantic understanding of the user's words in order to better answer the user's questions or provide relevant services. Traditional intent recognition methods are generally based on template matching or artificial feature sets, which are time-consuming and laborious and have poor scalability. If traditional neural network models such as Convolutional Neural Network (CNN) are used for intelligent recognition, there is a lack of context information, the accuracy of intent recognition is low, and a large amount of dialogue text information is required to improve the accuracy.

[0031] Therefore, the accuracy of existing related intent recognition technologies for recognizing the intent of dialogue texts is relatively low.

[0032] To solve the problem that the accuracy of existing related intent recognition technologies for recognizing the intent of dialogue texts is relatively low, this application proposes a dialogue intent recognition method, system, electronic device, and storage medium.

[0033] Refer to Figure 1 , the embodiment of this application provides a dialogue intent recognition method, which includes the following steps: Step S100: Obtain multiple historical dialogue texts and their corresponding first intent labels, multiple training dialogue texts and their corresponding second intent labels, and the dialogue text to be recognized; Step S200: Based on the first intent label, cluster multiple historical dialogue texts according to intent similarity to obtain a first clustering result, and, based on the second intent label, cluster multiple training dialogue texts according to intent similarity to obtain a second clustering result; Step S300: According to the second clustering result, grade the historical dialogue texts of each category by intention importance to obtain an intention grading result; Step S400: Extract multiple first keywords of the dialogue text to be recognized, and determine whether the information volume of the multiple first keywords is greater than or equal to a preset threshold according to the intention grading result; Step S500: If the information volume of the multiple first keywords is greater than or equal to the preset threshold, obtain multiple target dialogue texts corresponding to the multiple first keywords from the multiple historical dialogue texts according to the first clustering result; Step S600: Use the multiple target dialogue texts to adjust the pre-trained intention recognition model to obtain a tuned intention recognition model, and the pre-trained intention recognition model is trained based on multiple training dialogue texts and their corresponding second intention labels; Step S700: Use the tuned intention recognition model to perform intention recognition on all keywords corresponding to the dialogue text to be recognized to obtain an intention recognition result.

[0034] In this embodiment, by obtaining a plurality of historical dialogue texts and their corresponding first intent labels, a plurality of training dialogue texts and their corresponding second intent labels, and the dialogue text to be recognized; based on the first intent labels, clustering the plurality of historical dialogue texts according to similar intents to obtain a first clustering result, and, based on the second intent labels, clustering the plurality of training dialogue texts according to similar intents to obtain a second clustering result; according to the second clustering result, grading the intents of each category of historical dialogue texts according to intent importance to obtain an intent grading result; extracting a plurality of first keywords of the dialogue text to be recognized, and judging whether the information amount of the plurality of first keywords is greater than or equal to a preset threshold according to the intent grading result; if the information amount of the plurality of first keywords is greater than or equal to the preset threshold, then according to the first clustering result, obtaining a plurality of target dialogue texts corresponding to the plurality of first keywords from the plurality of historical dialogue texts; adjusting the pre-trained intent recognition model with the plurality of target dialogue texts to obtain a tuned intent recognition model, where the pre-trained intent recognition model is trained based on the plurality of training dialogue texts and their corresponding second intent labels; using the tuned intent recognition model to perform intent recognition on all the keywords corresponding to the dialogue text to be recognized to obtain an intent recognition result. In this way, by first clustering the plurality of historical dialogue texts according to similar intents, and then grading the intents of each category of historical dialogue texts according to intent importance, because the intent grading of the same category is better distinguished and the intent grading efficiency can be improved; then judging whether the information amount of the plurality of first keywords is greater than or equal to the preset threshold, because sufficient information amount can improve the accuracy of dialogue text intent recognition; and then adjusting the pre-trained intent recognition model with the plurality of target dialogue texts can improve the accuracy of the intent recognition model, thereby further improving the accuracy of dialogue text intent recognition.

[0035] The above-mentioned clustering the plurality of historical dialogue texts according to similar intents based on the first intent labels to obtain a first clustering result, and clustering the plurality of training dialogue texts according to similar intents based on the second intent labels to obtain a second clustering result can be to use well-known algorithms in the art such as the k-means algorithm and the DBSCAN algorithm to cluster the plurality of historical dialogue texts to obtain a first clustering result, and cluster the plurality of training dialogue texts to obtain a second clustering result.

[0036] The above-mentioned grading the intents of each category of historical dialogue texts according to intent importance can be for a person to distinguish the intent importance according to experience, so as to grade the intents of each category of historical dialogue texts according to intent importance.

[0037] The above-mentioned preset threshold can be set artificially and can be changed according to the actual situation, and this embodiment does not make specific limitations.

[0038] In some embodiments, the intent recognition model includes a word vector layer, normalization, a convolutional layer, a ReLU function, multiple semantic feature extraction modules, and a multi-layer perceptron. The tuned intent recognition model is used to perform intent recognition on all keywords corresponding to the conversation text to be recognized, and the intent recognition result is obtained, including: All keywords corresponding to the conversation text to be recognized are input into the tuned intent recognition model, and all keywords are vectorized through the word vector layer to obtain target keyword vectors; The target keyword vectors are subjected to feature vector extraction through normalization, a convolutional layer, and a ReLU function to obtain keyword features; The keyword features are input into the first semantic feature extraction module for semantic feature extraction to obtain a first semantic feature extraction result; The first semantic feature extraction result and the keyword features are input into the second semantic feature extraction module for semantic feature extraction to obtain a second semantic feature extraction result; The second semantic feature extraction result and the keyword features are input into the third semantic feature extraction module for semantic feature extraction to obtain a third semantic feature extraction result; The first semantic feature extraction result, the second semantic feature extraction result, and the third semantic feature extraction result are subjected to feature fusion to obtain fused semantic features; The fused semantic features are input into the multi-layer perceptron for intent recognition to obtain the intent recognition result.

[0039] In this embodiment, all keywords corresponding to the conversation text to be recognized are input into the tuned intent recognition model, and all keywords are vectorized through the word vector layer to obtain target keyword vectors; the target keyword vectors are subjected to feature vector extraction through normalization, a convolutional layer, and a ReLU function to obtain keyword features; the keyword features are input into the first semantic feature extraction module for semantic feature extraction to obtain a first semantic feature extraction result; the first semantic feature extraction result and the keyword features are input into the second semantic feature extraction module for semantic feature extraction to obtain a second semantic feature extraction result; the second semantic feature extraction result and the keyword features are input into the third semantic feature extraction module for semantic feature extraction to obtain a third semantic feature extraction result; the first semantic feature extraction result, the second semantic feature extraction result, and the third semantic feature extraction result are subjected to feature fusion to obtain fused semantic features; the fused semantic features are input into the multi-layer perceptron for intent recognition to obtain the intent recognition result. In this way, more refined semantic features are continuously extracted through multiple semantic feature extraction modules to improve the accuracy of intent recognition for conversation text.

[0040] In some embodiments, the pre-trained intent recognition model is obtained through the following method: Construct an intention recognition model including a word vector layer, normalization, a convolutional layer, a ReLU function, multiple semantic feature extraction modules, and a multi-layer perceptron; Extract the keywords of each training dialogue text from multiple training dialogue texts to obtain a keyword dataset; Filter the keywords in the keyword dataset to obtain a filtered dataset, and obtain the second intention label corresponding to the filtered dataset; Pre-train the constructed intention recognition model with the filtered dataset and its corresponding second intention label to obtain a pre-trained intention recognition model.

[0041] In this embodiment, by constructing an intention recognition model including a word vector layer, normalization, a convolutional layer, a ReLU function, multiple semantic feature extraction modules, and a multi-layer perceptron; extracting the keywords of each training dialogue text from multiple training dialogue texts to obtain a keyword dataset; filtering the keywords in the keyword dataset to obtain a filtered dataset, and obtaining the second intention label corresponding to the filtered dataset; pre-training the constructed intention recognition model with the filtered dataset and its corresponding second intention label to obtain a pre-trained intention recognition model. In this way, pre-training the constructed intention recognition model with the filtered dataset and its corresponding second intention label can not only improve the training efficiency of the intention recognition model, but also eliminate the influence of some irrelevant keywords during the training process of the intention recognition model, thereby improving the accuracy of the intention recognition model during recognition.

[0042] In some embodiments, filtering the keywords in the keyword dataset to obtain a filtered dataset includes: Vectorize the keywords in the keyword dataset to obtain a keyword vector dataset; Input the keyword vector dataset into the multi-head attention mechanism to obtain the weight corresponding to each keyword; Calculate the importance degree of each keyword according to the weight corresponding to each keyword; Filter the keywords in the keyword dataset according to the importance degree of each keyword to obtain a filtered dataset.

[0043] In this embodiment, the keywords in the keyword dataset are vectorized to obtain a keyword vector dataset; the keyword vector dataset is input into a multi-head attention mechanism to obtain the weight corresponding to each keyword; according to the weight corresponding to each keyword, the importance degree of each keyword is calculated; according to the importance degree of each keyword, the keywords in the keyword dataset are screened to obtain a screened dataset. In this way, by screening some important keywords and removing some irrelevant keywords, the influence of some irrelevant keywords can be eliminated during the training process of the intent recognition model, thereby improving the accuracy of the intent recognition model during recognition.

[0044] In some embodiments, determining whether the information amount of multiple first keywords is greater than or equal to a preset threshold according to the intent classification result includes: According to multiple first keywords, determining the category corresponding to the dialogue text to be recognized in the second clustering result; Selecting a relevant intent classification result according to the category corresponding to the dialogue text to be recognized to obtain a target intent classification result; Calculating the information amount value of multiple first keywords according to the target intent classification result; Comparing the information amount value of multiple first keywords with a preset threshold to determine whether the information amount of multiple first keywords is greater than or equal to the preset threshold.

[0045] In this embodiment, according to multiple first keywords, determining the category corresponding to the dialogue text to be recognized in the second clustering result; selecting a relevant intent classification result according to the category corresponding to the dialogue text to be recognized to obtain a target intent classification result; calculating the information amount value of multiple first keywords according to the target intent classification result; comparing the information amount value of multiple first keywords with a preset threshold to determine whether the information amount of multiple first keywords is greater than or equal to the preset threshold. In this way, by calculating the information amount value of keywords based on intent classification to see whether the keyword information meets the standard, because if a keyword involves more levels, it means that the information involved in the keyword is more detailed. By using some detailed information for subsequent intent recognition, the accuracy of intent recognition can be improved.

[0046] In some embodiments, calculating the information amount value of multiple first keywords according to the target intent classification result includes: ; Wherein, represents the information amount value of multiple first keywords, Indicates the relevance of the first keyword i to the intention level j in the target intention classification result. The relevance of the first keyword belonging to the intention level j is 1, and the relevance of the first keyword not belonging to the intention level j is 0. k represents the number of keywords belonging to the same intention level, n represents the number of intention levels corresponding to the current category of dialogue text, and m represents the number of first keywords.

[0047] In some embodiments, the dialogue intention recognition method further includes: If the information amount of all the first keywords corresponding to the dialogue text to be recognized is less than the preset threshold, then input the dialogue text to be recognized into the trained multi-round dialogue model to obtain the multi-round dialogue text related to the dialogue text to be recognized; Extract multiple second keywords from the multi-round dialogue text, and merge all the second keywords and all the first keywords to obtain a merged keyword set; According to the target intention classification result, calculate the information amount values of all the keywords in the merged keyword set; Compare the information amount values of all the keywords in the merged keyword set with the preset threshold to obtain a comparison result, until the comparison result indicates that the information amount values of all the keywords in the merged keyword set are greater than or equal to the preset threshold, so as to obtain all the keywords corresponding to the dialogue text to be recognized.

[0048] In this embodiment, if the information amount of all the first keywords corresponding to the dialogue text to be recognized is less than the preset threshold, then input the dialogue text to be recognized into the trained multi-round dialogue model to obtain the multi-round dialogue text related to the dialogue text to be recognized; extract multiple second keywords from the multi-round dialogue text, and merge all the second keywords and all the first keywords to obtain a merged keyword set; according to the target intention classification result, calculate the information amount values of all the keywords in the merged keyword set; compare the information amount values of all the keywords in the merged keyword set with the preset threshold to obtain a comparison result, until the comparison result indicates that the information amount values of all the keywords in the merged keyword set are greater than or equal to the preset threshold, so as to obtain all the keywords corresponding to the dialogue text to be recognized. In this way, when the information provided in the dialogue text to be recognized is insufficient, the keyword information amount can be enhanced by obtaining the multi-round dialogue text related to the dialogue text to be recognized, thereby laying a good data foundation for improving the intention recognition accuracy in the later stage.

[0049] For the convenience of those skilled in the art to understand, the following provides a set of best embodiments: Intent recognition is an important task in natural language processing (NLP). It aims to determine the intent or purpose expressed in the user's input statement. Simply put, intent recognition is to semantically understand the user's words in order to better answer the user's questions or provide relevant services. Traditional intent recognition methods are generally based on template matching or artificial feature sets, which are time-consuming, laborious, and not very scalable. If traditional neural network models such as convolutional neural network (CNN) are used for intelligent recognition, there is a lack of context information, the accuracy of intent recognition is low, and a large amount of dialogue text information is required to improve the accuracy. Therefore, the existing related intent recognition technologies have a relatively low accuracy in recognizing the intent of dialogue texts with relatively little information.

[0050] If there are multiple rounds of conversations, and intent recognition of each round of conversation is performed, this will waste some more time. And in the initial conversation, due to relatively few key information, the intent recognition of the conversation may be incorrect, thus affecting the intent recognition of subsequent conversations. Therefore, in this embodiment, it will be more accurate to first determine whether the keyword information meets the preset conditions before performing intent recognition, and it also saves some time. Retain the conversation text and intent recognition results that have undergone intent recognition as the historical conversation text dataset, and classify the historical conversation text dataset. When there is a conversation text with the same or similar problem next time, the intent recognition result of the conversation can be obtained faster and more accurately, which provides a data basis for the next conversation, reduces the search time, and improves the accuracy and efficiency of intent recognition.

[0051] The technical solution of this embodiment specifically includes the following steps: Step S1: Obtain a conversation text training dataset (including multiple training conversation texts and their corresponding second intent labels) and the conversation text to be recognized, extract the keywords in the conversation text to be recognized, and determine whether the information volume of all keywords corresponding to the conversation text to be recognized meets the preset conditions.

[0052] Cluster the conversation texts in the conversation text training dataset according to the intent to obtain a clustering result (i.e., the second clustering result), that is, this clustering result includes various types of conversation texts, and the intents of the conversation texts in the same category are relatively similar. Well-known algorithms such as the k-means algorithm and the DBSCAN algorithm in the art can be used, and this embodiment does not make specific descriptions and restrictions.

[0053] Grade the intent of each category of data in the conversation text training dataset according to the importance of the intent to construct an intent grading label. For example, for cases such as climate - time - location, book category - theme word - role - plot, where climate is the 1st-level intent, time is the 2nd-level intent, and location is the 3rd-level intent, perform intent grading to construct the intent level of each type of conversation text. The finer the intent level is divided, the better.

[0054] Extract multiple first keywords from the dialogue text to be recognized. Determine which category of dialogue text the dialogue text to be recognized belongs to based on all the first keywords corresponding to the dialogue text to be recognized, and select the intention level of the dialogue text corresponding to the category (i.e., the target intention classification result); Calculate the information quantity values of all the first keywords according to the intention level of the dialogue text corresponding to the category: ; Among them, represents the information quantity values of all the first keywords, represents the correlation between the first keyword i and the intention level j. The correlation of the first keyword belonging to the intention level is 1, and the correlation of the first keyword not belonging to the intention level is 0. k represents the number of multiple keywords belonging to the same intention level, n represents the number of intention levels corresponding to the dialogue text of the current category, and m represents the number of first keywords.

[0055] Compare the information quantity value of the keyword with the preset condition to obtain the first comparison result.

[0056] The preset conditions corresponding to different categories of dialogue texts are different and can be changed according to the actual situation. This embodiment does not make specific restrictions.

[0057] Step S2: If the first comparison result indicates that the information quantity of the keyword reaches the preset condition (reaching the preset condition means that the information quantity of the keyword is greater than or equal to the preset threshold), then directly jump to step S3. If the first comparison result indicates that the information quantity of the keyword does not reach the preset condition (not reaching the preset condition means that the information quantity of the keyword is less than the preset threshold), then obtain more key information in the dialogue mode.

[0058] Specifically, to obtain a multi-turn dialogue model, large language models such as LLM and ChatGPT models can be used. By classifying the intention of each category of data in the dialogue text training dataset according to the importance of the intention, intention classification labels are marked for the dialogue text corresponding to each level. The multi-turn dialogue model is trained with the dialogue text training dataset marked with intention classification labels to obtain a trained multi-turn dialogue model.

[0059] Input the dialogue text to be recognized into the trained multi-turn dialogue model to enable a human-machine dialogue related to the dialogue text to be recognized, so as to obtain a multi-turn dialogue text related to the dialogue text to be recognized.

[0060] Then extract the second keywords from the multi-turn dialogue text, and merge all the second keywords and all the first keywords to obtain a merged keyword set; Calculate the information value of the merged keyword set according to the intention level of the dialogue text corresponding to the category of the dialogue text to be recognized: ; Among them, represents the information value of the merged keyword set, represents the relevance between keyword i in the merged keyword set and intention level j. The relevance of a keyword belonging to an intention level is 1, and the relevance of a keyword not belonging to an intention level is 0. k represents the number of keywords belonging to the same intention level, n represents the number of intention levels corresponding to the dialogue text of the current category, and p represents the number of keywords in the merged keyword set.

[0061] Compare the information value of the keyword with the preset condition to obtain the second comparison result. If the second comparison result indicates that the information value of the merged keyword set has not reached the preset condition, continue to obtain more key information in the dialogue mode until the second comparison result indicates that the information of the merged keyword set reaches the preset condition, and then all the keywords corresponding to the dialogue text to be recognized are obtained.

[0062] Step S3: Train the constructed intention recognition model with the dialogue text training dataset to obtain a trained intention recognition model.

[0063] The specific training process of the intention recognition model includes: (1) Use any one of the algorithms such as the TF-IDF algorithm, the TextRank algorithm, and the KeyBERT algorithm to extract keywords from the dialogue text training dataset to obtain a keyword dataset. The algorithms such as the TF-IDF algorithm, the TextRank algorithm, and the KeyBERT algorithm are all well-known technical solutions to those skilled in the art, and this embodiment does not make specific limitations and descriptions.

[0064] (2) Screen the keyword dataset to obtain a screened dataset. That is, only retain the keywords that are relatively relevant to each dialogue text, and remove some less relevant keywords to improve the model training efficiency, and eliminate the influence of some irrelevant keywords during the model training process, thereby improving the accuracy of the model for intention recognition. The specific screening method is as follows: Represent the keywords in the keyword dataset as vectors to obtain a keyword vector dataset; Input the keyword vector dataset into the multi-head attention mechanism to obtain the weight corresponding to each keyword. Calculate the word frequency of each keyword in the corresponding dialogue text, and calculate the relevance between each keyword and the corresponding dialogue text according to the word frequency. Then, calculate the keyword importance (i.e., the importance degree of the keyword) according to the weight corresponding to each keyword and its corresponding relevance. Screen the keyword dataset according to the keyword importance, and retain the keywords that are relatively relevant to each dialogue text. The specific calculation process is as follows: ; ; Among them, S represents the keyword importance, represents the weight of the i-th keyword, n represents the number of all keywords in each dialogue text, represents the relevance between the i-th keyword and the corresponding dialogue text, k represents an adjustable parameter, represents the word frequency of the i-th keyword. Among them, the word frequency statistics of keywords can adopt the existing technologies well-known to those skilled in the art, and this embodiment does not make specific descriptions.

[0065] (3) Construct an intent recognition model.

[0066] The intent recognition model includes a word vector layer, normalization, a convolutional layer, a ReLU function, multiple semantic feature extraction modules, and a multi-layer perceptron MLP. Among them, the word vector layer can use Word2vec to vectorize all keywords. The semantic feature extraction module can use the encoder part of the Transformer model. Multiple semantic feature extraction modules can use the same network structure or different network structures. The word vector layer and the semantic feature extraction module can adopt the existing technologies well-known to those skilled in the art, and this embodiment does not make specific descriptions. The normalization in this embodiment can adopt layer normalization, and this embodiment does not make specific limitations. Refer to Figure 2 , the specific intent recognition process of the intent recognition model is as follows: Input the keyword into the intent recognition model, and vectorize the keyword through the word vector layer in the intent recognition model to obtain a keyword vector; Extract the feature vector from the keyword vector through normalization, a convolutional layer, and a ReLU function to obtain a keyword feature; Input the keyword feature into the first semantic feature extraction module for semantic feature extraction to obtain a first semantic feature extraction result; Input the first semantic feature extraction result and the keyword feature into the second semantic feature extraction module for semantic feature extraction to obtain a second semantic feature extraction result; Input the second semantic feature extraction result and the keyword feature into the third semantic feature extraction module for semantic feature extraction to obtain the third semantic feature extraction result; Perform feature fusion on the first semantic feature extraction result, the second semantic feature extraction result, and the third semantic feature extraction result to obtain the fused semantic feature; Input the fused semantic feature into a multi-layer perceptron for intention recognition to obtain the intention recognition result.

[0067] (4) Use the filtered keyword data set and its corresponding intention labels to pre-train the constructed intention recognition model to obtain a pre-trained intention recognition model.

[0068] Step S4: Obtain the historical dialogue text data set of the corresponding category through all the first keywords corresponding to the dialogue text to be recognized obtained in Step S1 or Step S2. Adjust the pre-trained intention recognition model through the historical dialogue text data set, that is, optimize the parameters of the pre-trained intention recognition model to obtain the optimized intention recognition model. Perform intention recognition on all the keywords corresponding to the dialogue text to be recognized through the optimized intention recognition model to obtain the intention recognition result. By further optimizing the parameters of the trained intention recognition model, the accuracy of intention recognition can be further improved. The longer the intention recognition model is used, the more accurate the prediction will be.

[0069] Specifically, save the user's each dialogue text (i.e., the historical dialogue text) and its corresponding dialogue text intention recognition result (i.e., the first intention label) as the historical dialogue text data set. The saved historical dialogue text data set is equivalent to the validation set used to adjust the parameters of the intention recognition model. Such a data set can change with the change of the topic, thereby enabling the intention recognition model to keep pace with the times, more accurately recognize the user's dialogue intention, and improve the user experience, while the data set in Step S1 can change or not. Cluster the historical dialogue text data set using the same clustering method as in Step S1 to obtain the historical dialogue text clustering result (i.e., the first clustering result). Calculate the similarity based on all the keywords corresponding to the dialogue text to be recognized and the historical dialogue text clustering result, and use the historical dialogue text data set of the category with the highest similarity as the historical dialogue text data set corresponding to the dialogue text to be recognized. Adjust the pre-trained intention recognition model through this historical dialogue text data set to obtain the optimized intention recognition model. Perform intention recognition on all the keywords corresponding to the dialogue text to be recognized through the optimized intention recognition model to obtain the intention recognition result. Specifically, it includes: Input all the keywords corresponding to the dialogue text to be recognized into the optimized intention recognition model, and vectorize all the keywords through the word vector layer to obtain the target keyword vector; Extract keyword features by normalizing the target keyword vector, passing it through a convolutional layer, and applying the ReLU function to obtain keyword features; Input the target keyword features into the first semantic feature extraction module for semantic feature extraction to obtain the first semantic feature extraction result; Input the first semantic feature extraction result and the keyword features into the second semantic feature extraction module for semantic feature extraction to obtain the second semantic feature extraction result; Input the second semantic feature extraction result and the keyword features into the third semantic feature extraction module for semantic feature extraction to obtain the third semantic feature extraction result; Fuse the first semantic feature extraction result, the second semantic feature extraction result, and the third semantic feature extraction result to obtain fused semantic features; Input the fused semantic features into a multi-layer perceptron for intent recognition to obtain an intent recognition result.

[0070] Refer to Figure 3 In addition, an embodiment of the present application also provides a dialogue intent recognition system, which includes a data acquisition unit 100, a data clustering unit 200, an intent classification unit 300, an information amount judgment unit 400, a historical data acquisition unit 500, a model adjustment unit 600, and an intent recognition unit 700, where: The data acquisition unit 100 is configured to acquire multiple historical dialogue texts and their corresponding first intent labels, multiple training dialogue texts and their corresponding second intent labels, and the dialogue text to be recognized; The data clustering unit 200 is configured to cluster multiple historical dialogue texts based on the first intent labels according to intent similarity to obtain a first clustering result, and cluster multiple training dialogue texts based on the second intent labels according to intent similarity to obtain a second clustering result; The intent classification unit 300 is configured to classify the historical dialogue texts of each category according to intent importance based on the second clustering result to obtain an intent classification result; The information amount judgment unit 400 is configured to extract multiple first keywords of the dialogue text to be recognized and judge whether the information amount of the multiple first keywords is greater than or equal to a preset threshold according to the intent classification result; The historical data acquisition unit 500 is configured to, if the information amount of the multiple first keywords is greater than or equal to the preset threshold, obtain multiple target dialogue texts corresponding to the multiple first keywords from the multiple historical dialogue texts according to the first clustering result; The model adjustment unit 600 is configured to adjust a pre-trained intent recognition model with the multiple target dialogue texts to obtain a tuned intent recognition model, and the pre-trained intent recognition model is trained based on the multiple training dialogue texts and their corresponding second intent labels; An intention recognition unit 700 is configured to perform intention recognition on all keywords corresponding to the dialogue text to be recognized by using the optimized intention recognition model, so as to obtain an intention recognition result.

[0071] It should be noted that since a dialogue intention recognition system in this embodiment and the above-mentioned dialogue intention recognition method are based on the same inventive concept, the corresponding content in the method embodiment is equally applicable to the system embodiment of the present application, and will not be elaborated here.

[0072] Refer to Figure 4 , an embodiment of the present application further provides an electronic device, which includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the dialogue intention recognition method described above in the present disclosure.

[0073] The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0074] The electronic device in the embodiment of the present application will be introduced in detail below.

[0075] The processor 1600 may be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solution provided by the embodiment of the present disclosure; The memory 1700 may be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 may store an operating system and other application programs. When implementing the technical solution provided by the embodiment of the present specification through software or firmware, the relevant program codes are stored in the memory 1700, and are called by the processor 1600 to execute the dialogue intention recognition method of the embodiment of the present disclosure.

[0076] The input / output interface 1800 is configured to implement information input and output; A communication interface 1900 for implementing communication interactions between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); A bus 2000 for transmitting information between various components of the device (such as a processor 1600, a memory 1700, an input / output interface 1800, and a communication interface 1900); Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.

[0077] The embodiments of the present disclosure also provide a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-mentioned dialogue intention recognition method.

[0078] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0079] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0080] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown, or combine certain steps, or different steps.

[0081] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0082] Those of ordinary skill in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.

[0083] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0084] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one)" or a similar expression thereof refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c may be single or multiple.

[0085] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other may be through some interfaces, and the indirect coupling or communication connection of devices or units may be in electrical, mechanical, or other forms.

[0086] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0087] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, can also exist as individual physical units, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0088] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs. The above has described the embodiments of this application in detail with reference to the drawings, but this application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the purpose of this application.

[0089] The above has described the embodiments of this application in detail with reference to the drawings, but this application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the purpose of this application.

Claims

1. A method for recognizing dialogue intent, characterized in that, The method includes: Obtaining a plurality of historical dialogue texts and their corresponding first intent labels, a plurality of training dialogue texts and their corresponding second intent labels, and a dialogue text to be recognized; Based on the first intent labels, clustering the plurality of historical dialogue texts according to intent similarity to obtain a first clustering result, and based on the second intent labels, clustering the plurality of training dialogue texts according to intent similarity to obtain a second clustering result; According to the second clustering result, grading the historical dialogue texts of each category according to intent importance to obtain an intent grading result; Extracting a plurality of first keywords of the dialogue text to be recognized, and judging whether the information amount of the plurality of first keywords is greater than or equal to a preset threshold according to the intent grading result; If the information amount of the plurality of first keywords is greater than or equal to the preset threshold, then according to the first clustering result, obtaining a plurality of target dialogue texts of the category corresponding to the plurality of first keywords from the plurality of historical dialogue texts; Adjusting a pre-trained intent recognition model with the plurality of target dialogue texts to obtain a tuned intent recognition model, where the pre-trained intent recognition model is trained based on the plurality of training dialogue texts and their corresponding second intent labels; Using the tuned intent recognition model to perform intent recognition on all keywords corresponding to the dialogue text to be recognized to obtain an intent recognition result.

2. The method for identifying dialogue intention according to claim 1, wherein The intent recognition model includes a word vector layer, normalization, a convolutional layer, a ReLU function, a plurality of semantic feature extraction modules, and a multi-layer perceptron. The using the tuned intent recognition model to perform intent recognition on all keywords corresponding to the dialogue text to be recognized to obtain an intent recognition result includes: Inputting all keywords corresponding to the dialogue text to be recognized into the tuned intent recognition model, and vectorizing all keywords through the word vector layer to obtain target keyword vectors; Performing feature vector extraction on the target keyword vectors through normalization, a convolutional layer, and a ReLU function to obtain keyword features; Inputting the keyword features into a first semantic feature extraction module for semantic feature extraction to obtain a first semantic feature extraction result; Inputting the first semantic feature extraction result and the keyword features into a second semantic feature extraction module for semantic feature extraction to obtain a second semantic feature extraction result; Inputting the second semantic feature extraction result and the keyword features into a third semantic feature extraction module for semantic feature extraction to obtain a third semantic feature extraction result; Performing feature fusion on the first semantic feature extraction result, the second semantic feature extraction result, and the third semantic feature extraction result to obtain a fused semantic feature; Inputting the fused semantic feature into a multi-layer perceptron for intent recognition to obtain an intent recognition result.

3. The dialogue intention recognition method according to claim 1, characterized in that The pre-trained intent recognition model is obtained by the following method: Constructing an intent recognition model including a word vector layer, normalization, a convolutional layer, a ReLU function, a plurality of semantic feature extraction modules, and a multi-layer perceptron; Extracting keywords of each training dialogue text in the plurality of training dialogue texts to obtain a keyword data set; Filter the keywords in the keyword dataset to obtain a filtered dataset, and obtain the second intent label corresponding to the filtered dataset; Use the filtered dataset and its corresponding second intent label to pre-train and construct a well-trained intent recognition model to obtain a pre-trained intent recognition model.

4. The method for identifying dialogue intention according to claim 3, wherein The filtering of the keywords in the keyword dataset to obtain a filtered dataset includes: Vectorize the keywords in the keyword dataset to obtain a keyword vector dataset; Input the keyword vector dataset into the multi-head attention mechanism to obtain the weight corresponding to each keyword; Calculate the importance degree of each keyword according to the weight corresponding to each keyword; Filter the keywords in the keyword dataset according to the importance degree of each keyword to obtain a filtered dataset.

5. The method for identifying dialogue intention according to claim 1, wherein The judging whether the information amount of the multiple first keywords is greater than or equal to a preset threshold according to the intent classification result includes: Judge the category corresponding to the to-be-recognized dialogue text in the second clustering result according to the multiple first keywords; Select a relevant intent classification result according to the category corresponding to the to-be-recognized dialogue text to obtain a target intent classification result; Calculate the information amount value of the multiple first keywords according to the target intent classification result; Compare the information amount value of the multiple first keywords with the preset threshold to judge whether the information amount of the multiple first keywords is greater than or equal to the preset threshold.

6. The method for identifying dialogue intention according to claim 5, wherein, The calculating the information amount value of the multiple first keywords according to the target intent classification result includes: ; Among them, represents the information quantity value of multiple first keywords, represents the correlation between the first keyword i and the intention level j in the target intention classification result. The correlation of the first keyword belonging to the intention level j is 1, and the correlation of the first keyword not belonging to the intention level j is 0. k represents the number of multiple keywords belonging to the same intention level, n represents the number of intention levels corresponding to the current category of dialogue text, and m represents the number of multiple first keywords.

7. The method for identifying dialogue intention according to claim 5, characterized in that, The dialogue intent recognition method further includes: If the information amount of all the first keywords corresponding to the to-be-recognized dialogue text is less than the preset threshold, input the to-be-recognized dialogue text into a trained multi-turn dialogue model to obtain a multi-turn dialogue text related to the to-be-recognized dialogue text; Extract multiple second keywords in the multi-turn dialogue text, and merge all the second keywords and all the first keywords to obtain a merged keyword set; Calculate the information amount value of all the keywords in the merged keyword set according to the target intent classification result; Compare the information amount value of all the keywords in the merged keyword set with the preset threshold to obtain a comparison result until the comparison result indicates that the information amount value of all the keywords in the merged keyword set is greater than or equal to the preset threshold to obtain all the keywords corresponding to the to-be-recognized dialogue text.

8. A dialogue intention recognition system, characterized in that, The system includes: A data acquisition unit, configured to acquire multiple historical dialogue texts and their corresponding first intent labels, multiple training dialogue texts and their corresponding second intent labels, and a to-be-recognized dialogue text; A data clustering unit, configured to cluster the multiple historical dialogue texts according to intent similarity based on the first intent label to obtain a first clustering result, and cluster the multiple training dialogue texts according to intent similarity based on the second intent label to obtain a second clustering result; An intention grading unit, configured to grade the historical dialogue texts of each category according to the importance of the intention based on the second clustering result, so as to obtain an intention grading result; An information quantity judgment unit, configured to extract a plurality of first keywords of the to-be-recognized dialogue text, and judge whether the information quantity of the plurality of first keywords is greater than or equal to a preset threshold according to the intention grading result; A historical data acquisition unit, configured to, if the information quantity of the plurality of first keywords is greater than or equal to the preset threshold, acquire a plurality of target dialogue texts corresponding to the plurality of first keywords from the plurality of historical dialogue texts according to the first clustering result; A model adjustment unit, configured to adjust a pre-trained intention recognition model by using the plurality of target dialogue texts to obtain a tuned intention recognition model, where the pre-trained intention recognition model is trained based on the plurality of training dialogue texts and their corresponding second intention labels; An intention recognition unit, configured to perform intention recognition on all keywords corresponding to the to-be-recognized dialogue text by using the tuned intention recognition model to obtain an intention recognition result.

9. An electronic device, characterized in that, It includes at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the dialogue intention recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the dialogue intention recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Enhancement method, system and equipment for intention recognition training sample data and medium

    CN114548313A

  • Intention recognition method and device, computer equipment and computer readable storage medium

    CN114678014A

  • Multi-level semantic intention recognition method and related equipment thereof

    CN115730597A

  • Intent recognition method and related device

    WO2024114186A1

Cited By

  • Dynamic intention understanding method based on multi-modal fusion

    CN120579009A

  • Voice-based user intention recognition method and device, equipment and medium

    CN121011179A

  • Voice-based user intent recognition methods, devices, equipment, and media

    CN121011179B

  • Conversation service method and system for user intention recognition in social interaction

    CN121117152A