A multi-stage intent recognition method and system with adaptive context learning

CN121233768BActive Publication Date: 2026-09-11SHENZHEN KUAIBEN ZHIYING TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511197777.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2026-09-11
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

[0002]随着生成式大语言模型的兴起,通过大模型进行零样本与少样本上下文学习,但是其调用代价高昂,延迟性高,实时性受限,并且难以规模化部署,针对于对话任务场景,意图识别往往是第一步,其推理输出结果需要实时,准确和聚焦,才能为下游任务提供准确有效的依据,否则会影响对话任务的整体时效性,导致中断用户体验差;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121233768B_ABST
    Figure CN121233768B_ABST
Patent Text Reader

Abstract

The application provides a multi-stage intention recognition method and system of adaptive context learning, comprising: performing semantic coding on an input sentence, processing the coding result, and outputting a first-stage intention prediction; evaluating the uncertainty of the first-stage intention prediction based on a Monte Carlo dropout method, and determining an applicable intention recognition route based on the relative size relationship between the uncertainty and a set threshold; retrieving top-k samples similar in semantics to the current input sentence in a vector semantic space to construct dynamic prompt words; identifying and determining the preset intention range of the input sentence based on a large language model; when within the preset intention range, reasoning based on the applicable intention recognition route according to the prompt words, context examples and thought chains, and analyzing the reasoning result based on a large language model decision tool to extract a top-p intention recognition candidate set and obtain a final intention category.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of natural language processing and artificial intelligence, and in particular to a multi-stage intent recognition method and system based on adaptive context learning. Background Technology

[0002] With the rise of generative large language models, zero-shot and few-shot context learning is carried out through large models. However, the cost of calling them is high, the latency is high, the real-time performance is limited, and it is difficult to deploy at scale. For dialogue task scenarios, intent recognition is often the first step. Its inference output needs to be real-time, accurate and focused in order to provide accurate and effective basis for downstream tasks. Otherwise, it will affect the overall timeliness of dialogue tasks, resulting in interruption and poor user experience.

[0003] CN117688947B proposes a scheme for intent recognition, feature extraction, and final answering based on a large language model. However, it requires multiple calls to different large models, which affects the response latency of downstream tasks. In particular, it is inefficient in high-concurrency scenarios, which seriously affects the user experience. In the actual implementation process, it is necessary to construct numerous and complex templates for prompt words such as "first prompt information", "second prompt information", "third prompt information", and "fourth prompt information". In actual deployment, continuous testing and optimization are required to bring out the best capabilities of the large model.

[0004] Current natural language text dialogue systems often rely on supervised training encoding models (such as BERT or Sentence-BERT) for intent recognition. However, when faced with open-domain queries and queries with multiple semantic overlaps that exceed the scope (OOS), they suffer from significant performance degradation, poor model generalization, and an inability to adapt to new intents in a timely manner.

[0005] Therefore, in order to overcome the above-mentioned shortcomings, the present invention provides a multi-stage intent recognition method and system with adaptive context learning. Summary of the Invention

[0006] This invention provides a multi-stage intent recognition method and system based on adaptive context learning. It determines the semantic encoding of the input statement and classifies the semantic encoding to achieve the first-stage intent prediction. Simultaneously, it evaluates the uncertainty of the first-stage intent prediction using the Monte Carlo dropout method, thereby determining the appropriate intent recognition route based on the uncertainty. This hybrid routing strategy structure maintains high performance while significantly reducing latency compared to pure large language model schemes, improving the overall real-time response capability of the system. Secondly, by retrieving top-k semantically similar samples from the vector semantic space, it effectively constructs dynamic prompts, facilitating intent recognition. Finally, it uses a large language model to identify and determine the input statement within a preset intent range, and then applies corresponding strategies based on the determination results to achieve the final intent recognition and determination. This method does not rely on large language model LLM fine-tuning, greatly reducing the requirements for fine-tuning training data samples and computing power. It also possesses cross-task transfer capabilities and can be seamlessly integrated into existing dialogue systems, demonstrating strong engineering feasibility and significantly improving the accuracy and reliability of multi-stage intent recognition.

[0007] This invention provides a multi-stage intent recognition method based on adaptive context learning, comprising:

[0008] Step 1: Semantically encode the input statement, process the encoding results, and output the first-stage intent prediction;

[0009] Step 2: Evaluate the uncertainty of the first-stage intent prediction based on the Monte Carlo dropout method, and determine the applicable intent recognition route based on the relative magnitude of the uncertainty and the set threshold;

[0010] Step 3: Retrieve the top-k semantically similar samples to the current input statement in the vector semantic space, and construct dynamic prompt words based on the top-k samples;

[0011] Step 4: Based on the large language model, identify and determine the pre-defined intent range of the input statement;

[0012] Step 5: When within the preset intent range, reason based on the applicable intent recognition route, prompt words, context examples, and thought chains, and parse the reasoning results based on the large language model decision tool to extract the top-p intent recognition candidate set and obtain the final intent category.

[0013] Preferably, in a multi-stage intent recognition method based on adaptive context learning, step 1, before semantically encoding the input statement and processing the encoding result to output the first-stage intent prediction, includes:

[0014] Training samples are retrieved from the sample database, and keywords are extracted from the training samples;

[0015] Randomly delete or replace keywords to generate a dataset that approximates the preset intent range. and The dataset size constraint is: Where D is the entire training sample dataset, |D| represents the size of the training sample dataset, and λ ranges from 0.1 to 0.3. This represents a newly created negative sample dataset based on the original training sample dataset D.

[0016] Construct a balanced training sample dataset:

[0017] Determine input sample pairs (x) from the constructed balanced training sample dataset. i x j And calculate the contrast loss:

[0018] L contrastive =∑ (i,j) y ij ·||f(x i )-f(x j )|| 2 +(1-y ij )·max(0,m-||f(x i )-f(x j )||) 2 ;

[0019] Where f(x) is the encoder output value, ||f(x) i )-f(x j )‖ represents sample x i and x j Euclidean distance, y ij Indicates sample x i and x j Indicators for similarity or matching; i and j represent the index values ​​of the training sample dataset, and m represents the boundary margin, with a default value of 0.5;

[0020] The encoder is conditionally trained based on contrastive loss.

[0021] Preferably, a multi-stage intent recognition method based on adaptive context learning, which conditionally trains the encoder based on contrastive loss, includes:

[0022] The encoder output is obtained based on the conditional training results, and the encoder output is used as data features to train a linear classifier.

[0023] The input statement is semantically encoded based on the encoder, and the semantically encoded input statement is linearly classified based on the trained linear classifier.

[0024] The first-stage intent prediction result is determined based on the linear classification results.

[0025] Preferably, a multi-stage intent recognition method based on adaptive context learning, which conditionally trains the encoder based on contrastive loss, includes:

[0026] The training samples are encoded sequentially based on the conditionally trained encoder to obtain the training sample encoding, and the training sample encoding is embedded into the vector semantic space.

[0027] Based on the embedding results, the storage address of each training sample encoding in the vector semantic space is determined. At the same time, the basic parameters of each training sample encoding are extracted, and the vector storage index of each training sample encoding in the vector semantic space is generated based on the storage address and basic parameters.

[0028] The obtained vector storage index is recorded and saved.

[0029] Preferably, in a multi-stage intent recognition method based on adaptive context learning, step 2 involves evaluating the uncertainty of the first-stage intent prediction using the Monte Carlo dropout method, and determining the applicable intent recognition route based on the relative magnitude of the uncertainty and a set threshold, including:

[0030] The SetFit fine-tuning model samples and predicts the output based on the user's input statement using the Monte Carlo dropout method:

[0031]

[0032] Where M represents the number of samples, and its value ranges from [5, 20]. This represents the output z of the p-th Monte Carlo dropout sampling of the SetFit fine-tuning model for the input statement X. (p) ;p represents the Monte Carlo dropout sampling sequence number;

[0033] Obtain the first-stage intention prediction results and calculate the prediction mean of the first-stage intention prediction results. Use the variance of the prediction mean as the prediction uncertainty. Specifically:

[0034]

[0035] U = Var[C(z) (p) )];

[0036] Among them, C(z) (p) The output z is predicted based on the p-th sample of the input statement by a trained linear classifier. (p) The first-stage intent prediction result obtained after processing; E[C(z) (p)Var[C(z)] represents the mean of the predicted classification of the intent prediction results in the first stage; (p) [)] represents the variance of the predicted classification of the intent prediction results in the first stage;

[0037] Variance is used as a measure of prediction uncertainty and compared with a set threshold.

[0038] If the prediction uncertainty is less than the preset threshold, the SetFit fine-tuning model is determined to be suitable for intent recognition routing; otherwise, the large language model is determined to be suitable for intent recognition routing.

[0039] Preferably, in a multi-stage intent recognition method based on adaptive context learning, step 3 involves retrieving the top-k semantically similar samples to the current input statement in the vector semantic space, and constructing dynamic prompt words based on the top-k samples, including:

[0040] Obtain the semantic encoding of the input statement, and retrieve the top-k samples that are semantically similar to the current input statement in the vector semantic space based on the semantic encoding;

[0041] Constructing prompt words based on top-k samples:

[0042]

[0043] CoTInstruction represents the thought chain instruction prompt; This represents the set of user intent and its corresponding description; user query: x → intent ? This indicates that if the user inputs a query of type x, the corresponding intent category is predicted.

[0044] Preferably, in a multi-stage intent recognition method based on adaptive context learning, step 4 involves identifying and determining the intent range of the input statement based on a large language model, including:

[0045] The semantic encoding of the user input sentence is decoded using the decoder in the large language model, and the vector representation of the last word of the user input sentence after decoding is obtained based on the decoding result. and the vector set H of the last word of the training samples under the user intent y class during the training phase. y ={h i};

[0046] Calculate the vector representation of the last word of the user input statement after it has passed through the decoder. The vector set H of the last word of the training samples under the user intent class y during the training phase. y ={h i The similarity of each training sample in} is as follows:

[0047]

[0048] in, This represents the vector representation of the last word of the user input statement x after it has passed through the decoder. Representing vectors sum vector μ y The dot product of the vectors; ||·|| denotes the norm of the corresponding vectors; H y This represents the set of vectors for the last word of the training samples under the user intent class y during the training phase; |·| corresponds to the size of the set, μ. y The average semantic vector representing the target intent category;

[0049] The calculated similarity is compared with the set similarity threshold;

[0050] If the calculated similarity is greater than the set similarity threshold, it is determined that the intent of the user's input statement exceeds the preset intent range; otherwise, it is determined that the intent of the user's input statement does not exceed the preset intent range.

[0051] Preferably, in a multi-stage intent recognition method based on adaptive context learning, step 5 involves, when within a preset intent range, reasoning based on prompts, contextual examples, and thought chains using an applicable intent recognition route, and parsing the reasoning results using a large language model decision tool to extract the top-p intent recognition candidate set, thus obtaining the final intent category, including:

[0052] When within the preset intent range, under the applicable intent recognition route, the user's input statement is inferred based on prompt words, context examples and thought chains, and the inference results are used as a coarse call candidate set.

[0053] Simultaneously, a standardized tool call instruction set is generated based on the coarse call candidate set, and a predefined tool list is loaded based on the standardized tool call instruction set;

[0054] Based on the loading results, the corresponding decision tools and target decision parameters are retrieved from the tool list, and the coarse candidate set is mapped into structured intent units based on the decision tools and target decision parameters.

[0055] The confidence level of structured intent units is determined based on the user's input statement, and the number of structured intent units with a confidence level greater than the preset confidence threshold is used as the fine recall candidate set;

[0056] Simultaneously, the application scenario of the user input statement is obtained, and the refined recall candidate set is summarized based on the application scenario to obtain the final intent recognition result corresponding to the user input statement.

[0057] Preferably, a multi-stage intent recognition method based on adaptive context learning retrieves the corresponding decision tools and target decision parameters from the tool list based on the loading result, and maps the coarse candidate set into structured intent units based on the decision tools and target decision parameters, including:

[0058] The coarse candidate set is obtained, and semantic parsing is performed on the coarse candidate set based on the large language model. Based on the semantic parsing results, multiple candidate intentions of the user input statement are determined.

[0059] Perform intent composition analysis on multiple candidate intents to identify existing composite intents, and then break down the composite intents to generate multiple independent candidate entries;

[0060] Based on multiple independent candidate items, the corresponding decision tools and target decision parameters are retrieved from the tool list, and each independent candidate item is analyzed in parallel based on the retrieval results;

[0061] Based on the parsing results, multiple intent branches corresponding to each independent candidate item are determined, and the intent is realized by combining the user's input statement to generate probabilities.

[0062] Based on probability, the structured intent units corresponding to the coarse call candidate set are obtained.

[0063] This invention provides a multi-stage intent recognition system based on adaptive context learning, comprising:

[0064] The preliminary intent determination module is used to semantically encode the input statement, process the encoding results, and output the first-stage intent prediction.

[0065] The intent recognition route determination module is used to evaluate the uncertainty of the first-stage intent prediction based on the Monte Carlo dropout method, and determine the applicable intent recognition route based on the relative magnitude of the uncertainty and a set threshold.

[0066] The prompt word construction module is used to retrieve the top-k semantically similar samples to the current input sentence in the vector semantic space, and construct dynamic prompt words based on the top-k samples;

[0067] The recognition range determination module is used to recognize and determine the preset intent range of the input statement based on the large language model;

[0068] The intent determination module is used to infer based on prompt words, context examples and thought chains when the intent is within the preset range, and to parse the inference results based on the applicable intent recognition route, extract the top-p intent recognition candidate set and obtain the final intent category.

[0069] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0070] The beneficial effects of the above technical solution are as follows: By determining the semantic encoding of the input statement and classifying the semantic encoding, the first-stage intent prediction can be determined. Simultaneously, the uncertainty of the first-stage intent prediction is evaluated using the Monte Carlo dropout method, thereby determining the appropriate intent recognition route based on the uncertainty. That is, through a hybrid routing strategy structure, while maintaining high performance, the latency is significantly lower than that of a pure large language model solution, improving the overall real-time response capability of the system. Secondly, by retrieving the top-k samples semantically similar to the current input statement from the vector semantic space, dynamic prompt words can be effectively constructed, facilitating intent recognition. Finally, a large language model is used to identify and determine the input statement within a preset intent range, enabling the final intent recognition and determination based on the determination result using appropriate strategies. This approach does not rely on large language model LLM fine-tuning, greatly reducing the requirements for fine-tuning training data samples and computing power, and possesses cross-task transfer capabilities. It can also be seamlessly integrated into existing dialogue systems, exhibiting strong engineering feasibility and significantly improving the accuracy and reliability of multi-stage intent recognition.

[0071] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.

[0072] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0073] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0074] Figure 1 This is a flowchart of a multi-stage intent recognition method based on adaptive context learning in an embodiment of the present invention;

[0075] Figure 2 This is a schematic diagram of a multi-stage intent recognition method based on adaptive context learning in an embodiment of the present invention.

[0076] Figure 3 This is a structural diagram of a multi-stage intent recognition system with adaptive context learning according to an embodiment of the present invention. Detailed Implementation

[0077] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0078] Example 1:

[0079] This embodiment provides a multi-stage intent recognition method based on adaptive context learning, such as... Figure 1 As shown, it includes:

[0080] Step 1: Semantically encode the input statement, process the encoding results, and output the first-stage intent prediction;

[0081] Step 2: Evaluate the uncertainty of the first-stage intent prediction based on the Monte Carlo dropout method, and determine the applicable intent recognition route based on the relative magnitude of the uncertainty and the set threshold;

[0082] Step 3: Retrieve the top-k semantically similar samples to the current input statement in the vector semantic space, and construct dynamic prompt words based on the top-k samples;

[0083] Step 4: Based on the large language model, identify and determine the pre-defined intent range of the input statement;

[0084] Step 5: When within the preset intent range, reason based on the applicable intent recognition route, prompt words, context examples, and thought chains, and parse the reasoning results based on the large language model decision tool to extract the top-p intent recognition candidate set and obtain the final intent category.

[0085] In this embodiment, the overall working principle diagram of a multi-stage intent recognition method based on adaptive context learning is shown below. Figure 2 As shown

[0086] In this embodiment, semantic encoding of the input statement is performed by a predefined encoder.

[0087] In this embodiment, the encoding result is processed and the first-stage intent prediction is output. This refers to classifying the user's input statement according to the semantic encoding by a constructed linear classifier to determine the user's intent direction, i.e., the first-stage intent prediction.

[0088] In this embodiment, the Monte Carlo dropout method enables the Dropout layer in the neural network multiple times (e.g., 50 times) during the inference phase. It quantifies the uncertainty by measuring the variance of the prediction results. The larger the variance, the lower the model's confidence in the input. Its function is to dynamically evaluate the reliability of the intended prediction and trigger subsequent processing routes.

[0089] In this embodiment, the uncertainty of the first-stage intent prediction is evaluated based on the Monte Carlo dropout method to assess the fuzziness of the first-stage intent prediction, thereby facilitating the determination of the specific intent recognition route used to analyze the input statement based on the evaluation results.

[0090] In this embodiment, the applicable intent recognition routing includes the SetFit fine-tuning model and the large language model.

[0091] In this embodiment, the vector semantic space is pre-constructed to store known and identifiable data samples.

[0092] In this embodiment, the top-k samples refer to the number of samples that are semantically similar to the input statement in the vector semantic space, where k is an integer.

[0093] In this embodiment, dynamic prompts refer to prompts that can guide the user's intent or determine the general direction during intent recognition.

[0094] In this embodiment, the identification and determination of the preset intent range of the input statement based on the large language model refers to calculating the similarity between the embedding vector of the last token of the current statement and the semantic class center of the intent predicted by the large language model. If the similarity is greater than the threshold, it is determined to be OOS (out of preset intent range), and if it is less than the threshold, it is determined to be within the range.

[0095] In this embodiment, extracting the top-p intent recognition candidate set means extracting the top P results from the recognition results as the final intent recognition result, where P is an integer.

[0096] The beneficial effects of the above technical solution are as follows: By determining the semantic encoding of the input statement and classifying the semantic encoding, the first-stage intent prediction can be determined. Simultaneously, the uncertainty of the first-stage intent prediction is evaluated using the Monte Carlo dropout method, thereby determining the appropriate intent recognition route based on the uncertainty. That is, through a hybrid routing strategy structure, while maintaining high performance, the latency is significantly lower than that of a pure large language model solution, improving the overall real-time response capability of the system. Secondly, by retrieving the top-k samples semantically similar to the current input statement from the vector semantic space, dynamic prompt words can be effectively constructed, facilitating intent recognition. Finally, a large language model is used to identify and determine the input statement within a preset intent range, enabling the final intent recognition and determination based on the determination result using appropriate strategies. This approach does not rely on large language model LLM fine-tuning, greatly reducing the requirements for fine-tuning training data samples and computing power, and possesses cross-task transfer capabilities. It can also be seamlessly integrated into existing dialogue systems, exhibiting strong engineering feasibility and significantly improving the accuracy and reliability of multi-stage intent recognition.

[0097] Example 2:

[0098] Based on Example 1, this example provides a multi-stage intent recognition method using adaptive context learning. Step 1 involves semantically encoding the input statement and processing the encoding results before outputting the first-stage intent prediction, including:

[0099] Training samples are retrieved from the sample database, and keywords are extracted from the training samples;

[0100] Randomly delete or replace keywords to generate a dataset that approximates the preset intent range. and The dataset size constraint is: Where D is the entire training sample dataset, |D| represents the size of the training sample dataset, and λ ranges from 0.1 to 0.3. This represents a newly created negative sample dataset based on the original training sample dataset D.

[0101] Construct a balanced training sample dataset:

[0102] Determine input sample pairs (x) from the constructed balanced training sample dataset. i x j And calculate the contrast loss:

[0103] L contrastive =∑ (i,j) y ij ·||f(x i )-f(x j )|| 2 +(1-y ij )·max(0,m-||f(x i )-f(x j )||) 2 ;

[0104] Where f(x) is the encoder output value, ||f(x) i )-f(x j )|| represents sample x i and x j Euclidean distance, y ij Indicates sample x i and x j Indicators for similarity or matching; i and j represent the index values ​​of the training sample dataset, and m represents the boundary margin, with a default value of 0.5;

[0105] The encoder is conditionally trained based on contrastive loss.

[0106] In this embodiment, keywords refer to data segments that can characterize the main idea and core content of the training samples.

[0107] In this embodiment, conditional training refers to monitoring the performance loss of the encoder during the training process, that is, stopping the training process after reaching the corresponding training standard or requirement. The training standard or requirement can be set in advance.

[0108] The beneficial effects of the above technical solution are: by retrieving training samples, the encoder can be effectively trained based on the training samples, thereby providing a reliable guarantee for effective semantic encoding of user input statements and facilitating multi-stage intent recognition.

[0109] Example 3:

[0110] Building upon Example 2, this example provides a multi-stage intent recognition method based on adaptive context learning, which conditionally trains the encoder based on contrastive loss, including:

[0111] The encoder output is obtained based on the conditional training results, and the encoder output is used as data features to train a linear classifier.

[0112] The input statement is semantically encoded based on the encoder, and the semantically encoded input statement is linearly classified based on the trained linear classifier.

[0113] The first-stage intent prediction result is determined based on the linear classification results.

[0114] In this embodiment, obtaining the encoder's output based on the conditional training results refers to the encoding results of the training samples.

[0115] The beneficial effects of the above technical solution are: by training the linear classifier based on the encoder's output, it is easier to perform linear classification of user input statements, thereby determining the general direction of the user's intent and providing direction and convenience for subsequent final intent recognition.

[0116] Example 4:

[0117] Building upon Example 2, this example provides a multi-stage intent recognition method based on adaptive context learning, which conditionally trains the encoder based on contrastive loss, including:

[0118] The training samples are encoded sequentially based on the conditionally trained encoder to obtain the training sample encoding, and the training sample encoding is embedded into the vector semantic space.

[0119] Based on the embedding results, the storage address of each training sample encoding in the vector semantic space is determined. At the same time, the basic parameters of each training sample encoding are extracted, and the vector storage index of each training sample encoding in the vector semantic space is generated based on the storage address and basic parameters.

[0120] The obtained vector storage index is recorded and saved.

[0121] In this embodiment, the basic parameters refer to the types of training sample encodings, etc.

[0122] In this embodiment, the vector storage index refers to the basis for locating the encoding of each training sample in the vector semantic space.

[0123] The beneficial effects of the above technical solution are: by storing the semantic encoding of the training samples by the encoder in the vector semantic space, reliable data support is provided for the scope of intent recognition, which makes it easier to determine whether the user's current input statement is within the preset range of consciousness, and thus provides convenience for determining the intent recognition strategy.

[0124] Example 5:

[0125] Based on Example 1, this example provides a multi-stage intent recognition method using adaptive context learning. In step 2, the uncertainty of the first-stage intent prediction is evaluated using the Monte Carlo dropout method, and an applicable intent recognition route is determined based on the relative magnitude of the uncertainty and a set threshold, including:

[0126] The SetFit fine-tuning model samples and predicts the output based on the user's input statement using the Monte Carlo dropout method:

[0127]

[0128] Where M represents the number of samples, and its value ranges from [5, 20]. This represents the output z of the p-th Monte Carlo dropout sampling of the SetFit fine-tuning model for the input statement X. (p) ;p represents the Monte Carlo dropout sampling sequence number;

[0129] Obtain the first-stage intention prediction results and calculate the prediction mean of the first-stage intention prediction results. Use the variance of the prediction mean as the prediction uncertainty. Specifically:

[0130]

[0131] U = Var[C(z) (p) )];

[0132] Among them, C(z) (p) The output z is predicted based on the p-th sample of the input statement by a trained linear classifier. (p) The first-stage intent prediction result obtained after processing; E[C(z) (p) Var[C(z)] represents the mean of the predicted classification of the intent prediction results in the first stage; (p) [)] represents the variance of the predicted classification of the intent prediction results in the first stage;

[0133] Variance is used as a measure of prediction uncertainty and compared with a set threshold.

[0134] If the prediction uncertainty is less than the preset threshold, the SetFit fine-tuning model is determined to be suitable for intent recognition routing; otherwise, the large language model is determined to be suitable for intent recognition routing.

[0135] The beneficial effects of the above technical solution are: by using the Monte Carlo dropout method to evaluate the uncertainty of the first-stage intent prediction, the uncertainty of the first-stage intent prediction can be accurately and effectively evaluated, and then the appropriate intent identification route can be determined according to the relative magnitude of the uncertainty and the set threshold, thereby ensuring the appropriateness of the intent analysis route and the reliability and accuracy of the final intent result.

[0136] Example 6:

[0137] Based on Example 1, this example provides a multi-stage intent recognition method with adaptive context learning. In step 3, the top-k samples that are semantically similar to the current input statement are retrieved in the vector semantic space, and dynamic prompt words are constructed based on the top-k samples, including:

[0138] Obtain the semantic encoding of the input statement, and retrieve the top-k samples that are semantically similar to the current input statement in the vector semantic space based on the semantic encoding;

[0139] Constructing prompt words based on top-k samples:

[0140]

[0141] CoTInstruction represents the thought chain instruction prompt; This represents the set of user intent and its corresponding description; user query: x → intent ? This indicates that if the user inputs a query of type x, the corresponding intent category is predicted.

[0142] In this embodiment, retrieving the top-k samples that are semantically similar to the current input statement in the vector semantic space based on semantic encoding refers to calculating and determining them according to a preset similarity calculation formula (such as the cosine value between vectors).

[0143] The beneficial effects of the above technical solution are: by comparing the similarity between the current input statement and the data stored in the vector semantic space, the top-k samples can be retrieved based on the similarity, thereby enabling the effective construction of dynamic prompt words based on the retrieved top-k samples, which facilitates intent recognition.

[0144] Example 7:

[0145] Based on Example 1, this example provides a multi-stage intent recognition method using adaptive context learning. In step 4, the input statement is identified and judged within a preset intent range based on a large language model, including:

[0146] The semantic encoding of the user input sentence is decoded using the decoder in the large language model, and the vector representation of the last word of the user input sentence after decoding is obtained based on the decoding result. and the vector set H of the last word of the training samples under the user intent y class during the training phase. y ={h i};

[0147] Calculate the vector representation of the last word of the user input statement after it has passed through the decoder. The vector set H of the last word of the training samples under the user intent class y during the training phase. y ={h i The similarity of each training sample in} is as follows:

[0148]

[0149] in, This represents the vector representation of the last word of the user input statement x after it has passed through the decoder. Representing vectors sum vector μ y The dot product of the vectors; ||·|| denotes the norm of the corresponding vectors; H y This represents the set of vectors for the last word of the training samples under the user intent class y during the training phase; |·| corresponds to the size of the set, μ. y The average semantic vector representing the target intent category;

[0150] The calculated similarity is compared with the set similarity threshold;

[0151] If the calculated similarity is greater than the set similarity threshold, it is determined that the intent of the user's input statement exceeds the preset intent range; otherwise, it is determined that the intent of the user's input statement does not exceed the preset intent range.

[0152] The beneficial effects of the above technical solution are: by recognizing the preset intent range of the input statement through a large language model, the range of the user's current query request can be accurately and effectively determined, which facilitates the adoption of corresponding recognition strategies for intent recognition and ensures the reliability of user intent recognition.

[0153] Example 8:

[0154] Based on Example 1, this example provides a multi-stage intent recognition method with adaptive context learning. In step 5, when the intent is within a preset range, reasoning is performed based on the applicable intent recognition route according to prompt words, context examples, and thought chains. The reasoning results are then analyzed using a large language model decision tool to extract the top-p intent recognition candidate set and obtain the final intent category, including:

[0155] When within the preset intent range, under the applicable intent recognition route, the user's input statement is inferred based on prompt words, context examples and thought chains, and the inference results are used as a coarse call candidate set.

[0156] Simultaneously, a standardized tool call instruction set is generated based on the coarse call candidate set, and a predefined tool list is loaded based on the standardized tool call instruction set;

[0157] Based on the loading results, the corresponding decision tools and target decision parameters are retrieved from the tool list, and the coarse candidate set is mapped into structured intent units based on the decision tools and target decision parameters.

[0158] The confidence level of structured intent units is determined based on the user's input statement, and the number of structured intent units with a confidence level greater than the preset confidence threshold is used as the fine recall candidate set;

[0159] Simultaneously, the application scenario of the user input statement is obtained, and the refined recall candidate set is summarized based on the application scenario to obtain the final intent recognition result corresponding to the user input statement.

[0160] In this embodiment, the standardized tool call instruction set is determined based on the coarse call candidate set and is used to call the tools required to analyze the contents of each coarse call candidate set.

[0161] In this embodiment, the coarse candidate set refers to the approximate range and purpose determined after analyzing the user's input statement based on prompts, contextual examples, and thought chains, which requires further detailed analysis.

[0162] In this embodiment, the target decision parameters are pre-set, specifically the standards and requirements for analyzing the coarse candidate set.

[0163] In this embodiment, the structured intent unit refers to the probability of each intent corresponding to each content in the coarse recall candidate set after detailed analysis of the coarse recall candidate set.

[0164] In this embodiment, the preset confidence threshold is set in advance.

[0165] In this embodiment, the application scenario refers to the specific query situation targeted by the current user's input statement, such as querying business inquiries.

[0166] The beneficial effects of the above technical solution are: by reasoning based on prompt words, contextual examples and thought chains, the coarse candidate set can be effectively determined, and then decision-making tools and target decision parameters can be loaded based on the coarse candidate set to further analyze the coarse candidate set, and the final intent recognition result can be accurately and effectively determined by combining confidence.

[0167] Example 9:

[0168] Building upon Example 8, this example provides a multi-stage intent recognition method based on adaptive context learning. It retrieves the corresponding decision tools and target decision parameters from the tool list based on the loading results, and maps the coarse candidate set into structured intent units based on the decision tools and target decision parameters, including:

[0169] The coarse candidate set is obtained, and semantic parsing is performed on the coarse candidate set based on the large language model. Based on the semantic parsing results, multiple candidate intentions of the user input statement are determined.

[0170] Perform intent composition analysis on multiple candidate intents to identify existing composite intents, and then break down the composite intents to generate multiple independent candidate entries;

[0171] Based on multiple independent candidate items, the corresponding decision tools and target decision parameters are retrieved from the tool list, and each independent candidate item is analyzed in parallel based on the retrieval results;

[0172] Based on the parsing results, multiple intent branches corresponding to each independent candidate item are determined, and the intent is realized by combining the user's input statement to generate probabilities.

[0173] Based on probability, the structured intent units corresponding to the coarse call candidate set are obtained.

[0174] In this embodiment, a composite intent refers to a single intent among candidate intents that contains multiple intents in different directions, such as "searching for travel routes and identifying cost-effective hotels along those routes".

[0175] In this embodiment, an independent candidate entry refers to the result of breaking down a composite intent into individual intents.

[0176] In this embodiment, the probabilistic realization intent is used to characterize the probability of the corresponding intent obtained by parsing the user's input statement in different directions.

[0177] The beneficial effects of the above technical solution are as follows: by performing semantic parsing on the coarse candidate set, the existing composite intents can be accurately and effectively determined. When composite intents exist, they are split into multiple independent candidate entries. Then, based on the independent candidate entries, the corresponding decision tools and target decision parameters are retrieved to simultaneously parse each independent candidate entry. Finally, the multiple intent branches corresponding to each independent candidate entry are determined, thereby generating probabilistic intents from the user's input statements and providing a reference for determining the final intent recognition result of the user.

[0178] Example 10:

[0179] This embodiment provides a multi-stage intent recognition system based on adaptive context learning, such as... Figure 3 As shown, it includes:

[0180] The preliminary intent determination module is used to semantically encode the input statement, process the encoding results, and output the first-stage intent prediction.

[0181] The intent recognition route determination module is used to evaluate the uncertainty of the first-stage intent prediction based on the Monte Carlo dropout method, and determine the applicable intent recognition route based on the relative magnitude of the uncertainty and a set threshold.

[0182] The prompt word construction module is used to retrieve the top-k semantically similar samples to the current input sentence in the vector semantic space, and construct dynamic prompt words based on the top-k samples;

[0183] The recognition range determination module is used to recognize and determine the preset intent range of the input statement based on the large language model;

[0184] The intent determination module is used to infer based on prompt words, context examples and thought chains when the intent is within the preset range, and to parse the inference results based on the applicable intent recognition route, extract the top-p intent recognition candidate set and obtain the final intent category.

[0185] The beneficial effects of the above technical solution are as follows: By determining the semantic encoding of the input statement and classifying the semantic encoding, the first-stage intent prediction can be determined. Simultaneously, the uncertainty of the first-stage intent prediction is evaluated using the Monte Carlo dropout method, thereby determining the appropriate intent recognition route based on the uncertainty. That is, through a hybrid routing strategy structure, while maintaining high performance, the latency is significantly lower than that of a pure large language model solution, improving the overall real-time response capability of the system. Secondly, by retrieving the top-k samples semantically similar to the current input statement from the vector semantic space, dynamic prompt words can be effectively constructed, facilitating intent recognition. Finally, a large language model is used to identify and determine the input statement within a preset intent range, enabling the final intent recognition and determination based on the determination result using appropriate strategies. This approach does not rely on large language model LLM fine-tuning, greatly reducing the requirements for fine-tuning training data samples and computing power, and possesses cross-task transfer capabilities. It can also be seamlessly integrated into existing dialogue systems, exhibiting strong engineering feasibility and significantly improving the accuracy and reliability of multi-stage intent recognition.

[0186] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A multi-stage intent recognition method based on adaptive context learning, characterized in that, include: Step 1: Semantically encode the input statement, process the encoding results, and output the first-stage intent prediction; Step 2: Evaluate the uncertainty of the first-stage intent prediction based on the Monte Carlo dropout method, and determine the applicable intent recognition route based on the relative magnitude of the uncertainty and the set threshold; Step 3: Retrieve the top-k semantically similar samples to the current input statement in the vector semantic space, and construct dynamic prompt words based on the top-k samples; Step 4: Based on the large language model, identify and determine the pre-defined intent range of the input statement; Step 5: When within the preset intent range, reason based on the applicable intent recognition route, prompt words, context examples, and thought chains, and parse the reasoning results based on the large language model decision tool to extract the top-p intent recognition candidate set and obtain the final intent category; In step 5, when the intent is within a preset range, reasoning is performed based on the applicable intent recognition route, prompts, contextual examples, and thought chains. The reasoning results are then analyzed using a large language model decision tool to extract the top-p intent recognition candidate set, resulting in the final intent category, including: When within the preset intent range, under the applicable intent recognition route, the user's input statement is inferred based on prompt words, context examples and thought chains, and the inference results are used as a coarse call candidate set. Simultaneously, a standardized tool call instruction set is generated based on the coarse call candidate set, and a predefined tool list is loaded based on the standardized tool call instruction set; Based on the loading results, the corresponding decision tools and target decision parameters are retrieved from the tool list, and the coarse candidate set is mapped into structured intent units based on the decision tools and target decision parameters. The confidence level of structured intent units is determined based on the user's input statement, and the number of structured intent units with a confidence level greater than the preset confidence threshold is used as the fine recall candidate set; Simultaneously, the application scenario of the user input statement is obtained, and the refined recall candidate set is summarized based on the application scenario to obtain the final intent recognition result corresponding to the user input statement.

2. The multi-stage intent recognition method based on adaptive context learning according to claim 1, characterized in that, In step 1, the input statement is semantically encoded, and the encoding result is processed before the first stage of intent prediction is output, including: Training samples are retrieved from the sample database, and keywords are extracted from the training samples; Randomly delete or replace keywords to generate a dataset that approximates the preset intent range. ,and The dataset size constraint is: ,in, For the entire training sample dataset, Indicates the size of the training sample dataset. The value range is 0.1-0.

3. This represents a newly created negative sample dataset based on the original training sample dataset D. Construct a balanced training sample dataset: ; Determine input sample pairs from the constructed balanced training sample dataset. And calculate the contrast loss: ; in, For encoder output value, Indicates sample and European distance, Indicates sample and Indicators indicating similarity or whether a match exists; and This represents the index value of the training sample dataset. Indicates the boundary interval; the default value is 0.

5. The encoder is conditionally trained based on contrastive loss.

3. The multi-stage intent recognition method based on adaptive context learning according to claim 2, characterized in that, Conditional training of the encoder based on contrastive loss includes: The encoder output is obtained based on the conditional training results, and the encoder output is used as data features to train a linear classifier. The input statement is semantically encoded based on the encoder, and the semantically encoded input statement is linearly classified based on the trained linear classifier. The first-stage intent prediction result is determined based on the linear classification results.

4. The multi-stage intent recognition method based on adaptive context learning according to claim 2, characterized in that, Conditional training of the encoder based on contrastive loss includes: The training samples are encoded sequentially based on the conditionally trained encoder to obtain the training sample encoding, and the training sample encoding is embedded into the vector semantic space. Based on the embedding results, the storage address of each training sample encoding in the vector semantic space is determined. At the same time, the basic parameters of each training sample encoding are extracted, and the vector storage index of each training sample encoding in the vector semantic space is generated based on the storage address and basic parameters. The obtained vector storage index is recorded and saved.

5. The multi-stage intent recognition method based on adaptive context learning according to claim 1, characterized in that, In step 2, the uncertainty of the first-stage intent prediction is evaluated based on the Monte Carlo dropout method, and the applicable intent recognition route is determined based on the relative magnitude of the uncertainty and a set threshold, including: The SetFit fine-tuning model samples and predicts the output based on the user's input statement using the Monte Carlo dropout method: ; in, This indicates the number of samples, with a value range of [5, 20]. This indicates that for input statement X, the SetFit fine-tuning model has the following parameters: Results of sub-Monte Carlo dropout sampling ; This represents the sampling sequence number value of the Monte Carlo dropout method; Obtain the first-stage intention prediction results and calculate the prediction mean of the first-stage intention prediction results. Use the variance of the prediction mean as the prediction uncertainty. Specifically: ; ; in, To perform a training-based linear classifier on the first... The input statement is sampled and predicted to output. The first-stage intent prediction results obtained after processing; This represents the mean of the predicted categories based on the intent prediction results of the first stage. This represents the variance of the predicted classification based on the intention prediction results of the first stage. Variance is used as a measure of prediction uncertainty and compared with a set threshold. If the prediction uncertainty is less than the preset threshold, the SetFit fine-tuning model is determined to be suitable for intent recognition routing; otherwise, the large language model is determined to be suitable for intent recognition routing.

6. The multi-stage intent recognition method based on adaptive context learning according to claim 1, characterized in that, In step 3, the top-k samples that are semantically similar to the current input statement are retrieved in the vector semantic space, and dynamic prompt words are constructed based on the top-k samples, including: Obtain the semantic encoding of the input statement, and retrieve the top-k samples that are semantically similar to the current input statement in the vector semantic space based on the semantic encoding; Constructing prompt words based on top-k samples: ; in, Indicates the prompt words for thought chain instructions; express A set of user intents and their corresponding descriptive information; This indicates that the user input query is Predict the corresponding intent category.

7. The multi-stage intent recognition method based on adaptive context learning according to claim 1, characterized in that, In step 4, the input statement is identified and judged within a preset intent range based on a large language model, including: The semantic encoding of the user input sentence is decoded using the decoder in the large language model, and the vector representation of the last word of the user input sentence after decoding is obtained based on the decoding result. and the training phase regarding user intent categories The vector set of the last word in the next training sample ; Calculate the vector representation of the last word of the user input statement after it has passed through the decoder. Compared with the training phase for user intent categories The vector set of the last word in the next training sample The similarity of each training sample is as follows: ; ; in, Indicates user input statement The vector representation of the last word after decoding; Representing vectors sum vector The dot product; This represents the norm of the corresponding vector; This indicates the training phase for user intent categories. The set of vectors of the last word in the next training sample; The corresponding set size, The average semantic vector representing the target intent category; The calculated similarity is compared with the set similarity threshold; If the calculated similarity is greater than the set similarity threshold, it is determined that the intent of the user's input statement exceeds the preset intent range; otherwise, it is determined that the intent of the user's input statement does not exceed the preset intent range.

8. The multi-stage intent recognition method based on adaptive context learning according to claim 1, characterized in that, Based on the loading results, the corresponding decision-making tools and target decision parameters are retrieved from the tool list. Then, based on the decision-making tools and target decision parameters, the coarse candidate set is mapped into structured intent units, including: The coarse candidate set is obtained, and semantic parsing is performed on the coarse candidate set based on the large language model. Based on the semantic parsing results, multiple candidate intentions of the user input statement are determined. Perform intent composition analysis on multiple candidate intents to identify existing composite intents, and then break down the composite intents to generate multiple independent candidate entries; Based on multiple independent candidate items, the corresponding decision tools and target decision parameters are retrieved from the tool list, and each independent candidate item is analyzed in parallel based on the retrieval results; Based on the parsing results, multiple intent branches corresponding to each independent candidate item are determined, and the intent is realized by combining the user's input statement to generate probabilities. Based on probability, the structured intent units corresponding to the coarse call candidate set are obtained.

9. A multi-stage intent recognition system with adaptive context learning, characterized in that, include: The preliminary intent determination module is used to semantically encode the input statement, process the encoding results, and output the first-stage intent prediction. The intent recognition route determination module is used to evaluate the uncertainty of the first-stage intent prediction based on the Monte Carlo dropout method, and determine the applicable intent recognition route based on the relative magnitude of the uncertainty and a set threshold. The prompt word construction module is used to retrieve the top-k semantically similar samples to the current input sentence in the vector semantic space, and construct dynamic prompt words based on the top-k samples; The recognition range determination module is used to recognize and determine the preset intent range of the input statement based on the large language model; The intent determination module is used to reason based on prompt words, context examples and thought chains when the intent is within the preset range, and to parse the reasoning results based on the applicable intent recognition route, extract the top-p intent recognition candidate set and obtain the final intent category based on the large language model decision tool. The intent determination module includes: When within the preset intent range, under the applicable intent recognition route, the user's input statement is inferred based on prompt words, context examples and thought chains, and the inference results are used as a coarse call candidate set. Simultaneously, a standardized tool call instruction set is generated based on the coarse call candidate set, and a predefined tool list is loaded based on the standardized tool call instruction set; Based on the loading results, the corresponding decision tools and target decision parameters are retrieved from the tool list, and the coarse candidate set is mapped into structured intent units based on the decision tools and target decision parameters. The confidence level of structured intent units is determined based on the user's input statement, and the number of structured intent units with a confidence level greater than the preset confidence threshold is used as the fine recall candidate set; Simultaneously, the application scenario of the user input statement is obtained, and the refined recall candidate set is summarized based on the application scenario to obtain the final intent recognition result corresponding to the user input statement.

Citation Information

Patent Citations

  • Large Model-Based Dialogue Processing Method, Device, Electronic Device, and Storage Medium

    CN117688947B

  • Intention recognition device, method and equipment based on hierarchical classification and storage medium

    CN111597320A

  • Intention recognition method and device, electronic equipment and storage medium

    CN114706945A

  • Speech quality average opinion score prediction method based on uncertainty perception

    CN118782102A

  • Intention recognition method, system and equipment

    CN119358563A