Intention recognition method, intention recognition device, and smart glasses

CN120783736BActive Publication Date: 2026-09-11HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511285887.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2026-09-11
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

[0005]在本实施例中提供了一种意图识别方法、意图识别装置和智能眼镜,以解决相关技术中意图识别准确度较低的问题

Benefits of technology

[0030]Compared with related technologies, this embodiment provides an intent recognition method, an intent recognition device, and smart glasses. The intent recognition method acquires user input information; inputs the user input information into a trained intent recognition model, enabling the trained intent recognition model to recognize the user input information's intent, and determines the tool corresponding to the recognition result from a preset toolset as the target tool, calling the target tool to execute the operation corresponding to the recognition result; wherein: the toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-constructed intent recognition samples combined with the preset toolset; intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multi-command samples; negative samples include at least speech misrecognition samples and non-user command samples. This improves the semantic discrimination capability of intent recognition, supports command interaction, and improves recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783736B_ABST
    Figure CN120783736B_ABST
Patent Text Reader

Abstract

The application relates to an intention recognition method, an intention recognition device and intelligent glasses, wherein the intention recognition method comprises the following steps: acquiring user input information; inputting the user input information into a trained intention recognition model to enable the trained intention recognition model to recognize the intention of the user input information, and determining a tool corresponding to a recognition result as a target tool from a preset tool set according to the recognition result, and calling the target tool to execute an operation corresponding to the recognition result; wherein: the trained intention recognition model is a large language model trained based on pre-constructed intention recognition samples and a preset tool set; the intention recognition samples comprise positive samples and negative samples for each intention category; different positive samples constitute multi-command samples; and the negative samples at least comprise voice misrecognition samples and non-user instruction samples. The application can improve the semantic distinguishing capability of intention recognition, support instruction interaction and improve recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction, and in particular to intent recognition methods, intent recognition devices, and smart glasses. Background Technology

[0002] Intent recognition is a key technology in Natural Language Processing (NLP), designed to identify the purpose or intent behind a user's input (such as text, voice, etc.). Intent recognition is applicable to scenarios such as intelligent customer service and dialogue systems, voice assistants, smart home control, and interaction with head-mounted displays.

[0003] Traditional intent recognition methods typically rely on rule matching, statistical machine learning, or deep learning models, requiring large amounts of labeled data and exhibiting limited generalization ability when faced with complex and varied user input. With the rapid development of artificial intelligence, Large Language Models (LLMs) have demonstrated powerful capabilities in Natural Language Processing (NLP) tasks, particularly in intent recognition. However, current applications of LLMs for intent recognition still suffer from poor semantic discrimination and intent recognition bias, leading to low accuracy.

[0004] There is currently no effective solution to the problem of low accuracy in intent recognition in related technologies. Summary of the Invention

[0005] This embodiment provides an intent recognition method, an intent recognition device, and smart glasses to address the problem of low intent recognition accuracy in related technologies.

[0006] Firstly, this embodiment provides an intent recognition method, including:

[0007] Obtain user input information;

[0008] The user input information is input into a trained intent recognition model, which then performs intent recognition on the user input information. Based on the recognition result, the model selects a tool from a preset toolset that corresponds to the recognition result as the target tool and calls the target tool to perform the operation corresponding to the recognition result. The toolset includes several tools, and different tools are functional components that perform different operations.

[0009] The trained intent recognition model is a large language model trained based on pre-constructed intent recognition samples combined with the preset toolset; the intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multiple command samples; the negative samples include at least speech misrecognition samples and non-user command samples.

[0010] In some embodiments, the process of constructing the speech misidentification samples includes:

[0011] Acquire preset voice data;

[0012] The voice data is input into a preset speech recognition model for text conversion, and the output of the speech recognition model is obtained.

[0013] The characters whose output probabilities satisfy the target conditions in the speech recognition model output are taken as target characters; the target conditions are that the output probability ranking is lower than the highest ranking preset position, and the difference between the output probability of the highest ranking and the highest ranking is less than a preset threshold.

[0014] The speech misrecognition sample is obtained by replacing the real text corresponding to the speech data with the target character.

[0015] In some embodiments, the process of constructing the non-user instruction sample includes:

[0016] Collect the response voice of the large language model to the user's voice in a preset voice interaction scenario;

[0017] The response voice is converted into text data to obtain the non-user command sample.

[0018] In some embodiments, the intent recognition samples further include hierarchical samples; the construction process of the hierarchical samples includes:

[0019] Based on the hierarchical rules of the intent categories, determine the parent intent categories and child intent categories in the intent categories;

[0020] Based on the hierarchical rules and preset conflict rules, multi-level prompt words are set for sample classification for the parent intent category and the child intent category;

[0021] Based on the multi-level prompt words, the intent recognition samples are hierarchically optimized using a large language model to obtain hierarchical samples.

[0022] In some embodiments, the training process of the intent recognition model includes: training a preset large language model using the intent recognition samples and the preset toolset as training data.

[0023] In some embodiments, the training process of the intent recognition model further includes: performing reinforcement learning training on the large language model based on preference samples; the preference samples are intent recognition samples that have been pre-predicted by different preset large language models and whose predicted categories are inconsistent with the actual intent categories.

[0024] In some embodiments, when performing intent recognition inference based on the trained intent recognition model, if a new tool is added to the toolset, the description information of the new tool is added to the input data of the trained intent recognition model so that the intent recognition model can determine the target tool from the updated toolset.

[0025] In some of these embodiments, during the inference process of the trained intent recognition model, after the toolset is first converted into a vector representation, the vector representation is cached based on an attention key-value caching mechanism.

[0026] Secondly, this embodiment provides an intent recognition device, including an acquisition module and an intent recognition module; wherein:

[0027] The acquisition module is used to acquire user input information;

[0028] The intent recognition module is used to input the user input information into the trained intent recognition model, so that the trained intent recognition model can recognize the intent of the user input information, and determine the tool corresponding to the recognition result from a preset toolset as the target tool according to the recognition result, and call the target tool to perform the operation corresponding to the recognition result; wherein: the toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-constructed intent recognition samples combined with the preset toolset; the intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multi-command samples; the negative samples include at least speech misrecognition samples and non-user command samples.

[0029] Thirdly, this embodiment provides a smart glasses that stores a computer program, which, when executed by a processor, implements the intent recognition method described in the first aspect above.

[0030] Compared with related technologies, this embodiment provides an intent recognition method, an intent recognition device, and smart glasses. The intent recognition method acquires user input information; inputs the user input information into a trained intent recognition model, enabling the trained intent recognition model to recognize the user input information's intent, and determines the tool corresponding to the recognition result from a preset toolset as the target tool, calling the target tool to execute the operation corresponding to the recognition result; wherein: the toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-constructed intent recognition samples combined with the preset toolset; intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multi-command samples; negative samples include at least speech misrecognition samples and non-user command samples. This improves the semantic discrimination capability of intent recognition, supports command interaction, and improves recognition accuracy.

[0031] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0032] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0033] Figure 1 This is an application environment diagram of an intent recognition method according to an embodiment of this application;

[0034] Figure 2 This is a flowchart of the intent recognition method according to an embodiment of this application;

[0035] Figure 3 This is a flowchart of an intent recognition method according to some embodiments of this application;

[0036] Figure 4 This is a structural block diagram of the intent recognition device according to an embodiment of this application;

[0037] Figure 5 This is an internal structural diagram of a computer device according to an embodiment of this application. Detailed Implementation

[0038] To better understand the purpose, technical solution, and advantages of this application, the application is described and explained below in conjunction with the accompanying drawings and embodiments.

[0039] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific order of objects.

[0040] Figure 1 This is an application environment diagram of an intent recognition method according to an embodiment of this application. The intent recognition method provided in this embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Server 104 acquires user input information; it inputs the user input information into a trained intent recognition model, enabling the trained intent recognition model to recognize the user input information and, based on the recognition result, determines the tool corresponding to the recognition result from a preset toolset as the target tool, and calls the target tool to execute the operation corresponding to the recognition result. The toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-constructed intent recognition samples combined with the preset toolset; the intent recognition samples include positive and negative samples for each intent category; different positive samples constitute multiple command samples; negative samples include at least speech misrecognition samples and non-user instruction samples. The execution control command for calling the tool is sent to terminal 102 via the communication network. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle systems. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0041] This embodiment provides an intent recognition method. Figure 2 This is a flowchart of the intent recognition method in this embodiment, as follows: Figure 2 As shown, the process includes the following steps:

[0042] Step S210: Obtain user input information. This intent recognition method is applicable to any intent recognition scenario. This intent recognition method can run on devices capable of providing computational support for large language model reasoning, such as cloud servers, edge devices, or embedded systems. The user input information can be text information directly entered by the user through the interactive interface, or text information obtained by speech recognition and conversion of the user's output speech information. For example, in a voice assistant scenario, if a user wants to lower the volume of music played on a mobile device, they issue the voice command "lower the volume." Then, the mobile device's speech recognition module recognizes this voice command and converts it into text information containing "lower the volume," which serves as the user input information.

[0043] Step S220: Input the user input information into the trained intent recognition model so that the trained intent recognition model can recognize the user input information and determine the tool corresponding to the recognition result from the preset toolset as the target tool according to the recognition result, and call the target tool to perform the operation corresponding to the recognition result; wherein: the toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-built intent recognition samples combined with the preset toolset; intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multi-command samples; negative samples include at least speech misrecognition samples and non-user command samples.

[0044] Specifically, the user input information obtained in step S210 is input into the trained intent recognition model. This intent recognition model, after training, can recognize the user input intent, associate the recognition result with the target tool to be invoked, and then execute the operation corresponding to the user's intent by invoking the target tool, such as turning down the volume. It should be noted that this toolset can include several tools for different intent categories, with one tool corresponding to one intent category. The toolset can include a tool description, parameter descriptions, and invocation method (call). After the intent recognition model matches the corresponding target tool through the tool description and parameter description, it then invokes the target tool's call to perform the operation.

[0045] Here, intent categories refer to the types of user intents; tools are pre-packaged functional components that can perform the operations corresponding to the intent categories. For example, the corresponding implementation logic for various preset intent categories can be pre-written in a computer program and packaged into tools for the intent recognition model to call. For instance, under the intent "volume adjustment", it can include intent categories such as "volume increase", "volume decrease", "mute", and "unmute", which can correspond to volume increase tools, volume decrease tools, mute tools, and unmute tools, respectively.

[0046] Before training the intent recognition model, positive and negative samples can be created for different intent categories to train the intent recognition capability of the large language model. Both positive and negative samples are text. Positive samples are text that is semantically consistent with the corresponding intent category, while negative samples are related to the corresponding intent category but do not conform to the description of that intent category or are semantically inconsistent with it. For example, for the intent category "volume adjustment," samples such as "volume increase," "volume decrease," "mute," and "unmute" are all positive samples, while "other volumes" can be negative samples. These four types of positive samples and one type of negative sample can collectively constitute the first-level type of the intent category "volume adjustment."

[0047] Specifically, this can be achieved by writing prompt words and using a large language model to match and generate the aforementioned positive and negative samples. For each intent category, positive and negative descriptions for sample generation are written, along with seed examples. For instance, for the intent category "mute," a positive seed example could be "turn off the sound," and a negative seed example could be "there is no sound." Providing seed examples as a small sample set in the prompt words facilitates the generation of positive and negative samples. Then, prompt words for sample generation (i.e., sample generation prompt words) are written based on the positive and negative descriptions and seed examples. These sample generation prompt words can include positive and negative descriptions of intent categories, a small number of seed examples, conflict handling for similar intent categories, and sample generation strategies.

[0048] Taking the intent category "mute" as an example, the corresponding sample-generated prompt words can take the following form:

[0049] Please generate 200 samples that contain the intent to mute.

[0050] #Silence refers to the sound of turning off the device.

[0051] #The following is an example of the "mute" intention:

[0052] Turn off the sound -> {"name":"Mute","parameters":{}}

[0053] #Type Conflict

[0054] Note the distinction between this and 'unmute, adjust volume' and other sound-related commands.

[0055] #Generation Strategy

[0056] Generate 100 simple instructions with fewer than 8 characters.

[0057] 100 instructions with a character count between 8 and 15 are generated, each containing the usage context.

[0058] Based on the generated prompts, positive and negative samples are generated using a large language model. Then, the generated positive and negative samples can be deduplicated. Specifically, text similarity detection can be used to remove duplicate or highly similar samples, avoiding sample redundancy. For example, the edit distance between two samples can be calculated, and samples with an edit distance less than 5% of the overall length, or less than 2, can be deduplicated. This edit distance refers to the minimum number of editing operations (such as adding, replacing, or deleting characters) required to transform one sample into another. The specific deduplication process described here is merely an example and does not constitute a limitation on the deduplication process in this step. Based on this, multi-command samples can be constructed. Specifically, multiple generated positive samples can be concatenated with commas as input, and the corresponding intent category can be used as output to form multi-command samples.

[0059] In the negative sample section, this step also includes speech misrecognition samples and non-user command samples. Speech misrecognition samples are added to address errors in the conversion of user voice commands to text in real-world scenarios. Non-user command samples are added to address the possibility that voice input in real-world scenarios may also contain model response speech that hasn't undergone echo cancellation; these are the text of the model's voice response, not the user command. Speech misrecognition samples aim to eliminate the impact of speech recognition errors on intent recognition, while non-user command samples aim to eliminate false recall caused by model response speech input during user conversations.

[0060] In related technologies, traditional intent recognition methods typically rely on rule matching, statistical machine learning, or deep learning models, requiring large amounts of labeled data and exhibiting limited generalization ability when faced with complex and varied user input. With the rapid development of artificial intelligence, Large Language Models (LLMs) have demonstrated powerful capabilities in Natural Language Processing (NLP) tasks, especially in the field of intent recognition. Currently, however, when applying LLMs for intent recognition, issues remain, including poor semantic discrimination and intent recognition bias, leading to low accuracy.

[0061] This embodiment, through steps S210 to S220, enables comprehensive and rich sample generation for intent recognition during the model training phase. Specifically, it constructs negative samples to address speech misrecognition and model speech response issues in voice interaction scenarios, and builds multi-command samples to meet the need for simultaneous issuance of multiple commands in real-world applications. Therefore, this embodiment enhances the semantic discrimination ability of the trained intent recognition model, supports command interaction, and improves recognition accuracy.

[0062] Therefore, through the above steps S210 to S220, user input information is obtained; the user input information is input into the trained intent recognition model, so that the trained intent recognition model can perform intent recognition on the user input information, and determine the tool corresponding to the recognition result from the preset toolset as the target tool according to the recognition result, and call the target tool to perform the operation corresponding to the recognition result; wherein: the toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-constructed intent recognition samples combined with the preset toolset; intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multi-command samples; negative samples include at least speech misrecognition samples and non-user command samples. It can improve the semantic discrimination ability of intent recognition, support command interaction, and improve recognition accuracy.

[0063] In one embodiment, the process of constructing speech misidentification samples may specifically include:

[0064] Acquire preset speech data; input the speech data into a preset speech recognition model for text conversion to obtain the speech recognition model output; take the characters whose output probabilities satisfy the target conditions in the speech recognition model output as target characters; the target conditions are that the output probability ranking is lower than the preset ranking of the highest ranking, and the difference between the output probability of the highest ranking and the target character is less than the preset threshold; replace the real text corresponding to the speech data according to the target characters to obtain speech misrecognition samples.

[0065] In voice interaction scenarios, user voice commands must first be converted into text before intent recognition is performed based on the converted text. Errors can occur during the voice-to-text conversion process. To address this, this embodiment incorporates voice misrecognition samples into the intent recognition samples to mitigate these errors.

[0066] Specifically, the batch-generated samples are converted into speech data using a Text-to-Speech (TTS) model, and then recognized as text again using an Automatic Speech Recognition (ASR) model. If the original text can be restored, the corresponding speech is retained; otherwise, it is discarded. When the ASR model outputs each character corresponding to the input speech data, the character whose output probability ranks in a preset position (e.g., the second highest) and whose difference from the highest probability is less than a preset threshold (e.g., 0.1) is retained as the target character. The original character at the corresponding position is replaced with the target character, thus obtaining a speech misrecognition sample (ASR error sample). In this way, one to three error characters can be retained in each sample.

[0067] In this embodiment, speech misidentification samples are constructed by combining text-to-speech and speech recognition methods. These samples are then used as negative samples to train the intent recognition model. This ensures that the final trained intent recognition model can handle speech conversion errors in real-world scenarios and accurately distinguish and recognize intents even when errors exist, thereby improving the robustness of intent recognition.

[0068] In another embodiment, the process of constructing a non-user instruction sample may specifically include:

[0069] In a preset voice interaction scenario, the large language model collects the response voice of the user's voice; the response voice is converted into text data to obtain non-user command samples.

[0070] In this embodiment, considering that in a voice interaction scenario, the received voice input includes not only the user's own voice commands but also, without echo cancellation, the voice information from the model's responses to the user. Therefore, in this embodiment, the batch-generated samples are used as input, and the output information of the large language model is used as negative samples of non-user commands. By adding negative samples of non-user commands to train the intent recognition model, the trained intent recognition model can remove erroneous recalls caused by non-command voice input in the voice interaction scenario during intent recognition. Therefore, this embodiment can improve the accuracy and robustness of intent recognition even when interference from model responses exists.

[0071] Furthermore, in some embodiments, additional negative samples can be included. For example, during the testing phase of the intent recognition model and the formal use phase of the intent recognition model for intent recognition, real user input is collected, and samples misidentified by the intent recognition model are used as negative samples. Misidentified samples can be verified using both sample check prompts and the fine-tuning model. Samples consistently verified as negative by both sample check prompts and the fine-tuning model, or samples whose verification results are inconsistent but include manually confirmed negative samples, are used as additional negative samples.

[0072] Specifically, the above-mentioned misidentified samples are checked using sample check prompts combined with a large language model, and the verification result of whether they are negative samples is output. A fine-tuning model is then used to check the above-mentioned misidentified samples, and the verification result of whether they are negative samples is output. By comparing the two sets of verification results, other negative samples are identified. This embodiment can further expand the richness of intent recognition samples and improve the accuracy of intent recognition.

[0073] In another embodiment, the intent recognition samples also include hierarchical samples; the construction process of hierarchical samples may specifically include: determining the parent intent category and child intent category in the intent category according to the hierarchical rules of the intent category; setting multi-level prompt words for sample classification for the parent intent category and child intent category based on the hierarchical rules and preset conflict rules; and using a large language model to hierarchically optimize the intent recognition samples based on the multi-level prompt words to obtain hierarchical samples.

[0074] Specifically, the generated prompts can be optimized by embedding decision branches and conflict rules into them, thus forming secondary classification prompts to reduce their length. This prompt concatenation method improves model inference speed. The hierarchical rule for intent categories can be to perform finer-grained classification within each category after a coarse initial classification. For example, parent intent categories might include "volume," "brightness," "app open," and "app close." Under the parent intent category of "volume," there might be child intent categories such as "mute," "unmute," "adjust volume," "increase volume," and "decrease volume." First, the samples corresponding to the parent intent category are determined based on the first-level classification prompts. Then, the samples corresponding to the child intent categories are obtained using the corresponding second-level classification prompts within that parent intent category.

[0075] The aforementioned decision branch specifically involves the logic for classifying the sample into intent categories. For example, it first checks whether the intent includes "mute" or "unmute." If not, it then checks whether the intent includes "adjust volume" or "increase / decrease volume." If still not, it is confirmed as an intent for "other volumes." The conflict rule determines which intent to classify a sample into when multiple volume intents exist simultaneously. For example, it prioritizes classifying the sample into the later intent.

[0076] The aforementioned decision branches and conflict rules are embedded into multi-level prompts used for sample classification. Combined with a large language model, these multi-level prompts are used to perform hierarchical optimization of intent recognition samples, resulting in hierarchical intent recognition samples, or tiered samples. In this embodiment, the construction of tiered samples is introduced to achieve hierarchical classification of intent categories, which can improve the model's inference speed and reduce the problem of slow inference caused by excessively long tool descriptions corresponding to intent categories when the model calls tools.

[0077] In one embodiment, the training process of the intent recognition model may specifically include: training a preset large language model using intent recognition samples and a preset toolset as training data.

[0078] This involves using tools corresponding to different intent categories as training data to train the large language model. For example, based on the parent and child intent categories mentioned above, corresponding tool sets are provided as training data to train the pre-defined large language model for intent recognition. This allows the trained intent recognition model to select tool sets for different intent categories. Additionally, samples from publicly available datasets can be combined with corresponding tools to train the large language model, further improving the semantic matching ability of the intent recognition model for tools. The intent categories in the publicly available datasets do not overlap with the intent categories corresponding to the pre-constructed samples.

[0079] In this embodiment, the intent recognition model is able to accurately match the input text with the tool that needs to be invoked, thereby using the matched tool to perform an operation corresponding to the user's intent.

[0080] In one embodiment, the training process of the intent recognition model further includes: performing reinforcement learning training on a large language model based on preference samples; the preference samples are intent recognition samples whose predicted categories are inconsistent with the actual intent categories obtained by pre-predicting categories using different preset large language models.

[0081] First, after generating samples in batches as described above, prediction prompts can be written for samples under all intent categories. These prompts include a description of each intent category, a description of the parameters required for that intent category, and a small number of input and output samples. The intent category of each generated sample is then predicted using a large language model. Specifically, the intent categories of the samples can be predicted using both the Deepseek and Qwen large language models, combined with the prompts, resulting in two sets of predicted categories. These predicted categories are then combined with the sample's actual intent category. Samples whose predicted categories do not match the actual intent category are manually verified to form preference samples. These preference samples are used for reinforcement learning training of the intent recognition model.

[0082] This method employs supervised fine-tuning (SFT), direct preference optimization (DPO), Kahneman-Tversky optimization (KTO), and generalized reinforcement preference optimization (GRPO) to fine-tune a lightweight large language model, such as 0.5B. After training, the model is then fine-tuned using DPO based on manually selected and rejected samples from the preference pool. This embodiment utilizes reinforcement learning training, which further enhances the model's stability and generalization ability.

[0083] In one embodiment, when performing intent recognition inference based on the trained intent recognition model, if a new tool is added to the toolset, the description information of the new tool is added to the input data of the trained intent recognition model so that the intent recognition model can determine the target tool from the updated toolset.

[0084] Once the trained intent recognition model is obtained, when performing inference based on the intent recognition model, if there are new intent categories and corresponding new tools, tool descriptions can be added directly to the input data of the intent recognition model during inference to achieve flexible configuration. This eliminates the need to rebuild the intent recognition samples used for training, improves the flexibility of model inference, and reduces development costs.

[0085] In one embodiment, during the inference process of the trained intent recognition model, after the toolset is initially converted into a vector representation, the vector representation is cached based on an attention key-value caching mechanism. Specifically, when the intent recognition model performs inference, the toolset needs to be input as part of the prompt words into the intent recognition model so that the intent recognition model can perform inference. After input, the descriptions of each tool and parameter in the toolset are converted into vector representations (kv_cache). Since the text to be converted is long, the process of calculating kv_cache each time is long, thereby reducing the inference speed of the model. To address this, this embodiment caches the corresponding vector representations after the relevant content of the toolset is initially converted into vector representations to avoid repeated vector representation calculations. In this way, the inference speed of the intent recognition model can be improved, thereby improving the efficiency of intent recognition. In addition, when the intent recognition model performs inference, it can first output the toolset corresponding to the parent intent category, and then input the toolset corresponding to the child intent category.

[0086] Figure 3 These are flowcharts of intent recognition methods in some embodiments, such as... Figure 3 As shown, the intent recognition method includes the following steps:

[0087] Step S301: Construct intent recognition samples; these intent recognition samples include positive samples and negative samples; negative samples include at least voice misrecognition samples and non-user command samples; different positive samples constitute multi-command samples. In addition, intent recognition samples also include hierarchical samples; the construction process of hierarchical samples includes: determining parent and child intent categories according to the hierarchical rules of intent categories; setting multi-level prompt words for sample classification for parent and child intent categories based on the hierarchical rules and preset conflict rules; and using a large language model to hierarchically optimize the intent recognition samples based on the multi-level prompt words to obtain hierarchical samples.

[0088] Step S302: Use intent recognition samples and a preset toolset as training data to train a preset large language model.

[0089] Step S303: Based on the training in step S302, reinforcement learning training is performed on the large language model based on preference samples; the preference samples are intent recognition samples that have been pre-predicted by different preset large language models and whose predicted categories are inconsistent with the actual intent categories.

[0090] Step S304: During the inference process of the trained intent recognition model, after the toolset is converted into a vector representation for the first time, the vector representation is cached based on the attention key-value caching mechanism.

[0091] Step S305: When a new tool is added to the toolset, the description information of the new tool is added to the input data of the trained intent recognition model so that the intent recognition model can determine the target tool from the updated toolset.

[0092] Step S306: Obtain user input information.

[0093] Step S307 involves inputting user input information into the trained intent recognition model, enabling the model to recognize the user's intent and, based on the recognition result, determine the target tool from a preset toolset that corresponds to the recognition result, and then invoke the target tool to perform the operation corresponding to the recognition result. The order of steps S304 to S305 and steps S306 to S307 can be interchanged.

[0094] Steps S301 to S307 described above implement a complete processing flow from the reconstruction of intent categories, the generation of intent recognition samples corresponding to the intent categories, training of the large language model, reinforcement learning training, to the deployment and inference of the intent recognition model. The large language model's prompts are used for batch generation and inspection of samples, and multi-command samples are added. Negative samples are enhanced for voice input scenarios. New intent categories and tools are incorporated into the model inference stage, avoiding complex redevelopment. This improves the semantic discrimination capability of intent recognition, supports multi-command interaction, increases the accuracy of intent recognition, achieves stable and highly generalized intent recognition, enhances the flexibility of intent recognition, and reduces development and maintenance costs.

[0095] This embodiment also provides an intent recognition device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. The terms "module," "unit," "subunit," etc., used below refer to combinations of software and / or hardware that implement a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0096] Figure 4 This is a structural block diagram of the intent recognition device 40 in this embodiment, as shown below. Figure 4 As shown, the intent recognition device 40 includes: an acquisition module 42 and an intent recognition module 44; wherein:

[0097] The acquisition module 42 is used to acquire user input information; the intent recognition module 44 is used to input the user input information into the trained intent recognition model, so that the trained intent recognition model can recognize the intent of the user input information, and determine the tool corresponding to the recognition result from the preset toolset as the target tool according to the recognition result, and call the target tool to perform the operation corresponding to the recognition result; wherein: the toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-built intent recognition samples combined with the preset toolset; the intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multi-command samples; negative samples include at least speech misrecognition samples and non-user command samples.

[0098] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0099] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated in this embodiment.

[0100] In one embodiment, a computer device is provided, which may be a terminal. Figure 5 This is an internal structural diagram of the computer device in this embodiment. (As shown...) Figure 5As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an intent recognition method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0101] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0102] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0103] Step 1: Obtain user input information;

[0104] Step 2: Input the user input information into the trained intent recognition model so that the trained intent recognition model can recognize the user input information and determine the tool corresponding to the recognition result from the preset toolset as the target tool, and call the target tool to perform the operation corresponding to the recognition result; wherein: the toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-built intent recognition samples and the preset toolset; intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multi-command samples; negative samples include at least speech misrecognition samples and non-user command samples.

[0105] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0106] Step 1: Obtain user input information;

[0107] Step 2: Input the user input information into the trained intent recognition model so that the trained intent recognition model can recognize the user input information and determine the tool corresponding to the recognition result from the preset toolset as the target tool, and call the target tool to perform the operation corresponding to the recognition result; wherein: the toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-built intent recognition samples and the preset toolset; intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multi-command samples; negative samples include at least speech misrecognition samples and non-user command samples.

[0108] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0109] Step 1: Obtain user input information;

[0110] Step 2: Input the user input information into the trained intent recognition model so that the trained intent recognition model can recognize the user input information and determine the tool corresponding to the recognition result from the preset toolset as the target tool, and call the target tool to perform the operation corresponding to the recognition result; wherein: the toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-built intent recognition samples and the preset toolset; intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multi-command samples; negative samples include at least speech misrecognition samples and non-user command samples.

[0111] In conjunction with the intent recognition methods provided in the above embodiments, this embodiment can also provide smart glasses to implement this method. The smart glasses store a computer program; when executed by a processor, the computer program implements any of the intent recognition methods described in the above embodiments.

[0112] It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. All other embodiments derived by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.

[0113] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0114] Obviously, the accompanying drawings are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application.

[0115] The term "embodiment" in this application refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply that it is mutually exclusive with or independent of other embodiments. It will be clearly or implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0116] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.

Claims

1. An intent recognition method, characterized in that, include: Obtain user input information; The user input information is input into a trained intent recognition model, which then performs intent recognition on the user input information. Based on the recognition result, the model selects a tool from a preset toolset that corresponds to the recognition result as the target tool and calls the target tool to perform the operation corresponding to the recognition result. The toolset includes several tools, and different tools are functional components that perform different operations. The trained intent recognition model is a large language model trained based on pre-constructed intent recognition samples combined with the preset toolset. The intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multiple command samples; the negative samples include at least speech misrecognition samples and non-user command samples; the positive samples are texts that are semantically consistent with the corresponding intent category, and the negative samples are samples that are related to the corresponding intent category but do not conform to the description of the intent category and are semantically inconsistent with the intent category; both the positive and negative samples are generated by writing prompt words and matching them using the large language model; the negative samples also include the speech misrecognition samples and non-user command samples; the speech misrecognition samples are speech recognition errors added to address the errors that exist in the process of converting user speech commands into text in real-world usage scenarios; the non-user command samples are response speech to the speech input, and are used to address model response speech that has not undergone echo cancellation in real-world usage scenarios. The process of constructing the speech misrecognition sample includes: acquiring preset speech data; inputting the speech data into a preset speech recognition model for text conversion to obtain the speech recognition model output; taking the characters whose output probabilities satisfy the target conditions in the speech recognition model output as target characters; the target conditions are that the output probability ranking is lower than the highest ranking preset position, and the difference between the output probability and the highest ranking is less than a preset threshold; replacing the real text corresponding to the speech data with the target characters to obtain the speech misrecognition sample; wherein, when the ASR model obtains the output of each character corresponding to the input speech data, the characters whose output probabilities rank preset position and whose difference from the highest probability is less than a preset threshold are retained as target characters, and the original characters at the corresponding positions are replaced with the target characters to obtain the speech misrecognition sample; The intent recognition samples also include other negative samples. Specifically, the misidentified samples are checked using sample check prompts combined with a large language model, and the verification result of whether they are negative samples is output. The misidentified samples are checked using a fine-tuning model, and the verification result of whether they are negative samples is output. The other negative samples are determined by comparing the two sets of verification results. The intent recognition samples also include hierarchical samples. The construction process of hierarchical samples specifically includes: determining the parent intent category and child intent category according to the hierarchical rules of intent categories; setting multi-level prompt words for sample classification for the parent intent category and child intent category based on the hierarchical rules and preset conflict rules; and using a large language model to perform hierarchical optimization on the intent recognition samples based on the multi-level prompt words to obtain hierarchical samples. Among them, decision branches and conflict rules are embedded in the prompt words to form multi-level prompt words.

2. The intent recognition method according to claim 1, characterized in that, The process of constructing the non-user instruction sample includes: Collect the response voice of the large language model to the user's voice in a preset voice interaction scenario; The response voice is converted into text data to obtain the non-user command sample.

3. The intent recognition method according to claim 1, characterized in that, The training process of the intent recognition model includes: using the intent recognition samples and the preset toolset as training data to train a preset large language model.

4. The intent recognition method according to claim 3, characterized in that, The training process of the intent recognition model further includes: performing reinforcement learning training on the large language model based on preference samples; the preference samples are intent recognition samples that have been pre-predicted by different preset large language models and whose predicted categories are inconsistent with the actual intent categories.

5. The intent recognition method according to claim 1, characterized in that, When performing intent recognition inference based on the trained intent recognition model, if a new tool is added to the toolset, the description information of the new tool is added to the input data of the trained intent recognition model so that the intent recognition model can determine the target tool from the updated toolset.

6. The intent recognition method according to any one of claims 1 to 5, characterized in that, During the inference process of the trained intent recognition model, after the toolset is first converted into a vector representation, the vector representation is cached based on an attention key-value caching mechanism.

7. An intent recognition device, characterized in that, It includes an acquisition module and an intent recognition module; wherein: The acquisition module is used to acquire user input information; The intent recognition module is used to input the user input information into the trained intent recognition model, so that the trained intent recognition model can recognize the intent of the user input information, and determine the tool corresponding to the recognition result from a preset toolset as the target tool, and call the target tool to perform the operation corresponding to the recognition result; wherein: the toolset includes several tools; different tools are functional components that perform different operations; the trained intent recognition model is a large language model trained based on pre-constructed intent recognition samples combined with the preset toolset; the intent recognition samples include: positive samples and negative samples for each intent category; different positive samples constitute multi-command samples; the negative samples include at least speech misrecognition samples and non-user pointers. The sample is defined as follows: a positive sample is text that is semantically consistent with the corresponding intent category; a negative sample is a sample that is related to the corresponding intent category but does not conform to the description of the intent category and is semantically inconsistent with the intent category. Both the positive and negative samples are generated by writing prompt words and matching them using a large language model. The negative sample also includes the speech misrecognition sample and the non-user command sample. The speech misrecognition sample is an error text added to address the errors that exist in the process of converting user speech commands into text in real-world usage scenarios. The non-user command sample is the response speech to the speech input, and it is used to address the model response speech that has not undergone echo cancellation in real-world usage scenarios. The process of constructing the speech misrecognition sample includes: acquiring preset speech data; inputting the speech data into a preset speech recognition model for text conversion to obtain the speech recognition model output; taking the characters whose output probabilities satisfy the target conditions in the speech recognition model output as target characters; the target conditions are that the output probability ranking is lower than the highest ranking preset position, and the difference between the output probability and the highest ranking is less than a preset threshold; replacing the real text corresponding to the speech data with the target characters to obtain the speech misrecognition sample; wherein, when the ASR model obtains the output of each character corresponding to the input speech data, the characters whose output probabilities rank preset position and whose difference from the highest probability is less than a preset threshold are retained as target characters, and the original characters at the corresponding positions are replaced with the target characters to obtain the speech misrecognition sample; The intent recognition samples also include other negative samples. Specifically, the misidentified samples are checked using sample check prompts combined with a large language model, and the verification result of whether they are negative samples is output. The misidentified samples are checked using a fine-tuning model, and the verification result of whether they are negative samples is output. The other negative samples are determined by comparing the two sets of verification results. The intent recognition samples also include hierarchical samples. The construction process of hierarchical samples specifically includes: determining the parent intent category and child intent category according to the hierarchical rules of intent categories; setting multi-level prompt words for sample classification for the parent intent category and child intent category based on the hierarchical rules and preset conflict rules; and using a large language model to perform hierarchical optimization on the intent recognition samples based on the multi-level prompt words to obtain hierarchical samples. Among them, decision branches and conflict rules are embedded in the prompt words to form multi-level prompt words.

8. A type of smart glasses, wherein a computer program is stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the intent recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Tool calling method in task-based dialogue, medium, equipment and product

    CN120523922A