Intention recognition method and device, electronic equipment and storage medium
By dynamically adjusting the balance between exploration and utilization through the MCTS algorithm, and combining the exploration coefficient and upper confidence bound, the optimal intent is selected and expanded, which solves the problem of insufficient intent recognition accuracy in existing technologies and achieves higher intent recognition accuracy and user experience.
Patent Information
- Application Number
- CN202510678899.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-05-21
AI Technical Summary
In the existing technology, the intent recognition method relies on a single confidence indicator, which leads to one-sided intent selection and reduces the accuracy of intent recognition. In particular, it is difficult to effectively handle complex user needs in multi-intent scenarios.
The Monte Carlo Tree Search (MCTS) algorithm is used to select the optimal intent by dynamically calculating the exploration coefficient and upper confidence limit, and then intent expansion and Monte Carlo simulation are performed to determine the target intent through comprehensive scoring.
It improves the accuracy and robustness of intent recognition, enhances the ability to understand complex user needs, and improves user experience.
Smart Images

Figure CN120688515A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to an intent recognition method, device, electronic device and storage medium. Background Art
[0002] During user conversations, intent recognition is a core technology in natural language processing (NLP) and dialogue systems. Its goal is to determine the underlying intent or need by analyzing user input (text, voice, etc.). Currently, multiple intents in user input can be identified based on rule matching or classification models. However, the optimal intent among these multiple intents is currently determined solely based on the confidence level of each intent. This reliance on a single confidence metric can lead to biased selection, thereby reducing the accuracy of intent recognition. Summary of the Invention
[0003] The present application provides an intent recognition method, apparatus, electronic device, and storage medium, which automatically select the optimal intent from multiple candidate intents by combining Monte Carlo Tree Search (MCTS), thereby improving the accuracy of intent recognition.
[0004] In a first aspect, the present application provides an intent recognition method, the method comprising: Determining, based on multiple first intentions corresponding to the first conversation, an exploration coefficient of each first intention at a first moment; Determine the upper confidence limit value corresponding to each first intention according to the exploration coefficient of each first intention; According to the upper confidence bound value corresponding to each first intent, a second intent is selected from the multiple first intents and the second intent is expanded to obtain multiple third intents; performing a Monte Carlo simulation based on the second intent and the plurality of third intents to determine a first score for the second intent and a first score for each of the plurality of third intents; determining a target intent from the second intent, the plurality of third intents, and the plurality of fourth intents based on the plurality of first scores and the second score of each fourth intent; The fourth intention is an intention in the first intention that is different from the second intention, and the second score is an average score of the fourth intention recorded before the first moment.
[0005] It can be seen that in this application, by dynamically calculating the exploration coefficient of each first intent in the first conversation to adjust the exploration and utilization balance in the MCTS algorithm, combining the exploration coefficient to calculate the upper confidence limit of each first intent and select the optimal second intent, further expanding the second intent to generate multiple third intents, evaluating the value of the second and third intents through Monte Carlo simulation, and determining the target intent by combining the scores of all the aforementioned intents, the technical effect of improving the accuracy and robustness of intent recognition can be achieved. This method enhances the ability to understand users with complex needs by dynamically adjusting the exploration coefficient, thereby improving the user experience.
[0006] In a second aspect, the present application provides an intention recognition device, the device comprising: a first processing unit, configured to determine, based on the plurality of first intentions corresponding to the first conversation, an exploration coefficient of each first intention at a first moment; The first processing unit is further configured to determine an upper confidence limit corresponding to each first intention according to the exploration coefficient of each first intention; a second processing unit, configured to select a second intent from the plurality of first intents and expand the second intent according to the upper confidence limit corresponding to each first intent, to obtain a plurality of third intents; a third processing unit, configured to perform a Monte Carlo simulation based on the second intent and the plurality of third intents to determine a first score for the second intent and a first score for each of the plurality of third intents; a fourth processing unit, configured to determine a target intent from the second intent, the plurality of third intents, and the plurality of fourth intents based on the plurality of first scores and the second score of each fourth intent; The fourth intention is an intention in the first intention that is different from the second intention, and the second score is an average score of the fourth intention recorded before the first moment.
[0007] In a third aspect, the present application provides an electronic device comprising a processor, a memory, and a communication interface. The processor, memory, and communication interface are interconnected and perform communication with each other. The memory stores executable program code, the communication interface is used for wireless communication, and the processor is used to retrieve the executable program code stored in the memory and execute some or all of the steps described in any method of the first aspect.
[0008] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements some or all of the steps described in the first aspect of the present application.
[0009] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when processed and executed, implements some or all of the steps described in the first aspect of the present application. The computer program product may be a software installation package. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0011] Figure 1 A schematic diagram of the structure of an intent recognition system provided in an embodiment of the present application; Figure 2 A flowchart of an intent recognition method provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of an intent tree provided in an embodiment of the present application; Figure 4 A flowchart of another intent recognition method provided in an embodiment of the present application; Figure 5 A flowchart of another method for identifying intent provided in an embodiment of the present application; Figure 6 A block diagram of the functional units of an intent recognition device provided in an embodiment of the present application; Figure 7 A block diagram of the functional units of another intention recognition device provided in an embodiment of the present application; Figure 8 This is a structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0012] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0013] The terms "first," "second," and so on, in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps is not limited to the listed steps but may optionally include steps not listed, or may optionally include other steps inherent to the process, method, product, or apparatus.
[0014] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0015] In practical applications, especially in the fields of intelligent customer service and dialogue systems, the text input by users may involve multiple potential intentions. Traditional intent recognition methods are mostly single intent recognition. Even if they can recognize multiple intent contents, there are still big problems in selection and they cannot fully handle complex multi-intent scenarios. For example: Figure 1 ,meaning Figure 2 ,meaning Figure 3 Wait, if you choose the one with the highest confidence Figure 1 , which may be necessary in subsequent conversations. Figure 2 Most importantly, the initial choice is not optimal. If all the user's intentions are answered one by one, the dialogue system will not be intelligent and humane enough, and will not be able to answer the user's intentions through an effective dialogue process.
[0016] Based on this, the present application provides an intent recognition method, which combines MCTS to automatically select the optimal intent from multiple candidate intents. Specifically, the exploration coefficient of each first intent in the current conversation is dynamically calculated to adjust the exploration and utilization balance in the MCTS algorithm, and the upper confidence limit of each first intent is calculated in combination with the exploration coefficient and the optimal second intention is selected. The second intention is further expanded to generate multiple third intentions, and the value of the second and third intentions is evaluated through Monte Carlo simulation. The target intent is determined by combining the scores of the multiple intentions obtained above, which can achieve the technical effect of significantly improving the accuracy and robustness of intent recognition. This method enhances the ability to understand users with complex needs by dynamically adjusting the exploration coefficient, thereby improving the user experience.
[0017] The following describes the scenarios involved in the embodiments of this application.
[0018] See also Figure 1 , Figure 1 A schematic diagram of the structure of an intention recognition system provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the intention recognition system 100 includes a terminal device 101 and a server 102 .
[0019] The terminal device 101 is used to implement a dialogue with the user, obtain the dialogue input by the user, and then send the dialogue input by the user to the server 102.
[0020] After receiving a conversation input by a user, server 102 analyzes the conversation and generates a corresponding response based on previous related conversations and a corresponding database. Server 102 can be integrated into terminal device 101. Optionally, terminal device 101 is not limited to a desktop computer, laptop computer, tablet computer, smartphone, etc. Server 102 is not limited to a server, server cluster, cloud server, cloud computing service center, or other device with computing capabilities.
[0021] Specifically, after obtaining the dialogue input by the user, the terminal device 101 sends the dialogue to the server 102. The server 102 first determines the multiple first intentions corresponding to the dialogue, and dynamically determines the exploration coefficient of each first intention at the first moment. Subsequently, the upper confidence limit value corresponding to each first intention is determined in combination with the exploration coefficient of each first intention at the first moment, and the optimal second intention is selected according to the upper confidence limit value. The second intention is further expanded to generate multiple third intentions. By performing Monte Carlo simulation on the second intention and multiple third intentions, the first score of the second intention and the first score of each intention in the multiple third intentions are determined. Finally, the scores of all intentions are combined to determine the target intention. In this way, the technical effect of significantly improving the accuracy and robustness of intent recognition can be achieved. By dynamically adjusting the exploration coefficient, the ability to understand users with complex needs is enhanced, thereby improving the user experience.
[0022] The following introduces the prior art involved in the embodiments of this application.
[0023] Monte Carlo Tree Search: Monte Carlo Tree Search, a branch of computational mathematics based on probability and statistical theory, uses random (or pseudo-random) numbers to solve complex decision-making problems. It uses a decision tree as a representation of the search space. By repeatedly simulating games or decision-making processes, it evaluates the value of different decisions and selects the most valuable one. Its core concept is to optimize the search tree by iteratively selecting, expanding, simulating, and updating nodes. It combines the generality of random simulation with the accuracy of tree search, enabling it to efficiently find optimal solutions even in scenarios with large search spaces.
[0024] The Monte Carlo tree search algorithm usually includes the following stages.
[0025] Selection: Starting from the root node, the optimal child nodes are recursively selected according to a specific strategy until a leaf node is reached. This selection strategy typically balances two factors: Exploration: Selecting previously unexplored nodes to gain new information. Exploitation: Selecting nodes that already have high scores or good performance to ensure stable system performance. For example, an upper confidence bound (UCB) strategy can be used to balance exploration and exploitation.
[0026] Expansion: If the leaf node is not a terminal node, create one or more child nodes for it and select one of them to expand.
[0027] Simulation: Starting from the expanded nodes, random simulation is performed. This involves simulating the game or decision-making process according to a strategy (such as random selection) until the game ends or a termination condition is reached. This step is often called a Monte Carlo simulation.
[0028] Backpropagation: Backpropagates simulation results (such as game wins and losses, reward values, etc.) into the search tree and updates node statistics (such as number of visits, average value, etc.) for subsequent selection and evaluation.
[0029] Based on this, an embodiment of the present application provides an intent recognition method, which is described in detail below in conjunction with the accompanying drawings.
[0030] In the first embodiment, the main process of the intention recognition method is described below.
[0031] See also Figure 2 , Figure 2 A flow chart of an intent recognition method provided in an embodiment of the present application is provided, and the method is applied to the above-mentioned server, such as Figure 2 As shown, the method includes the following steps.
[0032] Step S201 : determining an exploration coefficient of each first intention at a first moment according to multiple first intentions corresponding to the first conversation.
[0033] Among them, the first conversation can be text or voice information input by the user, which can serve as the original data for intent recognition by obtaining it from the user interaction interface or generating it through a speech-to-text module. For example, the first conversation can include but is not limited to the user's natural language expression, keyword combination or semantic vector representation, etc.
[0034] Multiple first intents can be a set of possible intents predicted by a multi-label classification model, each of which represents a possible need or goal of the user. In a specific embodiment, features (such as keywords, semantic vectors, etc.) can be extracted from user input, and a trained multi-label classification model can be used to predict all possible first intents and their confidence levels, thereby generating multiple candidate intents to provide a basis for subsequent selection. The multi-label classification model can be trained using traditional machine learning methods (such as multi-layer perceptrons (MLPs) and support vector machines (SVMs)) or deep learning methods (such as long short-term memory networks (LSTMs) or bidirectional encoder representations from transformers (BERT)).
[0035] The exploration coefficient refers to the parameter C used in the upper confidence bound strategy to balance exploration and exploitation. Larger C values favor more exploration, while smaller C values favor more exploitation. Specifically, when C is large, the second intent selected subsequently may be one that has been visited less frequently among multiple first intents; when C is small, the second intent selected subsequently may be one that has been visited more frequently among multiple first intents. The first moment can refer to the current moment, for example, the moment the first conversation was acquired.
[0036] Step S202: determining an upper confidence limit corresponding to each first intention according to the exploration coefficient of each first intention.
[0037] The upper confidence bound is the reference value corresponding to the upper confidence bound strategy. It can also be an indicator of node value, combining the current node score and the exploration factor to guide the MCTS selection strategy. For example, the dynamic exploration factor is substituted into the UCB formula to calculate the upper confidence bound for each first intent. This value comprehensively considers the node's historical performance (score) and exploration requirements, providing a quantitative basis for subsequent selection. As you can see, each node corresponds to an intent, and the intent corresponding to the current node is an extension of the intent corresponding to the current node's parent node.
[0038] The UCB formula is as follows:
[0039] Among them, average_reward_of_the_node represents the node score, in_parent_visits represents the number of times the parent node of the current node is visited, node_visits represents the number of times the current node is visited, and C is the exploration coefficient.
[0040] Step S203: Select a second intent from the multiple first intents according to the upper confidence limit value corresponding to each first intent, and expand the second intent to obtain multiple third intents.
[0041] Among them, the second intention is selected from multiple first intentions as the starting point of the expansion phase. For example, the upper confidence bounds of all first intentions can be compared, and the first intention corresponding to the largest upper confidence bound among all the upper confidence bounds of the first intentions can be selected as the second intention. This ensures that the intention selected is consistent with historical performance and takes into account exploration needs; the first intention corresponding to the second largest upper confidence bound among all the upper confidence bounds of the first intentions can also be selected as the second intention. This can avoid premature convergence and promote global exploration. The first intention corresponding to a single upper confidence bound can also be randomly selected from multiple higher upper confidence bounds as the second intention. This can balance resource allocation.
[0042] The third intent is a sub-intent derived from the second intent, representing a potential user demand path. For example, based on the second intent, a set of possible sub-intents is generated. These sub-intents expand the structure of the intent tree and provide more options for the simulation stage.
[0043] The aforementioned expansion can refer to semantic expansion. It can also be expanded based on sub-intent refinement, context association, user behavior prediction, etc.
[0044] For example, see Figure 3 , Figure 3 A schematic diagram of the structure of an intent tree provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, the root node includes user input. Starting from the root node, the multi-label classification model derives multiple intents corresponding to the user input, namely, querying loan amount, querying loan interest rate, and querying repayment method, which serve as three child nodes. After querying the loan amount, the user may further inquire about how the loan amount is calculated. This can be expanded to include the calculation method and maximum loan amount. For querying the loan interest rate, this can be expanded to include the types of loan interest rates and loan interest rate preferential indicators. For querying the repayment method, this can be expanded to include the conditions and fees for early repayment, as well as the repayment period.
[0045] Step S204 : performing a Monte Carlo simulation based on the second intent and the plurality of third intents to determine a first score for the second intent and a first score for each of the plurality of third intents.
[0046] The aforementioned Monte Carlo simulation primarily involves randomly selecting paths from the multiple paths formed by the expanded multiple third intents and the second intent, thereby simulating a conversation scenario based on the aforementioned intents and user feedback. Finally, the first score of the second intent, as well as the first score of each of the multiple third intents, is determined based on the simulated score corresponding to each intent. Exemplarily, a single intent from the multiple third intents is randomly selected to form a path with the second intent and perform a Monte Carlo simulation. Indicators such as user intent matching rate, interaction interest, and information gain are recorded, and a weighted formula is used to calculate the first score of each intent.
[0047] Step S205 : determining a target intent from the second intent, the plurality of third intents, and the plurality of fourth intents based on the plurality of first scores and the second score of each fourth intent.
[0048] The fourth intent is the first intent that differs from the second intent, and the second score is the average score of the fourth intent recorded before the first moment. The target intent is the final, optimal intent selected and serves as the basis for the system's response. For example, the scores of the second, third, and fourth intents are combined, and the highest-scoring intent is selected as the target intent. This process ensures a globally optimal selection.
[0049] As can be seen, this application provides an intent recognition method that dynamically calculates the exploration coefficient of each first intent to adjust the exploration and utilization balance in the MCTS algorithm, combines the exploration coefficient to calculate the upper confidence limit and select the optimal second intent, further expands the second intent to generate multiple third intents, evaluates the value of the second and third intents through Monte Carlo simulation, and determines the target intent based on the scores of multiple intents. This method can achieve the technical effect of significantly improving the accuracy and robustness of intent recognition. By dynamically adjusting the exploration coefficient and expanding the intent tree structure, this method enhances the ability to understand complex user needs, thereby improving the user experience.
[0050] In the second embodiment, the intention recognition method is described in detail below based on the determination details of the exploration coefficient.
[0051] See also Figure 4 , Figure 4 A flow chart of another intent recognition method provided in an embodiment of the present application is provided, which is applied to the above-mentioned server, such as Figure 4 As shown, the method includes the following steps.
[0052] Step S401: Determine an exploration coefficient for each first intention based on the conversation round corresponding to the first moment in the multi-round conversation, the behavioral entropy of the second conversation, and the intent confidence of each first intention in the multiple first intentions corresponding to the first conversation.
[0053] The conversation turn corresponding to the first moment in a multi-round conversation can be obtained by statistically analyzing historical conversation records. The behavioral entropy of the second conversation can be an indicator used to quantify the uncertainty and diversity of the user's conversational intentions in the second conversation. A higher value indicates a less stable user's intentions in the second conversation. The second conversation represents a conversation with a target user before the first moment, where the target user is the user corresponding to the first conversation. In a specific embodiment, the behavioral entropy can be obtained by calculating the probability-weighted information entropy of the user's intention distribution before the first moment. For example, the behavioral entropy can reflect the frequency and pattern of the user's switching between different intentions. Intent confidence can be the probability value corresponding to each intent output by a multi-label classification model, which is used to reflect the system's confidence in a particular intent. For example, intent confidence can include, but is not limited to, the probability value or similarity score of the system's prediction of a specific intent.
[0054] The exploration coefficient is determined by combining the three factors mentioned above. Specifically, in one embodiment, the number of conversation turns influences the exploration intensity through an exponential decay term, with the exploration requirement gradually decreasing as the conversation deepens. Behavioral entropy reflects the diversity of the user's intentions in the second conversation. When the user's intentions frequently switch, the exploration weight needs to be increased to accommodate uncertainty. Intent confidence acts as a compensation term. When the system lacks confidence in a particular intention, the exploration scope needs to be expanded to find a better solution. This process ensures that the exploration factor can be dynamically adjusted to meet the needs of different conversation stages, thereby improving the flexibility and accuracy of the selection strategy.
[0055] It can be seen that by comprehensively considering the conversation turn at the first moment in a multi-round conversation, the behavioral entropy of the second conversation, and the intent confidence of each first intent among the multiple first intents corresponding to the first conversation, and dynamically adjusting the exploration coefficient based on these factors, the system can flexibly balance the relationship between exploration and exploitation during the conversation. Compared with the traditional fixed-value exploration factor, this method can better adapt to changes in conversation scenarios. In particular, when user intentions are ambiguous or diverse, the system can proactively adjust the balance between exploration and exploitation to avoid premature convergence to incorrect intentions. In addition, the design of the dynamic exploration coefficient enhances the robustness of the system, enabling it to more accurately identify the user's true needs in complex conversation environments, thereby significantly improving the overall conversation quality and user experience.
[0056] Optionally, the exploration coefficient of each first intention is determined based on the conversation round corresponding to the first moment in the multi-round conversation, the behavioral entropy of the second conversation, and the intention confidence of each first intention in the multiple first intentions corresponding to the first conversation, including: constructing a first function based on the natural index and the conversation round corresponding to the first moment in the multi-round conversation; constructing a second function based on the difference between the first preset value and the intention confidence of each first intention in the multiple first intentions corresponding to the first conversation; determining a third function based on the first function, the second function, the behavioral entropy of the second conversation, and the second preset value; and determining the exploration coefficient of each first intention based on the preset exploration coefficient and the third function.
[0057] The natural exponential can be an exponential function with the mathematical constant e as its base, exhibiting smooth variations and can be generated through mathematical modeling. For example, the natural exponential can include, but is not limited to, a power operation with e as its base. In one specific embodiment, this can be achieved by using the natural exponential function to model the conversational turn corresponding to the first moment in a multi-round conversation. This can achieve the effect of gradually decreasing the value of the first function as the conversational turn increases, reflecting the trend of exploration intensity decreasing as the conversation deepens.
[0058] The first preset value can be a fixed benchmark value set to measure the degree of influence of the intent confidence, which can be obtained by manual setting or system default configuration. Exemplarily, the first preset value may include but is not limited to a value within a fixed numerical range. The difference between the first preset value and the intent confidence of each first intention in the multiple first intentions corresponding to the first conversation is calculated, and the second function is constructed based on this. In a specific embodiment, it can be achieved by difference calculation, so that when the intent confidence of each first intention is high, the difference is small and the second function value is also small, indicating that the exploration intensity should be reduced at this time; conversely, when the intent confidence of each first intention is low, the exploration range needs to be expanded to find a better solution.
[0059] The second preset value may be another fixed reference value set to balance the influence of other factors, and may be obtained through manual setting or system default configuration. Exemplarily, the second preset value may include but is not limited to a value within a fixed numerical range. The third function is constructed by adding the first function, the second function, the behavioral entropy of the second conversation, and the second preset value. In a specific embodiment, this can be achieved by direct summation, thereby achieving the effect of comprehensively considering the influence of the current conversation turn, the intent confidence of the first intent in the current conversation, and the diversity of user behavior on the exploration coefficient, ensuring that the dynamic adjustment mechanism can fully adapt to the needs of different conversation scenarios.
[0060] The exploration coefficient of each first intention is determined according to the product of the preset exploration coefficient and the third function. The preset exploration coefficient can be an initially set fixed exploration factor with a default value of 1, which can be obtained by manual setting or system default configuration as the basis for dynamic adjustment. Exemplarily, the preset exploration coefficient may include but is not limited to a value within a fixed numerical range. By taking the product of the preset exploration coefficient and the third function as the final exploration coefficient. In a specific embodiment, this can be achieved by multiplication operation, so that the preset value can be corrected by introducing the third function to achieve the effect of dynamic adjustment of the exploration factor.
[0061] The calculation formula of the exploration coefficient is as follows:
[0062] Where t represents the number of conversation turns. As the number of conversation turns increases, the exploration intensity gradually decreases. H represents the behavior entropy. S represents the confidence level of the current intent. If the confidence level is low, the exploration requirement is increased, and the exploration of alternative intents should be strengthened. α, β, γ, and δ are weight coefficients that can be adjusted based on the actual application. For example, they can be set to 0.5, 0.2, 0.3, and 0.4, respectively. C_base is the preset exploration coefficient, which can default to 1, indicating default exploration.
[0063] For example, the following scenario: In the fifth round of conversation, the user asked: "Can this plan be more favorable?"
[0064] Dialogue rounds: .
[0065] Historical entropy (users’ past intentions to frequently switch loan types, terms, etc.): .
[0066] Current confidence value: 0.6.
[0067] The dynamic C value calculation formula is as follows:
[0068] It can be seen that by constructing the first function based on the natural index and the number of conversation rounds, the trend of the exploration intensity decreasing as the conversation deepens is reflected; the second function is constructed based on the difference between the first preset value and the intention confidence of each first intention, and the exploration intensity is dynamically adjusted according to the intention confidence; the third function is constructed by determining the sum of the first function, the second function, the behavioral entropy and the second preset value, and the impact of the number of conversation rounds, intention confidence and user behavior diversity on the exploration coefficient is comprehensively considered; the exploration coefficient is determined according to the product of the preset exploration coefficient and the third function, and the dynamic adjustment of the exploration factor is realized, which can improve the flexibility and adaptability of the exploration strategy, avoid the problem of insufficient or excessive exploration caused by a single fixed value, and thus improve the accuracy of system intent recognition and the technical effect of user experience.
[0069] Optionally, determine multiple fifth intentions included in the second conversation, and the number of occurrences of each fifth intention in the multiple fifth intentions; determine the frequency of occurrence of each fifth intention based on the ratio between the number of occurrences of each fifth intention and the total number of occurrences corresponding to the multiple fifth intentions; determine the behavioral entropy based on the frequency of occurrence of each fifth intention.
[0070] The second conversation can refer to the conversation history that occurred before the current first conversation, including multiple rounds of interaction between the user and the system. By parsing the second conversation, all intents (i.e., fifth intents) can be extracted and the number of occurrences of each fifth intent can be counted. For example, the fifth intent can be a point of interest or need expressed by the user in past conversations, such as purchasing a product, checking the weather, or booking a service. By dividing the number of occurrences of each fifth intent by the total number of intents, its frequency of occurrence can be calculated. This process captures preference patterns in historical user behavior and provides basic data for subsequent calculations.
[0071] The logarithm is calculated using the natural logarithm as the base and is used to smooth the range of occurrence frequencies. For example, for each fifth intent, the fourth function value can be constructed by taking the natural logarithm of its frequency and multiplying it by the original frequency. This process leverages the core concept of information entropy theory, which states that low-frequency events carry more information, while high-frequency events carry relatively less. This approach allows for more accurate quantification of the information contribution of each fifth intent.
[0072] Behavioral entropy is an indicator used to describe the diversity of a user's historical behavior. Higher values indicate more uncertain or diverse user behavior. For example, the sum of the fourth function values corresponding to all fifth intents is used, and the negative of this sum is taken as the final behavioral entropy. This process comprehensively considers the information content of all fifth intents and reflects the overall distribution characteristics of the user's historical behavior. Higher behavioral entropy values indicate less fixed user intent choices, and the system's exploration requirements increase accordingly. Conversely, lower behavioral entropy values indicate more stable user behavior, and the system can focus more on leveraging known information.
[0073] The behavioral entropy formula is as follows:
[0074] in, Indicates user selection intent The frequency of user intention selection is variable, indicating a significant need for exploration. A low behavioral entropy value indicates a less significant need for exploration. N represents the number of fifth intents, and i represents the i-th fifth intent.
[0075] It can be seen that by determining the frequency of occurrence of each fifth intent in the second conversation, constructing the fourth function by combining the logarithmic value and the product of the frequency of occurrence, and further obtaining behavioral entropy by summing the fourth function values corresponding to multiple fifth intents and taking the negative value, the system's ability to characterize the diversity of users' historical behaviors can be enhanced. Compared with directly counting the frequency of intent occurrence, this method can more comprehensively reflect the uncertainty of user behavior, thereby providing a more reliable basis for dynamically adjusting the exploration factor. Especially in scenarios where user intent frequently switches or behavior patterns are complex, the introduction of behavioral entropy can significantly improve the adaptability and accuracy of the system, ensuring that the intent recognition process is more flexible and meets the needs of actual conversations.
[0076] Step S402: determining an upper confidence limit corresponding to each first intention according to the exploration coefficient of each first intention.
[0077] Step S403: Select a second intent from the multiple first intents according to the upper confidence limit value corresponding to each first intent, and expand the second intent to obtain multiple third intents.
[0078] Step S404: Perform a Monte Carlo simulation based on the second intent and the multiple third intents to determine a first score for the second intent and a first score for each of the multiple third intents.
[0079] Step S405 : determining a target intent from the second intent, the plurality of third intents, and the plurality of fourth intents based on the plurality of first scores and the second score of each fourth intent.
[0080] It is understandable that for the explanation of the aforementioned steps related to Example 1, please refer to the contents of Example 1 and will not be repeated here.
[0081] In the third embodiment, the intent recognition method is described in detail based on the scoring details of the second intent and the third intent.
[0082] See also Figure 5 , Figure 5 A flow chart of another method for identifying intent provided in an embodiment of the present application is provided, and the method is applied to the above-mentioned server, such as Figure 5 As shown, the method includes the following steps.
[0083] Step S501: Determine an exploration coefficient for each first intention based on the conversation round corresponding to the first moment in the multi-round conversation, the behavioral entropy of the second conversation, and the intent confidence of each first intention in the multiple first intentions corresponding to the first conversation.
[0084] Step S502: determining an upper confidence limit corresponding to each first intention according to the exploration coefficient of each first intention.
[0085] Step S503: Select a second intent from the multiple first intents according to the upper confidence limit value corresponding to each first intent, and expand the second intent to obtain multiple third intents.
[0086] Step S504: Determine multiple intention combinations, each intention combination including a second intention and a third intention.
[0087] The intention combination may be a combination of a second intention and a third intention, which is used to simulate the possibility of different dialogue paths. The third intention in the intention combination may be single or multiple.
[0088] Optionally, each of the multiple third intents generated in the expansion phase is paired with the second intent one by one to form multiple intent combinations. This process ensures that every possible dialogue path is considered, providing comprehensive input data for subsequent Monte Carlo simulations.
[0089] Optionally, multiple third intent combinations are determined from the multiple third intents generated in the expansion phase, each third intent combination including multiple third intents. In this case, each third intent combination is paired with the second intent one by one from the multiple third intent combinations to form multiple intent combinations.
[0090] Step S505 : Performing a Monte Carlo simulation based on each intention combination to determine a third score for each intention combination.
[0091] The third score can be an assessment of the effectiveness of each intent combination through simulated conversation scenarios, reflecting the combination's performance in actual conversations. For example, the third score may include, but is not limited to, metrics such as intent matching rate and interaction interest. In one specific embodiment, for each intent combination, the system randomly selects a path and simulates a conversation, records user feedback, and calculates the third score for the intent combination based on a comprehensive scoring formula. This process quantifies the value of each intent combination and provides a basis for subsequent score assignment.
[0092] For example, based on the above Figure 3 , the selectable intention combinations include root node → query interest rate → loan interest rate type.
[0093] Analog Dialogue: User: What is your interest rate?
[0094] System: Our loan interest rates may be adjusted due to market changes. Do you want to know the difference between fixed and floating interest rates?
[0095] User: What is the difference between fixed interest rate and floating interest rate?
[0096] System: Fixed interest rates remain unchanged during the loan period, suitable for a stable repayment plan; floating interest rates fluctuate with market interest rates and may be higher or lower.
[0097] System score: 10 points.
[0098] Based on the above simulated dialogue, it can be seen that the user asked in-depth questions, indicating that the path guides the user to further understand the types of interest rates, and the benefit of the path is positive (high score).
[0099] Based on the above Figure 3 ,The selectable intention combinations also include root node → query interest rate → preferential indicators of loan interest rate.
[0100] Analog Dialogue: User: What is your interest rate?
[0101] System: Currently our lowest annual interest rate is 4.5%, and we have many preferential indicators.
[0102] User: Oh, I'm not interested in the discount index, I'm not talking about the annual interest rate.
[0103] System: Our daily interest rate is 1%.
[0104] User: Hey.
[0105] System path assessment: 3 points.
[0106] Based on the above simulated dialogue, it can be seen that the user’s answer is not very satisfactory, the intent recognition is low, and the benefit of the path is low (low score).
[0107] Optionally, a Monte Carlo simulation is performed based on each intent combination to determine the third score of each intent combination, including: performing a Monte Carlo simulation based on each intent combination to determine multiple scoring factors for each intent combination; and performing weighted summation on the multiple scoring factors of each intent combination to determine the third score of each intent combination.
[0108] Scoring factors can be used to quantitatively evaluate the performance of intent combinations from different dimensions. They can be obtained by recording user feedback and extracting relevant information during Monte Carlo simulations. For example, scoring factors may include, but are not limited to, user intent matching rate, interaction interest, information gain, path completeness, and time and cost factors. These scoring factors can comprehensively capture the performance characteristics of each intent combination in actual conversations, providing multi-faceted data support for subsequent scoring calculations.
[0109] The third score can be the final score that quantitatively evaluates the overall performance of each intent combination. For example, the score can be obtained by integrating the results of multiple scoring factors. In a specific embodiment, the system performs a Monte Carlo simulation for each intent combination to simulate the conversation scenario, record user feedback, and extract the above-mentioned scoring factors from multiple dimensions. For example, the system can analyze whether the user's intent is accurately identified (intent matching rate), the user's level of participation in the conversation (interaction interest), the amount of new information provided during the conversation (information gain), whether the conversation path is complete and smooth (path completeness), and the time and resource consumption required to complete the conversation (time and cost factors). In this way, the system can obtain multiple scoring factors for each intent combination, thereby laying the foundation for subsequent scoring calculations.
[0110] For example, the evaluation criteria could be as follows: 1: User intent match rate (exact match: +10 points, partial match: +5 points, no match: +0 points).
[0111] 2: User interaction interest (user further asks questions or discusses in depth: +10 points, user simply confirms: +5 points, user shows no interest: +0 points).
[0112] 3: Information gain (providing new information that meets user needs: +10 points, providing partial new information: +5 points, no actual information gain: +0 points).
[0113] 4: Path completeness (successfully completed interaction: +10 points, uncompleted interaction: +0 points).
[0114] 5: Time and cost factors (few rounds (<3 rounds): +10 points, medium rounds (3-5 rounds): +5 points, many rounds (>5 rounds): +0 points).
[0115] Comprehensive scoring formula: Total score = w1*intent matching + w2*user interest + w3*information gain + w4*path completeness + w5*time efficiency.
[0116] The aforementioned weights can be adjusted flexibly based on actual needs, for example, w1 = 0.3, w2 = 0.3, w3 = 0.2, w4 = 0.1, and w5 = 0.1. This weighted summation ensures that the scoring results accurately reflect the actual value of the intent combination.
[0117] As can be seen, by performing Monte Carlo simulations on each intent combination to extract multiple scoring factors, and then calculating the third score for each intent combination through weighted summation of these scoring factors, the technical effect of significantly improving the scientific nature and accuracy of the third score can be achieved. Compared with single-dimensional evaluation methods, this method can more comprehensively capture the complex relationship between user needs and system responses, thereby providing a more reliable basis for the final target intent selection. In addition, by flexibly adjusting the weights of the scoring factors, the system can better adapt to the needs of different dialogue scenarios, further improving the accuracy of intent recognition and user experience.
[0118] Step S506: Based on the third score of each intention combination, determine the fourth score of the second intention and the first score of the third intention corresponding to each intention combination.
[0119] The fourth score of the second intent and the first score of the third intent corresponding to each intention combination may be the same score, or may be different scores determined based on different contribution values.
[0120] Optionally, the fourth score may be a value reflecting the contribution of the second intention to the current intention combination, and the first score may be a value reflecting the contribution of the third intention to the current intention combination. Exemplarily, the fourth score may include but is not limited to the proportion of the dominant role of the second intention in the entire combination. In a specific embodiment, the third score of each intention combination is decomposed into the contribution parts of the second intention and the third intention. Specifically, the fourth score of the second intention is assigned according to its dominant role in the entire combination, and the remaining part is used as the first score corresponding to the third intention. This process ensures that the score can accurately reflect the actual contribution of each intention.
[0121] It can be seen that by determining multiple intent combinations consisting of a second intent and a single third intent to simulate the possibility of different dialogue paths; performing Monte Carlo simulation based on each intent combination, recording user feedback and calculating the third score through a comprehensive scoring formula; decomposing the third score into the contribution parts of the second intent and the third intent, and determining the fourth score of the second intent and the first score of the third intent respectively, it is possible to significantly improve the scoring accuracy, enhance the level of refinement in the Monte Carlo simulation stage, and clearly identify the technical effect of the actual value of each intent.
[0122] Step S507: Determine the sum of the fourth score and the fifth score of the second intent corresponding to each intent combination.
[0123] The fifth score is the cumulative score of the second intention recorded before the first moment.
[0124] Step S508: Determine a first score for the second intention based on the ratio between the score and the first score.
[0125] Among them, the first number is the number of Monte Carlo simulations for the second intention.
[0126] Determining the score for the aforementioned intent corresponds to the backpropagation phase of the MCTS algorithm. For each node corresponding to the intent, after completing the Monte Carlo simulation, the number of simulations must be updated. At the same time, the score must be accumulated. This means the fourth score obtained from the current Monte Carlo simulation is accumulated based on the historical scores. Finally, the average score is calculated based on the accumulated final score and the updated number of simulations.
[0127] The following is based on Figure 3 The intent tree shown is used as an example.
[0128] Current intent tree structure: Root node (loan intention identification); ├── Query loan amount (N=3, W=15, Q=5); ├── Query loan interest rate (N=4, W=20, Q=5); └── Query repayment method (N=2, W=8, Q=4).
[0129] Among them, N represents the number of simulations, W represents the cumulative reward, and Q represents the average reward.
[0130] The results of the simulation phase are as follows.
[0131] Simulation path: root node → query loan interest rate → type of loan interest rate.
[0132] The back-propagation steps are as follows.
[0133] 1. Update the node: The parameters corresponding to the "Type of Loan Interest Rate" node are: N=1, W=10, Q=10.
[0134] 2. Update the parent node, that is, the "Query Loan Interest Rate" node (if there is a parent node, continue to update): N=5 (original value 4 increased by 1); W=30 (original value 20 increased by 10); Q=30 / 5=6 (recalculate the average reward).
[0135] Updated intent tree structure: Root node (loan intention identification); ├── Query loan amount (N=3, W=15, Q=5); ├── Query loan interest rate (N=5, W=30, Q=6); │└── Type of loan interest rate (N=1, W=10, Q=10); └── Query repayment method (N=2, W=8, Q=4).
[0136] As can be seen, by feeding the scoring results of the Monte Carlo simulation phase back to the relevant nodes in the tree, and by updating the number of simulations and accumulated rewards, the tree structure and selection strategy are optimized. Ultimately, the system can gradually converge to a tree structure that can better identify user intent.
[0137] Step S509 : determining a target intent from the second intent, the plurality of third intents, and the plurality of fourth intents according to the plurality of first scores and the second score of each fourth intent.
[0138] It is understandable that for the explanation of the aforementioned steps related to Example 1 and Example 2, please refer to the contents of Example 1 and Example 2, and no further details will be given here.
[0139] It can be seen that the present application provides an intent recognition method that extracts features from the first conversation input by the user and predicts multiple first intentions using a multi-label classification model, dynamically calculates the exploration coefficient of each first intention at the first moment to adjust the exploration and utilization balance in the MCTS algorithm, calculates the upper confidence limit value in combination with the exploration coefficient and selects the optimal second intention, further expands the second intention to generate multiple third intentions, evaluates the value of the second and third intentions through Monte Carlo simulation, determines the target intention by combining the scores of all the above intentions, and feeds the simulation results back to the intent tree to optimize the decision-making strategy, thereby achieving the technical effect of significantly improving the accuracy and robustness of intent recognition. This method enhances the ability to understand complex user needs by dynamically adjusting the exploration coefficient and expanding the intent tree structure, while optimizing the decision path through the backpropagation mechanism, thereby improving the user experience.
[0140] In accordance with the above-mentioned embodiment, please refer to Figure 6 , Figure 6 This is a functional unit block diagram of an intention recognition device provided in an embodiment of the present application. The intention recognition device is the above-mentioned server or a part of the server, such as Figure 6 As shown, the intention recognition device 60 includes: A first processing unit 601 is configured to determine, based on multiple first intentions corresponding to the first conversation, an exploration coefficient for each first intention at a first moment; The first processing unit 601 is further configured to determine an upper confidence limit corresponding to each first intention according to the exploration coefficient of each first intention; A second processing unit 602 is configured to select a second intent from the plurality of first intents according to the upper confidence limit corresponding to each first intent and expand the second intent to obtain a plurality of third intents; A third processing unit 603 is configured to perform a Monte Carlo simulation based on the second intent and the plurality of third intents to determine a first score for the second intent and a first score for each of the plurality of third intents; a fourth processing unit 604 for determining a target intent from the second intent, the plurality of third intents, and the plurality of fourth intents based on the plurality of first scores and the second score of each fourth intent; The fourth intention is an intention in the first intention that is different from the second intention, and the second score is an average score of the fourth intention recorded before the first moment.
[0141] In a feasible embodiment, in determining the exploration coefficient of each first intention at the first moment based on the multiple first intentions corresponding to the first conversation, the second processing unit 602 is specifically configured to: Determine an exploration coefficient for each first intent based on the conversation round corresponding to the first moment in the multi-round conversation, the behavioral entropy of the second conversation, and the intent confidence of each first intent in the multiple first intents corresponding to the first conversation; The second conversation is used to represent a conversation with a target user before the first moment, and the target user is the user corresponding to the first conversation.
[0142] In a feasible embodiment, in determining the exploration coefficient of each first intent based on the conversation turn corresponding to the first moment in the multi-round conversation, the behavioral entropy of the second conversation, and the intent confidence of each first intent in the multiple first intents corresponding to the first conversation, the second processing unit 602 is specifically configured to: Constructing a first function based on the natural index and the conversation turn corresponding to the first moment in the multi-round conversation; constructing a second function based on a difference between the first preset value and the intent confidence of each first intent in the plurality of first intents corresponding to the first conversation; determining a third function according to the first function, the second function, the behavioral entropy of the second conversation, and the second preset value; An exploration coefficient of each first intention is determined according to the preset exploration coefficient and the third function.
[0143] In a feasible embodiment, the second processing unit 602 is further configured to: determining a plurality of fifth intents included in the second conversation and a number of occurrences of each of the plurality of fifth intents; determining an occurrence frequency of each fifth intent based on a ratio between an occurrence count of each fifth intent and a total occurrence count corresponding to the plurality of fifth intents; Based on the occurrence frequency of each fifth intention, the behavioral entropy is determined.
[0144] In one feasible embodiment, in performing a Monte Carlo simulation based on the second intent and the plurality of third intents to determine the first score of the second intent and the first score of each of the plurality of third intents, the third processing unit 603 is specifically configured to: determining a plurality of intention combinations, each intention combination including a second intention and a third intention; Performing a Monte Carlo simulation based on each intention combination to determine a third score for each intention combination; Determine, based on the third score of each intention combination, a fourth score of the second intention corresponding to each intention combination and a first score of the third intention; A first score for the second intent is determined based on the fourth score for the second intent corresponding to each intent combination.
[0145] In a feasible embodiment, in determining the first score of the second intent based on the fourth score of the second intent corresponding to each intent combination, the third processing unit 603 is specifically configured to: Determine the sum of the fourth score and the fifth score of the second intention corresponding to each intention combination, where the fifth score is the cumulative score of the second intention recorded before the first moment; A first score for the second intent is determined based on a ratio between the score sum and the first score, where the first score is the number of Monte Carlo simulations for the second intent.
[0146] In a feasible embodiment, in performing a Monte Carlo simulation based on each intention combination to determine the third score of each intention combination, the third processing unit 603 is specifically configured to: Perform Monte Carlo simulation based on each intention combination to determine multiple scoring factors for each intention combination; A weighted sum is performed on the multiple scoring factors of each intent combination to determine a third score of each intent combination.
[0147] It can be understood that since the method embodiment and the device embodiment are different presentation forms of the same technical concept, the content of the method embodiment part in this application should be synchronously adapted to the device embodiment part and will not be repeated here.
[0148] In the case of integrated units, such as Figure 7 As shown, Figure 7 This is a block diagram of the functional units of another intention recognition device 60 provided in an embodiment of the present application. Figure 7 In the embodiment, the intention recognition device 60 includes: a processing module 712 and a communication module 711. The processing module 712 is used to control and manage the actions of the intention recognition device 60, for example, the steps of the first processing unit 601, the second processing unit 602, the third processing unit 603 and the fourth processing unit 604, and / or other processes for performing the technology described herein. The communication module 711 is used to support the interaction between the intention recognition device 60 and other devices. Figure 7 As shown, the intention recognition device 60 may further include a storage module 713 , which is used to store program codes and data of the intention recognition device 60 .
[0149] The processing module 712 may be a processor or controller, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The communication module 711 may be a transceiver, an RF circuit, or a communication interface, and the like. The storage module 713 may be a memory.
[0150] Among them, all relevant contents of each scenario involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and will not be repeated here. The above intention recognition device 60 can perform the above Figure 2 The intent recognition method shown.
[0151] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product comprises one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. A computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.
[0152] Figure 8 This is a structural block diagram of an electronic device provided in an embodiment of the present application. Figure 8 As shown, the electronic device 800 may include one or more of the following components: a processor 801, a memory 802 and a communication interface 803. The processor 801, the memory 802 and the communication interface 803 are interconnected and perform communication with each other. The memory 802 may store one or more computer programs, and the one or more computer programs may be configured to implement the methods described in the above embodiments when executed by one or more processors 801.
[0153] The processor 801 may include one or more processing cores. The processor 801 uses various interfaces and lines to connect the various parts of the entire electronic device 800, and executes various functions and processes data of the electronic device 800 by running or executing instructions, programs, code sets or instruction sets stored in the memory 802, and calling data stored in the memory 802. Optionally, the processor 801 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 801 can integrate one or more combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. It is understandable that the above-mentioned modem may not be integrated into the processor 801, but may be implemented separately through a communication chip.
[0154] The memory 802 may include a random access memory (RAM) or a read-only memory (ROM). The memory 802 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 802 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The data storage area may also store data created by the electronic device 800 during use.
[0155] It is understandable that the electronic device 800 may include more or fewer structural elements than those in the above structural block diagram, for example, a power module, physical buttons, a WiFi (Wireless Fidelity) module, a speaker, a Bluetooth module, a sensor, etc., which are not limited here.
[0156] The electronic device 800 may be a server or a part of a server.
[0157] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed by a processor, implements part or all of the steps of any one of the vehicle control unit diagnostic methods described in the above method embodiments.
[0158] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements some or all of the steps of any of the vehicle control unit diagnostic methods described in the above method embodiments. The computer program product may be a software installation package.
[0159] It should be noted that for any of the aforementioned method embodiments of the intent recognition method, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0160] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components or steps. The fact that certain measures are recited in different dependent claims does not mean that these measures cannot be combined to produce good results.
[0161] A person skilled in the art will understand that all or part of the steps in the various methods of any of the above-mentioned methods for identifying intentions can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0162] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application's method, device, electronic device, and storage medium for identifying intent. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those skilled in the art, based on the idea of the present application's method, device, electronic device, and storage medium for identifying intent, there may be changes in the specific implementation methods and scope of application. In summary, the content of this specification should not be understood as a limitation on the present application.
[0163] The present application is described with reference to the flowcharts and / or block diagrams of the methods, hardware products, and computer program products of the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0164] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0165] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0166] It can be understood that any product that is controlled or configured to execute the processing method of the flowchart described in the method embodiment of an intention recognition method of the present application, such as the terminal and computer program product in the above flowchart, falls within the scope of the related products described in the present application.
[0167] Obviously, those skilled in the art may make various modifications and variations to the intent recognition method, apparatus, electronic device, and storage medium provided in this application without departing from the spirit and scope of this application. Thus, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for identifying intention, characterized in that: The method comprises: Determining, based on multiple first intentions corresponding to the first conversation, an exploration coefficient of each first intention at a first moment; Determining an upper confidence limit value corresponding to each first intention according to the exploration coefficient of each first intention; Selecting a second intent from the multiple first intents and expanding the second intent according to the upper confidence limit corresponding to each first intent to obtain multiple third intents; performing a Monte Carlo simulation based on the second intent and the plurality of third intents to determine a first score for the second intent and a first score for each of the plurality of third intents; determining a target intent from among the second intents, the third intents, and the fourth intents based on the first scores and the second score of each fourth intent; The fourth intention is an intention in the first intention that is different from the second intention, and the second score is an average score of the fourth intention recorded before the first moment.
2. The method according to claim 1, characterized in that The determining, based on the multiple first intentions corresponding to the first conversation, an exploration coefficient of each first intention at the first moment includes: Determining an exploration coefficient for each first intent based on a conversation round corresponding to a first moment in the multi-round conversation, the behavioral entropy of the second conversation, and the intent confidence of each first intent in the multiple first intents corresponding to the first conversation; The second conversation is used to represent a conversation with a target user before the first moment, and the target user is the user corresponding to the first conversation.
3. The method according to claim 2, characterized in that The determining, based on the conversation round corresponding to the first moment in the multi-round conversation, the behavioral entropy of the second conversation, and the intent confidence of each first intent in the multiple first intents corresponding to the first conversation, includes: constructing a first function based on a natural index and a conversation round corresponding to a first moment in the multiple conversation rounds; constructing a second function based on a difference between the first preset value and the intent confidence of each first intent in the plurality of first intents corresponding to the first conversation; determining a third function according to the first function, the second function, the behavioral entropy of the second conversation, and the second preset value; The exploration coefficient of each first intention is determined according to the preset exploration coefficient and the third function.
4. The method according to claim 2 or 3, characterized in that The method further comprises: determining a plurality of fifth intents included in the second conversation and a number of occurrences of each of the plurality of fifth intents; determining an occurrence frequency of each fifth intent based on a ratio between the number of occurrences of each fifth intent and the total number of occurrences corresponding to the plurality of fifth intents; The behavior entropy is determined according to the occurrence frequency of each fifth intention.
5. The method according to claim 1, wherein The performing of a Monte Carlo simulation based on the second intent and the plurality of third intents to determine a first score for the second intent and a first score for each of the plurality of third intents includes: determining a plurality of intention combinations, each intention combination including the second intention and the third intention; Performing a Monte Carlo simulation based on each intention combination to determine a third score for each intention combination; Determining, based on the third score of each intention combination, a fourth score of the second intention corresponding to each intention combination and a first score of the third intention; Determine a first score for the second intent based on the fourth score for the second intent corresponding to each intention combination.
6. The method according to claim 5, characterized in that The determining, based on the fourth score of the second intent corresponding to each intention combination, the first score of the second intent includes: Determine a sum of a fourth score and a fifth score of the second intention corresponding to each intention combination, where the fifth score is a cumulative score of the second intention recorded before the first moment; A first score for the second intention is determined based on a ratio between the score and a first number, where the first number is a number of Monte Carlo simulations of the second intention.
7. The method according to claim 5, characterized in that The performing of a Monte Carlo simulation based on each intention combination to determine a third score for each intention combination includes: Performing a Monte Carlo simulation based on each intention combination to determine a plurality of scoring factors for each intention combination; A weighted sum is performed on the multiple scoring factors of each intention combination to determine a third score of each intention combination.
8. An intention recognition device, characterized in that: The device comprises: a first processing unit, configured to determine, based on the plurality of first intentions corresponding to the first conversation, an exploration coefficient of each first intention at a first moment; The first processing unit is further configured to determine an upper confidence limit corresponding to each first intention according to the exploration coefficient of each first intention; a second processing unit, configured to select a second intent from the plurality of first intents and expand the second intent according to the upper confidence limit corresponding to each first intent, to obtain a plurality of third intents; a third processing unit, configured to perform a Monte Carlo simulation based on the second intent and the plurality of third intents to determine a first score for the second intent and a first score for each of the plurality of third intents; a fourth processing unit, configured to determine a target intent from the second intents, the plurality of third intents, and the plurality of fourth intents based on the plurality of first scores and the second score of each fourth intent; The fourth intention is an intention in the first intention that is different from the second intention, and the second score is an average score of the fourth intention recorded before the first moment.
9. An electronic device comprising a processor, a memory, and an executable program code stored in the memory, wherein: The processor is used to call the executable program code stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and device for selecting decision
CN114615680A
Intelligent customer service language model optimization method and system for instruction enhancement based on Monte Carlo tree search
CN118170884A
Intelligent doorbell visitor identification method and system based on deep learning
CN118761034A
Large-scale data generation method, device and equipment
CN119883657A