Dynamic intention understanding method based on multi-modal fusion
Through the dynamic intention understanding method of multimodal fusion, the use of large models to identify user intentions is solved, and the limitations of single modal analysis are achieved, and high-accurate intention recognition and personalized services are achieved in online learning scenarios.
Patent Information
- Application Number
- CN202511088840.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-08-05
AI Technical Summary
Existing intention recognition technologies mainly rely on single modal analysis, making it difficult to effectively integrate multi-dimensional user behavior data, and fail to clarify the intention type in online learning scenarios.
The dynamic intention understanding method of multimodal fusion is adopted to obtain user multimodal behavior data, identify intent through large models, use rewritten dialogue text and multimodal data, and combine intent type description and understanding skills to design business processes to provide useful information.
It improves the accuracy of intention recognition, realizes automatic identification and personalized services of user intentions, and dynamically adjusts the intention type to adapt to changes in user intentions, and provides more convenient information services.
Smart Images

Figure CN120579009A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence, and in particular to a dynamic intent understanding method based on multimodal fusion. Background Art
[0002] Intent understanding refers to identifying the purpose behind individual behavior. Introducing intent understanding into intelligent dialogue is more conducive to providing users with accurate and useful information.
[0003] Existing intent recognition technologies primarily rely on single-modality analysis (e.g., text), making it difficult to effectively integrate multi-dimensional user behavior data and failing to clearly define the specific intent types in online learning scenarios. Patent applications CN115186076A, which provides a method for extracting intent from text understanding, and CN119886352A, which provides a deep learning-based method for understanding conversational intent and generating explanatory text, fail to address these issues. Summary of the Invention
[0004] An embodiment of the present invention provides a dynamic intent understanding method based on multimodal fusion to solve the above technical problems.
[0005] In a first aspect, an embodiment of the present invention provides a dynamic intent understanding method based on multimodal fusion, comprising:
[0006] Obtaining multimodal behavioral data of the user in this conversation, wherein the multimodal behavioral data includes voice, image, and conversation text;
[0007] Rewrite the text of this conversation based on whether it is the first conversation;
[0008] The rewritten conversation text and other multimodal behavior data are provided to the big model, prompting the big model to identify the intent of the current conversation. Specifically, the prompt specifies the big model's role as an expert in understanding user conversation intent, specifies understanding techniques under typical conversation patterns, specifies multiple intent types to be identified and their descriptions, and instructs the big model to understand the rewritten conversation and other multimodal behavior data based on the specified role and information, and determine to which of the multiple intent types the current conversation belongs.
[0009] Call the corresponding business process based on the intent recognition result to provide useful information for this conversation;
[0010] Among them, the multiple intent types include question-and-answer intent, search intent, to-do task intent, learning task intent, lesson plan creation intent and PPT creation intent.
[0011] In a second aspect, an embodiment of the present invention provides an electronic device, comprising:
[0012] one or more processors;
[0013] a memory for storing one or more programs,
[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the dynamic intent understanding method based on multimodal fusion described in any embodiment.
[0015] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, which, when executed by a processor, implements the dynamic intent understanding method based on multimodal fusion described in any embodiment.
[0016] An embodiment of the present invention provides a dynamic intent understanding method based on multimodal fusion, which not only integrates multimodal information and historical conversations in intent recognition to improve the accuracy of intent recognition, but also uses a large model to realize automatic recognition of user intent. The large model prompt specifies the role of the large model as an expert in user conversation intention understanding, specifies the understanding skills under typical conversation modes, specifies multiple intent types to be identified and their descriptions, and provides key information for the large model to identify user intent. The large model is then instructed to understand the rewritten conversation and other multimodal behavior data according to the specified role and information, and determine which of the multiple intent types this conversation belongs to, so that the large model can accurately identify basic user intent; finally, a corresponding business process is designed for each intent type to provide users with effective information more conveniently. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is a flow chart of a dynamic intent understanding method based on multimodal fusion provided by an embodiment of the present invention;
[0019] Figure 2 1 is a schematic diagram of the structure of a reinforcement learning network provided by an embodiment of the present invention;
[0020] Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.
[0022] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0023] In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0024] Figure 1 This is a flow chart of a dynamic intention understanding method based on multimodal fusion provided by an embodiment of the present invention. This method is applicable to online learning scenarios and is executed by electronic devices. Figure 1 As shown, the method specifically includes:
[0025] S110: Acquire multimodal behavior data of the user in this conversation, where the multimodal behavior data includes voice, image, and conversation text.
[0026] As described above, this embodiment is applicable to online learning scenarios, such as an online learning platform. The platform provides users with an intelligent dialogue window, through which users can obtain useful information and complete various tasks on the online learning platform more conveniently.
[0027] Optionally, intelligent conversations first capture the text data entered by the user during the conversation. Simultaneously, they capture the user's voice and image data through the camera and audio equipment of a learning device (such as a computer or mobile phone). These three types of data are combined as multimodal behavioral data for the conversation. Voice data can provide information about the user's voice, tone, and other semantic information not included in the text data, while image data can provide information about the user's facial expressions and mental state.
[0028] S120. Rewrite the text of the conversation based on whether the conversation is the first conversation.
[0029] Optionally, a "Start New Conversation" button is provided in the smart conversation window. When the user clicks the button or logs into the smart conversation window for the first time, the conversation is considered to be the first conversation.
[0030] Furthermore, if this conversation is the first conversation, the text of this conversation will be directly used as the rewritten conversation text and involved in subsequent operations; if this conversation is not the first conversation, the historical conversation text between this conversation and the most recent first conversation will be extracted first, and these historical conversations will be combined with the current conversation to serve as the rewritten conversation.
[0031] S130: Provide the rewritten conversation text and other multimodal behavior data to the big model, prompting the big model to recognize the intention of this conversation.
[0032] This embodiment divides the online learning scenario into multiple intent types, including question-and-answer intent, search intent, to-do task intent, learning task intent, lesson plan creation intent, and PPT creation intent. These intent types cover several basic needs in online learning scenarios.
[0033] This embodiment utilizes a large model to identify conversational intent. The aforementioned intent types and their descriptions play a crucial role in the large model's prompts. Furthermore, the prompts must include the following information to fully ensure accurate intent recognition: The large model is designated as an expert in understanding user conversational intent; comprehension techniques for typical conversational patterns are specified; and the large model is instructed to interpret the rewritten conversation and other multimodal behavioral data based on the designated role and information, determining which of the multiple intent types the conversation belongs to.
[0034] For example, a prompt is as follows:
[0035] ## Specify roles and tasks
[0036] You are an experienced and professional expert in understanding and judging the intention of user questions. Based on the questions raised by the user this time, you will understand the question intention and, based on the given intention classification, feedback the true intention of the user in this conversation.
[0037] ## Understanding Techniques in Specified Typical Conversation Modes
[0038] 1. If specific intention words are clearly mentioned in the user's question, return according to the intention content mentioned by the user;
[0039] 2. If the intention expressed in the user's question highly coincides with the meanings expressed by multiple intentions in the intention classification description, feedback multiple matching intention classifications;
[0040] 3. If there is no clear intention in the user's question, by default, return according to the "Q&A Intention".
[0041] ## Specified Output Format Requirements
[0042] Output the result according to the format requirements (separated by <Format Requirements>< / Format Requirements>), ensuring that the output result can be correctly parsed by Python json.loads.
[0043] ## Specified Multiple Intention Types to be Recognized and Their Descriptions
[0044] <Intention Classification>
[0045] Q&A Intention: This intention is mainly that the user hopes to quickly solve the problem and hopes that you can answer the question professionally;
[0046] Search Intention: This intention is mainly that the user wants to actively query certain resources and hopes that you conduct a business search and give the results according to the requirements;
[0047] To-Do Task Intention: This intention is mainly that the user wants to find the to-do items / work set that they need to complete in the future and hopes that you can help find the task items to be solved and the quantity;
[0048] Learning Task Intention: This intention is mainly that the user wants to find the learning content set that they need to complete in the future, and it emphasizes more on the search for learning resources "required";
[0049] Teaching Plan Creation Intention: This intention is mainly that the user wants to create article content and hopes that you can carefully create a teaching article for me;
[0050] PPT Creation Intention: This intention is mainly that the user wants to write a PPT and hopes that you can plan and create a beautiful PPT for me.
[0051] < / Intention Classification>
[0052] ## Specify other parameter variables:
[0053] <User question>
[0054] {question}
[0055] < / User question>
[0056] <Format requirements>
[0057] ["Return the intention of this conversation here"]
[0058] < / Format requirements>
[0059] Among them, ## is used to mark the annotations of the prompt, < > and <!-- --> is used to indicate the parameter slots. The content filled between <intention classification> and < / intention classification> is a variety of pre-set intention types and their descriptions. The modified user conversation is filled between <user question> and < / user question>.
[0060] Based on the above prompt, the large language model can relatively accurately determine which one of the above-mentioned multiple intention types this conversation belongs to.
[0061] S140. Call the corresponding business process according to the intention recognition result to provide useful information for this conversation.
[0062] In this embodiment, different business processes are designed for different intention types to provide faster and more accurate personalized services.
[0063] Exemplarily, for the Q&A intention, the business process is as follows: first, call the large model to analyze and trace the knowledge points, and then retrieve in the knowledge vector library according to the knowledge points and sources to obtain useful information. For the search intention, the business process is: directly query the knowledge vector library to retrieve matching information; at the same time, execute the business process of the Q&A intention, and the two processes run in parallel to provide both direct search results and analyzed search results for the user to choose. For the process of retrieving to-do tasks, call the large model to parse the to-do objects in the conversation text, and remotely schedule the to-do task statistical data according to the to-do objects to feedback the to-do task list of specific objects to the user. For the learning task intention, call the large model to parse the to-be-learned objects in the conversation text, and remotely schedule the learning resource data according to the to-be-learned objects to feedback specific learning resources to the user. For the lesson plan creation intention, call the large model to parse the creation key points, search for relevant texts of the creation key points in the knowledge vector library, and feedback them to the large model for processing into a complete manuscript. For the PPT creation intention: call the large model to parse the creation key points, and at the same time plan the page layout, search for relevant texts and multimedia materials of the creation key points in the knowledge vector library, and feedback them to the large model for processing into the corresponding PPT.
[0064] The above process is executed in every conversation. This entire method not only integrates multimodal information and historical conversations to improve the accuracy of intent recognition, but also leverages the big model to automatically identify user intent. The big model prompts specify the big model's role as an expert in understanding user conversation intent, specifies understanding techniques for typical conversation patterns, and specifies multiple intent types to be identified and their descriptions. This provides key information for the big model to identify user intent. The big model is then instructed to interpret the rewritten conversation and other multimodal behavior data based on the specified role and information, determining which of the multiple intent types the current conversation belongs to, enabling the big model to accurately identify basic user intent. Finally, a corresponding business process is designed for each intent type to more conveniently provide users with effective information.
[0065] Furthermore, as mentioned above, the multiple intent types pre-set in the prompt are key to ensuring the large model accurately recognizes intent. To adapt to the dynamic changes in user conversation intent and to promptly adjust the large model's intent recognition capabilities, the multiple pre-set intent types can be dynamically adjusted. In one specific embodiment, this process may include the following steps:
[0066] Step 1: Generate a user intent analysis chain based on the user's acceptance of useful information in previous conversations. This intent analysis chain consists of the intent recognition results corresponding to each conversation, including the user's acceptance of useful information or the termination of the conversation chain. This embodiment uses the intent analysis chain to reflect the large-scale model's exploration of the user's true intent. The shorter the intent analysis chain, the more accurate the large-scale model's intent recognition.
[0067] Optionally, first identify the chain nodes used to divide the intent analysis chain. Specifically, if the user clicks on or copies useful information from a conversation, it is determined that the user has accepted the useful information from the conversation, and the conversation is marked as a chain node. If the user fails to enter the conversation text within a set time limit, or exits the conversation interface, or clicks on the option to start a new conversation, it is determined that the user has terminated the conversation chain, and the conversation is also marked as a chain node. There is a special case at this time: after the user clicks on or copies useful information from a conversation, and then fails to enter the conversation text within a set time limit, or exits the conversation interface, or clicks on the option to start a new conversation, then the conversation only needs to be marked as a link point.
[0068] After the chain nodes are marked, the intent recognition results corresponding to all previous conversations between the two chain nodes are arranged in sequence to form the intent analysis chain for the same user. It should be noted that the intent analysis chain between two chain nodes includes the next chain node in the chain, but does not include the previous chain node in the chain. In other words, each chain node is counted as the end point of the previous intent analysis chain and is not included in the next intent analysis chain.
[0069] Step 2: Adjust the multiple intent types according to the changes in the intent recognition results in the intent analysis chain, and apply the adjusted multiple intent types to subsequent large model prompts.
[0070] As described above, this embodiment uses an intent analysis chain to reflect the large model's exploration of the user's true intent. A shorter intent analysis chain indicates a more accurate large model's intent recognition, allowing it to discover true intent and provide useful information in fewer conversations. A longer intent analysis chain indicates that the large model requires multiple rounds of conversation before it can identify true intent and provide useful information. This indicates that the intent classification provided to the large model is likely incomplete, affecting the accuracy of intent analysis. In this case, it is necessary to consider adjusting the intent classification. It is worth noting that the true intent here refers to the user's actual intent and is not limited to the multiple intent types specified in the large model prompt. It may be an intent type that is close to, more specific to, or completely different from these intent types. As long as one of the multiple intent types specified in the large model prompt is very close to the true intent, the large model can more accurately identify it. Therefore, the purpose of this embodiment is to improve the closeness between the intent type in the prompt and the true intent. Optionally, this embodiment provides the following two optional implementations for dynamically adjusting the multiple types specified in the large model prompt (hereinafter referred to as "intent classification adjustment") to achieve this goal:
[0071] The first optional implementation applies to intent analysis chains where the user accepts useful information from the last conversation. If the length of the intent analysis chain exceeds the set threshold, it indicates that the user's intent exploration process is too long and it takes multiple conversations to obtain trustworthy useful information. In this case, intent classification adjustments can be made based on the following two situations:
[0072] Case 1: All intent recognition results in the intent analysis chain are the same, indicating that the large model correctly identifies the user's intent but fails to provide credible and useful information in the early stage. This is most likely because the intent classification is not detailed enough, causing the user to have more in-depth conversations. In this case, if all intent recognition results in the intent analysis chain are the same, the same intent type corresponding to each intent recognition result can be subdivided according to the last two conversations in the intent analysis chain. Optionally, keyword analysis can be performed on the last two conversations separately to determine the keyword with the greatest semantic difference between the last conversation and the penultimate conversation, and the same intent type in the current intent analysis chain can be subdivided based on the keyword. For example, if all intent recognition results in the intent analysis chain are search intents, the content of the penultimate conversation is "Please help me search for content about rescue points in the flood control and disaster relief course," and the content of the last conversation is "Please help me search for content about rescue practice in the flood control and disaster relief course," then the keyword with the greatest semantic difference between the last conversation and the penultimate conversation is "practical." Based on this keyword, the "search intent" can be subdivided into "practical search intent" and "theoretical search intent," and descriptions of the subdivided intents and corresponding business processes can be added to each. In this way, the big model can pay attention to information about practical and theoretical aspects as early as possible in each conversation, identify the user's more detailed true intent as early as possible, or present both "practical search intent" and "theoretical search intent" to the user earlier, allowing the user to directly select the true subdivided intent. The various operations described above, from keyword segmentation to business process design, can be automatically implemented with the help of the big model, or can be implemented by developers based on experience, or implemented through a combination of automation and human experience. This embodiment does not impose specific limitations.
[0073] Case 2: All intent recognition results in the intent analysis chain are not identical, indicating that the large model, after multiple explorations of the user's intent, ultimately identified the correct intent and provided trustworthy and useful information. In this case, this embodiment focuses on the turning point from incorrect intent to correct intent to correct the most basic intent direction. Optionally, the last intent recognition result in the intent analysis chain can be extracted as a continuous repetition segment at the end of the intent analysis chain; and the intent type of the last conversation before the continuous repetition segment is subdivided based on the first conversation in the continuous repetition segment. For example, if the intent analysis chain is "Q&A, Search, To-Do Task, Learning Task, Learning Task," the last intent recognition result is a learning task, and this intent is repeated twice at the end of the intent analysis chain. Its continuous repetition segment is the final "Learning Task, Learning Task" segment; the first conversation in this segment is the second-to-last conversation, and the last conversation before this segment is the third-to-last conversation. The third-to-last intent type is subdivided into "To-Do Task" based on the content of the second-to-last conversation. Optionally, keyword analysis can be performed on each of these two conversations to determine the keywords with the greatest semantic difference between the penultimate and penultimate conversations, and the "to-do tasks" intent type can be subdivided based on these keywords. For example, the content of the penultimate conversation is "Please help me find the unfinished courses," and the content of the penultimate conversation is "Please help me preview the unfinished course content." The keywords with the greatest semantic difference between the penultimate conversation and the penultimate conversation are "preview" and "content." The "to-do tasks" intent can then be subdivided into "to-do task list" and "to-do task content preview." Corresponding descriptions can be added to the subdivided intents, such as a description for the to-do task list as "providing a task list to the user," and a description for the to-do task content preview as "providing a task list and a thumbnail preview of the content of each task to the user," to distinguish the meanings of the two intents. At the same time, corresponding business processes can also be subdivided and designed. For example, the business process for to-do task content preview includes a to-do task content search and thumbnail creation process compared to the business process for the to-do task list. Similarly, the above-mentioned operations from keyword classification to business process design can be automatically implemented with the help of a large model, or can be implemented by developers based on experience, or through a combination of automation and human experience. This embodiment does not impose specific restrictions.
[0074] The second optional implementation applies to intent analysis chains where the user terminates the conversation chain. If the length of the intent analysis chain exceeds the set threshold, it indicates that the user's exploration process is too long and no reliable and useful information has been obtained after multiple conversations. In this case, it is very likely that the intent recognition is inaccurate, resulting in a series of invalid questions and answers. In this case, the intent classification can be adjusted according to the following two situations:
[0075] Case 1: The intent recognition results in the intent analysis chain cover multiple existing intent types, indicating that none of the existing intent types touches on the user's true intention. In this case, new intent types can be added based on the content of the user's previous conversations. Optionally, all conversations in the intent analysis chain can be keyword parsed, and association rules can be used to determine the most frequently occurring keywords in all conversations, and new intent types can be added based on these keywords. For example, if the most frequently occurring keywords in all conversations are "bidding plan", "technical indicators", "business indicators", etc., a new intent type "bidding plan generation intention" can be added, and the corresponding business process can be designed. Similarly, the various operations from keyword division to business process design mentioned above can be automatically implemented with the help of a large model, or can be implemented by developers based on experience, or can be implemented through a combination of automation and human experience. This embodiment does not impose specific restrictions.
[0076] Case 2: If the intent recognition results in the intent analysis chain do not cover multiple existing intent types, the user can be asked questions based on the intent types not covered by the intent analysis chain to determine whether the user belongs to the uncovered intent types. For example, if "PPT creation intention" does not appear in the current intent analysis chain, the user can be asked "Do you want to create a PPT?" If the user answers in the affirmative, there may be useful information that can be trusted.
[0077] The above embodiment focuses on the intention analysis chain that is too long in the user exploration process. Based on the different exploration paths shown in the intention analysis, it analyzes the real intention identification obstacles that may be caused by intention classification, and proposes specific methods to eliminate such causes, providing improvement directions for the rationality and accuracy of intent classification.
[0078] However, when the above-mentioned intention exploration process is too long, in addition to the fact that intention classification may hinder the identification of true intent, factors such as inappropriate user conversation content and uncertainty in user intention may also hinder the identification of true intent. Therefore, whether the adjusted intention classification is truly beneficial to true intent analysis after being applied to the large model prompt requires further verification. To this end, this embodiment uses the multiple intent types after each adjustment as an intent classification scheme and constructs a reinforcement learning network to determine the optimal intent classification scheme. For example, in the prompt example of S130, the combination of the six intent types of question and answer intent, search intent, to-do task intent, learning task intent, lesson plan creation intent, and PPT creation intent is one intent classification scheme; in another embodiment, after refining the search intent into practical search intent and theoretical search intent, the combination of the seven intent types of question and answer intent, practical search intent, theoretical search intent, to-do task intent, learning task intent, lesson plan creation intent, and PPT creation intent is another intent classification scheme.
[0079] In a specific embodiment, the reinforcement learning network uses the user's intention analysis chain in actual conversation as the state variable and various intention classification schemes as the action variable, and constructs a reward function based on the final result and average length of the user's intention analysis chain to select the optimal intention classification scheme and apply it to the subsequent large model prompts. Optionally, the higher the coverage rate of the final results of the intention analysis chain of all users in a period of time for the used intention classification scheme, the higher the positive reward; the shorter the average length of the intention analysis chain of all users in a period of time, the higher the positive reward. Exemplarily, the following reward function can be constructed:
[0080]
[0081] in, Represents the reward value, Indicates the number of intent types in the currently used intent classification scheme covered by the final results of all user intent analysis chains within a period of time. Indicates the total number of intent types in the intent classification scheme, Indicates the average length of the intention analysis chain of all users in a period of time, and are positive proportional adjustment coefficients respectively.
[0082] Optionally, the user's intent analysis chain in the state variable can be encoded as a vector whose dimension is greater than the maximum length of the historical intent analysis chain. Each intent type corresponds to a non-zero value. The first dimension of the vector is the non-zero value corresponding to the first intent recognition result in the intent analysis chain, and the second dimension is the non-zero value corresponding to the second intent recognition result in the intent analysis chain. And so on, until all intent recognition results in the intent analysis chain have corresponding vector elements, and the remaining vector elements are filled with 0.
[0083] Optionally, the structure of the reinforcement learning network is as follows Figure 2As shown, the network takes the intention analysis chains (i.e., state variables) of all students in a period of time as batch input, and outputs the optimal intention classification scheme (i.e., action variables) to be adopted in the following period of time. The network includes a feature extraction layer and a Q-value calculation layer. The feature extraction layer is used to encode the input features to obtain the corresponding encoding features; the Q-value calculation layer is used to output the Q-value vector based on the encoding features. Each element in the Q-value vector is the probability of selecting each intention classification scheme, and the intention classification scheme with the highest probability is the action variable finally output by the reinforcement learning network. Optionally, both the feature extraction layer and the Q-value calculation layer adopt a fully connected structure, in which the number of nodes in the feature extraction layer gradually increases to enrich the feature information layer by layer; the number of nodes in the Q-value calculation layer gradually decreases to adapt to the dimension of the Q-value vector. The specific number of layers and nodes can be flexibly set as needed. The network updates the Q value in the following way:
[0084] First, calculate :
[0085]
[0086] in, Indicates that the status Next action Q value, Indicates that the status Take action The updated Q value, Representation status The largest of all Q values, Indicated by the state Take action arrive Rewards received after and All are adjustable coefficients.
[0087] Then, according to and The difference between them constructs the following loss function term :
[0088]
[0089] according to By minimizing , the network parameters of the feature extraction layer and the Q value calculation layer can be updated.
[0090] Based on this basic algorithm, the reinforcement learning network gradually learns the optimal intent classification scheme that truly improves intent recognition during continuous user conversations, observing the changes in the intent analysis chain brought about by each intent classification scheme. This optimal scheme is then applied to subsequent large-scale model prompts, resulting in even better intent recognition and user conversation results for the large model. Each new intent classification scheme is added by simply adding the corresponding node to the last layer of the Q-value calculation layer and performing a fully connected operation with the previous layer. The network will gradually update its parameters based on the conversational progress.
[0091] Of course, in addition to the above-mentioned reinforcement learning network, various intent classification schemes can also be evaluated based on user feedback to timely evaluate intent classification and avoid the effect of adjustment being degraded or unstable.
[0092] Furthermore, in another specific embodiment, taking into account the differences among users, the above-mentioned operations of adjusting and evaluating the intent classification scheme can also be performed separately for each user, that is, an intent analysis chain for each user is generated separately based on each user's acceptance of useful information in previous conversations; the multiple intent types for each user are adjusted separately based on the changes in the intent recognition results in each user's intent analysis chain, and the multiple intent types adjusted for each user are applied separately to each user's subsequent large model prompts.
[0093] Accordingly, for each user, the multiple intent types after each adjustment are regarded as an intent classification scheme; and a reinforcement learning network is constructed for each user, with the current user's intent analysis chain as the state variable and the various intent classification schemes of the current user as the action variable. A reward function is constructed according to the final result and average length of the current user's intent analysis chain to select the optimal intent classification scheme for the current user and apply it to the large model prompt for the current user.
[0094] The advantage of this approach is that the intent classification scheme can be adjusted for each user, adapting to their personalized learning habits while avoiding the potential inappropriate impact of applying intent classification adjustments for a single user to all users. The disadvantage is that it requires more storage space and computing resources. In actual applications, the implementation method can be flexibly selected based on the actual system conditions.
[0095] In addition, it should be noted that the user data involved in this application (including voice data, image data, and conversation text data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0096] Figure 3A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown, the device includes a processor 60, a memory 61, an input device 62 and an output device 63; the number of processors 60 in the device can be one or more. Figure 3 In the embodiment, a processor 60 is used as an example; the processor 60, the memory 61, the input device 62 and the output device 63 in the device can be connected by a bus or other means. Figure 3 The bus connection is taken as an example.
[0097] Memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the dynamic intent understanding method based on multimodal fusion in the embodiments of the present invention. Processor 60 executes the software programs, instructions, and modules stored in memory 61 to execute various functional applications and data processing of the device, thereby implementing the aforementioned dynamic intent understanding method based on multimodal fusion.
[0098] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, and these remote memories may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0099] The input device 62 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 63 may include a display device such as a display screen.
[0100] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the dynamic intent understanding method based on multimodal fusion of any embodiment.
[0101] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device.
[0102] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0103] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0104] Computer program code for performing the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.
Claims
1. A dynamic intent understanding method based on multimodal fusion, characterized in that: include: Obtaining multimodal behavioral data of the user in this conversation, wherein the multimodal behavioral data includes voice, image, and conversation text; Rewrite the text of this conversation based on whether it is the first conversation; The rewritten conversation text and other multimodal behavior data are fed to the big model to prompt it to recognize the intent of the conversation. Specifically, the prompt specifies the role of the large model as an expert in understanding user conversation intent, specifies understanding techniques in typical conversation modes, specifies multiple intent types to be identified and their descriptions, and instructs the large model to understand the rewritten conversation and other multimodal behavior data based on the specified role and information, and determine to which of the multiple intent types the current conversation belongs. Call the corresponding business process based on the intent recognition result to provide useful information for this conversation; Among them, the multiple intent types include question-and-answer intent, search intent, to-do task intent, learning task intent, lesson plan creation intent and PPT creation intent.
2. The method according to claim 1, characterized in that The above-mentioned rewriting of the text of the conversation based on whether this conversation is the first conversation includes: If this is the first conversation, the text of this conversation will be used as the rewritten text; If this conversation is not the first conversation, the text of this conversation and the text of the historical conversation will be combined to form the rewritten text of the conversation.
3. The method according to claim 1, characterized in that After calling the corresponding business process based on the intent recognition result to provide useful information for this conversation, it also includes: Generate a user intention analysis chain based on the user's acceptance of useful information in previous conversations, where the intention analysis chain is composed of the intention recognition results corresponding to the user's acceptance of useful information in each conversation or the termination of the conversation chain; According to the changes in the intention recognition results in the intention analysis chain, the multiple intention types are adjusted, and the adjusted multiple intention types are applied to subsequent large model prompts.
4. The method according to claim 3, characterized in that Generating a user intention analysis chain based on the user's acceptance of useful information in previous conversations also includes: If the user clicks or copies useful information from a conversation, it is determined that the user has accepted the useful information from the conversation, and the conversation is marked as a link node; If the user does not enter the conversation text within the set time limit, or exits the conversation interface, or clicks the option to start a new conversation, it is determined that the user has terminated the conversation chain and the conversation is marked as a chain node; The intention recognition results corresponding to all previous conversations of the same user between two chain nodes are arranged in sequence to form an intention analysis chain of the same user.
5. The method according to claim 3, characterized in that The adjusting of the multiple intent types according to the changes in the intent recognition results in the intent analysis chain includes: When the end point of the intent analysis chain is that the user has accepted the useful information from the last conversation: If all intent recognition results in the intent analysis chain are the same, the same intent types corresponding to the intent recognition results are subdivided according to the last two conversations in the intent analysis chain; Otherwise, extract the continuous repetition segment at the end of the intention analysis chain as the last intention recognition result in the intention analysis chain; and subdivide the intention type of the last conversation before the continuous repetition segment according to the first conversation of the continuous repetition segment.
6. The method according to claim 3, characterized in that The adjusting of the multiple intent types according to the changes in the intent recognition results in the intent analysis chain includes: When the end point of the intent analysis chain is the user terminating the conversation chain: If the intent recognition results in the intent analysis chain cover the multiple intent types, add new intent types based on the user's previous conversation content; Otherwise, questions are asked to the user based on the intent types not covered by the intent analysis chain to determine whether the user belongs to the uncovered intent types.
7. The method according to claim 3, after adjusting the multiple intent types based on changes in the intent recognition results in the intent analysis chain, further comprising: The multiple intention types after each adjustment are used as an intention classification scheme; Build a reinforcement learning network with the user's intention analysis chain as the state variable and various intention classification schemes as the action variables. Build a reward function based on the final result and average length of the user's intention analysis chain to select the optimal intention classification scheme and apply it to subsequent large-scale model prompts.
8. The method according to claim 3, characterized in that Generating a user intention analysis chain based on the user's acceptance of useful information in previous conversations includes: generating each user's intention analysis chain based on the user's acceptance of useful information in previous conversations; Correspondingly, the multiple intent types are adjusted according to the changes in the intent recognition results in the intent analysis chain, and the adjusted multiple intent types are applied to the subsequent large model prompts, including: adjusting the multiple intent types for each user according to the changes in the intent recognition results in the intent analysis chain of each user, and applying the adjusted multiple intent types for each user to the subsequent large model prompts of each user.
9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the dynamic intent understanding method based on multimodal fusion described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, which, when executed by a processor, implements the dynamic intent understanding method based on multimodal fusion as described in any one of claims 1-8.
Citation Information
Patent Citations
Agricultural intelligent question and answer method and system supporting multi-mode and multi-round dialogues
CN117370506A
Multi-robot dialogue method and system for software-as-a-service platform
CN117573834A
Question and answer method, device and equipment based on large model and storage medium
CN118246537A
User incoming call intention recognition method and device and readable storage medium
CN118899006A
Multi-modal intelligent knowledge retrieval and dialogue system and method
CN119003872A