Method and system for identifying implicit intentions in open-ended questions based on multi-agent reasoning
Through the multi-agent inference method, multilingual background knowledge and dynamic instantiation of multi-field expert agents for collaborative reasoning, the complex semantic understanding and deep inference problems of implicit intention recognition in open-ended problems are solved, and the identification of implicit intentions is achieved more accurately and comprehensively.
Patent Information
- Application Number
- CN202510138758.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-08
AI Technical Summary
The prior art has problems of complex semantic understanding, deep reasoning, and insufficient adaptability of dynamic environments when dealing with implicit intention recognition in open-ended problems.
Using a multi-agent reasoning method, multi-language background knowledge is constructed through multi-language translation agents and information retrieval agents, dynamically instantiate multi-field expert agents for collaborative reasoning, preside over management agents to guide discussions and conduct quality evaluations, and build an implicit intention library.
A more comprehensive and in-depth understanding and reasoning of open-ended problems has been achieved, which significantly improves the ability to adapt to the dynamic environment and accurately identify the implicit intentions in open-ended problems.
Smart Images

Figure CN119621916B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing technology and intelligent knowledge retrieval technology, and in particular to a method and system for identifying implicit intentions in open-ended questions based on multi-agent reasoning. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Implicit intent recognition aims to understand and identify the real intention expressed in the natural language information input by the user. In particular, for open-ended questions, that is, questions where the user's intention is not clear, the answer is not unique or not preset, accurately identifying its implicit intention is of great significance for predicting user behavior, exploring user needs, and then formulating effective decision-making plans. However, due to the diversity, ambiguity and uncertainty of implicit intentions in open-ended questions, existing methods are difficult to accurately identify and adapt in dynamic and complex application environments.
[0004] Implicit intent recognition based on knowledge graphs usually relies on pre-built, domain-specific structured knowledge bases. This approach has the following limitations when dealing with open-ended questions: First, the construction and update of knowledge graphs requires a lot of manpower and time costs, and it is difficult to reflect the rapid changes in the real world and the knowledge updates in emerging fields in real time. Second, the coverage of knowledge graphs is limited by the knowledge domains and data sources selected during construction, and it is difficult to handle open-ended questions that exceed the predefined knowledge scope, resulting in the inability to accurately identify user intent. Finally, such methods usually map user questions to fixed nodes or relationships in the knowledge graph, lacking a deep understanding of the semantics of the questions and multi-step reasoning capabilities, and are difficult to deal with the complex and changeable semantic expressions and implicit logical relationships in open-ended questions. In particular, when user intent requires multiple rounds of reasoning combined with context to be clear, knowledge graph-based methods are often powerless.
[0005] Methods based on BERT pre-trained language models usually require pre-definition of limited intent categories and rely on large-scale annotated data sets for training. However, when dealing with open-ended questions, user intent is often diverse and difficult to exhaust, and the preset intent categories are difficult to cover all possible intent expressions, resulting in a decrease in the model's recognition accuracy for novel or ambiguous intents. In addition, models such as BERT mainly "read" massive amounts of text data, namely "large-scale corpora", to find statistical regularities between words and sentences, namely "statistical patterns"; for example, it will find that the word "apple" often appears with words such as "eat" and "fruit", so it "remembers" this association. Although it has improved in understanding semantics, its recognition and understanding capabilities are obviously insufficient for open-ended questions that require complex semantic reasoning, logical reasoning, or judgment combined with common sense. In particular, when the user's intent is implicit and jumpy, BERT-based methods are difficult to accurately capture its deep semantics and logical relationships, and cannot perform effective multi-step reasoning to reveal the user's hidden true intentions.
[0006] In summary, existing methods have obvious limitations when dealing with implicit intent recognition in open-ended questions, especially in terms of reasoning ability. In actual application scenarios, the situation is often very complicated and the way users ask questions varies. It is difficult for existing methods to accurately identify what users want to express. Summary of the invention
[0007] In order to solve at least one technical problem existing in the above-mentioned background technology, the present invention provides a method and system for identifying implicit intentions in open-ended questions based on multi-agent reasoning, which effectively solves the deficiencies in processing complex semantic understanding, deep reasoning, and adaptability to dynamic environments, and more accurately and comprehensively identifies implicit intentions in open-ended questions.
[0008] In order to achieve the above object, the present invention adopts the following technical solution:
[0009] A first aspect of the present invention provides a method for identifying implicit intentions in open-ended questions based on multi-agent reasoning, comprising the following steps:
[0010] Use multilingual translation agents to translate input open-ended questions;
[0011] Using information retrieval agents to submit query requests using the translation results of open-ended questions as keywords, integrating the obtained query content to construct multilingual background knowledge for open-ended questions;
[0012] Based on the constructed multilingual background knowledge, we conduct topic analysis, generate multi-domain expert agent information according to the topic recognition results, and dynamically construct multi-domain expert agents;
[0013] According to the guiding questions raised by the host agent, multi-domain expert agents are used to generate their own implicit intentions and reasoning text content based on the guiding questions raised by the host agent;
[0014] The host agent is used to evaluate the quality of the implicit intention and reasoning text content. The speech of each screened expert agent is taken as a node to construct an implicit intention library.
[0015] Furthermore, the acquisition of query content includes: the information retrieval agent first obtains the relevant web page links returned by the search engine, then obtains the HTML content of these web pages through network requests, and uses the HTML parser to extract the main text data in the web pages.
[0016] Furthermore, the construction process of the multi-domain expert agent includes: using multilingual background knowledge as prompt words, utilizing a large language model, and dynamically generating a number of domain expert agents based on the topic analysis results of the multilingual background knowledge and predefined expert information templates, wherein the predefined expert information templates include the agent name, agent role, agent native language, and agent background description information.
[0017] Furthermore, the multi-domain expert agents are used to generate respective implicit intentions and reasoning text content according to the guiding questions raised by the host agent, including:
[0018] The host agent asks guiding questions based on the original question and the prompt words of the original question;
[0019] Expert agents in various fields express their own implicit intentions in response to the guiding questions. When each agent speaks, they conduct a comprehensive analysis of the current guiding question, the previous agent's implicit intentions, reasoning, their own role and background description information, and external knowledge retrieved through their native language. After each agent finishes speaking, it will generate its implicit intentions, reasoning, and links to cited external resources.
[0020] Furthermore, the method also includes updating the implicit intent library, specifically:
[0021] When a new node needs to be added to the temporary intent library, the implicit intent is converted into an implicit intent vector. The host agent calculates the cosine similarity between the implicit intent vector in the new node and the implicit intent vector in the existing node, and finds the N nodes with the highest similarity. According to the preset instructions and format requirements, based on the implicit intent vector in the new node and the information of the N nodes, the relationship between the new node and the N nodes is determined.
[0022] Furthermore, the Embedding technology is used to vectorize the implicit intent and convert each implicit intent text into a vector of fixed dimension.
[0023] Furthermore, after each round of speeches, the host management agent generates new guiding questions based on the current temporary intention library and its preset guiding question generation rules, and starts the next round of discussion. The discussion process will be iterated until the preset round limit is reached.
[0024] Furthermore, when using the host agent to evaluate the quality of each round of implicit intention and reasoning text content, multiple evaluation dimensions are set. For each evaluation dimension, the user's question, implicit intention, evaluation dimension and scoring criteria are taken as input, combined with the constructed LLM model, to evaluate the implicit intention, and output the score and reason.
[0025] Further, the scoring criteria are: 1: Very superficial; only a basic overview is provided, with obvious gaps in exploration; 2: Superficial; some details are provided, but many important aspects are omitted; 3: Moderate depth; key aspects are covered, but detailed exploration may be lacking in some areas; 4: Good depth; most aspects are explored in detail with only a few gaps; 5: Excellent depth; all relevant aspects are explored comprehensively and in depth, reflecting an in-depth and dynamic discussion.
[0026] A second aspect of the present invention provides an open-ended question implicit intention recognition system based on multi-agent reasoning, comprising:
[0027] A multilingual background knowledge construction module is used to translate the input open-ended questions using a multilingual translation agent; an information retrieval agent is used to submit a query request using the translation results of the open-ended questions as keywords, and the obtained query contents are integrated to construct multilingual background knowledge for the open-ended questions;
[0028] The multi-domain agent collaborative reasoning module is used to perform topic analysis based on the constructed multi-language background knowledge, generate multi-domain expert agent information according to the topic recognition results, and dynamically construct multi-domain expert agents;
[0029] The implicit intention recognition module is used to generate their own implicit intentions and reasoning text content according to the guiding questions raised by the host intelligent agent, using multi-domain expert intelligent agents to speak in turn according to the guiding questions raised by the host intelligent agent; the host intelligent agent is used to evaluate the quality of the implicit intentions and reasoning text content, and the speech of each screened expert intelligent agent is used as a node to construct an implicit intention library.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] The present invention aims to solve the problem that existing methods for identifying implicit intentions in open-ended questions are difficult to effectively handle complex semantics, conduct deep reasoning, adapt to dynamic environments, and identify multi-dimensional intentions. The invention constructs multi-language background knowledge for user open-ended questions; dynamically instantiates multiple domain expert agents based on multi-language background knowledge, and conducts multiple rounds of "conference"-style collaborative reasoning under the guidance of the host management agent to deeply explore the potential intentions of user questions; finally, the host management agent conducts quality demonstration based on multi-dimensional indicators on the potential intentions generated during the discussion to ensure the accuracy and reliability of the output results. The invention effectively solves the shortcomings of existing methods in handling complex semantic understanding, deep reasoning, and adaptability to dynamic environments, thereby more accurately and comprehensively identifying the implicit intentions in open-ended questions and significantly improving the ability to adapt to dynamic environments.
[0032] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0034] Figure 1 is a flow chart of a method for identifying implicit intentions in open-ended questions based on multi-agent reasoning provided by an embodiment of the present invention;
[0035] Figure 2 It is a schematic diagram of the internal process of the multilingual background knowledge construction module provided by an embodiment of the present invention;
[0036] Figure 3 It is a schematic diagram of the internal process of the multi-domain intelligent agent collaborative reasoning module provided by an embodiment of the present invention;
[0037] Figure 4 It is a schematic diagram of the internal data structure of the temporary intent library provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0039] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0040] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0041] In view of the difficulties of existing open-ended question implicit intention recognition methods in effectively handling complex semantics, conducting deep reasoning, adapting to dynamic environments, and identifying multi-dimensional intentions, the present invention proposes an open-ended question implicit intention recognition method based on multi-agent reasoning. The method achieves a more comprehensive and in-depth understanding and reasoning of open-ended questions by constructing three key modules: multilingual background knowledge, multi-domain agent collaborative reasoning, and implicit intention recognition module. Specifically, the present method first uses multilingual translation agents and information retrieval agents to collaborate to build multilingual background knowledge for user open-ended questions; then, based on the multilingual background knowledge, multiple domain expert agents are dynamically instantiated, and multiple rounds of "conference" collaborative reasoning are conducted under the guidance of the host management agent to deeply explore the potential intentions of user questions; finally, the host management agent conducts a quality demonstration based on multi-dimensional indicators on the potential intentions generated during the discussion to ensure the accuracy and reliability of the output results. The method proposed in the present invention can effectively solve the shortcomings of existing methods in handling complex semantic understanding, deep reasoning, and dynamic environment adaptability, thereby more accurately and comprehensively identifying the implicit intentions in open-ended questions and significantly improving the ability to adapt to dynamic environments.
[0042] Embodiment 1
[0043] like Figure 1 As shown, this embodiment provides a method for identifying implicit intentions of open-ended questions based on multi-agent reasoning, comprising the following steps:
[0044] Step 1: Use multilingual translation agents and information retrieval agents to collaborate to build multilingual background knowledge for open-ended questions input by users;
[0045] Step 101: construct a multilingual translation agent and an information retrieval agent;
[0046] Step 102: using a multilingual translation agent to translate the input open-ended question into Chinese, English and the target regional language;
[0047] For example, for a question containing "Spanish ham", the multilingual translation agent will translate it into Chinese, English, and Spanish. ” After that, it is first translated into Chinese: “What is the economic development trend of India in the next five years?” Then, based on the content of the question, it is inferred that the question is mainly concerned with “India”. Therefore, in addition to Chinese and English, the multilingual translation agent also translates it into Hindi, one of the main official languages of India: “ ” (Note: This is a Hindi example. In actual applications, multilingual translation intelligence may need to select a more appropriate local language based on the specific situation).
[0048] Based on the multilingual translation agent, open-ended questions input by users can be translated into Chinese, English and the language of the target region, which can enhance the adaptability to different languages.
[0049] Step 103: Using the information retrieval agent to submit the translation result of the open-ended question as a keyword to a query request, integrating the obtained content, and constructing multilingual background knowledge for the open-ended question;
[0050] like Figure 2 As shown, specifically including:
[0051] Step 1031: The information retrieval agent receives the translation results of the multilingual translation agent, i.e., the Chinese, English, and target regional language translations of the user's question. The information retrieval agent uses these translations as keywords to submit queries to general search engines (such as Baidu, Google, etc.).
[0052] Specifically, the information retrieval agent will first obtain the relevant web page links returned by the search engine, then obtain the HTML content of these web pages through network requests (such as HTTP GET requests), and use an HTML parser (such as Python's BeautifulSoup library) to extract the main text data in the web pages.
[0053] For example, the information retrieval agent uses English, Chinese and Hindi keywords to search: English keywords: ”, Chinese keywords: “What is the economic development trend of India in the next five years?”, Hindi keywords: ”.
[0054] The information retrieval agent uses the integrated search engine API to retrieve the HTML source code of these web pages through HTTP GET requests, and then uses Python The library extracts the main text data from the web page, such as title, text, release date, author and other information.
[0055] In this embodiment, in order to ensure the relevance of information, the information retrieval intelligent body will limit the acquisition of web page contents of the first 10 search results for each query.
[0056] Step 1032: The information retrieval agent integrates the retrieved webpage content and related information to construct multilingual background knowledge for the open-ended question;
[0057] In this embodiment, the information retrieval agent integrates the acquired multilingual webpage text data and other relevant information (such as documents, news, etc.) to construct a multilingual background description for the user's question. The background description is centered on the user's question and guided by the multilingual translation results. It covers relevant information collected from different languages, forming a multilingual background knowledge specifically for the user's question, providing support for the subsequent dynamic instantiation of multiple agents.
[0058] For example, an information retrieval agent would integrate text data in English, Chinese, and Hindi extracted from various web pages and combine it with other relevant information, such as: reports on the Indian economy obtained from institutions such as the World Bank and the International Monetary Fund (IMF) (which may include English, Chinese, or other language versions); economic policy documents obtained from official Indian government websites (which may include English and Hindi versions); and articles from mainstream Indian media such as the Times of India. ), Hindustan Times ( ) and other news reports (may include English and Hindi versions). Reports on the Indian economy obtained from Chinese media such as Xinhuanet and People's Daily Online. This information covers a variety of forms such as news reports, academic papers, government reports, forum discussions, etc., and constructs a multilingual background knowledge with user questions as the core, with special attention to information sources from local India, so as to have a deeper understanding of the local economic conditions and development trends.
[0059] Step 2: Based on the constructed multilingual background knowledge, multiple domain expert agents are dynamically instantiated for collaborative reasoning, and the host management agent guides and manages the entire "meeting" process to deeply explore the implicit intentions behind user questions.
[0060] like Figure 3 As shown, the specific steps include:
[0061] Step 201: construct a host management agent and a multi-domain expert agent;
[0062] Among them, the moderator management agent is a pre-defined agent with a fixed role. Its responsibilities include guiding multi-domain expert agents to discuss user issues, raising guiding questions, maintaining discussion order, summarizing discussion results, and evaluating the speeches of each agent.
[0063] The host management agent is configured during the code initialization phase, and is directly called to execute corresponding functions each time an intent recognition task is performed without the need for dynamic instantiation.
[0064] The construction process of multi-domain expert agents includes:
[0065] Based on the constructed multilingual background knowledge, the knowledge content is subject analyzed to identify the main topics or fields covered therein, and then several knowledge domain agents are dynamically instantiated according to the identified topics or fields for collaborative reasoning.
[0066] In this embodiment, the large language model (LLM) and prompt technology are used to dynamically generate specific information of several domain expert agents based on the results of topic analysis and predefined expert information templates. The expert information template contains four aspects: agent name, agent role, agent native language, and agent background description information. For example, a prompt similar to the following can be input to the LLM: "Based on the following background knowledge: [multilingual background knowledge generated in step 1], please generate information of 3 experts in different fields, including expert name, role, native language (for retrieving information) and detailed background description."
[0067] For example, the constructed multilingual background knowledge is analyzed by subject analysis to identify the main areas of interest such as "Indian economy", "macroeconomics", "international trade", "policies and regulations", and "emerging markets". According to the predefined expert information template, prompt words are input to the LLM, for example, "Based on the multilingual background knowledge constructed in step 1, please generate information about experts in three different fields, including the expert's name, role, native language, and detailed background description", and the LLM generates the following expert information:
[0068] Agent 1: Name: Raj; Role: Indian economist; Native language: Hindi; Background description: He has 20 years of experience in Indian economic research, focusing on India's macroeconomic policies, fiscal policies and financial market research, and has published many papers on India's economic growth model.
[0069] Agent 2: Name: Emily; Role: International trade expert; Native language: English; Background description: Has been engaged in international trade research for a long time, specializing in analyzing the global trade pattern, trade policies, and the status and role of emerging markets in global trade.
[0070] Agent 3: Name: Li Ming; Role: Emerging Market Analyst; Native Language: Chinese; Background Description: Focuses on research on economic development in emerging market countries, especially analysis of economic growth potential, investment environment and market risks of BRICS countries.
[0071] LLM will generate expert information that meets the requirements based on the prompt words and background knowledge, thereby instantiating a corresponding number of domain expert agents with relevant knowledge and professional capabilities. The number of agents instantiated here is preset to 3. Each instantiated domain expert agent is given a name, role, native language, and background description information, which comes from the specific content generated by LLM based on background knowledge and expert information templates. Among them, the name is used to identify the agent; the role is used to define the professional field and responsibilities of the agent, for example, the role of "economic expert agent" can be set as "expert proficient in economic market analysis, investment strategy and economic trend forecasting"; the native language is used as the default language for the agent when retrieving external knowledge; the background description information provides a more detailed background introduction of the agent, such as its field of expertise, research direction, etc., making it closer to the identity of a domain expert.
[0072] Step 202: using the host agent to raise guiding questions based on the open-ended questions and the current discussion status, and constructing a temporary intention library;
[0073] The specific steps include:
[0074] Step 221: At the beginning of the "meeting", the host management intelligent agent first proposes an initial guiding question based on the original question input by the user and the guiding question generation rules preset by its prompt technology.
[0075] For example, if the user's question is "How do you view the future development of artificial intelligence?", the host management agent may ask an initial guiding question such as "Please ask experts to analyze the possibilities of the future development of artificial intelligence from the three perspectives of technology, economy and society."
[0076] Step 2022: The agents in each knowledge domain express their own implicit intentions in response to the guiding questions. When speaking, each agent conducts a comprehensive analysis based on the current guiding question, the implicit intentions of the previous agent, reasoning and argumentation, its own role and background description information, and external knowledge retrieved in its native language, so as to simulate as much as possible the speech of a real expert based on his own knowledge and experience. After each agent finishes speaking, its implicit intentions, reasoning and argumentation, and referenced external resource links will be generated.
[0077] It should be noted that in this implementation, each knowledge domain agent expresses their own implicit intentions in response to the guiding questions, and can do so in sequence. Other embodiments also support random speeches, or the host agent makes autonomous judgments and selects the appropriate agent to speak, and the expert agent makes autonomous judgments and selects the next agent expert to speak, without limiting the conference communication method.
[0078] Step 223: The host management agent collects the structured content (implicit intent, reasoning, and referenced external resource links) of all agents, adds them to the temporary intent library, converts the implicit intent into a vector, and continuously updates and enriches the temporary intent library;
[0079] In this embodiment, Figure 4 As shown in the figure, the temporary intention library includes the implicit intentions, reasoning arguments and related references generated during the "meeting". Each node in the temporary intention library represents the speech of an expert agent, including: implicit intentions, reasoning arguments, and referenced external resource links. Among them, the implicit intentions are stored in the form of vectors for subsequent similarity calculation.
[0080] In this embodiment, the Embedding technology is used to vectorize the implicit intent, specifically: each implicit intent text is converted into a vector of a fixed dimension.
[0081] In this embodiment, the process of updating the temporary intent library includes: when a new node needs to be added to the temporary intent library, the host management intelligent body will first calculate the cosine similarity between the "implicit intent" vector in the new node and the "implicit intent" vector in the existing node, and find the N nodes with the highest similarity, where N can be set according to actual needs.
[0082] The host management agent decides whether to treat the new node as a child node of one of the N nodes (indicating expansion or supplement) or as a peer node (indicating a parallel relationship) based on its own judgment. This judgment process uses a large language model and provides it with a clear instruction and format requirement, so that it can judge the relationship between the new node and these nodes based on the vector of the new implicit intention and the information of the most similar N nodes, and output the judgment result.
[0083] For example, a prompt similar to the following can be input to the LLM: "Based on the vector of the new implicit intent and the information of the five most similar nodes, determine which node the new node should be the child of or the sibling node of, and explain the reason. New implicit intent vector: [vector]. Five most similar nodes: [node information]. Output format: Node relationship: [child node / sibling node]. Parent node: [node number]. Reason: [explain the reason]."
[0084] Step 2024: After each round of discussion, the host management agent will generate new guiding questions based on the current temporary intention library and its preset guiding question generation rules, and start the next round of discussion. The discussion process will be iterative until the preset round limit is reached. After each round of discussion, the host management agent evaluates the implicit intentions proposed by each agent.
[0085] The following is a practical example to further illustrate the entire "meeting" process.
[0086] First, the host management agent will propose an initial guiding question based on the user's question and the preset guiding question generation rules: "Experts, please discuss the possible direction of India's economic development in the next five years from the three perspectives of India's domestic economic policies, international trade environment and emerging market development trends."
[0087] Then, the agents in each knowledge domain will express their own implicit intentions in response to the guiding questions in turn. When speaking, each agent will conduct a comprehensive analysis based on the current guiding question, the implicit intention of the previous agent, reasoning, its own role and background description information, and external knowledge retrieved through its native language.
[0088] Raj (Indian economist, native speaker: Hindi):
[0089] Implicit Intent: The user may want to know the impact of the Indian government’s economic policies on economic development over the next five years.
[0090] Reasoning: In recent years, the Indian government has implemented a series of economic reform measures, such as the "Make in India" plan and the Goods and Services Tax (GST) reform, aiming to promote the development of the manufacturing industry, attract foreign investment and improve the business environment. The implementation of these policies will have a significant impact on the Indian economy in the next five years.
[0091] Search in your native language: To support his argument, Raj used his native Hindi to search for information on the official Indian government website and the Indian mainstream media. His search keywords included: ” (Economic Policy of the Government of India), “ ” (Make in India Program), ” (Goods and Services Tax (GST) Reform), etc. Through in-depth analysis of materials in his native language, Raj is able to more accurately grasp the details and latest developments of the Indian government’s economic policies, thereby providing more convincing support for his reasoning.
[0092] External resource links: [reference to official document released by the Indian Ministry of Finance on the “Make in India” program, in Hindi or English], and may also cite links to materials obtained from Hindi media and academic papers (for example: “ ” Analysis of Economic Policies of the Government of India).
[0093] Emily (International Trade Expert, Native Speaker: English):
[0094] Implicit intent: The user may want to understand the impact of the global trade environment on India's economic development.
[0095] Reasoning: As one of the world's important trading countries, India's economic development is closely related to the global trade environment. In the next five years, changes in the global trade pattern, such as the development of regional trade agreements, will have an impact on India's exports and economic growth.
[0096] Searching in her native language: Emily used her native English to search for reports and data from international organizations such as the World Trade Organization (WTO) and the International Monetary Fund (IMF). The keywords she searched for included: (Global Trade Environment)", " (India's Export Performance)"," (Regional Comprehensive Economic Partnership Agreement (RCEP))” etc. Through in-depth analysis of these English materials, Emily was able to have a more comprehensive understanding of the changing trends in the global trade environment and its impact on the Indian economy.
[0097] External resource links: [cites a link to a report on global trade trends released by the World Trade Organization (WTO)], and she may also cite links to materials obtained from English academic papers and industry reports (for example: " ”).
[0098] Li Ming (Emerging Markets Analyst, Native Speaker: Chinese):
[0099] Implicit intent: Users may want to know about India’s economic development prospects as an emerging market country.
[0100] Reasoning: As one of the BRICS countries, India has a huge demographic dividend and market potential. Compared with other emerging market countries, India's economic growth rate is faster, but it also faces challenges such as weak infrastructure and uneven education levels. Whether India's economy can continue to grow rapidly in the next five years depends on whether it can effectively solve these problems.
[0101] Searching in his native language: Li Ming used his native Chinese language to search for published reports on emerging markets and BRICS. His search keywords included: "Indian economic development", "BRICS cooperation", "emerging market risks", etc. Through in-depth analysis of these Chinese materials, Li Ming was able to gain a deeper understanding of India's opportunities and challenges as an emerging market country from a Chinese perspective.
[0102] External resource links: [cites a link to an analytical article on the economic development prospects of the BRICS countries, the link is in Chinese], and he may also cite links to materials obtained from Chinese academic papers and research reports.
[0103] After this round of discussion is completed, the host management agent will conduct a quality assessment on the speeches of the three experts (implicit intent, reasoning, and external resource links). If it exceeds the preset threshold, the implicit intent is identified as a high-quality intent, and the word embedding model is used to convert the implicit intent text into a vector representation.
[0104] Specifically, the moderator management agent utilizes the word embedding model to convert each implicit intent text into a vector of fixed dimension.
[0105] For example: Raj’s implicit intent: “The user may want to understand the impact of the Indian government’s economic policies on economic development in the next five years.” -> converted to vector V through the word embedding model Raj .
[0106] Emily’s implicit intent: “The user may want to understand the impact of the global trade environment on India’s economic development.” -> Converted to vector V through the word embedding model Emily .
[0107] Li Ming’s implicit intention: “The user may want to know about India’s economic development prospects as an emerging market country.” -> Converted to vector V through word embedding model 李明 .
[0108] Next, the moderator management agent decides how to add the new speech node to the temporary intention library by combining vector retrieval and LLM judgment.
[0109] Specifically, when a new speaking node (such as Emily’s speech) needs to be added, the host management agent first calculates the vector of the node’s implicit intention (V Emily ) and the implicit intention vector (V Raj ,V 李明 ) is the cosine similarity of .
[0110] For example:
[0111] Similarity (V Emily , V Raj ) = 0.88,
[0112] Similarity (V Emily , V 李明 ) =0.85,
[0113] Sort by the calculated cosine similarity values and find the Emily The k nodes with the highest similarity (for example, k=5, that is, find the 5 most similar nodes). Assume that after sorting, V Raj The similarity is the highest.
[0114] The host management agent judged through LLM that Emily's speech was a supplement and extension of Raj's speech, so it added Emily's speech node as a child node of Raj to the temporary intention library.
[0115] Specifically, the prompt words input to LLM are similar to: "Based on the vector of the new implicit intention and the information of the following five most similar nodes, determine which node the new node should be the child of or the sibling node of, and explain the reason. New implicit intention vector: [Emily / Li Ming's implicit intention vector]. The five most similar nodes: [Raj, Emily, Li Ming's existing node information]. Output format: Node relationship: [child node / sibling node]. Parent node: [node number]. Reason: [explain the reason]."
[0116] Each speech node contains: speaker (Raj / Emily / Li Ming), implicit intent text, implicit intent vector (V Raj / V Emily / V 李明 ), reasoning, and external resource links. This information is structured and stored in a temporary intent library.
[0117] Finally, the host management agent generates a new guiding question based on all the implicit intent contents of the current temporary intent library: "Please ask the experts to further discuss the economic risks and challenges that India may face in the next five years, as well as the possible countermeasures that the Indian government may take." The three experts continue to discuss around the new guiding question and generate new implicit intents, reasoning arguments, and external resource links. For example, Raj may mention India's high inflation risk and fiscal deficit problems, Emily may mention India's competitiveness in the global industrial chain and potential trade frictions, and Li Ming may mention the gap between the rich and the poor and social instability in India. The host management agent continues to add this information to the temporary intent library and judge the node relationship. The above discussion process will be iterative until the preset round limit (for example, 3 rounds) is reached.
[0118] Step 3: Conduct quality assessment and verification on the implicit intentions mined, and select the implicit intentions that meet the requirements;
[0119] This step evaluates and verifies the quality of the implicit intent generated in the "Multi-domain Agent Collaborative Reasoning" module to screen out high-quality implicit intent and guide the direction of the subsequent "meeting". The host management agent is responsible for implementation, and the core is to use the understanding and reasoning ability of the large language model (LLM) combined with carefully designed prompt words to achieve automatic evaluation of implicit intent.
[0120] The specific steps include:
[0121] After each round of discussion, the host management agent uses a large language model and prompt word engineering to evaluate the implicit intentions proposed by each agent based on the set evaluation dimensions.
[0122] The evaluation dimensions of this embodiment include four core dimensions: intent relevance, intent accuracy, analysis breadth, and analysis depth.
[0123] Among them, the intent relevance evaluates the relevance of the implicit intent to the user's input question and the degree of matching with the core semantics of the user's question. Specifically, it examines whether the implicit intent is closely centered around the user's question, whether it accurately captures the key information of the user's question, and whether it avoids deviating from the topic or introducing irrelevant information.
[0124] Intent accuracy evaluates whether the implicit intent accurately reflects the user's true intent and whether there is any misunderstanding or deviation, specifically examining the coverage of the keywords in the user's question;
[0125] Breadth of analysis: assess whether the implicit intention covers all aspects of the problem and whether different perspectives and possibilities are considered. Specifically, the number of topics involved in the implicit intention, the scope of the field, and the degree of connection with other implicit intentions are examined.
[0126] The depth of analysis evaluates whether the implicit intention has conducted in-depth analysis and reasoning on the problem and whether it has provided valuable insights. It specifically examines the length of the implicit intention's reasoning chain, the logical rigor of the argument, and the degree to which the essence of the problem is revealed.
[0127] For each evaluation dimension, the host management agent will use LLM and carefully designed prompt word engineering to perform the evaluation, providing user questions, implicit intentions, evaluation dimensions and scoring criteria as input to LLM.
[0128] For example, when evaluating "depth of analysis", you can input prompts similar to the following into the LLM: Please play the role of a professional evaluator and evaluate the depth of analysis of the following implicit intentions on the user problem based on the provided scoring criteria, and give a score of 1-5, and explain the reason for the score. User problem: [user problem]. Implicit intention: [implicit intention generated by the agent]. Evaluation dimension: Depth of analysis (evaluate the depth of exploration of the initial topic and its related areas by the implicit intention, and whether it reflects a dynamic discussion). Scoring criteria: 1 point: Very superficial; only a basic overview is provided, and there are obvious gaps in the exploration. 2 points: Superficial; some details are provided, but many important aspects are omitted. 3 points: Medium depth; key aspects are covered, but detailed exploration may be lacking in some areas. 4 points: Good depth; most aspects are explored in detail, with only a few gaps. 5 points: Excellent depth; all relevant aspects are explored comprehensively and in depth, reflecting an in-depth and dynamic discussion. Score: [LLM output score]. Reason: [LLM output score reason].
[0129] LLM will evaluate the implicit intent based on the requirements of the prompt words, combined with its own knowledge and reasoning ability, and output a score and reasoning.
[0130] Implicit intents with scores below the preset threshold will be excluded, while implicit intents with scores above the preset threshold will be retained in the temporary intent library and used to guide subsequent discussions.
[0131] In addition, the host management agent will generate feedback for each implicit intent based on the evaluation reasons output by the LLM. These feedbacks will be used to guide the domain expert agent to improve its speech content in the next round of discussion. For example, if the LLM believes that the analysis depth of a certain implicit intent is insufficient and gives corresponding reasons, the host management agent can generate feedback similar to the following: "The analysis depth of this implicit intent needs to be improved. LLM believes that [reason output by LLM]. It is recommended to further explore the root cause of the problem and provide more sufficient arguments."
[0132] The specific cases are as follows:
[0133] For example, for the implicit intent raised by Raj in the first round, "Users may want to know the impact of the Indian government's economic policies on economic development in the next five years", the host management agent input the following prompt words to the LLM: Please play the role of a professional evaluator, and evaluate the relevance of the following implicit intent to the user's question based on the provided scoring criteria, and give a score of 1-5 points, and explain the reason for the score. User question: What is the economic development trend of India in the next five years? Implicit intent: Users may want to know the impact of the Indian government's economic policies on economic development in the next five years. Evaluation dimension: Intent relevance (evaluate the relevance of the implicit intent to the user's question, and the degree of match with the core semantics of the user's question). Scoring criteria: 1 point: completely irrelevant; 2 points: very weak relevance; 3 points: average relevance; 4 points: relatively strong relevance; 5 points: highly relevant. Score: [LLM output score]. Reason: [LLM output score reason].
[0134] The host management agent generates feedback based on the LLM's evaluation results (e.g., score 5, reason: this implicit intent is directly related to the core of the user's problem, that is, India's economic development in the next five years, and policy is an important influencing factor). Since the implicit intent is highly relevant, the feedback can be a simple affirmation, such as: "This implicit intent is highly relevant to the user's problem and is a very good entry point."
[0135] Specifically, implicit intents with scores below a preset threshold (e.g., 3 points) will be excluded. For example, in a round of discussion, if an agent's implicit intent is too broad or not strongly related to the topic of discussion, and the score is below the threshold, the implicit intent will be excluded from the temporary intent library. Implicit intents with scores above the preset threshold will be retained in the temporary intent library and used to guide subsequent discussions.
[0136] Through the above steps, after multiple rounds of “meetings” and evaluations, multiple high-quality implicit intentions were eventually identified;
[0137] For example:
[0138] Users may want to know the impact of the Indian government’s economic policies on economic development over the next five years;
[0139] Users may want to understand the impact of the global trade environment on India’s economic development;
[0140] Users may want to know about India’s economic development prospects as an emerging market country;
[0141] Users may want to understand the economic risks and challenges that India may face in the next five years;
[0142] Users may want to know the possible response measures taken by the Indian government;
[0143] Users may want to know the development prospects of specific industries in India (such as manufacturing, services, and IT industries);
[0144] Users may want to know about the role of foreign investors in India's economic development.
[0145] These implicit intentions cover multiple aspects such as India's domestic economy, international trade, emerging markets, risks and challenges, government policies, and specific industries. They are analyzed and reasoned in depth, providing users with more comprehensive and accurate information services, and helping users to more deeply understand the complex semantics and multi-dimensional intentions behind open-ended questions. At the same time, through multi-language retrieval and analysis, especially the mining of local Indian language information, the analysis of the Indian economy is more in-depth and closer to the actual situation.
[0146] Embodiment 2
[0147] This embodiment provides an open-ended question implicit intention recognition system based on multi-agent reasoning, including:
[0148] A multilingual background knowledge construction module is used to translate the input open-ended questions using a multilingual translation agent; an information retrieval agent is used to submit a query request using the translation results of the open-ended questions as keywords, and the obtained query contents are integrated to construct multilingual background knowledge for the open-ended questions;
[0149] The multi-domain agent collaborative reasoning module is used to perform topic analysis based on the constructed multi-language background knowledge, generate multi-domain expert agent information according to the topic recognition results, and dynamically construct multi-domain expert agents;
[0150] The implicit intention recognition module is used to generate the implicit intention and reasoning text content of each speaker according to the guiding questions raised by the host agent by using the multi-domain expert agents to speak in turn according to the guiding questions raised by the host agent;
[0151] The host agent is used to evaluate the quality of the implicit intention and reasoning text content. The speech of each screened expert agent is taken as a node to construct an implicit intention library.
[0152] The steps involved in the system of the above embodiment 2 correspond to those of the method embodiment 1. For the specific implementation method, please refer to the relevant description part of the embodiment 1.
[0153] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for identifying implicit intentions in open-ended questions based on multi-agent reasoning, characterized by: The steps include: Using a multilingual translation agent to translate the input open-ended questions, wherein the multilingual translation agent translates the input open-ended questions into Chinese, English, and the target regional language; Using information retrieval agents to submit query requests using the translation results of open-ended questions as keywords, integrating the obtained query content to construct multilingual background knowledge for open-ended questions; Based on the constructed multilingual background knowledge, we conduct topic analysis, generate multi-domain expert agent information according to the topic recognition results, and dynamically construct multi-domain expert agents; The construction process of multi-domain expert agents includes: using multi-language background knowledge as prompt words, using a large language model, and dynamically generating several domain expert agents based on the results of the topic analysis of the multi-language background knowledge and predefined expert information templates, wherein the predefined expert information templates include agent name, agent role, agent native language, and agent background description information; According to the guiding questions raised by the host agent, the multi-domain expert agents are used to generate their own implicit intentions and reasoning text content according to the guiding questions raised by the host agent. The host agent is a pre-defined agent with a fixed role, whose responsibilities include guiding the multi-domain expert agents to discuss around the user's questions, raising guiding questions, maintaining the discussion order, summarizing the discussion results, and evaluating the speeches of each agent; The host agent is used to evaluate the quality of the implicit intention and reasoning text content. The speech of each screened expert agent is taken as a node to construct an implicit intention library.
2. The method for identifying implicit intentions in open-ended questions based on multi-agent reasoning as claimed in claim 1, characterized in that: The acquisition of query content includes: the information retrieval agent first obtains the relevant web page links returned by the search engine, then obtains the HTML content of these web pages through network requests, and uses the HTML parser to extract the main text data in the web pages.
3. The method for identifying implicit intentions in open-ended questions based on multi-agent reasoning as claimed in claim 1, characterized in that: The multi-domain expert agents are used to generate their respective implicit intentions and reasoning text content according to the guiding questions raised by the host agent, including: The host agent asks guiding questions based on the original question and the prompt words of the original question; Expert agents in various fields express their own implicit intentions in response to the guiding questions. When each agent speaks, they conduct a comprehensive analysis of the current guiding question, the previous agent's implicit intentions, reasoning, their own role and background description information, and external knowledge retrieved through their native language. After each agent finishes speaking, it will generate its implicit intentions, reasoning, and links to cited external resources.
4. The method for identifying implicit intentions in open-ended questions based on multi-agent reasoning as claimed in claim 1, characterized in that: The method also includes updating the implicit intent library, specifically: When a new node needs to be added to the temporary intent library, the implicit intent is converted into an implicit intent vector. The host agent calculates the cosine similarity between the implicit intent vector in the new node and the implicit intent vector in the existing node, and finds the N nodes with the highest similarity. According to the preset instructions and format requirements, based on the implicit intent vector in the new node and the information of the N nodes, the relationship between the new node and the N nodes is determined.
5. The method for identifying implicit intentions in open-ended questions based on multi-agent reasoning as claimed in claim 4, characterized in that: Embedding technology is used to vectorize implicit intent and convert each implicit intent text into a vector of fixed dimension.
6. The method for identifying implicit intentions in open-ended questions based on multi-agent reasoning as claimed in claim 1, characterized in that: After each round of speeches, the host management agent generates new guiding questions based on the current temporary intention library and its preset guiding question generation rules, and starts the next round of discussion. The discussion process will be iterated until the preset round limit is reached.
7. The method for identifying implicit intentions in open-ended questions based on multi-agent reasoning as claimed in claim 1, characterized in that: When using the host agent to evaluate the quality of each round of implicit intention and reasoning text content, multiple evaluation dimensions are set. For each evaluation dimension, the user's question, implicit intention, evaluation dimension and scoring criteria are used as input. Combined with the constructed LLM model, the implicit intention is evaluated and the score and reason are output.
8. The method for identifying implicit intentions in open-ended questions based on multi-agent reasoning as claimed in claim 7, characterized in that: The scoring criteria are: 1: Very superficial; only a basic overview is provided with obvious gaps in exploration; 2: Superficial; some details are provided but many important aspects are omitted; 3: Moderate depth; key aspects are covered but detailed exploration may be lacking in some areas; 4: Good depth; most aspects are explored in detail with only a few gaps; 5: Excellent depth; all relevant aspects are explored comprehensively and in depth, reflecting an in-depth and dynamic discussion.
9. An open-ended question implicit intention recognition system based on multi-agent reasoning, characterized by: include: A multilingual background knowledge building module, which is used to translate the input open-ended questions using a multilingual translation agent, wherein the multilingual translation agent translates the input open-ended questions into Chinese, English and the target regional language; Using information retrieval agents to submit query requests using the translation results of open-ended questions as keywords, integrating the obtained query content to construct multilingual background knowledge for open-ended questions; The multi-domain agent collaborative reasoning module is used to perform topic analysis based on the constructed multi-language background knowledge, generate multi-domain expert agent information based on the topic recognition results, and construct a multi-domain expert agent; The construction process of multi-domain expert agents includes: using multi-language background knowledge as prompt words, using a large language model, and dynamically generating several domain expert agents based on the results of the topic analysis of the multi-language background knowledge and predefined expert information templates, wherein the predefined expert information templates include agent name, agent role, agent native language, and agent background description information; The implicit intention recognition module is used to generate the implicit intention and reasoning text content of each subject according to the guiding questions raised by the host agent, using the multi-domain expert agents to speak in turn according to the guiding questions raised by the host agent; the host agent is a pre-defined agent with a fixed role, whose responsibilities include guiding the multi-domain expert agents to discuss around the user's questions, raising guiding questions, maintaining the discussion order, summarizing the discussion results, and evaluating the speeches of each agent; The host agent is used to evaluate the quality of the implicit intention and reasoning text content. The speech of each screened expert agent is taken as a node to construct an implicit intention library.
Citation Information
Patent Citations
Academic conference question-answering system based on large language model
CN118377867A
Large model intention reasoning method and device based on reasoning template, equipment and medium
CN119005153A
Cited By
Multi-agent intention recognition method and system under condition of incomplete multi-source environment information
CN120632566A
A multi-agent intention recognition method and system under the condition of multi-source environmental information incompleteness
CN120632566B