A large model intention arrangement dialogue response method and dialogue robot system

CN122838535APending Publication Date: 2026-09-29SICHUAN XW BANK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610889430.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]本发明公开了一种大模型意图编排的对话应答方法及对话机器人系统,拟解决现有技术中没有兼具流程刚性、响应敏捷与智能灵活的对话系统的技术问题

Benefits of technology

[0046](1)本方法通过设置对话图谱、节点,对节点类型进行判断,选择需要回复对话方的话术,并根据话术的类型,灵活调用机器学习语义引擎或大模型RAG语义引擎对话术进行改写或者编写,再根据对话方的回复对对话方的意图进行匹配,进行节点跳转。这种设置方式,一方面控制了对话的流程,使得对话精简、意图明确,降低对话轮数,提高对话效率;另一方面使得对话灵活、智能,带给对话方更高的体验感。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838535A_ABST
    Figure CN122838535A_ABST
Patent Text Reader

Abstract

The application discloses a kind of big model intention arrangement's dialog response method and dialog robot system, it is related to dialog robot technical field, it is proposed to solve the technical problem that existing does not have the dialog system of process rigidity, response agile and intelligent flexible.The present application includes S1: establish the dialog graph atlas containing several node types and different node dialogue type;S2: determine the type of current node in dialog graph atlas;S3: according to node type, call machine learning or big model RAG semantic engine to generate dialogue content;S4: output dialogue content, action type judgment is carried out;S5: when continue dialogue, according to the content of reply, based on big model RAG semantic engine generates customer intent and is matched after node jump, return to S1;S6: loop S1-S5, until the action type judgment in S4 is end dialogue.The present application sets dialog graph atlas, controls the process of dialogue, so that dialogue is simplified;On the other hand, it makes dialogue flexible and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of dialogue robot technology, specifically relating to a dialogue response method and dialogue robot system based on large model intent arrangement. Background Technology

[0002] Customer service in the banking sector demands extremely high levels of security, professionalism, and compliance, far exceeding that of typical retail or consumer industries. This is primarily due to the intangible and complex nature of financial products, which involve significant economic interests and privacy concerns for customers. The specialized and intensive nature of these services necessitates a correspondingly higher investment in customer service costs. Therefore, an increasing number of businesses are adopting voice or text chatbots to balance service effectiveness with cost.

[0003] Semantic engines are a core component of chatbot technology, responsible for converting natural language input into machine-understandable structured semantic information to achieve functions such as intent recognition, slot filling, and context management. In the era of large-scale models, two common types of semantic engines were used to implement chatbot functionality: machine learning semantic engines and large-scale model RAG semantic engines. Machine learning semantic engines rely on a combination of semantic vector retrieval technology and engineering to manage single-node and multi-turn conversations, resulting in fast responses but a mechanical feel and poor performance in multi-turn dialogues. Large-scale model RAG semantic engines, through customized large models or the addition of retrieval enhancement architectures, provide contextual information to independently handle all dialogue context and situational information generated by customers, resulting in excellent dialogue performance and the ability to understand contextual information, but with high training and inference costs. Furthermore, the financial industry is heavily regulated; if customer complaints arise, the internal workings must be clearly explained. However, large-scale model RAG semantic engines cannot explain detailed internal logic, resulting in poor explainability and difficulty in meeting regulatory compliance requirements in the financial sector.

[0004] In summary, existing technologies struggle to simultaneously meet the diverse requirements of high real-time response, robust process compliance control, deep semantic understanding and flexible generation, and high interpretability in high-pressure scenarios such as finance. A single search engine cannot handle complex contexts, while a single generation engine suffers from high latency and uncontrollable processes. Summary of the Invention

[0005] This invention discloses a dialogue response method and a dialogue robot system based on large model intent orchestration, which aims to solve the technical problem that there is no dialogue system in the prior art that combines process rigidity, response agility and intelligent flexibility.

[0006] To solve the aforementioned technical problems, the present invention adopts the following technical solution:

[0007] A dialogue response method based on large-scale model intent orchestration includes:

[0008] S1: Establish a dialogue graph, which contains several nodes. The nodes are divided into three types: start node, normal node and global node, and the corresponding dialogue type is set for each node.

[0009] S2: Determine the type of the current node in the dialogue graph based on the current dialogue;

[0010] S3: Based on the node type in S2, call the machine learning semantic engine or the large model RAG semantic engine to generate dialogue content through the corresponding speech type of the node type;

[0011] S4: Output the dialogue content generated by S3 to the other party and determine the action type, which is divided into ending the dialogue and continuing the dialogue.

[0012] S5: When the action type is to continue the dialogue, the dialogue intent is generated based on the content of the dialogue party's reply using the large model RAG semantic engine, and the dialogue intent is matched. If the intent match fails, it is processed as unrecognized. If the intent match succeeds, the node jump is performed, returning to S1.

[0013] S6: Loop through S1-S5 until the action type in S4 is determined to end the dialogue.

[0014] By adopting this technical solution, a dialogue graph and nodes are set up. The node type is determined, and the appropriate response is selected. Based on the response type, a machine learning semantic engine or a large-scale RAG semantic engine is flexibly invoked to rewrite or create the dialogue. Then, the respondent's intent is matched against the response, and node navigation is performed accordingly. This approach controls the dialogue flow, making it concise, clear in intent, reducing the number of dialogue rounds, and improving efficiency. Furthermore, it makes the dialogue flexible and intelligent, providing a superior user experience.

[0015] Preferably, in S1, when establishing the dialogue graph, the dialogue type of each node is preset based on human dialogue experience. The dialogue types include fixed dialogue, rewritten dialogue and free dialogue.

[0016] When the script type is a fixed script, the machine learning semantic engine is called to replace the variables and then output directly;

[0017] When the script type is script rewriting, the large model RAG semantic engine is called to replace variables and rewrite the output.

[0018] When the dialogue type is free dialogue, the large model RAG semantic engine is called to output directly.

[0019] After adopting this technical solution, the dialogue content of the responding party will be divided into three categories: fixed script, free script, and rewritten script. Depending on the type of script, the machine learning speech engine or the large model RAG semantic engine will be used to rewrite the script. This ensures the efficiency of script modification on the one hand, and improves the flexibility and intelligence of script expression on the other hand.

[0020] Preferably, S1 also includes configuring the dialogue parameters of each node. In S2, it is determined whether the information obtained at present meets the requirements of the dialogue parameters based on the dialogue parameters corresponding to the node. If not, S2 also includes parameter processing. The parameter processing is used to obtain relevant information of the dialogue party. The obtained information is returned by the large model RAG semantic engine in the form of key-value pairs and is used for speech variable replacement in S3.

[0021] By adopting this technical solution, parameters are configured for each node. Based on the information obtained from the current dialogue, it can determine whether the relevant parameters meet the requirements for generating a new dialogue. If not, parameter processing is performed to obtain the necessary parameters, and dialogue content is generated through methods such as variable substitution. This setup ensures the accuracy of each round of dialogue content, thereby improving the accuracy of the information received by the parties involved.

[0022] Preferably, parameter processing includes extracting key information based on historical dialogues and calling the large model RAG semantic engine to return the required parameter information.

[0023] After adopting this technical solution, parameter processing can also extract and return key information from historical dialogues, making it easier to retrieve or call it later.

[0024] Preferably, S1 further includes configuring operation functions for nodes as needed, and in S2, for nodes with configured operation functions, it is determined whether to call the operation functions based on the dialogue situation. The operation functions are used to realize interaction with the dialogue party or the outside world, and the output parameters obtained through the interaction content are used to generate dialogue content in S3.

[0025] After adopting this technical solution, operation functions are configured for the required nodes. When the system needs to interact with the other party or the outside world, the interaction can be carried out through the operation functions, such as sending verification codes, text messages, or making phone calls to the other party. This improves the functionality of this method and enhances the user experience for the other party.

[0026] Preferably, the intent matching result in S5 includes:

[0027] When the intent matches the associated node, jump to the matching node;

[0028] When the intent fails to match the associated node and an unidentified node exists, redirect to the unidentified node;

[0029] When the intent fails to match the associated node and there is no unidentified node, a fallback message is broadcast.

[0030] By adopting this technical solution, matching is performed based on the responses of the parties involved in the dialogue, and matching scenarios are set. This allows for accurate identification of the parties' intentions, reduces misunderstandings, improves the response efficiency of the method, and enables automated processing.

[0031] A large-scale model-based dialogue robot system that aims to orchestrate dialogue response methods, including

[0032] Dialogue Graph Building Module: Used to build a dialogue graph including data nodes, set the type of each node, and synchronously set the corresponding dialogue type for each node;

[0033] Node processing module: Used to determine the node type of the current dialogue in the dialogue graph;

[0034] The dialogue judgment module is used to determine the node type and corresponding dialogue type after the node processing module has judged the node, call the machine learning semantic engine or the large model RAG semantic engine to generate dialogue content, and broadcast it to the dialogue parties.

[0035] Intent matching module: Used to match the intent of the dialogue parties. After receiving the dialogue content, the dialogue parties take actions. Based on the dialogue parties' actions, the intent of the dialogue parties is matched. If the intent matching fails, it is treated as unrecognized. If the intent matching is successful, the node jump is performed.

[0036] After adopting this technical solution, various sub-modules are set up, each with specific functions. Performance optimization and capacity expansion can be performed individually for modules with heavy loads, resulting in more rational resource utilization. Examples include the dialogue judgment module and the intent matching module. Simultaneously, the modular system structure offers strong adaptability and high reusability. When adding new services or requiring modifications, new modules can be directly inserted or existing modules can be modified. Furthermore, the system can be reused in multiple scenarios, reducing redundant development.

[0037] Preferably, the dialogue graph establishment module further includes configuring dialogue parameters for each node. The node processing module determines whether the information obtained meets the dialogue parameter requirements based on the corresponding dialogue parameters. If not, the node processing module further includes calling parameters. The parameter processing is used to obtain relevant information of the dialogue parties. The obtained information is used to generate new dialogue content in the dialogue judgment module.

[0038] After adopting this technical solution, parameter processing can obtain basic information such as the name and gender of the dialogue parties, as well as relevant business information such as loans, which helps to improve the accuracy of the information presented to the other party in each round of dialogue.

[0039] Preferably, the dialogue graph establishment module further includes configuring operation functions for nodes as needed. When the node processing module processes a node with configured operation functions, it determines whether the operation function needs to be called based on the current dialogue situation. The operation function is used to realize interaction with the dialogue party or the outside world. The output parameters obtained through the interaction content are used to generate dialogue content in the dialogue judgment module.

[0040] After adopting this technical solution, the facility operation function makes function calls. When the system needs to interact with the other party or the outside world during the current dialogue, it can send text messages, make phone calls, etc. through function calls to prompt the other party to complete the information and accept the service, thus ensuring the functionality of the system and further improving its flexibility.

[0041] Preferably, an action judgment module is provided between the dialogue judgment module and the intent matching module. After the dialogue judgment module broadcasts the generated dialogue, the action judgment module judges the dialogue party's action based on the dialogue party's action. The dialogue party's action includes replying and not replying. When the dialogue party replies, the process jumps to the intent matching module; when the dialogue party does not reply, the process ends.

[0042] After adopting this technical solution, an action type judgment module is set up before matching the intent of the dialogue parties, which further standardizes the system process and helps to improve dialogue efficiency.

[0043] Preferably, each node dynamically invokes a machine learning semantic engine or a large model RAG semantic engine based on its type and configuration when establishing the dialogue graph. The machine learning semantic engine is used for parameter processing, function calls, and fast retrieval, while the large model RAG semantic engine is used for deep information extraction, speech rewriting, and intent deep matching.

[0044] By adopting this technical solution, the system abstracts the dialogue into a dialogue graph composed of nodes. Each node dynamically decides to call traditional machine learning tools (for parameter processing, function calls, and fast retrieval) or large model modules (for deep information extraction, speech rewriting, and deep intent matching) based on its type and configuration. The system then drives the transition between nodes through intent recognition results based on the large model, thereby achieving precise, flexible, and intelligent control over the dialogue process.

[0045] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0046] (1) This method sets up a dialogue graph and nodes, judges the node type, selects the dialogue script that needs to be replied to, and flexibly calls the machine learning semantic engine or the large model RAG semantic engine to rewrite or write the dialogue script according to the type of dialogue script. Then, it matches the dialogue party's intention according to the dialogue party's reply and performs node jump. This setting method controls the dialogue flow on the one hand, making the dialogue concise and the intention clear, reducing the number of dialogue rounds and improving the dialogue efficiency; on the other hand, it makes the dialogue flexible and intelligent, bringing a better experience to the dialogue party.

[0047] (2) The powerful contextual understanding capability of the large model RAG semantic engine is used for intent recognition. As the core decision-making mechanism for process jump in complex dialogue environment, it replaces the traditional rule matching or classification model (single sentence understanding) and significantly improves the accuracy of jump and scene adaptability.

[0048] (3) The system abstracts the dialogue into a dialogue graph composed of nodes. Each node dynamically decides to call traditional machine learning (for parameter processing, function calls, and fast retrieval) or a large model module (for deep information extraction, speech rewriting, and intent deep matching) based on its type and configuration, and drives the jump between nodes through the intent recognition results based on the large model, thereby achieving precise, flexible and intelligent control of the dialogue process. Attached Figure Description

[0049] The present invention will be described by way of example and with reference to the accompanying drawings, wherein:

[0050] Figure 1 A diagram illustrating the establishment of the dialogue graph, including start nodes, regular nodes, and global nodes;

[0051] Figure 2 This is a schematic diagram of the processing flow inside the node;

[0052] Figure 3 A schematic diagram of node parameter processing;

[0053] Figure 4 Set up a graph for scene variables;

[0054] Figure 5 Example diagram of assigning values ​​to variables when calling a function;

[0055] Figure 6 A diagram illustrating the rewritten script. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the embodiments and accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The modules of the embodiments of this application described and marked in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0057] In the description of the embodiments of this application, it should be noted that the terms "first", "second", "third", etc. are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0058] The following is combined Figures 1-6 The present invention will be described in detail below.

[0059] Example 1

[0060] A large-scale model intent orchestration method for dialogue responses, such as Figures 1-6 As shown, it includes:

[0061] S1: Establish a dialogue graph, which contains several nodes. The nodes are divided into three types: start node, normal node and global node, and the corresponding dialogue type is set for each node.

[0062] S2: Determine the type of the current node in the dialogue graph based on the current dialogue;

[0063] S3: Based on the node type in S2, call the machine learning semantic engine or the large model RAG semantic engine to generate dialogue content through the corresponding speech type of the node type;

[0064] S4: Output the dialogue content generated by S3 to the other party and determine the action type, which is divided into ending the dialogue and continuing the dialogue.

[0065] S5: When the action type is to continue the dialogue, the dialogue intent is generated based on the content of the dialogue party's reply using the large model RAG semantic engine, and the dialogue intent is matched. If the intent match fails, it is processed as unrecognized. If the intent match succeeds, the node jump is performed, returning to S1.

[0066] S6: Loop through S1-S5 until the action type in S4 is determined to end the dialogue.

[0067] In this embodiment, the dialogue flow is decomposed into nodes. For standardized, high-frequency processes (such as information confirmation and option selection), lightweight logical nodes (based on rules and retrieval) are used to achieve millisecond-level response with almost no computational cost. Furthermore, the large model is invoked only on demand: the large model in the generation node is only invoked when deep understanding, rewriting, or complex intent judgment is required (such as handling interlocutor complaints or generating personalized marketing scripts). This significantly reduces the frequency of large model invocations and context length, substantially lowering overall service latency and cloud computing costs, making it possible to deploy large models in real-time financial voice scenarios.

[0068] When the dialogue length and contextual knowledge information exceed 8000 tokens, the output response in the traditional large-scale model RAG format (assuming 60~100 tokens) takes 800ms~1400ms.

[0069] Because this method decomposes the modules, the intent understanding and session generation context are strictly controlled when the dialogue history grows, and the number of tokens will basically not exceed 3000. The time taken to output 60~100 tokens is 400ms~800ms.

[0070] In this embodiment, the dialogue graph described in S1 is designed based on the communication process between frontline business personnel and the dialogue parties. Business logic is visualized and described by graph-like nodes and edges. Product managers or business experts can adjust the process, modify the script, and add branches through configuration (rather than writing code), which greatly improves the speed of business iteration.

[0071] In this embodiment, the start node in S1 is initiated by me to the other party in the dialogue, including opening remarks such as greetings. The start node only exists at the beginning of the current round of dialogue.

[0072] In this embodiment, the ordinary node, global node, and start node in S1 are, for example, Figure 1 As shown, ordinary nodes need to be connected through the dialogue graph to establish jump logic, while global nodes can jump from any node without the need for dialogue graph connections. For example, if we want to call a party to notify them to check their text message, the first step is the opening node, and the second step is that the party may deny being the person in the conversation. Both steps can be established by connecting ordinary nodes. However, the party may file a complaint anywhere, and the complaint is a global node.

[0073] In this embodiment, when establishing the dialogue graph in S1, the dialogue type of each node is preset based on human dialogue experience. The dialogue types include fixed dialogue, rewritten dialogue and free dialogue.

[0074] When the script type is a fixed script, the machine learning semantic engine is called to replace the variables and then output directly;

[0075] When the script type is script rewriting, the large model RAG semantic engine is called to replace variables and rewrite the output.

[0076] When the dialogue type is free dialogue, the large model RAG semantic engine is called to output directly.

[0077] In this embodiment, the fixed script is designed by frontline sales personnel based on their communication with the other party, and the wording within the script remains unchanged. The rewritten script, however, involves frontline sales personnel pre-setting a basic short sentence, which is then rewritten by the large-scale model RAG semantic engine based on contextual information and requirements. Figure 6 As shown, free-form language does not require the use of basic short sentences and is generated directly from the context.

[0078] In this embodiment, S1 also includes configuring the dialogue parameters of each node. In S2, it is determined whether the information obtained at present meets the requirements of the dialogue parameters based on the dialogue parameters corresponding to the node. If not, S2 also includes parameter processing. The parameter processing is used to obtain relevant information of the dialogue party. The obtained information is returned by the large model RAG semantic engine in the form of key-value pairs and is used for speech variable replacement in S3.

[0079] In this embodiment, the parameter information processed in S1 includes dialogue party information, financial product information, and dialogue party business-related status obtained from the associated system.

[0080] In this embodiment, as Figure 3 As shown, the parameter processing is performed according to logical operations. That is, when the parameter belongs to a certain category, the corresponding value is output. For example, if the gender is female, the logical operation value is 'true' or 'false'; if it is true, the title is lady.

[0081] In this embodiment, the parameter processing in S1 also includes extracting key information based on historical dialogues and calling the large model RAG semantic engine to return the required parameter information.

[0082] In this embodiment, if the current node needs to extract key information based on historical dialogues in S1, an operation is triggered, such as extracting all entities such as the time and location of the user's description, and returning the required information in the form of key-value pairs through a specified large model. In this embodiment, the large model refers to Deepseek, qwen, chatgp, etc. These models are all industry-standard technologies and can be selected according to one's own situation.

[0083] In this embodiment, the key-value pairs are in JSON format.

[0084] In this embodiment, variables are assigned values ​​based on the returned information. There are two cases for information extraction, as shown in Table 1 below. When extraction is successful, new dialogue content is generated using real information. When extraction fails, a pre-set default value is used to ensure that the method will not be interrupted due to lack of data or recognition abnormalities, and the process can continue to run normally.

[0085] Table 1 Context Information Extraction Results

[0086]

[0087] In this embodiment, S1 also includes configuring operation functions for nodes as needed. In S2, for nodes with configured operation functions, it is determined whether to call the operation functions based on the dialogue situation. The operation functions are used to realize interaction with the dialogue party or the outside world. The output parameters obtained through the interaction content are used to generate dialogue content in S3.

[0088] In this embodiment, as Figures 4-5 As shown, the function call described in S1 involves setting up a corresponding function interface, extracting the corresponding output parameters based on the current dialogue information, assigning the output parameters to the corresponding interface, and calling the function interface to issue a response command, thereby completing the interaction with the dialogue party.

[0089] In this embodiment, the intent matching result in S5 includes:

[0090] When the intent matches the associated node, jump to the matching node;

[0091] When the intent fails to match the associated node and an unidentified node exists, redirect to the unidentified node;

[0092] When the intent fails to match the associated node and there is no unidentified node, a fallback message is broadcast.

[0093] In this embodiment, in S5, for intent matching, if the intent text returned by the large model RAG semantic engine model based on the analysis of historical chat records completely matches the configured intent, the process will jump directly. Alternatively, if the intent text returned based on the analysis of historical chat records does not completely match, the built-in model in the background will vectorize the text and perform vectorized similarity calculation (similarity is a floating-point number from 0 to 1). If it is greater than the manually configured threshold, such as 0.3, the most similar intent will be taken as the matching item; otherwise, it will be: intent not recognized.

[0094] In this embodiment, the long-context understanding capability of the large-scale RAG semantic engine is utilized, abandoning the single-sentence vector understanding mode and multi-turn conversation management mode of traditional machine learning semantic engines, enabling a deeper understanding of the latent semantics of the speakers. The accuracy of single-turn sentence understanding is approximately 85%~93% (the loss is mainly concentrated in the speaker's single-sentence responses such as 'um,' 'ah,' 'yes,' etc., which have unclear semantics). The accuracy of multi-turn sentences depends on the conversation management mode. The long-context understanding accuracy of the large-scale RAG semantic model in this method is consistently above 95%. As shown in Table 2:

[0095] Table 2 provides a comprehensive comparison of machine learning semantic engines, large-model RAG semantic engines, and this invention.

[0096]

[0097] Example 2

[0098] This embodiment is basically the same as Embodiment 1, except that: in this embodiment, as shown in the example... Figures 1-6 As shown, a large-scale model-based dialogue robot system that programs dialogue response methods includes:

[0099] Dialogue Graph Building Module: Used to build a dialogue graph including data nodes, set the type of each node, and synchronously set the corresponding dialogue type for each node;

[0100] Node processing module: Used to determine the node type of the current dialogue in the dialogue graph, and corresponding... Figure 1 Module 1;

[0101] The dialogue judgment module is used to determine the node type and corresponding dialogue type after the node processing module has judged the node. It then calls the machine learning semantic engine or the large-scale RAG semantic engine to generate the dialogue content and broadcasts it to the dialogue parties. Figure 1 Module 2;

[0102] The intent matching module is used to match the intents of the parties in a conversation. After receiving the conversation content, the parties take actions, and the module matches their intents based on these actions. If an intent match fails, it is treated as unrecognized; if an intent match is successful, the module redirects to the appropriate node. Figure 1 Module 3.

[0103] In this embodiment, the dialogue graph establishment module further includes configuring dialogue parameters for each node. The node processing module determines whether the information obtained meets the dialogue parameter requirements based on the corresponding dialogue parameters. If not, the node processing module further includes calling parameters. The parameter processing is used to obtain relevant information of the dialogue parties. The obtained information is used to generate new dialogue content in the dialogue judgment module.

[0104] In this embodiment, the dialogue graph establishment module further includes configuring operation functions for nodes as needed. When the node processing module processes a node with configured operation functions, it determines whether the operation function needs to be called based on the current dialogue situation. The operation function is used to realize interaction with the dialogue party or the outside world. The output parameters obtained through the interaction content are used to generate dialogue content in the dialogue judgment module.

[0105] In this embodiment, an action judgment module is set between the dialogue judgment module and the intent matching module. After the dialogue judgment module broadcasts the generated dialogue, the action judgment module judges the dialogue party's action based on the dialogue party's action. The dialogue party's action includes replying and not replying. When the dialogue party replies, the process jumps to the intent matching module; when the dialogue party does not reply, the process ends.

[0106] In this embodiment, each node dynamically invokes a machine learning semantic engine or a large model RAG semantic engine according to its type and configuration when establishing a dialogue graph. The machine learning semantic engine is used for parameter processing, function calls, and fast retrieval, while the large model RAG semantic engine is used for deep information extraction, speech rewriting, and intent deep matching.

[0107] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A dialogue response method based on large-scale model intent orchestration, characterized in that: include: S1: Establish a dialogue graph, which contains several nodes. The nodes are divided into three types: start node, normal node and global node, and the corresponding dialogue type is set for each node. S2: Determine the type of the current node in the dialogue graph based on the current dialogue; S3: Based on the node type in S2, call the machine learning semantic engine or the large model RAG semantic engine to generate dialogue content through the corresponding speech type of the node type; S4: Output the dialogue content generated by S3 to the other party and determine the action type, which is divided into ending the dialogue and continuing the dialogue. S5: When the action type is to continue the dialogue, the dialogue intent is generated based on the content of the dialogue party's reply using the large model RAG semantic engine, and the dialogue intent is matched. If the intent match fails, it is processed as unrecognized. If the intent match succeeds, the node jump is performed, returning to S1. S6: Loop through S1-S5 until the action type in S4 is determined to end the dialogue.

2. The dialogue response method for large-scale model intent orchestration according to claim 1, characterized in that: In S1, when building a dialogue graph, the dialogue type of each node is preset based on human dialogue experience. The dialogue types include fixed dialogue, rewritten dialogue and free dialogue. When the script type is a fixed script, the machine learning semantic engine is called to replace the variables and then output directly; When the script type is script rewriting, the large model RAG semantic engine is called to replace variables and rewrite the output. When the dialogue type is free dialogue, the large model RAG semantic engine is called to output directly.

3. The dialogue response method for large-scale model intent orchestration according to claim 2, characterized in that: S1 also includes configuring the dialogue parameters of each node. In S2, it is determined whether the information obtained at present meets the requirements of the dialogue parameters based on the dialogue parameters corresponding to the node. If not, S2 also includes parameter processing. The parameter processing is used to obtain relevant information of the dialogue party. The obtained information is returned in key-value pairs through the large model RAG semantic engine and used for speech variable replacement in S3.

4. The dialogue response method for large-scale model intent orchestration according to claim 3, characterized in that: Parameter processing includes extracting key information from historical dialogues and calling the large model RAG semantic engine to return the required parameter information.

5. The dialogue response method for large-scale model intent orchestration according to claim 3, characterized in that: S1 also includes configuring operation functions for nodes as needed. In S2, for nodes with configured operation functions, it is decided whether to call the operation functions based on the dialogue situation. The operation functions are used to realize the interaction with the dialogue party or the outside world. The output parameters obtained through the interaction content are used to generate dialogue content in S3.

6. The dialogue response method for large-scale model intent orchestration according to claim 1, characterized in that: The intent matching results in S5 include: When the intent matches the associated node, jump to the matching node; When the intent fails to match the associated node and an unidentified node exists, redirect to the unidentified node; When the intent fails to match the associated node and there is no unidentified node, a fallback message is broadcast.

7. A dialogue robot system that uses any one of the large model intentions from 1 to 5 to program dialogue response methods, characterized in that: include Dialogue Graph Building Module: Used to build a dialogue graph including data nodes, set the type of each node, and synchronously set the corresponding dialogue type for each node; Node processing module: Used to determine the node type of the current dialogue in the dialogue graph; The dialogue judgment module is used to determine the node type and corresponding dialogue type after the node processing module has judged the node, call the machine learning semantic engine or the large model RAG semantic engine to generate dialogue content, and broadcast it to the dialogue parties. Intent matching module: Used to match the intent of the dialogue parties. After receiving the dialogue content, the dialogue parties take actions. Based on the dialogue parties' actions, the intent of the dialogue parties is matched. If the intent matching fails, it is treated as unrecognized. If the intent matching is successful, the node jump is performed.

8. The dialogue robot system using the large model intent orchestration dialogue response method according to claim 6, characterized in that: The dialogue graph establishment module also includes configuring dialogue parameters for each node. The node processing module determines whether the information obtained meets the dialogue parameter requirements based on the corresponding dialogue parameters. If not, the node processing module also includes calling parameters. The parameter processing is used to obtain relevant information of the dialogue party. The obtained information is used to generate new dialogue content in the dialogue judgment module. The dialogue graph establishment module also includes configuring operation functions for nodes as needed. When a node is configured with an operation function, the node processing module determines whether the operation function needs to be called based on the current dialogue situation. The operation function is used to realize interaction with the dialogue party or the outside world. The output parameters obtained through the interaction content are used to generate dialogue content in the dialogue judgment module.

9. The dialogue robot system using the large model intent orchestration dialogue response method according to claim 6, characterized in that: An action judgment module is set between the dialogue judgment module and the intent matching module. After the dialogue judgment module broadcasts the generated dialogue, the action judgment module judges the dialogue party's action based on the dialogue party's actions. The dialogue party's action includes replying and not replying. When the dialogue party replies, the module jumps to the intent matching module. When the dialogue party does not reply, the process ends.

10. The dialogue robot system using the large model intent orchestration dialogue response method according to claim 6, characterized in that: When establishing a dialogue graph, each node dynamically invokes either a machine learning semantic engine or a large-scale RAG semantic engine based on its type and configuration. The machine learning semantic engine is used for parameter processing, function calls, and fast retrieval, while the large-scale RAG semantic engine is used for deep information extraction, speech rewriting, and intent deep matching.