Intelligent customer service method and system based on multi-agent cooperation
Patent Information
- Application Number
- CN202610897662.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]但是,诸如上述专利文献,现有的智能客服系统在处理用户的输入时,缺乏对问题类型的智能分流机制,无法根据问题特征选择最优的处理路径,造成响应效率低下
[0057]本发明通过构建问题路由智能体、RAG检索智能体、回复生成智能体和订单处理智能体等多个专业化智能体,根据问题类型动态调度不同的处理流程,并将多智能体协同的输出作为智能客服的控制信号,来控制信号驱动知识检索模块或订单处理模块动作,使得本发明实施例提供的智能客服方法能够较好地适应客服场景中复杂的用户需求、问题类型、业务逻辑及交互上下文的变化,不仅查询类问题的知识增强更加可靠,快回复类问题的高效处理以及下订单类问题的自动化闭环也得以实现,显著提升了智能客服系统的准确性、效率和业务处理能力。
Smart Images

Figure CN122840231A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent customer service technology. Specifically, this invention relates to an intelligent customer service method and system based on multi-agent collaboration. Background Technology
[0002] Patent document CN121388120A, published on 2026-01-23, discloses an intelligent question-answering method based on a pre-set multi-dimensional knowledge base and a large language model. The method includes: receiving a natural language query and converting it into a query semantic vector; identifying the corresponding intent category and a first confidence level; if the first confidence level is higher than an intent threshold, and the highest similarity between the query semantic vector and the standard question semantic vector in the pre-set multi-dimensional knowledge base is higher than a matching threshold, then returning a standard answer; otherwise, retrieving the generation context based on the query semantic vector, fusing the user query and the generation context, and inputting it into the large language model to generate a preliminary answer; verifying the factual consistency between the preliminary answer and the standard answer in the generation context to generate a credibility score; obtaining the target answer when the credibility score exceeds a credibility threshold; and dynamically optimizing the system based on user feedback data for the target answer.
[0003] However, existing intelligent customer service systems, such as those mentioned in the patent documents, lack an intelligent triage mechanism for question types when processing user input, and cannot select the optimal processing path based on question characteristics, resulting in low response efficiency.
[0004] Furthermore, for complex interactions involving business processes (such as order placement), existing systems struggle to effectively integrate multi-turn dialogue contexts, extract key information, and complete system integration, often requiring manual intervention and exhibiting low levels of automation. Summary of the Invention
[0005] This invention aims to overcome the shortcomings of existing technologies and proposes an intelligent customer service method and system based on multi-agent collaboration to achieve the following objectives: dynamically scheduling different processing flows according to the type of question, realizing knowledge enhancement for query questions, efficient processing for quick reply questions, and automated closed-loop processing for order placement questions, thereby significantly improving the accuracy, efficiency, and business processing capabilities of the intelligent customer service system.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] This invention provides an intelligent customer service method based on multi-agent collaboration, the method comprising:
[0008] Step S1: Construct a multi-agent collaborative architecture, including a problem routing agent, a RAG retrieval agent, a response generation agent, and an order processing agent;
[0009] Step S2: Collect and clean customer service Q&A data to build an external knowledge base;
[0010] Step S3: Construct a training dataset based on the collected data, including a question classification training set, a response generation training set, and an information extraction training set;
[0011] Step S4: Based on the constructed training dataset, fine-tune the training of each agent using a multi-task joint training architecture;
[0012] Step S5: For the input user question, the question routing agent determines the question type and routes the question to the query processing flow, quick reply processing flow, or order placement processing flow.
[0013] Step S6: For query-type questions, the RAG retrieval agent enhances the retrieval based on an external knowledge base, obtains relevant external knowledge, and inputs the retrieval results along with the user's question into the response generation agent to generate the response.
[0014] Step S7: For quick reply questions, directly input the user's question into the reply generation agent to generate the reply;
[0015] Step S8: For order placement issues, the order processing agent performs multiple rounds of information extraction based on the historical dialogue context to obtain the information required for placing an order. After forming structured data, it calls the order placement process API to complete the order placement operation.
[0016] Furthermore, in step S3:
[0017] The question classification training set is constructed in the form of (question, label) text pairs, where the question represents the original question text input by the user, and the label represents the question type label, which includes query type, quick reply type, and order placement type.
[0018] The response generation training set is constructed in the form of (context, response) text pairs, where the context includes a combination of the question context and the external knowledge to be recalled, and the response represents the standard response text for that context.
[0019] The information extraction training set is constructed in the form of text pairs (dialogue_history + user_input, structured_info). Dialogue_history represents the history of multi-turn dialogues, including role identifiers and dialogue content. User_input represents the user input text in the current turn. Structured_info represents the extracted structured information, represented in JSON or key-value pairs, including vehicle model and contact information fields.
[0020] Furthermore, in step S4, the multi-task joint training architecture includes three training tasks: task one is question classification training, task two is response generation training, and task three is information extraction training. The three training tasks share the same encoder to extract semantic features, and each task has its own output head and loss function.
[0021] Furthermore, for Task 1, firstly, the input text is encoded into a continuous vector representation using a pre-trained language model, and the hidden state marked [CLS] in the output of the pre-trained language model is taken as the semantic representation of the entire input sequence, as shown in Equation (1):
[0022] Formula (1): ; In formula (1), d represents the hidden layer dimension of the pre-trained language model PLM; question represents the question text input by the user.
[0023] Then, through a fully linked layer... Mapped to the category space, as shown in formula (2):
[0024] Formula (2): ;
[0025] In formula (2), It is the weight matrix of the classification layer. It is a bias term. It is the number of categories;
[0026] Finally, the cross-entropy loss function is used to calculate the gap between the model prediction and the true label, as shown in Equation (3):
[0027] Formula (3) ;
[0028] In formula (3), The value represents the loss for the problem classification task; N represents the number of samples in a training batch; k represents the total number of problem categories. This represents the one-hot encoding of the true label of the i-th sample; This represents the probability that the model predicts the i-th sample belongs to class j.
[0029] Furthermore, for Task 2, a pre-trained generative language model is used, which generates responses word by word in an autoregressive manner, as shown in Equation (4):
[0030] Formula (4): ;
[0031] In formula (4), represents all trainable parameters of the generative language model; r represents the word sequence of the target response; T represents the response length; c represents the context text, including the user question and external knowledge;
[0032] During the decoding process, conditional probability is used for decoding. When decoding t tokens, it is as shown in formula (5) and formula (6).
[0033] Formula (5): ;
[0034] Formula (6): ;
[0035] in, Let represent the hidden state of the decoder at step t, and Decoder represent the decoder; This represents the projection matrix of the hidden layer output;
[0036] The loss function is to minimize the negative log-likelihood of the generated response, as shown in Equation (7):
[0037] Formula (7): ;
[0038] In formula (7), This represents the loss value for the response generation task; N represents the number of samples in a training batch. Indicates the length of the target response in the i-th sample; This represents the t-th true word in the target response of the i-th sample; This represents the sequence of real words generated before the t-th word in the target response of the i-th sample; This represents the input context of the i-th sample; This means that the model predicts the t-th word as a true word given the context and the previously generated words. The probability of.
[0039] Furthermore, for Task 3, based on the pre-trained language model, the mapping from input to output is learned in a sequence-to-sequence manner. The pre-trained language model generates JSON strings token by token through autoregression, as shown in Formula (8): Formula (8): ;
[0040] In formula (8), It is the token sequence of the target JSON string, where T is the length of the JSON string. Represents all trainable parameters of the model;
[0041] During training, log-likelihood loss is used to maximize the probability of correctly generating JSON sequences. The loss function is shown in Equation (9):
[0042] Formula (9): ;
[0043] In formula (9), This represents the loss value for the information extraction task; N represents the number of samples in a training batch. This represents the length of the target JSON string in the i-th sample; This represents the t-th real token in the target JSON string of the i-th sample; This represents the sequence of real tokens generated before the t-th token in the i-th sample, representing the target JSON string. This represents the input of the i-th sample; This indicates that the model performs well under given input. and historically generated tokens In the case of predicting that the t-th token is a real token The probability of; This represents the trainable parameters of the pre-trained model.
[0044] Furthermore, in Task 3, the pre-trained language model also includes an auxiliary classification head to predict whether each field exists in the current dialogue, as shown in Equation (11):
[0045] Formula (11): :
[0046] In formula (11), This indicates the probability that field f exists in the input; The average pooling vector represents the hidden state of the input sequence; Represents the weights of the linear classifier; This represents the bias term of a linear classifier; This represents the sigmoid activation function;
[0047] The prediction of whether each field exists in the current dialogue is optimized by an auxiliary loss function, which is shown in formula (12);
[0048] Formula (12): ;
[0049] In formula (12), N represents the total number of samples in a training batch; i represents the index of the i-th sample in the batch; F represents a predefined set of target information fields, including two target fields: vehicle model and contact information; f represents a specific field in set F. Indicates the true label; Labels representing model predictions;
[0050] The final loss function for the training of information extraction in Task 3 is shown in Equation (13):
[0051] Formula (13): :
[0052] In formula (13), μ represents the weight coefficient of the auxiliary loss function.
[0053] This invention also provides an intelligent customer service system based on multi-agent collaboration. Using the aforementioned intelligent customer service method based on multi-agent collaboration, the system includes a question routing agent, a RAG retrieval agent, an external knowledge base module, a response generation agent, an order processing agent, and a historical dialogue management module. The question routing agent determines the type of question input by the user and routes it accordingly. The question types include query questions, quick reply questions, and order placement questions. For query questions, the RAG retrieval agent performs a retrieval based on the external knowledge base module and inputs the retrieval results along with the user input into the response generation agent to generate a response. For quick reply questions, the response generation agent directly generates the response. For order placement questions, the order processing agent handles the process to complete the order placement operation. During the processing, the response generation agent engages in question-and-answer sessions with the user to obtain and confirm the information needed for order placement. The historical dialogue management module stores and manages multi-turn dialogue contexts.
[0054] Furthermore, the RAG retrieval agent includes a query understanding submodule, a question expansion submodule, and a vector retrieval submodule. The query understanding submodule receives and confirms the query type question, and then the question expansion submodule expands it to generate a retrieval query vector. Based on the retrieval query vector, the vector retrieval submodule retrieves the K knowledge fragments with the highest vector similarity from the external knowledge base module, concatenates them with the user's question, and inputs them as a response to generate the agent.
[0055] Furthermore, the order processing intelligent agent includes an information extraction submodule, a structured data generation submodule, and an API call submodule. For order-related questions, the information extraction submodule obtains the information needed to place an order from the historical dialogue context and confirms it with the user. If the information is incomplete, it generates follow-up questions to obtain more information from the user. The structured data generation submodule converts the information needed to place an order into structured data. The API call submodule sends the structured data as a request parameter to the order system to perform the order placement operation.
[0056] The technical effects of this invention are as follows:
[0057] This invention constructs multiple specialized intelligent agents, such as a problem routing agent, a RAG retrieval agent, a response generation agent, and an order processing agent. It dynamically schedules different processing flows based on problem type and uses the collaborative output of these multiple agents as control signals to drive the actions of the knowledge retrieval module or the order processing module. This allows the intelligent customer service method provided by this invention to better adapt to the complex user needs, problem types, business logic, and changing interaction contexts in customer service scenarios. Not only is knowledge enhancement for query-type problems more reliable, but efficient processing of quick-response problems and automated closed-loop processing of order-placing problems are also achieved, significantly improving the accuracy, efficiency, and business processing capabilities of the intelligent customer service system. Attached Figure Description
[0058] Figure 1 A flowchart of an intelligent customer service method based on multi-agent collaboration provided in an embodiment of the present invention;
[0059] Figure 2 This is an architecture diagram of an intelligent customer service system based on multi-agent collaboration, provided for an embodiment of the present invention. Detailed Implementation
[0060] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. This is to help those skilled in the art to have a more complete, accurate, and in-depth understanding of the inventive concept and technical solutions of the present invention, and to facilitate its implementation. It should be noted that the terms "first," "second," etc., used in this application are only for the convenience of describing the technical solutions and to distinguish components; the corresponding component configurations may be the same or different, and are not intended to limit the scope of this application. To make the technical solutions of the present invention clearer, the present invention will be explained and illustrated through the following embodiments.
[0061] This invention proposes an intelligent customer service method and system based on multi-agent collaboration. This method constructs multiple specialized agents, such as a problem routing agent, a RAG retrieval agent, a response generation agent, and an order processing agent. It dynamically schedules different processing flows according to the problem type, achieving knowledge enhancement for query-type problems, efficient processing for quick-response problems, and automated closed-loop processing for order-placing problems. This significantly improves the accuracy, efficiency, and business processing capabilities of the intelligent customer service system. Figure 1 As shown, the method proposed in this embodiment of the invention includes the following steps S1 to S8.
[0062] Step S1: Construct a multi-agent collaborative architecture, including a problem routing agent, a RAG retrieval agent, a response generation agent, and an order processing agent;
[0063] Step S2: Collect and clean customer service Q&A data to build an external knowledge base;
[0064] Step S3: Construct a training dataset based on the collected data, including a question classification training set, a response generation training set, and an information extraction training set;
[0065] Step S4: Based on the constructed training dataset, fine-tune the training of each agent using a multi-task joint training architecture;
[0066] Step S5: For the input user question, the question routing agent determines the question type and routes the question to the query processing flow, quick reply processing flow, or order placement processing flow.
[0067] Step S6: For query-type questions, the RAG retrieval agent enhances the retrieval based on an external knowledge base, obtains relevant external knowledge, and inputs the retrieval results along with the user's question into the response generation agent to generate the response.
[0068] Step S7: For quick reply questions, directly input the user's question into the reply generation agent to generate the reply;
[0069] Step S8: For order placement issues, the order processing agent performs multiple rounds of information extraction based on the historical dialogue context to obtain the information required for placing an order. After forming structured data, it calls the order placement process API to complete the order placement operation.
[0070] Referring to step S1, the question routing agent is responsible for receiving user input, analyzing the semantic features of the question, and outputting question type labels. The question types include query, quick reply, and order placement. Query questions indicate that the user's input requires external knowledge to respond, including questions about prices, discounts, and vehicle parameters. Quick reply questions indicate that the user's input does not require external knowledge or is not an order placement question, including casual conversation questions. Order placement questions indicate that the user's input is a desire to place an order, and the user's contact information and intended vehicle model need to be sent as an order lead. The RAG retrieval agent is responsible for receiving query questions and retrieving relevant knowledge fragments from an external knowledge base based on vector retrieval technology. The response generation agent is responsible for generating natural language responses based on the input information, supporting the fusion of retrieved knowledge or direct generation based on internal knowledge. The order processing agent is responsible for managing the multi-turn dialogue process for order placement questions, performing information extraction, intent confirmation, and API calls.
[0071] Referring to step S2, collect seed data for customer service questions and answers, including automotive product manuals, FAQ documents, sales price policies for various car models, and historical customer service dialogue records. Perform deduplication, remove invalid data, and standardize the format of the data. Use a large model to augment the seed data to expand its scale. Simultaneously, construct an external knowledge base using the collected data as a directory index based on car model introductions, sales policies, and FAQ data. The car model introductions and sales policies are stored as unstructured text data. Knowledge documents are segmented into knowledge fragments and vector representations are generated and stored in a vector database. The FAQ data is stored in the form of QA pairs, and the documents are segmented by line to generate vector representations, which are also stored in the vector database.
[0072] Referring to step S3, the constructed training dataset includes a question classification training set, a response generation training set, and an information extraction training set. Among them:
[0073] The question classification training set contains user question text and corresponding question type labels, which are used to train the classification ability of the question routing agent. It is constructed in the form of (question, label) text pairs, where the question represents the original question text input by the user, and the label represents the question type label, which includes query type, quick reply type, and order placement type.
[0074] The response generation training set includes text pairs of question context, reference knowledge, and standard responses, used to train the response generation agent's generation ability. It is constructed in the form of (context, response) text pairs, where the context includes a combination of the question context and the recalled external knowledge, and the response represents the standard response text for that context.
[0075] The information extraction training set includes multi-turn dialogue history, user input, and extracted structured information (vehicle model, contact information), used to train the information extraction capability of the order processing agent. It is constructed in the form of text pairs (dialogue_history + user_input, structured_info), where dialogue_history represents the multi-turn dialogue history, including role identification and dialogue content, user_input represents the user input text in the current round, and structured_info represents the extracted structured information, represented in JSON or key-value pair format, including vehicle model and contact information fields.
[0076] Referring to step S4, based on the training sets obtained in step S3, this embodiment adopts a multi-task joint training architecture. This architecture includes three training tasks: Task 1 is question classification training, Task 2 is response generation training, and Task 3 is information extraction training. The three training tasks share the same encoder to extract semantic features, and each task has its own output head and loss function. During training, the gradients of multiple tasks jointly update the parameters of the shared encoder. During each training iteration, data from the three training sets are selected according to a preset ratio and combined into a batch as input for training.
[0077] Task 1 is a question classification training exercise. The input of the task is the user question text "question", and the output is the question type label y∈{query class, quick reply class, order placement class}. Specifically, firstly, a pre-trained language model (PLM) is used to encode the input text into a continuous vector representation. The hidden state marked [CLS] in the PLM output is taken as the semantic representation of the entire input sequence, as shown in formula (1):
[0078] Formula (1): ;In formula (1), d represents the hidden layer dimension of PLM; question represents the question text entered by the user;
[0079] Then, through a fully linked layer... Mapped to the category space, as shown in formula (2):
[0080] Formula (2): ;
[0081] In formula (2), It is the weight matrix of the classification layer. It is a bias term. It is the number of categories;
[0082] Finally, the cross-entropy loss function is used to calculate the gap between the model prediction and the true label, as shown in Equation (3):
[0083] Formula (3) ;
[0084] In formula (3), The value represents the loss for the problem classification task; N represents the number of samples in a training batch; k represents the total number of problem categories. This represents the one-hot encoding of the true label of the i-th sample; This represents the probability that the model predicts the i-th sample belongs to class j.
[0085] Task 2 is response generation training, using context as input and response as target output, and training is performed by maximizing the likelihood probability of the target sequence. The context is composed of the question context and external knowledge for recall, and a pre-trained generative language model is used. This model generates responses word by word in an autoregressive manner, as shown in formula (4):
[0086] Formula (4): ;
[0087] In formula (4), represents all trainable parameters of the generative language model; r represents the word sequence of the target response; T represents the response length; c represents the context text, including the user question and external knowledge;
[0088] During the decoding process, conditional probability is used for decoding. When decoding t tokens, it is as shown in formula (5) and formula (6).
[0089] Formula (5): ;
[0090] Formula (6): ;
[0091] in, Let represent the hidden state of the decoder at step t, and Decoder represent the decoder; This represents the projection matrix of the hidden layer output;
[0092] The loss function is to minimize the negative log-likelihood of the generated response, as shown in Equation (7):
[0093] Formula (7): ;
[0094] In formula (7), This represents the loss value for the response generation task; N represents the number of samples in a training batch. Indicates the length of the target response in the i-th sample; This represents the t-th true word in the target response of the i-th sample; This represents the sequence of real words generated before the t-th word in the target response of the i-th sample; This represents the input context of the i-th sample; This means that the model predicts the t-th word as a true word given the context and the previously generated words. The probability of.
[0095] Task 3 involves information extraction training, using the current dialogue and dialogue history as input information, and vehicle model and contact information as structured output information. Based on a pre-trained language model, the mapping from input to output is learned in a sequence-to-sequence manner. The pre-trained language model generates JSON strings token by token through an autoregressive approach, as shown in Formula (8): Formula (8): ;
[0096] In formula (8), It is the token sequence of the target JSON string, where T is the length of the JSON string. Represents all trainable parameters of the model;
[0097] During training, log-likelihood loss is used to maximize the probability of correctly generating JSON sequences. The loss function is shown in Equation (9):
[0098] Formula (9): ;
[0099] In formula (9), This represents the loss value for the information extraction task; N represents the number of samples in a training batch. This represents the length of the target JSON string in the i-th sample; This represents the t-th real token in the target JSON string of the i-th sample; This represents the sequence of real tokens generated before the t-th token in the i-th sample, representing the target JSON string. This represents the input of the i-th sample; This indicates that the model performs well under given input. and historically generated tokens In the case of predicting that the t-th token is a real token The probability of; This represents the trainable parameters of the pre-trained model.
[0100] Furthermore, to enhance the model's ability to perceive key information fields, a slot alignment mechanism is introduced in the attention layer. For each predefined field f, the model's attention weight for the corresponding value of that field is calculated, as shown in Equation (10).
[0101] Formula (10): ;
[0102] In formula (10), This represents the importance assessment of each token in the sequence to the value of the generated field f. This represents the query vector for field f, learned by the model. K represents the key vector matrix obtained after encoding the input sequence. This represents the dimension of the key vector.
[0103] Furthermore, in some cases, the input text may not contain information about all predefined fields. Therefore, in Task 3, the pre-trained language model also includes an auxiliary classification head to predict whether each field exists in the current dialogue, as shown in Equation (11):
[0104] Formula (11): :
[0105] In formula (11), This indicates the probability that field f exists in the input; The average pooling vector represents the hidden state of the input sequence; Represents the weights of the linear classifier; This represents the bias term of a linear classifier; This represents the sigmoid activation function;
[0106] The prediction of whether each field exists in the current dialogue is optimized by an auxiliary loss function, which is shown in formula (12);
[0107] Formula (12): ;
[0108] In formula (12), N represents the total number of samples in a training batch; i represents the index of the i-th sample in the batch; F represents a predefined set of target information fields, including two target fields: vehicle model and contact information; f represents a specific field in set F. Indicates the true label; Labels representing model predictions;
[0109] Finally, the final loss function for the training of information extraction in Task 3 is shown in Equation (13):
[0110] Formula (13): :
[0111] In formula (13), μ represents the weight coefficient of the auxiliary loss function.
[0112] In summary, the comprehensive loss function of the multi-task joint training framework of this application is shown in Equation (14).
[0113] Formula (14): ;
[0114] In formula (14), λ1, λ2, and λ3 represent the weight coefficients of the problem classification training task, the response generation training task, and the information extraction training task, respectively.
[0115] In addition, in the multi-task joint training architecture, while each task shares the encoder parameters, each task also has its own parameter weights, as shown in formula (15).
[0116] Formula (15): ;
[0117] In formula (15), This indicates shared encoder parameters, including the Transformer's Embedding layer, multi-head attention layer, feedforward network layer, and layer normalization parameters. The three training tasks share the same set of encoders to extract semantic features. This represents the task-specific parameters for problem classification, including the weight matrix of the classification heads. Bias terms . This indicates the parameters specific to the response generation task, including decoder parameters and the output projection matrix. Bias terms . This refers to parameters specific to the information extraction task, including the structured generation decoder and the field existence prediction layer.
[0118] Correspondingly, the multi-task joint training architecture includes shared encoder parameter updates and task-specific parameter updates during gradient updates. The shared encoder parameter updates need to serve three tasks simultaneously, and the update direction is a weighted average of the gradient directions of the three tasks, so that the features learned by the encoder are beneficial to all tasks, as shown in Equation (16).
[0119] Formula (16): ;
[0120] In formula (16), This represents the learning rate. This represents the task weight coefficient, which controls the proportion of contribution of task t to the update of shared parameters. Let represent the loss function for task t.
[0121] The task-specific parameter update only receives losses from the corresponding task and is not affected by other tasks, as shown in formula (17).
[0122] Formula (17): ;
[0123] In formula (17), This refers to a parameter specific to task t. This represents the gradient of the loss of task t with respect to its specific parameters. This represents the learning rate specific to each task.
[0124] In the specific execution of multi-task joint training, for the problem classification task, the number of training samples is 50,000, with 64 training samples per step, 5 training rounds, a learning rate of 1e−5, dropout of 0.1, a learning rate warm-up step of 500 steps, and a cosine annealing learning rate decay strategy. For the response generation task, the number of training samples is 100,000, with 32 training samples per step, 3 training rounds, a learning rate of 1e−6, dropout of 0.1, and a learning rate warm-up step of 0.1. For the information extraction task, the number of training samples is 30,000, with 16 training samples per step, 10 training rounds, a learning rate of 1e−5, dropout of 0.1, and a learning rate warm-up step of 300 steps.
[0125] Referring to step S5, for the input user question, the question routing agent first determines the question type and then routes the question to a query processing flow, a quick reply processing flow, or an order placement processing flow. The question routing agent receives the user's input question and determines the question type based on semantic understanding. If the question involves product information queries, policy consultations, or price inquiries that require real-time data support, it is classified as a query question; if the question is a greeting, expression of gratitude, or simple inquiry that does not require external knowledge support, it is classified as a quick reply question; if the question expresses a purchase intention, schedules a test drive, or submits an order and involves transaction processes, it is classified as an order placement question.
[0126] When the problem routing agent performs classification, it infers based on the model parameters obtained from the problem classification task training in step S4, as shown in formula (18).
[0127] Formula (18): ;
[0128] In formula (18), This indicates the model's classification decision for the input problem, which will determine the subsequent execution branch. The variable represents the question type label. This represents the original question text entered by the user. This represents the post-training parameters for the problem classification task, including shared encoder parameters and classification task-specific parameters.
[0129] Referring to step S6, for query-type questions, the RAG retrieval agent first expands the user's question to generate a retrieval query vector; it then retrieves the Top-K relevant knowledge fragments from an external knowledge base based on vector similarity; it concatenates the retrieved knowledge fragments with the user's question and inputs them into the response generation agent; the response generation agent generates an accurate and coherent natural language response based on the fused information.
[0130] The RAG retrieval agent expands the user's question and performs Top-K retrieval based on vector similarity. The specific process is as follows:
[0131] First, multiple expanded queries are generated based on the original problem, and each query in the expanded query set is converted into a vector representation, as shown in equations (19) and (20).
[0132] Formula (19): ;
[0133] Formula (20): ;
[0134] In formula (19), This represents the expanded query set generated based on the original question. This represents the original question. Let represent the template for the j-th generation extension problem. In formula (20), Encoder represents the pre-trained text encoder, which is the shared encoder trained in step S4.
[0135] Next, the external knowledge base K is divided into M knowledge segments, each segment is encoded as a vector and indexed, as shown in formulas (21) and (22).
[0136] Formula (21): ;
[0137] Formula (22): ;
[0138] In formula (21), K represents an external knowledge base. M represents the number of knowledge fragments contained in the knowledge base. In formula (22), This represents the vector representation of the i-th knowledge segment. Let i represent the i-th knowledge fragment.
[0139] Finally, the similarity score between the query vector and each knowledge fragment vector is calculated, and the preliminary retrieved results are sorted using the rerank model to finally select the Top-K knowledge fragments. The specific results are shown in formula (23).
[0140] Formula (23): ;
[0141] In formula (23), si represents the similarity score between the query set and the i-th knowledge fragment, and the initial retrieved knowledge fragment set R is obtained.
[0142] The results retrieved initially are reordered based on the rerank model, and the Top-K knowledge fragments are finally selected. As shown in formulas (24) and (25).
[0143] Formula (24): ;
[0144] Formula (25): ;
[0145] In formula (24), This represents the relevance score after reordering. This represents a reordering model that requires careful consideration of both query and document relevance. In formula (25), This indicates the final search results after reordering.
[0146] The response-generating agent performs inference based on the model parameters obtained from the response generation task training in step S4, as shown in formula (26).
[0147] Formula (26): ;
[0148] In formula (26), This represents the final output text generated by the model. This represents the context text, which is composed of the question and external knowledge. This indicates the candidate response text. This represents the post-training parameters for the response generation task in step S4, including shared encoder parameters and response generation task-specific parameters.
[0149] Referring to step S7, for quick reply questions, the user's question is directly combined with the preset system prompt words to input the reply generation agent; the reply generation agent directly generates the reply based on the reply generation model obtained in step S4, without going through the RAG retrieval process, thus shortening the response time.
[0150] Referring to step S8, for order-related questions, the order processing agent activates a multi-turn dialogue management process: First, it analyzes the historical dialogue context to extract the information needed for ordering (intended vehicle model, contact information), and determines the currently acquired information items. If the information is incomplete, it generates a follow-up question to ask the user for the missing information. If the information is complete, it generates an information confirmation statement to confirm the intended vehicle model and contact information with the user. After the user confirms, the intended vehicle model and contact information are extracted as structured information. The extraction of the intended vehicle model and contact information is based on the model parameters obtained from the information extraction task training in step S4, as shown in formula (27).
[0151] Formula (27): ;
[0152] In formula (27), This represents the final extracted structured information, in standard key-value pairs. This represents a structured parsing function that converts the linearized sequence output by the model into standard JSON format. This indicates the history of multiple rounds of dialogue. This represents the user input text for the current round. This represents candidate structured information. This represents the post-training parameters for the information extraction task, including shared encoder parameters and extraction task-specific parameters.
[0153] Next, the order placement API is invoked, sending structured data as request parameters to the order system. Finally, the API response is received, and an order completion notification or error handling reply is generated.
[0154] This invention also provides an intelligent customer service system based on multi-agent collaboration, such as... Figure 2 As shown, the system includes a question routing agent, a RAG retrieval agent, an external knowledge base module, a response generation agent, an order processing agent, and a historical dialogue management module. The question routing agent determines the type of question input by the user and routes it accordingly. The question types include query questions, quick reply questions, and order placement questions. For query questions, the RAG retrieval agent performs a retrieval based on the external knowledge base module and inputs the retrieval results along with the user input into the response generation agent to generate a response. For quick reply questions, the response generation agent directly generates the response. For order placement questions, the order processing agent processes the order to complete the order placement operation. During the processing, the response generation agent engages in question-and-answer sessions with the user to obtain and confirm the information needed to place the order. The historical dialogue management module is used to store and manage multi-turn dialogue contexts.
[0155] The RAG retrieval agent includes a query understanding submodule, a question expansion submodule, and a vector retrieval submodule. The query understanding submodule receives and confirms the query type question, and then the question expansion submodule expands it to generate a retrieval query vector. Based on the retrieval query vector, the vector retrieval submodule retrieves the K knowledge fragments with the highest vector similarity from the external knowledge base module, concatenates them with the user question, and inputs them as a response to generate the agent.
[0156] The order processing intelligent agent includes an information extraction submodule, a structured data generation submodule, and an API call submodule. For order-related questions, the information extraction submodule retrieves the necessary information from the historical dialogue context and confirms it with the user. If the information is incomplete, it generates follow-up questions to obtain more information from the user. The structured data generation submodule converts the necessary information into structured data. The API call submodule sends the structured data as a request parameter to the order system to perform the order placement operation.
[0157] In summary, this invention achieves intelligent question routing through a question routing agent, employing differentiated processing flows for query-type, quick-response, and order-placement-type questions, optimizing system resource allocation, and improving overall processing efficiency. For query-type questions, RAG (Retrieval Augmentation) technology is used, with the RAG retrieval agent obtaining real-time and accurate knowledge information from an external knowledge base, effectively solving the problems of knowledge lag and illusion in intelligent customer service Q&A, and significantly improving the accuracy and timeliness of responses. For order-placement-type questions, the order processing agent combines historical dialogues to perform multi-round information extraction, automatically extracting key business information (intended vehicle model, contact information) and forming structured data, realizing an automated closed loop for complex business processes and reducing reliance on human customer service. The multi-agent collaborative architecture adopted in this invention has clearly defined responsibilities and a high degree of specialization for each agent. Furthermore, through a multi-task joint training architecture, the capabilities of each agent are optimized for specific tasks, significantly improving the execution effect of specific tasks while maintaining the capabilities of the basic model.
[0158] The present invention has been described above by way of example with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited to the above-described manner. Any non-substantial improvements made using the inventive concept and technical solution; or the direct application of the inventive concept and technical solution to other situations without modification, are all within the protection scope of the present invention.
Claims
1. A smart customer service method based on multi-agent collaboration, characterized in that: The method includes: Step S1: Construct a multi-agent collaborative architecture, including a problem routing agent, a RAG retrieval agent, a response generation agent, and an order processing agent; Step S2: Collect and clean customer service Q&A data to build an external knowledge base; Step S3: Construct a training dataset based on the collected data, including a question classification training set, a response generation training set, and an information extraction training set; Step S4: Based on the constructed training dataset, fine-tune the training of each agent using a multi-task joint training architecture; Step S5: For the input user question, the question routing agent determines the question type and routes the question to the query processing flow, quick reply processing flow, or order placement processing flow. Step S6: For query-type questions, the RAG retrieval agent enhances the retrieval based on an external knowledge base, obtains relevant external knowledge, and inputs the retrieval results along with the user's question into the response generation agent to generate the response. Step S7: For quick reply questions, directly input the user's question into the reply generation agent to generate the reply; Step S8: For order placement issues, the order processing agent performs multiple rounds of information extraction based on the historical dialogue context to obtain the information required for placing an order. After forming structured data, it calls the order placement process API to complete the order placement operation.
2. The intelligent customer service method based on multi-agent collaboration according to claim 1, characterized in that: In step S3: The question classification training set is constructed in the form of (question, label) text pairs, where the question represents the original question text input by the user, and the label represents the question type label, which includes query type, quick reply type, and order placement type. The response generation training set is constructed in the form of (context, response) text pairs, where the context includes a combination of the question context and the external knowledge to be recalled, and the response represents the standard response text for that context. The information extraction training set is constructed in the form of text pairs (dialogue_history + user_input, structured_info). Dialogue_history represents the history of multi-turn dialogues, including role identifiers and dialogue content. User_input represents the user input text in the current turn. Structured_info represents the extracted structured information, represented in JSON or key-value pairs, including vehicle model and contact information fields.
3. The intelligent customer service method based on multi-agent collaboration according to claim 1, characterized in that: In step S4, the multi-task joint training architecture includes three training tasks: task one is question classification training, task two is response generation training, and task three is information extraction training. The three training tasks share the same encoder to extract semantic features, and each task has its own output head and loss function.
4. The intelligent customer service method based on multi-agent collaboration according to claim 3, characterized in that: For Task 1, firstly, the input text is encoded into a continuous vector representation using a pre-trained language model. The hidden state marked [CLS] in the output of the pre-trained language model is taken as the semantic representation of the entire input sequence, as shown in Equation (1): Official (1): ; In formula (1), d represents the hidden layer dimension of the pre-trained language model PLM; question represents the question text input by the user. Then, through a fully linked layer... Mapped to the category space, as shown in formula (2): Official (2): ; In formula (2), It is the weight matrix of the classification layer. It is a bias term. It is the number of categories; Finally, the cross-entropy loss function is used to calculate the gap between the model prediction and the true label, as shown in Equation (3): Official (3) ; In formula (3), The value represents the loss for the problem classification task; N represents the number of samples in a training batch; k represents the total number of problem categories. This represents the one-hot encoding of the true label of the i-th sample; This represents the probability that the model predicts the i-th sample belongs to class j.
5. The intelligent customer service method based on multi-agent collaboration according to claim 3, characterized in that: For Task 2, a pre-trained generative language model is used, which generates responses word by word in an autoregressive manner, as shown in Equation (4): Official (4): ; In formula (4), represents all trainable parameters of the generative language model; r represents the word sequence of the target response; T represents the response length; c represents the context text, including the user question and external knowledge; During the decoding process, conditional probability is used for decoding. When decoding t tokens, it is as shown in formula (5) and formula (6); Official (5): ; Official (6): ; in, Let represent the hidden state of the decoder at step t, and Decoder represent the decoder; This represents the projection matrix of the hidden layer output; The loss function is to minimize the negative log-likelihood of the generated response, as shown in Equation (7): Official (7): ; In formula (7), This represents the loss value for the response generation task; N represents the number of samples in a training batch. Indicates the length of the target response in the i-th sample; This represents the t-th true word in the target response of the i-th sample; This represents the sequence of real words generated before the t-th word in the target response of the i-th sample; This represents the input context of the i-th sample; This means that the model predicts the t-th word as a true word given the context and the previously generated words. The probability of.
6. The application method of an intelligent customer service system based on multi-agent collaboration according to claim 5, characterized in that: For Task 3, based on a pre-trained language model, the mapping from input to output is learned in a sequence-to-sequence manner. The pre-trained language model generates JSON strings token by token through an autoregressive approach, as shown in Formula (8): Formula (8): ; In formula (8), It is the token sequence of the target JSON string, where T is the length of the JSON string. Represents all trainable parameters of the model; During training, log-likelihood loss is used to maximize the probability of correctly generating JSON sequences. The loss function is shown in Equation (9): Official (9): ; In formula (9), This represents the loss value for the information extraction task; N represents the number of samples in a training batch. This represents the length of the target JSON string in the i-th sample; This represents the t-th real token in the target JSON string of the i-th sample; This represents the sequence of real tokens generated before the t-th token in the i-th sample, representing the target JSON string. This represents the input of the i-th sample; This indicates that the model performs well under given input. and historically generated tokens In the given case, predict that the t-th token is a real token. The probability of; This represents the trainable parameters of the pre-trained model.
7. The intelligent customer service method based on multi-agent collaboration according to claim 6, characterized in that: In Task 3, the pre-trained language model also has an auxiliary classification head to predict whether each field exists in the current dialogue, as shown in Equation (11): Official (11): : In formula (11), This indicates the probability that field f exists in the input; The average pooling vector represents the hidden state of the input sequence; Represents the weights of the linear classifier; This represents the bias term of a linear classifier; This represents the sigmoid activation function; The prediction of whether each field exists in the current dialogue is optimized by an auxiliary loss function, which is shown in formula (12); Official (12): ; In formula (12), N represents the total number of samples in a training batch; i represents the index of the i-th sample in the batch; F represents a predefined set of target information fields, including two target fields: vehicle model and contact information; f represents a specific field in set F. Indicates the true label; Labels representing model predictions; The final loss function for the training of information extraction in Task 3 is shown in Equation (13): Official (13): : In formula (13), μ represents the weight coefficient of the auxiliary loss function.
8. An intelligent customer service system based on multi-agent collaboration, using an intelligent customer service method based on multi-agent collaboration according to any one of claims 1-7, characterized in that: The system includes a problem routing agent, a RAG retrieval agent, an external knowledge base module, a response generation agent, an order processing agent, and a historical dialogue management module; The problem routing agent determines the type of question input by the user and routes it accordingly. The question types include query, quick reply, and order placement. For query-type questions, the RAG retrieval agent performs a search based on an external knowledge base module, and then inputs the search results along with the user's input to generate a response. For quick reply questions, the reply generation agent generates the reply directly; for order placement questions, the order processing agent processes the order to complete the order placement operation. During the processing, the reply generation agent engages in question-and-answer sessions with the user to obtain and confirm the information needed to place the order; the historical dialogue management module is used to store and manage the context of multi-turn dialogues.
9. The intelligent customer service system based on multi-agent collaboration according to claim 8, characterized in that: The RAG retrieval agent includes a query understanding submodule, a question expansion submodule, and a vector retrieval submodule. The query understanding submodule receives and confirms the query type question, and then the question expansion submodule expands it to generate a retrieval query vector. Based on the retrieval query vector, the vector retrieval submodule retrieves the K knowledge fragments with the highest vector similarity from the external knowledge base module, concatenates them with the user question, and inputs them as a response to generate the agent.
10. The intelligent customer service system based on multi-agent collaboration according to claim 8, characterized in that: The order processing intelligent agent includes an information extraction submodule, a structured data generation submodule, and an API call submodule. For order-related questions, the information extraction submodule retrieves the necessary information from the historical dialogue context and confirms it with the user. If the information is incomplete, it generates follow-up questions to obtain more information from the user. The structured data generation submodule converts the necessary information into structured data. The API call submodule sends the structured data as a request parameter to the order system to perform the order placement operation.
Citation Information
Patent Citations
Intelligent question answering method based on preset multi-dimensional knowledge base and large language model
CN121388120A