Active customer service dialogue, model training method and apparatus therefor, device, medium and product

WO2026195020A1PCT designated stage Publication Date: 2026-09-24BAIRONG ZHIXIN (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/084725
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-21
Filing Date
2026-03-20
Publication Date
2026-09-24

Smart Images

  • Figure CN2026084725_24092026_PF_FP_ABST
    Figure CN2026084725_24092026_PF_FP_ABST
Patent Text Reader

Abstract

An active customer service dialogue, a model training method and apparatus therefor, a device, a medium and a product. The active customer service dialogue model training method comprises: acquiring training data (201), wherein the training data comprises general dialogue data and real dialogue data, and the general dialogue data and the real dialogue data each comprise customer service dialogue data and customer dialogue data; performing feature extraction on the training data to obtain prompt information (202); combining the training data and the prompt information into mixed data (203); and using the mixed data to train a large language model so as to obtain an active customer service dialogue model (204).
Need to check novelty before this filing date? Find Prior Art

Description

Customer service proactive dialogue and its model training methods, devices, equipment, media and products

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese patent application No. CN202510337519.7, filed on March 21, 2025, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The embodiments disclosed herein relate to the field of artificial intelligence technology, specifically to proactive customer service dialogue and its model training methods, apparatus, devices, media, and products. Background Technology

[0004] Currently, customer service robots on the market can only passively respond to customer inquiries and cannot actively communicate with customers like human customer service representatives, resulting in reduced customer engagement.

[0005] In the realm of proactive dialogue, most systems are open-topic chatbots based on large models. While these systems can proactively communicate and excel in understanding and generating natural language, they lack specific dialogue topics and the ability to proactively communicate with customers in a targeted manner. Summary of the Invention

[0006] Embodiments of this disclosure provide methods, apparatus, devices, media, and products for proactive customer service dialogue and its model training.

[0007] In a first aspect, embodiments of this disclosure provide a method for training a customer service proactive dialogue model, the method comprising:

[0008] Acquire training data, which includes general dialogue data and real dialogue data, both of which include customer service dialogue data and customer dialogue data.

[0009] Feature extraction is performed on the training data to obtain prompt information;

[0010] The training data and the prompt information are combined into mixed data;

[0011] Using the mixed data, a large language model is trained to obtain a customer service proactive dialogue model.

[0012] As one possible implementation, the step of extracting features from the training data to obtain prompt information includes:

[0013] Get the preset prompt template;

[0014] Using each example of general dialogue data or real dialogue data in the training data as a unit, extract the prompt information according to the preset prompt template.

[0015] As one possible implementation, the step of training a large language model using the mixed data to obtain a customer service proactive dialogue model includes:

[0016] Using the large language model, predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data;

[0017] Calculate the loss between the predicted customer service dialogue data and the customer service dialogue data in the mixed data;

[0018] Based on the loss, the parameters of the large language model are adjusted;

[0019] Return to the step of using the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data until the large language model meets the preset conditions to obtain the customer service proactive dialogue model.

[0020] As one possible implementation, after obtaining the customer service proactive dialogue model, the method further includes:

[0021] The customer service proactive dialogue model was evaluated from multiple preset dimensions, resulting in multiple evaluation results;

[0022] Based on the multiple evaluation results, a comprehensive evaluation result is obtained;

[0023] If the comprehensive evaluation result does not meet the preset evaluation conditions, return to the step of using the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data.

[0024] As one possible implementation, the preset conditions include any combination of the following:

[0025] Return to the execution of the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data to reach the preset number of times;

[0026] The loss converges or is within a preset loss range;

[0027] The level of confusion in the customer service proactive dialogue model is within a preset confusion range;

[0028] The large-scale multi-task language understanding capability of the customer service proactive dialogue model is within the preset understanding capability range.

[0029] As one possible implementation, obtaining a comprehensive evaluation result based on the multiple evaluation results includes:

[0030] According to the preset reward mapping rules, each evaluation result is mapped to a preferred reward value;

[0031] A comprehensive evaluation result is obtained based on multiple preference reward values.

[0032] As one possible implementation, mapping each evaluation result to a preference reward value according to a preset reward mapping rule includes:

[0033] Each evaluation result is mapped to a preference reward value using formulas (1)-(2), where,

[0034] Formula (1) is

[0035] Formula (2) is

[0036] Where, r i It is the evaluation result, sign(r) i f(r) is an indicator function. i The evaluation result is r. i The preferred reward value, tanh is the hyperbolic tangent function, and f(r) is the preset control value. i The sensitivity parameter, x0, is the preset control f(r). i The parameters of the center point of the curve; and

[0037] The process of obtaining a comprehensive evaluation result based on multiple preference reward values ​​includes:

[0038] The comprehensive evaluation result is obtained using formula (3), where,

[0039] Formula (3) is

[0040] Where, r mo It is a comprehensive evaluation result, w k For the evaluation result r i The corresponding preset weights, where n is the number of evaluation results, and f(r) i The evaluation result is r. i The preference reward value.

[0041] Secondly, embodiments of this disclosure provide a method for proactive customer service dialogue, the method comprising:

[0042] Obtain instruction information;

[0043] The instruction information and real-time customer dialogue data are input into the trained customer service proactive dialogue model, and the real-time customer dialogue data is output. The customer service proactive dialogue model is trained using any of the methods described in the first aspect.

[0044] Thirdly, embodiments of this disclosure provide a training apparatus for a customer service proactive dialogue model, the apparatus comprising:

[0045] The training data acquisition module is used to acquire training data, which includes general dialogue data and real dialogue data. Both the general dialogue data and the real dialogue data include customer service dialogue data and customer dialogue data.

[0046] The feature extraction module is used to extract features from the training data to obtain prompt information;

[0047] The combining module is used to combine the training data and the prompt information into mixed data;

[0048] The training module is used to train the large language model using the mixed data to obtain a customer service proactive dialogue model.

[0049] Fourthly, embodiments of this disclosure provide a customer service proactive dialogue device, the device comprising:

[0050] Indication information acquisition module, used to acquire indication information;

[0051] The dialogue module is used to input the instruction information and real-time customer dialogue data into the trained customer service proactive dialogue model and output real-time customer service dialogue data. The customer service proactive dialogue model is trained using any of the methods described in the first aspect.

[0052] Fifthly, embodiments of this disclosure provide an electronic device comprising:

[0053] One or more processors;

[0054] Storage device, on which one or more programs are stored,

[0055] When the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in either the first aspect and / or the second aspect.

[0056] In a sixth aspect, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any of the first and / or second aspects.

[0057] In a seventh aspect, embodiments of this disclosure provide a computer program product including a computer program / instructions that, when executed by a processor, implement the method as described in any of the first and / or second aspects.

[0058] To enable proactive customer service dialogue systems based on large models in customer service scenarios, embodiments of this disclosure provide a method, apparatus, device, medium, and product for training proactive customer service dialogue models. This involves first acquiring training data, including general dialogue data and real dialogue data, both of which include customer service dialogue data and customer dialogue data. Then, feature extraction is performed on the training data to obtain prompt information. Next, the training data and the prompt information are combined into hybrid data. Finally, the hybrid data is used to train a large language model to obtain a proactive customer service dialogue model.

[0059] The model trained in this way can produce outputs that are closer to the customer service field based on prompts. Like a customer service representative, it can proactively guide customers to express their needs more clearly and accurately through multi-round dialogues to generate more reliable and accurate conversations. Attached Figure Description

[0060] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments, taken with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the scope of this disclosure. In the drawings:

[0061] Figure 1 is an exemplary system architecture diagram in which an embodiment of this disclosure can be applied;

[0062] Figure 2 is a flowchart of a customer service proactive dialogue model training method according to an embodiment of the present disclosure;

[0063] Figure 3A is a flowchart of a customer service proactive dialogue model training method according to another embodiment of this disclosure;

[0064] Figure 3B is a flowchart of step 302 of an embodiment of the present disclosure;

[0065] Figure 3C is a flowchart of step 304 of an embodiment of the present disclosure;

[0066] Figure 3D is a flowchart of step 3041 of an embodiment of this disclosure;

[0067] Figure 3E is a flowchart of step 3046 of an embodiment of this disclosure;

[0068] Figure 4 is an architecture diagram of a customer service proactive dialogue model training method according to an embodiment of this disclosure;

[0069] Figure 5 is a preference reward value curve of an embodiment of the present disclosure;

[0070] Figure 6 is a flowchart of a customer service proactive dialogue method according to an embodiment of the present disclosure;

[0071] Figure 7 is a structural diagram of a customer service proactive dialogue model training device according to an embodiment of the present disclosure;

[0072] Figure 8 is a structural diagram of a customer service proactive dialogue device according to an embodiment of the present disclosure;

[0073] Figure 9 is a schematic diagram of the structure of a computer system of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0074] As businesses grow and market competition intensifies, customer service has become a crucial means for companies to establish and maintain positive relationships with their customers. However, with expanding business scale and increasing customer traffic, especially with the deepening of digital transformation, the demand for intelligent and personalized services from both businesses and consumers has grown significantly. Relying solely on human customer service is insufficient to meet the needs of a large number of customers, resulting in low efficiency and high costs. To address these issues, automation technology has been introduced, giving rise to intelligent customer service chatbots.

[0075] Early intelligent customer service chatbots relied on preset rules and knowledge bases, which were highly efficient at handling common and standard questions. However, their comprehension capabilities were limited; they could only understand preset questions and keywords and struggled to handle non-standard expressions, such as those using dialects, industry jargon, or vague wording. If a customer's question did not match the correct keywords or phrases, the system might not be able to provide the correct answer.

[0076] With the development of machine learning technology, pattern-matching-based chatbots have replaced rule-based chatbots. These chatbots possess a certain level of intelligence and can handle complex dialogue scenarios, but their responses are fixed, lacking flexibility and a natural, fluid interactive experience. The rise of deep learning technology has ushered in a new era for chatbots, employing hybrid AI technologies such as RNNs and GANs to make dialogue more natural and human-like. However, they still have limitations in processing multi-turn dialogue context information, have limited generalization capabilities, and cannot provide personalized services.

[0077] With the development of artificial intelligence technology, chatbots based on large models have become a key tool in the customer service field. These large models, trained on ultra-large datasets, can accurately parse and understand customers' natural language input, generating fluent and natural language that makes the conversational experience closer to human-to-human communication.

[0078] However, most large-scale dialogue systems on the market currently offer open-topic systems without specified topics. The application of large-scale models in customer service is still under exploration. Traditional customer service dialogue systems can only passively respond to customer inquiries and cannot proactively engage in product-related communication like human customer service representatives, thus reducing customer participation in specific areas.

[0079] While some positive progress has been made in proactive dialogue within the social sphere—for example, generative language models have been used as primary response generators, and various innovative solutions have been introduced to dynamically select prompting strategies based on customer intent, personality, and emotion, thereby achieving personalized and proactive high-quality dialogue—some technologies have also explored proactive topic-switching mechanisms that can intelligently determine when to switch conversation topics and integrate external resources into the dialogue. However, although these technologies demonstrate a degree of proactivity in social scenarios, primarily optimized for social settings and adept at initiating new topics and guiding the conversation forward, they fail to effectively guide customers to express their specific needs in customer service scenarios.

[0080] In summary, the relevant technologies still have limitations. In the customer service field, they remain stuck in a question-and-answer format, passively responding to customer inquiries rather than proactively engaging with customers like human agents. Especially when customers don't ask questions, the technologies lack the ability to proactively inquire about their specific needs. In the marketing customer service field, where data volume is very limited, existing technologies cannot be applied to the complex scenarios, particularly when dealing with diverse marketing products and varying customer conversation styles, making effective adaptation difficult. Furthermore, due to the unique nature of marketing data, there are no specific optimization metrics to align with human preferences, and a lack of dedicated evaluation metrics to effectively assess the models.

[0081] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0082] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0083] Figure 1 illustrates an exemplary system architecture 100 of embodiments of customer service proactive dialogue and its model training methods, apparatus, devices, media and products that can be applied to this disclosure.

[0084] As shown in Figure 1, the system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as the medium for providing communication links between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0085] Customers can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as chat applications, voice recognition applications, short video social applications, audio and video conferencing applications, live video streaming applications, document editing applications, input method applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0086] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed on the terminal devices listed above. They can be implemented as multiple software programs or software modules (e.g., for providing dialogue services) or as a single software program or software module. No specific limitations are imposed here.

[0087] In some cases, the customer service proactive dialogue and its model training method provided in this disclosure can be executed by terminal devices 101, 102, and 103, and correspondingly, the customer service proactive dialogue and its model training device can be set in terminal devices 101, 102, and 103. In this case, the system architecture 100 may not include server 105.

[0088] In some cases, the customer service proactive dialogue and model training method provided in this disclosure can be jointly executed by terminal devices 101, 102, and 103 and server 105. For example, the step of "acquiring training data, which includes general dialogue data and real dialogue data, both of which include customer service dialogue data and customer dialogue data" can be executed by terminal devices 101, 102, and 103, and the step of "training a large language model using the mixed data to obtain a customer service proactive dialogue model" can be executed by server 105. This disclosure does not limit this. Correspondingly, the customer service proactive dialogue model training device can also be respectively set in terminal devices 101, 102, and 103 and server 105.

[0089] In some cases, the customer service proactive dialogue and its model training method provided in this disclosure can be executed by server 105. Accordingly, the customer service proactive dialogue model training device can also be set in server 105. In this case, the system architecture 100 may not include terminal devices 101, 102, and 103.

[0090] It should be noted that server 105 can be either hardware or software. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules (e.g., used to provide distributed services), or as a single software program or software module. No specific limitations are made here.

[0091] It should be understood that the number of terminal devices, networks, and servers shown in Figure 1 is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0092] All information, data, and signals disclosed herein are authorized by users / customers or by all parties, and the collection, use, and processing of such data comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0093] Referring again to Figure 2, which illustrates a flow 200 of an embodiment of a customer service proactive dialogue model training method according to the present disclosure, the method comprising at least the following steps 201 to 204.

[0094] Step 201: Obtain training data.

[0095] The training data includes general dialogue data and real dialogue data, both of which include customer service dialogue data and customer dialogue data.

[0096] The general dialogue data includes data from customer service interactions with customers, specifically question-and-answer dialogues. This data is used to prevent the model from losing general knowledge and expressive capabilities. For example, a typical example of general dialogue data might be:

[0097] Customer: How are xxx tiles?

[0098] Customer service: xxx is a well-known domestic company...

[0099] Client: How can we achieve this...?

[0100] Customer service: To achieve one...

[0101] and real conversation data, including conversation data initiatively initiated by customer service. It is conversation data between customers and customer service from real scenarios, which can not only include the question-and-answer format where customer service responds to conversations initiated by customers, but also include conversation data initiatively initiated by customer service. For example, a certain piece of real conversation data is:

[0102] Customer Service: Hello, is this Ms. Zhang?

[0103] Customer: Hello, who are you?

[0104] Customer Service: I am a customer service from XX Bank...

[0105] Customer Service: Do you understand my explanation?

[0106] Optionally, after obtaining the training data, it can be selectively cleaned, for example, illogical and repeated conversation content can be screened out or corrected. Furthermore, since conversations with too few turns are unstable and conversations with too many turns contain too little effective information, data whose conversation turns meet the preset elimination conditions are further eliminated. The preset elimination condition may be that the number of conversation turns is less than a turns or more than b turns (where a < b, and both a and b are natural numbers, which can be set according to the training data and are not limited herein), so as to obtain high-quality real conversation data. Taking a = 3 and b = 20 as an example, after the above processing, 755 pieces of real conversation data can be obtained.

[0107] Optionally, the training data may include general dialogue data and real dialogue data in a preset proportion. Preferably, the preset proportion can be 1:1. Thus, when there are 755 pieces of real dialogue data, more than 1500 pieces of training data can be obtained after mixing. Training the model with training data including a mixture of general dialogue data and real dialogue data can enhance the generalization ability of the model.

[0108] Step 202, performing feature extraction on the training data to obtain prompt information.

[0109] The prompt information may optionally include one or more of role information, customer information, product information, speech scripts, frequently asked questions (FAQ, Frequently Asked Questions), and conversation purpose, which are used to guide model training, reduce the generation of irrelevant or inaccurate content by the model, and better control the generation direction of the model. It can be well adapted when the amount of data in the customer service field is very small, especially when facing various marketing products and different users' conversation styles.

[0110] Wherein, the role information is the role of the customer service, such as product manager, top sales, and consultant.

[0111] Customer information refers to information about customers, which may include their basic identity information, business-related information, etc.

[0112] Product information refers to information about the target product, including but not limited to product name, functional description, components, and product specifications. In some marketing scenario examples, customer service representatives need to sell the product corresponding to this product information. It should be understood that this disclosure can be applied to marketing scenarios in various fields, including but not limited to e-commerce, catering, apparel, finance, and insurance. In other embodiments, this disclosure can also be applied to customer service scenarios other than marketing scenarios, such as product demonstrations and after-sales service for products or services.

[0113] A script is a dialogue guide or facilitator used to guide responses, inquiries, and / or the direction of the conversation. In some examples, scripts can be pre-set information. In some marketing scenario embodiments, scripts are dialogue guides or facilitators based on the marketing scenario. It should be understood that this disclosure can be applied to marketing scenarios in different fields, including but not limited to e-commerce, catering, apparel, finance, and insurance. In other embodiments, this disclosure can also be applied to customer service scenarios other than marketing scenarios, such as product demonstrations and after-sales service for products or services.

[0114] FAQs are frequently asked and answered questions by customer service representatives and customers.

[0115] The purpose of a conversation is the expected outcome of that conversation. For example, in some marketing scenarios, the primary purpose of the conversation is to facilitate a transaction, while in some consulting scenarios, the primary purpose is to gather information about the customer.

[0116] Optionally, a template for the prompt information can be pre-set according to the scenario, and then the corresponding features can be extracted according to the template. In some embodiments, the content of the prompt information can be adjusted according to the scenario and needs, without limitation. The prompt information of the training data can be obtained through 202, and this prompt information can be specific and detailed. Specific and detailed prompt information can significantly improve the reliability of the output of the customer service proactive dialogue model, making the model's output results more accurate and reliable.

[0117] Step 203: Combine the training data and the prompt information into mixed data.

[0118] By combining training data and prompts into a hybrid dataset, and then training a large language model, good generalization results can be achieved at minimal cost by simply adjusting the content of the prompts in different customer service scenarios.

[0119] The training data is used to overcome the limitations of traditional models that can only passively respond, enabling the model to move beyond a question-and-answer format. It can not only passively engage with customer inquiries but also actively communicate with customers like a human customer service representative. The prompts are used to assist training, providing guidance for model training and helping to better control the direction of model generation, making it more adaptable to different customers and customer service scenarios.

[0120] Step 204: Use the mixed data to train the large language model to obtain the customer service proactive dialogue model.

[0121] Large Language Model (LLM) can be a language model pre-trained on large-scale data. Mixed data introduces specific prompts into its training process, amplifying the guiding role of these prompts to guide and optimize the model's learning of customer service domain knowledge and dialogue skills, generating a customer service model with proactive dialogue capabilities.

[0122] Furthermore, the model's training data is based on real dialogue data. Through deep learning and optimization, the large language model can proactively communicate with customers. Even when customers do not ask questions, the large language model has the function of proactively inquiring about customer needs, realizing a more intelligent and proactive customer interaction mode, and providing an innovative customer service solution for various industries.

[0123] In summary, this embodiment incorporates prompts as guidance during the training of the customer service proactive dialogue model, significantly enhancing the relevance of the model's output and ensuring that it incorporates more prompts, thereby improving marketing effectiveness. Furthermore, the trained customer service proactive dialogue model can proactively communicate with customers, guiding them to clearly express their needs.

[0124] Continuing with Figures 3A and 4, a flow 300 of another embodiment of the customer service proactive dialogue model training method according to the present disclosure is shown, which includes at least the following steps 301 to 305.

[0125] Step 301: Obtain training data.

[0126] As one possible implementation, after acquiring the training data, the training data can be processed into the input format required by the model.

[0127] For example, convert the training data "Customer Service: 'Hello, is this Ms. Zhang?' Customer: 'Hello, who are you?' Customer Service: 'I am a customer service representative from xx bank...'... Customer Service: 'Do you understand my explanation?'" into the following message format:

[0128] [{"role":"assistant","content":"Hello, are you Ms. Zhang?"}

[0129] {"role":"user","content":"Hello, who are you?"},

[0130] {"role":"assistant","content":"I am a customer service representative from xx bank..."}

[0131] ...

[0132] {"role":"assistant","content":"Do you understand my explanation?"}]

[0133] Step 302: Extract features from the training data to obtain prompt information.

[0134] As one possible implementation, step 302 may include steps 3021 to 3022 as shown in FIG3B. Through this implementation, information related to the dialogue occurrence scenario can be set by a preset prompt template, so that the prompt template can guide the direction of model training, thereby making the model output more targeted.

[0135] Step 3021: Obtain the preset prompt template.

[0136] The preset prompt template can be a template pre-defined using Prompt Engineering. The content of the prompt template can be set according to the customer service and / or service scenario. For example, in a financial scenario, it may include sales-related information, and in a medical scenario, it may include information related to a patient's condition. For instance, a prompt template for a financial scenario might take the following form:

[0137] Step 3022: Using each example of general dialogue data or real dialogue data in the training data as a unit, extract the prompt information according to the preset prompt template.

[0138] Most existing large-scale models can perform various complex text generation and understanding tasks. Therefore, any large-scale model can be directly used to extract prompt information from the training data according to a preset prompt template. In the embodiments of this disclosure, a corresponding prompt message can be extracted from each example of general dialogue data or real dialogue data.

[0139] Optionally, the large model for extracting prompt information can be a generative pre-trained language model, such as GPT (Generative Pre-trained Transformer).

[0140] For example, using a specific example of real dialogue data as a unit, and following the template in step 3021, the following prompt information can be extracted:

[0141] For example, using a specific example of real-world dialogue data as a unit, and following the template in step 3021, the following prompt information can be extracted:

[0142] Step 303: Combine the training data and the prompt information into mixed data.

[0143] When combining training data and prompts into mixed data, the prompts need to be placed before the training data (e.g., the prompts should be positioned before the training data) to align with the large language model's pattern of predicting the following text based on the preceding text, thus enabling the prompts to guide the output.

[0144] As one possible implementation, the prompt information can be first converted into messages format and then appended to the corresponding messages format training data.

[0145] For example, the training data is information in the format of messages, as shown below:

[0146] [{"role":"assistant","content":"Hello, are you Ms. Zhang?"}

[0147] {"role":"user","content":"Hello, who are you?"},

[0148] {"role":"assistant","content":"I am a customer service representative from xx bank..."}

[0149] ...

[0150] {"role":"assistant","content":"Do you understand my explanation?"}]

[0151] After adding a cue message before the training data, the hybrid data combining the training data and the cue message will have the following message format:

[0152] [{"role":"system","content":"prompt message"},

[0153] [{"role":"assistant","content":"Hello, are you Ms. Zhang?"}

[0154] {"role":"user","content":"Hello, who are you?"},

[0155] {"role":"assistant","content":"I am a customer service representative from xx bank..."}

[0156] ...

[0157] {"role":"assistant","content":"Do you understand my explanation?"}).

[0158] Step 304: Use the mixed data to train the large language model to obtain the customer service proactive dialogue model.

[0159] As one possible implementation method, the training process can be divided into two stages: supervised fine-tuning and alignment reinforcement learning.

[0160] Supervised fine-tuning may include at least steps 3041 to 3047 as shown in Figure 3C, for training the model so that it can learn to engage in dialogue with customers based on prompts and achieve proactive dialogue with customers.

[0161] Step 3041: Use a large language model to predict customer service dialogues in mixed data and output predicted customer service dialogue data.

[0162] As one possible implementation, step 3041 may include at least steps 30411 to 30413 as shown in FIG3D.

[0163] Step 30411: Convert the input mixed data from messages format to string format.

[0164] The string format is a pattern that large language models can recognize. See below for the string format:

[0165] <|im_start|>system

[0166] **

Character Information

[0167] **Role:** Top Financial Product Salesperson

[0168] **Objective:** To recommend suitable fund products, encourage client purchases, and boost investment confidence.

[0169] …

[0170] <|im_end|>

[0171] <|im_start|>user

[0172] Hi, what's up?

[0173] <|im_start|>assistant

[0174] Hello, I am your sales manager.

[0175] Among them, <|im_end|> and <|im_start|> are special tokens used during the pre-training of large language models to distinguish different speaking roles.

[0176] Step 30412: Convert the input mixed data from string format to token format.

[0177] A token represents the basic unit that the model processes and generates; it can be a word, a subword, or a character. Large language models typically use complex tokenization to handle the vast diversity of human language while keeping the vocabulary size manageable.

[0178] When converting mixed data from string format to token format, you can first segment the mixed data into words, and then refer to the vocabulary list for conversion.

[0179] For example, the large language model Qwen has a relatively large vocabulary (vocabulary size, i.e., the total number of unique tokens the model recognizes, has a significant impact on the model's performance and versatility), with 151,646 tokens. Qwen employs a sub-word tokenization method called Byte Pair Encoding (BPE), which attempts to learn token combinations that can represent text using the fewest possible tokens.

[0180] This disclosure allows for the transformation of mixed data based on the vocabulary. For example, the string "tokenization" can be segmented into two words, "token" and "ization," using any word segmentation method. The corresponding codes for these two words in the Qwen vocabulary are then found, resulting in "123" and "4215." Finally, the word segments of each example of mixed data are combined into a token format such as [123,4251,856,..].

[0181] Step 30413: Using a large language model, predict the customer service dialogue data corresponding to the mixed data and output the predicted customer service dialogue data.

[0182] The mixed data sequence combines prompts and training data, while the large language model predicts the following text based on the preceding context. Therefore, the predicted customer service dialogue data output in step 30413 is predicted based on the prompts and the dialogue preceding the current predicted customer service dialogue data, and can be guided by the prompts. This prediction method trains a causal model of language, also known as an autoregressive language model or a decoder-only language model.

[0183] For example, mixed data

[0184] [Prompt Message]

[0185] Customer service: Hello, are you Ms. Zhang?

[0186] Customer: Hello, who are you?

[0187] Customer service: I am a customer service representative from XX Bank...

[0188] Customer service: Do you understand my explanation?

[0189] So, when predicting "Hello, is this Ms. Zhang?", it is based on the "[Prompt Information]"; when predicting "I am a customer service representative from xx bank...", it is based on the "[Prompt Information], and, Hello, is this Ms. Zhang? and, Hello, who are you?".

[0190] In the sentence "Hello, is this Ms. Zhang?", "you" was predicted based on the "[Prompt Information]", and "hello" was predicted based on the "[Prompt Information] and you".

[0191] Specifically, the mixed data in token format is input into a large language model, which only makes predictions for customer service dialogue data in token format. For each token in the customer service dialogue data, the model generates a corresponding prediction label, which contains the predicted next token. Each time a prediction label is generated, it is added to the generated sequence. Finally, the generated sequence constitutes the customer service dialogue data in token format.

[0192] It should be noted that the prediction of the first token in the customer service dialogue data is based on the token format prompt information and the preceding dialogue data. Subsequent token predictions, however, take into account the token format prompt information, the preceding dialogue data, and the tokens in the generated sequence corresponding to the customer service dialogue data generated in this prediction.

[0193] The "causal" aspect means that when predicting the next token, the model considers the past context (i.e., the generated tokens) and does not consider future tokens.

[0194] For example: the predicted dialogue is "The weather is nice today.", and the prediction process is:

[0195] Generate the prediction label "天" based on the token "今",

[0196] Generate the prediction label "天" based on the tokens "今" and "天",

[0197] Generate the prediction label "气" based on the tokens "今", "天", and "天",

[0198] Generate the prediction label "很" based on the tokens "今", "天", "天", and "气",

[0199] …

[0200] Generate the prediction label "EOS" based on the tokens "今", "天", "天", "气", "很", and "好".

[0201] It should be noted that EOS is a special symbol indicating the end of a sentence.

[0202] Combining the above prediction labels gives the predicted dialogue data "The weather is nice today.".

[0203] In summary, steps 30411 to 30413 can use a large language model to predict customer service dialogue data in mixed data and output predicted customer service dialogue data.

[0204] Step 3042: calculate the loss between the predicted customer service dialogue data and the customer service dialogue data in the mixed data.

[0205] As a possible implementation, in order to improve the measurement effect of model prediction accuracy, the Cross Entropy Loss can be used as the loss function. The cross entropy loss function can calculate the difference between the probability distribution predicted by the model and the real label, so as to guide the optimization direction of the model.

[0206] Specifically, the loss between the tokens corresponding to the predicted customer service dialogue data and the tokens corresponding to the actual customer service dialogue data can be calculated. Step 3041 obtains the tokens corresponding to the predicted customer service dialogue data, while the tokens corresponding to the customer service dialogue data in the training data are the actual tokens corresponding to the customer service dialogue data. Calculating the loss function of the two can be used to train the model so that its dialogue data is closer to the actual customer service dialogue data.

[0207] Step 3043: Adjust the parameters of the large language model based on the loss.

[0208] When adjusting the parameters of a large language model, a backpropagation algorithm can be used to calculate the gradient of the loss with respect to each layer's parameters, starting from the output layer. Then, an optimizer is used to update the weights and biases of each layer based on the calculated gradients and a preset learning rate. As one possible implementation, the optimizer can be AdamW, with a learning rate set to 5e-6, or 5 × 10⁻⁶. -6 .

[0209] Step 3044: Return to step 3041 until the large language model meets the preset conditions, and obtain the customer service proactive dialogue model.

[0210] The preset conditions can be combined with one or more of the following dimensions to determine whether the model can stop training.

[0211] The first option is to return to step 3041 and execute the preset number of times.

[0212] Since a single training session may not yield satisfactory results, the training data can be used to iterate and train a preset number of times. The preset number of times can be set according to needs and is not limited here.

[0213] The second scenario is that the loss converges or falls within the preset loss range.

[0214] The requirement is met if the loss converges or falls within the preset loss range. Generally, training can be stopped when the loss is around 0.1-0.5. In supervised fine-tuning, if the loss is too high (1-2 or even higher), the training is insufficient; if it is too low (close to 0), overfitting is likely, which will affect the model's ability to respond correctly.

[0215] The third type is where the level of confusion in the customer service proactive dialogue model is within a preset range.

[0216] The perplexity (PPL) of a model measures its predictive ability for text. A PPL within a preset perplexity range is sufficient. When obtaining the predicted label for each token, multiple possible predicted labels are first obtained, each with a corresponding probability. Then, the most probable predicted label is selected from among these multiple possible labels.

[0217] Therefore, each predicted label corresponds to a probability. For example, the probabilities of each predicted label in the sentence "The weather is nice today" are estimated to be 0.5, 0.25, 0.125, and 0.125, respectively. Substituting these probabilities into the preset PPL calculation formula, we can obtain the PPL of the customer service proactive dialogue model. The PPL of the customer service proactive dialogue model can then be used to reflect the uncertainty of its "average guess".

[0218] For language models, a lower PPL value indicates that the model's prediction of the sentence is more accurate, and the generated text is more fluent and natural; conversely, a higher PPL value indicates that the model's prediction of the sentence is not accurate enough.

[0219] Therefore, when the model's PPL is within the preset perplexity range, it means that the model's prediction has reached a certain level of accuracy.

[0220] Fourthly, the large-scale multi-task language understanding capability of the customer service proactive dialogue model is within the preset understanding capability range.

[0221] Large-scale multitask language understanding ability can be measured using the MMLU (Massive Multitask Language Understanding) test score. MMLU is a benchmark test used to evaluate the understanding and reasoning abilities of large language models across multiple subject areas. It measures the model's logical reasoning ability to ensure that the model does not overfit during training due to an overly simplistic metric. It measures the breadth and depth of the model's knowledge, as well as its logical reasoning ability across different subjects, through a series of carefully designed multiple-choice questions. If the MMLU test score is within the preset benchmark score range, then the customer service proactive dialogue model's large-scale multitask language understanding ability is within the preset understanding ability range.

[0222] Before training (i.e., supervised fine-tuning), a large language model can be tested using MMLU to obtain a score. After the above training (i.e., supervised fine-tuning), the model's output performance improves. Simultaneously, it's necessary to ensure that the MMLU test score of the trained model is essentially consistent with that before supervised fine-tuning. Generally, a fluctuation range of ±3 percentage points in the MMLU test score is considered reasonable.

[0223] For example, if the model's MMLU score is 84.2 before fine-tuning and 82.1 after fine-tuning, such a small change is within an acceptable range, indicating that the model has not overfitted due to fine-tuning.

[0224] The preset number of attempts, preset loss range, preset confusion range, and preset benchmark score range can all be set as needed, and there are no restrictions here.

[0225] In summary, steps 3041 to 3044 enabled the supervised fine-tuning of the customer service proactive dialogue model.

[0226] As one possible implementation, after supervised fine-tuning, alignment reinforcement learning can be selectively applied to the customer service proactive dialogue model. Alignment reinforcement learning can include steps 3045 to 3047, whereby it introduces multiple preset dimensions and uses the target effects of these dimensions as the model optimization direction (i.e., the optimization direction of multiple objectives, which can also be referred to as the target effects or evaluation results of each preset dimension). This makes the trained customer service proactive dialogue model more consistent with human-defined dialogue habits and better able to meet dialogue needs. This not only helps optimize the customer dialogue experience but also facilitates the achievement of the predetermined dialogue goals of the customer service proactive dialogue model.

[0227] Step 3045: Evaluate the customer service proactive dialogue model from multiple preset dimensions to obtain multiple evaluation results.

[0228] Multiple preset dimensions can include the number of dialogue rounds, customer emotional positivity, dialogue objective achievement rate (e.g., marketing success rate, consultation completion rate), scoring models, etc. The evaluation results are used to represent the ability of the customer service proactive dialogue model in different preset dimensions, and the evaluation results can be quantitative scores.

[0229] The more rounds of dialogue there are, the more willing the customer is to participate in the current conversation, and the higher the quality of the model's output, the higher the score for this dimension.

[0230] Customer emotional positivity is assessed through sentiment analysis to determine the customer's emotional tendency. The model's output dialogue data should guide customer emotions in a positive direction as much as possible, avoiding low customer emotions that could negatively impact the conversation experience. The more positive the user's emotions, the higher the customer emotional positivity score, and the better the quality of the model's output.

[0231] The dialogue objective achievement rate measures whether the dialogue needs have been met, such as whether marketing was successful or consultation was completed. It is the model's ultimate goal. The higher the dialogue objective achievement rate, the better the dialogue needs are met, the more the dialogue objective is achieved, and the better the quality of the model's output.

[0232] The scoring model (Reward model) is a model trained based on human ratings and dialogue data. It can simulate human automated scoring of dialogue data. Inputting prompts, real customer dialogue data, and predicted customer service dialogue data into the model will output a score. The higher the output score, the better the quality of the model's output. Since the scoring model is used to simulate human scoring, the scoring in the above dimensions can also be replaced by direct human scoring; there are no restrictions on this.

[0233] It should be noted that, to standardize the scoring range of the scoring model, the output can be connected to a sigmoid function to convert the score into a probability value between 0 and 1. This probability value is then multiplied by a preset coefficient to obtain the score within the preset range. For example, setting the preset coefficient to 2 will result in a floating-point score between 0 and 2. When using the scoring model, only the prompt message, real customer dialogue data, and predicted customer service dialogue data need to be input to obtain a floating-point score between 0 and 2.

[0234] The evaluation results from multiple dimensions are used to characterize the capabilities of the customer service proactive dialogue model in different dimensions, thereby controlling the training effect of the model. This allows the customer service proactive dialogue model to balance multiple key factors when outputting dialogue data, while increasing the likelihood of successfully converting the dialogue into a meaningful outcome.

[0235] Step 3046: Based on multiple evaluation results, obtain a comprehensive evaluation result.

[0236] By incorporating evaluation results from multiple pre-defined dimensions, the customer service proactive dialogue model can balance several key factors when outputting dialogue, ensuring successful conversion of the dialogue objective (such as marketing goals). However, if different pre-defined dimensions are to be influenced by the output to varying degrees, the multiple evaluation results can be weighted and summed according to pre-set weights when obtaining the comprehensive evaluation result. This allows the comprehensive evaluation result to take into account multiple objectives (i.e., evaluation results from multiple dimensions). In this way, the model subsequently trained based on the comprehensive evaluation result can achieve better performance across multiple pre-defined dimensions.

[0237] As one possible implementation, the comprehensive evaluation results can be calculated directly using a linear weighted sum method.

[0238] Specifically, the weighting coefficients for each evaluation result (i.e., number of dialogue turns, emotional positivity, marketing success rate, and scoring model) are represented as w1, w2, w3, and w4. Adjusting these weights ensures that the importance of different evaluation results is reasonably reflected in the overall evaluation. In this way, the model can achieve a balance between dialogue interaction, emotional inclination, marketing objectives, and human preferences during the optimization process, ultimately improving customer satisfaction, interaction efficiency, and the rate of achieving dialogue objectives.

[0239] For example, the weighting coefficient for the number of dialogue turns (w1) is 0.4, the weighting coefficient for customer sentiment positivity (w2) is 0.3, the weighting coefficient for marketing success rate (w3) is 0.3, and the weighting coefficient for the scoring model (w4) is 0.5. In this example, we make assumptions for each evaluation result. The number of dialogue turns is 4, meaning the model generated 4 responses, resulting in a score of 4; the sentiment positivity score is 8, meaning that the sentiment analysis score is high and the customer's sentiment tends to be positive; the marketing success rate is 1, representing a 100% marketing conversion rate; the Reward scoring model output is 1.1, indicating that the scoring model considers the output of the customer service proactive dialogue model to receive a comprehensive score of 1.1.

[0240] The overall evaluation result is equal to the sum of w1 multiplied by the number of dialogue rounds, w2 multiplied by the emotional positivity evaluation result, w3 multiplied by the marketing success rate evaluation result, and w4 multiplied by the scoring model evaluation result.

[0241] Substituting the above values ​​into a weighted summation, the comprehensive evaluation result is as follows:

[0242] 0.4×4+0.3×8+0.3×1+0.5×1.1=1.6+2.4+0.3+0.22=4.52.

[0243] The comprehensive evaluation results obtained in this way can combine the number of dialogue rounds, emotional positivity, marketing success rate, and scoring.

[0244] As another possible implementation, the traditional alignment algorithm can be improved by changing the single-objective preference value reward in the traditional alignment algorithm to a multi-objective preference value reward to obtain a comprehensive evaluation result. Specifically, step 3046 may include steps 30461 and 30462 as shown in Figure 3E.

[0245] Step 30461: Map each evaluation result to a preferred reward value according to the preset reward mapping rules.

[0246] To enhance sensitivity to assessment results and identify more relevant information when calculating comprehensive evaluation results—that is, to make the comprehensive evaluation results more sensitive to assessments that are too low or too high—a preset reward mapping rule can be set to map each assessment result to a preferred reward value. The preset reward mapping rule can specify a linear or non-linear mapping relationship, mapping assessment results to preferred reward values; the specific mapping rule can be set according to requirements.

[0247] For example, each of the evaluation results can be mapped to a preference reward value using formulas (1)-(2), where,

[0248] Formula (1) is

[0249] Formula (2) is

[0250] Where, r i It is the evaluation result, sign(r) i f(r) is an indicator function. i The evaluation result is r. i The preferred reward value, tanh is the hyperbolic tangent function, and f(r) is the preset control value. i The sensitivity parameter, x0, is the preset control f(r). i The parameters of the center point of the curve.

[0251] It should be noted that k and x0 can be preset as needed, and are not limited here. See Figure 5, where k controls f(r) i The steepness (i.e., sensitivity) of the curve increases with increasing k. i The higher the curve sensitivity, the more sensitive the curve becomes; x0 controls the center point of the curve, determining the value around which the curve changes most rapidly. Optionally, k=5, x0=0.5 can be set.

[0252] Step 30462: Obtain a comprehensive evaluation result based on multiple preference reward values.

[0253] After obtaining the preference reward values ​​corresponding to different evaluation results, the comprehensive evaluation result is obtained by weighted summation using formula (3), where,

[0254] Formula (3) is

[0255] Where, r mo It is a comprehensive evaluation result, w k For the evaluation result r i The corresponding preset weights, where n is the number of evaluation results, and f(r) i The evaluation result is r. i The preference reward value.

[0256] The comprehensive evaluation results calculated in steps 30461 to 30462 are more sensitive to evaluation results that are too high or too low, thereby controlling the training in step 3047 and improving the convergence speed and training efficiency.

[0257] Step 3047: In response to the fact that the comprehensive evaluation result does not meet the preset evaluation conditions, return to step 3041.

[0258] If the overall evaluation result does not meet the preset evaluation conditions, the system will return to step 3041 for retraining until the overall evaluation result meets the preset evaluation conditions, meaning the customer service proactive dialogue model performs well across multiple preset dimensions. The preset evaluation conditions can be set according to requirements and are not limited here.

[0259] In summary, steps 3045 to 3047 evaluate the capabilities of the customer service proactive dialogue model from multiple preset dimensions, and control the model for further training until it achieves good performance across these dimensions, better aligns with the dialogue habits of each dimension, and better meets the dialogue needs. This not only improves the customer dialogue experience but also helps achieve the dialogue objectives of the customer service proactive dialogue model.

[0260] In summary, the customer service proactive dialogue model obtained through the above training demonstrates good performance across multiple evaluation systems, including loss, PPL, MMLU, number of dialogue turns, customer emotional positivity, marketing success rate, and scoring model, achieving beneficial effects in at least the following three dimensions:

[0261] (1) It can automatically act as a customer service representative and proactively communicate with customers, breaking the limitations of the traditional model and realizing a more intelligent and proactive customer interaction mode, providing an innovative customer service solution for various industries;

[0262] (2) By using prompting engineering, key information is extracted for specific customer service scenarios to generate prompts to assist training, enabling the model to adapt to different products and scenarios. This provides an effective strategy to improve the generalization performance of the model with a small amount of data.

[0263] (3) The multi-objective preference value reward method based on weight allocation makes the intelligent customer service robot more inclined to human dialogue habits, providing an important reference for the design of optimization objectives in complex dialogue scenarios.

[0264] Referring again to Figure 6, which shows a flow 600 of an embodiment of the customer service proactive dialogue method according to the present disclosure, which is equivalent to the reasoning process of the trained model described above. The dialogue data obtained by reasoning can still be used as training data to iteratively train the model. The customer service proactive dialogue method includes the following steps 601-602.

[0265] Step 601: Obtain instruction information.

[0266] The prompt information is extracted from the features of the training data. However, there is no complete dialogue data in the actual use of the model. Therefore, the prompt information can be obtained from the system, including customer service information, specific customer information, specific information of the product to be recommended, specific excellent scripts for this scenario, and the question-answer pair candidate set (FAQ) retrieved by the embedding model vector. The prompt information is then used to instruct the model output.

[0267] In subsequent conversations, the aforementioned instructions can guide the model to output more accurate customer service responses, automatically specifying conversation content and response language style, thereby improving scenario relevance and / or marketing success rate.

[0268] Step 602: Input the instruction information and real-time customer dialogue data into the trained customer service proactive dialogue model, and output real-time customer service dialogue data.

[0269] Because the dialogue is continuous, customer conversation data can be continuously received and the output customer service conversation data can be recorded. The first customer service response can be based on the instruction information and / or the customer's initial conversation data, and thereafter, customer service responses can be based on the instruction information and all previous conversation data.

[0270] In summary, when the trained customer service proactive dialogue model is used, it can output information that is strongly related to the instructions and generate proactive dialogue, thereby improving the targeting of the dialogue and thus the marketing success rate, and ensuring the customer's user experience.

[0271] Referring further to Figure 7, as an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a customer service proactive dialogue model training device, which corresponds to the method embodiment shown in Figure 2, and the device can be specifically applied to various electronic devices.

[0272] As shown in Figure 7, the customer service proactive dialogue model training device 700 of this embodiment includes: a training data acquisition module 701, a feature extraction module 702, a combination module 703, and a training module 704.

[0273] Among them, the training data acquisition module 701 is used to acquire training data, which includes general dialogue data and real dialogue data, and both the general dialogue data and real dialogue data include customer service dialogue data and customer dialogue data.

[0274] Feature extraction module 702 is used to extract features from the training data to obtain prompt information;

[0275] Combined module 703 is used to combine the training data and the prompt information into mixed data;

[0276] Training module 704 is used to train the large language model using the mixed data to obtain a customer service proactive dialogue model.

[0277] As one possible implementation, the feature extraction module 702 includes:

[0278] Template preset unit 7021 (not shown in the figure) is used to obtain preset prompt templates;

[0279] Extraction unit 7022 (not shown in the figure) is used to extract prompt information according to the preset prompt template, taking each example of general dialogue data or real dialogue data in the training data as a unit.

[0280] As one possible implementation, the training module 704 includes:

[0281] Prediction unit 7041 (not shown in the figure) is used to predict customer service dialogue data corresponding to the mixed data using the large language model and output the predicted customer service dialogue data.

[0282] The loss calculation unit 7042 (not shown in the figure) is used to calculate the loss between the predicted customer service dialogue data and the customer service dialogue data in the mixed data;

[0283] Adjustment unit 7043 (not shown in the figure) is used to adjust the parameters of the large language model based on the loss;

[0284] Return to training unit 7044 (not shown in the figure), which is used to return the execution of using the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data until the large language model meets the preset conditions to obtain the customer service proactive dialogue model.

[0285] As one possible implementation, the customer service proactive dialogue model training device 700 further includes:

[0286] The multi-dimensional evaluation module 705 is used to evaluate the customer service proactive dialogue model from multiple preset dimensions and obtain multiple evaluation results.

[0287] The comprehensive evaluation module 706 is used to obtain a comprehensive evaluation result based on the multiple evaluation results;

[0288] Return to training module 707, which is used to respond to the fact that the comprehensive evaluation result does not meet the preset evaluation conditions, return to execute the step of using the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data.

[0289] As one possible implementation, the preset conditions include any combination of the following:

[0290] Return to the execution of the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data to reach the preset number of times;

[0291] The loss converges or is within a preset loss range;

[0292] The level of confusion in the customer service proactive dialogue model is within a preset confusion range;

[0293] The large-scale multi-task language understanding capability of the customer service proactive dialogue model is within the preset understanding capability range.

[0294] As one possible implementation, the comprehensive evaluation module 706 includes:

[0295] The preference calculation unit 7061 is used to map each evaluation result into a preference reward value according to a preset reward mapping rule;

[0296] The comprehensive calculation unit 7062 is used to obtain a comprehensive evaluation result based on multiple preference reward values.

[0297] As one possible implementation, the preference calculation unit 7061 includes:

[0298] The first formula calculation component 70611 (not shown in the figure) is used to map each of the evaluation results to a preference reward value using formulas (1)-(2), wherein,

[0299] Formula (1) is

[0300] Formula (2) is

[0301] Where, r i It is the evaluation result, sign(r) i f(r) is an indicator function. i The evaluation result is r. i The preferred reward value, tanh is the hyperbolic tangent function, and f(r) is the preset control value. i The sensitivity parameter, x0, is the preset control f(r). i The parameters of the center point of the curve; and

[0302] The integrated computing unit 7062 includes:

[0303] The second formula calculation component 70621 (not shown in the figure) is used to obtain the comprehensive evaluation result using formula (3), wherein,

[0304] Formula (3) is

[0305] Where, r mo It is a comprehensive evaluation result, w k For the evaluation result r i The corresponding preset weights, where n is the number of evaluation results, and f(r) i The evaluation result is r. i The preference reward value.

[0306] Referring further to FIG8, as an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a customer service proactive dialogue device, which corresponds to the method embodiment shown in FIG6, and the device can be specifically applied to various electronic devices.

[0307] As shown in Figure 8, the customer service proactive dialogue device 800 of this embodiment includes: an instruction information acquisition module 801 and a dialogue module 802.

[0308] Among them, the instruction information acquisition module 801 is used to acquire instruction information;

[0309] The dialogue module 802 is used to input the instruction information and real-time customer dialogue data into the trained customer service proactive dialogue model and output real-time customer service dialogue data. The customer service proactive dialogue model is trained using the customer service proactive dialogue model training method shown in Figure 2 or Figure 3A.

[0310] Referring now to FIG9, a schematic diagram of a computer system 900 suitable for implementing the electronic device of the present disclosure is shown. The computer system 900 shown in FIG9 is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present disclosure.

[0311] As shown in Figure 9, the computer system 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the computer system 900. The processing device 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0312] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows computer system 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 illustrates a computer system 900 with various electronic devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0313] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.

[0314] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0315] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0316] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to implement the methods shown in any of the embodiments and optional embodiments of FIG2 to FIG6.

[0317] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the client computer, partially on the client computer, as a standalone software package, partially on the client computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the client computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0318] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0319] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily limit the unit itself; for example, a training module can also be described as "used to return the execution of using the large language model to predict customer service dialogue data corresponding to the mixed data, output predicted customer service dialogue data, until the large language model meets preset conditions, thus obtaining a customer service proactive dialogue model."

[0320] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

Claims

1. A method for training a customer service proactive dialogue model, characterized in that, include: Acquire training data, which includes general dialogue data and real dialogue data, both of which include customer service dialogue data and customer dialogue data. Feature extraction is performed on the training data to obtain prompt information; The training data and the prompt information are combined into mixed data; Using the mixed data, a large language model is trained to obtain a customer service proactive dialogue model.

2. The method according to claim 1, characterized in that, The step of extracting features from the training data to obtain prompt information includes: Get the preset prompt template; Using each example of general dialogue data or real dialogue data in the training data as a unit, prompt information is extracted according to the preset prompt template, wherein the prompt information includes customer service domain knowledge and / or dialogue skills used to guide the training of the large language model.

3. The method according to claim 1 or 2, characterized in that, The process of training a large language model using the mixed data to obtain a customer service proactive dialogue model includes: Using the large language model, predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data; Calculate the loss between the predicted customer service dialogue data and the customer service dialogue data in the mixed data; Based on the loss, the parameters of the large language model are adjusted; Return to the step of using the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data until the large language model meets the preset conditions to obtain the customer service proactive dialogue model.

4. The method according to claim 3, characterized in that, After obtaining the customer service proactive dialogue model, the following is also included: The customer service proactive dialogue model was evaluated from multiple preset dimensions, resulting in multiple evaluation results; Based on the multiple evaluation results, a comprehensive evaluation result is obtained; If the comprehensive evaluation result does not meet the preset evaluation conditions, return to the step of using the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data.

5. The method according to claim 3, characterized in that, The preset conditions include any combination of the following: Return to the execution of the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data to reach the preset number of times; The loss converges or is within a preset loss range; The level of confusion in the customer service proactive dialogue model is within a preset confusion range; The large-scale multi-task language understanding capability of the customer service proactive dialogue model is within the preset understanding capability range.

6. The method according to claim 4, characterized in that, The process of obtaining a comprehensive evaluation result based on the multiple evaluation results includes: According to the preset reward mapping rules, each evaluation result is mapped to a preferred reward value; A comprehensive evaluation result is obtained based on multiple preference reward values.

7. The method according to claim 6, characterized in that, The step of mapping each evaluation result to a preferred reward value according to a preset reward mapping rule includes: Each evaluation result is mapped to a preference reward value using formulas (1)-(2), where, Formula (1) is Formula (2) is Where, r i It is the evaluation result, sign(r) i f(r) is an indicator function. i The evaluation result is r. i The preferred reward value, tanh is the hyperbolic tangent function, and k is the preset control f(r) value. i The sensitivity parameter, x0, is the preset control f(r). i The parameters of the center point of the curve; and The process of obtaining a comprehensive evaluation result based on multiple preference reward values ​​includes: The comprehensive evaluation result is obtained using formula (3), where, Formula (3) is Where, r mo It is a comprehensive evaluation result, w k For the evaluation result r i The corresponding preset weights, where n is the number of evaluation results, and f(r) i The evaluation result is r. i The preference reward value.

8. The method according to claim 1, characterized in that, The step of combining the training data and the prompt information into mixed data includes: The prompt information is incorporated into the training data as mixed data.

9. A method for proactive customer service dialogue, characterized in that, include: Obtain instruction information; The instruction information and real-time customer dialogue data are input into the trained customer service proactive dialogue model, and the real-time customer service dialogue data is output. The customer service proactive dialogue model is trained using the method described in any one of claims 1-8.

10. A training device for a customer service proactive dialogue model, characterized in that, include: The training data acquisition module is used to acquire training data, which includes general dialogue data and real dialogue data. Both the general dialogue data and the real dialogue data include customer service dialogue data and customer dialogue data. The feature extraction module is used to extract features from the training data to obtain prompt information; The combining module is used to combine the training data and the prompt information into mixed data; The training module is used to train the large language model using the mixed data to obtain a customer service proactive dialogue model.

11. The apparatus according to claim 10, characterized in that, The feature extraction module includes: The template preset unit is used to obtain preset prompt templates; The extraction unit is used to extract prompt information according to the preset prompt template, taking each example of general dialogue data or real dialogue data in the training data as a unit. The prompt information includes customer service domain knowledge and / or dialogue skills used to guide the training of the large language model.

12. The apparatus according to claim 10 or 11, characterized in that, The training module includes: The prediction unit is used to predict the customer service dialogue data corresponding to the mixed data using the large language model, and output the predicted customer service dialogue data. A loss calculation unit is used to calculate the loss between the predicted customer service dialogue data and the customer service dialogue data in the mixed data; An adjustment unit is used to adjust the parameters of the large language model based on the loss. Return to the training unit, which is used to return to the execution of using the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data until the large language model meets the preset conditions, thus obtaining the customer service proactive dialogue model.

13. The apparatus according to claim 12, characterized in that, The customer service proactive dialogue model training device also includes: The multi-dimensional evaluation module is used to evaluate the customer service proactive dialogue model from multiple preset dimensions and obtain multiple evaluation results. The comprehensive evaluation module is used to obtain a comprehensive evaluation result based on the multiple evaluation results; The system returns to the training module, which, in response to the fact that the comprehensive evaluation result does not meet the preset evaluation conditions, returns to the step of using the large language model to predict the customer service dialogue data corresponding to the mixed data and outputs the predicted customer service dialogue data.

14. The apparatus according to claim 13, characterized in that, The comprehensive evaluation module includes: A preference calculation unit is used to map each evaluation result into a preference reward value according to a preset reward mapping rule; The comprehensive calculation unit is used to obtain a comprehensive evaluation result based on multiple preference reward values.

15. The apparatus according to claim 10, characterized in that, The connecting module is further used for: The prompt information is incorporated into the training data as mixed data.

16. A customer service proactive dialogue device, characterized in that, include: Indication information acquisition module, used to acquire indication information; The dialogue module is used to input the instruction information and real-time customer dialogue data into the trained customer service proactive dialogue model and output real-time customer service dialogue data. The customer service proactive dialogue model is trained using the method described in any one of claims 1-8.

17. An electronic device, characterized in that, include: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1-8 and / or 9.

18. A computer-readable storage medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any of claims 1-8 and / or 9.

19. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, implements the method as described in any one of claims 1-8 and / or 9.