Customer service active dialogue and customer service active dialogue model training method, device, equipment, medium and product

By combining customer service conversation data and specific prompt information to train a large language model, the problem that customer service robots cannot communicate actively is solved, and intelligent and personalized customer service in the customer service field is achieved.

CN120336458APending Publication Date: 2025-07-18BAIRONG ZHIXIN (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337519.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing customer service robots cannot actively communicate with customers in the customer service field, and lack targeting, resulting in a decrease in customer participation, especially in marketing scenarios where the amount of data is small and cannot be adapted to complex scenarios.

Method used

By obtaining common and real customer service conversation data, combining feature extraction and specific prompt information, training large language models, generating customer service models with active conversation capabilities, and being able to actively ask for requirements when customers do not ask questions.

Benefits of technology

It realizes the proactive dialogue ability in the customer service field, improves the intelligence and pertinence of customer interaction, adapts to different scenarios, and improves marketing effectiveness and customer satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336458A_ABST
    Figure CN120336458A_ABST
Patent Text Reader

Abstract

The invention provides a customer service active dialogue and customer service active dialogue model training method and device, equipment, a medium and a product, and the customer service active dialogue model training method comprises the steps: obtaining training data which comprises general dialogue data and real dialogue data, the general dialogue data and the real dialogue data comprise customer service dialogue data and customer dialogue data; performing feature extraction on the training data to obtain prompt information; combining the training data and the prompt information into mixed data; and training a large language model by using the mixed data to obtain a customer service active dialogue model. The model obtained through training can obtain the output closer to the customer service field based on the prompt information, can actively guide the customer to express the requirements of the field more clearly like the customer service, accurately grasp the requirements through multiple rounds of dialogues, and generate more reliable and accurate dialogues.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of artificial intelligence technology, and particularly to a method, apparatus, device, medium, and product for proactive customer service conversations and their model training. Background Art

[0002] Currently, robots in the customer service field on the market can only passively respond to customer inquiries and cannot actively communicate with customers like human customer service, resulting in a decrease in customer engagement.

[0003] In the field of proactive conversations, most are open-topic conversation robots based on large models. Although they can communicate proactively and perform well in understanding and generating natural language, these systems lack specific conversation topics and do not have the ability to communicate proactively with customers in a targeted manner. Summary of the Invention

[0004] Embodiments of the present disclosure propose a method, apparatus, device, medium, and product for proactive customer service conversations and their model training.

[0005] In a first aspect, embodiments of the present disclosure provide a method for training a customer service proactive conversation model, the method comprising:

[0006] Obtaining training data, where the training data includes general conversation data and real conversation data, and both the general conversation data and real conversation data include customer service conversation data and customer conversation data;

[0007] Performing feature extraction on the training data to obtain prompt information;

[0008] Combining the training data and the prompt information into mixed data;

[0009] Using the mixed data to train a large language model to obtain a customer service proactive conversation model.

[0010] As a possible implementation, the performing feature extraction on the training data to obtain prompt information includes:

[0011] Obtaining a preset prompt template;

[0012] Taking each example of general conversation data or real conversation data in the training data as a unit, and extracting prompt information according to the preset prompt template.

[0013] As a possible implementation, the using the mixed data to train a large language model to obtain a customer service proactive conversation model includes:

[0014] Using the large language model to predict the customer service conversation data corresponding to the mixed data and outputting predicted customer service conversation data;

[0015] Calculate the loss between the predicted customer service conversation data and the customer service conversation data in the mixed data;

[0016] Based on the loss, adjust the parameters of the large language model;

[0017] Return to execute the prediction of the customer service conversation data corresponding to the mixed data using the large language model, output the predicted customer service conversation data, and until the large language model meets the preset conditions, obtain the customer service proactive conversation model.

[0018] As a possible implementation manner, after obtaining the customer service proactive conversation model, it further includes:

[0019] Evaluate the customer service proactive conversation model from multiple preset dimensions to obtain multiple evaluation results;

[0020] According to the multiple evaluation results, obtain a comprehensive evaluation result;

[0021] In response to the comprehensive evaluation result not meeting the preset evaluation conditions, return to execute the prediction of the customer service conversation data corresponding to the mixed data using the large language model, and output the predicted customer service conversation data.

[0022] As a possible implementation manner, the preset conditions include any combination of the following situations:

[0023] Return to execute the prediction of the customer service conversation data corresponding to the mixed data using the large language model, and output the predicted customer service conversation data reaching the preset number of times;

[0024] The loss converges or is within the preset loss range;

[0025] The perplexity of the customer service proactive conversation model is within the preset perplexity range;

[0026] The large-scale multi-task language understanding ability of the customer service proactive conversation model is within the preset understanding ability range.

[0027] As a possible implementation manner, the obtaining the comprehensive evaluation result according to the multiple evaluation results includes:

[0028] According to the preset reward mapping rule, map each evaluation result to a preference reward value;

[0029] According to the multiple preference reward values, obtain the comprehensive evaluation result.

[0030] As a possible implementation manner, the mapping each evaluation result to a preference reward value according to the preset reward mapping rule includes:

[0031] Map each of the evaluation results to a preference reward value using formulas (1)-(2), where

[0032] Formula (1) is

[0033] Formula (2) is

[0034] where r i is the evaluation result, sign(r i ) is the indicator function, f(r i ) is the preference reward value of the evaluation result r i , tanh is the hyperbolic tangent function, k is a parameter preset to control the sensitivity of f(r i ), and x0 is a parameter preset to control the center point of the f(r i ) curve; and

[0035] Obtain a comprehensive evaluation result according to the multiple preference reward values, including:

[0036] Obtain the comprehensive evaluation result using formula (3), where

[0037] Formula (3) is

[0038] where r mo is the comprehensive evaluation result, w k is the preset weight corresponding to the evaluation result r i , n is the number of the evaluation results, and f(r i ) is the preference reward value of the evaluation result r i .

[0039] In a second aspect, an embodiment of the present disclosure provides a customer service proactive dialogue method, and the method includes:

[0040] Obtain indication information;

[0041] Input the indication information and the customer real-time dialogue data into the trained customer service proactive dialogue model, and output the customer service real-time dialogue data, where the customer service proactive dialogue model is trained by using the method described in any one of the first aspects.

[0042] In a third aspect, an embodiment of the present disclosure provides a customer service proactive dialogue model training device, and the device includes:

[0043] A training data acquisition module, configured to acquire training data, where the training data includes general dialogue data and real dialogue data, and both the general dialogue data and the real dialogue data include customer service dialogue data and customer dialogue data;

[0044] A feature extraction module, configured to extract features from the training data to obtain prompt information;

[0045] A combination module, configured to combine the training data and the prompt information into hybrid data;

[0046] A training module, configured to use the hybrid data to train a large language model to obtain a customer service proactive dialogue model.

[0047] In a fourth aspect, an embodiment of the present disclosure provides a customer service proactive dialogue device, which includes:

[0048] An indication information acquisition module, configured to acquire indication information;

[0049] A dialogue module, configured to input the indication information and real-time customer dialogue data into the trained customer service proactive dialogue model, and output real-time customer service dialogue data, where the customer service proactive dialogue model is trained by using the method described in any one of the first aspect.

[0050] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, which includes:

[0051] One or more processors;

[0052] A storage device, on which one or more programs are stored,

[0053] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any one of the first aspect and / or the second aspect.

[0054] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, where the computer program, when executed by one or more processors, implements the method described in any one of the first aspect and / or the second aspect.

[0055] In a seventh aspect, an embodiment of the present disclosure provides a computer program product, including a computer program / instructions, where the computer program / instructions, when executed by a processor, implement the method described in any one of the first aspect and / or the second aspect.

[0056] To enable a large model-based dialogue system to be proactive in a customer service scenario, the embodiments of the present disclosure provide a customer service proactive dialogue and its model training method, device, equipment, medium and product. First, training data is obtained, where the training data includes general dialogue data and real dialogue data, and both the general dialogue data and real dialogue data include customer service dialogue data and customer dialogue data; then, feature extraction is performed on the training data to obtain prompt information; then, the training data and the prompt information are combined into mixed data; finally, the large language model is trained using the mixed data to obtain a customer service proactive dialogue model.

[0057] The model trained in this way can obtain an output closer to the customer service field based on the prompt information, and like a customer service representative, can actively guide customers to express their needs in this field more clearly, and accurately grasp the needs through multiple rounds of dialogue to generate more reliable and accurate conversations. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives and advantages of the present disclosure will become more apparent. The drawings are only for the purpose of showing the specific embodiments and are not considered to be a limitation of the present disclosure. In the drawings:

[0059] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied;

[0060] Figure 2 is a flowchart of the customer service proactive dialogue model training method according to an embodiment of the present disclosure;

[0061] Figure 3A is a flowchart of the customer service proactive dialogue model training method according to another embodiment of the present disclosure;

[0062] Figure 3B is a flowchart of step 302 according to an embodiment of the present disclosure;

[0063] Figure 3C is a flowchart of step 304 according to an embodiment of the present disclosure;

[0064] Figure 3D is a flowchart of step 3041 according to an embodiment of the present disclosure;

[0065] Figure 3E is a flowchart of step 3046 according to an embodiment of the present disclosure;

[0066] Figure 4 is an architecture diagram of the customer service proactive dialogue model training method according to an embodiment of the present disclosure;

[0067] Figure 5It is a preference reward value curve graph of an embodiment of the present disclosure;

[0068] Figure 6 It is a flowchart of a customer service initiative dialogue method of an embodiment of the present disclosure;

[0069] Figure 7 It is a structural diagram of a customer service initiative dialogue model training device of an embodiment of the present disclosure;

[0070] Figure 8 It is a structural diagram of a customer service initiative dialogue device of an embodiment of the present disclosure;

[0071] Figure 9 It is a schematic structural diagram of a computer system of an electronic device of an embodiment of the present disclosure. Detailed implementation manners

[0072] With the development of enterprises and the intensification of market competition, customer service has become an important means for enterprises to establish and maintain good relationships with customers. However, with the expansion of enterprise scale and the increase in customer flow, especially with the in-depth digital transformation, the demand for intelligent and personalized services from enterprises and consumers has increased significantly. Relying solely on human customer service is difficult to meet the needs of a large number of customers, with low efficiency and high costs. To solve these problems, automation technologies have been introduced, and intelligent customer service dialogue robots have emerged as the times require.

[0073] Early intelligent customer service dialogue robots were based on preset rules and knowledge bases, and were more efficient in handling common and standard questions. However, their understanding ability was limited, and they could only understand preset questions and keywords, and it was difficult to handle non-standard expressions, such as using dialects, industry terms, or vague statements. If the customer's question did not hit the correct keywords or phrases, the system might not be able to provide the correct answer.

[0074] With the development of machine learning technologies, pattern-matching-based dialogue robots have replaced rule-based dialogue robots, with a certain degree of intelligence and the ability to handle complex dialogue scenarios. However, their responses are fixed, lacking flexibility and a natural and smooth interaction experience. The rise of deep learning technologies has brought dialogue robots into a new stage. Using hybrid AI technologies such as RNN and GAN, the dialogue is more natural and user-friendly. However, there are limitations in processing multi-turn dialogue context information, and the generalization ability is limited, and it is unable to provide personalized services.

[0075] With the development of artificial intelligence technologies, large model-based dialogue robots have become key tools in the customer service field. These large models are trained through ultra-large-scale datasets, can accurately parse and understand the natural language input of customers, and generate smooth and natural language, making the dialogue experience close to human-to-human communication.

[0076] However, most of the large model dialogue systems on the current market are open-topic and have no designated topic. The application of large models in the customer service field is still under exploration. Traditional customer service dialogue systems can only passively respond to customer inquiries and cannot actively communicate around products like human customer service, reducing customer participation in the designated field.

[0077] Although some positive results have been achieved in the related technologies regarding proactive dialogue in the social field. For example, generative language models are utilized as the main response generator, and a variety of innovative solutions are introduced, which can dynamically select prompting strategies according to customer intentions, personalities, and emotions, thus achieving personalized and proactive high-quality conversations. Some related technologies also explore proactive topic switching mechanisms, which can intelligently determine when to switch conversation topics and integrate external resources into the conversation. However, although these technologies only show a certain degree of proactiveness in social scenarios, mainly optimized for social scenarios and good at starting new topics and guiding the conversation to continue, in the customer service scenario, they cannot effectively guide customers to express corresponding specific needs.

[0078] In summary, the related technologies still have technical limitations. In the customer service field, they have always been stuck in the form of question and answer, only able to passively reply to customer inquiries and unable to actively communicate with customers like human customer service. Especially when customers do not ask questions, the related technologies do not have the function of actively asking customers about their needs in the corresponding field. In the case of very little data in the marketing customer service field, the existing technologies cannot be applied to complex scenarios in the customer service field. Especially when facing various marketing products and different customer conversation styles, it is difficult to make good adaptations; due to the particularity of marketing data, there are no dedicated optimization metrics to align with human preferences and no dedicated evaluation metrics to evaluate the model well, etc.

[0079] The following further elaborates on the present disclosure in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are merely for explaining the related invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0080] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The following will detail the present disclosure with reference to the drawings and in conjunction with the embodiments.

[0081] Figure 1 An exemplary system architecture 100 is shown, which can apply the embodiments of the customer service proactive dialogue and its model training method, apparatus, device, medium, and product of the present disclosure.

[0082] As Figure 1As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0083] Customers can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as dialogue applications, speech recognition applications, short video social applications, audio and video conferencing applications, video live streaming applications, document editing applications, input method applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0084] The terminal devices 101, 102, 103 can be hardware or software. When the terminal devices 101, 102, 103 are hardware, they can be various electronic devices with a display screen, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they can be installed in the above-listed terminal devices. It can be implemented as multiple software or software modules (e.g., used to provide dialogue services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0085] In some cases, the customer service proactive dialogue and its model training method provided by the present disclosure can be executed by the terminal devices 101, 102, 103. Correspondingly, the customer service proactive dialogue and its model training device can be set in the terminal devices 101, 102, 103. In this case, the system architecture 100 may also not include the server 105.

[0086] In some cases, the customer service proactive dialogue and its model training method provided by the present disclosure can be jointly executed by the terminal devices 101, 102, 103 and the server 105. For example, the step of "obtaining training data, where the training data includes general dialogue data and real dialogue data, and both the general dialogue data and the real dialogue data include customer service dialogue data and customer dialogue data" can be executed by the terminal devices 101, 102, 103, and steps such as "using the mixed data to train a large language model to obtain a customer service proactive dialogue model" can be executed by the server 105. The present disclosure does not limit this. Correspondingly, the customer service proactive dialogue model training device can also be respectively arranged in the terminal devices 101, 102, 103 and the server 105.

[0087] In some cases, the customer service proactive dialogue and its model training method provided by the present disclosure can be executed by the server 105. Correspondingly, the customer service proactive dialogue model training device can also be arranged in the server 105. In this case, the system architecture 100 may not include the terminal devices 101, 102, 103 either.

[0088] It should be noted that the server 105 can be hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server 105 is software, it can be implemented as multiple software or software modules (such as for providing distributed services) or as a single software or software module. No specific limitation is made here.

[0089] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in

[0090] The information, data, and signals involved in the present disclosure are all authorized by the user / customer or fully authorized by all parties, and the collection, use, and processing of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0091] Continuing to refer to Figure 2 which shows a flow 200 of an embodiment of the customer service proactive dialogue model training method according to the present disclosure. This method includes at least the following steps 201 to 204.

[0092] Step 201, obtain training data.

[0093] The training data includes general dialogue data and real dialogue data, and both the general dialogue data and the real dialogue data include customer service dialogue data and customer dialogue data.

[0094] Among them, the general dialogue data, including the data of the conversation initiated by the customer service in response to the conversation initiated by the customer, can be the conversation data in the form of question and answer in response to the conversation initiated by the customer, which is used to prevent the model from losing general knowledge and expression ability. For example, a piece of general dialogue data is as follows:

[0095] Customer: How about xxx tiles?

[0096] Customer service: xxx is a well-known domestic...

[0097] Customer: How to achieve...?

[0098] Customer service: To achieve a...

[0099] The real conversation data, including the conversation data initiated by the customer service actively. It is the conversation data between the customer and the customer service from the real scenario, which can not only include the form of question and answer in response to the conversation initiated by the customer, but also include the conversation data initiated by the customer service actively. For example, a piece of real conversation data is as follows:

[0100] Customer service: Hello, is that Ms. Zhang?

[0101] Customer: Hello, who are you?

[0102] Customer service: I am the customer service of xx bank...

[0103] Customer service: Do you understand my explanation?

[0104] Optionally, after obtaining the training data, it can be selectively cleaned, for example, filtering out or correcting illogical and duplicate conversation content. Also, since the conversation is unstable with too few conversation turns and there is too little effective information with too many conversation turns, therefore, further removing the data whose conversation turns meet the preset removal conditions, and the preset removal conditions can be that the conversation turns are less than a rounds or more than b rounds (where a < b, and both a and b are natural numbers, which can be set according to the training data and are not limited here), so as to obtain high-quality real conversation data. Taking a = 3 and b = 20 as an example, after the above processing, 755 pieces of real conversation data can be obtained.

[0105] Optionally, the training data can include a preset proportion of general dialogue data and real conversation data. Preferably, the preset proportion can be 1:1. In this way, when the real conversation data is 755 pieces, the mixed training data is more than 1500 pieces. Using the training data including the mixture of general dialogue data and real conversation data to train the model can enhance the generalization ability of the model.

[0106] Step 202, perform feature extraction on the training data to obtain prompt information.

[0107] The prompt information can selectively include one or more of role information, customer information, product information, speech, FAQ (Frequently Asked Questions), and conversation purpose, which is used to guide model training, reduce irrelevant or inaccurate content generated by the model, and better control the generation direction of the model. In the case of very small amount of data in the customer service field, especially when facing a variety of marketing products and different user conversation styles, it can be well adapted.

[0108] Among them, the role information is the role of the customer service, such as product manager, sales champion, and consulting consultant.

[0109] Customer information refers to the customer's information, which may include their basic identity information, business-related information, etc.

[0110] Product information is information about the target product, including but not limited to product name, function introduction, components, product description, etc. In some examples of marketing scenarios, customer service needs to sell the product corresponding to the product information. It should be understood that the present disclosure can be applied to marketing scenarios in different fields, including but not limited to e-commerce, catering, clothing, finance, insurance, etc. In some other embodiments, the present disclosure can also be applied to scenarios applicable to customer service other than marketing scenarios, such as product demonstrations, after-sales service of products or services, etc.

[0111] The speech is a dialogue guide or guidance information, which is used to guide or guide the reply, inquiry and / or dialogue direction during the dialogue process. In some examples, the speech can be pre-set information. In some embodiments of marketing scenarios, the speech is a dialogue guide or guidance information based on the marketing scenario. It should be understood that the present disclosure can be applied to marketing scenarios in different fields, including but not limited to e-commerce, catering, clothing, finance, insurance, etc. In some other embodiments, the present disclosure can also be applied to scenarios applicable to customer service other than marketing scenarios, such as product demonstrations, after-sales of products or services, etc.

[0112] FAQ is the frequently asked questions and answers of customer service and clients.

[0113] The purpose of the conversation is the expected effect of the conversation. For example, in some marketing scenarios, the purpose of the conversation is mainly to facilitate transactions, and in some consulting scenarios, the purpose of the conversation is mainly to understand the customer's information.

[0114] Optionally, a template corresponding to the prompt information can be set in advance according to the scenario, and then corresponding feature extraction can be performed according to the template. In some embodiments, the content of the prompt information can be adjusted accordingly according to the scenario and needs, which is not limited here. Through 202, the prompt information of the training data can be obtained, and the prompt information can be specific and detailed. Specific and detailed prompt information can greatly improve the reliability of the output of the customer service active dialogue model, making the output result of the model more accurate and credible.

[0115] Step 203, combine the training data and the prompt information into mixed data.

[0116] Combine the training data and the prompt information into mixed data, and then train the large language model. In different customer service scenarios, only by adjusting the content of the prompt information can a good generalization effect be obtained at the lowest cost.

[0117] Among them, the training data is used to overcome the limitation that the traditional model can only respond passively, enabling the model to break out of the form of one question and one answer. It can not only passively answer the customer's questions, but also actively communicate with the customer like a human customer service. The prompt information is used to assist in training, provide guidance for model training, help better control the generation direction of the model, and make it better adapt to different customers and customer service scenarios.

[0118] Step 204, use the mixed data to train the large language model to obtain a customer service active dialogue model.

[0119] The large language model (Large Language Model—LLM) can be a language model pre-trained with a large amount of data. The mixed data introduces specific prompt information into its training process, amplifies the guiding role of the prompt information, and guides and optimizes the model's learning of customer service domain knowledge and dialogue skills to generate a customer service model with active dialogue capabilities.

[0120] In addition, the training data of this model is based on real dialogue data. Through deep learning and optimization, the large language model can actively communicate with customers. Even when the customer does not ask questions, the large language model also has the function of actively asking about the customer's needs, realizing a more intelligent and active customer interaction mode, and providing an innovative customer service solution for various industries.

[0121] In summary, in this embodiment, the training of the customer service active dialogue model adds prompt information as a guide, which fully improves the pertinence of the model output, makes its output consider more the content of the prompt information, and improves the marketing effect. In addition, the trained customer service active dialogue model can actively communicate with customers and guide customers to clearly express their needs.

[0122] Continue to combine Figure 3A andFigure 4 , which shows the process 300 of another embodiment of the customer service proactive dialogue model training method according to the present disclosure, which at least includes the following steps 301 to step 305.

[0123] Step 301, obtain training data.

[0124] As a possible implementation manner, after obtaining the training data, the training data can be processed into the input format required by the model.

[0125] For example, convert a certain training data "Customer service: 'Hello, is this Ms. Zhang?' Customer: 'Hello, who are you?' Customer service: 'I am the customer service of xx Bank...'... Customer service: 'Do you understand my explanation?' " into the following messages format:

[0126] [{"role":"assistant","content":"Hello, is this Ms. Zhang?"}

[0127] {"role":"user","content":"Hello, who are you?"},

[0128] {"role":"assistant","content":"I am the customer service of xx Bank..."}

[0129] ……

[0130] {"role":"assistant","content":"Do you understand my explanation?"}]

[0131] Step 302, perform feature extraction on the training data to obtain prompt information.

[0132] As a possible implementation manner, step 302 may include the following steps 3021 to step 3022 as Figure 3B shown. Through this implementation manner, it is possible to set the information associated with the dialogue occurrence scenario through a preset prompt template, so that the prompt template can guide the direction of model training, thereby making the output of the model more targeted.

[0133] Step 3021: Obtain a preset prompt template.

[0134] The preset prompt template can be a template preset by using Prompt Engineering. The content of the prompt template can be set according to the customer service and / or service scenario. For example, in the financial scenario, it can include sales-related information, and in the medical scenario, it can include disease-related information, etc. Exemplarily, the prompt template in the financial scenario can be in the following form:

[0135] Character information Customer information Product information Script (such as sales script) FAQ Purpose of conversation

[0136] Step 3022: Taking each piece of general dialogue data or real dialogue data in the training data as a unit, extract prompt information according to a preset prompt template.

[0137] Most existing large models can complete various complex text generation and understanding tasks. Therefore, any large model can be directly used to extract the prompt information in the training data according to the preset prompt template. In the embodiments of the present disclosure, a corresponding prompt information can be extracted from each piece of general dialogue data or real dialogue data.

[0138] Optionally, the large model for extracting prompt information can be a generative pre-trained language model, such as GPT (Generative Pre-trained Transformer).

[0139] For example, taking a certain piece of real dialogue data as a unit, according to the template in Step 3021, the following prompt information can be extracted:

[0140]

[0141]

[0142] For another example, taking a certain piece of real dialogue data as a unit, according to the template in Step 3021, the following prompt information can be extracted:

[0143]

[0144]

[0145] Step 303: Combine the training data and the prompt information into hybrid data.

[0146] When combining the training data and the prompt information into hybrid data, the prompt information needs to be combined before the training data (for example, the position of the prompt information is before the training data) so as to conform to the rule that the large language model predicts the following text based on the previous text and realize the guiding effect of the prompt information on the output.

[0147] As a possible implementation manner, the prompt information can be first converted into the messages format and then concatenated before the training data in its corresponding messages format.

[0148] For example, the training data is information in the messages format, specifically as follows:

[0149] [{"role":"assistant","content":"Hello, are you Ms. Zhang?"}

[0150] {"role":"user","content":"Hello, who are you?"},

[0151] {"role":"assistant","content":"I am the customer service of xx Bank..."}

[0152] ...

[0153] {"role":"assistant","content":"Do you understand my explanation?"}]

[0154] After adding prompt information in front of the training data, the mixed data combined with the training data and the prompt information is the following information in the messages format:

[0155] [{"role":"system","content":"Prompt information"},

[0156] [{"role":"assistant","content":"Hello, is this Ms. Zhang?"}

[0157] {"role":"user","content":"Hello, who are you?"},

[0158] {"role":"assistant","content":"I am the customer service of xx Bank..."}

[0159] ...

[0160] {"role":"assistant","content":"Do you understand my explanation?"}]。

[0161] Step 304: Use the mixed data to train the large language model to obtain a customer service proactive dialogue model.

[0162] As a possible implementation method, the training process can be divided into two stages: supervised fine-tuning and alignment reinforcement learning.

[0163] Supervised fine-tuning can at least include the following steps 3041 to 3047 as Figure 3C shown, which are used to train the model so that it can learn to have a dialogue as a customer service based on the prompt information and achieve proactive dialogue with customers.

[0164] Step 3041: Use the large language model to predict the customer service dialogue in the mixed data and output the predicted customer service dialogue data.

[0165] As a possible implementation, step 3041 may at least include the following steps 30411 to 30413 as shown in Figure 3D below.

[0166] Step 30411, convert the input mixed data from the messages format to the string format.

[0167] The string format is a pattern that can be recognized by the large language model. The string format is as follows:

[0168] <|im_start|>system

[0169] **

Role Information

[0170] **Role:** Financial product sales champion

[0171] **Goal:** Recommend suitable fund products, promote customer purchases, and enhance investment confidence.

[0172] …

[0173] <|im_end|>

[0174] <|im_start|>user

[0175] Hello, what's up<|im_end|>

[0176] <|im_start|>assistant

[0177] Hello, I'm your sales manager<|im_end|>

[0178] Among them, <|im_end|>, <|im_start|> are special tokens used in the pre-training of the large language model to distinguish different speaking roles.

[0179] Step 30412, convert the input mixed data from the string format to the token format.

[0180] A token represents the basic unit processed and generated by the model, which can be a word, sub-word, or character. Large language models typically use complex tokenization to handle the vast diversity of human languages while keeping the vocabulary size manageable.

[0181] When converting the mixed data from the string format to the token format, the mixed data can be tokenized first and then converted according to the vocabulary.

[0182] For example, the vocabulary of the large language model Qwen is relatively large (the vocabulary size, i.e., the total number of unique tokens recognized by the model, has a significant impact on the performance and versatility of the model), with 151,646 tokens. Qwen adopts a subword tokenization method called Byte Pair Encoding (BPE), which attempts to learn token combinations that can represent text with the fewest tokens.

[0183] The present disclosure can convert the mixed data based on this vocabulary. For example, the string "tokenization" can be decomposed into two tokens, "token" and "ization", using any tokenization method. Then, by corresponding to find the encodings of the two tokens in the Qwen vocabulary respectively, we get "123" and "4215". Finally, the tokens of each example of mixed data are combined into a token format in the form of [123, 4251, 856,..].

[0184] Step 30413: Use the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data.

[0185] The mixed data combines the prompt information and the training data in sequence, and the prediction of the large language model will predict the following text based on the previous text. Therefore, the predicted customer service dialogue data output in step 30413 is predicted based on the prompt information and the dialogue before the predicted customer service dialogue data output this time, and can be guided by the prompt information for prediction. This prediction method is the training of a causal model of language, and the causal model is also called an auto regressive language model or a decoder-only language model.

[0186] For example, the mixed data is

[0187] "

Prompt information

[0188] Customer service: Hello, is this Ms. Zhang?

[0189] Customer: Hello, who are you?

[0190] Customer service: I am the customer service of xx Bank...

[0191] Customer service: Do you understand my explanation like this?"

[0192] Then, when predicting "Hello, is this Ms. Zhang?", it is predicted based on "

Prompt information

Prompt information

[0193] In the sentence "Hello, may I ask if you are Ms. Zhang?", the word "you" is predicted based on the "

Prompt Information

Prompt Information

[0194] Specifically, the mixed data in token format is input into the large language model, and the large language model only makes predictions for the customer service dialogue data in token format. For each token in the customer service dialogue data, the model generates a corresponding prediction label, and the prediction label is the predicted next token. Each time a prediction label is generated, it is added to the generated sequence. Finally, the generated sequence forms the customer service dialogue data in token format.

[0195] It should be noted that the prediction of the first token in the customer service dialogue data is completed based on the prompt information in token format and the previous dialogue data. For the prediction of subsequent tokens, the prompt information in token format, the previous dialogue data, and the tokens in the generated sequence corresponding to the customer service dialogue data generated by this prediction are comprehensively considered.

[0196] In terms of "causality", it means that the model considers the past context (i.e., the generated tokens) when predicting the next token, rather than considering future tokens.

[0197] For example, the predicted dialogue is "The weather is very good today.", and its prediction process is as follows:

[0198] Generate the prediction label "tian" (day) based on the token "jin" (today).

[0199] Generate the prediction label "tian" (day) based on the tokens "jin" (today) and "tian" (day).

[0200] Generate the prediction label "qi" (weather) based on the tokens "jin" (today), "tian" (day), and "tian" (day).

[0201] Generate the prediction label "hen" (very) based on the tokens "jin" (today), "tian" (day), "tian" (day), and "qi" (weather).

[0202] …

[0203] Generate the prediction label "EOS" based on the tokens "jin" (today), "tian" (day), "tian" (day), "qi" (weather), "hen" (very), and "hao" (good).

[0204] It should be noted that EOS is a special symbol indicating the end of a sentence.

[0205] Combine the above prediction labels to obtain the predicted dialogue data "The weather is very good today.".

[0206] In summary, steps 30411 to 30413 can use a large language model to predict the customer service dialogue data in the mixed data and output the predicted customer service dialogue data.

[0207] Step 3042: Calculate the loss between the predicted customer service dialogue data and the customer service dialogue data in the mixed data.

[0208] As a possible implementation, in order to improve the measurement effect of the model prediction accuracy, the Cross Entropy Loss function can be used as the loss function. The Cross Entropy Loss function can calculate the difference between the probability distribution predicted by the model and the true label, thereby guiding the optimization direction of the model.

[0209] Specifically, the loss between the tokens corresponding to the predicted customer service dialogue data and the tokens corresponding to the actual customer service dialogue data can be calculated. What is obtained in step 3041 are the tokens corresponding to the predicted customer service dialogue data, while the tokens corresponding to the customer service dialogue data in the training data are the tokens corresponding to the actual customer service dialogue data. By calculating the loss function of the two, the model can be trained to make its dialogue data closer to the actual customer service dialogue data.

[0210] Step 3043: Adjust the parameters of the large language model based on the loss.

[0211] When adjusting the parameters of the large language model, first, through the backpropagation algorithm, calculate the gradient of the loss with respect to the parameters of each layer layer by layer starting from the output layer, and then use an optimizer to update the weights and bias parameters of the parameters of each layer in the network according to the calculated gradient and the preset learning rate. As a possible implementation, the optimizer can use AdamW, and the learning rate is set to 5e-6, that is, 5×10 -6 。

[0212] Step 3044: Return to execute step 3041 until the large language model meets the preset conditions, and then obtain the customer service initiative dialogue model.

[0213] The preset conditions can comprehensively consider one or more of the following multiple dimensions to determine whether the model can stop training.

[0214] First, return to execute step 3041 for a preset number of times.

[0215] Since the effect of a single training may not reach a good result, the training data can be cycled for a preset number of times. The preset number of times can be set according to requirements and is not limited here.

[0216] Second, the loss converges or is within the preset loss range.

[0217] If the loss converges or is within the preset loss range, the requirement is met. Generally, the training can be stopped when the loss is around 0.1 - 0.5. In supervised fine-tuning, if the loss is too high (1 - 2 or even higher), the training is insufficient; if it is too low (close to 0), it is prone to overfitting, which affects the normal answering of the model.

[0218] Thirdly, the perplexity of the customer service proactive dialogue model is within the preset perplexity range.

[0219] The perplexity (Perplexity, PPL) of the model can measure the model's prediction ability for text. If the PPL is within the preset perplexity range, the requirement is met. When obtaining the predicted labels corresponding to each token, actually multiple possible predicted labels are obtained first, and each predicted label has a corresponding possibility (which can be a probability), and then the one with the highest possibility among the multiple possible predicted labels is selected.

[0220] Therefore, the predicted label actually corresponds to a possibility. For example, the possibility estimates for each predicted label in the sentence "The weather is good today" are 0.5, 0.25, 0.125, and 0.125 respectively. Then, substituting the above possibilities into the preset PPL calculation formula, the PPL of the customer service proactive dialogue model can be obtained, and thus the "average guess" uncertainty degree of the customer service proactive dialogue model can be reflected by its PPL.

[0221] For a language model, if the PPL value is low, it means that the model's prediction of this sentence is relatively accurate, and the generated text is more fluent and natural; conversely, if the PPL value is high, it means that the model's prediction of this sentence is not accurate enough.

[0222] Therefore, when the PPL of the model is within the preset perplexity range, it represents that the model's prediction reaches a certain accuracy.

[0223] Fourthly, the large-scale multi-task language understanding ability of the customer service proactive dialogue model is within the preset understanding ability range.

[0224] The large-scale multi-task language understanding ability can be measured by the MMLU (Massive Multitask Language Understanding) test score. MMLU is a benchmark test used to evaluate the understanding and reasoning abilities of large language models in multiple academic fields, to measure the model's logical reasoning ability and ensure that the model does not overfit due to too single a measurement index during training. It measures the breadth and depth of the model's knowledge, as well as the model's logical reasoning ability in different academic disciplines through a series of carefully designed multiple-choice questions. If the MMLU test score is within the preset benchmark test score range, then the large-scale multi-task language understanding ability of the customer service proactive dialogue model is within the preset understanding ability range.

[0225] Before training (i.e., supervised fine-tuning), the large language model can be tested with MMLU to obtain a score. After the above training (i.e., supervised fine-tuning), the performance of the model in terms of output is improved. At the same time, it is necessary to ensure that the MMLU test score of the trained model is basically the same as that before supervised fine-tuning. Generally, it is reasonable to control the fluctuation range of the MMLU test score within ±3 percentage points.

[0226] For example, if the MMLU score of the model before fine-tuning is 84.2 and after fine-tuning is 82.1, this small range of change is within the acceptable range, indicating that the model has not overfitted due to fine-tuning.

[0227] The above preset number of times, preset loss range, preset perplexity range, and preset benchmark test score range can all be set as needed and are not limited here.

[0228] In summary, the supervised fine-tuning of the customer service proactive dialogue model is achieved through steps 3041 to 3044.

[0229] As a possible implementation, after supervised fine-tuning, the customer service proactive dialogue model can be selectively subjected to alignment reinforcement learning. Alignment reinforcement learning can include the following steps 3045 to 3047. Alignment reinforcement learning introduces multiple preset dimensions and uses the target effects of multiple preset dimensions as the model optimization direction (i.e., the optimization directions of multiple targets, and multiple targets can also be referred to as the target effects or evaluation results of each preset dimension), so that the trained customer service proactive dialogue model will be more in line with the artificially set dialogue habits and can better meet the dialogue needs. This not only helps to optimize the customer dialogue experience but also contributes to achieving the established dialogue goals of the customer service proactive dialogue model.

[0230] Step 3045: Evaluate the customer service proactive dialogue model from multiple preset dimensions to obtain multiple evaluation results.

[0231] The multiple preset dimensions can include the number of dialogue turns, customer emotional positivity, dialogue purpose achievement rate (such as marketing success rate, consultation completion rate), scoring model, etc. The evaluation results are used to represent the capabilities of the customer service proactive dialogue model in different preset dimensions, and the evaluation results can be quantitative scores.

[0232] Among them, the more the number of dialogue turns, the more it means that the customer is more willing to participate in the current dialogue, and the higher the quality of the model output, the higher the score for this dimension.

[0233] For customer emotional positivity, the tendency of the customer's emotions is judged through sentiment analysis. The output dialogue data of the model should try to guide the customer's emotions in a positive direction and avoid the customer's emotions being low, which may affect the dialogue experience. The more positive the user's emotions, the higher the score for customer emotional positivity and the better the quality of the model output.

[0234] The dialogue goal achievement rate is used to measure whether the dialogue requirements are met, such as whether the marketing is successful or the consultation is completed. It is the ultimate goal of the model. The higher the dialogue goal achievement rate, the more the dialogue requirements can be met, the more the dialogue goal can be achieved, and the better the quality of the model output.

[0235] The scoring model (Reward model) is a model trained based on manual scoring and dialogue data, which can simulate manual automation to score dialogue data. Inputting the prompt information, real customer dialogue data, and predicted customer service dialogue data into the model can output a score. The higher the output score, the better the quality of the model output. Since the scoring model is used to simulate manual scoring, the scoring of the above dimensions can also be replaced by direct manual scoring, which is not restricted here.

[0236] It should be noted that in order to unify the scoring range of the scoring model, the output of the scoring model can be connected to the sigmoid function to convert the score into a probability value between 0 and 1, and then multiply the probability value by a preset coefficient to obtain a score within the preset scoring range. For example, if the preset coefficient is set to 2, the score will be a floating-point score value between 0 and 2. When using the scoring model to score, only the prompt information, real customer dialogue data, and predicted customer service dialogue data need to be input to output a floating-point score value between 0 and 2.

[0237] The evaluation results of multiple dimensions are used to characterize the capabilities of the customer service proactive dialogue model in different dimensions, thereby controlling the training effect of the model, enabling the customer service proactive dialogue model to balance multiple key factors when outputting dialogue data, and simultaneously increasing the possibility of successful conversion of the ultimate dialogue goal.

[0238] Step 3046: Obtain a comprehensive evaluation result based on multiple evaluation results.

[0239] By introducing the evaluation results of multiple preset dimensions, the customer service proactive dialogue model can balance multiple key factors when outputting dialogue, ensuring the successful conversion of the ultimate dialogue goal (such as the marketing goal). However, if different preset dimensions are desired to affect the output to varying degrees, when obtaining the comprehensive evaluation result, the multiple evaluation results can be weighted and summed according to the preset weights, so that the comprehensive evaluation result can take into account multiple goals (i.e., the evaluation results of multiple dimensions). In this way, the model obtained by subsequent training based on the comprehensive evaluation result can achieve better performance in multiple preset dimensions.

[0240] As a possible implementation method, to calculate the comprehensive evaluation result, the method of linear weighted sum can be directly adopted.

[0241] Specifically, the weight coefficients corresponding to each evaluation result (i.e., the number of dialogue turns, emotional positivity, marketing success rate, and scoring model) are represented as w1, w2, w3, and w4. Adjusting the weights can ensure that the importance of different evaluation results is reasonably reflected in the comprehensive evaluation result. In this way, the model can achieve a balance among dialogue interaction, emotional tendency, marketing goals, and human preferences during the optimization process, ultimately improving customer satisfaction, enhancing interaction efficiency, and increasing the achievement rate of dialogue purposes.

[0242] For example, the weight coefficient of the number of dialogue turns (w1) is 0.4, the weight coefficient of customer emotional positivity (w2) is 0.3, the weight coefficient of marketing success rate (w3) is 0.3, and the weight coefficient of the scoring model weight (w4) is 0.5. In this example, we assume each evaluation result. The number of dialogue turns is 4 rounds, indicating that the model generated 4 responses and scored 4; the emotional positivity is 8 points, meaning a relatively high score through sentiment analysis and the customer's mood tending to be positive; the marketing success rate is 1, indicating a 100% marketing conversion rate; the Reward scoring model outputs 1.1 points, indicating that the scoring model believes the output of the customer service initiative dialogue model can obtain a comprehensive score of 1.1.

[0243] The comprehensive evaluation result is equal to the sum of w1 multiplied by the evaluation result of the number of dialogue turns, w2 multiplied by the evaluation result of emotional positivity, w3 multiplied by the evaluation result of marketing success rate, and w4 multiplied by the evaluation result of the scoring model.

[0244] Substituting the above values for weighted summation calculation, the comprehensive evaluation result is:

[0245] 0.4×4 + 0.3×8 + 0.3×1 + 0.5×1.1 = 1.6 + 2.4 + 0.3 + 0.22 = 4.52.

[0246] The comprehensive evaluation result obtained in this way can comprehensively evaluate the number of dialogue turns, emotional positivity, marketing success rate, and scoring evaluation.

[0247] As another possible implementation, the traditional alignment algorithm can be improved by changing the single-objective preference value reward in the traditional alignment algorithm to a multi-objective preference value reward to obtain the comprehensive evaluation result. Specifically, step 3046 can include steps 30461 and 30462 as shown in Figure 3E Figure 30461 and Figure 30462.

[0248] Step 30461, according to the preset reward mapping rule, map each evaluation result to a preference reward value.

[0249] When calculating the comprehensive evaluation result, in order to improve the sensitivity to the evaluation result and identify more effective information, that is, to make the comprehensive evaluation result more sensitive to the lower or higher evaluation results, a preset reward mapping rule can be set to map each evaluation result to a preference reward value. The preset reward mapping rule can specify a linear or non-linear mapping relationship to map the evaluation result to the preference reward value, and the specific mapping rule can be set according to requirements.

[0250] For example, use formulas (1)-(2) to map each of the said evaluation results to a preference reward value, where

[0251] Formula (1) is

[0252] Formula (2) is

[0253] where r i is the evaluation result, sign(r i ) is the indicator function, f(r i ) is the preference reward value of the evaluation result r i , tanh is the hyperbolic tangent function, k is a preset parameter for controlling the sensitivity of f(r i ), and x0 is a preset parameter for controlling the center point of the curve of f(r i ).

[0254] It should be noted that k and x0 can be preset according to needs and are not limited here. See Figure 5 , k controls the steepness (i.e., sensitivity) of the f(r i ) curve. The larger k is, the higher the sensitivity of the f(r i ) curve; x0 controls the center point of the curve and determines the value near which the curve changes fastest. Optionally, k = 5 and x0 = 0.5 can be set.

[0255] Step 30462, obtain the comprehensive evaluation result according to multiple preference reward values.

[0256] After obtaining the preference reward values corresponding to different evaluation results, then use formula (3) to perform weighted summation to obtain the comprehensive evaluation result, where

[0257] Formula (3) is

[0258] where r mo is the comprehensive evaluation result, w k is the preset weight corresponding to the evaluation result r i , n is the number of the said evaluation results, and f(r i ) is the preference reward value of the evaluation result r i .

[0259] The comprehensive evaluation result calculated according to steps 30461 to 30462 is more sensitive to the over-high or over-low evaluation results, thereby controlling the training in step 3047 and improving the convergence speed and training efficiency of the training.

[0260] Step 3047, in response to the comprehensive evaluation result not meeting the preset evaluation condition, return to execute step 3041.

[0261] When the comprehensive evaluation result does not meet the preset evaluation condition, it will return to 3041 for retraining until the comprehensive evaluation result meets the preset evaluation condition, that is, the customer service initiative dialogue model has good capabilities in multiple preset dimensions. The preset evaluation condition can be set according to requirements and is not limited here.

[0262] In summary, in steps 3045 to 3047, the capabilities of the customer service initiative dialogue model are evaluated from multiple preset dimensions, and the model is controlled to perform further training until the model can achieve good performance in multiple preset dimensions, which will be more in line with the dialogue habits of multiple preset dimensions and can better meet the dialogue needs. This can not only improve the customer dialogue experience but also help achieve the dialogue purpose of the customer service initiative dialogue model.

[0263] In summary, the customer service initiative dialogue model obtained through the above training can achieve good performance in multiple evaluation systems such as loss, PPL, MMLU, number of dialogue turns, customer emotion positivity, marketing success rate, and scoring model, achieving at least the following three-dimensional beneficial effects:

[0264] (1) It can automatically act as a customer service to communicate actively with customers, breaking the limitations of traditional models and realizing a more intelligent and proactive customer interaction mode, providing an innovative customer service solution for various industries;

[0265] (2) Using prompt engineering, for specific customer service scenarios, key information is extracted to generate prompt information to assist training, enabling the model to adapt to different products and scenarios, and providing an effective strategy to improve the generalization performance of the model with a small amount of data;

[0266] (3) The multi-objective preference value reward method based on weight allocation makes the intelligent customer service robot more inclined to human dialogue habits, providing an important reference for the optimization goal design in complex dialogue scenarios.

[0267] Continue to refer to Figure 6 , which shows the process 600 of an embodiment of the customer service initiative dialogue method according to the present disclosure, which is equivalent to the inference process of the above-trained model, and the dialogue data obtained by inference can still be used as training data for iterative training of the model. The customer service initiative dialogue method includes the following steps 601-602.

[0268] Step 601, obtain indication information.

[0269] The prompt information is obtained by extracting the features of the training data. However, in the actual use of the model, there is no complete conversation data. Therefore, the indication information can be obtained from the system, including customer service information, specific customer information, specific information of the product to be recommended, specific excellent conversation scripts for this scenario, and the candidate set of question-and-answer pairs (FAQ) retrieved by the embedding model vector. Then, the indication information is used to direct the model output.

[0270] In subsequent conversations, the above indication information can guide the model to output more accurate customer service responses, realizing the automatic specification of conversation content and reply language style, and improving the scenario targeting and / or marketing success rate.

[0271] Step 602, input the indication information and the real-time customer conversation data into the trained customer service proactive conversation model to output the real-time customer service conversation data.

[0272] Since the conversation is continuous, the customer conversation data can be continuously received, and the output customer service conversation data can be recorded. The first customer service reply can be based on the indication information and / or the customer's first conversation data. After that, the customer service reply can be based on the indication information and all the previous conversation data.

[0273] In summary, when using the trained customer service proactive conversation model, output information strongly related to the indication information can be output, and proactive conversations can be output, improving the conversation targeting and thus the marketing success rate, and ensuring the customer experience.

[0274] Further reference Figure 7 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a customer service proactive conversation model training device. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0275] As Figure 7 shown, the customer service proactive conversation model training device 700 of this embodiment includes: a training data acquisition module 701, a feature extraction module 702, a combination module 703, and a training module 704.

[0276] Among them, the training data acquisition module 701 is used to acquire training data, and the training data includes general conversation data and real conversation data. Both the general conversation data and the real conversation data include customer service conversation data and customer conversation data;

[0277] The feature extraction module 702 is used to extract features from the training data to obtain prompt information;

[0278] Combining module 703, configured to combine the training data and the prompt information into hybrid data;

[0279] Training module 704, configured to use the hybrid data to train a large language model to obtain a customer service proactive dialogue model.

[0280] As a possible implementation, the feature extraction module 702 includes:

[0281] Template preset unit 7021 (not shown in the figure), configured to obtain a preset prompt template;

[0282] Extraction unit 7022 (not shown in the figure), configured to take each piece of general dialogue data or real dialogue data in the training data as a unit, and extract prompt information according to the preset prompt template.

[0283] As a possible implementation, the training module 704 includes:

[0284] Prediction unit 7041 (not shown in the figure), configured to use the large language model to predict the customer service dialogue data corresponding to the hybrid data and output predicted customer service dialogue data;

[0285] Loss calculation unit 7042 (not shown in the figure), configured to calculate the loss between the predicted customer service dialogue data and the customer service dialogue data in the hybrid data;

[0286] Adjustment unit 7043 (not shown in the figure), configured to adjust the parameters of the large language model based on the loss;

[0287] Return training unit 7044 (not shown in the figure), configured to return and execute using the large language model to predict the customer service dialogue data corresponding to the hybrid data and output predicted customer service dialogue data until the large language model meets the preset conditions, and then obtain a customer service proactive dialogue model.

[0288] As a possible implementation, the customer service proactive dialogue model training device 700 further includes:

[0289] Multi-dimensional evaluation module 705, configured to evaluate the customer service proactive dialogue model from multiple preset dimensions to obtain multiple evaluation results;

[0290] Comprehensive evaluation module 706, configured to obtain a comprehensive evaluation result according to the multiple evaluation results;

[0291] Return training module 707, configured to, in response to the comprehensive evaluation result not meeting the preset evaluation conditions, return and execute using the large language model to predict the customer service dialogue data corresponding to the hybrid data and output predicted customer service dialogue data.

[0292] As a possible implementation manner, the preset conditions include any combination of the following situations:

[0293] Return and execute using the large language model to predict the customer service dialogue data corresponding to the mixed data, and output the predicted customer service dialogue data reaching a preset number of times;

[0294] The loss converges or is within a preset loss range;

[0295] The perplexity of the customer service proactive dialogue model is within a preset perplexity range;

[0296] The large-scale multi-task language understanding ability of the customer service proactive dialogue model is within a preset understanding ability range.

[0297] As a possible implementation manner, the comprehensive evaluation module 706 includes:

[0298] A preference calculation unit 7061, configured to map each of the evaluation results to a preference reward value according to a preset reward mapping rule;

[0299] A comprehensive calculation unit 7062, configured to obtain a comprehensive evaluation result according to multiple preference reward values.

[0300] As a possible implementation manner, the preference calculation unit 7061 includes:

[0301] A first formula calculation component 70611 (not shown in the figure), configured to map each of the evaluation results to a preference reward value by using formulas (1)-(2), where

[0302] Formula (1) is

[0303] Formula (2) is

[0304] where r i is the evaluation result, sign(r i ) is the indicator function, f(r i ) is the preference reward value of the evaluation result r i , tanh is the hyperbolic tangent function, k is a parameter preset to control the sensitivity of f(r i ), and x0 is a parameter preset to control the center point of the f(r i ) curve; and

[0305] The comprehensive calculation unit 7062 includes:

[0306] A second formula calculation component 70621 (not shown in the figure), configured to obtain a comprehensive evaluation result by using formula (3), where

[0307] Formula (3) is

[0308] where r mo is the comprehensive evaluation result, w k is the preset weight corresponding to the evaluation result r i , n is the number of the evaluation results, and f(r i ) is the preference reward value of the evaluation result r i .

[0309] Further referring to Figure 8 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a customer service active dialogue device, and this device embodiment corresponds to Figure 6 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0310] As Figure 8 shown, the customer service active dialogue device 800 in this embodiment includes: an indication information acquisition module 801 and a dialogue module 802.

[0311] Among them, the indication information acquisition module 801 is used to acquire indication information;

[0312] The dialogue module 802 is used to input the indication information and the customer real-time dialogue data into the trained customer service active dialogue model, and output the customer service real-time dialogue data, and the customer service active dialogue model is trained by using the customer service active dialogue model training method shown in Figure 2 or Figure 3A .

[0313] Next, referring to Figure 9 , which shows a schematic structural diagram of a computer system 900 of an electronic device suitable for implementing the present disclosure. Figure 9 The shown computer system 900 is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present disclosure.

[0314] As Figure 9 shown, the computer system 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage device 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the computer system 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0315] Typically, the following devices can be connected to the I / O interface 905: input devices 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, etc.; output devices 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and a communication device 909. The communication device 909 can allow the computer system 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the computer system 900 of the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.

[0316] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above functions defined in the methods of the embodiments of the present disclosure are executed.

[0317] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0318] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0319] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to implement the method shown in Figures 2 to 6 any of the embodiments and their optional implementation manners shown.

[0320] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the client computer, partially on the client computer, executed as a stand-alone software package, partially on the client computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the client computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0321] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions denoted in the blocks may occur in a different order than that denoted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0322] The units or modules involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the unit or module does not constitute a limitation on the unit itself in some cases. For example, the training module may also be described as "used to return and execute the use of the large language model to predict the customer service dialogue data corresponding to the mixed data, output the predicted customer service dialogue data, and obtain the customer service proactive dialogue model until the large language model meets the preset conditions."

[0323] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present disclosure that have similar functions.

Claims

1. A method for training a customer service proactive dialogue model, characterized in that, Including: Obtain training data, where the training data includes general conversation data and real conversation data, and both the general conversation data and the real conversation data include customer service conversation data and customer conversation data; Extract features from the training data to obtain prompt information; Combine the training data and the prompt information into mixed data; Use the mixed data to train a large language model to obtain a customer service proactive conversation model.

2. The method according to claim 1, wherein The extracting features from the training data to obtain prompt information includes: Obtain a preset prompt template; Taking each example of general conversation data or real conversation data in the training data as a unit, extract prompt information according to the preset prompt template.

3. The method according to claim 1, wherein The using the mixed data to train a large language model to obtain a customer service proactive conversation model includes: Using the large language model to predict the customer service conversation data corresponding to the mixed data, and output the predicted customer service conversation data; Calculate the loss between the predicted customer service conversation data and the customer service conversation data in the mixed data; Based on the loss, adjust the parameters of the large language model; Return to execute the using the large language model to predict the customer service conversation data corresponding to the mixed data, and output the predicted customer service conversation data until the large language model meets the preset conditions, then obtain the customer service proactive conversation model.

4. The method according to claim 3, wherein After obtaining the customer service proactive conversation model, it further includes: Evaluate the customer service proactive conversation model from multiple preset dimensions to obtain multiple evaluation results; Obtain a comprehensive evaluation result according to the multiple evaluation results; In response to the comprehensive evaluation result not meeting the preset evaluation conditions, return to execute the using the large language model to predict the customer service conversation data corresponding to the mixed data, and output the predicted customer service conversation data.

5. The method according to claim 3, characterized in that The preset conditions include any combination of the following situations: Return to execute the using the large language model to predict the customer service conversation data corresponding to the mixed data, and output the predicted customer service conversation data reaching the preset number of times; The loss converges or is within a preset loss range; The perplexity of the customer service proactive conversation model is within a preset perplexity range; The large-scale multi-task language understanding ability of the customer service proactive conversation model is within a preset understanding ability range.

6. The method according to claim 4, characterized in that, The obtaining a comprehensive evaluation result according to the multiple evaluation results includes: According to a preset reward mapping rule, map each evaluation result to a preference reward value; Obtain a comprehensive evaluation result according to the multiple preference reward values.

7. The method according to claim 6, wherein The according to a preset reward mapping rule, map each evaluation result to a preference reward value includes: Use formulas (1)-(2) to map each evaluation result to a preference reward value, where Formula (1) is Formula (2) is where r i is the evaluation result, sign(r i ) is the indicator function, f(r i ) is the preference reward value of the evaluation result r i , tanh is the hyperbolic tangent function, k is the parameter for presetting the sensitivity of f(r i ), and x0 is the parameter for presetting the center point of the curve of f(r i ); and The according to the multiple preference reward values, obtain a comprehensive evaluation result includes: Use formula (3) to obtain a comprehensive evaluation result, where Formula (3) is Among them, r mo is the comprehensive evaluation result, w k is the preset weight corresponding to the evaluation result r i , n is the number of the evaluation results, f(r i ) is the preference reward value of the evaluation result r i .

8. A customer service proactive dialogue method, characterized in that, Including: Obtain indication information; Input the indication information and customer real-time conversation data into the trained customer service proactive conversation model, and output customer service real-time conversation data, where the customer service proactive conversation model is trained by the method described in any one of claims 1-7.

9. A training device for a customer service proactive dialogue model, characterized in that, Including: A training data acquisition module for acquiring training data, where the training data includes general dialogue data and real dialogue data, and both the general dialogue data and the real dialogue data include customer service dialogue data and customer dialogue data; A feature extraction module for extracting features from the training data to obtain prompt information; A combination module for combining the training data and the prompt information into mixed data; A training module for using the mixed data to train a large language model to obtain a customer service proactive dialogue model.

10. A customer service proactive dialogue device, characterized in that, Comprising: An indication information acquisition module for acquiring indication information; A dialogue module for inputting the indication information and customer real-time dialogue data into the trained customer service proactive dialogue model and outputting customer service real-time dialogue data, where the customer service proactive dialogue model is trained by using the method described in any one of claims 1-7.

11. An electronic device, characterized in that, Comprising: One or more processors; A storage device on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of claims 1-7 and / or 8.

12. A computer-readable storage medium, characterized in that, On which a computer program is stored, where the computer program, when executed by one or more processors, implements the method described in any one of claims 1-7 and / or 8.

13. A computer program product, characterized in that, Comprising computer programs / instructions, where the computer programs / instructions, when executed by a processor, implement the method described in any one of claims 1-7 and / or 8.