Service conversation quality inspection model training method and device, equipment and storage medium

By generating preliminary analysis results through the teacher model and inputting them into the student model for learning and optimization, the problems of insufficient accuracy and low efficiency of the dialogue quality inspection system in the existing technology are solved, and the self-evolution and efficient adaptation of the dialogue quality inspection are achieved.

CN119646213BActive Publication Date: 2025-10-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411548215.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-10-10
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

Existing customer conversation quality inspection technology is difficult to quickly adapt to complex and diverse conversation content, and especially exhibits significant limitations when dealing with unknown conversations or sudden issues, resulting in insufficient accuracy and low efficiency.

Method used

The teacher model is used to conduct a preliminary analysis of the service dialogue content, and the preliminary analysis results are generated through the target prompt text. These results are then input into the student model for learning and optimization to form a service dialogue quality inspection model that can self-iterate and adapt to the needs of different scenarios.

Benefits of technology

It improves the accuracy and efficiency of conversation quality inspection, reduces the workload of manual review, enhances the generalization ability of the model, and can automatically adapt to changes in customer needs and complex conversation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646213B_ABST
    Figure CN119646213B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method and device of a service dialogue quality inspection model, equipment and a storage medium. The present disclosure relates to the technical field of artificial intelligence, in particular to the fields of large models, instruction optimization, data analysis, etc., and can be used in application scenarios such as intelligent customer service, satisfaction management, dialogue analysis, dialogue quality inspection, etc. The specific implementation scheme is as follows: obtaining a target prompt text; inputting service dialogue content and the target prompt text into a teacher model to generate a preliminary analysis result of the service dialogue content based on the target prompt text by the teacher model; inputting the preliminary analysis result into a student model to enable the student model to learn and optimize according to the preliminary analysis result; determining the student model after learning and optimization as a service dialogue quality inspection model, and the service dialogue quality inspection model is used to generate a service dialogue quality inspection result according to the target prompt text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of large models, instruction optimization, data analysis, etc., and can be used in application scenarios such as intelligent customer service, satisfaction management, dialogue analysis, and dialogue quality inspection. Specifically, it relates to a training method, device, equipment and storage medium for a service dialogue quality inspection model. Background Art

[0002] Existing customer conversation quality control technologies primarily rely on manual annotation rules or statistically based machine learning models. However, with the increasing complexity of conversation content and the diversification of customer needs, traditional quality control systems struggle to adapt and respond quickly, especially when handling unknown conversations or unexpected issues. Summary of the Invention

[0003] The present disclosure provides a training method, apparatus, device and storage medium for a service dialogue quality inspection model.

[0004] According to the first aspect of the present disclosure, a training method for a service conversation quality inspection model is provided, the method comprising: obtaining a target prompt text; inputting the service conversation content and the target prompt text into a teacher model, so that the teacher model generates a preliminary analysis result of the service conversation content based on the target prompt text; inputting the preliminary analysis result into a student model, so that the student model learns and optimizes according to the preliminary analysis result; and determining the student model after learning and optimization as a service conversation quality inspection model, which is used to generate a service conversation quality inspection result based on the target prompt text.

[0005] According to the second aspect of the present disclosure, a service conversation quality inspection method based on a large model is provided, the method comprising: obtaining the service conversation content to be quality inspected, inputting the service conversation content to be quality inspected into a service conversation quality inspection model, and obtaining the service conversation quality inspection result output by the service conversation quality inspection model based on the target prompt text; wherein the service conversation quality inspection model is obtained by training the method of the first aspect mentioned above.

[0006] According to a third aspect of the present disclosure, a training device for a service dialogue quality inspection model is provided, which includes: a first acquisition module for acquiring a target prompt text; a first control module for inputting the service dialogue content and the target prompt text into a teacher model, so that the teacher model generates a preliminary analysis result of the service dialogue content based on the target prompt text; a second control module for inputting the preliminary analysis result into a student model, so that the student model performs learning and optimization according to the preliminary analysis result; a determination module for determining the student model after learning and optimization as a service dialogue quality inspection model, and the service dialogue quality inspection model is used to generate a service dialogue quality inspection result based on the target prompt text.

[0007] According to a fourth aspect of the present disclosure, a large model-based service conversation quality inspection device is provided, the device comprising: a second acquisition module configured to acquire service conversation content to be inspected; and a quality inspection module configured to input the service conversation content to be inspected into a service conversation quality inspection model to obtain a service conversation quality inspection result output by the service conversation quality inspection model according to target prompt text; wherein the service conversation quality inspection model is obtained by training the method of the first aspect.

[0008] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising:

[0009] at least one processor; and

[0010] a memory in communication with the at least one processor; wherein

[0011] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any of the embodiments of the present disclosure.

[0012] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the method according to any of the embodiments of the present disclosure.

[0013] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any of the embodiments of the present disclosure.

[0014] The scheme of the present disclosure can enhance the generalization ability of the model and improve the accuracy and efficiency of conversation quality inspection.

[0015] It should be understood that the contents described in this part are not intended to identify the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0016] The accompanying drawings serve to better understand the present scheme and do not constitute limitations on the present disclosure. Among them:

[0017] Figure 1 is a flowchart of a training method of a service conversation quality inspection model according to an embodiment of the present disclosure.

[0018] Figure 2 is a schematic diagram of target prompt text according to an embodiment of the present disclosure.

[0019] Figure 3 is an evolution flowchart of a service conversation quality inspection model according to an embodiment of the present disclosure.

[0020] Figure 4 This is a schematic diagram of an example of a service dialogue quality inspection model according to an embodiment of the present disclosure.

[0021] Figure 5 is a flowchart of a service dialogue quality inspection method based on a large model according to an embodiment of the present disclosure;

[0022] Figure 6 This is a structural diagram of a training device for a service dialogue quality inspection model according to an embodiment of the present disclosure.

[0023] Figure 7 This is a structural diagram of a service dialogue quality inspection device based on a large model according to an embodiment of the present disclosure.

[0024] Figure 8 This is a scenario diagram of a training method for a service dialogue quality inspection model according to an embodiment of the present disclosure.

[0025] Figure 9 This is a scenario diagram of a service dialogue quality inspection method based on a large model according to an embodiment of the present disclosure.

[0026] Figure 10 3 is a schematic diagram of the structure of an electronic device used to implement the training method of the service dialogue quality inspection model and / or the service dialogue quality inspection method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0028] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.

[0029] In addition, for better illustration of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art will understand that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, elements and circuits well known to those skilled in the art are not described in detail in order to highlight the main ideas of the present disclosure.

[0030] Before introducing the technical solutions of the embodiments of the present disclosure, the technical terms that can be used in the present disclosure are further described:

[0031] Teacher model: a pre-trained model with strong analysis capability, used to generate preliminary analysis results of service conversation content.

[0032] Student model: a quality inspection model to be trained, which learns and optimizes by imitating the output of the teacher model.

[0033] Target prompt text: text information related to quality inspection standards, business specifications or user expectations, used to guide the teacher model and the student model to analyze the conversation content.

[0034] In the related art, the conversation quality inspection system usually analyzes the content in the conversation and gives a judgment through a pre-defined rule set or model. Many quality inspection systems can only rely on fixed model structures and cannot improve the quality inspection effect through model evolution. The conversation quality inspection technology mainly has the following disadvantages: relying on pre-defined rules and pre-trained models, unable to automatically adapt to new conversation modes or changes in customer needs, resulting in insufficient accuracy of conversation quality inspection; most quality inspection systems cannot improve the performance of conversation quality inspection through self-learning or model evolution, and need to rely on frequent manual intervention for model updating, which is inefficient; when involving multi-round conversations or complex problems, the classification and judgment ability of the system is weak, and misjudgment or omission is easy to occur.

[0035] The present disclosure proposes a model-based service conversation quality inspection method to at least partially solve one or more of the above problems and other potential problems. Through self-learning and model evolution, the accuracy and efficiency of conversation quality inspection can be improved, solving the problems of low efficiency, inaccurate judgment and lack of evolution ability of the current conversation quality inspection system in the customer service field, and adapting to the changing customer needs and complex conversation scenarios.

[0036] The embodiments of the present disclosure provide a training method of a service conversation quality inspection model, Figure 1This is a flow chart of a method for training a service conversation quality inspection model according to an embodiment of the present disclosure. The method for training a service conversation quality inspection model can be applied to a service conversation quality inspection model training device. The service conversation quality inspection model training device is located in an electronic device. The electronic device includes but is not limited to fixed devices and / or mobile devices. For example, fixed devices include but are not limited to servers, and servers can be cloud servers or ordinary servers. For example, mobile devices include but are not limited to: mobile phones, tablet computers. In some possible implementations, the method for training a service conversation quality inspection model can also be implemented by a processor calling computer-readable instructions stored in a memory. For example Figure 1 As shown in the figure, the training method of the service dialogue quality inspection model includes:

[0037] S101. Get target prompt text

[0038] S102: Inputting the service dialogue content and the target prompt text into the teacher model, so that the teacher model generates a preliminary analysis result of the service dialogue content based on the target prompt text;

[0039] S103, inputting the preliminary analysis results into the student model, so that the student model learns and optimizes according to the preliminary analysis results;

[0040] S104: Determine the student model after learning and optimization as a service dialogue quality inspection model, and the service dialogue quality inspection model is used to generate a service dialogue quality inspection result based on the target prompt text.

[0041] In the embodiment of the present disclosure, the target prompt text includes text for guiding the model to perform analysis, which contains key information and requirements required for the analysis.

[0042] In the disclosed embodiment, the teacher model is a pre-trained model with analytical capabilities, which can analyze the service dialogue content according to the target prompt text.

[0043] In the disclosed embodiment, the student model is a pre-trained quality inspection model that can learn and optimize based on the preliminary analysis results of the teacher model.

[0044] In the embodiment of the present disclosure, the target prompt text includes key information of multiple dimensions, and the multiple dimensions include: problem information, problem status, provided solution, solution result and service evaluation.

[0045] In some embodiments, after the service conversation content and target prompt text are input into the teacher model, the teacher model will generate a preliminary analysis result, which may include key information extraction, sentiment tendency judgment, compliance check and other levels of content in the conversation.

[0046] In some implementations, the student model learns and optimizes based on the initial analysis results. This includes mimicking the output of the teacher model, gradually mastering the ability to analyze conversation content, identify potential issues, and generate quality inspection results. Furthermore, the student model can continuously optimize its internal parameters and strategies through self-iteration and feedback mechanisms to improve the accuracy and efficiency of quality inspections.

[0047] In some embodiments, the student model after learning and optimization is determined as the service dialogue quality inspection model, including: after multiple rounds of learning and optimization, the performance of the student model will gradually stabilize and reach the expected quality inspection standards, and the student model can be determined as the final service dialogue quality inspection model and deployed in actual application scenarios.

[0048] For example, suppose the customer service department of an e-commerce platform needs to perform quality control on conversations between customer service representatives and users. This involves collecting and organizing historical conversation data between customer service representatives and users, as well as internal service specifications and user expectations, to form a set of target prompt texts. This historical conversation data and target prompt texts are then fed into a teacher model to generate preliminary analysis results. This preliminary analysis result is then fed into a student model, which learns and optimizes by mimicking the teacher model's output. After multiple rounds of learning and optimization, the student model is finalized as the final service conversation quality control model and deployed within the e-commerce platform's customer service system. As customer service representatives and users engage in conversations, the quality control model automatically performs quality control on the conversation content and generates quality control results. If there are potential issues or non-compliance with service specifications in the conversation, the quality control model will issue warnings or prompts, allowing customer service personnel to promptly correct the situation.

[0049] The technical solution of the disclosed embodiment utilizes a teacher model for preliminary analysis, which can greatly reduce the computational burden of the student model and improve the real-time and accuracy of quality inspection. By imitating the output of the teacher model, the student model can learn more about conversation content analysis, thereby enhancing its ability to handle different scenarios and problems. Deploying the trained service conversation quality inspection model into actual applications can automatically complete quality inspection tasks, greatly reducing the workload and cost of manual review and improving quality inspection efficiency.

[0050] In some embodiments, obtaining the target prompt text includes: obtaining an initial prompt text; optimizing and adjusting the initial prompt text according to the needs of the service industry to obtain the target prompt text.

[0051] In the disclosed embodiment, the target prompt text is used to guide the model to generate a specific output. It contains information and instructions that the model needs to understand and is an important basis for the model to generate output.

[0052] In the embodiments of the present disclosure, the service industry refers to an industry or department that provides service labor, which has a wide range and includes catering, tourism, finance, medical care, education, e-commerce and other fields.

[0053] In the disclosed embodiment, the initial prompt text is a basic, unoptimized text that may be derived from a preset template, historical data, or other relevant resources. Such text typically contains some basic instructions or information but has not been customized to the needs of a specific service industry.

[0054] In some embodiments, optimization and adjustment based on the needs of the service industry include: First, an in-depth analysis of the needs of the service industry is required. This includes understanding the specific scenarios, business processes, user preferences, and expected outputs of the service industry. Through demand analysis, it is possible to clarify what key information and instructions the target prompt text needs to include. Based on the results of the demand analysis, the initial prompt text is optimized and adjusted. This may include adding necessary contextual information, adjusting the wording of instructions, deleting redundant information, etc. The goal of optimization is to make the target prompt text more consistent with the actual needs of the service industry and to guide the model to generate more accurate and useful outputs.

[0055] In some implementations, the initial prompt text and related background information are input into a large model. Leveraging the model's natural language processing capabilities and understanding of large amounts of textual data, optimization suggestions are generated. Based on the optimization suggestions, the initial prompt text is optimized and adjusted to produce the target prompt text. Adjustments may include adjusting wording, sentence structure, and other factors to make the prompt text clearer, more accurate, and more engaging. Background information may include the background, characteristics, and customer needs of the service industry, allowing this information to be incorporated into the optimization process.

[0056] In some embodiments, after optimization is complete, the target prompt text may need to be tested and verified. For example, the target prompt text may be input into the model to see if the model's output meets expectations. If the model's output does not meet expectations, further adjustments to the target prompt text may be necessary.

[0057] Suppose a hotel industry wants to use a large model to answer employees' questions about the hotel's Property Management System (PMS). The initial prompt text might be a simple query: "Please explain what PMS is?" However, such a prompt text may not provide the specific information that employees need. To optimize this prompt text, we can make the following adjustments: Add contextual information: Adjust the query to: "In the tourism industry, what does PMS generally refer to?" This statement can provide employees with more background information and help them better understand the concept of PMS. Further adjust the prompt text to: "Please explain what PMS is in the tourism industry in a concise and clear manner and give its main functions." This statement clarifies the output requirements of the model and helps generate more accurate and useful answers. Through the above optimization adjustments, we can obtain target prompt text that better meets the needs of the hotel industry, thereby guiding the model to generate more accurate and useful outputs.

[0058] In this way, by optimizing and adjusting the target prompt text, it can be made more in line with the actual needs of the service industry, thereby improving the accuracy of the model output; the optimized target prompt text can more effectively guide the model to output, thereby reducing unnecessary iterations and adjustments and improving work efficiency.

[0059] In some embodiments, the target prompt text includes key information in multiple dimensions, including: problem information, problem status, provided solutions, solution results, and service evaluation.

[0060] In the embodiment of the present disclosure, the question information is the core question or demand extracted from the conversation, which clarifies the specific matters that the user wants to solve, and may include detailed information such as a description of the problem, the time and place of occurrence, etc.

[0061] In the disclosed embodiment, the problem status is the current state of the problem, such as reported, being processed, resolved, etc. The problem status reflects the changes of the problem during the dialogue process, such as from unknown to known, from pending to processed, etc.

[0062] In the embodiments of the present disclosure, the solutions provided are solutions or suggestions proposed for the problems, including the specific content of the solutions, implementation steps, expected effects, etc.

[0063] In the embodiment of the present disclosure, the solution result is to record the actual result after the implementation of the solution, such as whether the problem is solved and the degree of solution, etc., reflecting the effectiveness of the solution and providing a basis for subsequent improvements.

[0064] In the embodiment of the present disclosure, service evaluation is to collect user feedback on the service process and results, including evaluations in terms of satisfaction, service quality, response speed, etc.

[0065] In constructing the target prompt text, a templated approach can be adopted to provide fixed formats and expressions for each dimension. At the same time, according to the specific content of the dialogue and the characteristics of the service industry, the template is appropriately adjusted and supplemented to ensure the accuracy and completeness of the target prompt text.

[0066] Figure 2 A schematic diagram of the target prompt text is shown as Figure 2 The target prompt text includes task requirement description, extraction of problem information, solution, solution status, solution result, service evaluation, etc.

[0067] Specifically, the task requirement description is: understand the content of the dialogue, you need to read and understand the content of the dialogue between the customer and the artificial customer service.

[0068] Specifically, the extraction of problem information includes problem description (description): identify the key description of the problem or demand raised by the customer; phenomenon description (phenomenon): record the specific phenomenon or situation encountered by the customer; classification label (classification): select the most suitable label for the problem description or phenomenon from the given classification list, and if necessary, supplement the solution content judgment to ensure that the label belongs to the given list.

[0069] Specifically, the resolution includes: pay attention to all the solutions provided by the customer service, and objectively and in detail describe each step.

[0070] Specifically, the solution status includes: solved: the customer's response directly solves the customer's problem, or provides a clear solution, and the user does not show confusion or further inquiry about the problem in the subsequent dialogue; unsolved: the customer's response fails to solve the customer's problem, or the solution provided is not feasible and needs further follow-up processing (such as: directing the user to the information service center, information technology personnel on-site processing, arranging engineer service, waiting for other personnel to contact the user, or issuing a dispatch, etc.); not clear: the content of the dialogue is limited or the user does not give further feedback, and it is not possible to determine whether the problem has been solved (such as: the user does not respond to the customer service solution for a long time, expresses that they will try to suggest but there is no subsequent feedback, or the customer service suggests that the user consult other services).

[0071] Specifically, the solution result indicates whether the problem has been solved, and any challenges or difficulties in the solution process. If the problem is not solved or not clear, please explain the plan and way of further follow-up.

[0072] Specifically, the service evaluation can include the satisfaction score of the first object (such as the customer) and the service attitude score of the second object (such as the customer service).

[0073] For example, from the perspective of the customer, the satisfaction with the customer service is evaluated, including problem-solving rate and accuracy, service attitude, communication effectiveness, etc. The score of the customer satisfaction score reflects the customer's satisfaction with the service process. Specifically, the better the performance in problem-solving rate and accuracy, service attitude, communication effectiveness, etc., the higher the customer satisfaction score obtained after weighted calculation according to a scientific and reasonable proportion. Specifically, if the customer's tone is positive, the problem is solved quickly and accurately, and the communication process is smooth and unobstructed, the customer satisfaction score is high; if the customer's tone is strongly negative, the problem is completely unsolved or the communication is severely obstructed, the customer satisfaction score is low.

[0074] For example, according to the performance of the customer service in the service, including politeness, professionalism, initiative, patience, etc. The score of the service attitude score reflects the overall performance of the customer service in the service process. Specifically, the better the performance of the customer service in politeness, professionalism, initiative, patience, etc., the higher the service attitude score obtained after weighted calculation according to a scientific and reasonable proportion. If the customer service is always polite, respectful, professional, active, and patient, and responds quickly, the service attitude score is high; if the customer service has almost no politeness and professionalism, responds very slowly, and cannot effectively communicate with the customer, the service attitude score is low.

[0075] As can be seen, Figure 2 The prompt text shown gives the definition and design of the dialogue quality inspection, guiding the model to automatically inspect the dialogue between the customer and the artificial customer service, identify the problem, classify the phenomenon, judge the processing state, and give the corresponding quality feedback. This prompt text can adapt to intelligent customer service, customer feedback management and dialogue analysis scenarios, making the model more general and extensible.

[0076] Suppose in the customer service dialogue of a certain e-commerce platform, the user reflects that the received goods have quality problems. The following is an example of building a target prompt text: problem information: the user's purchased goods have quality problems such as damage, color inconsistency, etc. Problem status: the user has reported the problem to the customer service, and the customer service is verifying the situation. Provided solution: the customer service proposes a solution of return for refund or exchange, and informs the user of the specific operation steps. Solution result: the user chooses to return for refund, the goods have been returned, and the refund has been credited. Service evaluation: the user expresses satisfaction with the response speed and solution of the customer service, but expresses dissatisfaction with the quality of the goods. Through such a target prompt text, the key information in the dialogue can be clearly understood, providing strong support for subsequent quality inspection and analysis.

[0077] In this way, by including key information in multiple dimensions, the target prompt text can comprehensively reflect various aspects of the service dialogue, providing rich data support for subsequent quality inspection and analysis. The structured target prompt text makes the information clearer and more organized, facilitating quick understanding of the dialogue content by quality inspection personnel and improving quality inspection efficiency. By collecting service evaluation and problem solving results and other information, the enterprise can timely discover deficiencies and problems in the service process, providing a strong basis for service improvement.

[0078] In some embodiments, the teacher model generates a preliminary analysis result of the service dialogue content based on the target prompt text, including: using the teacher model to understand the service dialogue content to obtain specific content, context, and respective intentions and focus points of the first object and the second object of the service dialogue content; and generating the preliminary analysis result of the service dialogue content according to the specific content, context, and respective intentions and focus points of the first object and the second object of the service dialogue content based on the target prompt text.

[0079] In the embodiments of the present disclosure, the service dialogue content includes the dialogue between the customer and the customer service, including the customer's question, demand, and the customer service's response, suggestion, etc.

[0080] In the embodiments of the present disclosure, the first object usually refers to the customer, who is the receiver of the service; and the second object usually refers to the customer service, who is the provider of the service.

[0081] In the embodiments of the present disclosure, the intention refers to the action target or demand of the object in the dialogue; and the focus point refers to the aspect that the object in the dialogue focuses on or emphasizes.

[0082] In some embodiments, the teacher model first analyzes the service dialogue content and extracts key information such as specific events, time, place, goods or services, etc. Based on the extracted specific content, the teacher model constructs the context environment of the dialogue, understands the background, historical communication and the stage of the current dialogue. The teacher model further analyzes the respective intentions and focus points of the first object (such as the customer) and the second object (such as the customer service) in the dialogue. This includes identifying the actual demand, emotional state of the customer, and response strategy, service attitude, etc. of the customer service.

[0083] In some embodiments, based on the understanding of the service dialogue content, the teacher model matches the key information in the dialogue with the target prompt text to determine whether the dialogue content meets the preset quality inspection standard or business specification. The teacher model analyzes the dialogue content from multiple dimensions such as problem information, problem status, provided solution, solution result, and service evaluation according to the requirements of the target prompt text.

[0084] In some embodiments, the teacher model integrates the analyzed information into a preliminary analysis result according to the target prompt text, including the evaluation results of compliance, customer satisfaction, service efficiency, etc.

[0085] Suppose in a customer service conversation on an e-commerce platform, the customer reports to the customer service that the purchased goods have not been delivered on time. The following is an example of the preliminary analysis result generated by the teacher model:

[0086] Service conversation content understanding includes:

[0087] 1. Specific content: The customer's purchased goods have not been delivered on time, and the customer asks about the delivery time and reason.

[0088] 2. Context: The customer has repeatedly urged delivery before, and the customer service has indicated that it is being handled.

[0089] 3. Intent and focus: The customer wants to receive the goods as soon as possible and is concerned about the delivery time and reason; the customer service wants to solve the customer's problem and is concerned about how to appease the customer and provide a solution.

[0090] Preliminary analysis results include:

[0091] 1. Problem information: The customer's purchased goods have not been delivered on time.

[0092] 2. Problem status: The customer has reported to the customer service, and the customer service is handling it.

[0093] 3. Provided solution: The customer service promises to speed up the delivery and provides certain compensation measures.

[0094] 4. Solution result: Not yet clear, need to follow up later.

[0095] 5. Service evaluation: The customer is dissatisfied with the delay in delivery and expresses general satisfaction with the customer service response speed and attitude.

[0096] Through such preliminary analysis results, the key information and problems in the conversation can be clearly understood, providing strong support for subsequent service improvement and quality inspection work.

[0097] In this way, by deeply understanding the specific content, context, and intent and focus of the first and second objects of the service conversation content, the teacher model can more accurately judge the quality of the conversation and the problems. The teacher model analyzes the conversation content from multiple dimensions to ensure the comprehensiveness and depth of the analysis, which helps to discover potential service problems and improvement points. The teacher model can quickly generate preliminary analysis results, reducing the workload of manual quality inspection and improving the efficiency and accuracy of quality inspection, thereby providing accurate data support for the learning and optimization of subsequent student models, which helps to improve the efficiency and accuracy of student model quality inspection.

[0098] In some embodiments, a preliminary analysis result of the service dialogue content is generated according to the target prompt text, including: guiding the teacher model to gradually infer the logic behind the problem according to the target prompt text through chain reasoning: combining the logic behind the problem and generating a preliminary analysis result of the service dialogue content according to the target prompt text.

[0099] In the disclosed embodiments, chain reasoning is a method of gradually reasoning out the logic behind things. By identifying the cause-effect relationship, time sequence and other relationships between things, a logical chain is constructed to deeply understand the essence of things.

[0100] In some embodiments, a chain reasoning approach is used to gradually reason out the logic behind the problem and generate analysis results based on this logic, including: the teacher model first understands the content of the service conversation and extracts key information from the conversation, such as the identities of the two parties in the conversation, the topic of the conversation, the questions or needs raised, etc. Based on the extracted key information, the teacher model gradually analyzes the logical relationships in the conversation content through chain reasoning. This logical relationship includes identifying the root cause of the problem, the impact of the problem on both parties, and the attitudes and reactions of both parties to the problem. After identifying the logic of the problem, the teacher model constructs this logic into a logical chain according to chronological order or causal relationship, forming a comprehensive and in-depth understanding of the problem.

[0101] In some embodiments, the preliminary analysis results are generated in conjunction with logic, including: after constructing the logical chain, the teacher model matches the logical information in the conversation with the target prompt text to determine which logical information meets the requirements of the target prompt text. Based on the matching results, the teacher model generates preliminary analysis results of the service conversation content according to the format and requirements of the target prompt text. For example, the preliminary analysis results include an assessment of the compliance of the conversation, a determination of the nature of the problem, and the division of responsibilities between the two parties. After generating the preliminary analysis results, the teacher model can conduct further review and revision to ensure the accuracy and completeness of the analysis results.

[0102] In this way, through chain reasoning, the teacher model can gradually delve into the essence of the problem and reveal the logical relationships behind it, thereby improving the depth and accuracy of the analysis. Constructing a logical chain helps the teacher model more clearly present the ins and outs of the problem, enhancing the logic and persuasiveness of the analysis results. Chain reasoning guides the teacher model to conduct analysis in a more orderly manner, avoiding the interference of invalid information and improving the efficiency of analysis.

[0103] In some embodiments, the student model learns and optimizes according to the preliminary analysis results, including: in the first stage, the student model is trained by imitating the preliminary analysis results output by the teacher model to learn to handle various problems and scenarios in the dialogue; in the second stage, the student model is optimized for target dialogue phenomena by introducing sample selection and feedback enhancement mechanisms; in the third stage, the student model iterates and optimizes itself based on its own data accumulation and feedback during the dialogue process to achieve comprehensive autonomous evolution of the student model.

[0104] In the embodiments of the present disclosure, imitation learning is a machine learning method that learns to handle specific tasks by imitating the behavior of an expert or teacher model.

[0105] In the embodiments of the present disclosure, feedback enhancement is a machine learning method that improves the training effect of a model by introducing additional supervision information (such as scores, labels, etc.).

[0106] In the embodiments of the present disclosure, self-iteration refers to the process of continuously optimizing and upgrading the model based on new data and feedback in actual application.

[0107] In the embodiments of the present disclosure, the first stage is the imitation learning stage.

[0108] In some implementations, the student model is trained by imitating the preliminary analysis results output by the teacher model to learn to handle various problems and scenarios in the dialogue, including: collecting a large amount of service dialogue content and its corresponding teacher model preliminary analysis results to form a training data set; the student model takes the preliminary analysis results of the teacher model as the learning goal, imitates the processing logic and output results of the teacher model through supervised learning, compares the output results of the student model with the preliminary analysis results of the teacher model, evaluates the imitation learning effect of the student model, and makes necessary adjustments and optimizations.

[0109] In the embodiments of the present disclosure, the second stage is the optimization and improvement stage.

[0110] In some implementations, in the second stage, the student model optimizes for target dialogue phenomena by introducing sample selection and feedback enhancement mechanisms, including: in the training data set, high-quality samples related to target dialogue phenomena are selected for further optimization of the student model; an artificial or automatic feedback mechanism is introduced to score or label the output results of the student model to provide additional supervision information; based on the selected samples and feedback information, the student model is further trained and optimized to improve its ability to handle specific dialogue phenomena.

[0111] In the embodiments of the present disclosure, the third stage is the self-iteration stage.

[0112] In some embodiments, in the third stage, the student model performs self-iteration and optimization based on its own data accumulation and feedback during the dialogue process to achieve comprehensive autonomous evolution of the student model, including: in actual applications, the student model continuously collects new dialogue data and user feedback to form its own data accumulation; regularly evaluates the student model, analyzes its performance in actual applications, and identifies existing problems and deficiencies; based on the evaluation results and new data accumulation, the student model performs self-iteration and optimization, including adjusting model parameters, updating the knowledge base, etc., to achieve comprehensive autonomous evolution.

[0113] Figure 3 The evolution process diagram of the service dialogue quality inspection model is shown in Figure 2. Figure 3 As shown in the figure, the evolution process is divided into L0, L1, L2, and L3 stages. The L0 stage is the instruction optimization stage, which is used to optimize the target prompt text. By optimizing the prompt, the teacher model reaches the theoretical optimal performance and provides a stable baseline. This stage does not require annotation. The L1 stage (i.e., the first stage) is the imitation learning stage. The student model is trained by imitating the preliminary analysis results of the teacher model's output to learn to handle various questions and scenarios in the conversation. The smaller and more responsive student model learns the teacher model's output, achieving over 70% of the teacher model's performance. This stage does not require annotation. The L2 stage (i.e., the second stage) is the optimization and improvement stage. The student model optimizes for the target conversational phenomena by introducing sample selection and feedback enhancement mechanisms. This second stage overcomes the marginal effects of the imitation stage and continuously improves the student model's performance, reaching 100% of the teacher model's performance. This stage does not require annotation. The L3 stage (i.e., the third stage) is the self-iteration stage. The student model uses its own data accumulation and feedback from the conversation process to self-iterate and optimize, achieving comprehensive and autonomous evolution. Through the third stage, the model effect can be continuously improved, and the model performance can be continuously improved through the method of model self-evolution; this stage may require one-time labeling, such as manually correcting which results output by the student model are good and which are bad.

[0114] Through the above four stages, the model evolution mechanism is adopted to enable the service dialogue quality inspection system to automatically adapt to changes in different customer needs, significantly improving the accuracy and efficiency of dialogue quality inspection.

[0115] In this way, through imitation learning and optimization, the student model can quickly master the ability to handle various conversational issues and scenarios, improving processing efficiency. By introducing sample selection and feedback enhancement mechanisms, the student model can optimize for specific conversational phenomena and enhance its adaptability. Based on its own data accumulation and feedback mechanism, the student model can continuously iterate and optimize itself, achieving continuous evolution and improving model quality inspection results.

[0116] In some embodiments, the student model is optimized for the target dialogue phenomenon by introducing a sample selection and feedback reinforcement mechanism. This includes gradually improving the student model's ability to recognize and handle the target dialogue phenomenon through multiple iterations of training on selected samples, and adjusting the student model's parameters based on rewards and penalties provided by the feedback reinforcement mechanism after each iteration.

[0117] In some implementations, sample selection involves filtering high-quality samples closely related to the target dialogue phenomenon from a large amount of dialogue data. These samples should be representative and fully reflect the characteristics of the target dialogue phenomenon. The selected samples are labeled to clearly indicate their class or processing results, providing supervision information during training. As training progresses, new samples are continuously introduced to replace old samples that have been fully learned, maintaining the diversity and timeliness of the training data.

[0118] In some implementations, a reward and penalty mechanism is designed to evaluate the student model's performance in handling the target dialogue phenomenon. When the model correctly identifies and handles the target dialogue phenomenon, it is rewarded; when the model performs poorly, it is penalized. Reward and penalty signals can be based on various factors such as dialogue fluency, user satisfaction, and accuracy of processing results. These signals can be obtained through manual labeling, user feedback, or automatic evaluation systems. Based on the reward and penalty signals, the student model's parameters are adjusted to optimize its ability to handle the target dialogue phenomenon, which can be achieved through gradient descent, reinforcement learning, and other algorithms.

[0119] In some implementations, a pre-trained model or randomly initialized parameters are used as the starting point for the student model. Multiple iterations of training are performed on selected samples. In each iteration, the student model is trained using the current sample set and the model parameters are adjusted based on the reward and penalty signals provided by the feedback reinforcement mechanism. After each iteration, the student model's performance is evaluated using a validation set. If the performance improves, continue iterating; if the performance decreases or stabilizes, consider adjusting the training strategy or stopping training.

[0120] Suppose, in an intelligent customer service system, the student model needs to optimize its ability to recognize and handle the conversational phenomenon of "users inquiring about product features." The following is an example of specific implementation and results: A large number of conversation samples related to user inquiries about product features are filtered from conversation logs and annotated. These samples include users asking about product features and customer service staff answering their questions. A feedback enhancement mechanism based on user satisfaction is designed. When users express satisfaction with the results of their product feature inquiries, the model is rewarded; when users express dissatisfaction, the model is penalized. Multiple iterations of training are performed on the selected samples. In each iteration, the student model is trained using the current sample set, and model parameters are adjusted based on user satisfaction feedback. After multiple iterations of training, the student model's ability to recognize and handle user inquiries about product features is significantly improved. In actual application, user satisfaction has significantly increased, and customer service efficiency has also improved.

[0121] In this way, through sample selection and feedback enhancement mechanisms, the student model can more accurately identify target conversational phenomena and improve processing accuracy. By introducing diverse samples and feedback signals, the student model can better adapt to different scenarios and user needs, enhancing robustness. Through iterative training and parameter adjustment, the student model can converge to the optimal solution more quickly, accelerating the training process.

[0122] In some embodiments, the method further includes: evaluating the performance of the student model at the end of each stage, and deciding whether to proceed to the next stage based on the evaluation results.

[0123] In order to ensure that the student model can steadily improve in each optimization stage and smoothly enter the next stage after reaching the predetermined goal, the performance of the student model is evaluated at the end of each stage.

[0124] In some implementations, developing evaluation criteria includes: developing clear evaluation indicators based on the characteristics of the target conversational phenomenon, such as accuracy, recall, precision, user satisfaction, etc.; setting reasonable thresholds for each evaluation indicator as a basis for determining whether to proceed to the next stage.

[0125] In some embodiments, the evaluation process includes: preparing an independent validation dataset at the end of each stage to evaluate the performance of the student model; using the validation dataset to evaluate the student model and calculate the value of each evaluation indicator; comparing the evaluation results with the set threshold and analyzing whether the performance of the student model meets the requirements.

[0126] In some embodiments, the evaluation results and analysis determine whether to proceed to the next stage. If the performance meets the requirements, the training proceeds to the next stage. If the performance does not meet the requirements, the training strategy is adjusted or additional training data is added, and the training of the current stage is repeated.

[0127] In some implementations, the evaluation results are promptly fed back to the model developer to understand the model's performance improvements and existing issues. Based on the evaluation results and feedback, the training strategy, sample selection, feedback enhancement mechanism, etc. are adjusted to improve model performance.

[0128] Suppose, in an intelligent customer service system, the student model needs to optimize its ability to recognize and process conversations involving users inquiring about product features. The following is an example of specific implementation and results:

[0129] Evaluation standard formulation: Accuracy, recall rate and user satisfaction are established as evaluation indicators, and reasonable thresholds are set, such as accuracy not less than 85%, recall rate not less than 80%, and user satisfaction not less than 4.5 points (out of 5 points).

[0130] Evaluation Process: At the end of each phase, the student model was evaluated using an independent validation dataset. After the first phase, the evaluation results showed an accuracy of 82%, a recall of 78%, and a user satisfaction score of 4.3, which did not meet the set threshold.

[0131] Feedback and Adjustment: Evaluation results were provided to model developers to analyze the reasons why the model's performance fell short of expectations. This analysis revealed that the model had difficulty identifying complex user questions. Therefore, the training strategy was adjusted to increase the number of complex user question samples and optimize the feedback enhancement mechanism to improve the model's ability to handle complex questions.

[0132] Retraining and Evaluation: We retrained the model using the adjusted training strategy and conducted another evaluation at the end of Phase 2. The evaluation results showed an accuracy of 87%, a recall of 83%, and a user satisfaction score of 4.7, all of which met or exceeded the set thresholds. Therefore, we decided to proceed to the next phase to further optimize the student model's performance.

[0133] In this way, through periodic evaluations, model problems can be promptly identified, preventing them from accumulating and causing performance degradation. Evaluation at the end of each stage ensures that the model advances to the next stage after achieving its intended goals, avoiding unnecessary retraining and improving training efficiency. Adjusting training strategies based on evaluation results can rationally allocate computing resources and time, improving resource utilization efficiency.

[0134] In some embodiments, the student model learns and optimizes based on the preliminary analysis results, including: determining a target reinforcement learning method that matches the service conversation content from multiple candidate reinforcement learning methods; the student model learns and optimizes based on the preliminary analysis results based on the target reinforcement learning method.

[0135] In the embodiments of the present disclosure, the reinforcement learning method is a machine learning algorithm that learns the optimal behavior policy through the interaction between the agent and the environment. In the dialogue system, the reinforcement learning method can be used to optimize the dialogue policy and improve the dialogue quality.

[0136] In the embodiments of the present disclosure, the target reinforcement learning method is determined from multiple candidate reinforcement learning methods that best match the service dialogue content, which is used to train and optimize the student model.

[0137] In the embodiments of the present disclosure, the preliminary analysis result is the result obtained after the dialogue content is preliminarily processed and analyzed, including user intent, dialogue state, historical information, etc. These results will be used as input data for training and optimizing the student model.

[0138] In some embodiments, the target reinforcement learning method is determined, including: screening candidate methods related to the service dialogue content from the existing reinforcement learning method library. These candidate methods should be able to process key information in the dialogue, such as user intent, dialogue state, historical information, etc. According to the preliminary analysis result, the applicability of each candidate method is evaluated. This includes considering the performance, stability, computational complexity of the method, and whether it is suitable for processing the characteristics of the current dialogue content. Through comparative analysis, the target reinforcement learning method that best matches the target dialogue content is determined. According to the characteristics of the target reinforcement learning method, the corresponding parameters are set. These parameters may include learning rate, discount factor, exploration rate, etc., which will affect the learning speed and effect of the model.

[0139] In some embodiments, the learning and optimization of the student model include: taking the preliminary analysis result as input data, including dialogue content, user intent, dialogue state, etc. These data will be used to train and optimize the student model. Based on the target reinforcement learning method, the input data is used to train the student model. During the training process, the model will generate responses according to the dialogue content and adjust its strategy according to the reward signal (such as user satisfaction, dialogue success rate, etc.). During the training process, the performance of the student model is continuously evaluated, and the strategy is adjusted according to the evaluation result, such as adjusting the parameters, improving the model structure or introducing new features, etc. Through multiple iterations of training, the performance of the student model is gradually optimized. After each iteration, the model is evaluated using new data, and whether to continue training or enter the next stage is determined according to the evaluation result.

[0140] Assuming in an intelligent customer service system, the student model needs to optimize the processing of the dialogue content of "user consulting product functions". First, from the reinforcement learning method library, select candidate methods related to service dialogue content, such as Q-learning, Deep Q-Network (DQN), etc. Then, according to the preliminary analysis results (such as user intent, dialogue state, etc.), evaluate the applicability of each candidate method. Finally, determine DQN as the target reinforcement learning method because it can handle high-dimensional state space and has good stability and performance. Use the preliminary analysis results as input data to train the student model based on the DQN method. During training, the model generates responses based on dialogue content and adjusts its strategy based on user satisfaction and other reward signals. Through multiple iterations of training, the performance of the student model is gradually optimized. Finally, the student model can more accurately identify user intent and dialogue state, improving the accuracy and fluency of the dialogue.

[0141] In this way, by introducing reinforcement learning methods, the student model can more accurately identify user intent and dialogue state, thereby improving the accuracy and fluency of the dialogue. The matching and parameter setting of the target reinforcement learning method enable the student model to better adapt to different dialogue scenarios and user needs, improving the model's generalization ability. Through iterative training and strategy optimization, the student model can learn effective dialogue strategies more quickly, improving learning efficiency.

[0142] Figure 4 An example diagram of a service dialogue quality inspection model is shown, as shown in Figure 4 The task description (i.e., target prompt text) and dialogue example in Figure 4 are input to the teacher model, and in the L1 stage, the output of the teacher model is as shown in Figure 4 After training of the student model imitating the content of 401, the output of the student model is as shown in Figure 4 It can be seen that the output of the student model is missing fields compared to the output of the teacher model. In the L2 stage, the output of the teacher model is as shown in Figure 4 The student model introduces a sample selection and feedback enhancement mechanism, and when dealing with complex or rare dialogue phenomena, the student model can rely on the feedback of the teacher model for targeted optimization, thereby further improving performance. The output of the student model is as shown in Figure 4 It can be seen that the output of the student model is not missing fields compared to the output of the teacher model, but the answer is not accurate. In the L3 stage, the student model no longer relies on the teacher model, but optimizes itself through its own data and dialogue feedback, realizing the complete autonomous evolution of the model. The output of the current version of the student model is as shown in Figure 4 The output of the next version of the student model is as shown in Figure 4As shown in the middle 406, the output field of the next version of the student model is complete and accurate.

[0143] The present disclosure combines the student model with the teacher model by introducing the concept of "model evolution", and gradually improves the quality inspection effect through an iterative optimization process, thereby realizing the self-optimization and evolution of the dialogue quality inspection system. Through the dynamic evolution mechanism, the student model can perform well in complex dialogue scenarios and gradually surpass the performance of the teacher model, solving the technical bottleneck that the existing quality inspection system cannot adapt to diversified needs. The gradual evolution process of the model can reduce manual intervention and improve the automation level and work efficiency of the system.

[0144] The algorithms and models of the present disclosure can be applied to multiple fields, and are particularly suitable for the following scenarios: 1. Intelligent customer service: can be used for dialogue quality in intelligent customer service systems, helping enterprises better understand customer service performance and customer feedback. 2. Customer satisfaction management: through dialogue analysis, the system can timely find and handle key factors affecting customer satisfaction, improving customer experience. 3. Dialogue analysis: can be used to analyze the dialogue content between customers and customer service, providing targeted service improvement suggestions for enterprises. 4. Multi-round dialogue quality inspection: particularly suitable for problem identification and handling in complex multi-round dialogues, which can help customer service quickly respond to complex customer needs.

[0145] The present disclosure provides a service dialogue quality inspection method based on a large model, Figure 5 is a flowchart of a service dialogue quality inspection method based on a large model according to an embodiment of the present disclosure. The service dialogue quality inspection method can be applied to electronic devices, including but not limited to fixed devices and / or mobile devices. For example, fixed devices include but are not limited to servers, which can be cloud servers or ordinary servers. For example, mobile devices include but are not limited to: mobile phones, tablet computers, vehicle-mounted devices, personal computers, etc. In some possible implementation manners, the service dialogue quality inspection method can also be realized by a processor calling computer readable instructions stored in a memory. As Figure 5 As shown, the service dialogue quality inspection method includes:

[0146] S501: Obtain service dialogue content to be inspected;

[0147] S502: Input the service dialogue content to be inspected into a service dialogue quality inspection model to obtain a service dialogue quality inspection result output by the service dialogue quality inspection model according to the target prompt text.

[0148] In an embodiment of the present disclosure, the target prompt text is a text used to guide the model to output the quality inspection result in the service dialogue quality inspection. It may contain specific rules, standards or requirements, and the model will evaluate the dialogue content according to these prompt texts.

[0149] In the disclosed embodiment, the service dialogue quality inspection model is used to automatically process and analyze the service dialogue content and output quality inspection results.

[0150] In some implementations, the content of the service dialogue to be inspected may be sourced from call records of a customer service center, chat records of an online customer service system, user feedback on a social media platform, and the like.

[0151] In some embodiments, the acquired conversation content is preprocessed, including removing irrelevant information (such as noise, advertisements, etc.), standardizing text format (such as unifying capitalization and punctuation, etc.), and natural language processing steps such as word segmentation and stop word removal.

[0152] In some embodiments, the preprocessed conversation content is used as input and transmitted to the service conversation quality inspection model through an application programming interface (API) interface or other data transmission methods. After receiving the input, the service conversation quality inspection model will parse and evaluate the conversation content based on the target prompt text, and finally output the quality inspection results.

[0153] In some embodiments, the quality inspection results generally include Figure 2 The content of the prompt text displayed may also include evaluation results on the compliance of the conversation content, service quality, customer satisfaction, etc., as well as specific violation information or improvement suggestions.

[0154] Consider a company with a customer service center that generates a large number of call logs daily. To improve quality inspection efficiency, the company decided to implement a service conversation quality inspection method. The company randomly selected 1,000 conversations from the customer service center's call logs for quality inspection and transmitted the pre-processed conversations to a model via an API. After receiving the input, the model evaluated each conversation and output quality inspection results. The results showed that 30 conversations had service quality issues, such as poor customer service attitude and inaccurate answers. Based on the quality inspection results, the company promptly implemented corresponding improvement measures, such as strengthening customer service training and optimizing service processes, thereby improving service quality and increasing customer satisfaction.

[0155] This automated service conversation quality inspection model significantly improves quality inspection efficiency and reduces the time and cost of manual quality inspections. Trained and fine-tuned with extensive data, the model accurately identifies key information and potential issues in conversations, improving the accuracy of quality inspections. Through feedback from quality inspection results, companies can promptly identify service issues and implement appropriate improvement measures, thereby optimizing service quality and increasing customer satisfaction.

[0156] In some embodiments, the service dialogue content to be inspected includes at least one of the following: content in text form, content in voice form.

[0157] In the embodiment of the present disclosure, the text content is service conversation content presented in text form, such as emails, online customer service chat records, etc.

[0158] In the embodiment of the present disclosure, the content in voice form is service conversation content presented in voice form, such as telephone call records, voice messages, etc.

[0159] In some implementations, textual service conversation content is collected from various channels (e.g., online customer service systems, email, social media, etc.). The text is cleaned to remove irrelevant characters (e.g., tags, special symbols, etc.), and natural language processing operations such as word segmentation and part-of-speech tagging are performed. The preprocessed text content is then fed into a trained service conversation quality inspection model, which then performs a quality assessment of the text based on pre-set rules or algorithms.

[0160] In some implementations, speech recognition technology is used to convert spoken content into text. This step may require the use of specialized speech recognition engines or tools. The converted text undergoes the same preprocessing as text content. The preprocessed text is then input into the service dialogue quality inspection model for quality assessment.

[0161] This automated quality inspection process significantly reduces the time and cost of manual quality inspections, improving efficiency. Leveraging advanced natural language processing technology and machine learning algorithms, the service conversation quality inspection model accurately identifies issues within conversations, improving quality inspection accuracy. This method can process both text and voice service conversations, meeting quality inspection requirements in diverse scenarios.

[0162] It should be understood that Figures 2 to 4 The schematic diagram shown is only exemplary and not restrictive, and it is scalable, and those skilled in the art can Figures 2 to 4 Various obvious changes and / or substitutions can be made to the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0163] The present disclosure provides a training device for a service dialogue quality inspection model, such as Figure 6As shown, the apparatus can comprise: comprising: a first acquisition module 601, configured to acquire a target prompt text; a first control module 602, configured to input the service dialogue content and the target prompt text into a teacher model, so that the teacher model generates a preliminary analysis result of the service dialogue content based on the target prompt text; a second control module 603, configured to input the preliminary analysis result into a student model, so that the student model learns and optimizes according to the preliminary analysis result; a determination module 604, configured to determine the student model after learning and optimization as a service dialogue quality inspection model, and the service dialogue quality inspection model is used to generate a service dialogue quality inspection result according to the target prompt text.

[0164] In some embodiments, the first acquisition module 601 comprises: an acquisition submodule, configured to acquire an initial prompt text; an adjustment submodule, configured to optimize and adjust the initial prompt text according to the demand of the service industry, to obtain the target prompt text.

[0165] In some embodiments, the target prompt text comprises key information of multiple dimensions, and the multiple dimensions comprise: question information, question state, provided solution, solution result and service evaluation.

[0166] In some embodiments, the first control module 602 comprises: an understanding submodule, configured to understand the service dialogue content by using the teacher model, to obtain the specific content, context and respective intentions and points of attention of the first object and the second object of the service dialogue content; a first control submodule, configured to generate the preliminary analysis result of the service dialogue content according to the target prompt text based on the specific content, context and respective intentions and points of attention of the first object and the second object of the service dialogue content.

[0167] In some embodiments, the first control submodule 602 is configured to: guide the teacher model to gradually infer the logic behind the question according to the target prompt text by a chain reasoning manner; and generate the preliminary analysis result of the service dialogue content according to the target prompt text in combination with the logic behind the question.

[0168] In some embodiments, the second control module 603 comprises: a second control submodule, configured to train the student model by imitating the preliminary analysis result output by the teacher model in a first stage, so as to learn to process various problems and scenes in the dialogue; a third control submodule, configured to optimize the student model for the target dialogue phenomenon by introducing a sample selection and feedback enhancement mechanism in a second stage; and a fourth control submodule, configured to perform self iteration and optimization of the student model according to its own data accumulation and feedback in the dialogue process in a third stage, so as to realize the overall autonomous evolution of the student model.

[0169] In some embodiments, the third control submodule is configured to: gradually improve the recognition and processing capability of the student model for the target dialogue phenomenon on the selected samples through multiple iterations of training; and adjust the parameters of the student model according to the rewards and penalties provided by the feedback enhancement mechanism after each iteration.

[0170] In some embodiments, the second control module 603 further includes: an evaluation submodule configured to evaluate the performance of the student model at the end of each stage; and a decision submodule configured to decide whether to enter the next stage according to the evaluation result.

[0171] In some embodiments, the second control module 603 includes: a determination submodule configured to determine a target reinforcement learning method that matches the service dialogue content from a plurality of candidate reinforcement learning methods; and a fifth control submodule configured to enable the student model to learn and optimize based on the target reinforcement learning method according to the preliminary analysis result.

[0172] The specific functions and examples of the modules and submodules of the device of the embodiments of the present disclosure are described above in the corresponding steps of the method embodiments, and will not be described here again.

[0173] The training device of the service dialogue quality inspection model of the embodiments of the present disclosure can enhance the generalization capability of the model and improve the efficiency and accuracy of service dialogue quality inspection.

[0174] The embodiments of the present disclosure provide a service dialogue quality inspection device based on a large model, as shown in Figure 7 The device can include: a second acquisition module 701 configured to acquire service dialogue content to be inspected; and a quality inspection module 702 configured to input the service dialogue content to be inspected into a service dialogue quality inspection model to obtain a service dialogue quality inspection result output by the service dialogue quality inspection model according to a target prompt text. The service dialogue quality inspection model is obtained by training the method described above.

[0175] In some embodiments, the service dialogue content to be inspected includes at least one of: content in a text form and content in a voice form.

[0176] The specific functions and examples of the modules and submodules of the device of the embodiments of the present disclosure are described above in the corresponding steps of the method embodiments, and will not be described here again.

[0177] The service dialogue quality inspection device based on a large model of the embodiments of the present disclosure can not only greatly improve the quality inspection efficiency, reduce the time and cost of manual quality inspection, but also improve the accuracy of quality inspection through the service dialogue quality inspection model.

[0178] The embodiments of the present disclosure provide a scene schematic diagram of a training method of a service dialogue quality inspection model, as shown in Figure 8 .

[0179] As previously mentioned, the training method for the service conversation quality inspection model provided in the embodiments of the present disclosure is applied to electronic devices. The term "electronic device" is intended to refer to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The term "electronic device" may also refer to various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0180] Specifically, the electronic device can perform the following operations:

[0181] Get the target prompt text;

[0182] Inputting the service conversation content and the target prompt text into the teacher model, so that the teacher model generates a preliminary analysis result of the service conversation content based on the target prompt text;

[0183] Inputting the preliminary analysis results into the student model so that the student model learns and optimizes according to the preliminary analysis results;

[0184] The student model after learning and optimization is determined as the service dialogue quality inspection model, which is used to generate service dialogue quality inspection results based on the target prompt text.

[0185] It should be understood that Figure 8 The scene diagram shown is only illustrative and not restrictive. Those skilled in the art can Figure 8 Various obvious changes and / or substitutions can be made to the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0186] The embodiment of the present disclosure provides a scenario diagram of a service dialogue quality inspection method based on a large model, such as Figure 9 shown.

[0187] As previously mentioned, the service conversation quality inspection method provided in the embodiments of the present disclosure is applied to electronic devices. The term "electronic device" is intended to refer to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The term "electronic device" may also refer to various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0188] Specifically, the electronic device can perform the following operations:

[0189] Get the content of the service dialogue to be inspected;

[0190] The service conversation quality inspection content to be inspected is input into the service conversation quality inspection model, and a service conversation quality inspection result output by the service conversation quality inspection model according to the target prompt text is obtained.

[0191] It should be understood that, Figure 9 The scene diagram shown is merely illustrative and non-limiting, and a person skilled in the art can make various obvious changes and / or replacements based on the examples Figure 9 The resulting technical solutions still belong to the disclosure range of the embodiments of the present disclosure.

[0192] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information involved are in line with relevant laws and regulations and do not violate public order and good customs.

[0193] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0194] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown in the figure, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0195] As shown in Figure 10 The device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a Read-Only Memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a Random Access Memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.

[0196] A number of the components in the device 1000 are connected to the I / O interface 1005, including an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through computer networks, such as the Internet, and / or various telecommunication networks.

[0197] The computing unit 1001 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a CPU, a Graphics Processing Unit (GPU), various special-purpose Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 performs various methods and processes described above, such as the training method of a service conversation quality inspection model and / or the large model-based service conversation quality inspection method. For example, in some embodiments, the training method of a service conversation quality inspection model and / or the large model-based service conversation quality inspection method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded to the RAM 1003 and executed by the computing unit 1001, one or more steps of the training method of a service conversation quality inspection model and / or the large model-based service conversation quality inspection method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the training method of a service conversation quality inspection model and / or the large model-based service conversation quality inspection method by any other appropriate means, such as by means of firmware.

[0198] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system on chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0199] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general or special purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / operations specified in the flowcharts and / or block diagrams. The program code can execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0200] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0201] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a Cathode Ray Tube (CRT) or Liquid Crystal Display (LCD) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0202] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.

[0203] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0204] It should be understood that the various forms of flow shown above can be re-ordered, added to, or have steps deleted, using the steps described above. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure can be achieved, which is not limited herein.

[0205] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A training method for a service dialogue quality inspection model, comprising: Get the target prompt text; Inputting the service dialogue content and the target prompt text into a teacher model, so that the teacher model generates a preliminary analysis result of the service dialogue content based on the target prompt text; The preliminary analysis results are input into the student model so that the student model learns and optimizes according to the preliminary analysis results; wherein the student model learns and optimizes according to the preliminary analysis results, including: in a first stage, the student model is trained by imitating the preliminary analysis results output by the teacher model to learn to handle various problems and scenarios in the conversation; in a second stage, the student model optimizes the target conversation phenomenon by introducing a sample selection and feedback enhancement mechanism; in a third stage, the student model performs self-iteration and optimization based on its own data accumulation and feedback during the conversation to achieve comprehensive autonomous evolution of the student model, wherein the student model optimizes the target conversation phenomenon by introducing a sample selection and feedback enhancement mechanism, including: on selected samples, through multiple iterative training, gradually improving the student model's ability to recognize and process the target conversation phenomenon; after each iteration, adjusting the parameters of the student model according to the rewards and penalties provided by the feedback enhancement mechanism; at the end of each stage, evaluating the performance of the student model, and deciding whether to enter the next stage based on the evaluation results; The student model after learning and optimization is determined as a service dialogue quality inspection model, and the service dialogue quality inspection model is used to generate a service dialogue quality inspection result based on the target prompt text.

2. The method according to claim 1, wherein The obtaining target prompt text includes: Get the initial prompt text; According to the needs of the service industry, the initial prompt text is optimized and adjusted to obtain the target prompt text.

3. The method according to claim 1 or 2, wherein: The target prompt text includes key information of multiple dimensions, including: problem information, problem status, provided solutions, solution results and service evaluation.

4. The method according to claim 1, wherein The teacher model generates preliminary analysis results of the service dialogue content based on the target prompt text, including: Using the teacher model to understand the service conversation content, obtain the specific content and context of the service conversation content, and the intentions and concerns of the first object and the second object respectively; Based on the specific content and context of the service dialogue content and the respective intentions and concerns of the first object and the second object, a preliminary analysis result of the service dialogue content is generated according to the target prompt text.

5. The method according to claim 4, wherein Generating a preliminary analysis result of the service dialogue content according to the target prompt text includes: Through chain reasoning, the teacher model is guided to gradually infer the logic behind the question according to the target prompt text; Combined with the logic behind the question, a preliminary analysis result of the service dialogue content is generated according to the target prompt text.

6. The method according to claim 1, wherein The student model learns and optimizes according to the preliminary analysis results, including: Determining a target reinforcement learning method that matches the service conversation content from a plurality of candidate reinforcement learning methods; The student model is based on the target reinforcement learning method and is learned and optimized according to the preliminary analysis results.

7. A service dialogue quality inspection method based on a large model, comprising: Get the content of the service dialogue to be inspected; The service conversation content to be quality inspected is input into a service conversation quality inspection model to obtain a service conversation quality inspection result output by the service conversation quality inspection model based on the target prompt text; wherein, the service conversation quality inspection model is obtained by training using the method described in any one of claims 1 to 6.

8. The method according to claim 7, wherein: The content of the service dialogue to be inspected includes at least one of the following: content in text form and content in voice form.

9. A training device for a service dialogue quality inspection model, comprising: The first acquisition module is used to obtain the target prompt text; a first control module, configured to input the service dialogue content and the target prompt text into a teacher model, so that the teacher model generates a preliminary analysis result of the service dialogue content based on the target prompt text; a second control module for inputting the preliminary analysis results into the student model so that the student model learns and optimizes according to the preliminary analysis results; wherein the second control module includes: a second control submodule for, in the first stage, the student model is trained by imitating the preliminary analysis results output by the teacher model to learn to handle various problems and scenarios in the conversation; a third control submodule for, in the second stage, the student model is optimized for the target conversation phenomenon by introducing a sample selection and feedback enhancement mechanism; a fourth control submodule for, in the third stage, the student model is self-iterated and optimized based on its own data accumulation and feedback during the conversation to achieve comprehensive autonomous evolution of the student model; wherein the third control submodule is used to: on selected samples, through multiple iterative training, gradually improve the student model's ability to recognize and process the target conversation phenomenon; after each iteration, adjust the parameters of the student model according to the rewards and penalties provided by the feedback enhancement mechanism; evaluate the performance of the student model at the end of each stage, and decide whether to enter the next stage based on the evaluation results; A determination module is used to determine the student model after learning and optimization as a service dialogue quality inspection model, and the service dialogue quality inspection model is used to generate a service dialogue quality inspection result based on the target prompt text.

10. The device according to claim 9, wherein The first acquisition module includes: Get submodule, used to get the initial prompt text; The adjustment submodule is used to optimize and adjust the initial prompt text according to the needs of the service industry to obtain the target prompt text.

11. The device according to claim 9 or 10, wherein: The target prompt text includes key information of multiple dimensions, including: problem information, problem status, provided solutions, solution results and service evaluation.

12. The device according to claim 9, wherein The first control module includes: An understanding submodule, configured to use the teacher model to understand the service dialogue content, and obtain the specific content and context of the service dialogue content, as well as the intentions and concerns of the first and second objects; The first control submodule is configured to generate a preliminary analysis result of the service dialogue content according to the target prompt text based on the specific content and context of the service dialogue content and the intentions and concerns of the first object and the second object respectively.

13. The device according to claim 12, wherein The first control submodule is configured to: Through chain reasoning, the teacher model is guided to gradually infer the logic behind the question according to the target prompt text; Combined with the logic behind the question, a preliminary analysis result of the service dialogue content is generated according to the target prompt text.

14. The device according to claim 9, wherein The second control module includes: a determination submodule, configured to determine a target reinforcement learning method that matches the service conversation content from a plurality of candidate reinforcement learning methods; The fifth control submodule is used for the student model to learn and optimize according to the preliminary analysis results based on the target reinforcement learning method.

15. A service dialogue quality inspection device based on a large model, comprising: The second acquisition module is used to obtain the content of the service dialogue to be inspected; A quality inspection module is used to input the service dialogue content to be quality inspected into a service dialogue quality inspection model to obtain the service dialogue quality inspection result output by the service dialogue quality inspection model based on the target prompt text; wherein, the service dialogue quality inspection model is obtained by training using the method described in any one of claims 1 to 6.

16. The device according to claim 15, wherein The content of the service dialogue to be inspected includes at least one of the following: content in text form and content in voice form.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are for causing a computer to execute the method according to any one of claims 1-8.

19. A computer program product comprising a computer program stored on a storage medium, the computer program implementing the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Natural language processing task execution method and device, natural language processing task model training method and device, and medium

    CN118536605A

  • Intelligent voice customer service quality inspection method and device based on multi-modal large model

    CN118631939A