An electric pin dialog evaluation weight calibration method and system

CN122820005APending Publication Date: 2026-09-25BEIJING WEIJUZHIHUI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611077237.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-20
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

并未充分解决评价权重可信度问题

Benefits of technology

[0020]大语言模型自动生成评价维度、评分、证据和置信度,固定人工评分表滞后,维度覆盖不足,提高维度发现能力,并为评分提供可复核证据。减少人为指定权重带来的主观性。客户先验转化模型,分离客户基础转化倾向,减少名单质量偏差。在相似客户先验条件下评价坐席对话行为贡献,降低名单质量、客户天然意向和活动批次对评分权重的干扰。残差建模、倾向评分或因果校准,直接用转化标签训练权重会把混杂因素学进去,得到更接近坐席对话行为净贡献的权重。既能利用大语言模型自动生成评价维度和证据,又能基于真实转化标签学习可信权重,并能控制客户先验和其它混杂因素影响。权重可信度和稳定性筛选,避免人工权重缺少依据,自动维度可能不稳定;形成具有样本覆盖率、置信度和版本记录的净贡献权重。净贡献权重具有可追溯的数据来源和可复核的计算依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820005A_ABST
    Figure CN122820005A_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a method and system for calibrating evaluation weight of electric sales dialogue, the method comprises the following steps: constructing a training sample set of electric sales, reading the transcription text of the dialogue data of each sample through a large language model, automatically generating a candidate evaluation dimension set, for each sample, outputting a score, an evidence fragment, an evidence strength, a score confidence and an abnormality label on each evaluation dimension of the evaluation dimension set; obtaining the customer basic conversion probability of each dialogue based on the customer prior data, the agent and the operation data of each sample, constructing a decontamination contribution model based on the conversion label of each sample, the customer basic conversion probability, and the evaluation dimension score, the evidence strength and the score confidence output by the large language model, determining the net contribution weight of each evaluation dimension to the conversion label; retaining the evaluation dimension and the corresponding net contribution weight that meet the condition. The evaluation weight in the quality inspection score can have traceable data sources and reviewable calculation basis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent quality inspection and conversion feedback modeling for telemarketing dialogues, specifically to a method and system for calibrating evaluation weights in telemarketing dialogues. Background Technology

[0002] Enterprise telemarketing systems typically accumulate call recordings, speech-to-text transcripts, speaker separation results, customer information, list sources, agent information, campaign batches, reach records, conversion results, and manual quality inspection results. Traditional quality inspection mainly relies on manual sampling, keyword rules, fixed scoring sheets, or traditional natural language processing models. While these methods can identify some compliance or service issues, they are insufficient in evaluating whether agent conversational behavior effectively drives conversions.

[0003] In a fixed scoring system, the weights of each evaluation dimension are typically determined by experience. For example, the weighting of dimensions such as opening, needs assessment, selling point introduction, objection handling, and follow-up actions often relies on expert judgment or management experience. While this fixed scoring system is easy to implement, the lack of verification of actual conversion results for the weighting sources makes it easy for situations to arise where a high score doesn't necessarily translate to a mediocre conversion, or where actions truly crucial to conversion receive insufficient weight in the scoring.

[0004] Large language models can perform semantic understanding, summarization, behavior recognition, objection identification, and structured scoring on telemarketing conversations, providing new technical conditions for automated quality inspection. However, if the large language model is only responsible for providing multi-dimensional scores, while the weights of each dimension are still manually specified, the quality inspection score still lacks quantifiable basis. If the weights are directly learned from the conversion results, they are easily affected by factors such as the customer's natural intent, the quality of the contact list, the source of the channel, the activity batch, and the seat allocation strategy, causing the model to misjudge customers who are inherently easy to convert as contributing significantly to a certain conversation dimension.

[0005] Existing technologies typically focus on predicting whether a customer will convert or identifying the differences between successful and failed sessions, rather than calculating the net contribution of each conversation evaluation dimension to the conversion outcome after controlling for customer preconditions. This fails to adequately address the issue of the reliability of evaluation weights. Summary of the Invention

[0006] This invention provides a method and system for calibrating the evaluation weight of telemarketing conversations, which can solve the above-mentioned technical problems in the prior art.

[0007] To achieve the above objectives, in one aspect, embodiments of the present invention provide a method for calibrating the evaluation weights of telemarketing conversations, comprising:

[0008] Construct a training sample set for telemarketing, with each sample including dialogue data and conversion tags;

[0009] The system reads the transcribed text of each sample's dialogue data using a large language model, automatically generates a set of candidate evaluation dimensions, and outputs a score, evidence fragment, evidence strength, score confidence, and anomaly marker for each sample on each evaluation dimension of the set.

[0010] For each sample, the customer prior data, agent data, and operational data are used to obtain the basic customer conversion probability of each conversation. The basic customer conversion probability is used to describe the customer conversion tendency of the sample without using the current conversation data.

[0011] Based on the conversion tag, customer base conversion probability, and the evaluation dimension scores, evidence strength, and score confidence output by the large language model for each sample, a confounding contribution model is constructed, and the net contribution weight of each evaluation dimension to the conversion tag is determined according to the confounding contribution model.

[0012] The evaluation dimensions are filtered, and only those that meet the criteria and their corresponding net contribution weights are retained.

[0013] On the other hand, embodiments of the present invention provide a telemarketing dialogue evaluation weight calibration system, comprising:

[0014] The data acquisition module is used to build a training sample set for telemarketing, with each sample including dialogue data and conversion tags;

[0015] The large language model structuring module is used to read the transcribed text of each sample's dialogue data through the large language model, automatically generate a set of candidate evaluation dimensions, and output a score, evidence fragment, evidence strength, score confidence, and anomaly marker for each sample on each evaluation dimension of the evaluation dimension set.

[0016] The customer prior conversion module is used to obtain the basic customer conversion probability of each dialogue based on the customer prior data, agent data and operational data of each sample. The basic customer conversion probability is used to describe the customer conversion tendency of the sample without using the current dialogue data.

[0017] The decontamination contribution module is used to construct a decontamination contribution model based on the conversion tag, customer base conversion probability, and evaluation dimension scores, evidence strength, and score confidence output by the large language model for each sample. The net contribution weight of each evaluation dimension to the conversion tag is determined based on the decontamination contribution model.

[0018] The evaluation weight generation module is used to filter evaluation dimensions and retain the evaluation dimensions that meet the conditions and their corresponding net contribution weights.

[0019] The above technical solution has the following beneficial effects:

[0020] The large language model automatically generates evaluation dimensions, scores, evidence, and confidence levels, eliminating the lag and insufficient dimensional coverage of manual scoring sheets, improving dimensional discovery capabilities, and providing verifiable evidence for scoring. It reduces the subjectivity introduced by manually assigned weights. A customer prior conversion model separates customer base conversion tendencies, reducing list quality bias. It evaluates the contribution of agent dialogue behavior under similar customer prior conditions, reducing the interference of list quality, customer natural intent, and activity batches on scoring weights. Residual modeling, propensity scoring, or causal calibration directly trains weights using conversion labels, incorporating confounding factors to obtain weights that more closely approximate the net contribution of agent dialogue behavior. It can automatically generate evaluation dimensions and evidence using the large language model, learn credible weights based on real conversion labels, and control for the influence of customer priors and other confounding factors. Weight credibility and stability screening avoids the lack of basis for manual weights and the potential instability of automatically generated dimensions; it forms net contribution weights with sample coverage, confidence levels, and version records. Net contribution weights have traceable data sources and verifiable calculation basis. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a method for calibrating the evaluation weight of telemarketing dialogue according to an embodiment of the present invention;

[0023] Figure 2 This is a structural diagram of a telemarketing dialogue evaluation weight calibration system according to an embodiment of the present invention;

[0024] Figure 3 This is a system overall architecture diagram according to an embodiment of the present invention;

[0025] Figure 4 This is a flowchart of the method according to an embodiment of the present invention;

[0026] Figure 5 This is a flowchart of the decontamination contribution modeling process according to an embodiment of the present invention;

[0027] Figure 6 This is a flowchart of the continuous calibration closed-loop process according to an embodiment of the present invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] like Figure 1 As shown, in conjunction with embodiments of the present invention, a method for calibrating the evaluation weight of telemarketing conversations is provided, comprising:

[0030] S110: Construct a training sample set for telemarketing, with each sample including dialogue data and conversion tags;

[0031] S120: Read the transcribed text of the dialogue data of each sample through a large language model, automatically generate a set of candidate evaluation dimensions, and output the score, evidence fragment, evidence strength, score confidence and anomaly mark for each sample on each evaluation dimension of the evaluation dimension set.

[0032] S130: For each sample, the customer prior data, agent data and operational data are used to obtain the customer basic conversion probability of each dialogue. The customer basic conversion probability is used to describe the customer conversion tendency of the sample without using the current dialogue data.

[0033] S140: Based on the conversion tag, customer base conversion probability, and the evaluation dimension scores, evidence strength, and score confidence output by the large language model for each sample, construct a decontamination contribution model, and determine the net contribution weight of each evaluation dimension to the conversion tag according to the decontamination contribution model.

[0034] S150: Filter the evaluation dimensions and retain the evaluation dimensions that meet the conditions and their corresponding net contribution weights.

[0035] In this embodiment of the invention, the large language model automatically generates evaluation dimensions, scores, evidence, and confidence levels, fixing the lag and insufficient dimension coverage of manual scoring tables, improving dimension discovery capabilities, and providing verifiable evidence for scoring. This reduces the subjectivity caused by manually assigning weights.

[0036] A customer prior conversion model separates customer base conversion tendencies, reducing list quality bias. It evaluates the contribution of agent dialogue behavior under similar customer prior conditions, minimizing the interference of list quality, customer natural intent, and activity batches on scoring weights. Residual modeling, propensity scoring, or causal calibration are used; directly training weights with conversion labels incorporates confounding factors, resulting in weights that more closely approximate the net contribution of agent dialogue behavior. This model can automatically generate evaluation dimensions and evidence using a large language model, learn credible weights based on real conversion labels, and control for the influence of customer priors and other confounding factors. Weight credibility and stability screening avoids the lack of basis for manual weights and the potential instability of automatically generated dimensions; it forms net contribution weights with sample coverage, confidence levels, and version records. Net contribution weights have traceable data sources and verifiable calculation basis.

[0037] Preferably, S110: Constructing a training sample set for telemarketing, including:

[0038] Acquire telemarketing data. Each piece of telemarketing data includes dialogue data, customer prior data, agent and operational data, and conversion tags. Each piece of telemarketing data serves as a sample. Dialogue data includes call recordings, transcribed text after speech recognition, speaker roles, call duration, call rounds, and call time. Customer prior data includes channel source, list batches, historical behavior, historical outreach, and customer segmentation. Agent and operational data includes agent identifiers, teams, experience levels, shifts, activity batches, and allocation strategies. Conversion tags are obtained by setting attribution windows based on business configurations and mapping target events within those windows. Conversion tags include multi-level conversion tags, conversion time, conversion stage, target event, and completion stage.

[0039] By using a preset granularity, the dialogue data, customer prior data, agent and operational data, and conversion tags in the telemarketing data are associated and aligned according to customer identifier, call identifier, agent identifier, batch identifier, and timestamp to obtain a training sample set.

[0040] Dialogue data is used by the large language model to generate evaluation dimensions, scores, evidence, and confidence levels; prior customer data is used to estimate the customer's basic conversion tendency without considering the content of the current dialogue; agent and operational data are used as control variables for confounding factors or scenario correction variables. Conversion tags represent the actual conversion result of one or more target events completed within the attribution window, including multi-level conversion tags, conversion time, conversion stage, and whether the target event was completed. These tags record whether the customer completed the target event within the attribution window and to which stage; they also serve as targets for supervised learning, causal calibration, or weight estimation. Conversion tags can be multi-level tags. For example, events such as successful appointment, document submission, application processing, credit approval, order placement, payment, loan disbursement, repeat purchase, and renewal can be single-level or multi-level. For instance, the first level could indicate whether the customer has completed initial intent confirmation, the second level could indicate whether document submission or application processing has been completed, and the third level could indicate whether the final target event has been completed. The above levels are just examples; the actual levels and event names are determined by business configuration.

[0041] The attribution window refers to the time frame, starting from the end of a single or group of telemarketing conversations, used to determine whether subsequent conversion events belong to that conversation or group of conversations. It aligns conversation samples with conversion tags and can be flexibly configured based on business, customer group, product, and channel. It can determine whether a specific conversion tag belongs to a preceding telemarketing conversation. For example, starting with the end time of a call, the end time of a customer session, or the last contact time, it checks whether a target event occurred within a preset time frame; if so, the target event is used as the conversion tag for the corresponding sample. The attribution window can be configured based on business objectives, product characteristics, customer group characteristics, or channel characteristics.

[0042] The granularity of training samples can be a single valid call, a single customer session, or an aggregated record of multiple outreaches to the same customer within the attribution window. During sample alignment, conversation data, prior customer data, agent and operational data, and conversion tags are associated according to customer identifier, call identifier, agent identifier, batch identifier, and timestamp. This enables the association of data from different sources by customer, call, agent, batch, and time.

[0043] Preferably, S120: The transcribed text of the dialogue data for each sample is read using a large language model, and a candidate evaluation dimension set is automatically generated. For each sample, a score, evidence fragment, evidence strength, score confidence, and anomaly marker are output for each evaluation dimension in the evaluation dimension set, including:

[0044] For each sample in the training sample set, the transcribed text corresponding to its dialogue data is input into the large language model, or the text communication record is directly input. The large language model identifies semantic units in the dialogue data. The semantic units include agent behavior, customer response, change of intent, objection type, advancement action, and information disclosure. When conditions permit, speaker separation results, call time, call rounds, business objective description, historical sample summary, and necessary business constraints can also be input simultaneously.

[0045] Candidate evaluation dimensions are automatically generated based on semantic units, and each candidate evaluation dimension is generated with a name, definition, positive examples, negative examples, scoring rules, and evidence extraction rules.

[0046] The candidate evaluation dimensions are semantically clustered, deduplicated, merged, and hierarchically processed to form the set of evaluation dimensions for the current period.

[0047] For each sample of dialogue data, output the score, evidence fragment, evidence strength, score confidence, and anomaly marker for each evaluation dimension in the evaluation dimension set.

[0048] Telemarketing conversations refer to call or text communication records between agents and customers regarding marketing, sales, conversion, or service goals. These may include recordings, transcripts, speaker separation results, and call metadata. Evaluation dimensions are structured evaluation items used to describe agent behavior, customer reactions, communication progress, and objection handling in the telemarketing conversation. Evidence fragments are fragments of the original conversation or their location identifiers that support a score for a specific evaluation dimension. They are used to verify the scoring basis and reduce purely black-box scoring. This embodiment of the invention does not require all evaluation dimensions to be fixed in advance; instead, the large language model automatically generates candidate evaluation dimensions based on conversation data, business objectives, and historical feedback. Unlike fixed manual scoring tables, the large language model's structuring module transforms unstructured telemarketing conversations into modelable structured data, ensuring that subsequent weight learning is not based on simple keywords or human impressions, but on verifiable conversational behavior features.

[0049] Preferably, in S130, the customer prior conversion module is used to estimate the basic conversion probability of a customer based on prior customer data and agent and operational data, without considering the specific dialogue data of this conversation. Prior customer data, agent and operational data are used to describe the customer's basic conversion tendency and business allocation conditions (i.e., agent and operational data). Evaluation dimensions generated from the content of this conversation are not used to reduce information leakage. This is to distinguish between the potential conversion of the customer and the incremental impact of agent dialogue behavior. Specifically, it is expressed as: P0=F(X_prior,X_channel,X_batch,X_agent_base,X_operation), where P0 represents the basic conversion probability of the customer; X_prior represents the customer's prior characteristics; X_channel represents the channel characteristics; X_batch represents the list or activity batch characteristics; X_agent_base represents the agent or team basic variables; and X_operation represents the historical contact time and operational strategy. Customer prior conversion models can employ logistic regression, gradient boosting trees, generalized linear models, random forests, neural networks, calibration models, hierarchical models, or other models suitable for probability estimation.

[0050] Preferably, S140: Based on the conversion tag, customer base conversion probability, and evaluation dimension scores, evidence strength, and score confidence output by the large language model for each sample, a decontamination contribution model is constructed. The net contribution weight of each evaluation dimension to the conversion tag is determined according to the decontamination contribution model, including:

[0051] For each sample, the actual conversion result is compared with the customer's basic conversion probability to obtain the conversion residual. The conversion residual represents the portion of the actual conversion result that is not explained by the customer's prior variables and scenario variables. It is used to estimate the incremental contribution of this dialogue behavior to the target conversion. The customer's prior variables come from customer prior data, and the scenario variables come from agent and operational data. For example, if a customer's conversion probability was low under the customer's prior conditions but actually converted, then the sample forms a positive incremental signal; if a customer's conversion probability was high under the customer's prior conditions but did not actually convert, then the sample forms a negative incremental signal.

[0052] The output of the large language model on each evaluation dimension, including the score, evidence strength, score confidence, and scenario variables, are used as input variables. The conversion residual or actual conversion result is used as the learning objective to construct and train the dialogue behavior contribution model. The actual conversion result can be binary classification, multi-class classification, ordinal label, continuous benefit label, or multi-stage label.

[0053] Each sample's dialogue data is input into the dialogue behavior contribution model, which outputs the model parameters, contribution values, feature importance, or equivalent interpretable outputs corresponding to each evaluation dimension. The contribution value represents the degree of influence of a certain evaluation dimension on the conversion label.

[0054] The net contribution weight of each evaluation dimension is determined based on the model parameters, contribution value, or feature importance corresponding to each evaluation dimension.

[0055] Confounding factors refer to factors that simultaneously affect both dialogue performance and conversion results, but are not the target dialogue behavior itself. After controlling for customer priors and confounding factors, a deconfounding contribution modeling module is used to determine the net contribution of evaluation dimensions to conversion tags, such as... Figure 5 As shown. Net contribution weight represents the incremental direction and magnitude of this evaluation dimension's impact on conversion results, considering prior customer data and business allocation conditions (agent and operational data). Net contribution weight can be output as a single evaluation dimension or as a semantically merged family of dimensions.

[0056] Preferably, S140: Based on the conversion tag, customer base conversion probability, and evaluation dimension scores, evidence strength, and score confidence output by the large language model for each sample, a decontamination contribution model is constructed. The net contribution weight of each evaluation dimension to the conversion tag is determined according to the decontamination contribution model, including:

[0057] The corresponding dialogue behaviors of a certain type of agent behavior are regarded as dialogue behaviors to be evaluated.

[0058] For the dialogue behavior to be evaluated, the corresponding customer prior variables and scenario variables are regarded as confounding variables that need to be controlled, and the conversion tag in the attribution window where the dialogue behavior to be evaluated is located is regarded as the result variable. The customer prior variables come from customer prior data, and the scenario variables come from agent and operation data.

[0059] Based on the corresponding customer prior variables and scenario variables, control samples with similar conditions are extracted from the telemarketing data; specifically, the degree of tendency to accept the dialogue behavior is estimated for the sample based on the confounding variables, and then control samples with similar customer prior conditions are constructed through matching, weighting or dual robust estimation.

[0060] In the control sample, the differences in conversion results or conversion rates between sample groups possessing the dialogue behavior to be evaluated and sample groups not possessing the dialogue behavior to be evaluated are compared, and the net contribution weight of the evaluation dimension is estimated based on the differences in conversion results or conversion rates. The differences in conversion results or conversion rates characterize the incremental conversion contribution of the dialogue behavior to be evaluated under similar customer prior conditions and scenario conditions, and are used to estimate the net contribution weight of the corresponding evaluation dimension. The differences in conversion results or conversion rates represent the incremental conversion contribution of the evaluation dimension or dimension family under similar customer prior conditions. The evaluation dimension is used for the net contribution weight of online quality inspection scoring.

[0061] Preferably, S140: Based on the conversion tag, customer base conversion probability, and evaluation dimension scores, evidence strength, and score confidence output by the large language model for each sample, a decontamination contribution model is constructed. The net contribution weight of each evaluation dimension to the conversion tag is determined according to the decontamination contribution model, including:

[0062] Three types of variables are simultaneously input into the joint constraint model. The first type of variable is the customer prior variable, which comes from the customer prior data. The second type of variable is the scenario variable, which comes from the agent and operation data. The third type of variable is the evaluation dimension variable, which includes the rating, evidence strength, rating confidence, behavioral label and customer reaction label output by the large language model.

[0063] By using customer prior variables and scenario variables as basic conversion contribution items, and evaluation dimension variables as dialogue behavior contribution items, the net contribution weight of each evaluation dimension is obtained.

[0064] To avoid miscounting customer prior contributions as dialogue behavior contributions, independent parameter groups, grouping regularization constraints, monotonic constraints, hierarchical parameters, or interpretable decomposition structures can be set for different variable groups. The joint constraint model output can separately provide the customer's basic conversion tendency, the impact of operational scenarios, and the impact of dialogue behavior, thereby extracting the net contribution weight of the evaluation dimension from the dialogue behavior impact component. The relationship between the above three types of variables and the model output is shown in Table 1.

[0065] Table 1

[0066]

[0067] Preferably, S140: Based on the conversion tag, customer base conversion probability, and evaluation dimension scores, evidence strength, and score confidence output by the large language model for each sample, a decontamination contribution model is constructed. The net contribution weight of each evaluation dimension to the conversion tag is determined according to the decontamination contribution model, including:

[0068] For each sample, the actual conversion result is compared with the customer's basic conversion probability to obtain the conversion residual. The conversion residual represents the part of the actual conversion result that is not explained by the customer's prior variables and scenario variables. It is used to estimate the incremental contribution of this dialogue behavior to the target conversion. The customer's prior variables come from the customer's prior data, and the scenario variables come from the agent and operation data.

[0069] The output of the large language model on each evaluation dimension, the score, evidence strength, score confidence, and scene variables are used as input variables. The conversion residual or actual conversion result is used as the learning objective to construct and train the dialogue behavior contribution model.

[0070] Each sample's dialogue data is input into the dialogue behavior contribution model, which then outputs the model parameters, contribution values, feature importance, or equivalent interpretable outputs corresponding to each evaluation dimension.

[0071] The net contribution weight of each evaluation dimension is determined based on the model parameters, contribution value, or feature importance corresponding to each evaluation dimension.

[0072] The net contribution weight of each evaluation dimension output by the dialogue behavior contribution model is used as the initial net contribution weight.

[0073] The corresponding dialogue behaviors of a certain type of agent behavior are regarded as dialogue behaviors to be evaluated.

[0074] For the dialogue behavior to be evaluated, the corresponding customer prior variables and scenario variables are regarded as confounding variables that need to be controlled, and the conversion tag in the attribution window where the dialogue behavior to be evaluated is located is regarded as the result variable. The customer prior variables come from customer prior data, and the scenario variables come from agent and operation data.

[0075] Based on the corresponding prior customer variables and scenario variables, extract similar control samples from the telemarketing data;

[0076] In the control sample, the differences in conversion results or conversion rates between the sample group with the dialogue behavior to be evaluated and the sample group without the dialogue behavior to be evaluated are compared.

[0077] For evaluation dimensions where the net contribution weight reaches the preset quality threshold, the incremental contribution is reviewed by the difference in conversion results or the difference in conversion rate to obtain the incremental conversion contribution.

[0078] Four types of variables are simultaneously input into the joint constraint model. The first type of variable is the customer prior variable, which comes from the customer prior data. The second type of variable is the scenario variable, which comes from the agent and operation data. The third type of variable is the evaluation dimension variable, which includes the rating, evidence strength, rating confidence, behavioral label and customer reaction label of the evaluation dimension output by the large language model. The fourth type of variable is the incremental conversion contribution.

[0079] Customer prior variables and scenario variables are used as basic conversion contribution items, and evaluation dimension variables are used as dialogue behavior contribution items. The incremental conversion contribution is used to calibrate the dialogue behavior contribution items to obtain the net contribution weight of each evaluation dimension.

[0080] The three methods described above can be used individually or in combination. When used in combination, a two-stage residual model is employed to obtain the initial net contribution weights; for evaluation dimensions with high contributions or significant controversy, propensity score scoring or causal calibration methods are used to verify their incremental contributions; the verified evaluation dimensions and net contribution weights are then incorporated into the joint constraint model and the subsequent evaluation weight generation module. Alternatively, the joint constraint model can be used to complete the overall decomposition first, followed by causal calibration verification of key evaluation dimensions. The goal is to output the net contribution weights of the evaluation dimensions to the conversion label after controlling for or reducing the influence of confounding factors.

[0081] Preferably, S150: Filter the evaluation dimensions, retaining the evaluation dimensions that meet the conditions and their corresponding net contribution weights, including:

[0082] For each evaluation dimension or family of evaluation dimensions, calculate sample coverage, score variance, evidence confidence, stability across time windows, and correlation stability with transformation residuals.

[0083] For each evaluation dimension, the sample coverage rate is used to determine whether the evaluation dimension appears in a sufficient number of samples; if the number of samples containing valid evidence fragments of the evaluation dimension is lower than the sample number threshold, the sample coverage of the evaluation dimension is determined to be insufficient.

[0084] The rating variance is used to determine whether the evaluation dimension has the ability to distinguish different dialogue qualities; if the rating variance is less than the variance threshold, then the evaluation dimension cannot distinguish the quality of the dialogue.

[0085] The confidence level of evidence is used to determine whether the score is supported by sufficient evidence; if the confidence level of evidence is less than the confidence level threshold, the scoring basis for that evaluation dimension is invalid.

[0086] The stability of the evaluation dimension across time windows is used to determine whether it performs consistently across different periods. If the contribution direction of the evaluation dimension is inconsistent across different time windows, or if the fluctuation of the net contribution weight exceeds the weight fluctuation threshold, the stability of the evaluation dimension across time windows is deemed insufficient. An example of insufficient stability in relation to the conversion residual is: assuming that a certain evaluation dimension performs a positive contribution in some time windows and a negative contribution in other time windows, or if the contribution size fluctuates beyond a preset range.

[0087] The correlation stability with conversion residuals is used to determine whether the contribution direction and net contribution weight of the evaluation dimension are stable after controlling for customer prior variables and scenario variables; if the absolute value of the net contribution weight is less than the weight threshold, the contribution of the evaluation dimension is determined to be insufficient.

[0088] Evaluation dimensions with sample coverage below the coverage threshold, evaluation dimensions with evidence confidence below the confidence threshold, evaluation dimensions that highly overlap with other evaluation dimensions, or evaluation dimensions with an absolute value of net contribution weight less than the weight threshold are removed to obtain the retained evaluation dimensions and their corresponding net contribution weights.

[0089] For the retained evaluation dimensions, the net contribution weight, contribution direction, confidence interval, or correlation stability level with the transformation residual is output, and a candidate weight set is formed based on the net contribution weight, contribution direction, and correlation stability level with the transformation residual.

[0090] The candidate weights in the candidate weight set are normalized, truncated, smoothed, or hierarchically corrected to obtain the net contribution weights.

[0091] Normalization is used to convert the net contribution weights of each evaluation dimension to a unified dimension; truncation is used to limit abnormally large net contribution weights; hierarchical correction is used to combine global and local weights in different customer groups, channels, or business scenarios. Global weights refer to the evaluation dimension weights calculated uniformly based on a larger range of historical samples, rather than being trained separately for a specific business scenario, customer group, channel, or agent. They reflect the average net contribution of each evaluation dimension to the conversion result within the overall sample scope. Table 2 shows an example of global weights calculated based on all telemarketing samples.

[0092] Table 2 Global weights calculated based on all telemarketing samples

[0093]

[0094] Smoothing refers to the practice of not directly using the estimated contribution weight of a certain evaluation dimension on the current small sample size when the sample coverage of that evaluation dimension is insufficient or the net contribution weight fluctuates significantly. Instead, the current estimated contribution weight is combined with the global weight, the weight of similar customer groups, or the historical weight in a proportional manner to form the quality inspection weight. For example, when the sample size in the current period is small, the global weight or historical weight can be used more. As the sample coverage, evidence confidence, and stability across time windows of the evaluation dimension improve, the proportion of the estimated weight in the current period is gradually increased. Thus, the quality inspection weight can gradually stabilize with the accumulation of data. Historical weights refer to the net contribution weights of evaluation dimensions that the system has calculated and published in previous training periods, previous versions, or previous time windows. Examples are shown in Table 3, where V3 is the current online version, and V1 and V2 can be used as historical weights.

[0095] Table 3 Examples of Historical Weights

[0096]

[0097] Preferably, the telemarketing dialogue evaluation weight calibration method further includes S160:

[0098] The retained evaluation dimensions and their corresponding net contribution weights will be published to the online quality inspection system or the agent performance system.

[0099] For new telemarketing dialogues, the large language model outputs scores, evidence fragments, and score confidence for each evaluation dimension.

[0100] Based on the aforementioned scores, net contribution weights, and score confidence levels, the weighted contribution values ​​of each evaluation dimension in the newly added telemarketing dialogue are calculated. Then, based on the weighted contribution values ​​of each evaluation dimension, the effective score of the new telemarketing dialogue is calculated. The effective score, along with the score, score confidence level, weighted contribution value, and corresponding evidence fragments for each evaluation dimension, as well as the current weight version, score confidence level, agent improvement suggestions, and quality inspection review entry point, are output. This can be used for quality inspection review, agent training, performance evaluation assistance, and script improvement. Real-time feedback is also provided.

[0101] The effective score is calculated as follows: EffectiveScore=H(ΣW_k×S_k×C_k, scenario correction parameter, version parameter), where S_k is the score of the k-th evaluation dimension, C_k is the corresponding evidence confidence or score confidence, W_k is the net contribution weight, and H is the normalization, segmentation mapping, business threshold or other score transformation function.

[0102] First, the scores for each evaluation dimension are combined with the net contribution weight and the confidence level of the evidence to obtain the weighted contribution of that evaluation dimension. Then, the weighted contributions of each evaluation dimension are summarized and mapped according to the current scenario correction parameters and version parameters to form an effective online score that can be reviewed by quality inspectors. Compared with the weight of human experience, the net contribution weight calibrated by conversion feedback in this embodiment of the invention can better reflect the relationship between each dialogue behavior and the actual conversion effect.

[0103] Preferably, such as Figure 6 As shown, the method for calibrating the evaluation weights of telemarketing conversations also includes S170:

[0104] We continuously collect new dialogue data, new customer prior data, new agent data, and new conversion tags to form new training samples; and align the new conversion tags with the corresponding dialogue samples according to the attribution window.

[0105] When the number of new training samples reaches a preset amount, a new round of evaluation dimension and weight generation is triggered; or

[0106] When the differences between new dialogue data, new customer prior data, new agent data, and new conversion tags in the newly added training samples and historical periodic samples exceed their respective thresholds, a new round of evaluation dimension and weight generation is triggered; or

[0107] When the preset update cycle is reached, a new round of evaluation dimension and weight generation is triggered; or

[0108] When the rating calibration error, which measures the consistency between the effective rating and the actual conversion result, exceeds the error threshold, a new round of evaluation dimension generation and weight generation is triggered.

[0109] The stability and calibration effect of the new and old weight versions are compared in historical playback samples, grayscale samples, or online samples. Once the new version meets the preset conditions, it is released to the quality inspection system and the agent performance system. This allows the continuously increasing real business samples to be transformed into more accurate weight estimates, gradually strengthening the basis for quality inspection scoring and thus improving the effectiveness of quality inspection review, agent training, and performance evaluation.

[0110] like Figure 2 As shown, in conjunction with embodiments of the present invention, a telemarketing dialogue evaluation weight calibration system is provided, comprising:

[0111] The data acquisition module 21 is used to build a training sample set for telemarketing, and each sample includes dialogue data and conversion tags;

[0112] The large language model structuring module 22 is used to read the transcribed text of the dialogue data of each sample through the large language model, automatically generate a set of candidate evaluation dimensions, and output the score, evidence fragment, evidence strength, score confidence and anomaly label for each sample on each evaluation dimension of the evaluation dimension set.

[0113] The customer prior conversion module 23 is used to obtain the basic customer conversion probability of each dialogue based on the customer prior data, agent and operation data of each sample. The basic customer conversion probability is used to describe the customer conversion tendency of the sample without using the current dialogue data.

[0114] The decontamination contribution module 24 is used to construct a decontamination contribution model based on the conversion tag, customer base conversion probability, and evaluation dimension scores, evidence strength and score confidence output by the large language model for each sample, and to determine the net contribution weight of each evaluation dimension to the conversion tag based on the decontamination contribution model.

[0115] The evaluation weight generation module 25 is used to filter the evaluation dimensions and retain the evaluation dimensions that meet the conditions and their corresponding net contribution weights.

[0116] Preferably, the data acquisition module 21 is specifically used for:

[0117] Acquire telemarketing data. Each piece of telemarketing data includes dialogue data, customer prior data, agent and operational data, and conversion tags. Each piece of telemarketing data serves as a sample. Dialogue data includes call recordings, transcribed text after speech recognition, speaker roles, call duration, call rounds, and call time. Customer prior data includes channel source, list batches, historical behavior, historical outreach, and customer segmentation. Agent and operational data includes agent identifiers, teams, experience levels, shifts, activity batches, and allocation strategies. Conversion tags are obtained by setting attribution windows based on business configurations and mapping target events within those windows. Conversion tags include multi-level conversion tags, conversion time, conversion stage, target event, and completion stage.

[0118] By using a preset granularity, the dialogue data, customer prior data, agent and operational data, and conversion tags in the telemarketing data are associated and aligned according to customer identifier, call identifier, agent identifier, batch identifier, and timestamp to obtain a training sample set.

[0119] Preferably, the large language model structuring module 22 is used for:

[0120] For each sample in the training sample set, the transcribed text corresponding to its dialogue data is input into the large language model. The large language model identifies semantic units in the dialogue data. The semantic units include agent behavior, customer reaction, change of intent, objection type, advancement action, and information disclosure.

[0121] Candidate evaluation dimensions are automatically generated based on semantic units, and each candidate evaluation dimension is generated with a name, definition, positive examples, negative examples, scoring rules, and evidence extraction rules.

[0122] The candidate evaluation dimensions are semantically clustered, deduplicated, merged, and hierarchically processed to form the set of evaluation dimensions for the current period.

[0123] For each sample of dialogue data, output the score, evidence fragment, evidence strength, score confidence, and anomaly marker for each evaluation dimension in the evaluation dimension set.

[0124] Preferably, the decontamination contribution module 24 includes a two-stage residual modeling sub-unit, specifically used for:

[0125] For each sample, the actual conversion result is compared with the customer's basic conversion probability to obtain the conversion residual. The conversion residual represents the part of the actual conversion result that is not explained by the customer's prior variables and scenario variables. It is used to estimate the incremental contribution of this dialogue behavior to the target conversion. The customer's prior variables come from the customer's prior data, and the scenario variables come from the agent and operation data.

[0126] The output of the large language model on each evaluation dimension, the score, evidence strength, score confidence, and scene variables are used as input variables. The conversion residual or actual conversion result is used as the learning objective to construct and train the dialogue behavior contribution model.

[0127] Each sample's dialogue data is input into the dialogue behavior contribution model, which then outputs the model parameters, contribution values, feature importance, or equivalent interpretable outputs corresponding to each evaluation dimension.

[0128] The net contribution weight of each evaluation dimension is determined based on the model parameters, contribution value, or feature importance corresponding to each evaluation dimension.

[0129] Preferably, the decontamination contribution module 24 includes a propensity score modeling subunit, specifically used for:

[0130] The corresponding dialogue behaviors of a certain type of agent behavior are regarded as dialogue behaviors to be evaluated.

[0131] For the dialogue behavior to be evaluated, the corresponding customer prior variables and scenario variables are regarded as confounding variables that need to be controlled, and the conversion tag in the attribution window where the dialogue behavior to be evaluated is located is regarded as the result variable. The customer prior variables come from customer prior data, and the scenario variables come from agent and operation data.

[0132] Based on the corresponding prior customer variables and scenario variables, extract similar control samples from the telemarketing data;

[0133] In the control sample, the difference in conversion results or conversion rate between the sample group with the dialogue behavior to be evaluated and the sample group without the dialogue behavior to be evaluated is compared, and the net contribution weight of the evaluation dimension is estimated based on the difference in conversion results or conversion rate. The difference in conversion results or conversion rate is used to characterize the incremental conversion contribution of the dialogue behavior to be evaluated under similar customer prior conditions and scenario conditions, and is used to estimate the net contribution weight of the corresponding evaluation dimension.

[0134] Preferably, the decontamination contribution module 24 includes a joint constraint model sub-unit, specifically used for:

[0135] Three types of variables are simultaneously input into the joint constraint model. The first type of variable is the customer prior variable, which comes from the customer prior data. The second type of variable is the scenario variable, which comes from the agent and operation data. The third type of variable is the evaluation dimension variable, which includes the rating, evidence strength, rating confidence, behavioral label and customer reaction label output by the large language model.

[0136] By using customer prior variables and scenario variables as basic conversion contribution items, and evaluation dimension variables as dialogue behavior contribution items, the net contribution weight of each evaluation dimension is obtained.

[0137] Preferably, the decontamination contribution module 24 includes a combined model subunit, specifically used for:

[0138] For each sample, the actual conversion result is compared with the customer's basic conversion probability to obtain the conversion residual. The conversion residual represents the part of the actual conversion result that is not explained by the customer's prior variables and scenario variables. It is used to estimate the incremental contribution of this dialogue behavior to the target conversion. The customer's prior variables come from the customer's prior data, and the scenario variables come from the agent and operation data.

[0139] The output of the large language model on each evaluation dimension, the score, evidence strength, score confidence, and scene variables are used as input variables. The conversion residual or actual conversion result is used as the learning objective to construct and train the dialogue behavior contribution model.

[0140] Each sample's dialogue data is input into the dialogue behavior contribution model, which then outputs the model parameters, contribution values, feature importance, or equivalent interpretable outputs corresponding to each evaluation dimension.

[0141] The net contribution weight of each evaluation dimension is determined based on the model parameters, contribution value, or feature importance corresponding to each evaluation dimension.

[0142] The net contribution weight of each evaluation dimension output by the dialogue behavior contribution model is used as the initial net contribution weight.

[0143] The corresponding dialogue behaviors of a certain type of agent behavior are regarded as dialogue behaviors to be evaluated.

[0144] For the dialogue behavior to be evaluated, the corresponding customer prior variables and scenario variables are regarded as confounding variables that need to be controlled, and the conversion tag in the attribution window where the dialogue behavior to be evaluated is located is regarded as the result variable. The customer prior variables come from customer prior data, and the scenario variables come from agent and operation data.

[0145] Based on the corresponding prior customer variables and scenario variables, extract similar control samples from the telemarketing data;

[0146] In the control sample, the differences in conversion results or conversion rates between the sample group with the dialogue behavior to be evaluated and the sample group without the dialogue behavior to be evaluated are compared.

[0147] For evaluation dimensions where the net contribution weight reaches the preset quality threshold, the incremental contribution is reviewed by the difference in conversion results or the difference in conversion rate to obtain the incremental conversion contribution.

[0148] Four types of variables are simultaneously input into the joint constraint model. The first type of variable is the customer prior variable, which comes from the customer prior data. The second type of variable is the scenario variable, which comes from the agent and operation data. The third type of variable is the evaluation dimension variable, which includes the rating, evidence strength, rating confidence, behavioral label and customer reaction label of the evaluation dimension output by the large language model. The fourth type of variable is the incremental conversion contribution.

[0149] By taking customer prior variables and scenario variables, and the evaluation dimension variables as dialogue behavior contribution items, and using the incremental conversion contribution to calibrate the dialogue behavior contribution items, the net contribution weight of each evaluation dimension is obtained.

[0150] Preferably, the evaluation weight generation module 25 is specifically used for:

[0151] For each evaluation dimension, the sample coverage rate is used to determine whether the evaluation dimension appears in a sufficient number of samples; if the number of samples containing valid evidence fragments of the evaluation dimension is lower than the sample number threshold, the sample coverage of the evaluation dimension is determined to be insufficient.

[0152] The rating variance is used to determine whether the evaluation dimension has the ability to distinguish different dialogue qualities; if the rating variance is less than the variance threshold, then the evaluation dimension cannot distinguish the quality of the dialogue.

[0153] The confidence level of evidence is used to determine whether the score is supported by sufficient evidence; if the confidence level of evidence is less than the confidence level threshold, the scoring basis for that evaluation dimension is invalid.

[0154] The stability of the evaluation dimension across time windows is used to determine whether it performs consistently across different periods. If the contribution direction of the evaluation dimension is inconsistent across different time windows, or if the fluctuation of the net contribution weight exceeds the weight fluctuation threshold, the stability of the evaluation dimension across time windows is deemed insufficient.

[0155] The correlation stability with conversion residuals is used to determine whether the contribution direction and net contribution weight of the evaluation dimension are stable after controlling for customer prior variables and scenario variables; if the absolute value of the net contribution weight is less than the weight threshold, the contribution of the evaluation dimension is determined to be insufficient.

[0156] Evaluation dimensions with sample coverage below the coverage threshold, evaluation dimensions with evidence confidence below the confidence threshold, evaluation dimensions that highly overlap with other evaluation dimensions, or evaluation dimensions with an absolute value of net contribution weight less than the weight threshold are removed to obtain the retained evaluation dimensions and their corresponding net contribution weights.

[0157] For the retained evaluation dimensions, the net contribution weight, contribution direction, confidence interval, or correlation stability level with the transformation residual is output, and a candidate weight set is formed based on the net contribution weight, contribution direction, and correlation stability level with the transformation residual.

[0158] The candidate weights in the candidate weight set are normalized, truncated, smoothed, or hierarchically corrected to obtain the net contribution weights.

[0159] Preferably, the telemarketing dialogue evaluation weight calibration system further includes an online quality inspection and feedback module, used for:

[0160] The retained evaluation dimensions and their corresponding net contribution weights will be published to the online quality inspection system or the agent performance system.

[0161] For new telemarketing dialogues, the large language model outputs scores, evidence fragments, and score confidence for each evaluation dimension.

[0162] Based on the scores, net contribution weights, and score confidence levels, calculate the weighted contribution values ​​of each evaluation dimension in the new telemarketing dialogue, and calculate the effective scores of the new telemarketing dialogue based on the weighted contribution values ​​of each evaluation dimension; and output the effective scores, as well as the scores, score confidence levels, weighted contribution values, and corresponding evidence fragments for each evaluation dimension.

[0163] Preferably, the telemarketing dialogue evaluation weight calibration system further includes an update module for:

[0164] Continuously collect new dialogue data, new customer prior data, new agent data, and new conversion tags to form new training samples;

[0165] When the number of new training samples reaches a preset amount, a new round of evaluation dimension and weight generation is triggered; or

[0166] When the differences between new dialogue data, new customer prior data, new agent data, and new conversion tags in the newly added training samples and historical periodic samples exceed their respective thresholds, a new round of evaluation dimension and weight generation is triggered; or

[0167] When the preset update cycle is reached, a new round of evaluation dimension and weight generation is triggered; or

[0168] When the rating calibration error, which measures the consistency between the effective rating and the actual conversion result, exceeds the error threshold, a new round of evaluation dimension generation and weight generation is triggered.

[0169] In summary, as Figure 3 and Figure 4 As shown, a method for calibrating the evaluation weights of telemarketing conversations is summarized as follows:

[0170] S101 and S102: Obtain telemarketing conversation data, conversion tags, customer prior characteristics, agent characteristics, and batch list characteristics.

[0171] S103: Utilize a large language model to perform structured processing on telemarketing dialogue data, automatically generate or update evaluation dimensions, and output the definition, score, evidence fragments, and confidence level corresponding to each evaluation dimension.

[0172] S104: Construct a customer prior conversion model that does not use the content of this conversation based on customer prior characteristics, agent characteristics, and list batch characteristics, and obtain the basic customer conversion probability.

[0173] S105: Train a confounding contribution model based on customer base conversion probability, conversion tags, and evaluation dimension scores, evidence strength, and confidence levels output by the large language model.

[0174] S106: Determine the net contribution weight of each automatically generated evaluation dimension to the conversion label based on the decontamination contribution model, and generate a weight credibility or stability index.

[0175] S107: Feed the net contribution weight and its version information back to the online quality inspection system or agent performance system to weight and score subsequent telemarketing conversations and output dimension contribution explanations and evidence.

[0176] S108: When new samples meet the preset conditions, regenerate or update the evaluation dimensions and net contribution weights, and release them in a versioned manner.

[0177] Example 1: Evaluation Weight Learning Based on Two-Stage Residual Modeling

[0178] Collect historical telemarketing call texts, customer prior data, agent data, batch list data, and multi-level conversion tags. Set up an attribution window based on business configuration and map target events within the attribution window to conversion tags. A large language model reads the dialogue text, automatically generates candidate evaluation dimensions, and outputs a score, evidence, and confidence level for each dialogue. Train a prior conversion model using customer priors, channels, batch lists, agent base variables, and operational variables to obtain P0 for each sample. The system compares actual conversion results with the customer's base conversion probability, forming an incremental signal not explained by customer prior conditions, and uses this incremental signal to train a dialogue behavior contribution model. Based on the model parameters, contribution values, or feature importance corresponding to each evaluation dimension in the contribution model, extract the net contribution weight of the evaluation dimensions and filter dimensions with insufficient coverage, insufficient evidence, or insufficient correlation stability. Publish the weighted version to the online quality inspection system to weight and score subsequent dialogues and display evidence.

[0179] Example 2: Dimensional Family Contribution Estimation Based on Causal ...

[0180] High-scoring states corresponding to a specific evaluation dimension family generated by the large language model are defined as processing variables. Customer priors, channels, batches, agent seats, and operational variables are treated as confounding variables. Control samples under similar prior conditions are constructed using propensity score matching, inverse probability weighting, or dual robust estimation. The conversion difference between those exhibiting and not exhibiting this high-quality dialogue behavior is calculated as the net contribution of this evaluation dimension family. This net contribution is normalized with the contributions of other dimension families to form online scoring weights.

[0181] Example 3: Continuous calibration after sample increase

[0182] During the first training cycle, an initial set of evaluation dimensions and weights is generated based on existing samples. After the first set of weights goes live, the system continues to collect new dialogues and conversion tags. When new samples meet preset conditions, the system re-executes evaluation dimension merging, confounding contribution modeling, and weight calibration. The differences between the first and second sets of weights in terms of scoring stability, conversion calibration error, and dimension coverage are compared. If the second set of weights meets the preset conditions, it is adopted as the new online scoring version; otherwise, the old version is retained or samples are collected again. The more comprehensive the sample, the more stable the weight estimation; the more stable the weights, the more reliable the quality inspection scoring and performance evaluation.

[0183] The beneficial technical effects achieved by the embodiments of the present invention are as follows:

[0184] The technical problem this invention aims to solve is: in telemarketing dialogue quality inspection and agent performance evaluation, how to transform unstructured dialogues into interpretable evaluation dimensions, scores, and evidence; and, under the feedback of real conversion tags, control confounding factors such as customer prior intent, list source, activity batch, and agent allocation, learning the net contribution weight of each evaluation dimension to the conversion result, thereby providing a quantifiable, verifiable, and sustainably updated weight basis for quality inspection scoring. Furthermore, as the dialogue samples and conversion tags increase, how to continuously update the evaluation dimensions and weights, so that the quality inspection scoring system strengthens with the accumulation of business data, rather than remaining on manual experience weights or one-time model results.

[0185] In this embodiment of the invention, the large language model automatically generates evaluation dimensions, scores, evidence, and confidence levels, fixing the lag and insufficient dimension coverage of manual scoring tables, improving dimension discovery capabilities, and providing verifiable evidence for scoring. This reduces the subjectivity caused by manually assigning weights.

[0186] A customer prior conversion model separates customer base conversion tendencies, reducing list quality bias. It evaluates the contribution of agent conversational behavior under similar customer prior conditions, minimizing the interference of list quality, customer natural intent, and activity batches on scoring weights. Residual modeling, propensity scoring, or causal calibration, by directly training weights with conversion labels, incorporates confounding factors, resulting in weights that more closely approximate the net contribution of agent conversational behavior. This model can automatically generate evaluation dimensions and evidence using large language models, learn credible weights based on real conversion labels, and control for the influence of customer priors and other confounding factors.

[0187] The weights are screened for reliability and stability to avoid the lack of basis for manual weights and the potential instability of automatically generated dimensions; resulting in net contribution weights with sample coverage, confidence level, and version records. The net contribution weights have traceable data sources and verifiable calculation basis.

[0188] It can learn the net contribution weight of each evaluation dimension from real business conversion feedback and feed it back into the quality inspection and agent performance system, making subsequent scoring more stable and interpretable, avoiding the difficulty of using offline analysis results in practical management, and realizing a closed loop of online scoring, review, training and performance auxiliary evaluation.

[0189] As new dialogue samples and conversion tags accumulate, evaluation dimensions and their net contribution weights can be recalibrated. The net contribution weight calibration results are used to improve online quality inspection and performance evaluation, gradually enhancing subsequent quality inspection scores, agent performance evaluations, and training improvement suggestions, forming a continuous improvement technology loop. The improved quality inspection results can then be used as data for subsequent training and review, creating a positive cycle of sample accumulation, weight calibration, quality inspection enhancement, and data re-accumulation. The continuous update mechanism improves the stability of weight estimation and the effectiveness of online quality inspection evaluation. This closed loop allows net contribution weights to be updated according to changes in sample size, customer distribution, and business stage, avoiding long-term reliance on fixed manual experience in the quality inspection system.

[0190] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process may be rearranged without departing from the scope of this disclosure. The appended method claims provide elements of various steps in an exemplary order and are not intended to limit the scope to the specific order or hierarchy described.

[0191] In the above detailed description, various features are combined together in a single embodiment to simplify this disclosure. This approach to disclosure should not be construed as reflecting an intention that embodiments of the claimed subject matter require more features than are explicitly stated in each claim. Rather, as reflected in the appended claims, the invention is presented with fewer features than all of the features of the single disclosed embodiment. Therefore, the appended claims are hereby explicitly incorporated into the detailed description, wherein each claim stands alone as a preferred embodiment of the invention.

[0192] The disclosed embodiments have been described above to enable any person skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the spirit and scope of this disclosure. Therefore, this disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.

[0193] The foregoing description includes examples of one or more embodiments. It is certainly impossible to describe all possible combinations of components or methods in order to describe the above embodiments, but those skilled in the art will recognize that further combinations and arrangements of the various embodiments are possible. Therefore, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. Furthermore, the term "comprising" as used in the specification or claims is interpreted in a manner similar to the term "including," as interpreted when used as a conjunction in the claims. Additionally, the use of any term "or" in the specification of the claims is intended to mean "non-exclusive or."

[0194] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for calibrating the evaluation weights of telemarketing conversations, characterized in that, include: Construct a training sample set for telemarketing, with each sample including dialogue data and conversion tags; The system reads the transcribed text of each sample's dialogue data using a large language model, automatically generates a set of candidate evaluation dimensions, and outputs a score, evidence fragment, evidence strength, score confidence, and anomaly marker for each sample on each evaluation dimension of the set. For each sample, the customer prior data, agent data, and operational data are used to obtain the basic customer conversion probability of each conversation. The basic customer conversion probability is used to describe the customer conversion tendency of the sample without using the current conversation data. Based on the conversion tag, customer base conversion probability, and the evaluation dimension scores, evidence strength, and score confidence output by the large language model for each sample, a confounding contribution model is constructed, and the net contribution weight of each evaluation dimension to the conversion tag is determined according to the confounding contribution model. The evaluation dimensions are filtered, and only those that meet the criteria and their corresponding net contribution weights are retained.

2. The method for calibrating the evaluation weight of telemarketing dialogues according to claim 1, characterized in that, Construct a training sample set for telemarketing, including: Acquire telemarketing data. Each piece of telemarketing data includes dialogue data, customer prior data, agent and operational data, and conversion tags. Each piece of telemarketing data serves as a sample. Dialogue data includes call recordings, transcribed text after speech recognition, speaker roles, call duration, call rounds, and call time. Customer prior data includes channel source, list batches, historical behavior, historical outreach, and customer segmentation. Agent and operational data includes agent identifiers, teams, experience levels, shifts, activity batches, and allocation strategies. Conversion tags are obtained by setting attribution windows based on business configurations and mapping target events within those windows. Conversion tags include multi-level conversion tags, conversion time, conversion stage, target event, and completion stage. By using a preset granularity, the dialogue data, customer prior data, agent and operational data, and conversion tags in the telemarketing data are associated and aligned according to customer identifier, call identifier, agent identifier, batch identifier, and timestamp to obtain a training sample set.

3. The method for calibrating the evaluation weight of telemarketing dialogues according to claim 1, characterized in that, The system reads the transcribed text of each sample's dialogue data using a large language model, automatically generates a set of candidate evaluation dimensions, and outputs a score, evidence fragment, evidence strength, score confidence, and anomaly marker for each evaluation dimension in the set for each sample, including: For each sample in the training sample set, the transcribed text corresponding to its dialogue data is input into the large language model. The large language model identifies semantic units in the dialogue data. The semantic units include agent behavior, customer reaction, change of intent, objection type, advancement action, and information disclosure. Candidate evaluation dimensions are automatically generated based on semantic units, and each candidate evaluation dimension is generated with a name, definition, positive examples, negative examples, scoring rules, and evidence extraction rules. The candidate evaluation dimensions are semantically clustered, deduplicated, merged, and hierarchically processed to form the set of evaluation dimensions for the current period. For each sample of dialogue data, output the score, evidence fragment, evidence strength, score confidence, and anomaly marker for each evaluation dimension in the evaluation dimension set.

4. The method for calibrating the evaluation weight of telemarketing dialogues according to claim 1, characterized in that, Based on the conversion tags, customer base conversion probabilities, and the evaluation dimension scores, evidence strength, and score confidence levels output by the large language model for each sample, a decontamination contribution model is constructed. The net contribution weight of each evaluation dimension to the conversion tag is determined according to the decontamination contribution model, including: For each sample, the actual conversion result is compared with the customer's basic conversion probability to obtain the conversion residual. The conversion residual represents the part of the actual conversion result that is not explained by the customer's prior variables and scenario variables. It is used to estimate the incremental contribution of this dialogue behavior to the target conversion. The customer's prior variables come from the customer's prior data, and the scenario variables come from the agent and operation data. The output of the large language model on each evaluation dimension, the score, evidence strength, score confidence, and scene variables are used as input variables. The conversion residual or actual conversion result is used as the learning objective to construct and train the dialogue behavior contribution model. Each sample's dialogue data is input into the dialogue behavior contribution model, which then outputs the model parameters, contribution values, feature importance, or equivalent interpretable outputs corresponding to each evaluation dimension. The net contribution weight of each evaluation dimension is determined based on the model parameters, contribution value, or feature importance corresponding to each evaluation dimension.

5. The method for calibrating the evaluation weight of telemarketing dialogues according to claim 1, characterized in that, Based on the conversion tags, customer base conversion probabilities, and the evaluation dimension scores, evidence strength, and score confidence levels output by the large language model for each sample, a decontamination contribution model is constructed. The net contribution weight of each evaluation dimension to the conversion tag is determined according to the decontamination contribution model, including: The corresponding dialogue behaviors of a certain type of agent behavior are regarded as dialogue behaviors to be evaluated. For the dialogue behavior to be evaluated, the corresponding customer prior variables and scenario variables are regarded as confounding variables that need to be controlled, and the conversion tag in the attribution window where the dialogue behavior to be evaluated is located is regarded as the result variable. The customer prior variables come from customer prior data, and the scenario variables come from agent and operation data. Based on the corresponding prior customer variables and scenario variables, extract similar control samples from the telemarketing data; In the control sample, the difference in conversion results or conversion rate between the sample group with the dialogue behavior to be evaluated and the sample group without the dialogue behavior to be evaluated is compared, and the net contribution weight of the evaluation dimension is estimated based on the difference in conversion results or conversion rate. The difference in conversion results or conversion rate is used to characterize the incremental conversion contribution of the dialogue behavior to be evaluated under similar customer prior conditions and scenario conditions, and is used to estimate the net contribution weight of the corresponding evaluation dimension.

6. The method for calibrating the evaluation weight of telemarketing dialogues according to claim 5, characterized in that, Based on the conversion tags, customer base conversion probabilities, and the evaluation dimension scores, evidence strength, and score confidence levels output by the large language model for each sample, a decontamination contribution model is constructed. The net contribution weight of each evaluation dimension to the conversion tag is determined according to the decontamination contribution model, including: Three types of variables are simultaneously input into the joint constraint model. The first type of variable is the customer prior variable, which comes from the customer prior data. The second type of variable is the scenario variable, which comes from the agent and operation data. The third type of variable is the evaluation dimension variable, which includes the rating, evidence strength, rating confidence, behavioral label and customer reaction label output by the large language model. By using customer prior variables and scenario variables as basic conversion contribution items, and evaluation dimension variables as dialogue behavior contribution items, the net contribution weight of each evaluation dimension is obtained.

7. The method for calibrating the evaluation weight of telemarketing dialogues according to claim 6, characterized in that, Based on the conversion tags, customer base conversion probabilities, and the evaluation dimension scores, evidence strength, and score confidence levels output by the large language model for each sample, a decontamination contribution model is constructed. The net contribution weight of each evaluation dimension to the conversion tag is determined according to the decontamination contribution model, including: For each sample, the actual conversion result is compared with the customer's basic conversion probability to obtain the conversion residual. The conversion residual represents the part of the actual conversion result that is not explained by the customer's prior variables and scenario variables. It is used to estimate the incremental contribution of this dialogue behavior to the target conversion. The customer's prior variables come from the customer's prior data, and the scenario variables come from the agent and operation data. The output of the large language model on each evaluation dimension, the score, evidence strength, score confidence, and scene variables are used as input variables. The conversion residual or actual conversion result is used as the learning objective to construct and train the dialogue behavior contribution model. Each sample's dialogue data is input into the dialogue behavior contribution model, which then outputs the model parameters, contribution values, feature importance, or equivalent interpretable outputs corresponding to each evaluation dimension. The net contribution weight of each evaluation dimension is determined based on the model parameters, contribution value, or feature importance corresponding to each evaluation dimension. The net contribution weight of each evaluation dimension output by the dialogue behavior contribution model is used as the initial net contribution weight. The corresponding dialogue behaviors of a certain type of agent behavior are regarded as dialogue behaviors to be evaluated. For the dialogue behavior to be evaluated, the corresponding customer prior variables and scenario variables are regarded as confounding variables that need to be controlled, and the conversion tag in the attribution window where the dialogue behavior to be evaluated is located is regarded as the result variable. The customer prior variables come from customer prior data, and the scenario variables come from agent and operation data. Based on the corresponding prior customer variables and scenario variables, extract similar control samples from the telemarketing data; In the control sample, the differences in conversion results or conversion rates between the sample group with the dialogue behavior to be evaluated and the sample group without the dialogue behavior to be evaluated are compared. For evaluation dimensions where the net contribution weight reaches the preset quality threshold, the incremental contribution is reviewed by the difference in conversion results or the difference in conversion rate to obtain the incremental conversion contribution. Four types of variables are simultaneously input into the joint constraint model. The first type of variable is the customer prior variable, which comes from the customer prior data. The second type of variable is the scenario variable, which comes from the agent and operation data. The third type of variable is the evaluation dimension variable, which includes the rating, evidence strength, rating confidence, behavioral label and customer reaction label of the evaluation dimension output by the large language model. The fourth type of variable is the incremental conversion contribution. Customer prior variables and scenario variables are used as basic conversion contribution items, and evaluation dimension variables are used as dialogue behavior contribution items. The incremental conversion contribution is used to calibrate the dialogue behavior contribution items to obtain the net contribution weight of each evaluation dimension.

8. The method for calibrating the evaluation weight of telemarketing dialogues according to claim 4 or 7, characterized in that, The evaluation dimensions are filtered, retaining those that meet the criteria and their corresponding net contribution weights, including: For each evaluation dimension, the sample coverage rate is used to determine whether the evaluation dimension appears in a sufficient number of samples; if the number of samples containing valid evidence fragments of the evaluation dimension is lower than the sample number threshold, it is determined that the sample coverage of the evaluation dimension is insufficient. The rating variance is used to determine whether the evaluation dimension has the ability to distinguish different dialogue qualities; if the rating variance is less than the variance threshold, then the evaluation dimension cannot distinguish the quality of the dialogue. The confidence level of evidence is used to determine whether the score is supported by sufficient evidence; if the confidence level of evidence is less than the confidence level threshold, the scoring basis for that evaluation dimension is invalid. The evaluation dimension is judged to be consistent across different periods by using cross-time window stability assessment. If the contribution direction of the evaluation dimension is inconsistent in different time windows, or the fluctuation of the net contribution weight exceeds the weight fluctuation threshold, the cross-time window stability of the evaluation dimension is determined to be insufficient. The correlation stability with conversion residuals is used to determine whether the contribution direction and net contribution weight of the evaluation dimension are stable after controlling for customer prior variables and scenario variables; if the absolute value of the net contribution weight is less than the weight threshold, the contribution of the evaluation dimension is determined to be insufficient. Evaluation dimensions with sample coverage below the coverage threshold, evaluation dimensions with evidence confidence below the confidence threshold, evaluation dimensions that highly overlap with other evaluation dimensions, or evaluation dimensions with an absolute value of net contribution weight less than the weight threshold are removed to obtain the retained evaluation dimensions and their corresponding net contribution weights. For the retained evaluation dimensions, the net contribution weight, contribution direction, confidence interval, or correlation stability level with the transformation residual is output, and a candidate weight set is formed based on the net contribution weight, contribution direction, and correlation stability level with the transformation residual. The candidate weights in the candidate weight set are normalized, truncated, smoothed, or hierarchically corrected to obtain the net contribution weights.

9. The method for calibrating the evaluation weight of telemarketing dialogues according to claim 1, characterized in that, Also includes: The retained evaluation dimensions and their corresponding net contribution weights will be published to the online quality inspection system or the agent performance system. For new telemarketing dialogues, the large language model outputs scores, evidence fragments, and score confidence for each evaluation dimension. Based on the scores, net contribution weights, and score confidence levels, calculate the weighted contribution values ​​of each evaluation dimension in the new telemarketing dialogue, and calculate the effective scores of the new telemarketing dialogue based on the weighted contribution values ​​of each evaluation dimension; and output the effective scores, as well as the scores, score confidence levels, weighted contribution values, and corresponding evidence fragments for each evaluation dimension. The aforementioned method for calibrating the evaluation weights of telemarketing conversations also includes: Continuously collect new dialogue data, new customer prior data, new agent data, and new conversion tags to form new training samples; When the number of new training samples reaches a preset amount, a new round of evaluation dimension and weight generation is triggered; or When the differences between new dialogue data, new customer prior data, new agent data, and new conversion tags in the newly added training samples and historical periodic samples exceed their respective thresholds, a new round of evaluation dimension and weight generation is triggered; or When the preset update cycle is reached, a new round of evaluation dimension and weight generation is triggered; or When the rating calibration error, which measures the consistency between the effective rating and the actual conversion result, exceeds the error threshold, a new round of evaluation dimension generation and weight generation is triggered.

10. A telemarketing dialogue evaluation weight calibration system, characterized in that, include: The data acquisition module is used to build a training sample set for telemarketing, with each sample including dialogue data and conversion tags; The large language model structuring module is used to read the transcribed text of each sample's dialogue data through the large language model, automatically generate a set of candidate evaluation dimensions, and output a score, evidence fragment, evidence strength, score confidence, and anomaly marker for each sample on each evaluation dimension of the evaluation dimension set. The customer prior conversion module is used to obtain the basic customer conversion probability of each dialogue based on the customer prior data, agent data and operational data of each sample. The basic customer conversion probability is used to describe the customer conversion tendency of the sample without using the current dialogue data. The decontamination contribution module is used to construct a decontamination contribution model based on the conversion tag, customer base conversion probability, and evaluation dimension scores, evidence strength, and score confidence output by the large language model for each sample. The net contribution weight of each evaluation dimension to the conversion tag is determined based on the decontamination contribution model. The evaluation weight generation module is used to filter evaluation dimensions and retain the evaluation dimensions that meet the conditions and their corresponding net contribution weights.