Quality evaluation method and device for high-quality evaluation set of intelligent marketing intelligent agent

CN122286230BActive Publication Date: 2026-09-22BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610759133.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-09-22
Estimated Expiration
2046-05-28

AI Technical Summary

Technical Problem

目前,对评测集的质量评估主要依赖于人为评估,需要投入大量的人力和时间,且评估结果通常受到评估人的主观因素影响,质量评估结果的一致性难以保证

Benefits of technology

[0009]通过上述技术方案,获取待评估的评测集以及与评测集关联的参考信息,通过参考信息提供用于对与评测集进行质量评估的客观参考,避免主观因素的影响。对于评测集的每一样本,利用多个预设的评估维度,以及各评估维度关联的参考信息,确定样本在不同评估维度各自的维度分数,其中评估维度包括与问题之间相关的至少一个第一类维度以及与答案质量相关的至少一个第二类维度,由此,通过针对问题和答案分别设置多样化的评估维度,对样本在每个维度的质量表现通过分数进行量化,初步确定样本在各评估维度的质量分数。之后,通过样本在各评估维度的维度分数,确定评测集中样本的质量评估结果,即,通过对样本在多个评估维度的分数进行融合汇总,得到样本的综合性的质量评估结果,以实现对样本的全面质量评估。其中,在预设的评估维度中,设置有用于评估样本答案受知识支持程度的第一评估维度,能够有效评估样本答案在可信度方面的表现,利于甄别样本答案中可能存在的幻觉问题。这样,能够自动地对评测集中的每一样本进行全面的质量评估,既能够提升评测集的质量评估效率,也能够保证质量评估结果的一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122286230B_ABST
    Figure CN122286230B_ABST
Patent Text Reader

Abstract

A quality evaluation method and device for a high-quality evaluation set of an intelligent marketing intelligent agent. The method comprises: obtaining an evaluation set to be evaluated; obtaining reference information associated with the evaluation set; for each sample, determining the dimension score of the sample in each evaluation dimension using a plurality of preset evaluation dimensions and the reference information associated with each evaluation dimension, the evaluation dimensions including at least one first type of dimension and at least one second type of dimension; determining the quality evaluation result of each sample in the evaluation set according to the dimension score of the sample in each evaluation dimension; wherein the second type of dimension includes a first evaluation dimension, and the dimension score of the sample in the first evaluation dimension is determined by the following method: performing text splitting processing on the answer of the sample to obtain a plurality of splitting units; for each splitting unit, determining whether the splitting unit can be supported by associated knowledge information; and determining the dimension score of the sample in the first evaluation dimension according to the number of splitting units that can be supported by the associated knowledge information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical solution relates to the field of computer technology, specifically to a method and apparatus for evaluating the quality of an evaluation set for intelligent marketing agents. Background Technology

[0002] In building and optimizing large language models, a high-quality evaluation set is crucial for measuring model performance and guiding model iteration. Currently, the quality assessment of the evaluation set mainly relies on human evaluation, which requires a significant investment of manpower and time. Furthermore, the evaluation results are often influenced by the subjective factors of the evaluators, making it difficult to guarantee the consistency of the quality assessment results. Summary of the Invention

[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] Firstly, a method for evaluating the quality of a test set for intelligent marketing agents is provided, including: Obtain the evaluation set to be evaluated, the evaluation set including at least one sample, the sample including questions and answers; Obtain reference information associated with the evaluation set, the reference information including related knowledge information; For each sample, using multiple preset evaluation dimensions and reference information associated with each evaluation dimension, the dimension score of the sample in each evaluation dimension is determined. The evaluation dimensions include at least one first type dimension and at least one second type dimension. The first type dimension is related to the quality of the question, and the second type dimension is related to the quality of the answer. Based on the dimensional scores of the samples in each evaluation dimension, the quality assessment result of each sample in the evaluation set is determined; The second type of dimension includes a first evaluation dimension, which characterizes the degree to which the sample's answer is supported by the associated knowledge information. The reference information associated with the first evaluation dimension includes the associated knowledge information. The sample's dimensional score in the first evaluation dimension is determined in the following way: The answers to the sample are processed by text splitting to obtain multiple split units; for each split unit, it is determined whether the split unit can be supported by the associated knowledge information; based on the number of split units that can be supported by the associated knowledge information, the dimension score of the sample in the first evaluation dimension is determined.

[0005] Secondly, a device for evaluating the quality of a test set for intelligent marketing agents is provided, comprising: The first acquisition module is used to acquire the evaluation set to be evaluated, the evaluation set including at least one sample, the sample including questions and answers; The second acquisition module is used to acquire reference information associated with the evaluation set, the reference information including associated knowledge information; An evaluation module is used to determine the dimension score of each sample in each of the evaluation dimensions using multiple preset evaluation dimensions and reference information associated with each evaluation dimension. The evaluation dimensions include at least one first type dimension and at least one second type dimension. The first type dimension is related to the quality of the question, and the second type dimension is related to the quality of the answer. The determination module is used to determine the quality assessment result of each sample in the evaluation set based on the dimensional scores of the sample in each evaluation dimension. The second type of dimension includes a first evaluation dimension, which is used to characterize the degree to which the sample's answer is supported by the associated knowledge information. The reference information associated with the first evaluation dimension includes the associated knowledge information. The dimension score of the sample in the first evaluation dimension is determined by the following method: performing text splitting on the sample's answer to obtain multiple split units; for each split unit, determining whether the split unit can be supported by the associated knowledge information; and determining the dimension score of the sample in the first evaluation dimension based on the number of split units that can be supported by the associated knowledge information.

[0006] Thirdly, a computer-readable medium is provided having a computer program stored thereon, wherein the computer program, when executed by a processing device, implements the steps of the method described in the first aspect.

[0007] Fourthly, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect above.

[0008] Fifthly, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, it implements the steps of the method described in the first aspect.

[0009] The above technical solution obtains the evaluation set to be evaluated and the reference information associated with it. This reference information provides an objective basis for quality assessment of the evaluation set, avoiding the influence of subjective factors. For each sample in the evaluation set, multiple pre-defined evaluation dimensions and the reference information associated with each dimension are used to determine the sample's dimensional scores in different evaluation dimensions. These evaluation dimensions include at least one first-type dimension related to the question and at least one second-type dimension related to the answer quality. Thus, by setting diverse evaluation dimensions for both questions and answers, the quality performance of the sample in each dimension is quantified through scores, initially determining the sample's quality score in each evaluation dimension. Subsequently, the quality assessment result of the samples in the evaluation set is determined based on the dimensional scores of the samples in each evaluation dimension. That is, by aggregating and summarizing the scores of the samples in multiple evaluation dimensions, a comprehensive quality assessment result is obtained, achieving a holistic quality assessment of the samples. Among the pre-defined evaluation dimensions, a first evaluation dimension is set to assess the degree of knowledge support for the sample answers, effectively evaluating the credibility of the sample answers and facilitating the identification of potential illusion problems in the sample answers. This allows for a comprehensive quality assessment of each sample in the evaluation set, which not only improves the efficiency of the quality assessment but also ensures the consistency of the assessment results.

[0010] Other features and advantages of the technical solution will be described in detail in the following detailed implementation section. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the technical solution will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a flowchart illustrating a method for evaluating the quality of a test set for intelligent marketing agents, based on certain scenarios. Figure 2 This is a block diagram illustrating a quality assessment device for evaluation sets used in intelligent marketing agents, based on certain scenarios. Figure 3 A schematic diagram of an electronic device suitable for implementing the above-mentioned technical solution is shown. Detailed Implementation

[0012] The technical solution will now be described in more detail with reference to the accompanying drawings. Although certain scenarios are shown in the drawings, it should be understood that the technical solution can be implemented in various forms and should not be construed as limited to the scenarios described herein. Rather, these scenarios are provided to provide a more thorough and complete understanding of the technical solution. It should be understood that the accompanying drawings and the scenarios described are for illustrative purposes only and are not intended to limit the scope of protection of the technical solution.

[0013] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of the technical solution is not limited in this respect.

[0014] The term "comprising" and its variations as used herein can be open-ended, meaning "including but not limited to". The term "based on" can mean "at least partially based on". The term "one case" means "at least one case"; the term "another case" means "at least one additional case"; the term "some cases" means "at least some cases". Definitions of other terms will be given in the following description.

[0015] It should be noted that the concepts of "first" and "second" mentioned here are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.

[0016] It should be noted that the terms "one" and "more" used here are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0017] The names of messages or information exchanged between the multiple devices in the implementation are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0018] It is understandable that before using the technical solutions provided here, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in accordance with relevant laws and regulations, and their authorization should be obtained through appropriate means.

[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations described herein.

[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the technical solution. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the technical solution.

[0022] At the same time, it is understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and related provisions.

[0023] With the increasingly widespread application of large language models, high-quality evaluation sets are becoming increasingly important for model optimization. Typically, the quality assessment of evaluation sets relies on human evaluation methods, such as reading through the entire set according to evaluation criteria and determining its quality. Alternatively, scripts can be written to check the format of the evaluation set, or keyword comparison can be used to roughly assess its quality. However, these methods all have significant drawbacks. For example, human evaluation requires substantial human resources and time, has a long evaluation cycle, and is inefficient. Furthermore, because the evaluation is highly dependent on the evaluator's subjective understanding and experience, differences in evaluation criteria between different evaluators, and even among the same evaluator at different times, can easily arise, leading to unstable and inconsistent quality assessment results. Similarly, script writing or keyword comparison methods only provide a superficial and simple quality assessment, lacking depth and failing to achieve true quality evaluation. Moreover, when quality evaluation criteria need to be adjusted, the human evaluation process must be redesigned or the scripts rewritten, which is time-consuming and lacks flexibility.

[0024] To address the aforementioned issues, a method and apparatus for evaluating the quality of evaluation sets for intelligent marketing agents are provided.

[0025] Figure 1 This is a flowchart illustrating a method for evaluating the quality of an evaluation set for an intelligent marketing agent, illustrating several scenarios. The method presented in this paper can be applied to the quality evaluation of an evaluation set for an intelligent marketing agent, a dialogue system based on a large language model that can output corresponding answers based on input questions. Figure 1 As shown, the method may include steps 11 to 14.

[0026] In step 11, the evaluation set to be evaluated is obtained.

[0027] The evaluation set may include at least one sample, and the sample may include a question and an answer. The question may include a question text described in natural language, and the answer may include the correct answer corresponding to that question.

[0028] Optionally, in addition to questions and answers, samples may include other content depending on the specific context, such as related knowledge information, category labels, and difficulty level labels. Related knowledge information may include the background knowledge upon which the answers to the questions in the sample depend, such as data table structures, business rules, and knowledge bases. This related knowledge information may include at least one of structured data, unstructured text, or semi-structured data. Category labels may include the category to which the sample belongs, such as, but not limited to, scenario categories or domain categories. Difficulty level labels can be used to characterize the difficulty of the questions in the sample, such as easy or difficult.

[0029] In some cases, the evaluation set can be generated by a large language model from new samples based on a seed set and related knowledge information relevant to the intelligent marketing agent (e.g., domain-related, scenario-related, etc.). The seed set can include at least one seed sample, and the seed sample has the same sample structure as the one in the evaluation set. It can include questions and answers, and may further include, for example, related knowledge information, category labels, difficulty level labels, etc.

[0030] In some cases, the evaluation set can be obtained by generating new samples from a large language model based on the associated knowledge information related to the intelligent marketing agent, and the generation of new samples can be independent of the seed set.

[0031] In other cases, the evaluation set may also include new samples generated by each of the two methods mentioned above.

[0032] In step 12, reference information associated with the evaluation set is obtained.

[0033] Reference information associated with the evaluation set can be used to provide a reference for the quality assessment of the evaluation set and assist in the quality assessment of the samples in the evaluation set.

[0034] Optionally, the reference information may include information related to the evaluation set. For example, related knowledge information used to generate the evaluation set, a seed set (e.g., all questions in the seed set), etc. As another example, if the samples include category labels, they may also include predefined categories related to the evaluation set. As yet another example, if the samples include difficulty level labels, they may also include predefined difficulty levels related to the evaluation set.

[0035] In step 13, for each sample, multiple preset evaluation dimensions and reference information associated with each evaluation dimension are used to determine the dimension score of the sample in each evaluation dimension. The evaluation dimensions include at least one first-class dimension and at least one second-class dimension. The first-class dimension is related to the quality of the question, and the second-class dimension is related to the quality of the answer.

[0036] The evaluation dimensions can be flexibly set according to the actual quality evaluation needs to achieve quality evaluation in different situations.

[0037] For each evaluation dimension, there are usually reference information set up to achieve the quality evaluation for that evaluation dimension. Therefore, based on the reference information associated with that evaluation dimension and the corresponding quality evaluation method, the quality score of the sample corresponding to the current evaluation dimension can be determined, that is, the dimension score of that evaluation dimension.

[0038] For example, if the evaluation dimension includes an evaluation dimension used to characterize the relevance between the questions and knowledge in the sample, then in order to achieve a quality evaluation of this evaluation dimension, it is necessary to refer to the related knowledge information in the reference information, and then use the set method for determining the relevance to determine the relevance between the questions and related knowledge information in the sample, so as to obtain the dimension score of this evaluation dimension.

[0039] Optionally, for each evaluation dimension, a score range corresponding to that evaluation dimension can be preset so that the obtained dimension scores all fall within that score range, or are ultimately mapped to that score range, facilitating management and use. For example, the score range can be set to 0-1.

[0040] Therefore, the dimensional scores of any sample in the evaluation set in each of the multiple evaluation dimensions can be determined using the above method.

[0041] In some cases, the second dimension may include the first evaluation dimension. The first evaluation dimension can be used to characterize the degree to which the sample's answer is supported by related knowledge information. Correspondingly, the reference information associated with the first evaluation dimension may include related knowledge information. The relevant explanations of related knowledge information have been given above and will not be repeated here.

[0042] In such cases, the dimensional score of the sample in the first evaluation dimension can be determined in the following way: The sample answers are processed by text splitting to obtain multiple split units; For each segmented unit, determine whether the segmented unit can be supported by associated knowledge information; The dimensional score of a sample in the first evaluation dimension is determined based on the number of split units that can be supported by associated knowledge information.

[0043] Optionally, the answers to the sample can be split into multiple sub-units using natural language processing methods, such as large language models. For example, the answer can be split into multiple minimal semantic units, which are semantically indivisible.

[0044] For each segmented unit, it can be determined whether the segmented unit can be supported by associated knowledge information, that is, whether there is knowledge in the associated knowledge information that can support the segmented unit. In this paper, the segmented unit being supported by associated knowledge information means that the semantics expressed by the segmented unit can be directly or indirectly inferred from a part of the associated knowledge information. For example, the associated knowledge information contains text with the same or similar semantics as the segmented unit, or multiple fragments in the associated knowledge information can be combined to infer the content of the segmented unit. The method for determining whether a segmented unit can be supported by associated knowledge information can be flexibly selected according to actual needs, such as using semantic matching or natural language inference.

[0045] Optionally, the above judgment can be achieved using a large language model. For example, the split unit and related knowledge information can be concatenated and input into the large language model together, and the large language model can be guided by preset prompt words to judge whether the split unit can be supported by the related knowledge information.

[0046] Optionally, a natural language reasoning model can be used to determine whether the split unit is included in the associated knowledge information. If the split unit is included in the associated knowledge information, it is determined that the split unit can be supported by the associated knowledge information.

[0047] Optionally, the similarity of semantic vectors between the split unit and the associated knowledge information can be determined. If the similarity exceeds a specified threshold, it is determined that the split unit can be supported by the associated knowledge information.

[0048] Using the above method, it can be determined whether each segmentation unit can be supported by associated information. In this way, after processing each segmentation unit, the number of segmentation units that can be supported by associated knowledge information and the number of segmentation units that cannot be supported by associated knowledge information can be determined. Furthermore, based on the number of segmentation units that can be supported by associated knowledge information, the dimensional score of the sample in the first evaluation dimension can be determined.

[0049] Optionally, a function or mapping relationship related to the first quantity can be designed to make the dimensional score positively correlated with the first quantity, where the first quantity is the number of split units that can be supported by associated knowledge information. That is, the larger the first quantity, the higher the dimensional score of the sample in the first evaluation dimension. For example, if the total number of split units is denoted as the second quantity, the ratio of the first quantity to the second quantity can be determined as the dimensional score of the sample in the first evaluation dimension.

[0050] Optionally, after obtaining multiple split units, and before determining whether the split units can be supported by associated knowledge information, the split units can be preprocessed by filtering, and the dimension score of the first evaluation dimension can be determined by referring to the above method using the filtered split units.

[0051] For example, the semantic relevance between each split unit and the sample's question text can be determined separately (e.g., using a large language model). A relevance threshold is set, and split units with semantic relevance less than the threshold are filtered out. Split units with semantic relevance reaching the threshold are selected as the filtered split units. Here, when calculating the dimensional score of the sample in the first evaluation dimension, the total number of split units refers to the number of filtered split units.

[0052] This method measures how faithfully a sample's answers are to related knowledge information. Samples with lower levels of support from related knowledge information, i.e., those more likely to be illusions, are quantified with lower scores, while samples with higher levels of support from related knowledge information are quantified with higher scores, thus reflecting the credibility of the sample's answers.

[0053] In step 14, the quality assessment result of each sample in the evaluation set is determined based on the dimensional scores of the sample in each evaluation dimension.

[0054] After obtaining the dimensional scores of the sample in each evaluation dimension, these dimensional scores can be summarized, for example by averaging or weighted summation, to obtain the overall quality score of the sample, which serves as the quality assessment result of the sample.

[0055] Optionally, after obtaining the quality scores of the samples in the evaluation set, these quality scores can be output in a structured format, or they can be displayed in a visual form such as a list or graph. The visual presentation can serve as the quality assessment result of the evaluation set.

[0056] The above technical solution obtains the evaluation set to be evaluated and the reference information associated with it. This reference information provides an objective basis for quality assessment of the evaluation set, avoiding the influence of subjective factors. For each sample in the evaluation set, multiple pre-defined evaluation dimensions and the reference information associated with each dimension are used to determine the sample's dimensional scores in different evaluation dimensions. These evaluation dimensions include at least one first-type dimension related to the question and at least one second-type dimension related to the answer quality. Thus, by setting diverse evaluation dimensions for both questions and answers, the quality performance of the sample in each dimension is quantified through scores, initially determining the sample's quality score in each evaluation dimension. Subsequently, the quality assessment result of the samples in the evaluation set is determined based on the dimensional scores of the samples in each evaluation dimension. That is, by aggregating and summarizing the scores of the samples in multiple evaluation dimensions, a comprehensive quality assessment result is obtained, achieving a holistic quality assessment of the samples. Among the pre-defined evaluation dimensions, a first evaluation dimension is set to assess the degree of knowledge support for the sample answers, effectively evaluating the credibility of the sample answers and facilitating the identification of potential illusion problems in the sample answers. This allows for a comprehensive quality assessment of each sample in the evaluation set, which not only improves the efficiency of the quality assessment but also ensures the consistency of the assessment results.

[0057] In some cases, the reference information may include at least one reference question. The first dimension may include a second evaluation dimension, which can be used to characterize the degree of difference between the sample question and the reference question. Correspondingly, the reference information associated with the second evaluation dimension may include at least one reference question.

[0058] Optionally, at least one reference problem may include all or some of the problems in the seed set associated with the evaluation set as described above.

[0059] The second evaluation dimension measures the difference between the sample questions and the reference questions, thereby reflecting the degree of repetition between the sample questions and the existing reference questions. The greater the difference between the sample questions and the reference questions, the more novel the sample is compared to the existing reference questions, which is more conducive to the diversity of the evaluation set, and therefore the higher the corresponding quality score.

[0060] In such cases, the dimensional score of the sample in the second evaluation dimension can be determined in the following way: For each reference problem, determine the similarity between the sample problem and the reference problem; Determine the maximum similarity among the similarity scores; Based on the maximum similarity, the dimensional score of the sample in the second evaluation dimension is determined, and the dimensional score of the second evaluation dimension is negatively correlated with the maximum similarity.

[0061] For the currently evaluated sample, text similarity can be calculated between the sample's question text and each reference question text. For example, text similarity can be characterized by any of edit distance, cosine similarity, or semantic similarity. Semantic similarity can be determined based on a model (e.g., a large language model).

[0062] After determining the similarity between the sample question and each reference text, the maximum similarity can be taken as the maximum similarity, and the sample's dimensional score in the second evaluation dimension can be determined based on this maximum similarity. The dimensional score is negatively correlated with the maximum similarity; that is, the higher the maximum similarity, the greater the degree of repetition between the sample and the reference question, the lower the quality, and the lower the score. For example, if the score range is set between 0 and 1, the maximum similarity can be subtracted from 1 to obtain the dimensional score for the second evaluation dimension. Alternatively, if the score range is also set between 0 and 1, the reciprocal of the maximum similarity can be determined, and then normalized to the 0-1 range to obtain the dimensional score for the second evaluation dimension.

[0063] Optionally, after determining the similarity between the sample's question and the reference question for each reference question, a similarity reference value can be determined based on the similarity. For example, this can be done by taking the mean or median of the similarity. Then, the similarity reference value is used to determine the dimensional score of the sample in the second evaluation dimension. The dimensional score of the second evaluation dimension is negatively correlated with the similarity reference value. The method for determining the dimensional score of the second evaluation dimension based on the similarity reference value is the same as the method for determining the dimensional score of the second evaluation dimension based on the maximum similarity, and will not be described again.

[0064] This method allows us to measure the degree of duplication between a sample and existing samples, quantifying similar or duplicated samples as lower scores and newer samples as higher scores, thereby reflecting the sample's performance in terms of data diversity.

[0065] In some cases, the first dimension may include a third evaluation dimension, which can be used to characterize the completeness of the questions in the sample. The reference information may include at least one evaluation rule, and correspondingly, the reference information associated with the third evaluation dimension includes at least one evaluation rule.

[0066] At least one evaluation rule can be set according to actual needs, and different evaluation rules are used to evaluate the completeness of the sample questions from different aspects. Optionally, the evaluation rules may include, but are not limited to, the requirement that the questions should include the required key fields, the requirement that the question length should meet the length requirements, and the requirement that the question structure should be reasonable.

[0067] In such cases, the sample's score in the third evaluation dimension can be determined in the following way: Obtain the initial score for the third evaluation dimension, as well as the deduction value for each evaluation rule; Based on each evaluation rule, determine whether the problem in the sample meets the evaluation rule; If a problem in a sample does not meet the evaluation rules, the total deduction value is determined based on the deduction value corresponding to the evaluation rules. Based on the initial score and the total deduction, determine the sample's dimensional score in the third evaluation dimension.

[0068] The third evaluation dimension has an initial score, which can be set to the maximum value of the score range for the third evaluation dimension. For example, if the score range for the third evaluation dimension is 0-1, then the initial score is set to 1. Each evaluation rule can have a deduction value set as needed. The deduction values ​​for different evaluation rules can be the same or different. For example, evaluation rules that have a greater impact on the completeness of the problem can have a larger deduction value, while evaluation rules that have a less significant impact on the completeness of the problem can have a smaller deduction value.

[0069] Using each evaluation rule, determine whether the sample's problem meets the rule. For evaluation rules that cannot be met, record the corresponding deduction value; for evaluation rules that can be met, no deduction is required. In this way, after each evaluation rule has been evaluated, the total deduction value can be obtained by summing the recorded deduction values, and the difference between the initial score and the total deduction value is determined as the sample's dimension score in the third evaluation dimension.

[0070] This method can automatically detect incompleteness in the samples, quantifying samples that are less complete to a lower score and samples that are more complete to a higher score, thereby reflecting the samples' performance in terms of data integrity.

[0071] In some cases, the first dimension may include a fourth evaluation dimension. This fourth evaluation dimension can be used to characterize the degree of relevance between the sample's problem and related knowledge information. Correspondingly, the reference information associated with the fourth evaluation dimension may include related knowledge information. The relevant explanations of related knowledge information have been given above and will not be repeated here.

[0072] In such cases, the sample's score in the fourth evaluation dimension can be determined in the following way: Using a large language model, the relevance score between the sample's question and related knowledge information is determined, which serves as the sample's dimensional score in the fourth evaluation dimension.

[0073] Optionally, the sample questions and related knowledge information can be concatenated, and a score range can be specified. These can be input into a large language model, and the model can be guided by preset prompts to determine the degree of relevance between the questions and related knowledge information. The degree of relevance can then be quantified into a score within the specified score range.

[0074] This method can measure the semantic relevance between a sample and related knowledge information, quantifying samples with higher semantic relevance to higher scores and samples with lower semantic relevance to lower scores, thereby reflecting the sample's performance in terms of knowledge relevance.

[0075] In some cases, the first dimension may include the seventh evaluation dimension. The seventh evaluation dimension can be used to characterize the semantic accuracy of the sample's question. Semantic accuracy can include the degree of semantic fluency and grammatical correctness. Correspondingly, the reference information associated with the seventh evaluation dimension may include rules for evaluating semantic accuracy. The above rules can be flexibly set according to needs, such as semantic fluency, grammatical clarity, reasonable expression, and absence of obvious problems.

[0076] In such cases, the sample's score in the seventh evaluation dimension can be determined as follows: Using a large language model, based on the rules used to evaluate semantic accuracy, the semantic accuracy score of the sample's question is determined, which serves as the sample's dimensional score in the seventh evaluation dimension.

[0077] Optionally, the sample question can be concatenated with the rules used to evaluate semantic accuracy, and a score range can be specified. The two are then input into a large language model, and the large language model can be guided by preset prompts to evaluate the semantic accuracy of the question in accordance with the rules used to evaluate semantic accuracy. The semantic accuracy is then quantified into a score within the specified score range.

[0078] This method allows us to measure the semantic accuracy of a sample, quantifying samples with higher semantic accuracy into higher scores and samples with lower semantic accuracy into lower scores, thereby reflecting the sample's performance in terms of semantic correctness.

[0079] In some cases, the second type of dimension may include an eighth evaluation dimension, which can be used to characterize the relevance between the sample's question and its answer. In such cases, the sample's score on the eighth evaluation dimension can be determined as follows: Determine the semantic similarity score between the sample's question and the sample's answer, and use this score as the sample's dimensional score in the eighth evaluation dimension.

[0080] For example, a large language model can be used to take the sample questions and answers as input, determine the degree of relevance between the questions and answers through guide words, and quantify it into the score range corresponding to the eighth evaluation dimension.

[0081] For example, we can first vectorize the questions and answers of the samples to obtain question vectors and answer vectors, determine the vector similarity between the two, and then measure the vector similarity into the score range corresponding to the eighth evaluation dimension.

[0082] This method measures the relevance between the questions and answers in a sample, quantifying samples with a higher degree of relevance to the questions and answers into higher scores, and samples with a lower degree of relevance to the questions and answers, i.e., those that are irrelevant to the question, into lower scores, thereby reflecting the performance of the samples in terms of question-answer relevance.

[0083] In some cases, if the sample also includes category labels, the reference information may include at least one reference category, and the evaluation dimension may also include a fifth evaluation dimension. The reference information associated with the fifth evaluation dimension may include at least one reference category. Reference categories can be set as needed, and the sample's category labels can be taken from the reference categories.

[0084] In such cases, the sample's score in the fifth evaluation dimension can be determined in the following way: Using a large language model, determine the recommendation category to which the sample's question belongs in at least one reference category; Determine the first degree of matching between the recommended category and the category label of the sample; Based on the first degree of matching, determine the dimensional score of the sample in the fifth evaluation dimension.

[0085] Using a large language model and reference categories, a guiding word is set to guide the large language model in determining the recommendation category to which the sample belongs within the reference categories. This then determines the degree of matching between the sample's category label and the recommendation category, i.e., the first degree of matching. For example, the large language model can determine the recommendation category of the sample's question and simultaneously determine the confidence level of the recommendation category. If the recommendation category matches the category label, the confidence level is determined as the first degree of matching. Subsequently, the first degree of matching is further mapped to the score range corresponding to the fifth evaluation dimension, obtaining the dimension score of the fifth evaluation dimension.

[0086] This method allows us to measure the reasonableness of the category labels on a sample and reflects the sample's performance in terms of classification accuracy.

[0087] In some cases, if the sample also includes difficulty level labels, the reference information may include at least one reference difficulty level, and the evaluation dimension may also include a sixth evaluation dimension. The reference information associated with the sixth evaluation dimension may include at least one reference difficulty level. The reference difficulty level can be set as needed, and the difficulty level label of the sample can be taken from the reference category.

[0088] In such cases, the sample's score in the sixth evaluation dimension can be determined in the following way: Using a large language model, determine the recommended difficulty level of the sample's question within at least one reference difficulty level; Determine the second degree of matching between the recommended difficulty level and the difficulty level label of the sample; Based on the second degree of matching, determine the dimensional score of the sample in the sixth evaluation dimension.

[0089] Using a large language model and a reference difficulty level, a guiding word is set to guide the large language model in determining the recommended difficulty level of a sample within the reference difficulty level. This then determines the degree of matching between the sample's difficulty level label and the recommended difficulty level, i.e., the second degree of matching. For example, the large language model can determine the recommended difficulty level of a sample's question and simultaneously determine the confidence level of that recommendation. When the recommended difficulty level matches the difficulty level label, the confidence level is determined as the second degree of matching. Subsequently, the second degree of matching is further mapped to the score range corresponding to the sixth evaluation dimension to obtain the dimension score for the sixth evaluation dimension.

[0090] This method allows us to measure the reasonableness of the difficulty level labels on a sample, reflecting the accuracy of the labeling in terms of difficulty.

[0091] It should be noted that when determining the dimensional scores of a sample, any one of the aforementioned evaluation dimensions can be selected as needed, or any combination of the aforementioned evaluation dimensions can be flexibly combined as needed to conduct a multi-dimensional quality assessment of the sample and obtain the dimensional scores of each evaluation dimension.

[0092] In some cases, step 14 may include the following steps: For each sample, a quality score is obtained by weighted summation based on the sample's dimensional scores in each evaluation dimension and the corresponding weight values ​​of each evaluation dimension. This quality score serves as the sample's quality assessment result.

[0093] Optionally, different weight values ​​can be set for different evaluation dimensions according to the actual needs of quality assessment. Then, based on the sample's dimensional scores in each evaluation dimension and the weight values ​​of each evaluation dimension, a weighted sum is obtained to obtain the sample's quality score. In this way, evaluation dimensions can be used flexibly according to actual needs to obtain a quality score that better meets the requirements.

[0094] Figure 2 This is a block diagram illustrating a quality assessment device for evaluation sets used in intelligent marketing agents, based on various scenarios. For example... Figure 2 As shown, the device 20 may include: The first acquisition module 21 is used to acquire the evaluation set to be evaluated, the evaluation set including at least one sample, the sample including questions and answers; The second acquisition module 22 is used to acquire reference information associated with the evaluation set, the reference information including associated knowledge information; Evaluation module 23 is used to determine the dimension score of each sample in each evaluation dimension by using multiple preset evaluation dimensions and reference information associated with each evaluation dimension. The evaluation dimensions include at least one first type dimension and at least one second type dimension. The first type dimension is related to the quality of the question, and the second type dimension is related to the quality of the answer. The determination module 24 is used to determine the quality assessment result of each sample in the evaluation set based on the dimensional scores of the sample in each evaluation dimension. The second type of dimension includes a first evaluation dimension, which is used to characterize the degree to which the sample's answer is supported by the associated knowledge information. The reference information associated with the first evaluation dimension includes the associated knowledge information. The dimension score of the sample in the first evaluation dimension is determined by the following method: performing text splitting on the sample's answer to obtain multiple split units; for each split unit, determining whether the split unit can be supported by the associated knowledge information; and determining the dimension score of the sample in the first evaluation dimension based on the number of split units that can be supported by the associated knowledge information.

[0095] Optionally, the reference information includes at least one reference question; the first category dimension includes a second evaluation dimension, which is used to characterize the degree of difference between the question of the sample and the reference question, and the reference information associated with the second evaluation dimension includes the at least one reference question; The dimensional score of the sample in the second evaluation dimension is determined by the following sub-modules: The first determining submodule is used to determine the similarity between the sample problem and the reference problem for each reference problem; The second determining submodule is used to determine the maximum similarity among the similarities; The third determining submodule is used to determine the dimensional score of the sample in the second evaluation dimension based on the maximum similarity, wherein the dimensional score of the second evaluation dimension is negatively correlated with the maximum similarity.

[0096] Optionally, the first type of dimension includes a third evaluation dimension, which is used to characterize the completeness of the questions in the sample; the reference information includes at least one evaluation rule, and the reference information associated with the third evaluation dimension includes the at least one evaluation rule; The dimensional score of the sample in the third evaluation dimension is determined by the following sub-modules: The acquisition submodule is used to acquire the initial score of the third evaluation dimension and the deduction value corresponding to each evaluation rule; The fourth determination submodule is used to determine whether the problem of the sample satisfies the evaluation rule based on each of the evaluation rules; The fifth determination submodule is used to determine the total deduction value based on the deduction value corresponding to the evaluation rule in response to the problem of the sample not meeting the evaluation rule; The sixth determining submodule is used to determine the dimension score of the sample in the third evaluation dimension based on the initial score and the total deduction value.

[0097] Optionally, the reference information includes associated knowledge information; the first type of dimension includes a fourth evaluation dimension, which is used to characterize the degree of correlation between the problem of the sample and the associated knowledge information, and the reference information associated with the fourth evaluation dimension includes the associated knowledge information; The sample's dimensional score in the fourth evaluation dimension is determined by the following sub-modules: The seventh determination submodule is used to determine the relevance score between the problem of the sample and the associated knowledge information using a large language model, which is used as the dimension score of the sample in the fourth evaluation dimension.

[0098] Optionally, the sample further includes category labels; the reference information includes at least one reference category; the evaluation dimension further includes a fifth evaluation dimension, and the reference information associated with the fifth evaluation dimension includes the at least one reference category; The sample's dimensional score in the fifth evaluation dimension is determined by the following sub-modules: The eighth determination submodule is used to determine the recommendation category to which the question of the sample belongs in the at least one reference category by utilizing a large language model; The ninth determining submodule is used to determine the first degree of matching between the recommended category and the category label of the sample; The tenth determining submodule is used to determine the dimension score of the sample in the fifth evaluation dimension based on the first matching degree.

[0099] Optionally, the sample further includes a difficulty level label; the reference information includes at least one reference difficulty level; the evaluation dimension further includes a sixth evaluation dimension, and the reference information associated with the sixth evaluation dimension includes the at least one reference difficulty level; The sample's score in the sixth evaluation dimension is determined by the following sub-modules: The eleventh determination submodule is used to determine the recommended difficulty level to which the question of the sample belongs in the at least one reference difficulty level using a large language model; The twelfth determination submodule is used to determine the second degree of matching between the recommended difficulty level and the difficulty level label of the sample; The thirteenth determining submodule is used to determine the dimension score of the sample in the sixth evaluation dimension based on the second matching degree.

[0100] Optionally, the determining module 24 is used to perform weighted summation on each sample based on the sample's dimensional scores in each evaluation dimension and the corresponding weight values ​​of each evaluation dimension to obtain the sample's quality score, which is then used as the quality evaluation result of the sample.

[0101] The following is for reference. Figure 3 The diagram illustrates a structural schematic of an electronic device 600 suitable for implementing the above-described technical solution. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs (Televisions), desktop computers, etc. Figure 3 The electronic device shown is merely an example and should not be construed as limiting its functionality or scope of use.

[0102] like Figure 3As shown, electronic device 600 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. The random access memory 603 also stores various programs and data required for the operation of electronic device 600. The processing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0103] Typically, the following devices can be connected to the input / output interface 605: input devices 606 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 608 including, for example, magnetic tape, hard disk, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0104] In particular, depending on certain circumstances, the processes described in the flowchart above can be implemented as computer software programs. For example, a computer program product is provided, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. This computer program can be downloaded and installed from a network via communication device 609, or installed from storage device 608, or installed from read-only memory 602. When the computer program is executed by processing device 601, it performs the functions defined in the above-described methods.

[0105] It should be noted that the aforementioned computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In one case, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In another case, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.

[0106] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), the internet (e.g., the Internet), and peer-to-peer networks (e.g., ad-hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0107] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0108] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: Obtain the evaluation set to be evaluated, the evaluation set including at least one sample, the sample including questions and answers; Obtain reference information associated with the evaluation set, the reference information including related knowledge information; For each sample, using multiple preset evaluation dimensions and reference information associated with each evaluation dimension, the dimension score of the sample in each evaluation dimension is determined. The evaluation dimensions include at least one first type dimension and at least one second type dimension. The first type dimension is related to the quality of the question, and the second type dimension is related to the quality of the answer. Based on the dimensional scores of the samples in each evaluation dimension, the quality assessment result of each sample in the evaluation set is determined; The second type of dimension includes a first evaluation dimension, which characterizes the degree to which the sample's answer is supported by the associated knowledge information. The reference information associated with the first evaluation dimension includes the associated knowledge information. The sample's dimensional score in the first evaluation dimension is determined in the following way: The answers to the sample are processed by text splitting to obtain multiple split units; for each split unit, it is determined whether the split unit can be supported by the associated knowledge information; based on the number of split units that can be supported by the associated knowledge information, the dimension score of the sample in the first evaluation dimension is determined.

[0109] Computer program code for performing the above operations can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages, as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0110] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0111] The modules mentioned above can be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself; for example, the first acquisition module can also be described as "the module that acquires the evaluation set to be evaluated".

[0112] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Parts (ASSPs), Systems on Chips (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0113] In this context, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0114] The above description is merely illustrative and explains the technical principles employed. Those skilled in the art should understand that the scope of the technical solution is not limited to specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features provided herein that have similar functions.

[0115] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limitations on the scope of the technical solution. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.

[0116] Although the technical solution has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the aforementioned apparatus, the specific manner in which each module performs its operation has already been described in detail in the section concerning the method, and will not be elaborated upon here.

Claims

1. A method for evaluating the quality of a high-quality evaluation set for intelligent marketing agents, comprising: Obtain the evaluation set to be evaluated, the evaluation set including at least one sample, the sample including questions and answers; Obtain reference information associated with the evaluation set, the reference information including related knowledge information; For each sample, using multiple preset evaluation dimensions and reference information associated with each evaluation dimension, the dimension score of the sample in each evaluation dimension is determined. The evaluation dimensions include at least one first type dimension and at least one second type dimension. The first type dimension is related to the quality of the question, and the second type dimension is related to the quality of the answer. Based on the dimensional scores of the samples in each evaluation dimension, the quality assessment result of each sample in the evaluation set is determined; the quality assessment result is used to determine whether the evaluation set is a high-quality evaluation set. The second type of dimension includes a first evaluation dimension, which characterizes the degree to which the sample's answer is supported by the associated knowledge information. The reference information associated with the first evaluation dimension includes the associated knowledge information. The sample's dimensional score in the first evaluation dimension is determined in the following way: The answers to the sample are processed by text splitting to obtain multiple split units; for each split unit, it is determined whether the split unit can be supported by the associated knowledge information; based on the number of split units that can be supported by the associated knowledge information, the dimension score of the sample in the first evaluation dimension is determined. The reference information includes at least one reference question; the first category of dimensions includes a second evaluation dimension, which is used to characterize the degree of difference between the sample's question and the reference question, and the reference information associated with the second evaluation dimension includes the at least one reference question; the dimensional score of the sample in the second evaluation dimension is determined in the following way: For each of the reference questions, the similarity between the sample question and the reference question is determined; the maximum similarity is determined from the similarity; and the dimensional score of the sample in the second evaluation dimension is determined based on the maximum similarity, wherein the dimensional score of the second evaluation dimension is negatively correlated with the maximum similarity.

2. The method according to claim 1, wherein the first type of dimension includes a third evaluation dimension, the third evaluation dimension being used to characterize the completeness of the problem in the sample; the reference information includes at least one evaluation rule, and the reference information associated with the third evaluation dimension includes at least one evaluation rule; The dimensional score of the sample in the third evaluation dimension was determined in the following way: Obtain the initial score for the third evaluation dimension, and the deduction value corresponding to each evaluation rule; Based on each of the evaluation rules, determine whether the problem of the sample satisfies the evaluation rule; If the problem of the sample does not meet the evaluation rules, the total deduction value is determined according to the deduction value corresponding to the evaluation rules; Based on the initial score and the total deduction, the dimensional score of the sample in the third evaluation dimension is determined.

3. The method according to claim 1, wherein the first type of dimension includes a fourth evaluation dimension, the fourth evaluation dimension being used to characterize the degree of correlation between the problem of the sample and the associated knowledge information, and the reference information associated with the fourth evaluation dimension includes the associated knowledge information; The dimensional score of the sample in the fourth evaluation dimension was determined in the following manner: Using a large language model, the relevance score between the sample's question and the associated knowledge information is determined, which is then used as the sample's dimensional score in the fourth evaluation dimension.

4. The method according to claim 1, wherein the sample further includes category labels; the reference information includes at least one reference category; the evaluation dimension further includes a fifth evaluation dimension, and the reference information associated with the fifth evaluation dimension includes the at least one reference category; The dimensional score of the sample in the fifth evaluation dimension was determined in the following way: Using a large language model, determine the recommendation category to which the question of the sample belongs in the at least one reference category; Determine the first degree of matching between the recommended category and the category label of the sample; Based on the first matching degree, the dimensional score of the sample in the fifth evaluation dimension is determined.

5. The method according to claim 1, wherein the sample further includes a difficulty level label; the reference information includes at least one reference difficulty level; the evaluation dimension further includes a sixth evaluation dimension, and the reference information associated with the sixth evaluation dimension includes the at least one reference difficulty level; The sample's score in the sixth evaluation dimension was determined in the following way: Using a large language model, determine the recommended difficulty level to which the question of the sample belongs in the at least one reference difficulty level; Determine the second degree of matching between the recommended difficulty level and the difficulty level label of the sample; Based on the second matching degree, the dimensional score of the sample in the sixth evaluation dimension is determined.

6. The method according to claim 1, wherein determining the quality assessment result of each sample in the evaluation set based on the dimensional scores of the sample in each evaluation dimension includes: For each sample, a quality score is obtained by weighted summation based on the sample's dimensional scores in each evaluation dimension and the corresponding weight values ​​of each evaluation dimension, which serves as the quality evaluation result of the sample.

7. A quality assessment device for a high-quality evaluation set for intelligent marketing agents, comprising: The first acquisition module is used to acquire the evaluation set to be evaluated, the evaluation set including at least one sample, the sample including questions and answers; The second acquisition module is used to acquire reference information associated with the evaluation set, the reference information including associated knowledge information; An evaluation module is used to determine the dimension score of each sample in each of the evaluation dimensions using multiple preset evaluation dimensions and reference information associated with each evaluation dimension. The evaluation dimensions include at least one first type dimension and at least one second type dimension. The first type dimension is related to the quality of the question, and the second type dimension is related to the quality of the answer. The determination module is used to determine the quality assessment result of each sample in the evaluation set based on the dimensional scores of the sample in each evaluation dimension; the quality assessment result is used to determine whether the evaluation set is a high-quality evaluation set. The second type of dimension includes a first evaluation dimension, which characterizes the degree to which the sample's answer is supported by the associated knowledge information. The reference information associated with the first evaluation dimension includes the associated knowledge information. The dimension score of the sample in the first evaluation dimension is determined by: performing text segmentation on the sample's answer to obtain multiple segmentation units; for each segmentation unit, determining whether the segmentation unit can be supported by the associated knowledge information; and determining the dimension score of the sample in the first evaluation dimension based on the number of segmentation units that can be supported by the associated knowledge information. The reference information includes at least one reference question; the first category of dimensions includes a second evaluation dimension, which is used to characterize the degree of difference between the sample's question and the reference question, and the reference information associated with the second evaluation dimension includes the at least one reference question; the dimensional score of the sample in the second evaluation dimension is determined in the following way: For each of the reference questions, the similarity between the sample question and the reference question is determined; the maximum similarity is determined from the similarity; and the dimensional score of the sample in the second evaluation dimension is determined based on the maximum similarity, wherein the dimensional score of the second evaluation dimension is negatively correlated with the maximum similarity.

8. A computer-readable medium having a computer program stored thereon, wherein, When executed by a processing device, the computer program performs the steps of the method according to any one of claims 1-6.

9. An electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-6.

10. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.