Evaluation method, evaluation device, electronic equipment, storage medium and program product
By performing out-of-distribution detection and label prediction separately and employing a multi-dimensional evaluation method, the problems of task interference and incomplete evaluation in the evaluation of closed-set label systems are solved, achieving an accurate and comprehensive evaluation of the label system and providing a basis for optimization decisions.
Patent Information
- Application Number
- CN202511692965.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies suffer from task interference and incomplete evaluation methods in the evaluation of closed-set labeling systems, leading to inaccurate evaluation results.
By employing a method that separates out-of-distribution detection and label prediction, a multi-dimensional evaluation is conducted using the out-of-distribution detection results and label prediction results to obtain a comprehensive evaluation result, thereby inferring the root cause of defects in the labeling system.
It enables accurate and comprehensive evaluation of the closed-set labeling system, provides targeted optimization decision-making basis, and improves the completeness and accuracy of the labeling system.
Smart Images

Figure CN121502260A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to an evaluation method, evaluation apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] In practical applications of machine learning and artificial intelligence, labeling systems are a crucial foundation for algorithm training and evaluation. In related technologies, large language models can be used to evaluate closed-set labeling systems to determine whether they meet the application requirements. However, these evaluation methods are prone to interference when performing out-of-distribution detection and closed-set classification tasks, and their evaluation dimensions are not comprehensive enough to obtain accurate and comprehensive evaluation results. Summary of the Invention
[0003] This disclosure provides an evaluation method, an evaluation apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0004] In a first aspect, this disclosure provides an evaluation method applied to a preset label evaluation model. The evaluation method includes: performing out-of-distribution detection on multiple first sample data based on a preset label set to obtain an out-of-distribution detection result for each first sample data, wherein the out-of-distribution detection result is used to characterize whether the predicted label corresponding to the first sample data belongs to the label set; determining second sample data and a first evaluation result for the label set based on the out-of-distribution detection result, wherein the second sample data is the first sample data whose predicted label belongs to the label set; evaluating the label set using the second sample data to obtain a second evaluation result for the label set; and determining a comprehensive evaluation result for the label set based on the first evaluation result and the second evaluation result.
[0005] Secondly, this disclosure provides an evaluation device applied to a preset label evaluation model. The evaluation device includes: a detection module, used to perform out-of-distribution detection on multiple first sample data based on a preset label set, to obtain an out-of-distribution detection result for each first sample data, wherein the out-of-distribution detection result is used to characterize whether the predicted label corresponding to the first sample data belongs to the label set; a first evaluation module, used to determine second sample data and a first evaluation result for the label set based on the out-of-distribution detection result, wherein the second sample data is the first sample data whose predicted label belongs to the label set; a second evaluation module, used to evaluate the label set using the second sample data, to obtain a second evaluation result for the label set; and a determination module, used to determine a comprehensive evaluation result for the label set based on the first evaluation result and the second evaluation result.
[0006] Thirdly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the evaluation method described above.
[0007] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described evaluation method.
[0008] Fifthly, this disclosure provides a computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the evaluation method described above.
[0009] The evaluation method provided in this embodiment separates out-of-distribution detection and label prediction into separate operations, ensuring that they do not interfere with each other. Furthermore, it can perform multi-dimensional evaluation of the label set based on the out-of-distribution detection results and label prediction results, obtaining corresponding first and second evaluation results. Then, a comprehensive evaluation result is obtained based on the first and second evaluation results. This allows for the inference of the root causes of label system defects based on the comprehensive evaluation results, providing a basis for decision-making in the subsequent optimization of the label system.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0012] Figure 1 This diagram illustrates an application scenario of the evaluation method provided in the embodiments of this disclosure.
[0013] Figure 2 A flowchart of an evaluation method provided for an embodiment of this disclosure.
[0014] Figure 3 This is a schematic diagram illustrating the processing steps of an evaluation method provided in an embodiment of this disclosure.
[0015] Figure 4 This is a block diagram of an evaluation apparatus provided in an embodiment of the present disclosure.
[0016] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0017] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0018] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0019] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0021] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0022] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information in this technical solution comply with relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution follows relevant national laws and regulations (e.g., the "Information Security Technology - Personal Information Security Specification"). For example, appropriate measures are taken for personal information access control; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; and explicit identity targeting is eliminated when using personal information to avoid precisely identifying specific individuals.
[0023] In practical applications of machine learning and artificial intelligence, labeling systems are a crucial foundation for algorithm training and evaluation. Closed-set label taxonomy is a common form of label organization, consisting of a predefined, finite set of labels. Its advantage lies in ensuring the explicitness of the classification task. However, closed-set label taxonomy often suffers from problems such as incompleteness, non-orthogonality, and difficulty in dynamic expansion in real-world applications.
[0024] In related technologies, closed-set labeling systems can be evaluated to determine whether they meet usage requirements. For example, Large Language Models (LLMs), which have emerged in recent years, possess powerful semantic understanding and reasoning capabilities, providing new approaches for the design and optimization of labeling systems. However, directly using LLMs for annotation and evaluation on closed-set labeling systems still faces the following problems: 1. Task interference problem: When performing out-of-distribution (OOD) and closed-set classification tasks, LLMs are prone to reducing the reliability of labeling results due to target conflicts between the two tasks; 2. Limitations of evaluation methods: Most current evaluation methods are based on a single metric (such as overall accuracy), which cannot comprehensively reflect the shortcomings of the labeling system (such as low coverage or poor orthogonality).
[0025] In view of this, the present disclosure provides an evaluation method, evaluation apparatus, electronic device, computer-readable storage medium, and computer program product that can accurately and comprehensively evaluate a closed-set labeling system, thereby providing targeted decision-making basis for the optimization of the closed-set labeling system.
[0026] The evaluation method according to the embodiments of this application can be executed by an electronic device such as a terminal device or a server. The terminal device can be an in-vehicle device, user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in memory. Alternatively, the method can be executed by a server.
[0027] Figure 1 This diagram illustrates an application scenario of the evaluation method provided in the embodiments of this disclosure.
[0028] like Figure 1 As shown, the application scenario of this disclosure embodiment may include terminal device 101, network 103, and server 102. Network 103 is used as a medium to provide a communication link between terminal device 101 and server 102. Network 103 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0029] Users can use terminal device 101 to interact with server 102 via network 103 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0030] Terminal device 101 can be various electronic devices with a display screen and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0031] Server 102 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal device 101 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0032] It should be noted that the evaluation method provided in this embodiment can be executed by server 102. Accordingly, the evaluation method and apparatus provided in this embodiment can be located in server 102. The evaluation method and apparatus provided in this embodiment can also be executed by a server or server cluster that is different from server 102 and capable of communicating with terminal device 101 and / or server 102. Accordingly, the evaluation method and apparatus provided in this embodiment can also be located in a server or server cluster that is different from server 102 and capable of communicating with terminal device 101 and / or server 102.
[0033] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0034] In a first aspect, embodiments of this disclosure provide an evaluation method.
[0035] Figure 2 A flowchart illustrating an evaluation method provided in an embodiment of this disclosure. (Refer to...) Figure 2 This method can be applied to a pre-defined label evaluation model and may include the following steps:
[0036] Step S21: Perform out-of-distribution detection on multiple first sample data based on a preset label set to obtain the out-of-distribution detection result for each first sample data. The out-of-distribution detection result is used to characterize whether the predicted label corresponding to the first sample data belongs to the label set.
[0037] Step S22: Based on the out-of-distribution detection results, determine the second sample data and the first evaluation result for the label set. The second sample data is the first sample data predicting that the label belongs to the label set.
[0038] Step S23: Evaluate the label set using the second sample data to obtain the second evaluation result of the label set;
[0039] Step S24: Determine the comprehensive evaluation result of the tag set based on the first evaluation result and the second evaluation result.
[0040] Therefore, in this embodiment, out-of-distribution detection and label prediction are performed separately, and the two will not affect each other. Furthermore, the label set can be evaluated in multiple dimensions based on the out-of-distribution detection results and the label prediction results to obtain the corresponding first evaluation result and second evaluation result. Then, a comprehensive evaluation result is obtained based on the first evaluation result and the second evaluation result. Thus, the root cause of the label system defects can be inferred from the comprehensive evaluation result, providing a decision basis for the subsequent optimization of the label system.
[0041] The evaluation method of the embodiments of this disclosure will be described in detail below.
[0042] In some alternative embodiments, the tag set is a pre-built closed tag system that includes multiple tags and may also include auxiliary information such as tag definitions.
[0043] For example, the tag set includes the following tags: cat, dog, cow, sheep, horse, mouse, rabbit.
[0044] Furthermore, the tag set is a closed-set tag system to be evaluated. Through subsequent evaluation, it is necessary to clarify whether the tags in the tag set are complete and whether the boundaries between tags are clear, so that the tag set can be optimized in a more targeted manner to obtain a more complete and accurate tag set.
[0045] In some alternative embodiments, a sample set may be constructed, which includes multiple first sample data, which are sample data to be labeled or to be predicted by labels.
[0046] In some alternative embodiments, a sample set containing multiple first sample data can be constructed from the original sample set through a sampling method.
[0047] It should be noted that, to obtain comprehensive and accurate evaluation results, it is best to construct a sample set that covers all labels in the label set. In other words, the labels of multiple first sample data should at least cover all labels in the label set. Of course, there may be cases where one or more labels of the first sample data do not belong to the label set.
[0048] In some optional embodiments, out-of-distribution detection can be performed on multiple first sample data based on a preset label set to obtain an out-of-distribution detection result for each first sample data. The out-of-distribution detection result is used to characterize whether the predicted label corresponding to the first sample data belongs to the label set. In other words, out-of-distribution detection can determine whether the first sample data belongs to the label system of the label set. Out-of-distribution detection is mainly used to identify uncovered data samples; corresponding to the embodiments of this disclosure, out-of-distribution detection mainly identifies first sample data that does not belong to the label set.
[0049] For example, if the label set includes the following labels: cat, dog, cow, sheep, horse, mouse, rabbit, and the animal corresponding to the first sample data is a snake, then out-of-distribution detection can determine that the first sample data does not belong to the label set, or in other words, the first sample data does not belong to the label system corresponding to the label set, and the first sample data belongs to the other category of sample data.
[0050] In some alternative embodiments, out-of-distribution detection can be achieved using a classifier approach. For example, training data (such as first sample data) can be input into a preset classifier to obtain the classification result of the test sample. If the test sample is classified into a non-existent category, the sample may be an OOD sample. Each classifier corresponds to one label.
[0051] In some alternative embodiments, out-of-distribution detection can be achieved based on feature analysis methods, which primarily utilize the feature differences between training data (such as first sample data) and test data (such as any label in a label set) to identify OOD samples. For example, the Euclidean distance between the test sample and the training data can be calculated, and whether the test sample belongs to the OOD sample can be determined based on the Euclidean distance.
[0052] It should be noted that the above implementation methods for off-distribution detection are merely illustrative examples, and the embodiments disclosed herein do not impose any limitations on them.
[0053] In some optional embodiments, after obtaining the out-of-distribution detection results of the first sample data, the second sample data can be selected from the first sample data based on the out-of-distribution detection results, thereby obtaining the first evaluation result for the label set.
[0054] In some optional embodiments, the first evaluation result includes the label coverage ratio; correspondingly, determining the second sample data and the first evaluation result for the label set based on the out-of-distribution detection result includes: selecting the second sample data from multiple first sample data based on the out-of-distribution detection result, wherein the out-of-distribution detection result of the second sample data indicates that the predicted label corresponding to the second sample data belongs to the label set, and the number of second sample data is multiple; determining the label coverage ratio based on the ratio of the first total number of second sample data to the second total number of first sample data.
[0055] Therefore, we can select second sample data that belong to the label set distribution from multiple first sample data. Based on this, we determine the first total number of second sample data and the second total number of first sample data, and calculate the ratio of the first total number to the second total number. This ratio is used as the label coverage ratio.
[0056] For example, the label coverage ratio can be represented by the following formula:
[0057] ratio=(N-N_ood) / N=1-N_ood / N
[0058] Where N represents the second total number of the first sample data, N_ood represents the out-of-distribution detection samples in the first sample data, and (N-N_ood) represents the first total number of the second sample data.
[0059] It should be noted that after determining multiple first sample data and labeling each first sample data with a true label, the true label coverage percentage of these first sample data is a fixed value. This fixed value may be the same as or different from the label coverage percentage in the first evaluation result. In other words, the label coverage percentage in the first evaluation result is determined through out-of-distribution detection. If it is completely accurate, its value is the same as the aforementioned fixed value; if it is not completely accurate, its value is different from the aforementioned fixed value.
[0060] Therefore, the first evaluation result reflects the coverage or completeness of the label set; in other words, it assesses the label set from the perspective of its completeness. Theoretically, the higher the completeness of the label set as indicated by the first evaluation result, the higher the probability of a successful label recognition when using that label set.
[0061] Furthermore, the second sample data can be used to evaluate the label set in other dimensions, thereby obtaining a second evaluation result for the label set.
[0062] In some optional embodiments, the label set is evaluated using the second sample data to obtain a second evaluation result for the label set, including: performing label identification on the second sample data based on the label set to determine the predicted label for each second sample data, with the preset label being the label in the label set; and determining the second evaluation result for the label set based on the predicted labels and real labels of multiple second sample data.
[0063] The process involves label identification of the second sample data based on a label set, specifically selecting the label that best matches the second sample data from multiple labels in the label set as its predicted label. Further, to determine the accuracy of the predicted label, the true label for each second sample data point needs to be determined beforehand. The true label for the second sample data can be determined manually or by using a pre-defined label recognition model followed by manual verification to ensure the accuracy of the true labels.
[0064] Since the label set may have problems such as ambiguous label definitions and unclear label boundaries, there is a possibility of misidentifying and predicting labels. The label set can be evaluated by analyzing the misidentification situation, and the corresponding evaluation result is the second evaluation result.
[0065] In some optional embodiments, the second evaluation result includes at least one of overall accuracy, category accuracy, and obfuscated label results, which evaluate the label set from different perspectives.
[0066] In some optional embodiments, the second evaluation result includes overall accuracy, category accuracy, and confusion label result; correspondingly, based on the predicted labels and true labels of multiple second sample data, a second evaluation result for the label set is determined, including: for any first true label among the multiple true labels, based on the predicted labels and true labels of the multiple second sample data, determining a first number of true examples and a second number of false negative examples corresponding to the first true label, where a true example indicates that the predicted label of the second sample data is consistent with the first true label, and a false negative example indicates that the true label of the second sample data is consistent with the first true label, but the predicted label of the second sample data is inconsistent with the first true label. The actual labels are inconsistent; the overall accuracy is determined by the ratio of the sum of multiple first quantities to the total number of multiple second sample data; the category accuracy corresponding to each first true label is determined by the first quantity and the second quantity; candidate label pairs are constructed based on the second sample data where the predicted label is inconsistent with the true label, and the occurrence frequency of each candidate label pair is determined. Each candidate label pair includes the true label and the incorrectly predicted label; for any candidate label pair, it is determined whether the candidate label pair belongs to the target label pair by the ratio of the occurrence frequency of the candidate label pair to the total number of true labels corresponding to the candidate label pair. The obfuscated label results include the target label pair.
[0067] Therefore, for any first true label Lt1 among multiple true labels, assuming the true label corresponding to the second sample data d1 is also Lt1, if the predicted label of the second sample data d1 is also Lt1, it is a true example; otherwise, if the predicted label is not Lt1, it is a false negative. Furthermore, if the true label corresponding to the second sample data d2 is Lt2, but its predicted label is Lt1, it is a false positive. Further, we can count the number of true examples to obtain the first count, and count the number of false negative examples to obtain the second count.
[0068] Based on this, for each first true label, there exists a first quantity. By summing the first quantities of all first true labels, we can obtain the sum of multiple first quantities. Then, by calculating the ratio between this sum and the total number of all second sample data, we can obtain the overall accuracy.
[0069] For example, the overall accuracy (Acc) can be calculated using the following formula:
[0070]
[0071] Where N represents the total number of second sample data, and n represents the sequence number of the second sample data. This indicates the predicted label for the second sample data. This represents the true label of the second sample data. This represents the total number of second sample data whose predicted labels match the true labels, which is also the sum of multiple first quantities.
[0072] Based on this, the category accuracy corresponding to each first true label can be determined according to the first quantity and the second quantity.
[0073] For example, the class accuracy Acc_Li of any first true label Li can be characterized in the following form:
[0074] Acc_Li = TP_i / (TP_i + FN_i)
[0075] Where TP_i represents the first number of true instances corresponding to the first true label Li, and FN_i represents the second number of false negative instances corresponding to the first true label Li.
[0076] For example, the class accuracy Acc_Li of any first true label Li can also be characterized in the following form:
[0077] Acc_Li = TP_i / (TP_i + FN_i + a)
[0078] Where TP_i represents the first number of true examples corresponding to the first true label Li, FN_i represents the second number of false negative examples corresponding to the first true label Li, and a is a numerically stable term, which is usually a very small positive number.
[0079] It should be noted that category accuracy can be used to reflect the differences in predictions for different labels, and can help to discover problems such as ambiguous label boundaries, unclear definitions, or interference between similar labels in a label set.
[0080] Moreover, if the overall accuracy has reached the set evaluation baseline, and the category accuracy of a certain real label in the label set is significantly lower than the average analog accuracy, it can be reasonably inferred that the real label has problems such as relative definition ambiguity or unclear boundaries in the current label system. In subsequent optimization, the real label can be optimized first.
[0081] In some optional embodiments, the portion of the second sample data where the predicted labels do not match the true labels can be filtered out, and corresponding candidate label pairs can be constructed based on these incorrectly predicted labels. For example, candidate label pairs may include <first label, second label>. The frequency of occurrence of each candidate label pair is counted, and those candidate label pairs with higher frequency of occurrence are selected as target label pairs. In other words, the target label pairs are high-frequency error label pairs, and during identification, the first label is usually incorrectly identified as the second label.
[0082] For example, a tag-tag confusion matrix can be generated, which includes multiple candidate tag pairs and marks the tag pairs with a high confusion ratio. These tag pairs are the target tag pairs.
[0083] For example, for a certain candidate label pair If a large proportion of the second sample data with the true label Li is predicted as Lj (i.e., the predicted label is Lj), then this candidate label pair can be determined as the target label pair. This can be achieved by setting a preset threshold. Filtering target label pairs can be represented in the following form:
[0084]
[0085] Therefore, if the ratio of the frequency of occurrence of a candidate tag pair to the number of actual tags in the first candidate tag pair is greater than a preset threshold, then the candidate tag pair is considered to belong to the target tag pair, which is a high-frequency and easily confused tag pair. In other words, for the target tag pair (Li, Lj), it is considered that there is semantic confusion between Li and Lj. In subsequent problem analysis and optimization suggestions, it can be suggested to further clarify the definition of the target tag pair or define its boundaries, or, based on actual business needs, consider whether to merge, rename, or supplement examples of the two tags in the target tag pair.
[0086] In summary, the embodiments of this disclosure have designed a set of structured evaluation metrics, including overall accuracy, category accuracy, confusion label results, and label coverage ratio, to systematically measure the labeling effect and assist in identifying problems such as ambiguous label boundaries and redundant definitions.
[0087] In some optional embodiments, after determining the comprehensive evaluation result of the tag set based on the first evaluation result and the second evaluation result, the evaluation method may further include: determining problem analysis information and / or optimization suggestion information for the tag set based on the comprehensive evaluation result.
[0088] For example, if the value of the comprehensive evaluation result representing the percentage of label coverage is low, it indicates that the label set does not cover the entire range and has poor completeness.
[0089] For example, if the overall accuracy of the comprehensive evaluation result is low, it indicates that the label setting method of the label set is unreasonable or inaccurate, including problems such as significant semantic overlap between labels.
[0090] For example, category accuracy can reflect the difference in prediction for different labels. If the overall evaluation result indicates that the accuracy of a certain category is low, it means that the corresponding label may have problems such as ambiguous label boundaries, unclear definition, and interference between similar labels.
[0091] For example, if the comprehensive evaluation results indicate the presence of easily confused label pairs (such as target label pairs), it indicates that the definitions of these label pairs are unclear or the boundaries are ambiguous.
[0092] Based on this, targeted optimization suggestions can be provided for the aforementioned problems that may exist in the tag set, including but not limited to adding tags, improving tag definitions, and improving tag boundaries.
[0093] In some optional embodiments, the evaluation method can be applied to a preset label evaluation model; wherein the label evaluation model performs out-of-distribution detection processing based on a first preset prompt information and performs label prediction processing based on a second preset prompt information.
[0094] In other words, the evaluation method of this embodiment can be executed by a preset label evaluation model, and the label evaluation model is implemented by a first preset prompt information when performing out-of-distribution detection, and by a second preset prompt information when performing label prediction.
[0095] It should be noted that by comparing the actual label coverage ratio with the label coverage ratio obtained from out-of-distribution detection, the ability of the label evaluation model to exclude samples from non-labeled systems can be measured, reflecting its ability to "refuse recognition" or "determine fuzzy boundaries." Furthermore, this comparison result can evaluate the design quality of the first prompt information, which is beneficial for subsequent optimization of the OOD judgment rules.
[0096] For example, the label evaluation model can perform several of the following tasks:
[0097] (1) Provide a few positive examples and "other class" examples to the label evaluation model through the few-shot prompt method, guide the label evaluation model to identify samples outside the label system and label them as "other class". The output structure can be set to "Label: Other class; Reason: XXX" to facilitate subsequent cluster analysis and label expansion suggestions.
[0098] (2) Under the premise of knowing the entire set of labels (i.e. all real labels), construct a closed set prompt (i.e. the second preset prompt information) and clarify the task objective as "select the best matching item from the given set of labels". The given set of labels is the set of labels composed of real labels. The label style and granularity can be controlled by few-shot examples, thereby improving the consistency and interpretability of the model's judgment.
[0099] It should be noted that, in order to improve the consistency and robustness of the inference model, multiple Prompt templates can be constructed and executed for multiple first sample data to obtain multiple rounds of output results. By comparing the consistency of the multiple rounds of output results, the most stable Prompt template can be selected as the template used in subsequent inference, thereby effectively mitigating systematic errors caused by the design deviation of the prompt.
[0100] (3) After the predicted label of each second sample data is output by inference and aligned with the real label, the category accuracy index of each real label can be calculated by label dimension to generate a mapping table between the real label and the category accuracy, which can be used to help identify label categories with obvious performance differences.
[0101] (4) Construct a label-label confusion matrix. For target label pairs where the proportion of the second sample data whose true label is A and which is predicted as B exceeds a set threshold, a label confusion prompt message can be generated to prompt the user that there may be boundary intersection or definition ambiguity.
[0102] (5) It can automatically generate optimization suggestions for the structured label system by combining the distribution of category accuracy, label-label confusion matrix and OOD sample clustering results, including but not limited to: definition clarification suggestions, label splitting / merging suggestions, label mutual exclusion suggestions, new label candidate suggestions, etc. The generated suggestions can be output in data formats such as JSON.
[0103] In summary, the evaluation method (or label evaluation model) in this embodiment adopts a two-stage labeling structure, which divides the entire evaluation process into two sequential tasks: the first stage is the "out-of-systems sample identification task (OOD identification)," which determines whether the first sample data belongs to the existing label set (i.e., the label system composed of the real labels of multiple first sample data); the second stage is the "closed-set classification task," which only performs label attribution judgment on samples that are not "other classes" (i.e., the second sample data), and selects the most matching real label from the label set as the predicted label for the second sample data. This sequential processing structure can effectively avoid problems such as label drift and misclassification, improve the stability of the overall judgment, and thus improve the accuracy and stability of label evaluation.
[0104] It should be noted that, compared to the traditional approach that relies entirely on human experience to build and maintain a tag system, this invention, by introducing the capabilities of a large language model, achieves tag structure analysis, problem localization, and suggestion generation driven by a validation set. This significantly reduces the cost of manual tag design and iteration, and improves the efficiency of data asset construction. Furthermore, since the evaluation method of this disclosure does not depend on a specific model architecture when executing based on the tag evaluation model, it is applicable to mainstream general-purpose large language models (such as GPT series, Claude, GLM, Wenxin, etc.) and can be deployed in various industry text scenarios (such as financial collection, customer service, product review analysis, etc.), possessing strong practical implementation capabilities, versatility, and platform compatibility.
[0105] Figure 3 This is a schematic diagram illustrating the processing steps of an evaluation method provided in an embodiment of this disclosure. (Refer to...) Figure 3 The label set includes multiple labels and belongs to the set to be evaluated; the original sample set includes sample data of multiple labels to be labeled. Multiple first sample data are selected from the original sample set through sampling so that the label set can be evaluated subsequently using multiple first sample data.
[0106] like Figure 3 As shown, firstly, out-of-distribution detection is performed on multiple first sample data based on the label set to obtain the out-of-distribution detection result for each first sample data, thereby characterizing whether the first sample data belongs to the label system covered by the label set. Further, based on the out-of-distribution detection results, second sample data is selected from the multiple first sample data. The out-of-distribution detection result of the second sample data indicates that the predicted label corresponding to the second sample data belongs to the label set. Then, the label coverage ratio is determined based on the ratio of the first total number of second sample data to the second total number of first sample data.
[0107] Based on this, for each second sample data, label identification is performed on the second sample data according to the label set, and the label that best matches the second sample data is selected from the label set as the predicted label for that second sample data.
[0108] It should be noted that for each first or second sample data point, its true label has been pre-determined through methods such as manual annotation. This true label is the accurate label of the second sample data, and it belongs to the label set. Furthermore, to ensure a comprehensive evaluation of the label set, the first sample data selected when sampling the original sample set should cover all labels in the label set.
[0109] Furthermore, for any first true label among multiple true labels, based on the predicted labels and true labels of multiple second sample data, determine the first number of true examples and the second number of false negative examples corresponding to the first true label; determine the overall accuracy based on the ratio of the sum of the multiple first numbers to the total number of multiple second sample data; determine the category accuracy corresponding to each first true label based on the first and second numbers; construct candidate label pairs based on the second sample data where the predicted labels and true labels are inconsistent, and determine the occurrence frequency of each candidate label pair, where each label pair includes the true label and the incorrectly predicted label; for any candidate label pair, determine whether the candidate label pair belongs to the target label pair based on the ratio of the occurrence frequency of the candidate label pair to the total number of true labels corresponding to the candidate label pair, and the obfuscated label results include the target label pair.
[0110] In some optional embodiments, the evaluation method can be applied to a preset label evaluation model; in other words, the evaluation of a set of labels can be performed through the label evaluation model.
[0111] In some optional embodiments, the label evaluation model performs out-of-distribution detection processing based on a first preset prompt and performs label prediction processing based on a second preset prompt.
[0112] It should be noted that the reason for using different preset prompts to perform out-of-distribution detection and label prediction is mainly because performing both out-of-distribution detection and label prediction simultaneously within the same preset prompt can easily lead to cognitive conflict or inference chain confusion in the label evaluation model, thereby affecting the stability and reliability of the processing results. Therefore, in this embodiment, the out-of-distribution detection task and the label prediction task are explicitly separated and a serialized structural process is constructed to avoid task conflicts at the model level and improve the reliability of labeling.
[0113] Based on the label coverage ratio, overall accuracy, category accuracy, and confusion label results, the comprehensive evaluation result of the label set can be determined. Based on the comprehensive evaluation result, problem analysis information and / or optimization suggestions for the label set can be determined. Then, based on the problem analysis information and / or optimization suggestions, the label set can be optimized accordingly to obtain a more complete and reasonable label set.
[0114] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0115] In addition, this disclosure also provides evaluation apparatus, electronic equipment, and computer-readable storage medium, all of which can be used to implement any of the evaluation methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding section of the method and will not be repeated here.
[0116] Figure 4 This is a block diagram of an evaluation apparatus provided in an embodiment of the present disclosure.
[0117] Reference Figure 4 This disclosure provides an evaluation apparatus for use with a preset label evaluation model. The evaluation apparatus 400 includes:
[0118] The detection module 401 is used to perform out-of-distribution detection on multiple first sample data based on a preset label set, and obtain the out-of-distribution detection result for each first sample data. The out-of-distribution detection result is used to characterize whether the predicted label corresponding to the first sample data belongs to the label set.
[0119] The first evaluation module 402 is used to determine the second sample data and the first evaluation result for the label set based on the out-of-distribution detection result. The second sample data is the first sample data that predicts that the label belongs to the label set.
[0120] The second evaluation module 403 is used to evaluate the label set using the second sample data to obtain a second evaluation result for the label set;
[0121] The determination module 404 is used to determine the comprehensive evaluation result of the tag set based on the first evaluation result and the second evaluation result.
[0122] In summary, in this embodiment of the disclosure, the label evaluation model separates out-of-distribution detection and label prediction, and the two do not affect each other. Furthermore, the label set can be evaluated in multiple dimensions based on the out-of-distribution detection results and the label prediction results to obtain the corresponding first evaluation result and second evaluation result. Then, a comprehensive evaluation result is obtained based on the first evaluation result and the second evaluation result. Thus, the root cause of the label system defects can be inferred from the comprehensive evaluation result, providing a decision basis for the subsequent optimization of the label system.
[0123] Each module in the aforementioned evaluation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can invoke and execute the operations corresponding to each module.
[0124] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0125] Reference Figure 5This disclosure provides an electronic device comprising: at least one processor 501; at least one memory 502; and one or more I / O interfaces 503 connected between the processor 501 and the memory 502; wherein the memory 502 stores one or more computer programs executable by the at least one processor 501, the one or more computer programs being executed by the at least one processor 501 to enable the at least one processor 501 to perform the above-described evaluation method.
[0126] The modules in the aforementioned electronic devices can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0127] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the evaluation method described above. The computer-readable storage medium may be volatile or non-volatile.
[0128] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described evaluation method.
[0129] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0130] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0131] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0132] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0133] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0134] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0135] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0136] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0138] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. An evaluation method, characterized in that, The method, applied to a pre-defined label evaluation model, includes: Based on a preset label set, out-of-distribution detection is performed on multiple first sample data to obtain the out-of-distribution detection result for each first sample data. The out-of-distribution detection result is used to characterize whether the predicted label corresponding to the first sample data belongs to the label set. Based on the out-of-distribution detection results, a second sample data and a first evaluation result for the label set are determined, wherein the second sample data is the first sample data predicting that the label belongs to the label set; The tag set is evaluated using the second sample data to obtain a second evaluation result for the tag set; Based on the first evaluation result and the second evaluation result, a comprehensive evaluation result for the tag set is determined.
2. The method according to claim 1, characterized in that, The step of evaluating the tag set using the second sample data to obtain a second evaluation result for the tag set includes: The second sample data is labeled according to the label set to determine the predicted label for each second sample data, wherein the preset label is a label in the label set; A second evaluation result for the label set is determined based on the predicted labels and real labels of multiple second sample data.
3. The method according to claim 2, characterized in that, The second evaluation result includes at least one of overall accuracy, category accuracy, and confusion label results; The step of determining the second evaluation result of the label set based on the predicted labels and real labels of multiple second sample data includes: For any first real label among multiple real labels, based on the predicted labels and real labels of multiple second sample data, determine the first number of true examples and the second number of false negative examples corresponding to the first real label. The true examples indicate that the predicted label of the second sample data is consistent with the first real label, and the false negative examples indicate that the real label of the second sample data is consistent with the first real label, but the predicted label of the second sample data is inconsistent with the first real label. The overall accuracy is determined based on the ratio of the sum of the plurality of first quantities to the total number of the plurality of second sample data; Based on the first quantity and the second quantity, determine the category accuracy corresponding to each first true label; Based on the second sample data where the predicted labels are inconsistent with the true labels, alternative label pairs are constructed, and the occurrence frequency of each alternative label pair is determined. Each alternative label pair includes the true label and the predicted label that was incorrectly predicted. For any candidate tag pair, the ratio of the number of occurrences of the candidate tag pair to the total number of real tags corresponding to the candidate tag pair is used to determine whether the candidate tag pair belongs to the target tag pair. The obfuscated tag result includes the target tag pair.
4. The method according to claim 1, characterized in that, The first assessment result includes the percentage of label coverage; The step of determining the second sample data and the first evaluation result for the label set based on the out-of-distribution detection results includes: Based on the out-of-distribution detection results, the second sample data is selected from multiple first sample data. The out-of-distribution detection results of the second sample data indicate that the predicted label corresponding to the second sample data belongs to the label set, and the number of second sample data is multiple. The label coverage ratio is determined based on the ratio of the first total quantity of the second sample data to the second total quantity of the first sample data.
5. The method according to claim 1, characterized in that, After determining the comprehensive evaluation result of the tag set based on the first evaluation result and the second evaluation result, the method further includes: Based on the comprehensive evaluation results, problem analysis information and / or optimization suggestions are determined for the tag set.
6. The method according to any one of claims 1 to 5, characterized in that, The label evaluation model performs out-of-distribution detection processing based on the first preset prompt information and label prediction processing based on the second preset prompt information.
7. An evaluation device, characterized in that, The device, applied to a preset label evaluation model, includes: The detection module is used to perform out-of-distribution detection on multiple first sample data based on a preset label set, and to obtain the out-of-distribution detection result for each first sample data. The out-of-distribution detection result is used to characterize whether the predicted label corresponding to the first sample data belongs to the label set. The first evaluation module is used to determine the second sample data and the first evaluation result for the label set based on the out-of-distribution detection result, wherein the second sample data is the first sample data predicting that the label belongs to the label set; The second evaluation module is used to evaluate the tag set using the second sample data to obtain a second evaluation result for the tag set; The determination module is used to determine the comprehensive evaluation result of the tag set based on the first evaluation result and the second evaluation result.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, such that the at least one processor is able to perform the evaluation method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the evaluation method as described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the evaluation method as described in any one of claims 1-6.