Method, device and electronic equipment for optimizing a label system

CN120994777BActive Publication Date: 2026-08-21北京中关村科金技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511074307.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2026-08-21
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

[0003]本申请实施例提供一种标签体系的优化方法、装置及电子设备,以解决标签体系的质量较差的问题

Benefits of technology

[0018] In this embodiment, historical session records are labeled using tags in a tagging system, and quality parameters of the tagging system are obtained based on the labeled historical session records. The tagging system is then optimized based on the quality parameters. By quantitatively evaluating and optimizing the tags in the tagging system, the quality of the tagging system can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994777B_ABST
    Figure CN120994777B_ABST
Patent Text Reader

Abstract

The application provides an optimization method and device of a label system and electronic equipment, and relates to the technical field of large models. The method comprises the following steps: obtaining a pre-constructed label system and a conversation record, the conversation record being a history record of user interaction with an artificial intelligence device; labeling a label in the conversation record, the label being at least part of the labels in the label system; determining a quality parameter of the label system according to the conversation labeled with the label, the quality parameter comprising at least one of overlapping degree, ambiguity, confidence and coverage of the label; and optimizing the label system in the case of determining that the label system needs to be optimized according to the quality parameter. According to the application, the quality parameter of the label system is obtained based on the labeled history conversation record, so that the label system is optimized according to the quality parameter. Through quantitative evaluation and optimization of the labels in the label system, the quality of the label system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large model technology, and in particular to a method, apparatus and electronic device for optimizing a labeling system. Background Technology

[0002] Currently, the construction of tagging systems mainly relies on manual experience for design. Due to the lack of unified design standards, tag definitions are inconsistent across different business scenarios. This tagging system directly affects the accuracy of data classification, thereby impacting business processing efficiency. Clearly, the current tagging system suffers from poor quality. Summary of the Invention

[0003] This application provides a method, apparatus, and electronic device for optimizing a labeling system to solve the problem of poor labeling system quality.

[0004] To solve the above-mentioned technical problems, this application is implemented as follows:

[0005] In a first aspect, embodiments of this application provide a method for optimizing a labeling system, the method comprising:

[0006] Acquire a pre-built tag system and session records, wherein the session records are the historical records of user interactions with artificial intelligence devices;

[0007] Label the sessions in the session records with tags, where the tags are at least a portion of the tags in the tagging system;

[0008] The quality parameters of the tagging system are determined based on the sessions that are tagged with the tags. The quality parameters include at least one of the following: tag overlap, ambiguity, confidence, and coverage.

[0009] If the labeling system is determined to be in need of optimization based on the quality parameters, the labeling system is optimized.

[0010] Secondly, embodiments of this application provide a labeling system optimization apparatus, the apparatus comprising:

[0011] The acquisition module is used to acquire a pre-built tag system and session records, wherein the session records are the historical records of user interactions with artificial intelligence devices;

[0012] The annotation module is used to annotate the sessions in the session records with tags, wherein the tags are at least some of the tags in the tagging system;

[0013] The determination module is used to determine the quality parameters of the tagging system based on the sessions labeled with the tags, wherein the quality parameters include at least one of tag overlap, ambiguity, confidence, and coverage;

[0014] An optimization module is used to optimize the labeling system when the quality parameters indicate that the labeling system needs to be optimized.

[0015] Thirdly, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the tag system optimization method described in the first aspect.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the tag system optimization method described in the first aspect.

[0017] Fifthly, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the optimization method for the tagging system as described in the first aspect.

[0018] In this embodiment, historical session records are labeled using tags in a tagging system, and quality parameters of the tagging system are obtained based on the labeled historical session records. The tagging system is then optimized based on the quality parameters. By quantitatively evaluating and optimizing the tags in the tagging system, the quality of the tagging system can be improved. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is one of the flowcharts of a labeling system optimization method provided in the embodiments of this application;

[0021] Figure 2 This is a flowchart illustrating a supplementary labeling system provided in an embodiment of this application;

[0022] Figure 3 This is a second flowchart of a labeling system optimization method provided in an embodiment of this application;

[0023] Figure 4 This is a schematic diagram of the structure of an optimization device for a labeling system provided in an embodiment of this application;

[0024] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] This application provides a method, apparatus, and electronic device for optimizing a labeling system to solve the problem of poor labeling system quality.

[0027] See Figure 1 , Figure 1 This is a flowchart of a labeling system optimization method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:

[0028] Step 101: Obtain a pre-built tag system and session records, wherein the session records are the historical records of user interactions with artificial intelligence devices;

[0029] Step 102: Tag the sessions in the session records, wherein the tags are at least some of the tags in the tagging system;

[0030] Step 103: Determine the quality parameters of the tagging system based on the sessions labeled with the tags. The quality parameters include at least one of the following: tag overlap, ambiguity, confidence, and coverage.

[0031] Step 104: If the labeling system is determined to be in need of optimization based on the quality parameters, the labeling system is optimized.

[0032] The pre-built tagging system is a pre-established tagging system that can be used to classify and label data (including user behavior, session records, product information, etc.) for analysis or management. The tagging system includes multiple tags, which can be manually defined or generated through other methods.

[0033] The aforementioned tagging system can be a collection of tags, specifically including user tags used to describe user characteristics, tags for classifying user questions or customer service conversations, tags used to classify products or services, and general tags for other scenarios, which are not limited here.

[0034] To evaluate the quality of the tags in this tagging system, historical conversation records are obtained. These records can include conversations between users and AI devices (such as intelligent customer service or voice assistants), as well as other conversation records. Conversation records can include text, speech-to-text, or structured data (such as user questions or system responses).

[0035] Label the sessions in the session log.

[0036] In some implementations, a conversation summary is extracted, along with semantic and business keywords. These keywords are then compared with tags in a tagging system. This process can be handled using a large model. For example, the conversation summary "user inquires about loan application process" is extracted, and keywords in the dialogue include "loan" and "application." The extracted keywords are then matched with a tagging system, and the tag "loan consultation" is assigned.

[0037] In some implementations, each session is labeled according to the session semantics.

[0038] The quality parameters of the tagging system can be determined by adding semantics to the sessions or the tags themselves.

[0039] In an established tagging system, the following situations may exist:

[0040] Semantic overlap in tags. For example, "car" and "sedan" may have a high degree of semantic overlap. If they are not distinguished, the tag system will become bloated and inefficient.

[0041] The same label can have multiple meanings in different contexts. For example, "apple" can refer to both fruit and technology companies. This ambiguity can directly affect the accuracy of subsequent data classification or retrieval.

[0042] The label definitions in the labeling system are vague and the boundaries are unclear. For example, if the scope of the label definition is too broad or too narrow, it may lead to overgeneralization during classification or the omission of important features.

[0043] Labels cannot fully cover the key concepts of the target domain, and some data may not be accurately classified due to the lack of appropriate labels, affecting subsequent analysis and application.

[0044] Based on the above issues, the quality of the tag system is quantitatively evaluated by obtaining quality parameters of the tag system, according to the semantics of the session records with added tags and the semantics of the tags.

[0045] Overlap is used to characterize the degree of semantic overlap of tags in the tag system. In some implementations, the overlap can be determined by obtaining the frequency of multiple tags appearing together in the same session and the semantics of the tag descriptions. Alternatively, it can be determined by constructing a tag association graph and utilizing the tightness of the graph.

[0046] Ambiguity can be measured by the semantic similarity between multiple sessions labeled with the same tag. The lower the semantic similarity, the higher the ambiguity of the tag. Alternatively, it can be determined based on ambiguity reported by users.

[0047] Confidence can be determined based on the dialogue generated from the tag description and the relevance of that dialogue to the tag description. Additionally, the confidence level of the classification can be assessed by calculating the probability of a tag being correctly classified and the probability of it being complained about based on historical data.

[0048] Coverage can be determined by calculating the number of uncovered sessions in the session logs. To improve coverage, the tagging system can be improved by acquiring a richer set of session logs.

[0049] If the label system is found to have any of the above-mentioned situations, i.e., it needs to be optimized, based on the quality parameters of the label, the label system is optimized. After optimization, the process of obtaining the above-mentioned quality parameters, evaluating the quality and optimizing can be repeated until the quality parameters of the optimized label system meet the preset conditions.

[0050] In addition, when a new label is detected in the labeling system, the steps of obtaining quality parameters, evaluating quality, and optimizing the labeling system described above can be performed to improve the quality of the labeling system.

[0051] In this embodiment of the application, historical session records are labeled using tags in a tagging system, and the quality parameters of the tagging system are obtained by adding labels to the historical session records. Optimization is then performed based on the quality parameters. By using the above method to quantitatively evaluate and optimize the quality of the tagging system, the quality of the tagging system can be improved.

[0052] Optionally, the quality parameter includes overlap; determining the quality parameter of the tagging system based on the sessions that label the tags includes:

[0053] Obtain N tags that annotate the same session in the session record and the total number of sessions in the session record, where N is an integer greater than 1;

[0054] Determine the co-occurrence frequency of each pair of tags among the N tags in the session record, where the co-occurrence frequency is the proportion of the number of times each pair of tags is labeled for the same session relative to the total number of sessions;

[0055] The overlap of the N tags is determined based on the co-occurrence frequency and the semantic similarity of each pair of tags.

[0056] The session records include the original content of the session records and the corresponding tagging results.

[0057] If N tags are used to label the same session, the co-occurrence frequency of each pair of tags in the session records is obtained. That is, the proportion of times the same pair of tags appears simultaneously in multiple sessions out of the total number of sessions. The higher the co-occurrence frequency of two tags and the higher the semantic similarity of their descriptions, the higher the tag overlap.

[0058] In some implementations, when N is greater than 2, the co-occurrence frequency of each pair (i.e., two) of the N tags can be obtained. For example, when N = 3, the co-occurrence frequency of the three pairs of tags in a session can be obtained based on the three tags.

[0059] For example, if "loan consultation" and "repayment issues" are both tagged in 20 sessions, the co-occurrence count is 20. If the total number of sessions is 100, the co-occurrence frequency is 20 / 100 = 20%.

[0060] Co-occurrence frequency can be calculated using the following formula:

[0061]

[0062] Among them, f i,j This represents the co-occurrence frequency of tags i and j, where N represents the total number of sessions. i,j This indicates the number of sessions that are simultaneously labeled as both label i and label j.

[0063] The overlap of tags can be calculated based on their co-occurrence frequency. i,j :

[0064]

[0065] Among them, overlap i,j The overlap of the labels is represented by ; a and b are weight coefficients, and sim(·) represents the semantic similarity between the two label descriptions. The "label name: label description" is encoded into a vector through word embedding models such as BAAI GeneralEmbedding (BGE), and then the semantic similarity is calculated using cosine similarity.

[0066] The overlap is compared with a preset overlap threshold to determine whether the overlap exceeds the preset threshold (which can be preset and adjusted according to business scenarios). When it exceeds the preset threshold, the overlapping tags are optimized.

[0067] For example, suppose the tags "loan problem" and "early repayment" appear together 100 times in historical sessions, with a total of 1000 sessions. The co-occurrence frequency is 0.1. Using the BGE model, the semantic similarity between "loan problem" and "early repayment" is calculated to be 0.8. Assuming a = 0.5 and b = 0.5, calculate the tag overlap:

[0068] overlap i,j =0.5·e0.1+0.5·0.8≈0.5·1.105+0.4=0.5525+0.4=0.9525

[0069] According to the calculation results, the labels "loan issues" and "early repayment" have a high degree of overlap and semantic overlap.

[0070] In some implementations, when N is greater than 2, the co-occurrence frequency of N tags in the session can also be obtained simultaneously.

[0071] The overlap is quantitatively analyzed using the above methods. A composite overlap calculation model that integrates co-occurrence frequency and semantic similarity is used, and adjustable parameters are used to adapt to different business scenarios. The accuracy of identifying implicit semantic overlap and overlapping business scenarios (such as "request for early repayment" and "request for early settlement") is significantly improved.

[0072] Optionally, the quality parameter includes ambiguity, which indicates that the label includes at least two semantics; determining the quality parameter of the labeling system based on the session that labels the label includes:

[0073] In the session records, a first session record and a second session record labeled with a first tag are obtained, wherein the first tag is at least a portion of the tags labeled in the session records;

[0074] Determine the expected value of the semantic similarity between the first session record and the second session record;

[0075] The ambiguity of the first tag is determined based on the expected value of the semantic similarity.

[0076] The ambiguity of a label represents the semantic similarity between multiple sessions labeled with the same label. For each session in the session record, a label is assigned. When there are at least two sessions labeled with the same label, the semantic similarity between each pair of sessions can be obtained, and the expected value of the semantic similarity can be obtained.

[0077] The ambiguity of label i can be calculated using the following formula:

[0078]

[0079] A iLet D represent the ambiguity of label i, where i,j∈T represents the label set. i This represents the set of sessions labeled with tag i. Represents the session set D i The expected semantic similarity between pairs of objects.

[0080] Specifically, for semantic similarity of conversations, the semantic similarity between pairs of conversations can be obtained by first extracting conversation summaries using a large model, then encoding them into vectors using word embedding models such as BGE, and finally calculating cosine similarity.

[0081] After obtaining the ambiguity value, the ambiguity can be compared with a preset threshold to assess whether the label is ambiguous.

[0082] For example, assuming the expected semantic similarity between "loan problem" and "early repayment" is 0.6, the ambiguity value calculated based on this expected value is 0.53 (greater than the preset value of 0.5), indicating that there is some semantic ambiguity.

[0083] By using the above quantitative calculation method to calculate the ambiguity of labels, the accuracy of label ambiguity recognition can be improved.

[0084] Optionally, the quality parameter includes ambiguity; determining the quality parameter of the tagging system based on the session that labels the tag includes:

[0085] A first session is generated based on the semantics of a second tag using a large model, wherein the second tag includes at least a portion of the tags labeled for the session record;

[0086] The confidence level of the second label is determined using a large model, and the confidence level is used to characterize the relevance between the second label and the first session;

[0087] The ambiguity of the second label is determined based on the correlation.

[0088] For the second label of the session record, a large model is used to generate the first session based on the semantics of the label description. Then, the large model is asked to answer whether the generated session is related to the label description, i.e., to determine the confidence of the second label. The large model answers either "yes" or "no". The confidence can be measured by the logit of the "yes" token, i.e., calculating the probability p(yes|d,i) that the session is related to the label. This can be expressed by the formula:

[0089]

[0090] Where s = LLM(d,i) represents the probability that the large model predicts all tokens are associated with the session.

[0091] The ambiguity Fi of the labels is calculated using the formula for all generated sessions:

[0092]

[0093] Among them, |D LM | represents the set of sessions generated by the large model.

[0094] A ambiguity threshold can be set, and the calculated ambiguity value can be compared with the threshold to determine whether the label is ambiguous. If most of the generated dialogues are judged as "irrelevant," it indicates that the label description is too broad or vague. If most of the generated dialogues are judged as "relevant," it indicates that the label definition is clear.

[0095] For example, the large model generates dialogues based on the "loan process" tag description, such as: "How do I apply for a loan?"

[0096] The probability (i.e., p(yes|d,i)) of the large model judging the generated dialogue as "relevant" is 0.7. Assuming that 60% of the generated dialogues are judged as "relevant", the ambiguity of the label is calculated to be 0.4 according to the above formula. This value is compared with the preset threshold and is greater than the preset value, indicating that it is not clear enough and needs to be optimized.

[0097] Calculating the ambiguity of a label using the above method can improve the accuracy of identifying ambiguous labels.

[0098] Optionally, the quality parameter includes coverage; determining the quality parameters of the tagging system based on the sessions labeled with the tag includes:

[0099] The maximum confidence score between the second session and the third label in the session record is determined using a large model. The third label includes at least some of the labels labeled for the session record. The maximum confidence score is used to indicate that the second session has the greatest correlation with the third label in the label system.

[0100] The coverage of the tag system is determined based on the difference between the maximum confidence level and the mean maximum confidence level, where the mean maximum confidence level is the average of the maximum tag confidence levels corresponding to all sessions in the session record.

[0101] To determine whether a session d in the session record has been overwritten, the maximum confidence score of the session tag can be compared with the mean of the maximum confidence scores.

[0102] In some implementations, each session d (e.g., "How to repay early?") in the session record and the tags in the tagging system (e.g., "Loan consultation", "Repayment issues") can be input into a large model (e.g., LLM). The large model outputs the relevance confidence score (e.g., "yes" or "no") between each session and the tag. For each session, the maximum confidence score between it and all tags in the tagging system is taken. This maximum confidence score indicates that the session has the highest probability of relevance to that tag, and the confidence score is used to represent the degree of matching between the session and the tag.

[0103] For example, the confidence level of the conversation "How to repay early?" is 0.8 with the label "Loan Consultation" and 0.7 with "Repayment Issues", so the maximum confidence level is 0.8.

[0104] In addition, the mean of the maximum confidence scores for all sessions and tags was calculated.

[0105] If the difference between the maximum confidence level and the mean confidence level of session d is greater than the preset value, it indicates that the confidence level of the session is significantly higher than the average level and is determined to be covered; otherwise, it is determined to be uncovered.

[0106] The probability p that session d is covered d Expressed as a formula:

[0107]

[0108] in σ represents the mean of the maximum confidence scores for each session tag, and σ represents the standard deviation. |D| represents the set of sessions in the session record.

[0109] The coverage C(D,T) of the tagging system for the entire conversation in the conversation records is calculated as follows:

[0110]

[0111] H(·) is the step function, p d When -τ>0, H(·)=1; otherwise, H(·)=0. A larger coverage value indicates stronger coverage capability.

[0112] For example, suppose the total number of dialogues |D| = 100, of which 80 dialogues have a p d If the value is greater than 0, then the coverage C(D,T) = 80 / 100 = 0.8, indicating that 20% of sessions are not covered. Coverage can be assessed by setting a threshold to ultimately achieve the target coverage, i.e., a preset threshold (e.g., 90% or 95%).

[0113] By designing a dynamic threshold judgment model, the model automatically identifies uncovered session samples through statistical significance testing. It can adaptively adjust the coverage standard according to the data distribution, thereby reducing the false judgment rate by more than 40% and improving the judgment accuracy.

[0114] Optionally, optimizing the labeling system when it is determined that the labeling system needs optimization based on the quality parameters includes at least one of the following:

[0115] In the case where the labeling system includes overlapping labels, the overlapping labels are integrated;

[0116] In cases where the labeling system includes ambiguous labels, the ambiguous labels are corrected.

[0117] In cases where the labeling system includes ambiguous labels, the names of the labels shall be modified;

[0118] If the tag system does not cover the session records, a new tag is determined based on the uncovered third session, and the new tag is added to the tag system.

[0119] Through the above quantitative calculations, we can obtain ambiguous, vaguely defined, and overlapping candidate tag sets, as well as uncovered dialogue sets. Targeted optimizations can be performed on these candidate tag sets.

[0120] First, for overlapping tags, determine the correct tag by integrating or renaming them. For example, integrate "car" and "sedan" into a more reasonable tag.

[0121] During optimization, a large model can be used for fine-tuning to optimize the labeling system. Optimization instructions are as follows:

[0122] #Tag overlap optimization

[0123] ## Role Definition

[0124] You are a tag system optimization expert, specializing in the analysis and optimization of tag semantic overlap.

[0125] ##Task Description

[0126] The input candidate set of labels contains several label pairs. You need to determine whether there are overlapping or inclusive relationships between the label concepts. If so, you should integrate them.

[0127] ## Output Requirements

[0128] 1. **Output Content**: Output tags after optimizing semantic overlap.

[0129] 2. **Output Format**: Output tags directly, separated by semicolons (;).

[0130] 3. **Output Example**: Tag 1; Tag 2

[0131] ## Tag Candidate Set

[0132] {(i,j)|overlap| i,j >τ o ;i,j∈T)}

[0133] Return:

[0134] By integrating overlapping tags in the above manner, the overlap of tags is reduced, and the efficiency of tag matching is improved.

[0135] Second, for ambiguous tags, consolidation or renaming can be used to correct them. A large model is used to optimize the tags, with the following instructions:

[0136] # Tag Ambiguity Optimization

[0137] ## Role Definition

[0138] You are a tag system optimization expert, specializing in the analysis and optimization of tag ambiguity.

[0139] ##Task Description

[0140] Check if the labels in the input candidate set have multiple meanings that could lead to ambiguity. By integrating and renaming the labels, ensure that each label has a unique and clear meaning.

[0141] ## Output Requirements

[0142] 1. **Output Content**: Output the labels after optimizing for ambiguity issues.

[0143] 2. **Output Format**: Output tags directly, separated by semicolons (;).

[0144] 3. **Output Example**: Tag 1; Tag 2

[0145] ## Tag Candidate Set

[0146] {T(A i >τ a ;i∈T)}

[0147] Return:

[0148] Third, for ambiguous tags, the semantic boundaries of the tags can be improved by optimizing the tag naming.

[0149] The instructions for optimizing fuzzy labels using a large model are as follows:

[0150] #Fuzzy Optimization

[0151] ## Role Definition

[0152] You are a tag system optimization expert, specializing in the analysis and optimization of tag ambiguity.

[0153] ##Task Description

[0154] To address issues such as unclear or ambiguous label definitions and boundaries, we will optimize label naming to ensure that labels are semantically accurate and actionable.

[0155] ## Output Requirements

[0156] 1. **Output Content**: Output the labels after optimizing the fuzziness issue.

[0157] 2. **Output Format**: Output tags directly, separated by semicolons (;).

[0158] 3. **Output Example**: Tag 1; Tag 2

[0159] ## Tag Candidate Set

[0160] {T(F i >τ F ;i∈T)}

[0161] Return:

[0162] IV. For uncovered session records, extract keywords from these sessions to obtain tags (i.e., tags to be added). Construct a hierarchy from these tags, compare them with tags in the existing tag system, remove duplicates, and then add them to the existing tag system. See the above process for details. Figure 2 As shown.

[0163] In addition, during the optimization process, optimization can be achieved by, for example, splitting ambiguous tags into more specific sub-tags, adjusting tag names, or supplementing tag descriptions.

[0164] To facilitate understanding of this embodiment, specific scenarios are used as examples below.

[0165] In intelligent conversation scenarios between users and intelligent assistants, historical conversation records between the intelligent assistant and all users are acquired and labeled using a tagging system. A large-scale model analyzes the conversation records (e.g., "How to repay early?"), assessing the match between the dialogue and the tagging system, calculating the maximum confidence score for each conversation, and comparing it with the overall mean to determine coverage. If some conversations are found to have classification biases due to insufficient tag coverage or high overlap, the system automatically adjusts the tag definitions or adds new tags. For example, it might merge "repayment issues" and "early repayment" into "loan repayment," or add a new tag like "loan limit adjustment," thereby improving classification accuracy, enhancing the precision of user question categorization, and ultimately improving response efficiency.

[0166] The entire process of optimizing the labeling system in this application can be found in [reference needed]. Figure 3 As shown, the annotation module is used to annotate historical dialogues based on the initialized tagging system. The evaluation module calculates various indicators based on the annotated historical dialogue data, assessing the tagging system's overlap, ambiguity, vagueness, and coverage, identifying tags that need optimization. The optimization module then specifically optimizes the tags and supplements those with insufficient coverage. The optimized tagging system can be annotated with historical dialogues again, evaluated, and optimized repeatedly until the tagging system's indicators meet the requirements.

[0167] This application's embodiments, by establishing quantitative indicators, unify the optimization processes of the three dimensions of ambiguity, semantic overlap, and coverage within a single framework, achieving systematic and complete optimization of the tagging system. Simultaneously, quantitative calculations improve the reliability of identifying tag ambiguity, vague definitions, semantic overlap, and coverage issues. This significantly enhances the semantic clarity, structural rationality, and application effectiveness of the tagging system, providing more reliable semantic support for tasks such as data annotation, classification, and retrieval.

[0168] By analyzing the confidence level of dialogue tagging through a coverage evaluation algorithm, it automatically identifies areas with insufficient coverage and recommends supplementary tags. This approach is suitable for rapidly developing emerging fields to quickly establish tagging systems, effectively solving the problem of incomplete coverage caused by limitations in human experience in traditional methods.

[0169] The optimization process can be automated, significantly improving the efficiency of tag system construction and updates. Compared with manual methods, it can save time and costs and is applicable to large-scale, dynamically changing tag system management scenarios.

[0170] See Figure 4 , Figure 4 This is a schematic diagram of the structure of an optimization device for a labeling system provided in an embodiment of this application, as shown below. Figure 4 As shown, the labeling system optimization device 400 includes:

[0171] The acquisition module 401 is used to acquire a pre-built tag system and session records, wherein the session records are the historical records of user interactions with artificial intelligence devices;

[0172] The annotation module 402 is used to annotate the sessions in the session records with tags, wherein the tags are at least some of the tags in the tagging system;

[0173] The determination module 403 is used to determine the quality parameters of the tagging system based on the sessions labeled with the tags, wherein the quality parameters include at least one of the following: tag overlap, ambiguity, confidence, and coverage.

[0174] The optimization module 404 is used to optimize the label system when the label system is determined to be to be optimized based on the quality parameters.

[0175] Optionally, the quality parameter includes overlap; the determining module includes:

[0176] The first acquisition submodule is used to acquire N tags that mark the same session in the session record and the total number of sessions in the session record, where N is an integer greater than 1;

[0177] The first determining submodule is used to determine the co-occurrence frequency of each pair of tags among the N tags in the session record, wherein the co-occurrence frequency is the proportion of the number of times each pair of tags is labeled for the same session to the total number of sessions;

[0178] The second determining submodule is used to determine the overlap of the N tags based on the co-occurrence frequency and the semantic similarity of each pair of tags.

[0179] Optionally, the quality parameter includes ambiguity, and the ambiguity characterization label includes at least two semantics; the determining module includes:

[0180] The second acquisition submodule is used to acquire, from the session records, a first session record and a second session record marked with a first tag, wherein the first tag is at least a portion of the tags marked on the session records;

[0181] The third determining submodule is used by the second determining submodule to determine the expected value of the semantic similarity between the first session record and the second session record;

[0182] The fourth determining submodule is used to determine the ambiguity of the first tag based on the expected value of the semantic similarity.

[0183] Optionally, the quality parameter includes ambiguity; the determining module includes:

[0184] A generation submodule is used to generate a first session based on the semantics of a second label using a large model, wherein the second label includes at least a portion of the labels labeled for the session record;

[0185] The fifth determination submodule is used to determine the confidence level of the second label using a large model, wherein the confidence level is used to characterize the correlation between the second label and the first session;

[0186] The sixth determining submodule is used to determine the ambiguity of the second label based on the correlation.

[0187] Optionally, the quality parameter includes coverage; the determining module includes:

[0188] The seventh determination submodule is used to determine the maximum confidence of the second session and the third label in the session record using a large model. The third label includes at least some of the labels labeled for the session record. The maximum confidence is used to indicate that the second session has the greatest correlation with the third label in the label system.

[0189] The eighth determination submodule is used to determine the coverage of the tag system based on the difference between the maximum confidence level and the mean maximum confidence level, wherein the mean maximum confidence level is the average of the maximum tag confidence levels corresponding to all sessions in the session record.

[0190] Optionally, the optimization module is configured to perform at least one of the following:

[0191] In the case where the labeling system includes overlapping labels, the overlapping labels are integrated;

[0192] In cases where the labeling system includes ambiguous labels, the ambiguous labels are corrected.

[0193] In cases where the labeling system includes ambiguous labels, the names of the labels shall be modified;

[0194] If the tag system does not cover the session records, a new tag is determined based on the uncovered third session, and the new tag is added to the tag system.

[0195] The tagging system optimization device can achieve Figure 1 The various processes implemented in the method embodiments can achieve the same technical effect, and will not be described again here to avoid repetition.

[0196] It should be noted that the electronic device provided in this application embodiment is an optimization device capable of executing the above-described tagging system. Therefore, all implementation methods in the above-described tagging system optimization method embodiments are applicable to this electronic device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.

[0197] like Figure 5 As shown, this application embodiment also provides an electronic device 500, including: a processor 501, a memory 502, and a program stored in the memory 502 and executable on the processor 501. When the program is executed by the processor 501, it implements the various processes of the above-described tag system optimization method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0198] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described optimized tagging system method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0199] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0200] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0201] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0202] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for optimizing a labeling system, characterized in that, include: Acquire a pre-built tag system and session records, wherein the session records are the historical records of user interactions with artificial intelligence devices; Label the sessions in the session records with tags, where the tags are at least a portion of the tags in the tagging system; The quality parameters of the tagging system are determined based on the sessions labeled with the tags. The quality parameters include at least one of the following: tag overlap, ambiguity, vagueness, and coverage. The overlap is used to characterize the degree of semantic overlap of tags in the tagging system. The ambiguity is used to characterize that the tags include at least two semantics. The vagueness is used to characterize the ambiguity of the tag definition. The coverage is used to represent the ability to cover the sessions. If the labeling system is determined to be in need of optimization based on the quality parameters, the labeling system is optimized. When the labeling system is determined to be in need of optimization based on the quality parameters, the optimization of the labeling system includes at least one of the following: In the case where the labeling system includes overlapping labels, the overlapping labels are integrated; In cases where the labeling system includes ambiguous labels, the ambiguous labels are corrected. In cases where the labeling system includes ambiguous labels, the names of the labels shall be modified; If the tag system does not cover the session records, the tag to be added is determined based on the uncovered third session, and the tag to be added is added to the tag system.

2. The method according to claim 1, characterized in that, The quality parameters include overlap; determining the quality parameters of the tagging system based on the sessions that label the tags includes: Obtain N tags that annotate the same session in the session record and the total number of sessions in the session record, where N is an integer greater than 1; Determine the co-occurrence frequency of each pair of tags among the N tags in the session record, where the co-occurrence frequency is the proportion of the number of times each pair of tags is labeled for the same session relative to the total number of sessions; The overlap of the N tags is determined based on the co-occurrence frequency and the semantic similarity of each pair of tags.

3. The method according to claim 1, characterized in that, The quality parameters include ambiguity, which indicates that the tag includes at least two semantics; determining the quality parameters of the tagging system based on the sessions that annotate the tags includes: In the session records, a first session record and a second session record labeled with a first tag are obtained, wherein the first tag is at least a portion of the tags labeled in the session records; Determine the expected value of the semantic similarity between the first session record and the second session record; The ambiguity of the first tag is determined based on the expected value of the semantic similarity.

4. The method according to claim 1, characterized in that, The quality parameters include ambiguity; determining the quality parameters of the tagging system based on the sessions that label the tags includes: A first session is generated based on the semantics of a second tag using a large model, wherein the second tag includes at least a portion of the tags labeled for the session record; The confidence level of the second label is determined using a large model, and the confidence level is used to characterize the relevance between the second label and the first session; The ambiguity of the second label is determined based on the correlation.

5. The method according to claim 1, characterized in that, The quality parameters include coverage; determining the quality parameters of the tagging system based on the sessions labeled with the tags includes: The maximum confidence score between the second session and the third label in the session record is determined using a large model. The third label includes at least some of the labels labeled for the session record. The maximum confidence score is used to indicate that the second session has the greatest correlation with the third label in the label system. The coverage of the tag system is determined based on the difference between the maximum confidence level and the mean maximum confidence level, where the mean maximum confidence level is the average of the maximum tag confidence levels corresponding to all sessions in the session record.

6. An optimization device for a labeling system, characterized in that, include: The acquisition module is used to acquire a pre-built tag system and session records, wherein the session records are the historical records of user interactions with artificial intelligence devices; The annotation module is used to annotate the sessions in the session records with tags, wherein the tags are at least some of the tags in the tagging system; The determination module is used to determine the quality parameters of the tag system based on the sessions labeled with the tags. The quality parameters include at least one of the following: tag overlap, ambiguity, vagueness, and coverage. The overlap is used to characterize the degree of semantic overlap of tags in the tag system. The ambiguity characterizes that the tags include at least two semantics. The vagueness is used to characterize the ambiguity of the tag definition. The coverage is used to represent the ability to cover the sessions. An optimization module is used to optimize the labeling system when the quality parameters determine that the labeling system needs to be optimized. The optimization module is used to perform at least one of the following: In the case where the labeling system includes overlapping labels, the overlapping labels are integrated; In cases where the labeling system includes ambiguous labels, the ambiguous labels are corrected. In cases where the labeling system includes ambiguous labels, the names of the labels shall be modified; If the tag system does not cover the session records, the tag to be added is determined based on the uncovered third session, and the tag to be added is added to the tag system.

7. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method for optimizing the tagging system as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method for optimizing the labeling system as described in any one of claims 1 to 5.

9. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the method for optimizing the tagging system as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Enterprise portrait label intelligent generation method

    CN119168504A

  • Systems and methods for label versioning for machine learning input data

    US20240202572A1