Label setting method and device, equipment, storage medium and product
By identifying unlabeled descriptive dimensions and analyzing based on large models, multi-dimensional descriptive labels are set for the corpus, solving the problem of single information dimensions in corpus annotation, achieving more efficient multi-dimensional annotation and label coverage, and improving the accuracy and efficiency of annotation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING QIHOOD TECHNOLOGY CO LTD
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-12
AI Technical Summary
Current technologies only perform corpus annotation from a single dimension, resulting in limited information dimensions and insufficient label coverage, which affects the effectiveness of the model.
By identifying unlabeled descriptive dimensions, selecting descriptive dimensions to be labeled, and setting multi-dimensional descriptive labels for the corpus based on large-scale model analysis, including consistency and similarity matching, and using the label setting expert confirmation results, the accuracy and coverage of the labels are ensured.
It enables multi-dimensional annotation of corpora, improves information dimensions and label coverage, enhances annotation accuracy and efficiency, and reduces manual annotation errors.
Smart Images

Figure CN122020198A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to label setting methods, apparatus, devices, storage media and products. Background Technology
[0002] In fields such as natural language processing, text analysis, and artificial intelligence, high-quality annotation of corpora is a key prerequisite for training data preparation. However, traditional corpus platforms use a single dimension for annotation, such as only classifying labels or only quality scores, which has problems such as limited information dimensions, insufficient label coverage, and restrictions on subsequent model use. Summary of the Invention
[0003] The main purpose of this application is to provide a label setting method, apparatus, device, storage medium and product, which aims to solve the technical problem that when related technologies annotate corpora, they only annotate from a single dimension, resulting in poor actual use effect.
[0004] To achieve the above objectives, this application proposes a label setting method, the method comprising: Determine the unlabeled descriptive dimensions of the corpus to be labeled, where the descriptive dimension is an independent direction that limits or describes the corpus; Select the dimension to be labeled from the unlabeled description dimensions; In response to the annotation action corresponding to the dimension to be annotated, set the corresponding dimension description label for the corpus to be labeled.
[0005] Optionally, the step of responding to the annotation action corresponding to the dimension to be labeled, and setting corresponding dimension description labels for the corpus to be labeled, includes: In response to the annotation action corresponding to the dimension to be annotated, determine the corpus description tag corresponding to the annotation action; The corpus to be labeled is analyzed using a large model to determine the optional descriptive tags corresponding to the descriptive dimensions to be labeled. Match the corpus description tags with the optional description tags; When a match is successful, the dimension description label corresponding to the corpus to be labeled is set as the corpus description label.
[0006] Optionally, matching the corpus description tags with the optional description tags includes: Perform consistency matching between the corpus description tags and each optional description tag; If there is no optional description tag that matches the corpus description tag, then the corpus description tag is matched with each optional description tag to determine the tag similarity of each optional description tag; If there is an optional descriptive tag with a corresponding tag similarity greater than or equal to the preset similarity threshold, then the match is considered successful.
[0007] Optionally, if no alternative description tag matches the corpus description tag, then after performing similarity matching between the corpus description tag and each alternative description tag to determine the tag similarity corresponding to each alternative description tag, the method further includes: If there is no corresponding optional descriptive label with a similarity greater than or equal to a preset similarity threshold, then a label confirmation task is constructed based on the corpus descriptive labels and the corpus to be labeled. Obtain at least one tag confirmation result from the tag setting expert for the tag confirmation task; Determine the label correctness confidence rate based on the confirmation results of at least one label; If the correct confidence rate of the label is greater than the preset confidence threshold, then the match is considered successful.
[0008] Optionally, before analyzing the corpus to be labeled using a large model to determine the optional description tags corresponding to the description dimension to be labeled, the method further includes: Obtain the set description tags of the corpus to be labeled; Detect whether the corpus description tags conflict with the already set description tags; If there is no conflict, then proceed with the step of analyzing the corpus to be labeled using a large model to determine the optional description tags corresponding to the description dimension to be labeled.
[0009] Optionally, detecting whether the corpus description tags conflict with the already set description tags includes: Based on the set description tags and the search conditions of the corpus description tag construction rules; Check if a conflicting rule corresponding to the search condition of the specified rule exists in the conflict rule base; If there is a conflicting rule corresponding to the rule search condition, then a conflict is determined to exist; If there is no conflicting rule corresponding to the rule search condition, then it is determined that there is no conflict.
[0010] Optionally, selecting the dimension to be labeled from the unlabeled description dimensions includes: Statistical analysis was performed on the dimension description tags corresponding to each corpus in the corpus to determine the total amount of corpus and the total number of tags set for each description dimension. The annotation coverage rate corresponding to each description dimension is determined based on the total amount of the corpus and the total amount of the tag settings. The unlabeled descriptive dimensions are sorted based on the labeled coverage, and the descriptive dimensions to be labeled are determined based on the sorting results.
[0011] Optionally, selecting the dimension to be labeled from the unlabeled description dimensions includes: Obtain the task backlog ratio corresponding to each description dimension; Based on the task stacking ratio, a description dimension to be labeled is selected from the unlabeled description dimensions.
[0012] Optionally, obtaining the task backlog ratio corresponding to each description dimension includes: The task list for setting tags is statistically analyzed to obtain the total backlog of tasks and the number of tag setting tasks corresponding to each description dimension. The task accumulation ratio corresponding to each descriptive dimension is determined based on the task quantity set by the label and the total task accumulation.
[0013] Optionally, obtaining the task backlog ratio corresponding to each description dimension includes: Retrieve the task allocation thread pool corresponding to each description dimension; Get the number of waiting sequence elements corresponding to the task allocation thread pool; The task backlog ratio corresponding to each descriptive dimension is determined based on the number of waiting sequence elements.
[0014] Optionally, before determining the unlabeled descriptive dimensions of the corpus to be labeled, the method further includes: Statistical analysis was performed on the dimension description tags corresponding to each corpus in the corpus to determine the total number of tags set for each corpus. The target corpus is the corpus in which the total number of labels set is less than the total number of preset dimensions. Select the text to be annotated from the target text; Accordingly, after setting corresponding dimension description labels for the corpus to be labeled in response to the annotation action corresponding to the dimension to be labeled, the method further includes: Detect whether there are target corpora that were not selected as corpora to be labeled; If it exists, return to the step of selecting the corpus to be annotated from the target corpus.
[0015] Furthermore, to achieve the above objectives, this application also proposes a label setting device, which includes: The determination module is used to determine the unlabeled descriptive dimensions of the corpus to be labeled, where the descriptive dimension is an independent direction that limits or describes the corpus; The selection module is used to select a dimension to be labeled from the unlabeled description dimensions. The setting module is used to respond to the annotation action corresponding to the dimension to be labeled and set the corresponding dimension description label for the corpus to be labeled.
[0016] Optionally, the setting module is further configured to respond to the annotation action corresponding to the dimension to be annotated, determine the corpus description label corresponding to the annotation action; analyze the corpus to be annotated through a large model to determine the optional description label corresponding to the dimension to be annotated; match the corpus description label with the optional description label; and when the match is successful, set the dimension description label corresponding to the corpus to be annotated as the corpus description label.
[0017] Optionally, the setting module is further configured to perform consistency matching between the corpus description tag and each optional description tag; if there is no optional description tag that matches the corpus description tag, then perform similarity matching between the corpus description tag and each optional description tag to determine the tag similarity corresponding to each optional description tag; if there is an optional description tag with a corresponding tag similarity greater than or equal to a preset similarity threshold, then the matching is determined to be successful.
[0018] Optionally, the setting module is further configured to: if there is no corresponding optional descriptive tag with a tag similarity greater than or equal to a preset similarity threshold, construct a tag confirmation task based on the corpus descriptive tags and the corpus to be labeled; obtain at least one tag confirmation result from a tag setting expert for the tag confirmation task; determine the tag correctness confidence rate based on the at least one tag confirmation result; and determine a successful match if the tag correctness confidence rate is greater than a preset confidence threshold.
[0019] Optionally, the setting module is further configured to obtain the set description tags of the corpus to be labeled; detect whether the corpus description tags conflict with the set description tags; if there is no conflict, then execute the step of analyzing the corpus to be labeled through a large model to determine the optional description tags corresponding to the description dimension to be labeled.
[0020] Optionally, the setting module is further configured to construct rule search conditions based on the set description tags and the corpus description tags; detect whether there are conflicting rules corresponding to the rule search conditions in the conflict rule base; if there are conflicting rules corresponding to the rule search conditions, it is determined that there is a conflict; if there are no conflicting rules corresponding to the rule search conditions, it is determined that there is no conflict.
[0021] In addition, to achieve the above objectives, this application also proposes a label setting device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the label setting method as described above.
[0022] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the label setting method as described above.
[0023] In addition, to achieve the above objectives, this application also proposes a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the label setting method as described above.
[0024] One or more technical solutions proposed in this application have at least the following technical effects: Because it can automatically determine the unlabeled descriptive dimensions of the corpus to be labeled and select the descriptive dimensions to be labeled from them, and then automatically set the corresponding dimension description labels for the corpus to be labeled based on the annotation actions of the annotators for the descriptive dimensions to be labeled, it realizes multi-dimensional labeling of the corpus, thereby improving the information dimensions and label coverage of the corpus. Attached Figure Description
[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart illustrating the label setting method of this application in Embodiment 1; Figure 2 This is a flowchart illustrating Embodiment 2 of the label setting method for this application. Figure 3 This is a flowchart illustrating Embodiment 3 of the label setting method for this application; Figure 4 This is a schematic diagram of the module structure of the label setting device according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the tag setting method in the embodiments of this application.
[0028] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0029] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0030] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0031] Based on this, the embodiments of this application provide a label setting method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the label setting method for this application.
[0032] In this embodiment, the label setting method includes steps S10 to S30: Step S10: Determine the unlabeled descriptive dimensions of the corpus to be labeled.
[0033] It should be noted that the execution subject of this embodiment may be the label setting device or a platform composed of at least one label setting device. The label setting device may be a personal computer, a server or other electronic device, or other devices that can achieve the same or similar functions. This embodiment does not limit this. In this embodiment and the following embodiments, the label setting device is used as an example to describe the label setting method of this application.
[0034] It should be noted that the corpus to be labeled can be corpus that requires label settings, which can be specified by the administrator of the label setting device. Unlabeled description dimensions can be description dimensions that do not have corresponding dimension description labels set. Here, the description dimension is an independent direction that limits or describes the corpus.
[0035] For example, if a corpus needs to be labeled from two descriptive dimensions: classification and grading, and if a certain corpus to be labeled has not been labeled, then the unlabeled descriptive dimensions of the corpus to be labeled include both classification and grading; however, if a dimension description label for the classification dimension has been set for the corpus to be labeled, but a dimension description label for the grading dimension has not been set, then the unlabeled descriptive dimensions of the corpus to be labeled include the grading dimension.
[0036] Step S20: Select the dimension to be labeled from the unlabeled description dimensions.
[0037] In practical use, at least one of the unannotated description dimensions can be selected as the description dimension to be annotated.
[0038] When selecting the unlabeled description dimensions, the higher priority unlabeled description dimensions can be selected as the unlabeled description dimensions based on their respective priorities.
[0039] For example: Suppose there are three unlabeled description dimensions: K1, K2, and K3. The priority is K1 > K3 > K2. If there are two description dimensions to be selected, then K1 and K3 will be selected as the description dimensions to be labeled.
[0040] Of course, depending on actual needs, the administrators of the label setting equipment can also pre-set mandatory dimensions, that is, dimensions that must be labeled. In this case, the unlabeled description dimensions that have been set as mandatory dimensions can be selected first as the description dimensions to be labeled. If there are still remaining selections, further selections can be made according to priority.
[0041] In practical applications, after determining the dimensions of description to be labeled, a label setting task can be constructed based on the corpus to be labeled and the dimensions of description to be labeled, and the label setting task can be distributed to the labelers so that they can perform label labeling based on the label setting task.
[0042] The labeler can be a user using a label setting device, or a user who can communicate with the label setting device using other identification methods to perform the labeling. This embodiment does not impose any restrictions on this.
[0043] In practical implementation, in order to ensure the efficiency and effectiveness of label setting, before assigning label setting tasks, labelers can be grouped according to their historical labeling accuracy or labeling preferences, so that different description dimensions to be labeled correspond to different groups. Then, the label setting tasks are distributed to the labelers in the groups corresponding to the description dimensions to be labeled, ensuring that they can be assigned to more suitable labelers for labeling.
[0044] Step S30: In response to the annotation action corresponding to the dimension to be annotated, set the corresponding dimension description label for the corpus to be labeled.
[0045] In practical use, after the label setting task is distributed to the labelers, the label setting device can monitor the labeling actions performed by the labelers according to the label setting task, and set corresponding dimension description labels for the corpus to be labeled based on the labeling actions.
[0046] For example, after generating a tag setting task, the tag setting device distributes the task to the annotators. The annotator displays the annotation page on their device according to the tag setting task. Then, it monitors the annotator's annotation actions on the annotation page (such as entering a tag and clicking submit), and sets the corresponding dimension description tags for the corpus to be labeled based on the annotation actions.
[0047] In practical applications, after setting the corresponding dimension description labels for the corpus to be labeled, it is also possible to detect whether there are still unlabeled description dimensions in the corpus to be labeled. If so, the process can return to step S20 until each description dimension of the corpus to be labeled has been set with the corresponding dimension description labels.
[0048] In practice, the administrator of the tag setting device can also set the tag setting status for each description dimension of each corpus to indicate whether a corresponding dimension description tag is set for a certain description dimension of the corpus. At this time, when setting the corresponding dimension description tag for the corpus to be labeled, the tag setting status of the description dimension corresponding to the dimension description tag can also be changed to set at the same time.
[0049] Understandably, if a label setting state exists, it is possible to check whether the corpus to be labeled still has an unset label setting state. If so, it is possible to return to step S20 until each description dimension of the corpus to be labeled has a corresponding dimension description label set.
[0050] For example: Suppose that the description dimension has two dimensions, classification and grading. If a piece of corpus has only been classified and labeled by user C (such as "medical"), and its grading status is still displayed as "unlabeled", then we can return to step S20, treat the grading dimension of the corpus as the unlabeled description dimension, and continue to set the dimension description label of the grading dimension for the corpus.
[0051] In a specific implementation, to ensure that annotation can be performed in batches, the following steps may be included before step S10 in this embodiment: Statistical analysis was performed on the dimension description tags corresponding to each corpus in the corpus to determine the total number of tags set for each corpus. The target corpus is the corpus in which the total number of labels set is less than the total number of preset dimensions. Select the text to be annotated from the target text; Accordingly, after step S30, the following may also be included: Detect whether there are target corpora that were not selected as corpora to be labeled; If it exists, return to the step of selecting the corpus to be annotated from the target corpus.
[0052] It should be noted that the corpus can be a database pre-stored with a large amount of data, and the corpus can be pre-configured by the administrator of the label setting device. The preset total number of dimensions can be the total number of descriptive dimensions to be labeled pre-set by the administrator of the label setting device. The total number of set labels can be the total number of dimension description labels for different dimensions that have been set for the corpus. For example, if corpus A has 3 dimension description labels, namely A, B, and C, but A and B belong to the same descriptive dimension, then the total number of set labels is 2.
[0053] It is understandable that if the total number of set labels for the corpus is less than the total number of preset dimensions, it means that the corpus has at least one description dimension without set dimension description labels, and therefore, it can be used as the target corpus.
[0054] In practical use, the amount of data used for training is generally large, ranging from thousands to tens of thousands or even millions. To avoid confusion or excessive resource consumption, the label setting device can set labels for each corpus in batches. Therefore, at least a preset number of corpora can be selected from the target corpus as the corpus to be labeled. The preset number can be set in advance by the administrator of the label setting device according to the performance of the label setting device, such as setting the preset number to 100.
[0055] Understandably, in order to ensure that each corpus can be tagged, after setting the corresponding dimension description tags for the corpus to be tagged, it is also possible to detect whether there are target corpora that have not been selected as corpora to be tagged. If so, the step of selecting corpora to be tagged from the target corpora is returned, and tagging is continued. If the label does not exist, it can be determined that the label setting is complete. At this time, a prompt can be given so that the administrator of the label setting device can perform subsequent processing in a timely manner, such as model training.
[0056] This embodiment provides a label setting method. Since it can automatically determine the unlabeled descriptive dimensions of the corpus to be labeled and select the descriptive dimensions to be labeled from them, and then, based on the annotation actions of the annotators for the descriptive dimensions to be labeled, the corresponding dimension description labels are automatically set for the corpus to be labeled, thereby realizing multi-dimensional labeling of the corpus and improving the information dimensions and label coverage of the corpus.
[0057] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S30 includes steps S301 to S304: Step S301: In response to the annotation action corresponding to the dimension to be annotated, determine the corpus description tag corresponding to the annotation action.
[0058] It should be noted that the label setting device can determine the dimension description labels, i.e., corpus description labels, set by the labeler for the corpus to be labeled based on the labeling actions of the labeler.
[0059] For example: based on the user's submission action, read the tags that the user entered or selected in the tag settings bar, and use those tags as corpus description tags.
[0060] Step S302: Analyze the corpus to be labeled using a large model to determine the optional description tags corresponding to the description dimension to be labeled.
[0061] Step S303: Match the corpus description tags with the optional description tags.
[0062] Step S304: When a match is successful, set the dimension description label corresponding to the corpus to be labeled as the corpus description label.
[0063] In practical use, since annotators usually annotate a large amount of data, annotation errors may occur. In order to deal with such annotation errors, a large model can be used to analyze the corpus to be annotated, and the large model can deduce the dimension description labels that the corpus to be annotated may correspond to in the dimension to be annotated, thereby obtaining optional description labels. There can be multiple optional description labels.
[0064] Next, the corpus description tags are matched with the optional description tags to determine whether there are any annotation errors, in order to avoid the mistakes caused by a large amount of manual annotation.
[0065] In practical use, if the corpus description label and the optional description label match successfully, it means that the corpus description label set by the annotator is within a reasonable range. Therefore, the dimension description label corresponding to the corpus to be labeled can be set to the corpus description label set by the annotator.
[0066] In a specific implementation, in order to reasonably determine whether the corpus description tags set by the annotators are reasonable, step S303 in this embodiment may include: Perform consistency matching between the corpus description tags and each optional description tag; If there is no optional description tag that matches the corpus description tag, then the corpus description tag is matched with each optional description tag to determine the tag similarity of each optional description tag; If there is an optional descriptive tag with a corresponding tag similarity greater than or equal to the preset similarity threshold, then the match is considered successful.
[0067] It should be noted that consistency matching can be a comparison of whether the corpus description labels and optional description labels are completely consistent.
[0068] It is understandable that if there are optional description tags that are consistent with the corpus description tags, it means that the dimension description tags set by the annotator are consistent with a certain dimension description tag derived by the large model. In this case, it can be determined that the corpus description tags set by the annotator are within a reasonable range, and the match can be determined to be successful. If there are no optional description tags that match the corpus description tags, it means that the dimension description tags set by the annotators are not completely consistent with the dimension description tags derived by the large model. This may be due to differences in synonyms or expressions. Therefore, the corpus description tags and each optional description tag can be matched for similarity to determine the tag similarity corresponding to each optional description tag.
[0069] In the similarity matching process, semantic similarity comparison can be used. For example, semantic similarity comparison can be performed using word embedding-based methods, and semantic similarity can be calculated using the Word2Vec model. Tag similarity represents the semantic similarity between optional description tags and corpus description tags. The greater the tag similarity, the closer the semantics between the optional description tags and corpus description tags.
[0070] It is understandable that if there are optional descriptive tags with a corresponding tag similarity greater than or equal to the preset similarity threshold, it means that the dimension descriptive tags set by the annotator and the dimension descriptive tags derived by the large model are not completely consistent. However, the overall semantics are similar. Therefore, it can be considered that the corpus descriptive tags set by the annotator are within a reasonable range, and the match can be determined to be successful.
[0071] The preset similarity threshold can be set in advance by the administrator of the tag setting device, for example, setting the preset similarity threshold to 85%.
[0072] In specific implementations, to minimize annotation errors, this embodiment, if no alternative description tag matches the corpus description tag, performs similarity matching between the corpus description tag and each alternative description tag to determine the tag similarity corresponding to each alternative description tag. The implementation may further include: If there is no corresponding optional descriptive label with a similarity greater than or equal to a preset similarity threshold, then a label confirmation task is constructed based on the corpus descriptive labels and the corpus to be labeled. Obtain at least one tag confirmation result from the tag setting expert for the tag confirmation task; Determine the label correctness confidence rate based on the confirmation results of at least one label; If the correct confidence rate of the label is greater than the preset confidence threshold, then the match is considered successful.
[0073] It should be noted that the label setting expert can be a human expert or a pre-trained neural network model. If a neural network model is used as the label setting expert, different label setting experts can be pre-trained using different model structures and / or different model training sets.
[0074] It is understandable that if there are no corresponding optional descriptive tags with a similarity greater than or equal to the preset similarity threshold, it means that the dimension description tags set by the annotators are not completely consistent with the dimension description tags derived by the large model, and the overall semantics are also quite different, requiring further judgment. Therefore, a tag confirmation task can be constructed and distributed to at least one tag setting expert for further analysis.
[0075] In practical use, the label confirmation results include two types: incorrect setting and correct setting. The proportion of correct setting in all label confirmation results can be counted, and this proportion can be used as the label correctness confidence rate. If the label correctness confidence rate is greater than the preset confidence threshold, it means that most label setting experts believe that the setting is correct. Therefore, the match can be determined to be successful.
[0076] The preset information threshold can be set in advance by the administrator of the tag setting device, for example, setting the preset information threshold to 70%.
[0077] If the label correctness confidence rate is less than or equal to the preset confidence threshold, it means that a large number of label setting experts have identified the setting as incorrect. In this case, it can be determined that the labeling personnel have made a mistake. Instead of setting the corresponding dimension description labels for the corpus to be labeled, the constructed label setting task can be recycled and redistributed to other labeling personnel for re-labeling.
[0078] In a specific implementation, in order to minimize unnecessary performance consumption, the following steps may be included before step S302 in this embodiment: Obtain the set description tags of the corpus to be labeled; Detect whether the corpus description tags conflict with the already set description tags; If there is no conflict, then proceed with the step of analyzing the corpus to be labeled using a large model to determine the optional description tags corresponding to the description dimension to be labeled.
[0079] It should be noted that tag matching involves a relatively complex algorithm and requires calling a large model, resulting in high overall performance consumption. Therefore, it is necessary to minimize unnecessary performance consumption. Thus, a tag conflict detection engine can be pre-set to detect whether the corpus description tags conflict with the pre-set description tags. If a conflict exists, it means that the dimension description labels set by the annotator at this time are significantly different from the dimension description labels set previously. Therefore, it can be determined that the label setting by the annotator has made a mistake. At this time, the label setting task can be reclaimed and redistributed to other annotators for re-annotation. If there is no conflict, further analysis is required. Therefore, step S302 and subsequent steps can be executed to make a further judgment.
[0080] The conflict detection engine can be pre-configured by the administrator of the tag setting device.
[0081] In practical implementation, setting up a conflict detection engine is quite difficult, and running a unified engine is relatively performance-intensive. Therefore, a method with lower performance requirements can be used to implement early conflict verification. In this case, the detection of whether the corpus description tag conflicts with the set description tag in this embodiment can include: Based on the set description tags and the search conditions of the corpus description tag construction rules; Check if a conflicting rule corresponding to the search condition of the specified rule exists in the conflict rule base; If there is a conflicting rule corresponding to the rule search condition, then a conflict is determined to exist; If there is no conflicting rule corresponding to the rule search condition, then it is determined that there is no conflict.
[0082] It should be noted that multiple conflict rules can be pre-stored in the conflict rule base, and these conflict rules can be pre-configured by the administrator of the tag setting device.
[0083] It is understandable that if a conflicting rule exists in the conflict rule base corresponding to the rule search condition, it means that the description tag and the corpus description tag have been set at this time and there is a contradiction. Therefore, it can be determined that there is a conflict. If there is no conflicting rule corresponding to the rule search condition in the conflict rule base, it means that at least for now, it can be determined that the set description label and the corpus description label are not obviously contradictory. Therefore, it can be determined that there is no conflict.
[0084] In practical use, the administrators of the label setting device can also collect conflicting description labels in real time, construct new conflict rules, and add the new conflict rules to the conflict rule library.
[0085] Understandably, instead of building an engine directly, a dynamically updatable conflict rule base is used, which makes the overall setup simpler, more flexible, and allows for updates at any time.
[0086] This embodiment provides a label setting method. By introducing a large model to derive optional description labels, this embodiment verifies whether the labeling personnel have made mistakes in their labeling. This ensures that labeling errors can be automatically avoided without the need for additional manual processing, thereby improving the reliability of label setting.
[0087] Based on the first embodiment of this application, in the third embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 Step S20 includes steps S201 to S202: Step S201: Obtain the task backlog ratio corresponding to each description dimension.
[0088] It should be noted that the task backlog ratio can be the proportion of tasks that have been built but not yet completed under the label setting in this description dimension.
[0089] In a specific implementation, the proportion of tasks to the total number of tasks can be used as the task backlog ratio. In this case, step S201 in this embodiment may include: The task list for setting tags is statistically analyzed to obtain the total backlog of tasks and the number of tag setting tasks corresponding to each description dimension. The task accumulation ratio corresponding to each descriptive dimension is determined based on the task quantity set by the label and the total task accumulation.
[0090] It should be noted that the tag setting task list can be a data table used to record tag setting tasks that have been generated but not yet completed.
[0091] In practical use, the number of tag setting tasks in the tag setting task list can be counted to determine the total backlog of tasks. Then, based on the description dimension to which they belong, further classification and statistics can be performed to determine the number of tag setting tasks corresponding to each description dimension. Finally, the ratio between the number of tag setting tasks corresponding to the description dimension and the total backlog of tasks can be used as the task backlog ratio corresponding to the description dimension.
[0092] In a specific implementation, in order to more reasonably determine the task backlog ratio, step S201 in this embodiment may include: Retrieve the task allocation thread pool corresponding to each description dimension; Get the number of waiting sequence elements corresponding to the task allocation thread pool; The task backlog ratio corresponding to each descriptive dimension is determined based on the number of waiting sequence elements.
[0093] It should be noted that the number of elements in the waiting sequence can be the total number of elements filled in the waiting sequence of the task allocation thread pool corresponding to the dimension.
[0094] When actually executing tag setting tasks, tag setting tasks for different description dimensions may be executed by different threads, or even by different devices or clusters. In this case, a corresponding task allocation thread pool can be set for each description dimension. The task allocation thread pool is responsible for distributing and managing the tag setting tasks for the corresponding description dimension.
[0095] In this case, the task backlog ratio corresponding to the description dimension can be determined based on the number of waiting sequence elements corresponding to the task allocation thread pool. For example, if the total number of waiting sequence elements corresponding to the task allocation thread pool is SA and the number of waiting sequence elements is SB, then the task backlog ratio of this description dimension is P=SB / SA.
[0096] Step S202: Select the description dimension to be labeled from the unlabeled description dimensions based on the task stacking ratio.
[0097] In practical use, the unlabeled description dimensions can be sorted from low to high based on the task backlog ratio corresponding to each description dimension. Then, the top N unlabeled description dimensions are selected as the description dimensions to be labeled.
[0098] N can be preset by the administrator of the tag setting device, for example, N can be set to 1 or 2.
[0099] In a specific implementation, the dimension to be labeled can also be selected based on the annotation coverage. Step S20 in this embodiment may include: Statistical analysis was performed on the dimension description tags corresponding to each corpus in the corpus to determine the total amount of corpus and the total number of tags set for each description dimension. The annotation coverage rate corresponding to each description dimension is determined based on the total amount of the corpus and the total amount of the tag settings. The unlabeled descriptive dimensions are sorted based on the labeled coverage, and the descriptive dimensions to be labeled are determined based on the sorting results.
[0100] It should be noted that when labeling data for model training, in some cases, the labeling may not be completed before stopping. To accommodate this situation, it is necessary to ensure that the label coverage of each description dimension is balanced as much as possible. Therefore, the description dimension to be labeled can be selected based on the label coverage.
[0101] In practical use, the dimension description tags corresponding to each corpus in the corpus can be counted to determine the total number of existing corpora, and the total number of corpora with dimension description tags for each set description dimension can be counted to obtain the total number of tag settings. Then, the ratio of the total number of tag settings to the total number of corpora is used as the annotation coverage rate corresponding to the description dimension.
[0102] For example: Suppose the total amount of corpus is K, and there are 3 description dimensions. The total number of corpus with dimension description labels corresponding to each description dimension is K1, K2, and K3. That is, the total number of labels set for each description dimension is K1, K2, and K3, respectively. Then, the label coverage for each description dimension is K1 / K, K2 / K, and K3 / K, respectively.
[0103] In practical applications, the label setting method of this embodiment can be executed collaboratively by multiple functional modules, for example: Assuming there are two types of description dimensions: categorical and hierarchical, then the label setting method can be executed collaboratively by the following modules: Corpus storage module: Stores raw corpus data, i.e., the corpus; The classification and annotation module: provides dimensional description labels for the corpus based on semantic or topical classification dimensions; The hierarchical annotation module adds dimensional description labels to the corpus based on dimensions such as quality or grade. State tracking module: records the label annotation status of each corpus in two descriptive dimensions; Management console: Provides functions for task assignment, progress monitoring, and result review; Among them, the classification dimensions can include labels describing the dimensions of the corpus content, such as "finance", "law", "medicine", and "entertainment". The grading dimensions can include descriptive labels such as "high quality", "medium quality", "low quality" or "A / B / C", which reflect the value or usability of the corpus. Furthermore, the dimension description labels for the two dimensions are stored separately, independently controllable, and do not affect each other. The setting of dimension description labels allows for asynchronous execution.
[0104] In actual annotation, annotators can receive either classification tasks or hierarchical tasks, and only one description dimension is bound when receiving a task; Once the labeling is complete, the status of that dimension is set to "Completed," while the other dimension remains "Unlabeled" and can continue to be claimed. It supports cross-annotation, meaning that different annotators can operate on different dimensions simultaneously without interfering with each other's annotation process; It supports auxiliary capabilities such as tag conflict detection, dimension status reminders, and tag log tracking. For example, each generation, execution, and revocation of a tag setting task is logged accordingly.
[0105] In addition, the backend management terminal of the tag setting device supports dual-dimensional task assignment, tagging task collection, and dimension coverage analysis; it can configure the priority, mandatory fields, and default value mechanism for each tag; and it supports viewing the dual-tag status, historical change records, and user contributions for each corpus.
[0106] This embodiment provides a label setting method. When selecting the description dimension to be labeled, this embodiment takes into account the task accumulation ratio of each description dimension to ensure that the task accumulation of a certain description dimension is not too large, thus avoiding the lag caused by excessive resource consumption due to loading too many tasks at once.
[0107] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the label setting method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0108] This application also provides a label setting device, please refer to... Figure 4 The label setting device includes: The determination module 10 is used to determine the unlabeled descriptive dimensions of the corpus to be labeled, wherein the descriptive dimension is an independent direction that limits or describes the corpus; Selection module 20 is used to select a dimension to be labeled from the unlabeled description dimensions; Setting module 30 is used to set corresponding dimension description labels for the corpus to be labeled in response to the labeling action corresponding to the dimension to be labeled.
[0109] The tag setting device provided in this application, employing the tag setting method in the above embodiments, can solve the technical problem that related technologies, when annotating corpora, only perform annotation from a single dimension, resulting in poor practical application effects. Compared with the prior art, the beneficial effects of the tag setting device provided in this application are the same as those of the tag setting method provided in the above embodiments, and other technical features in the tag setting device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0110] This application provides a label setting device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the label setting method in the first embodiment described above.
[0111] The following is for reference. Figure 5The diagram illustrates a structural schematic of a tag setting device suitable for implementing embodiments of this application. The tag setting device in this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The label setting device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0112] like Figure 5 As shown, the tag setting device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the tag setting device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the tag setting device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows tag setting devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0113] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0114] The tag setting device provided in this application, employing the tag setting method described in the above embodiments, can solve the technical problem that related technologies, when annotating corpora, only perform annotation from a single dimension, resulting in poor actual usage effects. Compared with the prior art, the beneficial effects of the tag setting device provided in this application are the same as those of the tag setting method provided in the above embodiments, and other technical features of this tag setting device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0115] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0116] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0117] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the label setting method in the above embodiments.
[0118] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0119] The aforementioned computer-readable storage medium may be included in the tag setting device; or it may exist independently and not be assembled into the tag setting device.
[0120] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a tag setting device, cause the tag setting device to: determine unlabeled descriptive dimensions of the corpus to be labeled, wherein the descriptive dimension is an independent direction for limiting or describing the corpus; select a descriptive dimension to be labeled from the unlabeled descriptive dimensions; and, in response to a labeling action corresponding to the descriptive dimension to be labeled, set a corresponding dimension description tag for the corpus to be labeled.
[0121] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Python, Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0123] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0124] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described tag setting method. This solves the technical problem that related technologies, when tagging corpora, only perform tagging from a single dimension, resulting in poor practical effectiveness. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the tag setting method provided in the above embodiments, and will not be elaborated upon here.
[0125] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the label setting method described above.
[0126] The computer program product provided in this application can solve the technical problem that related technologies, when annotating corpora, only annotate from a single dimension, resulting in poor practical effects. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the tag setting method provided in the above embodiments, and will not be repeated here.
[0127] All user-related data involved in this application (such as user privacy data, user behavior data, etc.) were obtained with the user's permission or consent; that is to say, when this application is used in a specific product or technology, user permission is required to obtain and process the relevant data, and the processing of the relevant data must comply with the relevant laws, regulations and regulatory standards of the relevant countries and regions.
[0128] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.
[0129] This application also discloses A1, a label setting method, the label setting method comprising: Determine the unlabeled descriptive dimensions of the corpus to be labeled, where the descriptive dimension is an independent direction that limits or describes the corpus; Select the dimension to be labeled from the unlabeled description dimensions; In response to the annotation action corresponding to the dimension to be annotated, set the corresponding dimension description label for the corpus to be labeled.
[0130] A2. The label setting method as described in A1, wherein the step of setting corresponding dimension description labels for the corpus to be labeled in response to the labeling action corresponding to the dimension to be labeled includes: In response to the annotation action corresponding to the dimension to be annotated, determine the corpus description tag corresponding to the annotation action; The corpus to be labeled is analyzed using a large model to determine the optional descriptive tags corresponding to the descriptive dimensions to be labeled. Match the corpus description tags with the optional description tags; When a match is successful, the dimension description label corresponding to the corpus to be labeled is set as the corpus description label.
[0131] A3. The tag setting method as described in A2, wherein matching the corpus description tags with the optional description tags includes: Perform consistency matching between the corpus description tags and each optional description tag; If there is no optional description tag that matches the corpus description tag, then the corpus description tag is matched with each optional description tag to determine the tag similarity of each optional description tag; If there is an optional descriptive tag with a corresponding tag similarity greater than or equal to the preset similarity threshold, then the match is considered successful.
[0132] A4. The tag setting method as described in A3, further comprising, after determining the tag similarity between the corpus description tag and each optional description tag respectively, if no optional description tag is found that matches the corpus description tag, and then performing similarity matching on each optional description tag, the method also includes: If there is no corresponding optional descriptive label with a similarity greater than or equal to a preset similarity threshold, then a label confirmation task is constructed based on the corpus descriptive labels and the corpus to be labeled. Obtain at least one tag confirmation result from the tag setting expert for the tag confirmation task; Determine the label correctness confidence rate based on the confirmation results of at least one label; If the correct confidence rate of the label is greater than the preset confidence threshold, then the match is considered successful.
[0133] A5. The label setting method as described in A2, before analyzing the corpus to be labeled using a large model to determine the optional description labels corresponding to the description dimension to be labeled, further includes: Obtain the set description tags of the corpus to be labeled; Detect whether the corpus description tags conflict with the already set description tags; If there is no conflict, then proceed with the step of analyzing the corpus to be labeled using a large model to determine the optional description tags corresponding to the description dimension to be labeled.
[0134] A6. The tag setting method as described in A5, wherein detecting whether the corpus description tag conflicts with the already set description tag includes: Based on the set description tags and the search conditions of the corpus description tag construction rules; Check if a conflicting rule corresponding to the search condition of the specified rule exists in the conflict rule base; If there is a conflicting rule corresponding to the rule search condition, then a conflict is determined to exist; If there is no conflicting rule corresponding to the rule search condition, then it is determined that there is no conflict.
[0135] A7. The label setting method as described in A1, wherein selecting the description dimension to be labeled from the unlabeled description dimensions includes: Statistical analysis was performed on the dimension description tags corresponding to each corpus in the corpus to determine the total amount of corpus and the total number of tags set for each description dimension. The annotation coverage rate corresponding to each description dimension is determined based on the total amount of the corpus and the total amount of the tag settings. The unlabeled descriptive dimensions are sorted based on the labeled coverage, and the descriptive dimensions to be labeled are determined based on the sorting results.
[0136] A8. The label setting method as described in A1, wherein selecting the description dimension to be labeled from the unlabeled description dimensions includes: Obtain the task backlog ratio corresponding to each description dimension; Based on the task stacking ratio, a description dimension to be labeled is selected from the unlabeled description dimensions.
[0137] A9. The tag setting method as described in A8, wherein obtaining the task backlog ratio corresponding to each description dimension includes: The task list for setting tags is statistically analyzed to obtain the total backlog of tasks and the number of tag setting tasks corresponding to each description dimension. The task accumulation ratio corresponding to each descriptive dimension is determined based on the task quantity set by the label and the total task accumulation.
[0138] A10. The tag setting method as described in A8, wherein obtaining the task backlog ratio corresponding to each description dimension includes: Retrieve the task allocation thread pool corresponding to each description dimension; Get the number of waiting sequence elements corresponding to the task allocation thread pool; The task backlog ratio corresponding to each descriptive dimension is determined based on the number of waiting sequence elements.
[0139] A11. The label setting method as described in any one of A1-A10, before determining the unlabeled descriptive dimensions of the corpus to be labeled, further includes: Statistical analysis was performed on the dimension description tags corresponding to each corpus in the corpus to determine the total number of tags set for each corpus. The target corpus is the corpus in which the total number of labels set is less than the total number of preset dimensions. Select the text to be annotated from the target text; Accordingly, after setting corresponding dimension description labels for the corpus to be labeled in response to the annotation action corresponding to the dimension to be labeled, the method further includes: Detect whether there are target corpora that were not selected as corpora to be labeled; If it exists, return to the step of selecting the corpus to be annotated from the target corpus.
[0140] This application also discloses B12, a label setting device, the label setting device comprising: The determination module is used to determine the unlabeled descriptive dimensions of the corpus to be labeled, where the descriptive dimension is an independent direction that limits or describes the corpus; The selection module is used to select a dimension to be labeled from the unlabeled description dimensions. The setting module is used to respond to the annotation action corresponding to the dimension to be labeled and set the corresponding dimension description label for the corpus to be labeled.
[0141] B13. The tag setting device as described in B12, wherein the setting module is further configured to, in response to the annotation action corresponding to the dimension to be annotated, determine the corpus description tag corresponding to the annotation action; analyze the corpus to be annotated through a large model to determine the optional description tag corresponding to the dimension to be annotated; match the corpus description tag with the optional description tag; and, when the match is successful, set the dimension description tag corresponding to the corpus to be annotated as the corpus description tag.
[0142] B14. The tag setting device as described in B13, wherein the setting module is further configured to perform consistency matching between the corpus description tag and each optional description tag; if there is no optional description tag consistent with the corpus description tag, then perform similarity matching between the corpus description tag and each optional description tag to determine the tag similarity corresponding to each optional description tag; if there is an optional description tag with a corresponding tag similarity greater than or equal to a preset similarity threshold, then the matching is determined to be successful.
[0143] B15. The tag setting device as described in B14, wherein the setting module is further configured to: if there is no corresponding optional descriptive tag with a tag similarity greater than or equal to a preset similarity threshold, construct a tag confirmation task based on the corpus descriptive tags and the corpus to be labeled; obtain at least one tag confirmation result from a tag setting expert corresponding to the tag confirmation task; determine the tag correctness confidence rate based on the at least one tag confirmation result; and determine a successful match if the tag correctness confidence rate is greater than a preset confidence threshold.
[0144] B16. The tag setting device as described in B13, wherein the setting module is further configured to acquire the set description tags of the corpus to be labeled; detect whether the corpus description tags conflict with the set description tags; if there is no conflict, perform the step of analyzing the corpus to be labeled through a large model to determine the optional description tags corresponding to the description dimension to be labeled.
[0145] B17. The tag setting device as described in B16, wherein the setting module is further configured to construct rule search conditions based on the set description tags and the corpus description tags; detect whether there is a conflicting rule in the conflict rule base corresponding to the rule search conditions; if there is a conflicting rule corresponding to the rule search conditions, then determine that there is a conflict; if there is no conflicting rule corresponding to the rule search conditions, then determine that there is no conflict.
[0146] This application also discloses C18, a label setting device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the label setting method as described above.
[0147] This application also discloses D19, a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the tag setting method as described above.
[0148] This application also discloses E20, a computer program product comprising a computer program that, when executed by a processor, implements the steps of the label setting method as described above.
Claims
1. A label setting method, characterized in that, The label setting method includes: Determine the unlabeled descriptive dimensions of the corpus to be labeled, where the descriptive dimension is an independent direction that limits or describes the corpus; Select the dimension to be labeled from the unlabeled description dimensions; In response to the annotation action corresponding to the dimension to be annotated, set the corresponding dimension description label for the corpus to be labeled.
2. The label setting method as described in claim 1, characterized in that, The step of responding to the annotation action corresponding to the dimension to be annotated, and setting corresponding dimension description labels for the corpus to be labeled, includes: In response to the annotation action corresponding to the dimension to be annotated, determine the corpus description tag corresponding to the annotation action; The corpus to be labeled is analyzed using a large model to determine the optional descriptive tags corresponding to the descriptive dimensions to be labeled. Match the corpus description tags with the optional description tags; When a match is successful, the dimension description label corresponding to the corpus to be labeled is set as the corpus description label.
3. The label setting method as described in claim 2, characterized in that, The step of matching the corpus description tags with the optional description tags includes: Perform consistency matching between the corpus description tags and each optional description tag; If there is no optional description tag that matches the corpus description tag, then the corpus description tag is matched with each optional description tag to determine the tag similarity of each optional description tag; If there is an optional descriptive tag with a corresponding tag similarity greater than or equal to the preset similarity threshold, then the match is considered successful.
4. The label setting method as described in claim 3, characterized in that, If no alternative description tag matches the corpus description tag, then after performing similarity matching between the corpus description tag and each alternative description tag to determine the tag similarity corresponding to each alternative description tag, the method further includes: If there is no corresponding optional descriptive label with a similarity greater than or equal to a preset similarity threshold, then a label confirmation task is constructed based on the corpus descriptive labels and the corpus to be labeled. Obtain at least one tag confirmation result from the tag setting expert for the tag confirmation task; Determine the label correctness confidence rate based on the confirmation results of at least one label; If the correct confidence rate of the label is greater than the preset confidence threshold, then the match is considered successful.
5. The label setting method as described in claim 2, characterized in that, Before analyzing the corpus to be labeled using a large model to determine the optional description tags corresponding to the description dimension to be labeled, the method further includes: Obtain the set description tags of the corpus to be labeled; Detect whether the corpus description tags conflict with the already set description tags; If there is no conflict, then proceed with the step of analyzing the corpus to be labeled using a large model to determine the optional description tags corresponding to the description dimension to be labeled.
6. The label setting method as described in claim 5, characterized in that, The detection of whether the corpus description tags conflict with the already set description tags includes: Based on the set description tags and the search conditions of the corpus description tag construction rules; Check if a conflicting rule corresponding to the search condition of the specified rule exists in the conflict rule base; If there is a conflicting rule corresponding to the rule search condition, then a conflict is determined to exist; If there is no conflicting rule corresponding to the rule search condition, then it is determined that there is no conflict.
7. A label setting device, characterized in that, The label setting device includes: The determination module is used to determine the unlabeled descriptive dimensions of the corpus to be labeled, where the descriptive dimension is an independent direction that limits or describes the corpus; The selection module is used to select a dimension to be labeled from the unlabeled description dimensions. The setting module is used to respond to the annotation action corresponding to the dimension to be labeled and set the corresponding dimension description label for the corpus to be labeled.
8. A label setting device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the tag setting method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the label setting method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the label setting method as described in any one of claims 1 to 6.