A microtask corpus data cleaning method

Through the micro-task corpus data cleaning method, the corpus classification task and corpus editing task are separated, and the translator and system algorithms of different levels are used for cleaning, which solves the problems of high cost and low efficiency of manual cleaning in the existing technology, and achieves efficient and low-cost corpus data cleaning effect.

CN114564972BActive Publication Date: 2025-06-06IOL WUHAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210206766.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-03
Publication Date
2025-06-06
Estimated Expiration
2042-03-03

AI Technical Summary

Technical Problem

The existing corpus data cleaning methods rely on a large amount of manual screening, resulting in high cost, low efficiency, and high requirements for interpreter capabilities.

Method used

The micro-task corpus data cleaning method is adopted to separate corpus classification tasks and corpus editing tasks, and the cleaning is carried out using translators of different levels, and automatic review and confirmation is carried out through system algorithms to reduce workload and improve efficiency.

Benefits of technology

It improves the efficiency and quality of corpus data cleaning, reduces cleaning costs, and reduces manual work burden through automatic audit and confirmation mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114564972B_ABST
    Figure CN114564972B_ABST
Patent Text Reader

Abstract

The present invention discloses a micro-task corpus data cleaning method, which specifically includes the following steps: S1, pre-embed the corpus data of known results into the corpus data to be cleaned to form the corpus embedding data and then start cleaning; S2, configure the cleaning parameters of the corpus data; S3, clean the corpus classification task; S4, calculate the classification result of the corpus classification task: 1) obtain the corpus whose classification result can be confirmed; 2) calculate the credibility of the first-level translator processing the corpus classification task; 3) confirm the classification result of the corpus classification task; 4) review the classification result of the corpus S5, clean the corpus editing task; S6, quality check the edited corpus. The present invention includes the cleaning of corpus classification tasks and corpus editing tasks, and uses translators of different levels to clean different tasks, which is highly targeted and improves the cleaning efficiency. At the same time, the cleaning tasks are automatically audited and confirmed by the system algorithm, which can reduce the cleaning workload and save the cleaning cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of corpus data cleaning, and in particular relates to a micro-task corpus data cleaning method. Background Art

[0002] Large-scale, high-quality pre-training corpus can allow the model to learn more knowledge expressions and better understand the various representations and meanings of words, thus becoming more intelligent. Currently, most machine translations are based on neural networks, and neural network-based translation methods require a large amount of corpus data to train the machine translation engine. How to select high-quality corpus from a large amount of corpus data to achieve corpus cleaning.

[0003] The existing corpus data cleaning method mainly obtains the target corpus through a large amount of manual screening, but the cleaning cost of this cleaning method is too high. In addition, the manual checking and modification of the corpus data one by one has low work efficiency and also requires high ability of the translator. Summary of the invention

[0004] In order to solve the above-mentioned technical problems, the present invention provides a micro-task corpus data cleaning method, which includes the cleaning of corpus classification tasks and corpus editing tasks. Different tasks are cleaned with the help of translators of different levels, which is highly targeted and improves the cleaning efficiency. At the same time, the cleaning tasks are automatically reviewed and confirmed with the help of system algorithms, which can reduce the cleaning workload and save cleaning costs.

[0005] The technical solution adopted by the present invention is:

[0006] A microtask corpus data cleaning method specifically comprises the following steps:

[0007] S1. The corpus data to be cleaned is pre-embedded with the corpus data of known results to form corpus embedding data and then the cleaning begins;

[0008] S2. Configure the cleaning parameters of the corpus data;

[0009] S3. Cleaning corpus classification tasks: The system assigns corpus classification tasks to the first-level translator, who classifies and processes one or more corpus classification tasks, where each corpus classification task includes one or more task items, and each task item corresponds to a piece of corpus;

[0010] S4. Calculate the classification results of the corpus classification task:

[0011] After completing the corpus classification task, the first-level translator automatically calculates the classification results and obtains the corpus data that can be directly used and the corpus data that needs to be edited. The specific steps include the following:

[0012] 1) Obtain corpus for which classification results can be confirmed:

[0013] When a piece of corpus is processed and classified by multiple first-level translators, if the classification results of all first-level translators are the same, then the piece of corpus can be confirmed; when the corpus classification task is corpus embedded data, the classification result of the corpus classification task processing can be confirmed;

[0014] 2) Calculate the credibility of the first-level translator in handling the corpus classification task:

[0015] Obtain all the corpus with known results from the corpus classification task in which the first-level translator participates, denoted as A; calculate all the correct classification results, denoted as C; let RE = the credibility of the first-level translator in the corpus classification task, then After the calculation is completed, the credibility of the corpus classification task processed this time is included in the historical credibility of the first-level translator;

[0016] Let RE 1 ,RE 2 ,RE 3 ,...RE n is the historical credibility of the first-level translator. Excluding the highest historical credibility record and the lowest historical credibility of the first-level translator, let REA = the final credibility of the first-level translator = the average credibility, then

[0017] 3) Confirm the classification results of the corpus classification task:

[0018] A corpus is classified by multiple first-level translators. The attributes that need to be modified and the modified attribute values ​​can be obtained from the configuration cleaning parameters. Let TV = modified attribute value, TVP = attribute value obtained by each first-level translator, REA is the average credibility of the first-level translators calculated in step 2), then TVP = TV*REA, and then calculate the attribute values ​​obtained by each first-level translator when the attribute definition of whether to be modified is "yes" and "no" respectively;

[0019] Let a = the sum of the attribute values ​​obtained by the first-level translator when the attribute definition of whether to modify is "yes", b = the sum of the attribute values ​​obtained by the first-level translator when the attribute definition of whether to modify is "no", and y = the attribute difference, then

[0020] Calculate y and compare the attribute difference y with the classification confirmation difference threshold in the configuration cleaning parameters: if y> classification confirmation difference threshold, the classification result of the corpus is confirmed to need to be modified; if y≦classification confirmation difference threshold, the classification result of the corpus is confirmed to not need to be modified, and the rest of the classification results cannot be confirmed;

[0021] 4) Review the classification results of the corpus:

[0022] The system automatically extracts the most difficult-to-confirm classification results from the corpus and generates the first batch of review tasks, which are manually reviewed by secondary translators.

[0023] After the secondary translator completes the first batch of review tasks, it automatically returns to step 3 in S4 to recalculate the classification results of the corpus classification task: if the classification results of the corpus that cannot be confirmed are still obtained after recalculation, the second batch of review tasks will continue to be automatically generated, and this cycle will be repeated until the classification results of all corpus tasks are confirmed;

[0024] S5. Corpus cleaning and editing tasks:

[0025] Based on the classification results obtained in S4, the translator of level 2 or above shall modify and improve the corpus data: 1) Determine whether the corpus should be edited based on the classification results of the reviewed corpus: if the corpus is confirmed to need to be modified, the corpus editing process shall be carried out; if the corpus is confirmed to not need to be modified, the corpus data cleaning is completed; if the corpus cannot be confirmed, the corpus shall be subject to the manual review process;

[0026] 2) After the corpus editing task is completed, it is determined whether the edited corpus needs to be classified again: if necessary, the edited corpus re-enters step S3 for corpus classification task cleaning; if not, the edited corpus undergoes a quality inspection process;

[0027] S6. Corpus after quality inspection and editing:

[0028] The corpus that does not need to be reclassified after editing will be manually quality-checked by a Level 2 translator or above: if the manual quality check passes, the corpus data is cleaned; if the manual quality check fails, return to S5 to re-do the corpus editing task and fill in the editing comments.

[0029] Furthermore, in S4, the classification results of the corpus are reviewed. If the classification results that cannot be confirmed by the corpus are not manually reviewed, the number of compensation times needs to be calculated. Specifically, one task processing person is added. After the task processing person configuration is completed, the classification results of the corpus classification task are recalculated. When the task processing person> the maximum number of classifications, the classification is terminated.

[0030] Furthermore, the cleaning parameters described in S2 include the number of corpus classifications, the number of classification task items, the configuration of corpus embedding data, whether to edit, manual review, manual quality inspection after editing, classification task category, whether attributes need to be modified, modified attribute values, classification confirmation difference threshold and maximum number of classifications.

[0031] Furthermore, the cleaning parameters in S2 also include task handlers, and the first-level translators, second-level translators, and translators above the second-level translators are distinguished by defining the levels of the task handlers.

[0032] The beneficial effects of the present invention are:

[0033] 1) The present invention pre-embeds corpus data with known results into the corpus data to be cleaned to form corpus embedding data, providing data support for the subsequent calculation of classification results;

[0034] 2) The present invention provides two modes, manual review and automatic review, in the result review of the corpus classification task and the result quality inspection process of the corpus editing task, so as to meet different cleaning requirements and improve the cleaning quality of the corpus data;

[0035] 3) Use the credibility of translators in handling corpus classification tasks to increase the accuracy and efficiency of corpus cleaning. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flow chart of a micro-task corpus data cleaning method of the present invention;

[0037] Figure 2 It is a flow chart of corpus classification task cleaning in a micro-task corpus data cleaning method of the present invention; DETAILED DESCRIPTION

[0038] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0039] It should be noted that in the description of the present invention, it should be noted that the orientation or position relationship indicated by the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc. is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0040] Example 1

[0041] like Figure 1 to Figure 2 As shown, a micro-task corpus data cleaning method specifically includes the following steps:

[0042] S1. The corpus data to be cleaned is pre-embedded with the corpus data of known results to form corpus embedding data and then the cleaning begins;

[0043] S2. Configure the cleaning parameters of the corpus data;

[0044] S3. Cleaning corpus classification tasks: The system assigns corpus classification tasks to the first-level translator, who classifies and processes one or more corpus classification tasks, where each corpus classification task includes one or more task items, and each task item corresponds to a piece of corpus;

[0045] S4. Calculate the classification results of the corpus classification task:

[0046] After completing the corpus classification task, the first-level translator automatically calculates the classification results and obtains the corpus data that can be directly used and the corpus data that needs to be edited. The specific steps include the following:

[0047] 1) Obtain corpus for which classification results can be confirmed:

[0048] When a piece of corpus is processed and classified by multiple first-level translators, if the classification results of all first-level translators are the same, then the piece of corpus can be confirmed; when the corpus classification task is corpus embedded data, the classification result of the corpus classification task processing can be confirmed;

[0049] 2) Calculate the credibility of the first-level translator in handling the corpus classification task:

[0050] Obtain all the corpus with known results from the corpus classification task in which the first-level translator participates, denoted as A; calculate all the correct classification results, denoted as C; let RE = the credibility of the first-level translator in the corpus classification task, then After the calculation is completed, the credibility of the corpus classification task processed this time is included in the historical credibility of the first-level translator;

[0051] Let RE 1 ,RE 2 ,RE 3 ,...RE n is the historical credibility of the first-level translator. Excluding the highest historical credibility record and the lowest historical credibility of the first-level translator, let REA = the final credibility of the first-level translator = the average credibility, then

[0052] 3) Confirm the classification results of the corpus classification task:

[0053] A corpus is classified by multiple first-level translators. The attributes that need to be modified and the modified attribute values ​​can be obtained from the configuration cleaning parameters. Let TV = modified attribute value, TVP = attribute value obtained by each first-level translator, REA is the average credibility of the first-level translators calculated in step 2), then TVP = TV*REA, and then calculate the attribute values ​​obtained by each first-level translator when the attribute definition of whether to be modified is "yes" and "no" respectively;

[0054] Let a = the sum of the attribute values ​​obtained by the first-level translator when the attribute definition of whether to modify is "yes", b = the sum of the attribute values ​​obtained by the first-level translator when the attribute definition of whether to modify is "no", and y = the attribute difference, then

[0055] Calculate y and compare the attribute difference y with the classification confirmation difference threshold in the configuration cleaning parameters: if y> classification confirmation difference threshold, the classification result of the corpus is confirmed to need to be modified; if y≦classification confirmation difference threshold, the classification result of the corpus is confirmed to not need to be modified, and the rest of the classification results cannot be confirmed;

[0056] 4) Review the classification results of the corpus:

[0057] The system automatically extracts the most difficult-to-confirm classification results from the corpus and generates the first batch of review tasks, which are manually reviewed by secondary translators.

[0058] After the secondary translator completes the first batch of review tasks, it automatically returns to step 3 in S4 to recalculate the classification results of the corpus classification task: if the classification results of the corpus that cannot be confirmed are still obtained after recalculation, the second batch of review tasks will continue to be automatically generated, and this cycle will be repeated until the classification results of all corpus tasks are confirmed;

[0059] S5. Corpus cleaning and editing tasks:

[0060] Based on the classification results obtained in S4, the translator of level 2 or above shall modify and improve the corpus data: 1) Determine whether the corpus should be edited based on the classification results of the reviewed corpus: if the corpus is confirmed to need to be modified, the corpus editing process shall be carried out; if the corpus is confirmed to not need to be modified, the corpus data cleaning is completed; if the corpus cannot be confirmed, the corpus shall be subject to the manual review process;

[0061] 2) After the corpus editing task is completed, it is determined whether the edited corpus needs to be classified again: if necessary, the edited corpus re-enters step S3 for corpus classification task cleaning; if not, the edited corpus undergoes a quality inspection process;

[0062] S6. Corpus after quality inspection and editing:

[0063] The corpus that does not need to be reclassified after editing will be manually quality-checked by a Level 2 translator or above: if the manual quality check passes, the corpus data is cleaned; if the manual quality check fails, return to S5 to re-do the corpus editing task and fill in the editing comments.

[0064] As a further step of this embodiment, the classification results of the corpus are reviewed in S4. If the classification results that cannot be confirmed by the corpus are not subject to manual review process, it is necessary to calculate the number of compensation times, specifically: add a task processing person, and recalculate the classification results of the corpus classification task after the task processing person configuration is completed. When the task processing person>the maximum number of classifications, the classification is ended.

[0065] As a further step of this embodiment, the cleaning parameters described in S2 include the number of corpus classifications, the number of classification task items, the configuration of corpus embedding data, whether to edit, manual review, manual quality inspection after editing, classification task category, whether the attribute needs to be modified, the modification attribute value, the classification confirmation difference threshold and the maximum number of classifications. The configuration of the cleaning parameters in this embodiment is as follows:

[0066]

[0067] As a further feature of this embodiment, the cleaning parameters in S2 also include task handlers, and the first-level translators, second-level translators, and translators above the second-level translators are distinguished by defining the levels of the task handlers.

[0068] The working principle of the present invention is as follows: the present invention sets two cleaning processes, namely, a corpus classification task and a corpus editing task, for cleaning micro-task corpus data. The cleaning of the corpus classification task is a task for classifying corpus, and the purpose is to screen out corpus that is not accurate enough and needs manual editing. Since no corpus modification is involved, the processing speed is fast. At the same time, a corpus will be classified by multiple people, and the system will automatically calculate the reliability of the classification result, so the ability of the translator is not required to be high. The cleaning of the corpus editing task is a task for editing the corpus. After the classification task classifies the corpus, the corpus that needs to be modified will generate a corpus editing task. After receiving the task, the translator needs to modify the corpus. After the modification is completed, the corpus classification task can be performed again or manual quality inspection can be performed until the corpus does not need to be modified and the corpus cleaning is completed.

[0069] The embodiments described above are merely descriptions of preferred implementation modes of the present invention and are not intended to limit the scope of the present invention. Without departing from the principles and essence of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the scope of protection determined by the claims of the present invention.

Claims

1. A micro-task corpus data cleaning method, It is characterized in that The specific steps include: S1. The corpus data to be cleaned is pre-embedded with the corpus data of known results to form corpus embedding data and then the cleaning begins; S2. Configure the cleaning parameters of the corpus data; S3. Cleaning corpus classification tasks: The system assigns corpus classification tasks to the first-level translator, who classifies and processes one or more corpus classification tasks, where each corpus classification task includes one or more task items, and each task item corresponds to a piece of corpus; S4. Calculate the classification results of the corpus classification task: After completing the corpus classification task, the first-level translator automatically calculates the classification results and obtains the corpus data that can be directly used and the corpus data that needs to be edited. The specific steps include the following: 1) Obtain corpus for which classification results can be confirmed: When a piece of corpus is processed and classified by multiple first-level translators, if the processing and classification results of all first-level translators are the same, then the corpus can be confirmed; When the corpus classification task is corpus buried data, the classification result of the corpus classification task processing can be confirmed; 2) Calculate the credibility of the first-level translator in handling the corpus classification task: Obtain all the corpus with known results from the corpus classification task in which the first-level translator participates, denoted as A; calculate all the correct classification results, denoted as C; let RE = the credibility of the first-level translator in the corpus classification task, then After the calculation is completed, the credibility of the corpus classification task processed this time is included in the historical credibility of the first-level translator; Let RE 1 ,RE 2 ,RE 3 ,...RE n is the historical credibility of the first-level translator. Excluding the highest historical credibility record and the lowest historical credibility of the first-level translator, let REA = the final credibility of the first-level translator = the average credibility, then 3) Confirm the classification results of the corpus classification task: A corpus is classified by multiple first-level translators. The attributes that need to be modified and the modified attribute values ​​can be obtained from the configuration cleaning parameters. Let TV = modified attribute value, TVP = attribute value obtained by each first-level translator, REA is the average credibility of the first-level translators calculated in step 2), then TVP = TV*REA, and then calculate the attribute values ​​obtained by each first-level translator when the attribute definition of whether to be modified is "yes" and "no" respectively; Let a = the sum of the attribute values ​​obtained by the first-level translator when the attribute definition of whether to modify is "yes", b = the sum of the attribute values ​​obtained by the first-level translator when the attribute definition of whether to modify is "no", and y = the attribute difference, then Calculate y and compare the attribute difference y with the classification confirmation difference threshold in the configuration cleaning parameters: if y> classification confirmation difference threshold, the classification result of the corpus is confirmed to need to be modified; if y≦classification confirmation difference threshold, the classification result of the corpus is confirmed to not need to be modified, and the rest of the classification results cannot be confirmed; 4) Review the classification results of the corpus: The system automatically extracts the most difficult-to-confirm classification results from the corpus and generates the first batch of review tasks, which are manually reviewed by secondary translators. After the secondary translator completes the first batch of review tasks, it automatically returns to step 3 in S4 to recalculate the classification results of the corpus classification task: if the classification results of the corpus that cannot be confirmed are still obtained after recalculation, the second batch of review tasks will continue to be automatically generated, and this cycle will be repeated until the classification results of all corpus tasks are confirmed; S5. Corpus cleaning and editing tasks: Based on the classification results obtained in S4, the second-level translator or above will modify and improve the corpus data: 1) Determine whether the corpus needs to be edited based on the classification results of the reviewed corpus: If the corpus is confirmed to need to be modified, the corpus editing process will be carried out; If the corpus is confirmed to be in no need of modification, the corpus data cleaning is completed; If the corpus cannot be confirmed, the corpus will undergo a manual review process; 2) After the corpus editing task is completed, it is determined whether the edited corpus needs to be classified again. If necessary, the edited corpus reenters step S3 for corpus classification task cleaning; If not necessary, the edited corpus will be subjected to a quality control process; S6. Corpus after quality inspection and editing: The corpus that does not need to be reclassified after editing will be manually quality-checked by a Level 2 translator or above: if the manual quality check passes, the corpus data is cleaned; if the manual quality check fails, return to S5 to re-do the corpus editing task and fill in the editing comments.

2. According to the microtask corpus data cleaning method of claim 1, It is characterized in that In S4, the classification results of the corpus are reviewed. If the classification results that cannot be confirmed by the corpus are not manually reviewed, the number of compensation times needs to be calculated. Specifically, one task processing person is added. After the task processing person configuration is completed, the classification results of the corpus classification task are recalculated. When the task processing person> the maximum number of classifications, the classification is terminated.

3. According to the microtask corpus data cleaning method of claim 1, It is characterized in that The cleaning parameters described in S2 include the number of corpus classifications, the number of classification task items, the configuration of corpus embedding data, whether to edit, manual review, manual quality inspection after editing, classification task category, whether the attribute needs to be modified, the modification attribute value, the classification confirmation difference threshold and the maximum number of classifications.

4. A microtask corpus data cleaning method according to claim 1, It is characterized in that The cleaning parameters in S2 also include task handlers, and the first-level translators, second-level translators, and translators above the second-level translators are distinguished by defining the levels of the task handlers.

Citation Information

Patent Citations

  • Text classification method and device, computer equipment and storage medium

    CN108334605A

  • Knowledge graph construction method and device, storage medium and computing equipment

    CN113849658A