A co - construction process method for rapid iterative upgrade of a machine translation engine

Through the online machine translation engine co-construction process, the problems of data redundancy and low manual processing efficiency during the machine translation engine iteration process are solved, and rapid iteration and data quality are guaranteed, and user participation and benefits are improved.

CN115470802BActive Publication Date: 2025-07-29IOL WUHAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211201805.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-07-29
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

During the iterative upgrade of existing machine translation engines, there is redundant and invalid data in corpus acquisition and processing, high cost, long cycle, low efficiency, and difficult to guarantee data quality.

Method used

The training process of the machine translation engine is online. By opening up co-construction permissions, users upload corpuses and perform pre-processing, cleaning, quality inspection and public comments online, and use blockchain to confirm rights to realize process and standardized data management.

Benefits of technology

It improves the efficiency of corpus acquisition, reduces manual processing costs, shortens the iteration cycle, ensures data quality, and provides users with a sense of participation and actual benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115470802B_ABST
    Figure CN115470802B_ABST
Patent Text Reader

Abstract

The present invention discloses a co - construction process method for the rapid iterative upgrade of a machine translation engine, which enables the acquisition, pre - processing, cleaning, quality inspection, training, and crowd evaluation of corpus data to be online, process - oriented, and standardized, and the data generated in each link can be traced with a record. The beneficial effects of the present invention are as follows: it solves the problems of a large amount of invalid redundant data, high manual processing costs, long cycles, and low efficiency encountered in the iterative process of the machine translation engine. Moreover, this co - construction mode of the engine can bring a sense of participation and actual benefits to users participating in the engine co - construction. At the same time, all qualified data will be permanently stored in the blockchain and bound to users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for co-building a process, in particular to a method for co-building a process for rapid iterative upgrading of a machine translation engine, and belongs to the technical field of machine translation. Background Art

[0002] The development of machine translation technology has always closely followed the development of disciplines such as computer technology, information theory, and linguistics. It is one of the important directions of artificial intelligence.

[0003] Among machine translation systems, one type is corpus-based. This uses large amounts of corpus data to train the machine translation engine, resulting in more accurate translation results. However, the sheer volume of data and the difficulty in controlling the degree of compatibility between the data and the machine translation engine can negatively impact the quality of the machine translation engine. Therefore, to train a high-quality machine translation engine, it may be helpful to consider data acquisition and processing.

[0004] Machine translation engines are at the core of machine translation. The engine's iteration and upgrade cycle, as well as the quality of the corpus data used for model training, significantly impact machine translation results. In existing technologies, traditional processes such as corpus acquisition, preprocessing, cleaning, quality inspection, training, and public review are largely based on offline manual processing, which can lead to problems such as incorrect cleaning, missed cleaning, incorrect sampling, and missed sampling. Furthermore, limited manual capacity hinders the rapid flow and connection between various processes, which also affects the iteration and upgrade speed of machine translation engines. Summary of the Invention

[0005] The purpose of the present invention is to provide a co-construction process method for rapid iteration and upgrading of machine translation engines in order to solve at least one of the above technical problems. This method makes the entire process of machine translation engine training online, thereby improving training efficiency, shortening iteration cycle and reducing production costs.

[0006] The present invention achieves the above-mentioned object through the following technical solution: a co-construction process method for rapid iterative upgrade of a machine translation engine, characterized in that the co-construction process method includes the following steps:

[0007] Step 1: Create a machine translation engine and set its industry and language direction to clarify the vertical attributes of the machine translation engine. The machine translation engine is open to co-construction permissions and published to the engine market for users to upload corpus terms;

[0008] Step 2: The user uploads the corpus to the machine translation engine with open co-construction permissions;

[0009] Step 3: The machine translation engine preprocesses the corpus uploaded by the user, and the status of the preprocessed corpus transitions to data to be cleaned.

[0010] Step 4: The creator of the machine translation engine grants corresponding permissions to the users participating in the co-construction. The users with the permission to clean the corpus receive the data to be cleaned in the form of task taking, and perform online cleaning through the corpus cleaning tool. The data to be cleaned that fails the cleaning is discarded, and the data to be cleaned that passes the cleaning has its status transition to data to be sampled.

[0011] Step 5: The users with the sampling permission manually judge the data to be sampled in the form of task taking. The machine translation engine determines the overall quality of the data to be sampled according to the manual judgment result in proportion. If it fails, the data to be sampled is discarded. If it passes, the data to be sampled is stored in the database, confirmed on the blockchain, and rewards are distributed to the co-constructing users. And the data to be sampled is converted into qualified data.

[0012] Step 6: The engine creator initiates a training application for the qualified data, which is reviewed by the system administrator. After the review passes, the qualified data is placed into the machine translation algorithm model for training, and the status of the qualified data transitions to training data. After the training data is completed, if the BLEU value meets the requirements, the status of the qualified data transitions to data to be peer-reviewed.

[0013] Step 7: The data to be peer-reviewed is realized through tasks. The translation teachers receive the peer-review tasks. After the data to be peer-reviewed passes, the engine creator initiates an online application for the training engine of the data to be peer-reviewed. After the system administrator's review passes, the training engine will be called in various machine translation scenarios.

[0014] As a further solution of the present invention: In the second step, the corpus is uploaded through a Word source text and translation document, an Excel corpus pair, or an associated online corpus. If the user uploads a Word source text and translation document, the machine translation engine will automatically perform corpus alignment. If the user uploads an Excel corpus, the corresponding corpus pair of the source text and translation must be uploaded. If the user uploads an associated online corpus, the user needs to create the corpus in advance in the machine translation engine.

[0015] As a further solution of the present invention: In the third step, the corpus preprocessing includes language identification, corpus alignment, duplicate removal, and abandonment actions.

[0016] As a further solution of the present invention: In the fourth step, the corresponding permissions include administrator permission, sampling permission, and cleaning permission.

[0017] As a further solution of the present invention: in the fifth step, the manual determination of the data to be sampled and inspected includes equidistant sampling logic sampling configured according to a ratio and online sampling through a corpus sampling tool.

[0018] The beneficial effects of the present invention are as follows:

[0019] 1) Online creation of a machine translation engine. The open mode of co-building the machine translation engine allows more online users to participate, increasing the channels for obtaining corpus and improving the acquisition efficiency;

[0020] 2) Users upload corpus files online, supporting the upload of multiple formats, especially Word source text and translation files. After the users upload, the machine translation engine performs corpus alignment, greatly saving the time for manual alignment;

[0021] 3) After the users upload the corpus, the machine translation engine filters the corpus through language identification, corpus alignment, and deduplication logic, etc., removing invalid corpus and greatly improving the efficiency of corpus preprocessing;

[0022] 4) Through the permission management mode of the engine co-building team, on the one hand, it is convenient for the team to carry out standardized management, and on the other hand, it ensures the quality and security of the corpus data to a certain extent; users clean the corpus through the online corpus cleaning tool. The main convenience brought is that it supports querying unqualified corpus according to conditions, quickly modifying them, and giving corresponding prompts to avoid users from mis-cleaning or missing cleaning. All data is saved in real time, and the cleaning results of each time are generated and stored online. The cleaning records are traceable, and users can receive co-building rewards, greatly improving the enthusiasm of users to participate;

[0023] 5) The sampling inspection process is online. The data for sampling inspection is the data after cleaning. The system supports users with sampling inspection permissions to reject the cleaning results or replace the person for cleaning. In this way, the quality of the data can be guaranteed to a certain extent. For data that is unqualified itself, it has the right to be discarded. The data that passes the sampling inspection will be stored in the database and confirmed on the blockchain. The data is bound to the corpus uploader, which can bring lasting benefits to the corpus uploader;

[0024] 6) The data training process is online. All the data that users can select when initiating a training application are the data that have passed the sampling inspection, which ensures the quality of the data used for engine training. And the review of the training application can view the data to be trained, which adds another guarantee to the data quality;

[0025] 7) The multi-review process is that multiple senior translation teachers conduct quality reviews on the results of engine training. Multiple reviews can make the review results more accurate and credible. After the multi-review passes, the engine creator can initiate an online application for the engine trained with this part of the corpus. After the system administrator approves, the engine containing the latest training results will be launched and will be called in various machine translation scenarios;

[0026] 8) The processes of obtaining, preprocessing, cleaning, quality inspection, training, and crowd evaluation of corpus data are made process-based and standardized, solving the problems of a large amount of invalid redundant data encountered in the iterative process of the machine translation engine, high manual processing costs, long cycles, and low efficiency. Moreover, this co-construction model can bring a sense of participation and actual benefits to users participating in the co-construction of the engine. At the same time, all qualified data will be permanently stored in the blockchain and bound to users. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a flowchart of the co-construction task of the engine of the present invention;

[0028] Figure 2 It is a diagram of uploading a corpus file in the third embodiment of the present invention;

[0029] Figure 3 It is a diagram of an online corpus cleaning tool in the third embodiment of the present invention;

[0030] Figure 4 It is a diagram of an online corpus sampling inspection tool in the third embodiment of the present invention;

[0031] Figure 5 It is a diagram of initiating a training application in the third embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0033] Embodiment 1

[0034] As Figure 1 shown, a co-construction process method for rapid iterative upgrade of a machine translation engine includes the following steps:

[0035] First: Create a machine translation engine, set the industry and language direction to which the machine translation engine belongs, so as to clarify the vertical attributes of the engine.

[0036] The newly created machine translation engine can automatically filter out corpus data that does not meet the requirements according to the language direction, making the quality of the corpus data more accurate. The open co-construction permission of the machine translation engine is released to the engine market, and users can upload corpus data in the engine market that supports co-construction, which can not only improve the engine iteration speed but also bring rewards to the engine co-builders.

[0037] Second: Users can upload corpora in the engines with open co-construction permissions. They can upload through Word source-translation documents, Excel corpus data pairs, or associated online corpora.

[0038] If users upload Word source-translation documents, the machine translation engine will automatically perform corpus alignment, which can save the time of manual alignment and improve efficiency.

[0039] If users upload Excel corpus data, they must upload the corresponding corpus pairs of the source and translation.

[0040] If users upload an associated online corpus, they need to create the corpus in the system in advance. The associated online corpus can be used not only for engine co-construction but also for associating with the corpus during webcat translation.

[0041] Third: The machine translation engine preprocesses the corpus data uploaded by users. The preprocessing includes language identification, corpus alignment, deduplication, and discarding actions.

[0042] The preprocessing is completely automatically completed by the machine translation engine without manual intervention. After identifying the language of the uploaded document in the machine translation engine, it is compared with the set language direction of the machine translation engine. If the comparison results are the same, corpus alignment will be performed on the Word file, and then the corpus that cannot be paired, duplicate corpus, and corpus with a large number of punctuation marks that cannot form sentences will be filtered and discarded. The corpus data that passes the preprocessing will be transferred to the data to be cleaned.

[0043] Fourth: The engine creator assigns corresponding permissions to the users participating in co-construction. The corresponding permissions include administrator permissions, sampling inspection permissions, and cleaning permissions.

[0044] The main functions of the corresponding permissions are to distinguish user identities and restrict user behaviors (such as downloading corpora). On the one hand, it facilitates the standardized management of the team, and on the other hand, it ensures the quality and security of the corpus data to a certain extent. At the same time, the permissions are associated with the overall process. Users with corpus cleaning permissions receive the data to be cleaned in the form of tasks and perform online cleaning through the corpus cleaning tool. The corpus that fails to pass the cleaning will be discarded, and the corpus data that passes the cleaning will be transferred to the data to be sampled.

[0045] The cleaning process is the most critical step in the iterative upgrade process of the machine translation engine because the attitude and results of the cleaners will directly affect the subsequent processes. If the cleaning fails, it will lead to an extended or even interrupted iterative cycle, affecting the corpus quality and thus the machine translation quality.

[0046] Fifth: Users with sampling inspection permissions receive the data to be sampled in the form of tasks. Sampling inspection is carried out online through the corpus sampling inspection tool according to the equidistant sampling logic configured by proportion.

[0047] The manually cleaned corpus is judged. During this process, the sampler can reject the sampled data for re-cleaning (this function is equivalent to supporting modification and resubmission). After the manual sampling passes, the system determines the overall quality of the current batch of data according to the results marked manually. If it is unqualified, the data will be discarded. If it is qualified, the data will be stored in the database, confirmed on the blockchain, and rewards will be given to the co-builders, and the data status will be changed to qualified data;

[0048] Sixth: The engine creator initiates a training application for the sampled and qualified corpus data, which is reviewed by the system administrator. After the review passes, the corpus data will be placed into the machine translation algorithm model for training, and the data status will be changed to data in training. After the training is completed, if the BLEU value meets the requirements (generally speaking, the result of machine translation is highly similar to the standard answer), the data status will be changed to data to be reviewed by the public.

[0049] Seventh: The public review is still achieved through tasks. Senior translation teachers receive public review tasks and evaluate the quality of the results of the engine's machine translation that applies the co-built corpus data. Multiple reviews can make the review results more accurate and credible. After the public review passes, the engine creator can initiate an online application for the engine trained by this part of the corpus. After the system administrator's review passes, the engine containing the latest training results will be launched and will be called in various machine translation scenarios.

[0050] Embodiment 2

[0051] As Figures 2 to 5 shown, a co-construction process method for rapid iterative upgrade of a machine translation engine includes the following steps:

[0052] 1. Translation company A created a vertical machine translation engine for the mechanical engineering industry in the Chinese-English direction and expected to iterate a version of the engine in the short term. Therefore, company A opened the engine co-construction permission;

[0053] 2. User Zhang has some corpus data in the Chinese-English direction of the mechanical engineering industry accumulated in his work. Zhang hopes to make the corpus in his hand more valuable. Zhang saw the engine created by company A on the Twinslator website. Therefore, Zhang uploaded the corpus data in his hand and completed the cleaning through the online corpus cleaning tool;

[0054] 3. Company A found that the corpus uploaded by Zhang has been cleaned, so it invited professional translators to join the engine team. The professional translators conducted an online sampling inspection on the corpus data cleaned by Zhang ( Figures 1-4 ), and Zhang's data quality is very high, and the sampling inspection passed smoothly;

[0055] 4. Company A initiated an engine training application with the qualified random inspection data provided by Xiao Zhang ( Figures 1-5 ). The platform party saw Company A's application through the background, checked the quality of the corpus data, and approved the training application. The machine translation model started training with this part of the corpus data;

[0056] 5. After the training results were transmitted back, Company A commissioned professional translators to review the training results. The review passed smoothly, and Company A initiated an online application for the engine;

[0057] 6. The platform party approved the online application initiated by Company A, and the machine translation engine for the mechanical engineering industry in the Chinese-English direction will be called in the machine translation scenario.

[0058] Working principle: When creating a machine translation engine, set the industry to which the machine translation engine belongs and the supported language direction and then publish and list it. After users participate in co-construction and upload corpus data, the system will preprocess the data (mainly including language identification, corpus alignment, duplicate removal, and discard operations). Invalid data will be automatically discarded. The qualified preprocessed data will then go through cleaning and random inspection and review processes to control the quality of the corpus data. High-quality data also needs to pass algorithm model training and manual evaluation and review. After passing the online application review, a latest version of the machine translation engine will appear and be called in various machine translation scenarios.

[0059] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed claims.

[0060] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A co - construction process method for rapid iterative upgrading of a machine translation engine, characterized in that, The co-construction process method includes the following steps: Step 1: Create a machine translation engine and set its industry and language direction to clarify the vertical attributes of the machine translation engine. The machine translation engine is open to co-construction permissions and published to the engine market for users to upload corpus terms; Step 2: The user uploads the corpus to the machine translation engine with open co-construction permissions; Step 3: The machine translation engine preprocesses the corpus uploaded by the user, and the preprocessed corpus is converted into data to be cleaned; Step 4: The creator of the machine translation engine grants corresponding permissions to users participating in the co-construction. Users with corpus cleaning permissions receive tasks to clean data and perform online cleaning using corpus cleaning tools. Data that fails the cleaning process is discarded, while data that passes the cleaning process is transferred to data for sampling inspection. Step 5: Users with sampling inspection authority manually judge the data to be inspected by accepting tasks. The machine translation engine determines the overall quality of the data to be inspected in proportion to the manual judgment results. If the data fails, the data to be inspected is discarded. If the data passes, the data to be inspected is stored and the ownership is confirmed on the blockchain. A reward is issued to the co-construction user, and the data to be inspected is converted to qualified data. Step 6: The engine creator initiates a training application for qualified data, which is reviewed by the system administrator. Once the review is passed, the qualified data will be placed in the machine translation algorithm model for training. The qualified data status will be transferred to training data. After the training data is completed, if the BLEU value meets the requirements, the qualified data status will be transferred to data for public review. Step 7: After the public review data is completed through tasks, the translator will receive the public review task. After the public review data is passed, the engine creator will initiate an online application for the training engine of the public review data. After the system administrator reviews and approves it, the training engine will be called in various machine translation scenarios.

2. The co-construction process method according to claim 1, wherein: In step 2, the corpus is uploaded via a Word source-translation document, an Excel corpus pair, or an associated online corpus. If the user uploads a Word source-translation document, the machine translation engine will automatically perform corpus alignment. If the user uploads an Excel corpus, the corpus pair corresponding to the source-translation must be uploaded. If the user uploads an associated online corpus, the user needs to create a corpus in the machine translation engine in advance.

3. The co-construction process method according to claim 1, wherein: In the step three, the corpus preprocessing includes language identification, corpus alignment, deduplication and discarding.

4. The co-construction process method according to claim 1, characterized in that: In step 4, the corresponding permissions include administrator permissions, sampling permissions, and cleaning permissions.

5. The co-construction process method according to claim 1, characterized in that: In the step 5, the manual determination of the data to be sampled includes logical sampling based on equidistant sampling after proportional configuration and online sampling using a corpus sampling tool.

Citation Information

Patent Citations

  • A machine translation engine recommendation method and device

    CN109710948A

  • Text processing method and device, electronic equipment and computer readable storage medium

    CN113011126A