A method, storage medium and electronic device for reducing the scale of a machine translation database
Through the method of classifying and determining the boundary value of knowledge, the machine translation database is reduced, and the problems of degradation in the existing technology of database quality and increased computing overhead are solved, and efficient database reduction and applicability are achieved.
Patent Information
- Application Number
- CN202210566109.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-05-23
AI Technical Summary
The prior art lacks interpretability when reducing the scale of machine-translated databases, resulting in a decline in database quality and an increase in computing overhead, and is unable to effectively remove redundant entries.
By building a database, the classification entries are types that are mastered and not mastered, and the knowledge boundary values in the local space are determined, the entries that meet the conditions are added to the candidate set, and the entries are randomly discarded according to the preset ratio to form the final reduced database.
While ensuring database quality, it significantly reduces storage space and computing costs, improves the interpretability and applicability of the database, and is suitable for different languages and fields.
Smart Images

Figure CN114970570B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and specifically to a method, a storage medium, and an electronic device for reducing the scale of a machine translation database. Background Art
[0002] Domain adaptation is an important topic in Neural Machine Translation (NMT). Its purpose is to adapt an NMT model in a general domain to a target domain so that it can handle translation tasks in the target domain.
[0003] Recently, the retrieval-based kNN-MT framework [1] has become a new paradigm for domain adaptation. Different from the traditional fine-tune [2] method, this framework can quickly complete domain adaptation without updating the parameters of the general domain NMT model, greatly alleviating the problem of catastrophic forgetting.
[0004] Specifically, the kNN-MT framework extracts translation knowledge from bilingual data in the target domain and constructs a database (datastore). Each entry in the database is a key-value pair, where the key is the hidden state of the context at the current position, and the value is the answer word that should be generated at the current position. By retrieving relevant translation knowledge from this target domain database, the general domain NMT model can be capable of handling translation tasks in the target domain. However, in the kNN-MT framework, all target language words in the bilingual data are saved into the database, which results in a large database size and may have a great deal of redundancy. In addition, storing a large-scale database occupies a large amount of disk space, and retrieving a large-scale database incurs a large amount of computational overhead. These are all problems brought about by the large-scale database.
[0005] In order to reduce the number of entries in the database, researchers have made some preliminary attempts. Currently, there are two methods for reducing the database scale:
[0006] The first one: randomly discard entries in the database. Experimental results show that randomly discarding database entries will cause a significant decline in translation performance. Random discarding is obviously not a feasible approach. This also indirectly shows that reducing the database scale is a very challenging problem.
[0007] Second type: For some entries with the same value, if the translation costs of some of these entries are the same, then there are redundant entries that can be discarded among these entries. Here, the translation cost is defined as the perplexity (PPL) of the general-domain NMT model for translating the value words.
[0008] Existing technical solutions for reducing the database size have a serious defect: they do not fully consider the capabilities of the general-domain NMT model. Intuitively, only when the translation effect of the general-domain NMT model is not good, is it necessary to retrieve the database and use the retrieved target-domain knowledge to assist in correcting the original output distribution of the general-domain NMT model. When the translation effect of the general-domain NMT model is good, the retrieved entries can no longer play any positive role. Obviously, the existing solutions do not fully consider the capabilities of the general-domain NMT model in each local area of the target-domain semantic space. This shortcoming will lead to the inability of the existing technical solutions for reducing the database size to effectively remove redundant entries in the database, and the reduction process lacks interpretability and cannot guarantee the quality of the database after reduction.
[0009] Specifically, the method in the first type randomly deletes entries in the database and does not consider the translation ability of the general-domain NMT model at all during the process of reducing the database size. This is the simplest reduction method, without any interpretability, and the database after reduction also seriously affects the translation effect. The method in the second type clusters the entries in the database through the translation cost, and then deletes a certain proportion of entries within each cluster. This approach also has risks: on the one hand, the translation cost cannot directly reflect the translation ability of the general-domain NMT model. Because whether the general-domain NMT model can accurately translate has no direct connection with the perplexity; on the other hand, the designed translation cost index of this method is too heuristic, and the process of deleting entries also lacks interpretability and cannot effectively ensure that the deleted entries are useless for the NMT model to process the target-domain translation task. Experimental results also show that after applying this method to reduce the database, the translation performance of the model in the target domain will significantly decline. Summary of the Invention
[0010] To overcome the shortcomings in the above background technology that the reduction process lacks interpretability and cannot guarantee the quality of the database after reduction, the purpose of the present invention is to provide a method, a storage medium, and an electronic device for reducing the size of a machine translation database.
[0011] To achieve the above purpose, in the first aspect of the present invention, a method for reducing the size of a machine translation database is provided, including the following steps:
[0012] S1: Construct a database and classify all entries according to the mastery of each entry in the database;
[0013] S2: Determine knowledge boundary values for different entries according to the distribution of entries in the local space;
[0014] S3: Analyze the types of each entry and the corresponding knowledge boundary values, and add the entries that meet the conditions to the candidate set;
[0015] S4: Randomly discard a certain number of entries from the candidate set according to a preset ratio to obtain the finally reduced database.
[0016] In some possible implementation manners, in S1, the "construct a database and classify all entries according to the mastery of each entry in the database" specifically includes the following steps:
[0017] S11: The specific process of constructing the database is as follows: Input the bilingual parallel sentence pair (x, y) into the general-domain NMT model. The NMT model will encode the source language sentence x and the sentence fragment y before the t-th word in the target language sentence <t as a whole into a hidden layer state h(x, y <t ) in the form of a high-dimensional vector, where h represents the mapping from text to a high-dimensional vector. In this way, the hidden layer state h(x, y <t ) and the t-th word y t of the target language sentence constitute a translation knowledge entry;
[0018] When the NMT model is in the hidden layer state h(x, y <t ), the answer word y t should be generated. Save the hidden layer state and the answer word at each position in the parallel corpus in the form of key-value pairs to complete the construction of the database;
[0019] S12: Judge whether y t and are consistent, and classify the knowledge entries in the database into two categories: the mastered entries known and the unmastered entries unknown, specifically as follows:
[0020]
[0021] In some possible implementation manners, in order to judge the mastery of the database entries by the NMT model, during the construction of the database, in addition to retaining the key-value pair: the hidden layer state h(x, y <t ) and the answer word y t , the predicted word predicted by the model at this position is also retained
[0022] In some possible implementation manners, in S2, the step of "determining knowledge boundary values for different entries according to the distribution status of the entries in the local space" specifically includes the following steps:
[0023] S21: Determine the local space, where the local space N k is the k-nearest neighbor of the entry (key, val) in the database, and is specifically represented as:
[0024]
[0025] S22: Determine knowledge boundary values for different entries according to the distribution status of the entries in the local space. The knowledge boundary values are specifically represented as:
[0026]
[0027] In some possible implementation manners, the "condition" is that the entry in the database belongs to known and its km value is greater than the threshold k p and add the entry into a candidate set.
[0028] In some possible implementation manners, the "predetermined ratio" is: pruning ratio.
[0029] In a second aspect of the present invention, there is provided a computer-readable storage medium for storing program codes for executing the method for reducing the scale of a machine translation database as described above.
[0030] In a third aspect of the present invention, there is provided an electronic device, which includes a processor and a memory: the memory is used for storing program codes and transmitting the program codes to the processor; the processor is used for executing the method for reducing the scale of a machine translation database as described above according to the instructions in the program codes.
[0031] The beneficial effects of the present invention are as follows:
[0032] Technically speaking, this method designs two concepts, namely the entry type and the knowledge boundary, to describe the ability of the general-domain NMT model in the local semantic space of the target domain; starting from the perspective of the general-domain NMT ability, entries in the database are discarded based on local accuracy; while reducing the scale of the database as much as possible, the quality of the reduced database is also ensured; it has stronger interpretability and better demonstrates the cooperation relationship between the general-domain NMT model and the database within the kNN-MT framework; the storage space occupied by the reduced database is significantly reduced, and the computational cost brought is also reduced, which makes the kNN-MT framework more practical.
[0033] From the application level, this method is easy to reproduce, has a simple process, and does not rely on any additional complex neural network structures; it has high compatibility, and the reduced database can be used in any kNN-MT framework; it has a wide range of adaptability and can reduce the scale of databases in different languages and different fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is the overall step flowchart of a method for reducing the scale of a machine translation database according to an embodiment of the present invention;
[0035] Figure 2 It is the specific example step flowchart of a method for reducing the scale of a machine translation database according to an embodiment of the present invention;
[0036] Figure 3 It is the schematic diagram of constructing a database according to parallel sentence pairs according to an embodiment of the present invention;
[0037] Figure 4 It is the schematic diagram of analyzing the types of database entries and the knowledge boundary values according to an embodiment of the present invention;
[0038] Figure 5 It is the schematic diagram of setting the pruning ratio according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The following elaborates on the preferred embodiments of the present invention in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0040] This embodiment provides a method for reducing the scale of a machine translation database. Referring to the attached Figure 1 as shown, it includes the following steps:
[0041] S1: Construct a database, and classify all entries according to the mastery of each entry in the database; S1 specifically includes the following steps:
[0042] S11: Constructing a database is to convert the discrete translation knowledge in the parallel corpus into continuous translation knowledge that can be utilized by the NMT model. The specific construction process of the database is as follows: Input the bilingual parallel sentence pair (x, y) into the general domain NMT model, and the NMT model will encode the source language sentence x and the sentence fragment y before the t-th word in the target language sentence <t as a whole into a hidden layer state h(x, y <t ) in the form of a high-dimensional vector, where h represents the mapping from text to a high-dimensional vector. In this way, the hidden layer state h(x, y <t ) and the t-th word y t of the target language sentence constitute a translation knowledge entry;
[0043] When the NMT model is in the hidden layer state h(x, y <t ), the answer word y should be generated. t By saving the hidden layer states and answer words at each position in the parallel corpus in the form of key-value pairs, the construction of the database can be completed.
[0044] To judge the NMT model's mastery of the database entries, during the construction of the database, in addition to retaining the key-value pair: hidden layer state h(x, y <t ) and answer word y t , the predicted word predicted by the model at this position is also retained.
[0045] Intuitively speaking, if the predicted word predicted by the general domain NMT model at the hidden layer state position is the same as the answer word, it means that the general domain NMT model has mastered this piece of knowledge. The database constructed using the general domain NMT model can be directly combined with the general domain NMT model and can quickly enhance the translation performance of the general domain NMT model in the target domain.
[0046] S12: Judge whether y t and are the same, and classify the knowledge entries in the database into two categories: mastered entries known and unmastered entries unknown, specifically as follows:
[0047]
[0048] This judgment method is simple and convenient and also has strong interpretability.
[0049] S2: Determine the knowledge boundary values for different entries according to the distribution of entries in the local space; S2 specifically includes the following steps:
[0050] S21: Determine the local space, where the local space N k is the k-nearest neighbor of the entries (key, val) in the database, specifically expressed as:
[0051]
[0052] S22: Determine the knowledge boundary values for different entries according to the distribution of entries in the local space, and the knowledge boundary values are specifically expressed as:
[0053]
[0054] S3: Analyze the type of each entry and the corresponding knowledge boundary value, and add the entries that meet the conditions to the candidate set. The "condition" is: the entry in the database is both known and its km value is greater than the threshold k p The entries are added into a candidate set. The form of this condition is very simple and very convenient in practical application.
[0055] S4: randomly discard certain entries from the candidate set according to a preset ratio to obtain a final reduced database. The "preset ratio" is: pruning ratio, which is set to determine the number of discarded entries. The pruning ratio multiplied by the original database size is the number of discarded entries.
[0056] For specific examples, see the attached Figure 2 As shown, a method for reducing the size of a machine translation database comprises the following steps:
[0057] S1: Building a database, classifying all items in the database by understanding the status of each item in the database, specifically including the following steps:
[0058] S11: The purpose of building a database is to transform the discrete translation knowledge in the parallel corpus into continuous translation knowledge that can be used by the NMT model. The specific construction process of the database is as follows: Take the Chinese-English parallel sentence pair ("His solution has flaws") as an example, and input it into the general domain NMT model. The NMT model will encode the source language sentence "His solution has flaws" and the sentence fragment "His solution" before the tth word (taking t=3 as an example) in the target language sentence as a whole into a hidden state h ("His solution has flaws", "His solution") in the form of a high-dimensional vector, where h represents the mapping from text to high-dimensional vector. In this way, the hidden state h ("His solution has flaws", "His solution") and the third word "has" of the target language sentence constitute a translation knowledge entry (h ("His solution has flaws", "Hissolution"), "has").
[0059] Refer to the attached Figure 3 As shown, when the NMT model is in the hidden state h(x,y <t ) should generate the answer word y t , save the hidden state and answer words of each position in the parallel corpus in the form of key-value pairs, and the database construction can be completed. In order to judge the mastery of the database entries by the NMT model, in the process of building the database, in addition to retaining the key-value pairs: such as the hidden state
[0060] h(“His solution has flaws”, “His solution”) and the answer word “has”, and also retain the predicted word, such as “have”, predicted by the model at this position.
[0061] Intuitively speaking, if the predicted word by the general-domain NMT model at the hidden layer state position is the same as the answer word, it indicates that the general-domain NMT model has mastered this piece of knowledge. The database constructed using the general-domain NMT model can be directly combined with the general-domain NMT model and can quickly enhance the translation performance of the general-domain NMT model in the target domain.
[0062] S12: Judge y t and Whether they are the same, classify the knowledge entries in the database into two categories: mastered entries known and unmastered entries unknown, as follows:
[0063]
[0064] This judgment method is simple and convenient, and also has strong interpretability. Refer to the appendix Figure 4 As shown, for example, for the entry (h(“His solution has flaws”, “His solution”), “has”), if the predicted word generated by the NMT model at this hidden layer state is “have”, which is inconsistent with the answer word “has”, then this entry is an unmastered entry for the NMT model.
[0065] S2: According to the distribution of entries in the local space, determine the knowledge boundary value (knowledge margin, km) for different entries, specifically including the following steps:
[0066] S21: Determine the local space, where the local space is the k-nearest neighbors of the entries (key, val) in the database, specifically expressed as:
[0067]
[0068] S22: According to the distribution of entries in the above local space, determine the knowledge boundary value for different entries, and the knowledge boundary value is specifically expressed as:
[0069]
[0070] Refer to the appendix Figure 4 As shown, for example, for the entry
[0071] For (h(“His solution has flaws”, “His solution”), “has”), if the first 4 entries among its 8 nearest neighbor entries are all entries that have been mastered, while the 5th entry is an entry that has not been mastered, then the knowledge boundary value of this entry is 4). Designing the knowledge boundary value is to characterize the generalization ability of the general-domain NMT model in the local space (neighborhood) of each entry.
[0072] S3: Analyze the types of each entry and the corresponding knowledge boundary values, and add the entries that meet the conditions to the candidate set. The conditions are: the entries in the database belong to the mastered entries known, and their km values are greater than the threshold k p Add the entries that meet this condition to a candidate set. The form of this condition is very simple and very convenient in actual applications. Refer to the appendix Figure 4 As shown in the figure, for example, for the entry (h(“His solution has flaws”, “His”), “solution”), it is both an entry mastered by the NMT model and its km value is greater than the threshold. Then this entry will be added to the candidate set.
[0073] Traverse the entries in the database. When a certain entry in the database belongs to the mastered entries known and its km value is very large, the local accuracy of the general-domain NMT model on this entry is very high. This means that in the local space where this entry is located, the general-domain NMT model itself can output the correct translation. Therefore, it is not necessary to save the entries located in these local spaces.
[0074] S4: Randomly discard a certain number of entries from the candidate set according to a preset ratio to obtain the finally reduced database. The “preset ratio” is: the pruning ratio. Setting the pruning ratio is to determine the number of discarded entries. The number of discarded entries is the pruning ratio multiplied by the scale of the original database. Refer to the appendix Figure 5 As shown in the figure, for example, if the original database has 4 entries and the pruning ratio is set to 75%, 3 entries will be deleted, and the scale of the reduced database becomes 1 entry.
[0075] This embodiment also provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the method for reducing the scale of the machine translation database described above.
[0076] The storage medium stores program instructions capable of implementing all the above methods. Among them, the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.
[0077] This embodiment further provides an electronic device, which includes a processor and a memory: the memory is used to store program codes and transmit the program codes to the processor; the processor is used to execute the above method for reducing the scale of a machine translation database according to the instructions in the program codes.
[0078] The processor can also be referred to as a CPU (Central Processing Unit). The processor may be an integrated circuit chip with the ability to process signals. The processor can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0079] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for reducing the scale of a machine translation database, characterized in that, It includes the following steps: S1: Construct a database and classify all entries based on the mastery of each entry in the database; S2: Determine knowledge boundary values for different entries according to the distribution of entries in the local space; S3: Analyze the types of each entry and the corresponding knowledge boundary values, and add the eligible entries to the candidate set; S4: Randomly discard a certain number of entries from the candidate set according to a preset ratio to obtain the finally reduced database; In S1, the construction of the database and the classification of all entries based on the mastery of each entry in the database specifically include the following steps: S11: The specific construction process of the database is as follows: Input the bilingual parallel sentence pair (x, y) into the general-domain NMT model. The NMT model will encode the source language sentence x and the sentence fragment y before the t-th word in the target language sentence <t as a whole into a hidden layer state h(x, y <t ) in the form of a high-dimensional vector, where h represents the mapping from the text to the high-dimensional vector. In this way, the hidden layer state h(x, y <t ) and the t-th word y t of the target language sentence form an entry; When the NMT model is in the hidden layer state h(x, y <t ), it should generate the answer word y t , and save the hidden layer states and answer words at each position in the parallel corpus in the form of key-value pairs to complete the construction of the database; S12: Determine whether y t and are consistent, and classify the knowledge entries in the database into two categories: mastered entries known and unmastered entries unknown, as follows: In S2, the determination of knowledge boundary values for different entries according to the distribution of entries in the local space specifically includes the following steps: S21: Determine a local space, where the local space N k is the k-nearest neighbor of the entry (key, val) in the database, specifically represented as: S22: Determine knowledge boundary values for different entries according to the distribution of entries in the local space, and the knowledge boundary values are specifically expressed as:
2. The method for reducing the scale of a machine translation database according to claim 1, wherein To determine the NMT model's mastery of database entries, during the process of constructing the said database, in addition to retaining the key-value pairs: hidden layer state h(x,y <t ) and answer word y t , the predicted word predicted by the model at this position is also retained 3. A method for reducing the scale of a machine translation database according to claim 1, characterized in that, The conditions are as follows: the entries in the database belong to known and their km values are greater than the threshold k p The entries are added to a candidate set.
4. A method for reducing the scale of a machine translation database according to claim 1, 2 or 3, characterized in that The preset ratio is: the pruning ratio.
5. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program codes, and the program codes are used to execute the method for reducing the scale of a machine translation database according to any one of claims 1-4.
6. An electronic device, characterized in that, The electronic device includes a processor and a memory: The memory is used to store program codes and transmit the program codes to the processor; the processor is used to execute the method for reducing the scale of a machine translation database according to the instructions in the program codes according to any one of claims 1-4.
Citation Information
Patent Citations
Data storage method, querying method and device
CN101968806A
Man-machine collaborative construction method for domain term semantic knowledge base
CN110765781A