Corpus Classification Method, Vertical Industry Machine Translation Method and Device
By using the target parallel corpus to calculate sentence vector similarity, classifying corpus and training in the vertical industry machine translation model, the problem of poor performance of machine translation model in the vertical industry is solved, and simple and efficient corpus classification and translation are achieved.
Patent Information
- Application Number
- CN202210089423.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-01-25
AI Technical Summary
In the prior art, the vertical industry machine translation model has poor training performance and poor interpretability of classification results, requiring a large number of manual markers and neural network screening corpus.
By obtaining original word segmentation and translated word segmentation based on the target parallel corpus, calculating sentence vector similarity, and classifying the corpus into specific types when the similarity reaches the threshold, using expected data from other industries for training to improve model performance.
It realizes simple corpus classification and machine translation, improves model performance and interpretability, and simplifies the training process.
Smart Images

Figure CN114461799B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a corpus classification method, a vertical industry machine translation method and device. Background Art
[0002] With the rapid development of the fields of artificial intelligence and machine learning, machine translation technology has grown rapidly, being able to meet the translation needs of high timeliness requirements or massive texts in various industries, greatly reducing the labor cost.
[0003] A vertical industry refers to a small field vertically segmented under a comprehensive industry field. In the machine translation model adapted to the practical field, the quality of the corpus greatly affects the translation quality of the model. Usually, it is necessary to manually label or combine neural networks to screen out the corpus belonging to this vertical industry, and then perform fine-tuning to generate the model. This method requires using neural networks to classify a large number of corpus samples, and the interpretability of the classification results is poor, and the finally trained model is often not the optimal solution. Summary of the Invention
[0004] The present invention provides a corpus classification method, a vertical industry machine translation method and device, to solve the defect that the performance of the model trained simply using vertical industry corpus data in the prior art is poor, and to achieve classifying other industry corpus data, mixing the corresponding type of corpus data into the training set for training, and improving the model performance.
[0005] The present invention provides a corpus classification method, including:
[0006] Based on a target parallel corpus, obtaining the original text segmentation and the translated text segmentation of each target corpus;
[0007] Based on the original text segmentation and the translated text segmentation, obtaining a first original text sentence vector and a first translated text sentence vector;
[0008] Embedding the first original text sentence vector and the first translated text sentence vector respectively to obtain a second original text sentence vector and the second translated text sentence vector;
[0009] Based on the first original text sentence vector, the first translated text sentence vector, the second original text sentence vector and the second translated text sentence vector, calculating to obtain a target similarity;
[0010] When the target similarity is greater than or equal to a target threshold, the type of the target corpus is a first target type;
[0011] Among them, the target threshold is set based on the target parallel corpus. The number of the target parallel corpora can be one or more, and the target corpus is the text data in the target parallel corpus.
[0012] According to a corpus classification method provided by the present invention, after calculating the target similarity, it further includes: when the target similarity is less than the target threshold, the type of the target corpus is the second target type;
[0013] Among them, the first target type and the second target type are mutually exclusive.
[0014] According to a corpus classification method provided by the present invention, calculating the target similarity based on the first source sentence vector, the first translation sentence vector, the second source sentence vector, and the second translation sentence vector includes:
[0015] Obtaining a first similarity based on the first source sentence vector and the second translation sentence vector;
[0016] Obtaining a second similarity based on the first translation sentence vector and the second source sentence vector;
[0017] Performing weighted summation based on the first similarity and the second similarity to obtain the target similarity.
[0018] According to a corpus classification method provided by the present invention, obtaining the first source sentence vector and the first translation sentence vector based on the source word segmentation and the translation word segmentation includes:
[0019] Embedding the source word segmentation and the translation word segmentation respectively to obtain source word vectors and translation word vectors;
[0020] Summing the source word vectors and the translation word vectors respectively to obtain the first source sentence vector and the first translation sentence vector.
[0021] The present invention also provides a vertical industry machine translation method, including:
[0022] Obtaining a source language text to be translated in a target vertical industry;
[0023] Inputting the source language text to be translated into a pre-established target translation model to obtain a target language text corresponding to the source language text to be translated;
[0024] Among them, the target translation model uses the initial training corpus pair of the target vertical industry, and the enhanced training corpus pair with the target type obtained from the target parallel corpus by using any one of the above-mentioned corpus classification methods. The target type matches the target vertical industry, and the enhanced training corpus pair has a target proportional relationship with the initial training corpus pair.
[0025] The present invention also provides a corpus classification device, including:
[0026] A word segmentation module, configured to obtain the original text word segmentation and the translated text word segmentation of each target corpus based on the target parallel corpus;
[0027] A first acquisition module, configured to obtain a first original text sentence vector and a first translated text sentence vector based on the original text word segmentation and the translated text word segmentation;
[0028] A second acquisition module, configured to respectively embed the first original text sentence vector and the first translated text sentence vector to obtain a second original text sentence vector and the second translated text sentence vector;
[0029] A similarity calculation module, configured to calculate a target similarity based on the first original text sentence vector, the first translated text sentence vector, the second original text sentence vector, and the second translated text sentence vector;
[0030] A first classification module, configured to determine that the type of the target corpus is the first target type when the target similarity is greater than or equal to a target threshold;
[0031] Among them, the target threshold is set based on the target parallel corpus. The number of the target parallel corpora can be one or more, and the target corpus is the text data in the target parallel corpus.
[0032] The present invention also provides a vertical industry machine translation device, including:
[0033] A source text acquisition module, configured to acquire the source language text to be translated in the target vertical industry;
[0034] A translation module, configured to input the source language text to be translated into a pre-established target translation model to obtain the target language text corresponding to the source language text to be translated;
[0035] Among them, the target translation model uses the initial training corpus pair of the target vertical industry, and the enhanced training corpus pair with the target type obtained from the target parallel corpus by using any one of the above-mentioned corpus classification methods. The target type matches the target vertical industry, and the enhanced training corpus pair has a target proportional relationship with the initial training corpus pair.
[0036] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above-mentioned corpus classification methods and the steps of any one of the above-mentioned vertical industry machine translation methods are implemented.
[0037] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned corpus classification methods and the steps of any one of the above-mentioned vertical industry machine translation methods are implemented.
[0038] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned corpus classification methods and the steps of any one of the above-mentioned vertical industry machine translation methods are implemented.
[0039] The corpus classification method, vertical industry machine translation method and device provided by the present invention respectively perform simple combinations on the original text segmentation and translated text segmentation of the target corpus to form a first original text sentence vector and a first translated text sentence vector, obtain a second original text sentence vector and a second translated text sentence vector through sentence embedding, cross-combine the sentence vectors obtained by different methods, calculate the target similarity between the vectors, and classify the target corpus as the first target type when the target similarity is greater than or equal to the target threshold. It can rely on simple vector calculations to realize the classification of different corpora, with simple operation and high interpretability. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 is a schematic flowchart of the corpus classification method provided by the present invention;
[0042] Figure 2 is a schematic flowchart of the vertical industry machine translation method provided by the present invention;
[0043] Figure 3 is a schematic structural diagram of the corpus classification device provided by the present invention;
[0044] Figure 4 is a schematic structural diagram of the vertical industry machine translation device provided by the present invention;
[0045] Figure 5It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0046] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without any creative efforts shall fall within the protection scope of the present invention.
[0047] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or more.
[0048] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification and claims of the present invention, unless otherwise clearly specified in the context, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0049] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0050] Figure 1 It is a schematic flowchart of the corpus classification method provided by the present invention. As Figure 1 shown, the corpus classification method provided by the embodiments of the present invention includes: Step 101, based on a target parallel corpus, obtain the original text segmentation and the translated text segmentation of each target corpus.
[0051] Among them, the number of target parallel corpora can be one or more, and the target corpus is the text data in the target parallel corpus.
[0052] It should be noted that the execution subject of the corpus classification method provided by the embodiments of the present invention is a corpus classification device.
[0053] The application scenario of the corpus classification method provided by the embodiments of the present invention is to determine the training set before training a machine translation model in a certain field.
[0054] In the prior art, its training set usually consists of sample source data and sample target data in this field.
[0055] In the embodiments of the present invention, its training set includes not only the sample source data and sample target data in this field, but also the sample source data and sample target data in other related fields, and there are certain commonalities in the translation styles between this field and its related fields.
[0056] Therefore, the corpus classification method provided by the embodiments of the present invention is applicable to classifying a large amount of sample data in the whole field, screening out the sample data in other related fields that conform to the translation style of this field, and adding them to the training set.
[0057] It should be noted that before step 101, the corpus classification device needs to determine the target parallel corpus according to the actual training task of the machine translation model.
[0058] The target parallel corpus is the operation object of the corpus classification device and includes two monolingual corpora, and one corpus is the translation of the other corpus. Among them, the corresponding segments (usually sentences or paragraphs) need to match, and the fields involved are different from the field corresponding to the machine translation model.
[0059] The embodiments of the present invention do not make a specific limitation on the number of target parallel corpora.
[0060] Optionally, in the case where the actual model training task is that the source text and the target text have a one-to-one relationship in terms of language or translation content, selecting one target parallel corpus can represent this relationship.
[0061] Optionally, in the case where the actual model training task is that the source text and the target text have a one-to-n relationship in terms of language or translation content, selecting n target parallel corpora can represent the relationship between the source text and any one of the target texts. Among them, n is a positive integer greater than 1.
[0062] The target corpus refers to the text pair written in different languages in the target parallel corpus and having a "translation relationship" with each other, that is, it includes the source text and the target text.
[0063] Specifically, in step 101, the corpus classification device performs word segmentation on the source text and the target text included in each target corpus in the target parallel corpus respectively to obtain the corresponding source text word segmentation and target text word segmentation.
[0064] Among them, the corresponding word segmentation method can be determined in advance according to the languages of the source text and the target text or the translation task of the target text, and the embodiments of the present invention do not make a specific limitation on this.
[0065] Exemplarily, for Chinese text, the tokenizers adopted include but are not limited to IK tokenizer, jieba tokenizer, THULAC tokenizer, or the tokenizer open-sourced by Peking University.
[0066] For English text, the tokenizers adopted include but are not limited to the English tokenizer in ElasticSearch, the multi-lingual tokenizer toolkit built into Lucene, or the Natural Language Toolkit (NLTK).
[0067] Step 102: Based on the original text tokenization and the translated text tokenization, obtain the first original sentence vector and the first translated sentence vector.
[0068] Specifically, in Step 102, the corpus classification device performs vector representation on the original text tokenization obtained in Step 101 and then combines and adds them to obtain the first original sentence vector. At the same time, it also performs vector representation on the translated text tokenization obtained in Step 101 and then combines and adds them to obtain the first translated sentence vector.
[0069] The first original sentence vector refers to the sum of the original text tokenization vectors. The first original sentence vector is used as the text representation of a sentence simply composed of the original text tokenization and can be represented by V S representation.
[0070] The first translated sentence vector refers to the sum of the translated text tokenization vectors. The first translated sentence vector is used as the text representation of a sentence simply composed of the translated text tokenization and can be represented by V T representation.
[0071] Step 103: Embed the first original sentence vector and the first translated sentence vector respectively to obtain the second original sentence vector and the second translated sentence vector.
[0072] Specifically, in Step 103, the corpus classification device uses the sentence embedding algorithm on the first original sentence vector and the first translated sentence vector obtained in Step 102 respectively to obtain the second original sentence vector and the second translated sentence vector.
[0073] The second original sentence vector refers to the first original sentence vector and its semantic information in the corresponding paragraph. The second original sentence vector is used as the text representation of the original sentence group containing the semantic information of the paragraph context and can be represented by V SS representation. V SS Compared with V S , it has richer semantic information.
[0074] The second translated sentence vector refers to the first translated sentence vector and its semantic information in the corresponding paragraph. The second translated sentence vector is used as the text representation of the translated sentence group containing the semantic information of the paragraph context and can be represented by VST Characterization. V ST Compared with V T , it has richer semantic information.
[0075] The embodiments of the present invention do not specifically limit the sentence embedding algorithm.
[0076] Exemplarily, the sentence embeddings used by the corpus classification device include but are not limited to technologies such as Doc2Vec, SentenceBERT, InferSent, and Universal Sentence Encoder.
[0077] Preferably, the corpus classification device uses a multi-sentence expression model to embed the first original sentence vector and the first translated sentence vector.
[0078] Among them, the multi-sentence expression model includes but is not limited to the Baidu multilingual model ERNIE-M, the LASER (Language-Agnostic Sentence Representations) toolkit extended and enhanced by Facebook researchers, etc.
[0079] Step 104: Calculate the target similarity based on the first original sentence vector, the first translated sentence vector, the second original sentence vector, and the second translated sentence vector.
[0080] Specifically, in step 104, the corpus classification device cross-combines the two types of sentence vectors generated in step 102 and step 103 in the way that one type of original sentence vector corresponds to the other type of translated sentence vector, calculates the similarity corresponding to each combination, and then merges them to obtain the target similarity.
[0081] The embodiments of the present invention do not specifically limit the similarity calculation method between vectors.
[0082] Optionally, the distance between two vectors can be calculated. The closer the distance between two vectors, the more similar the two vectors are. The distance representation methods include but are not limited to Euclidean distance, Manhattan distance, Chebyshev distance, Minkowski distance, standardized Euclidean distance, Mahalanobis distance, or Lance distance, etc. The corresponding target similarity value range can be normalized to [0,1].
[0083] Optionally, the cosine of the included angle in geometry can be used to measure the difference in the directions of two vectors. The specific implementation methods include but are not limited to the cosine of the included angle or the Tanimoto coefficient, and the corresponding target similarity value range is [0,1].
[0084] Step 105: When the target similarity is greater than or equal to the target threshold, the type of the target corpus is the first target type.
[0085] Among them, the target threshold is set based on the target parallel corpus.
[0086] It should be noted that the target threshold refers to the similarity threshold corresponding to the target parallel corpus. The target threshold is used to compare with the target similarity of each target corpus, and from a more fundamental perspective, all target corpora in the target parallel corpus are divided into "generalized engineering" category and "generalized literature" category.
[0087] Among them, the "generalized engineering" category is characterized by giving priority to the accuracy of translation during translation, requiring strict and correct expression of the original meaning, and having a preference for "literal translation" in terms of translation style. For example, international engineering, automobile manufacturing, etc.
[0088] The "generalized literature" category is characterized by giving priority to readability and fluency of expression during translation, not requiring strict expression of the original meaning, and even freely re-creating according to the context in some cases, and having a preference for "free translation" in terms of translation style. For example, movie dialogues, online novels, etc.
[0089] The embodiments of the present invention do not specifically limit the value of the target threshold. Exemplarily, for different target parallel corpora, the corresponding target threshold can be any value in [0.5, 0.8].
[0090] Specifically, in step 105, the corpus classification device compares the target similarity of each target corpus with the corresponding target threshold. There are two comparison results: meeting the target threshold and not meeting the target threshold.
[0091] Among them, meeting the target threshold means that the target similarity is greater than or equal to the target threshold, indicating that the similarity between the word and sentence vectors and the embedded sentence vector is relatively high, that is, the sentence can be interpreted literally, and it is determined that the target corpus is more inclined to the "generalized engineering" category and belongs to the first target type.
[0092] Not meeting the target threshold means that the target similarity is less than the target threshold, indicating that the similarity between the word and sentence vectors and the embedded sentence vector is relatively low, that is, the sentence cannot be interpreted literally, and it is determined that the target corpus does not belong to the first target type.
[0093] In the prior art, it is usually necessary to establish corresponding samples and corresponding labels for different parallel corpora to train a classification model, and directly classify the corpora in the model parallel corpus during actual application, that is, every time the parallel corpus is changed, re-training is required, and the process is cumbersome and the interpretability is poor.
[0094] In an embodiment of the present invention, a first source sentence vector and a first target sentence vector are respectively formed by simply combining the word segmentation of the source text and the word segmentation of the target text based on the target corpus. The second source sentence vector and the second target sentence vector are obtained through sentence embedding. By cross-combining the sentence vectors obtained in different ways, the target similarity between the vectors is calculated, and when the target similarity is greater than or equal to the target threshold, the target corpus is classified into the first target type. It can rely on simple vector calculations to achieve the classification of different corpora, with simple operation and high interpretability.
[0095] Based on any of the above embodiments, after calculating the target similarity, it further includes: when the target similarity is less than the target threshold, the type of the target corpus is the second target type.
[0096] Wherein, the first target type and the second target type are mutually exclusive.
[0097] Specifically, after step 104, the corpus classification device compares the target similarity of each target corpus with the corresponding target threshold. When the comparison result does not meet the target threshold, it indicates that the similarity between the word and sentence vectors and the embedded sentence vectors is low, and the sentence cannot be directly interpreted literally. Then it is determined that the target corpus is more inclined to the "general literature" category, belonging to the second target type that is completely mutually exclusive with the first target type.
[0098] It can be understood that when the target similarity is less than the target threshold and belongs to a certain preset range interval, it indicates that the similarity between the word and sentence vectors and the embedded sentence vectors is extremely low. Then it is determined that the target corpus has no reference value, that is, the data is discarded. The embodiment of the present invention does not specifically limit the preset range interval. Exemplarily, for different target parallel corpora, the preset range interval can be generalized to [0, 0.3].
[0099] The embodiment of the present invention classifies the target corpus into the second target type when the target similarity is less than the target threshold. It can rely on simple vector calculations to achieve the classification of different corpora, with simple operation and high interpretability.
[0100] Based on any of the above embodiments, based on the first source sentence vector, the first target sentence vector, the second source sentence vector, and the second target sentence vector, calculating the target similarity includes: obtaining a first similarity based on the first source sentence vector and the second target sentence vector.
[0101] Specifically, in step 104, the corpus classification device combines the first source sentence vector and the second target sentence vector as a corpus pair, and uses the similarity calculation method between vectors to obtain the first similarity.
[0102] Preferably, the corpus classification device will calculate the first source sentence vector and the second translated sentence vector using the Euclidean dot product formula, and the calculation formula for the obtained first similarity is as follows:
[0103]
[0104] where SIM1 is the first similarity, V S is the first source sentence vector, V ST is the second translated sentence vector, and θ1 is the angle between V S and V ST When SIM1 is 1, θ1 is 0°, indicating that the similarity between vector V S and V ST is 100% (i.e., exactly the same). Conversely, when SIM1 is 0, θ1 is 90°, indicating that the similarity between vector V S and V ST is 0% (i.e., completely different).
[0105] Based on the second translated sentence vector and the second source sentence vector, obtain the second similarity.
[0106] Specifically, the corpus classification device takes the second source sentence vector and the first translated sentence vector as a corpus pair combination, and uses the similarity calculation method between vectors to obtain the second similarity.
[0107] Preferably, the corpus classification device will calculate the second source sentence vector and the first translated sentence vector using the Euclidean dot product formula, and the calculation formula for the obtained second similarity is as follows:
[0108]
[0109] where SIM2 is the second similarity, V T is the first translated sentence vector, V SS is the second source sentence vector, and θ2 is the angle between V T and V SS between them.
[0110] Based on the first similarity and the second similarity, perform weighted summation to obtain the target similarity.
[0111] Specifically, the corpus classification device sets the weight values corresponding to the first similarity and the second similarity, and performs summation calculation to obtain the target similarity.
[0112] The calculation formula for the target similarity is as follows:
[0113] SIM = a·SIM1 + b·SIM2
[0114] Among them, SIM is the target similarity, SIM1 is the first similarity, a is the weight value corresponding to the first similarity, SIM2 is the second similarity, b is the weight value corresponding to the second similarity, and the sum of the weight values is 1. In the embodiments of the present invention, the values of the weight values a and b are not specifically limited.
[0115] Preferably, a and b are equal and both take the value of 0.5.
[0116] In the embodiments of the present invention, based on the cross-combination of sentence vectors obtained by different methods, the first similarity and the second similarity between vectors are calculated respectively, and the target similarity is obtained through weighted operations on the first similarity and the second similarity. Furthermore, classification is performed through the target similarity. It is possible to rely on simple vector calculations to achieve the classification of different corpora, with simple operation and high interpretability.
[0117] Based on any of the above embodiments, based on the original text word segmentation and the translated text word segmentation, the first original text sentence vector and the first translated text sentence vector are obtained, including: embedding the original text word segmentation and the translated text word segmentation respectively to obtain the original text word vector and the translated text word vector.
[0118] Specifically, in step 102, the corpus classification device embeds the original text word segmentation and the translated text word segmentation in step 101 respectively to generate the original text word vector and the translated text word vector correspondingly.
[0119] The embodiments of the present invention do not specifically limit this process.
[0120] Preferably, the corpus classification device uses multilingual BERT to perform word embedding on the original text word segmentation, and the obtained original text word vector is denoted as V Si ={V S1 ,V S2 ,…,V Sn}, where n is the dimension of the original text word vector. Similarly, the translated text word vector obtained by embedding the translated text word segmentation is denoted as V Ti ={V T1 ,V T2 ,…,V Tm}, where m is the dimension of the translated text word vector.
[0121] Sum the original text word vector and the translated text word vector respectively to obtain the first original text sentence vector and the first translated text sentence vector.
[0122] Specifically, the corpus classification device performs addition combination on the original text word vectors respectively to obtain the first original text sentence vector, and at the same time, performs addition combination on the translated text word vectors to obtain the first translated text sentence vector.
[0123] The embodiments of the present invention do not specifically limit this process.
[0124] Exemplarily, the corpus classification device calculates the sum vector of each word vector, and its calculation formula is as follows:
[0125]
[0126]
[0127] In the embodiments of the present invention, based on the embedding of the source text segmentation and the target text segmentation, the source text word vectors and the target text word vectors are obtained. By simply combining and summing the source text word vectors and the target text word vectors, the first source sentence vector and the first target sentence vector are obtained. Furthermore, classification is performed based on the calculated target similarity. It is possible to rely on simple vector calculations to achieve the classification of different corpora, with simple operation and high interpretability.
[0128] Figure 2 It is a schematic flow chart of the vertical industry machine translation method provided by the present invention. As Figure 2 shown, the vertical industry machine translation method includes: Step 201, obtaining the source language text to be translated in the target vertical industry.
[0129] It should be noted that the execution subject of the vertical industry machine translation method is the vertical industry machine translation device.
[0130] The application scenario of the vertical industry machine translation device is in a machine translation model used in a certain vertical field. In the original samples, a certain type of sample corpus selected from a parallel corpus of other related fields is mixed in for training the model. And when applying the model, the corresponding target language text is translated from the source text in this vertical field.
[0131] The vertical industry machine translation device may include mobile terminals such as mobile phones, smart phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), navigation devices, smart bracelets, smart watches, etc., and fixed terminals such as digital TVs, desktop computers, etc. Hereinafter, it is assumed that the electronic device is a mobile terminal. However, those skilled in the art will understand that, except for components specifically for mobile purposes, the structure according to the embodiments of the present application can also be applied to fixed-type terminals.
[0132] The target vertical industry refers to a certain sub-industry field.
[0133] Specifically, in step 201, the vertical industry machine translation device determines the source language text to be translated according to the translation task for the target vertical industry.
[0134] Step 202, inputting the source language text to be translated into a pre-established target translation model to obtain the target language text corresponding to the source language text to be translated.
[0135] Among them, the target translation model adopts the initial training corpus pair of the target vertical industry, and the enhanced training corpus pair with the target type obtained from the target parallel corpus by using the corpus classification method of any of the above embodiments. The target type matches the target vertical industry, and the enhanced training corpus pair has a target proportional relationship with the initial training corpus pair.
[0136] It should be noted that the target translation model is obtained after training based on the source text and the translated text in the initial training corpus pair and the enhanced training corpus pair of the target vertical industry.
[0137] Among them, the initial training corpus pair is the source text-translated text selected from the field where the target vertical industry is located as the initial training sample.
[0138] The enhanced training corpus pair refers to the sample corpus pair added on the basis of the initial training sample. The specific implementation process is to classify the parallel corpus of other industries except the target vertical industry according to the corpus classification method in the above embodiments to obtain the corpus pair of the target type.
[0139] Among them, the target type is consistent with the type corresponding to the vertical industry, and can be the first target type or the second target type.
[0140] The corresponding training process is that the vertical industry machine translation device initializes the weight coefficients between the layers of the constructed target translation model, and then mixes the initial training corpus pair and the enhanced training corpus pair in a target proportional relationship to form a training set. The embodiments of the present invention do not make specific limitations on the target proportional relationship.
[0141] Preferably, the target proportional relationship is that the ratio of the initial training corpus pair to the enhanced training corpus pair is 1:x, where the value range of x is [0.5, 3].
[0142] Input a group of training samples in the training set into the neural network under the current weight coefficients, and calculate the outputs of the nodes of the input layer, hidden layer, and output layer in turn. According to the gradient descent method, the cumulative error between the final output result of the output layer and the actual connection position state type is used to correct the weight coefficients between the nodes of the input layer and the hidden layer. According to the above process, until all the training samples in the training set are traversed, the weight coefficients of the input layer and the hidden layer can be obtained.
[0143] Specifically, in step 202, the vertical industry machine translation device restores the target translation model in step 202 according to the weight coefficients of the input layer and the hidden layer of the neural network, and inputs the source language text to be translated into the trained target translation model, and a target language text with a certain translation relationship with the source language text to be translated can be obtained.
[0144] The embodiments of the present invention do not make specific limitations on this process.
[0145] Exemplarily, if the target translation model is a translation model for automobile manufacturing, since the translation style in the field of automobile manufacturing generally tends to be literal translation, the user can select the same amount of data as the automobile manufacturing corpus from the corpora in other fields that meet the first target type (i.e., general engineering) and mix them in as enhanced data to improve the performance of the model.
[0146] Based on the enhanced training corpus pairs with the target type screened by the corpus classification method, the embodiments of the present invention mix them into the initial training corpus pairs of the vertical industry to train the target translation model and improve the performance of the target translation model. Furthermore, in application, the source language text to be translated is used as the input of the target translation model, and the output result is the target language text, which can improve the accuracy of the target translation model.
[0147] Figure 3 It is a schematic structural diagram of the corpus classification device provided by the present invention. Based on the content of any of the above embodiments, as Figure 3 shown, the corpus classification device includes a word segmentation module 310, a first acquisition module 320, a second acquisition module 330, a similarity calculation module 340, and a first classification module 350, where:
[0148] The word segmentation module 310 is used to obtain the original text word segmentation and the translated text word segmentation of each target corpus based on the target parallel corpus.
[0149] The first acquisition module 320 is used to obtain the first original sentence vector and the first translated sentence vector based on the original text word segmentation and the translated text word segmentation.
[0150] The second acquisition module 330 is used to embed the first original sentence vector and the first translated sentence vector respectively to obtain the second original sentence vector and the second translated sentence vector.
[0151] The similarity calculation module 340 is used to calculate the target similarity based on the first original sentence vector, the first translated sentence vector, the second original sentence vector, and the second translated sentence vector.
[0152] The first classification module 350 is used to determine that the type of the target corpus is the first target type when the target similarity is greater than or equal to the target threshold.
[0153] Among them, the target threshold is set based on the target parallel corpus. The number of target parallel corpora can be one or more, and the target corpus is the text data in the target parallel corpus.
[0154] Specifically, the word segmentation module 310, the first acquisition module 320, the second acquisition module 330, the similarity calculation module 340, and the first classification module 350 are electrically connected in sequence.
[0155] The word segmentation module 310 performs word segmentation on the original text and the translation included in each target corpus in the target parallel corpus, and obtains the corresponding original text word segmentation and translation word segmentation.
[0156] The first acquisition module 320 uses vector representation for the original text word segmentation obtained in the word segmentation module 310, and then performs combined addition to obtain the first original text sentence vector. At the same time, it also uses vector representation for the translation word segmentation obtained in the word segmentation module 310, and performs combined addition to obtain the first translation sentence vector.
[0157] The second acquisition module 330 uses the sentence embedding algorithm for the first original text sentence vector and the first translation sentence vector obtained in the first acquisition module 320 respectively, to obtain the second original text sentence vector and the second translation sentence vector.
[0158] The similarity calculation module 340 cross - combines the two types of sentence vectors generated by the first acquisition module 320 and the second acquisition module 330 in the way that one type of original text sentence vector corresponds to the other type of translation sentence vector, calculates the similarity corresponding to each combination, and then combines them to obtain the target similarity.
[0159] The first classification module 350 compares the target similarity of each target corpus with the corresponding target threshold. There are two comparison results: meeting the target threshold and not meeting the target threshold.
[0160] Among them, meeting the target threshold means that the target similarity is greater than or equal to the target threshold. This indicates that the similarity between the word and sentence vectors and the embedded sentence vectors is relatively high, that is, the sentence can be literally explained, and it is determined that the target corpus is more inclined to the "generalized engineering" category and belongs to the first target type.
[0161] Optionally, the device further includes a second classification module, where:
[0162] The second classification module is used to determine that the type of the target corpus is the second target type when the target similarity is less than the target threshold.
[0163] Among them, the first target type and the second target type are mutually exclusive.
[0164] Optionally, the similarity calculation module 340 includes a first calculation unit, a second calculation unit, and a total calculation unit, where:
[0165] The first calculation unit is used to obtain the first similarity based on the first original text sentence vector and the second translation sentence vector.
[0166] A second calculation unit, configured to obtain a second similarity based on the first translated sentence vector and the second original sentence vector.
[0167] A total calculation unit, configured to perform weighted summation based on the first similarity and the second similarity to obtain a target similarity.
[0168] Optionally, the first obtaining module 320 includes a word vector unit and a sentence vector unit, where:
[0169] The word vector unit is configured to respectively embed the original text word segmentation and the translated text word segmentation to obtain an original text word vector and a translated text word vector.
[0170] The sentence vector unit is configured to respectively sum the original text word vector and the translated text word vector to obtain a first original sentence vector and a first translated sentence vector.
[0171] The corpus classification device provided by an embodiment of the present invention is configured to execute the above-mentioned corpus classification method of the present invention. Its implementation manner is consistent with the implementation manner of the corpus classification method provided by the present invention and can achieve the same beneficial effects, which will not be elaborated here.
[0172] In an embodiment of the present invention, the first original sentence vector and the first translated sentence vector are respectively formed by simply combining the original text word segmentation and the translated text word segmentation of the target corpus. The second original sentence vector and the second translated sentence vector are obtained through sentence embedding. By cross-combining the sentence vectors obtained in different ways, the target similarity between the vectors is calculated, and when the target similarity is greater than or equal to the target threshold, the target corpus is classified as the first target type. It can rely on simple vector calculations to realize the classification of different corpora, with simple operations and high interpretability.
[0173] Figure 4 It is a schematic structural diagram of a vertical industry machine translation device provided by the present invention. Based on the content of any of the above embodiments, as Figure 4 shown, the vertical industry machine translation device includes a source text obtaining module 410 and a translation module 420, where:
[0174] The source text obtaining module 410 is configured to obtain the source language text to be translated in the target vertical industry.
[0175] The translation module 420 is configured to input the source language text to be translated into a pre-established target translation model to obtain the target language text corresponding to the source language text to be translated.
[0176] Wherein, the target translation model uses the initial training corpus pair of the target vertical industry, and the enhanced training corpus pair with the target type obtained from the target parallel corpus by using any of the corpus classification methods. The target type matches the target vertical industry, and the enhanced training corpus pair has a target proportional relationship with the initial training corpus pair.
[0177] Specifically, the source text acquisition module 410 and the translation module 420 are electrically connected in sequence.
[0178] The source text acquisition module 410 determines the source language text to be translated according to the translation task for the target vertical industry.
[0179] The translation module 420 restores the target translation model according to the weight coefficient between the input layer and the hidden layer of the neural network, and inputs the source language text to be translated into the trained target translation model, and can obtain the target language text having a certain translation relationship with the source language text to be translated.
[0180] The vertical industry machine translation device provided by the embodiment of the present invention is used to execute the above vertical industry machine translation method of the present invention. Its implementation manner is consistent with the implementation manner of the vertical industry machine translation method provided by the present invention, and the same beneficial effects can be achieved, which will not be elaborated here.
[0181] The embodiment of the present invention is based on the enhanced training corpus pair with the target type screened from the corpus classification method, and is mixed into the initial training corpus pair of the vertical industry to train the target translation model and improve the performance of the target translation model. Furthermore, in application, the source language text to be translated is used as the input of the target translation model, and the output result is the target language text, which can improve the accuracy of the target translation model.
[0182] Figure 5 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the corpus classification method, which includes: based on the target parallel corpus, obtaining the original text word segmentation and the translated text word segmentation of each target corpus; based on the original text word segmentation and the translated text word segmentation, obtaining the first original sentence vector and the first translated sentence vector; embedding the first original sentence vector and the first translated sentence vector respectively to obtain the second original sentence vector and the second translated sentence vector; calculating the target similarity based on the first original sentence vector, the first translated sentence vector, the second original sentence vector, and the second translated sentence vector; when the target similarity is greater than or equal to the target threshold, the type of the target corpus is the first target type; where the target threshold is set based on the target parallel corpus, the number of target parallel corpora can be one or more, and the target corpus is the text data in the target parallel corpus. It can also execute the vertical industry machine translation method, which includes: obtaining the source language text to be translated in the target vertical industry; inputting the source language text to be translated into the pre-established target translation model to obtain the target language text corresponding to the source language text to be translated; where the target translation model uses the initial training corpus pair of the target vertical industry, and the enhanced training corpus pair obtained from the target parallel corpus by using any of the above corpus classification methods, the target type matches the target vertical industry, and the enhanced training corpus pair has a target proportional relationship with the initial training corpus pair.
[0183] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0184] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the corpus classification method provided by each of the above methods. The method includes: based on a target parallel corpus, obtaining the original text word segmentation and the translated text word segmentation of each target corpus; based on the original text word segmentation and the translated text word segmentation, obtaining a first original sentence vector and a first translated sentence vector; embedding the first original sentence vector and the first translated sentence vector respectively to obtain a second original sentence vector and a second translated sentence vector; calculating a target similarity based on the first original sentence vector, the first translated sentence vector, the second original sentence vector and the second translated sentence vector; in the case where the target similarity is greater than or equal to a target threshold, the type of the target corpus is a first target type; wherein, the target threshold is set based on the target parallel corpus, the number of target parallel corpora can be one or more, and the target corpus is the text data in the target parallel corpus. The vertical industry machine translation method can also be executed. The method includes: obtaining a source language text to be translated in a target vertical industry; inputting the source language text to be translated into a pre-established target translation model to obtain a target language text corresponding to the source language text to be translated; wherein, the target translation model uses an initial training corpus pair of the target vertical industry, and an enhanced training corpus pair with a target type obtained from the target parallel corpus by using any one of the above corpus classification methods, the target type matches the target vertical industry, and the enhanced training corpus pair has a target proportional relationship with the initial training corpus pair.
[0185] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the corpus classification method provided by the above various methods. The method includes: based on a target parallel corpus, obtaining the original text word segmentation and the translated text word segmentation of each target corpus; based on the original text word segmentation and the translated text word segmentation, obtaining a first original text sentence vector and a first translated text sentence vector; embedding the first original text sentence vector and the first translated text sentence vector respectively to obtain a second original text sentence vector and a second translated text sentence vector; calculating a target similarity based on the first original text sentence vector, the first translated text sentence vector, the second original text sentence vector and the second translated text sentence vector; in the case where the target similarity is greater than or equal to a target threshold, the type of the target corpus is a first target type; wherein, the target threshold is set based on the target parallel corpus, the number of target parallel corpora can be one or more, and the target corpus is the text data in the target parallel corpus. It can also execute a vertical industry machine translation method, which includes: obtaining a source language text to be translated in a target vertical industry; inputting the source language text to be translated into a pre-established target translation model to obtain a target language text corresponding to the source language text to be translated; wherein, the target translation model uses an initial training corpus pair of the target vertical industry, and an enhanced training corpus pair with the target type obtained from the target parallel corpus by using any one of the above corpus classification methods, the target type matches the target vertical industry, and the enhanced training corpus pair has a target proportional relationship with the initial training corpus pair.
[0186] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0187] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A corpus classification method, characterized in that, including: Based on the target parallel corpus, obtain the original text word segmentation and the translated text word segmentation of each target corpus; Based on the original text word segmentation and the translated text word segmentation, obtain the first original sentence vector and the first translated sentence vector; Embed the first original sentence vector and the first translated sentence vector respectively to obtain the second original sentence vector and the second translated sentence vector; the first original sentence vector is composed of the sum of the vectors of the original text word segmentation, the first translated sentence vector is composed of the sum of the vectors of the translated text word segmentation, the second original sentence vector is the text representation of the original sentence combination containing the semantic information of the paragraph context, and the second translated sentence vector is the text representation of the translated sentence combination containing the semantic information of the paragraph context; Based on the first original sentence vector, the first translated sentence vector, the second original sentence vector and the second translated sentence vector, calculate the target similarity; When the target similarity is greater than or equal to the target threshold, the type of the target corpus is the first target type; When the target similarity is less than the target threshold, the type of the target corpus is the second target type; wherein, the first target type and the second target type are mutually exclusive; wherein, the target threshold is used to divide the translation style type of the target corpus, the first target type is the "literal translation" translation style type, the second target type is the "free translation" translation style type, the target threshold is set based on the target parallel corpus, the number of the target parallel corpora can be one or more, and the target corpus is the text data in the target parallel corpus.
2. The corpus classification method according to claim 1, wherein The calculating the target similarity based on the first original sentence vector, the first translated sentence vector, the second original sentence vector and the second translated sentence vector includes: Based on the first original sentence vector and the second translated sentence vector, obtain the first similarity; Based on the first translated sentence vector and the second original sentence vector, obtain the second similarity; Based on the first similarity and the second similarity, perform weighted summation to obtain the target similarity.
3. The corpus classification method according to claim 1, wherein The obtaining the first original sentence vector and the first translated sentence vector based on the original text word segmentation and the translated text word segmentation includes: Embed the original text word segmentation and the translated text word segmentation respectively to obtain the original word vector and the translated word vector; Sum the original word vector and the translated word vector respectively to obtain the first original sentence vector and the first translated sentence vector.
4. A method for machine translation in a vertical industry, characterized in that, including: Obtain the source language text to be translated in the target vertical industry; Input the source language text to be translated into the pre-established target translation model to obtain the target language text corresponding to the source language text to be translated; wherein, the target translation model uses the initial training corpus pair of the target vertical industry, and the enhanced training corpus pair with the target type obtained from the target parallel corpus by using the corpus classification method as described in any one of claims 1 to 3, the target type matches the target vertical industry, and the enhanced training corpus pair has a target proportional relationship with the initial training corpus pair.
5. A corpus classification device, characterized in that, including: A word segmentation module, configured to obtain the source text word segmentation and the target text word segmentation of each target corpus based on a target parallel corpus; A first acquisition module, configured to obtain a first source sentence vector and a first target sentence vector based on the source text word segmentation and the target text word segmentation; A second acquisition module, configured to respectively embed the first source sentence vector and the first target sentence vector to obtain a second source sentence vector and a second target sentence vector; the first source sentence vector is composed of the sum of the vectors of the source text word segmentation, the first target sentence vector is composed of the sum of the vectors of the target text word segmentation, the second source sentence vector is a text representation of the source text sentence combination containing paragraph context semantic information, and the second target sentence vector is a text representation of the target text sentence combination containing paragraph context semantic information; A similarity calculation module, configured to calculate a target similarity based on the first source sentence vector, the first target sentence vector, the second source sentence vector, and the second target sentence vector; A first classification module, configured to determine that the type of the target corpus is a first target type when the target similarity is greater than or equal to a target threshold; A second classification module, configured to determine that the type of the target corpus is a second target type when the target similarity is less than the target threshold; Wherein, the first target type and the second target type are mutually exclusive; Wherein, the target threshold is used to divide the translation style type of the target corpus, the first target type is the "literal translation" translation style type, the second target type is the "free translation" translation style type, the target threshold is set based on the target parallel corpus, the number of the target parallel corpora can be one or more, and the target corpus is the text data in the target parallel corpus.
6. A vertical industry machine translation device, characterized in that, Including: A source text acquisition module, configured to acquire a source language text to be translated in a target vertical industry; A translation module, configured to input the source language text to be translated into a pre-established target translation model to obtain a target language text corresponding to the source language text to be translated; Wherein, the target translation model uses an initial training corpus pair in the target vertical industry, and an enhanced training corpus pair with a target type obtained from the target parallel corpus by using the corpus classification method described in any one of claims 1 to 3, the target type matches the target vertical industry, and the enhanced training corpus pair has a target proportional relationship with the initial training corpus pair.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the corpus classification method described in any one of claims 1 to 3 or the steps of the vertical industry machine translation method described in claim 4.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the corpus classification method described in any one of claims 1 to 3 or the steps of the vertical industry machine translation method described in claim 4.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the corpus classification method described in any one of claims 1 to 3 or the steps of the vertical industry machine translation method described in claim 4.
Citation Information
Patent Citations
Corpus classification method and device
CN110196910A
Translation model training method, device and equipment and storage medium
CN112560510A