A model training method, device, computer device, and readable storage medium
By identifying and activation the original sample set, dividing the sample set and generating the activation sample set, the translation model is trained, and the problem of noisy data affecting the performance of the translation model is solved, and the accuracy and data volume of the translation model are guaranteed.
Patent Information
- Application Number
- CN202011261074.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-11-12
AI Technical Summary
The prior art in the training of translation model is due to the existence of noisy data, resulting in a reduction in the amount of training data, affecting the performance and accuracy of the translation model.
By obtaining the original sample set, using the original recognition model for sample recognition, dividing it into the target sample set and the first activation sample set, and calling the activation model to activate the target sample, generating the second activation sample set, and finally training the recognition processing model to ensure the amount of training data while improving the accuracy of the translation model.
Without adding additional data and changing the model architecture, the translation accuracy of the translation model is effectively improved and training time is shortened.
Smart Images

Figure CN112257470B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a model training method, apparatus, computer device, and readable storage medium. Background Art
[0002] A translation model is a data-driven task. To obtain a translation model with good performance, a large amount of training data is often required for training. However, the presence of noisy data in these large-scale training data makes the training of the machine translation model very difficult, thereby affecting the performance of the translation model.
[0003] Currently, to solve the above problems, generally, a large amount of original training data is denoised, and then the remaining original training data after denoising is used to train the translation model. This results in a small amount of data for training the translation model, and the translation accuracy of the translation model trained based on a small amount of sample data is low. Summary of the Invention
[0004] Embodiments of the present application provide a model training method, apparatus, computer device, and readable storage medium, which can reasonably utilize training data to train a translation model, and while ensuring the amount of data for training the translation model, effectively improve the translation accuracy of the trained translation model.
[0005] On the one hand, an embodiment of the present application provides a model training method, including:
[0006] Obtain an original sample set, and train a sample recognition model according to the original sample set to obtain an original recognition model; the original sample set includes a plurality of original samples;
[0007] Call the original recognition model to perform recognition processing on each original sample respectively to obtain an original recognition result of each original sample;
[0008] According to the original recognition result of each original sample, divide the original sample set into a target sample set and a first activation sample set, where the target sample set includes at least one target sample;
[0009] Call an activation model to perform activation processing on each target sample to obtain a second activation sample set; wherein, the activation model is trained based on the first activation sample set;
[0010] Train a recognition processing model according to the first activation sample set and the second activation sample set to obtain a target recognition model.
[0011] On the one hand, an embodiment of the present application provides a model training apparatus, including:
[0012] An acquisition module, configured to acquire an original sample set, train a sample recognition model according to the original sample set, and obtain an original recognition model; the original sample set includes a plurality of original samples;
[0013] A calling module, configured to call the original recognition model to perform recognition processing on each original sample respectively, and obtain the original recognition result of each original sample;
[0014] A determination module, configured to divide the original sample set into a target sample set and a first activation sample set according to the original recognition result of each original sample, where the target sample set includes at least one target sample;
[0015] The calling module is further configured to call an activation model to perform activation processing on each target sample, and obtain a second activation sample set; wherein, the activation model is trained based on the first activation sample set;
[0016] The determination module is further configured to train a recognition processing model according to the first activation sample set and the second activation sample set, and obtain a target recognition model.
[0017] One aspect of an embodiment of the present application provides a computer device, including a processor and a memory, the processor and the memory are connected to each other, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the model training method described above.
[0018] One aspect of an embodiment of the present application provides a computer-readable storage medium, in which program instructions are stored, and when the program instructions are executed, they are used to implement the model training method described above.
[0019] One aspect of an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium, and when the computer instructions are executed by a processor of a computer device, the model training method described above is executed.
[0020] In an embodiment of the present application, the computer device activates the target sample set in the original sample set to obtain a first activation sample set, and trains a recognition processing model according to the first activation sample set and the second activation sample set in the original sample set. Without changing the training model and adding extra data, there is no need to remove noise data from the training data, and the training data is reasonably utilized to train the translation model. While ensuring the data volume of the trained translation model, the translation accuracy of the trained translation model is effectively improved. Description of the Drawings
[0021] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0022] Figure 1 It is a schematic structural diagram of a model training provided by an embodiment of the present application;
[0023] Figure 2 It is a schematic flowchart of a model training method provided by an embodiment of the present application;
[0024] Figure 3 It is a schematic flowchart of a model training method provided by an embodiment of the present application;
[0025] Figure 4 It is a schematic flowchart of a model training method provided by an embodiment of the present application;
[0026] Figure 5 It is the performance of an NMT model when different target ratios are regarded as inactive samples provided by an embodiment of the present application;
[0027] Figure 6 It is a schematic structural diagram of a model training device provided by an embodiment of the present application;
[0028] Figure 7 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0030] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0031] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0032] Among them, natural language processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0033] The model training method provided by the embodiments of this application involves technologies such as natural language processing in artificial intelligence, and can distinguish a large amount of original data, that is, divide a large amount of original data into original inactive data and original active data. Further, the computer device activates the inactive data in the original data, so as to train the recognition processing model according to the activated data and the original active data to obtain a target recognition model, thereby realizing that there is no need to perform denoising processing on the original data, ensuring the data volume for training the translation model, shortening the model training time, and effectively improving the translation accuracy of the trained translation model.
[0034] In a feasible embodiment, such as Figure 1Schematic diagram of the architecture for model training shown. The architecture for model training mainly involves two models, namely the original recognition model and the activation model. When wanting to train a translation model, the computer device can first obtain a large number of original samples, where the original samples include source-side data X and the corresponding target-side data Y of the source-side data, and call the pre-trained recognition model to recognize the original samples, thereby obtaining a plurality of inactive samples and a plurality of active samples. Then the computer device trains the preset model according to the plurality of active samples to obtain the activation model. Further, the computer device calls the activation model to process the data of each inactive sample, that is, translates the source-side data X of the inactive sample to obtain the target-side data Y1, synthesizes the source-side data X and the target-side data Y1 into a parallel corpus (X, Y1), and constructs an activation sample with the parallel corpus (X, Y1); activation samples corresponding to each inactive sample can be obtained according to the above method. Then, the recognition processing model is trained according to each activation sample and each active sample, thereby completing the training of the recognition processing model to obtain the target recognition model.
[0035] In specific applications, the model training method provided in the embodiments of the present application can be applied to the training of all machine translation models. For example, the machine translation model can be a neural machine translation model (Neural Machine Translation, NMT). In the embodiments of the present application, when training a translation model using this model training method, the architecture of the translation model does not need to be changed. The model training method is illustrated by the following embodiments. Further, since the model training method provided in the embodiments of the present application mainly involves the processing of sample data, the solution of the present application can be applied to the training of all artificial intelligence models. For example, the training of image recognition models, the training of semantic segmentation models, etc.
[0036] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a model training method provided in an embodiment of the present application. The model training method can be executed by a computer device. The model training method described in this embodiment includes the following steps S201 - S205:
[0037] S201, obtain an original sample set, and train a sample recognition model according to the original sample set to obtain an original recognition model.
[0038] Among them, the original sample set includes a plurality of original samples. Each original sample can be a text-type sample, a voice-type sample, etc. For example, each original sample can include a first source text and a first target text corresponding to the first source text, that is, each original sample contains sample data and the label of the sample data.
[0039] In a specific implementation, the computer device can obtain multiple original samples from a preset source. For example, it can obtain multiple original samples from the storage space of the computer device or from the network. Then, based on the multiple original samples, a sample recognition model is trained to obtain an original recognition model.
[0040] S202. Invoke the original recognition model to perform recognition processing on each original sample respectively, and obtain the original recognition result of each original sample.
[0041] Among them, the original recognition result may include the first predicted text, or a set of phrase probabilities corresponding to the first target text.
[0042] S203. According to the original recognition result of each original sample, divide the original sample set into a target sample set and a first activation sample set.
[0043] Among them, the target sample set may include at least one target sample. It can be understood that each of these target samples can be considered as samples that have little or negative contribution to the training of the subsequent recognition processing model, that is, they can be called inactive samples. The first activation sample set may include at least one first activation sample set. Each of the first activation samples can be considered as text that has a greater contribution to the training of the subsequent recognition processing model, that is, they can be called active samples. In a specific implementation, dividing the original sample set into a target sample set and a first activation sample set can be determined according to the first predicted text of each original sample or the set of phrase probabilities corresponding to the first target text. In a feasible embodiment, the original recognition result includes the set of phrase probabilities corresponding to the first target text. The computer device adds up the phrase probabilities in the set of phrase probabilities to obtain a target probability, and divides the original sample set into a target sample set and a first activation sample set according to the target probability. In another feasible embodiment, the original recognition result includes the first predicted text. The computer device can determine the recognition error of each original sample according to the first predicted text of each original sample, and divide the original sample set into a target sample set and a first activation sample set according to the recognition error of each original sample.
[0044] S204. Invoke the activation model to perform activation processing on each target sample to obtain a second activation sample set.
[0045] Among them, the activation model is trained based on the first activation sample set. The activation model can be a forward activation model or a reverse activation model. In a specific implementation, forward translation and reverse translation can be used to implement the activation model. After the computer device obtains the first activation sample set, it can train the forward model (i.e., the forward translation model) or the reverse model (i.e., the reverse translation model) according to each first activation sample, so as to obtain a forward activation model or a reverse activation model.
[0046] In a feasible embodiment, each target sample includes a second source text and a second target text. It can be known that the second source text here is the first source text in the aforementioned original sample, and the second target text is the first target text in the aforementioned original sample. The computer device can extract the second source text or the second target text or both the second source text and the second target text from the second source text and the second target text of each target sample as the text to be processed for each target sample. Then, the computer device calls the activation model to perform activation processing on the text to be processed for each target sample, obtains the second predicted text for each target sample, combines the text to be processed for each target sample and the second predicted text into each second activation sample, and combines each second activation sample into a second activation sample set. In specific implementation, the following three methods can be selected to generate the second activation sample set according to requirements or experience.
[0047] (1) The computer device can extract the second source text from the second source text and the second target text of each target sample as the text to be processed for each target sample. Then, the computer device calls the forward activation model to perform activation processing on the second source text of each target sample, obtains the second predicted text for each target sample, and combines the second source text of each target sample and the second predicted text of each target sample into each second activation sample. Exemplarily, the target sample is (x, y), the text to be processed is x (i.e., the aforementioned second source text), the computer device calls the forward activation model to perform activation processing on x, obtains the second predicted text y1 of the target sample, and the computer device combines x and y1 into the activation sample (x, y1).
[0048] (2) The computer device can extract the second target text from the second source text and the second target text of each target sample as the text to be processed for each target sample. Then, the computer device calls the reverse activation model to perform activation processing on the second target text of each target sample, obtains the second predicted text for each target sample, and combines the second target text of each target sample and the second predicted text of each target sample into each second activation sample. Exemplarily, the target sample is (x, y), the text to be processed is y (i.e., the aforementioned second target text), the computer device calls the reverse activation model to perform activation processing on y, obtains the second predicted text x1 of the target sample, and the computer device combines x1 and y into the activation sample (x1, y).
[0049] (3) The computer device can extract the second source text and the second target text from the second source text and the second target text of each target sample as the text to be processed for each target sample. At this time, the computer device calls the forward activation model to perform activation processing on the second source text of each target sample to obtain the third predicted text of each target sample; calls the reverse activation model to perform activation processing on the second target text of each target sample to obtain the fourth predicted text of each target sample. Finally, the computer device combines the second source text and the third predicted text of each target sample into each second activation sample, and combines the second target text and the fourth predicted text of each target sample into each second activation sample, so that the subsequent generated second activation sample set can maximize the performance of the target recognition model.
[0050] Exemplarily, the target sample is (x, y), the text to be processed is x, the computer device calls the forward activation model to perform activation processing on x (i.e., the above-mentioned second source text) to obtain y1 of the target sample (i.e., the above-mentioned third predicted text); calls the reverse activation model to process y (i.e., the second target text) to obtain x1 of the target sample. According to the user's needs, the computer determines both x1 and y1 as the second predicted text. Then, x1 and y of the target sample are combined into the activation sample (x1, y), and x and y1 of the target sample are combined into the activation sample (x, y1).
[0051] It should be noted that the above process of activating each target sample is a process of re-labeling the second source text or the second target text of the target sample. For example, if the second source text is x, then x is re-labeled, and this label can be understood as y1. Generating each second activation sample actually means synthesizing parallel corpora, and the parallel corpora can be the above-mentioned (x1, y), (x, y1), etc.
[0052] S205. Train the recognition processing model according to the first activation sample set and the second activation sample set to obtain the target recognition model.
[0053] In a feasible embodiment, after obtaining the target recognition model, the user can submit a request on the web page or translation interface provided by the computer device. The computer device can respond to the translation request for the text to be translated, call the target recognition model to perform text translation on the text to be translated to obtain the translation text of the text to be translated, and then output the translation text. By translating through this target recognition model, a relatively accurate translation text can be obtained.
[0054] In the embodiments of the present application, the computer device activates the target sample set in the original sample set to obtain the first activation sample set, and trains the recognition processing model according to the first activation sample set and the second activation sample set in the original sample set. Without changing the training model and adding extra data, there is no need to remove the noise data. By reasonably using the training data to train the translation model, while ensuring the amount of data for training the translation model, the translation accuracy of the trained translation model is effectively improved. And during the process of activating the target sample set, different activation models can exist to ensure the accuracy of the obtained first activation sample set, further improving the translation accuracy of the trained translation model.
[0055] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of another model training method provided by the embodiments of the present application. This model training method can be executed by a computer device. The model training method described in this embodiment mainly introduces the case where the original recognition result includes the first predicted text, and includes the following steps S301 - S307:
[0056] S301. Obtain the original sample set, and train a sample recognition model according to the original sample set to obtain an original recognition model.
[0057] Among them, the original sample set includes multiple original samples.
[0058] It should be noted that the specific implementation manner of step S301 can refer to the specific implementation manner of step S201.
[0059] S302. Invoke the original recognition model to perform recognition processing on each original sample respectively to obtain the first predicted text of each original sample.
[0060] Among them, each original sample includes a first source text and a first target text corresponding to the first source text.
[0061] In a feasible embodiment, the computer device can invoke the original recognition model to perform recognition processing on the first source text of each original sample respectively to obtain the first predicted text of each original sample. In another feasible embodiment, the first predicted text includes a first unit text and a second unit text, and the original recognition model includes a forward recognition model and a reverse recognition model. The computer device can train two preset recognition models respectively according to a large amount of training data to obtain the forward recognition model and the reverse recognition model. Then the computer device can invoke the forward recognition model to perform recognition processing on the first source text to obtain the first unit text; the computer device can invoke the reverse recognition model to perform recognition on the first target text to obtain the second unit text.
[0062] S303: Determine a recognition error of each original sample according to the first predicted text corresponding to each original sample.
[0063] In a feasible embodiment, the present application proposes two methods of determining the recognition error of each original sample, one of which is to determine the recognition error of each original sample through the first target text and the first predicted text corresponding to each original sample; the other is to determine the recognition error of each original sample based on the first predicted text, the first source text and the first target text of each original sample.
[0064] First, the first determination method is described in detail:
[0065] After obtaining the first predicted text of each original sample, a polling priority can be set for multiple original samples. The computer device selects a target original sample from multiple original samples according to the polling priority, and aligns the first target text of the target original sample with the first predicted text of the target original sample, counts the overlapping characters of the first target text and the first predicted text after the alignment, and determines the recognition error of the target original sample according to the overlapping characters and the total number of characters. Wherein, the total number of characters refers to the total number of characters of the first target text of the target original sample or the first predicted text of the target original sample. In a specific implementation, in order to accurately determine the recognition error of the target original sample, the first target text of the target original sample and the first predicted text of the target original sample can be aligned, so that it can be determined whether the characters at the same position are the same. The computer device counts the number of identical characters, and determines the sample accuracy of the recognition processing of the target original sample according to the ratio of the number of identical characters to the total number of characters, so that the recognition error of the target original sample can be determined according to the sample accuracy, and when each original sample is used as the target original sample, the polling is stopped.
[0066] For example, the target original sample includes a first source text "have lunch", and the first target text corresponding to the first source text is "吃饭". The computer device calls the original recognition model to recognize the first source text to obtain the first predicted text of the original sample as "吃瓜". The computer aligns "吃饭" and "吃瓜". The computer device counts the number of identical characters as 1, that is, only the character "吃" is the same, and the total number of characters is 2. According to the ratio of the number of identical characters 1 to the total number of characters 2, the sample accuracy of the recognition processing of the target original sample is determined to be 0.5. The computer device determines that the recognition error of the target original sample is 0.5 based on the sample accuracy.
[0067] In a feasible embodiment, after aligning the first target text of the target original sample and the first predicted text of the target original sample, if the characters at the same positions are the same, the probability at that position is 1; if the characters at the same positions are different, the probability at that position is 0. Then, the probabilities at all positions are added and averaged to determine the accuracy of the target original sample. According to the accuracy of the target original sample, the recognition error of the target original sample can be determined.
[0068] It should be noted that when each original sample is used as the target sample, the recognition error of each original sample can be determined according to the implementation process of the above target original sample.
[0069] A detailed description of the second determination method is as follows:
[0070] In a feasible embodiment, after the computer device obtains the first unit text and the second unit text, cross-entropy can be introduced to determine the recognition error of text recognition. In a specific implementation, the computer device can determine the first cross-entropy of each original sample according to the first unit text of each original sample and the first target text of each original sample; and determine the second cross-entropy of each original sample according to the second unit text of each original sample and the first source text of each original sample. Then, the first cross-entropy and the second cross-entropy of each original sample are superimposed as the recognition error of each original sample. Or the first cross-entropy and the second cross-entropy of each original sample are superimposed and averaged, and the averaged cross-entropy is used as the recognition error of each original sample.
[0071] S304, according to the recognition error of each original sample, determine at least one target sample that meets the first preset condition from each original sample, and determine at least one first activation sample that does not meet the first preset condition from each original sample.
[0072] In a specific implementation, the first preset condition can be set in advance according to requirements or experience. The computer device determines whether the recognition error of each original sample meets the first preset condition, and determines the original sample that meets the first preset condition as the target sample; and determines the original sample that does not meet the first preset condition as the first activation sample.
[0073] In a feasible embodiment, the first preset condition is that the recognition error is greater than the error threshold. The computer device can determine whether the recognition error of each original sample is greater than the error threshold, and determine the original sample corresponding to the recognition error greater than the error threshold as the target sample, and determine the original sample corresponding to the recognition error less than or equal to the error threshold as the first activation text.
[0074] In a feasible embodiment, if the first preset condition is to obtain the target proportion of original samples from each original sample from high to low, the computer device can sort each original sample according to the recognition error of each original sample, and obtain the original samples that meet the target proportion from high to low from each original sample, and use the obtained original samples that meet the target proportion as target samples, and determine the remaining original samples that do not meet the target proportion as the first activation samples. Exemplarily, if the number of original samples is 100 and the first preset condition is to obtain 10% of the original samples from 100 original samples from high to low, the computer device sorts according to the recognition errors of the 100 original samples, and then obtains the original samples that meet 10% from high to low from the 100 original samples as target samples, that is, obtains 10 original samples from high to low from the 100 original samples, and determines these 10 original samples as target samples; the remaining original samples that do not meet 10% are determined as the first activation samples, that is, each of the remaining 90 original samples is determined as the first activation sample.
[0075] S305, combine at least one target sample into a target sample set, and combine at least one first activation sample into a first activation sample set.
[0076] S306, call the activation model to perform activation processing on each target sample to obtain a second activation sample set; wherein, the activation model is trained based on the first activation sample set.
[0077] S307, train the recognition processing model according to the first activation sample set and the second activation sample set to obtain a target recognition model.
[0078] Among them, for the specific implementation manners of steps S306 - S307, reference may be made to the specific implementation manners of steps S204 - S205.
[0079] In the embodiments of the present application, the computer device determines the recognition errors of each original text according to the first predicted text, and divides each original sample according to whether the recognition error meets the first preset condition to obtain a reasonable target sample set and a first activation sample set. This ensures the amount of data required for model training, and after the subsequent target sample set is activated, it is used to train the recognition processing model together with each first activation sample, which more effectively improves the translation accuracy of the target recognition model.
[0080] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of another model training method provided by the embodiments of the present application. This model training method can be executed by a computer device. In the embodiments of the present application, the model training method described mainly introduces the case where the original recognition result includes the phrase probability set corresponding to the first target text, including the following steps S401 - S407:
[0081] S401, obtaining an original sample set, training a sample recognition model according to the original sample set, and obtaining an original recognition model.
[0082] In the specific implementation, the original sample set obtained can be used {[x n ,y n ]} indicates that the computer device is based on the original sample set {[x n ,y n ]} The training sample recognition model is actually to maximize the original sample set {[x n ,y n ]}, and the resulting original recognition model can subsequently determine the phrase probability of each original sample.
[0083] S402, calling the original recognition model to perform recognition processing on each original sample respectively, and obtaining a phrase probability set corresponding to the first target text of each original sample.
[0084] Each of the original samples includes a first source text and a first target sample corresponding to the first source text.
[0085] In a specific implementation, the probabilities of different phrases corresponding to the original recognition model, the computer device calls the original recognition model to recognize the first source text of each original sample respectively, and the phrase probability set corresponding to the first target text of each original sample can be determined according to the probabilities of different phrases corresponding to the original recognition model. Exemplarily, the original sample is (I am going to have lunch), and the probability of the phrase "have lunch" is 0.2 and the probability of the phrase "eat lunch" is 0.1. The computer device calls the original recognition model to recognize "I am going to have lunch". Among them, the phrase "eat" in the first source text corresponds to the phrase "have lunch" in the first target text of the original sample. The computer device can determine that the probability of the phrase "have lunch" is 0.2. Similarly, the probabilities of the phrases "I", "am", and "going to" can be determined in turn. The phrase probabilities in the first target text "I am going to have lunch" of the original sample are combined to obtain the phrase probability set corresponding to the first target text.
[0086] S403, superimposing the phrase probabilities in the phrase probability set as the target probability.
[0087] In a specific implementation, the computer device can directly superimpose the phrase probabilities in the phrase probability set as the target probability. Alternatively, the computer device can calculate the target probability from the phrase probabilities in the phrase probability set using maximum likelihood estimation, that is, the calculation formula of this maximum likelihood estimation is: I(y|x) = ∏p(y t |x,y<t). Where p(y t |x) is the phrase probability, I(y|x) represents the probability of obtaining y under x, and this target probability represents the confidence from the source - side sentence (i.e., the first source text) x to the target - side sentence y (i.e., the first target text). In a feasible embodiment, in order to be able to handle the influence brought by the text length subsequently, after determining the target probability, the first target sample can be normalized according to the target probability.
[0088] S404. According to the target probability of each original sample, determine at least one target sample that meets the second preset condition from each original sample, and determine at least one first activation sample that does not meet the second preset condition from each original sample.
[0089] Among them, the second preset condition can be set according to requirements or experience. Empirically speaking, if an original sample obtains a very low target probability, then it is unlikely to provide useful information for improving the performance of the target recognition model, and thus can be considered a target sample (or an inactive sample). Therefore, the target probability can be used as an indicator to measure the activity degree of each original sample, and then study the characteristic differences (activity or not) of each original sample. In a specific implementation, the computer device can sort each original sample according to the target probability of each original sample, and then determine whether the target probability of each original sample meets the second preset condition, divide the original samples whose target probabilities meet the second preset condition into target samples, and divide the original samples whose target probabilities do not meet the second preset condition into first activation samples.
[0090] In a feasible embodiment, the second preset condition is to obtain a target proportion of the original samples from the original samples in ascending order. Then, the computer device can sort each original sample according to the target probability of each original sample, obtain the original samples that meet the target proportion from the original samples in ascending order, and use the obtained original samples that meet the target proportion as the target samples, and determine the remaining original samples that do not meet the target proportion as the first activation samples. Exemplarily, if the number of original samples is 50 and the second preset condition is to obtain 10% of the original samples from 50 original samples in ascending order, the computer device sorts the 50 original samples according to the recognition error, and then obtains the original samples that meet 10% from the 50 original samples in ascending order as the target samples, that is, obtains 5 original samples from the 50 original samples in ascending order, and determines these 5 original samples as the target samples; the remaining original samples that do not meet 10% are determined as the first activation samples, that is, each of the remaining 45 original samples is determined as the first activation sample.
[0091] It should be noted that the target proportion can be determined according to the characteristic differences of the original sample set. Among them, the characteristic differences of the original sample set refer to the activity level. For a specific original sample set, the target proportion can be determined according to specific verification tests. For example, for the original sample set of "English - French", after specific verification tests, it can be known that it is most reasonable that the original samples with a target proportion of 10% are the target samples, which can improve the performance of the target recognition model.
[0092] In a feasible embodiment, the first preset condition is that the target probability is less than the sentence - level threshold. The computer device can determine whether the target probability of each original sample is less than the sentence - level threshold, and determine the original samples corresponding to less than the sentence - level threshold as the target samples, and determine the original samples corresponding to greater than or equal to the sentence - level threshold as the first activation texts. Among them, the sentence - level threshold can be set according to experience.
[0093] S405, Combine at least one target sample into a target sample set, and combine at least one first activation sample into a first activation sample set.
[0094] S406, Invoke the activation model to perform activation processing on each target sample to obtain a second activation sample set; wherein, the activation model is trained based on the first activation sample set;
[0095] S407, Train the recognition processing model according to the first activation sample set and the second activation sample set to obtain a target recognition model.
[0096] Among them, the specific implementation manners of steps S405 - S407 can refer to the specific implementation manners of steps S305 - S307.
[0097] In the embodiments of the present application, the computer device sets a second preset condition according to the performance requirements of the trained translation model, and divides each original text into a target sample and a first activation sample according to whether the phrase probability set corresponding to the first target text meets the second preset condition, so as to obtain an accurate and reasonable target sample set and a first activation sample set, ensuring the data volume required for model training, and making the subsequent target sample set be activated and used to train the recognition processing model together with each first activation sample, which more effectively improves the translation accuracy of the target recognition model.
[0098] In summary, the model training method of the embodiments of the present application mainly includes the following steps:
[0099] (1) First, train an NMT model based on the training data (corresponding to the above original sample set) as the recognition model M identification (corresponding to the above original recognition model).
[0100] (2) Calculate the sentence-level probability I(y|x) (corresponding to the above target probability) for each sample (x, y) of the training data through the recognition model M identification so as to sort all the training data according to I(y|x).
[0101] (3) Divide the training data into S = S inactive + S active , and consider that 10% of the training data with the lowest I(y|x) is inactive samples S inactive (corresponding to the above target samples), and the remaining 90% are active samples S active (corresponding to the above first activation samples).
[0102] (4) Train another NMT model based on the active samples S inactive as the activation model M rejuvenation .
[0103] (5) Translate the source-side sentence x of the inactive samples (corresponding to the above second source text) through the activation model M rejuvenation to obtain the target-side translation result y' (the second predicted text), and these synthetic parallel sentence pairs (x, y') form the activation samples (corresponding to the above second activation samples), that is
[0104] (6) Combine the activation samples with the original active samples to obtain S’ = S rejuvenated + S active , and train the final NMT model (target recognition model) based on the dataset S’.
[0105] For the model training method provided above, the present application verifies this model training method through the following three groups of experiments.
[0106] (1) When using the original sample set of "English - German" for experiments and selecting samples at different target ratios (the second preset condition) as inactive samples according to the target probability, the final performance improvement brought by this model training method to the NMT model. Figure 5 shows the performance of the NMT model when samples at different target ratios are regarded as inactive samples, and at the same time shows the final NMT model performance results when samples of the same ratio are randomly sampled and activated. Obviously, activating inactive samples is significantly better than activating an equal amount of randomly sampled samples. From Figure 5 it can be seen that the ratio of 10% is a reasonable target ratio. At the same time, as the ratio of inactive samples increases, the final performance improvement of the model gradually decreases. Intuitively, such a phenomenon is also reasonable, mainly because samples with relatively high target probabilities can already provide useful information to the NMT model. Forcibly activating them will not only not result in obvious improvement, but may even harm the final performance. Therefore, it is verified in the "English - German" experiment that 10% is a more appropriate target ratio for inactive samples.
[0107] (2) The model training method is verified on different model frameworks and language pairs. As shown in Table 1, this model training method has achieved consistent and significant improvements on powerful baseline models and two language pairs of WMT14 (a machine translation competition) "English - German" and "English - French", which demonstrates the effectiveness and generality of the model training method. It should be noted that this model training method can achieve significant improvements without changing the model and adding additional data, which enables this model training method to be robustly applicable to various existing NMT models.
[0108] Table 1
[0109]
[0110] (3) The proposed method for activating inactive samples is compared with past related work, including data diversification and data denoising. As shown in Table 2, both this model training method and the related work can individually improve the performance of the NMT model, but the model training method provided in the embodiments of the present application can further improve on the basis of other methods, which shows that the model training method and these related works are complementary. And we calculated the overlap between the noise samples identified in data denoising and our inactive samples, and found that only 32% of the inactive samples are also noise samples. This shows that our inactive samples are not completely composed of noisy data.
[0111] Table 2
[0112]
[0113] Regarding the verification of the model training method provided in the embodiments of the present application as described above, it can be seen that the embodiments of the present application only focus on the problem of translation data and do not need to adjust the framework of the target recognition model. Moreover, experiments have found that removing 10% of the least active samples hardly affects the performance of the NMT model. Further, we found that the least active samples obtained through different random seeds, model sizes, and model frameworks have a very high coincidence rate (80%). It can be seen that there are least active samples in large-scale data sets, and it depends not on the model but on the data distribution itself (i.e., the characteristics of the data itself). We conducted experiments on two standard translation data sets, WMT14 English-German and English-French, and found that data activation can consistently and significantly improve the performance of all NMT models.
[0114] Further, please refer to Figure 6 , which is a schematic structural diagram of a model training device provided in the embodiments of the present application. As Figure 6 shown, the model training device can be applied to the computer device in the above Figure 2 or Figure 3 or Figure 4 corresponding embodiments. Specifically, the model training device can be a computer program (including program code) running in the computer device. For example, the model training device is an application software; this model training device can be used to execute the corresponding steps in the method provided in the embodiments of the present application.
[0115] An acquisition module 601, configured to acquire an original sample set, train a sample recognition model according to the original sample set, and obtain an original recognition model; the original sample set includes a plurality of original samples;
[0116] A calling module 602, configured to call the original recognition model to perform recognition processing on each original sample respectively, and obtain an original recognition result of each original sample;
[0117] A determination module 603, configured to divide the original sample set into a target sample set and a first activation sample set according to the original recognition result of each original sample, where the target sample set includes at least one target sample;
[0118] The calling module 602 is further configured to call an activation model to perform activation processing on each target sample, and obtain a second activation sample set; wherein, the activation model is trained based on the first activation sample set;
[0119] The determination module 603 is further configured to train a recognition processing model according to the first activation sample set and the second activation sample set, and obtain a target recognition model.
[0120] In a feasible embodiment, the original recognition result includes a first predicted text; the determining module 603 is specifically configured to:
[0121] Determine the recognition error of each original sample according to the first predicted text corresponding to each original sample;
[0122] According to the recognition error of each original sample, determine at least one target sample that meets the first preset condition from the original samples, and determine at least one first activation sample that does not meet the first preset condition from the original samples;
[0123] Combine the at least one target sample into a target sample set, and combine the at least one first activation sample into a first activation sample set.
[0124] In a feasible embodiment, each original sample includes a first source text and a first target text corresponding to the first source text, and the first predicted text is the text obtained by the original recognition model after recognizing the first source text; the determining module 603 is specifically configured to:
[0125] Set a polling priority for the multiple original samples, and select a target original sample from the multiple original samples according to the polling priority;
[0126] Align the first target text of the target original sample with the first predicted text of the target original sample;
[0127] Count the overlapping characters of the first target text and the first predicted text after alignment processing, and determine the recognition error of the target original sample according to the overlapping characters and the total number of characters;
[0128] Stop polling when each original sample serves as the target original sample.
[0129] In a feasible embodiment, each original sample includes a first source text and a first target text corresponding to the first source text, the first predicted text includes a first unit text and a second unit text, the original recognition model includes a forward recognition model and a reverse recognition model, the first unit text is the text obtained by the forward recognition model recognizing the first source text, and the second unit text is the text obtained by the reverse recognition model recognizing the first target text; the determining module 603 is specifically configured to:
[0130] Determine the first cross entropy of each original sample according to the first unit text of each original sample and the first target text of each original sample;
[0131] Determine the second cross-entropy of each original sample according to the second unit text of each original sample and the first source text of each original sample;
[0132] Superimpose the first cross-entropy and the second cross-entropy of each original sample as the recognition error of each original sample.
[0133] In a feasible embodiment, each original sample includes a first source text and a first target text corresponding to the first source text, and the original recognition result includes a set of phrase probabilities corresponding to the first target text, and the set of phrase probabilities is obtained by the original recognition model performing recognition processing on the phrases in the first target text; the determining module 603 is specifically configured to:
[0134] Superimpose the phrase probabilities in the set of phrase probabilities as the target probability;
[0135] According to the target probability of each original sample, determine at least one target sample that meets the second preset condition from each original sample, and determine at least one first activation sample that does not meet the second preset condition from each original sample;
[0136] Combine the at least one target sample into a target sample set, and combine the at least one first activation sample into a first activation sample set.
[0137] In a feasible embodiment, each target sample includes a second source text and a second target text; the determining module 603 is configured to extract the text to be processed of each target sample from the second source text and the second target text of each target sample;
[0138] The calling module 602 is configured to call an activation model to perform activation processing on the text to be processed of each target sample to obtain a second predicted text of each target sample;
[0139] The determining module 603 is configured to combine the text to be processed of each target sample and the second predicted text into each second activation sample;
[0140] The determining module 603 is configured to combine the each second activation sample into a second activation sample set.
[0141] In a feasible embodiment, the target recognition model is a model for text translation; the apparatus further includes: an output module 604, wherein,
[0142] The determining module 603 is configured to respond to a translation request for the text to be translated;
[0143] The calling module is configured to call the target recognition model to perform text translation on the text to be translated, so as to obtain the translated text of the text to be translated;
[0144] The output module 604 is configured to output the translated text.
[0145] It can be understood that the functions of the functional modules of the model training device in this embodiment can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the relevant descriptions of the above method embodiments Figure 2 、 Figure 3 or Figure 4 and will not be elaborated herein.
[0146] Further, please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. The above Figure 2 or Figure 3 or Figure 4 The computer device in the corresponding embodiment may be Figure 7 the computer device shown. As Figure 7 shown, the computer device may include: a processor 701, an input device 702, an output device 703, and a memory 704. The above processor 701, input device 702, output device 703, and memory 704 are connected through a bus 705. The memory 704 is used to store a computer program, and the computer program includes program instructions. The processor 701 is configured to execute the program instructions stored in the memory 704.
[0147] In the embodiment of the present application, the processor 701 performs the following operations by running the executable program code in the memory 704: obtaining an original sample set, training a sample recognition model according to the original sample set to obtain an original recognition model; the original sample set includes a plurality of original samples; calling the original recognition model to perform recognition processing on each original sample respectively to obtain the original recognition result of each original sample; according to the original recognition result of each original sample, dividing the original sample set into a target sample set and a first activation sample set, the target sample set includes at least one target sample; calling an activation model to perform activation processing on each target sample to obtain a second activation sample set; wherein, the activation model is trained based on the first activation sample set; training a recognition processing model according to the first activation sample set and the second activation sample set to obtain a target recognition model.
[0148] In a feasible embodiment, the original recognition result includes a first predicted text; the processor 701 is specifically configured to: determine the recognition error of each original sample according to the first predicted text corresponding to each original sample; determine at least one target sample that meets a first preset condition and at least one first activation sample that does not meet the first preset condition from the original samples according to the recognition error of each original sample; combine the at least one target sample into a target sample set, and combine the at least one first activation sample into a first activation sample set.
[0149] In a feasible embodiment, each original sample includes a first source text and a first target text corresponding to the first source text, and the first predicted text is a text obtained by the original recognition model after performing recognition processing on the first source text; the processor 701 is specifically configured to: set a polling priority for the multiple original samples, select a target original sample from the multiple original samples according to the polling priority; perform alignment processing on the first target text of the target original sample and the first predicted text of the target original sample; count the overlapping characters of the first target text and the first predicted text after the alignment processing, and determine the recognition error of the target original sample according to the overlapping characters and the total number of characters; stop polling when each original sample serves as the target original sample.
[0150] In a feasible embodiment, each original sample includes a first source text and a first target text corresponding to the first source text, the first predicted text includes a first unit text and a second unit text, the original recognition model includes a forward recognition model and a reverse recognition model, the first unit text is a text obtained by the forward recognition model performing recognition on the first source text, and the second unit text is a text obtained by the reverse recognition model performing recognition on the first target text; the processor 701 is specifically configured to: determine the first cross entropy of each original sample according to the first unit text of each original sample and the first target text of each original sample; determine the second cross entropy of each original sample according to the second unit text of each original sample and the first source text of each original sample; superimpose the first cross entropy and the second cross entropy of each original sample as the recognition error of each original sample.
[0151] In a feasible embodiment, each of the original samples includes a first source text and a first target text corresponding to the first source text. The original recognition result includes a set of phrase probabilities corresponding to the first target text, and the set of phrase probabilities is obtained by the original recognition model through recognition processing of the phrases in the first target text. The processor 701 is specifically configured to: superimpose the phrase probabilities in the set of phrase probabilities to obtain a target probability; determine at least one target sample that meets a second preset condition from the original samples according to the target probability of each original sample, and determine at least one first activation sample that does not meet the second preset condition from the original samples; combine the at least one target sample into a target sample set, and combine the at least one first activation sample into a first activation sample set.
[0152] In a feasible embodiment, each of the target samples includes a second source text and a second target text. The processor 701 is specifically configured to: extract the text to be processed of each target sample from the second source text and the second target text of each target sample; call the activation model to perform activation processing on the text to be processed of each target sample to obtain a second predicted text of each target sample; combine the text to be processed and the second predicted text of each target sample into each second activation sample; combine the each second activation sample into a second activation sample set.
[0153] In a feasible embodiment, the target recognition model is a model for text translation. The processor 701 is specifically configured to: respond to a translation request for a text to be translated; call the target recognition model to perform text translation on the text to be translated to obtain a translated text of the text to be translated; output the translated text.
[0154] It should be understood that in the embodiments of the present application, the so-called processor 701 may be a central processing unit (CPU), and this processor 701 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0155] The memory 704 may include a read-only memory and a random access memory, and provide instructions and data to the processor 701. A part of the memory 704 may also include a non-volatile random access memory.
[0156] The input device 702 may include a keyboard and the like, and input a translation request to the processor 701; the output device 703 may include a display and the like.
[0157] In a specific implementation, the processor 701, the input device 702, the output device 703, and the memory 704 described in the embodiments of the present application may execute the implementation manners described in all the above embodiments, and may also execute the implementation manners described in the above devices, which will not be elaborated herein.
[0158] The embodiments of the present application provide a computer-readable storage medium, and the computer-readable storage medium stores a computer program. The computer program includes program instructions, and when the program instructions are executed by a processor, the steps executed in all the above embodiments can be executed.
[0159] The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium, and when the computer instructions are executed by a processor of a computer device, the methods in all the above embodiments are executed.
[0160] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0161] The above-disclosed is only a preferred embodiment of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.
Claims
1. A model training method, characterized in that, Including: Obtain an original sample set, train a sample recognition model according to the original sample set to obtain an original recognition model; the original sample set includes a plurality of original samples, and each original sample includes a first source text and a first target text corresponding to the first source text; Call the original recognition model to perform recognition processing on each original sample respectively to obtain the original recognition result of each original sample; According to the original recognition result of each original sample, divide the original sample set into a target sample set and a first activation sample set, and the target sample set includes at least one target sample; wherein, when the original recognition result of each original sample includes a phrase probability set corresponding to the corresponding first target text, and the phrase probability set is obtained by the original recognition model performing recognition processing on the phrases in the first target text, the dividing the original sample set into a target sample set and a first activation sample set according to the original recognition result of each original sample includes: adding the phrase probabilities in the phrase probability set in the original recognition result of any original sample as a target probability; according to the target probability of each original sample, determine at least one target sample that meets the second preset condition from each original sample, and determine at least one first activation sample that does not meet the second preset condition from each original sample, combine the at least one target sample into a target sample set, and combine the at least one first activation sample into a first activation sample set; Call an activation model to perform activation processing on each target sample to obtain a second activation sample set; wherein, the activation model is trained based on the first activation sample set; Train a recognition processing model according to the first activation sample set and the second activation sample set to obtain a target recognition model.
2. The method according to claim 1, wherein The original recognition result includes a first predicted text; the dividing the original sample set into a target sample set and a first activation sample set according to the original recognition result of each original sample includes: Determine the recognition error of each original sample according to the first predicted text corresponding to each original sample; According to the recognition error of each original sample, determine at least one target sample that meets the first preset condition from each original sample, and determine at least one first activation sample that does not meet the first preset condition from each original sample; Combine the at least one target sample into a target sample set, and combine the at least one first activation sample into a first activation sample set.
3. The method according to claim 2, wherein The first predicted text is the text obtained by the original recognition model performing recognition processing on the first source text; The determining the recognition error of each original sample according to the first predicted text corresponding to each original sample includes: Set a polling priority for the plurality of original samples, and select a target original sample from the plurality of original samples according to the polling priority; Align the first target text of the target original sample and the first predicted text of the target original sample; Count the overlapping characters between the first target text and the first predicted text after statistical alignment, and determine the recognition error of the target original sample according to the overlapping characters and the total number of characters; When each original sample is used as the target original sample, stop polling.
4. The method according to claim 2, wherein The first predicted text includes a first unit text and a second unit text. The original recognition model includes a forward recognition model and a reverse recognition model. The first unit text is the text obtained by the forward recognition model recognizing the first source text, and the second unit text is the text obtained by the reverse recognition model recognizing the first target text; The determining the recognition error of each original sample according to the first predicted text corresponding to each original sample includes: Determine the first cross-entropy of each original sample according to the first unit text of each original sample and the first target text of each original sample; Determine the second cross-entropy of each original sample according to the second unit text of each original sample and the first source text of each original sample; Superimpose the first cross-entropy and the second cross-entropy of each original sample as the recognition error of each original sample.
5. The method according to claim 1, wherein Each target sample includes a second source text and a second target text; The calling the activation model to perform activation processing on each target sample to obtain a second activation sample set includes: Extract the text to be processed of each target sample from the second source text and the second target text of each target sample; Call the activation model to perform activation processing on the text to be processed of each target sample to obtain the second predicted text of each target sample; Combine the text to be processed and the second predicted text of each target sample into each second activation sample; Combine each second activation sample into a second activation sample set.
6. The method according to claim 1, wherein The target recognition model is a model for text translation; the method further includes: Respond to a translation request for the text to be translated; Call the target recognition model to perform text translation on the text to be translated to obtain the translated text of the text to be translated; Output the translated text.
7. A model training device, characterized in that, Including: An acquisition module, configured to acquire an original sample set, and train a sample recognition model according to the original sample set to obtain an original recognition model; the original sample set includes a plurality of original samples, and each original sample includes a first source text and a first target text corresponding to the first source text; A calling module, configured to call the original recognition model to perform recognition processing on each original sample respectively to obtain the original recognition result of each original sample; A determination module, configured to divide the original sample set into a target sample set and a first activation sample set according to the original recognition results of each of the original samples, where the target sample set includes at least one target sample; wherein, when the original recognition result of each original sample includes a set of phrase probabilities corresponding to a corresponding first target text, and the set of phrase probabilities is obtained by the original recognition model through recognition processing of phrases in the first target text, the dividing the original sample set into a target sample set and a first activation sample set according to the original recognition results of each of the original samples includes: superimposing the phrase probabilities in the set of phrase probabilities in the original recognition result of any original sample as a target probability; according to the target probabilities of each of the original samples, determining at least one target sample that satisfies a second preset condition from the original samples, and determining at least one first activation sample that does not satisfy the second preset condition from the original samples, combining the at least one target sample into a target sample set, and combining the at least one first activation sample into a first activation sample set; The calling module is further configured to call an activation model to perform activation processing on each target sample to obtain a second activation sample set; wherein, the activation model is trained based on the first activation sample set; The determination module is further configured to train an identification processing model according to the first activation sample set and the second activation sample set to obtain a target identification model.
8. A computer device, characterized in that, It includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1-6.
9. A computer storage medium, characterized in that, The computer storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the method according to any one of claims 1-6 is executed.
10. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions are executed by a processor of a computer device, the method according to any one of claims 1-6 is executed.
Citation Information
Patent Citations
Text processing method and device based on deep learning, equipment and medium
CN111209377A