Data processing method and device, electronic equipment and computer readable storage medium
The source text is generated by pre-training the language model, and high-quality training corpus is generated using multiple compilation and scoring models. The problem of manual annotation dependence in the existing technology is solved, and the training efficiency and performance capabilities of the machine translation model are improved.
Patent Information
- Application Number
- CN202510105836.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
Existing machine translation technology relies on manual annotation of translation corpus, with limited coverage, small training samples and low generation efficiency.
The source text is automatically generated through the pre-trained language model, and the source text is compiled using multiple reference compilation models. The compiled translated text is scored in combination with the scoring model to generate training corpus pairs for preference optimization training, which are used to train the target compiled model.
Rapidly generate high-quality training corpus, improves the efficiency of corpus generation and improves the compilation performance capabilities of the compiled model.
Smart Images

Figure CN119990158A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data processing method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid development of globalization and informatization, the demand for machine translation technology is increasing. Whether it is communication among multinational companies, international academic research, or the acquisition of daily multi-language information, high-quality translation services are becoming more and more important in modern society.
[0003] Usually, in machine translation technology, the initial text is obtained through manual search, and then the translator translates the initial text into the target language to obtain manually annotated translation corpus. Then, the translation corpus is used as a training sample, and the compilation model is trained using a supervised learning method, thereby obtaining a trained compilation model. The compilation model can be applied to text translation.
[0004] However, current machine translation technology relies on translators to manually annotate translation corpora, which has limited coverage, a small number of training samples, and low efficiency in generating training samples. Summary of the invention
[0005] Based on this, the present application at least provides a data processing method, device, electronic device, computer-readable storage medium and computer program product.
[0006] In a first aspect, the present application provides a data processing method, comprising:
[0007] Generate source text based on the pre-trained language model and target text information;
[0008] Compile the source text by referring to the compilation model to obtain a first target text;
[0009] Scoring the first target text by using a scoring model to obtain a scoring result;
[0010] Based on the scoring result, a training corpus pair for preference optimization training is generated, and the training corpus pair is used to train a target compilation model.
[0011] In this embodiment, a source text is automatically generated with the help of a pre-trained language model, and the source text is compiled through multiple reference compilation models to obtain a first target text. The first target text is scored by a scoring model to obtain a scoring result. Based on the scoring result, the compiled translation is generated into a training corpus pair, and the training corpus pair contains the compilation preference accurately reflected by the scoring result. In this way, by adopting this method, high-quality training corpus can be quickly generated, thereby improving the efficiency of corpus generation.
[0012] In one embodiment, after generating a training corpus pair for preference optimization training based on the scoring result, the method further includes:
[0013] Get the initial target compilation model;
[0014] Performing supervised fine-tuning on the initial target compilation model to obtain a supervised fine-tuned target compilation model;
[0015] Based on the training corpus, preference optimization training is performed on the supervised fine-tuned target compilation model to obtain a trained target compilation model; the trained target compilation model is used to implement the compilation task corresponding to the target text information.
[0016] In this embodiment, based on a sufficient amount of high-quality training corpus pairs, the compilation model is trained in a manner combining supervised fine-tuning and preference optimization to obtain a trained compilation model, and the compilation performance capability of the compilation model is improved.
[0017] In one embodiment, the performing supervised fine-tuning on the initial target compilation model to obtain the supervised fine-tuned target compilation model includes:
[0018] Selecting a target compiled translation with the highest score among the multiple versions of the compiled translations;
[0019] Based on the target compiled translation and the source text, the initial compilation model is trained to obtain a supervised fine-tuned target compilation model.
[0020] In this embodiment, the compiled translation with the highest score among multiple versions of compiled translations is used as a reference compiled translation, and the initial compiled model is trained by supervised learning-style supervised fine-tuning in combination with the source text, thereby improving the compilation accuracy of the compiled model.
[0021] In one embodiment, generating a source text based on a pre-trained language model and target text information includes:
[0022] Inputting target text information into the pre-trained language model, wherein the target text information includes text content information, stylistic form, and language conversion type compiled by the target compilation model;
[0023] Processing the text content information, the stylistic form and the language conversion type through the pre-trained language model to obtain corpus material information that meets the requirements of the target text information;
[0024] The corpus material information is processed according to the information synthesis algorithm in the pre-trained language model to obtain the source text.
[0025] In this embodiment, the source text is automatically generated through the pre-trained language model and the pre-set target text information, without relying on manual information search and information integration, thereby improving the efficiency of source text generation.
[0026] In one embodiment, the first target text includes compiled translations output by each compilation model for the same source text, and compiling the source text by referring to the compilation model to obtain the first target text includes:
[0027] Obtain multiple reference compiled models;
[0028] The source text is compiled by using the multiple reference compilation models to obtain compiled translations output by the reference compilation models for the same source text.
[0029] In this embodiment, multiple reference compilation models are called to compile the source text, and translations with different styles and structures are generated through the multiple reference compilation models to improve the compilation diversity of the compiled translations, and further enhance the flexibility and adaptability of the compilation results.
[0030] In one embodiment, the first target text includes compiled translations output by various compilation models for the same source text, and the scoring of the first target text by the scoring model to obtain a scoring result includes:
[0031] Scoring the compiled translation using the plurality of scoring models to obtain an initial scoring result of the compiled translation;
[0032] Based on the initial scoring results obtained by the same scoring model, the compiled translations are sorted to obtain the scoring sorting results corresponding to the compiled translations;
[0033] Based on the score ranking results of the same score model corresponding to each of the compiled translations, each of the compiled translations is re-scored to obtain score results of the multiple score models for each of the compiled translations.
[0034] In this embodiment, the compiled translations are scored using multiple scoring models, and the scoring criteria of the scoring models are unified to obtain the scoring results of the compiled translations, thereby achieving optimization of the ranking of the compiled translations.
[0035] In one embodiment, generating a training corpus pair for preference optimization training based on the scoring result includes:
[0036] Pairing the compiled translations of the multiple versions in pairs to obtain compiled translation pairs;
[0037] Comparing the scoring results corresponding to the compiled translations in each compiled translation pair, and adding a preference tag to each compiled translation in the compiled translation pair based on the comparison result, to obtain an updated compiled translation pair;
[0038] The source text is added to the compiled translation in the updated compiled translation pair to generate a training corpus pair for preference optimization training.
[0039] In this embodiment, the user feedback preference of the compiled corpus in each compiled corpus pair is determined by ranking the compiled translations of each version given by the scoring results, and then a training corpus pair carrying a preference label is generated for preference optimization training of the target compilation model to improve the compilation performance capability of the target compilation model.
[0040] In a second aspect, the present application further provides a data processing device, comprising:
[0041] A first generation module, used to generate a source text based on a pre-trained language model and target text information;
[0042] A compiling module, used for compiling the source text by referring to the compiling model to obtain a first target text;
[0043] A scoring module, used to score the first target text using multiple scoring models to obtain a scoring result;
[0044] The second generating module is used to generate a training corpus pair for preference optimization training based on the scoring result, and the training corpus pair is used to train the target compilation model.
[0045] In a third aspect, the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0046] Generate source text based on the pre-trained language model and target text information;
[0047] Compile the source text by referring to the compilation model to obtain a first target text;
[0048] Scoring the first target text by using a scoring model to obtain a scoring result;
[0049] Based on the scoring result, a training corpus pair for preference optimization training is generated, and the training corpus pair is used to train a target compilation model.
[0050] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0051] Generate source text based on the pre-trained language model and target text information;
[0052] Compile the source text by referring to the compilation model to obtain a first target text;
[0053] Scoring the first target text by using a scoring model to obtain a scoring result;
[0054] Based on the scoring result, a training corpus pair for preference optimization training is generated, and the training corpus pair is used to train a target compilation model.
[0055] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:
[0056] Generate source text based on the pre-trained language model and target text information;
[0057] Compile the source text by referring to the compilation model to obtain a first target text;
[0058] Scoring the first target text by using a scoring model to obtain a scoring result;
[0059] Based on the scoring result, a training corpus pair for preference optimization training is generated, and the training corpus pair is used to train a target compilation model.
[0060] The above-mentioned data processing method, device, electronic device, computer-readable storage medium and computer program product automatically generate source text with the help of a pre-trained language model, and compile the source text through multiple compilation models. The scoring model scores multiple versions of compiled translations to obtain scoring results for each compiled translation. Based on the scoring results corresponding to the compiled translations, the compiled translations are generated into training corpus pairs, and the training corpus pairs contain the compilation preferences accurately reflected by the scoring results. In this way, by adopting this method, high-quality training corpus can be quickly generated, thereby improving the efficiency of corpus generation.
[0061] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions are briefly introduced below. Obviously, the drawings herein are incorporated into the specification and constitute a part of the specification. These drawings show embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure.
[0063] Figure 1An application environment diagram of a data processing method in an embodiment;
[0064] Figure 2 is a flow chart of a data processing method in one embodiment;
[0065] Figure 3 A flowchart of optimizing the training compilation model steps based on the training corpus pair preference in one embodiment;
[0066] Figure 4 A schematic diagram of a flow chart of a supervised fine-tuning step for a compiled model in one embodiment;
[0067] Figure 5 A schematic diagram of a process for generating a source language model based on a pre-trained language model in one embodiment;
[0068] Figure 6 A schematic diagram of a flow chart of steps of translating a source text based on a compilation model in one embodiment;
[0069] Figure 7 A flowchart of the steps of optimizing training corpus based on translation sorting in one embodiment;
[0070] Figure 8 A flowchart of the steps of generating training corpus pairs for preference optimization training in one embodiment;
[0071] Fig. 9 is a structural block diagram of a data processing device in one embodiment;
[0072] Fig.10 FIG. 4 is a diagram showing the internal structure of an electronic device in one embodiment. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0074] The data processing method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 generates a source text based on the pre-trained language model and the target text information. Taking the source text as the basis for generating the training corpus, the source text is compiled by referring to the compilation model to obtain the first target text. Then, the terminal 102 scores the first target text through the scoring model to obtain a scoring result. Furthermore, based on the scoring result, the terminal clarifies the preference of each compiled translation, marks the preference of each compiled translation, and generates a training corpus pair for preference optimization training. The target compilation model is trained by the generated large data volume and high-quality training corpus, so that the trained target compilation model can achieve accurate translation of the text to be translated.
[0075] The terminal 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. Head-mounted devices may be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.
[0076] In an exemplary embodiment, Figure 2 As shown, a data processing method is provided, which is applied to Figure 1 The terminal in is taken as an example to illustrate, including the following steps 202 to 206. Among them:
[0077] Step 202: Generate source text based on the pre-trained language model and target text information.
[0078] In the implementation, a pre-trained language model is integrated on the terminal side, and the pre-trained language model completes model training through pre-training, fine-tuning and context learning. Pre-training refers to the initial training of the model using a large-scale data set unsupervised learning method before the target task. During the pre-training process, the model will be exposed to a large amount of unlabeled text data, such as books, articles, websites, etc. A large amount of unlabeled text data is used to train the language model. The model acquires knowledge and features by learning the internal representation of the input data so as to perform fine-tuning or transfer learning on subsequent specific tasks. The goal of pre-training is to capture the underlying patterns, structures, and semantic knowledge present in the text corpus. After the pre-training is completed, the model is further fine-tuned and contextually learned based on the sample data, so that the language model after training can achieve intelligent text output.
[0079] Furthermore, the language model after training can output the corresponding text results based on the target text information input by the user. Specifically, in the process of corpus generation, the text of different compilation targets is output for different target text information. For example, taking the translation in the compilation form as an example, when the user wants to obtain the target compilation model in the form of a news release translated between Chinese and English, when generating the training corpus of the target compilation model, the target text information needs to require the output of the source text in the form of a news release in Chinese and English for training the target compilation model. Furthermore, if the target compilation model requires specialization in the translation of text in a certain field, the training corpus can also limit the content of a specific field, such as the financial field, the financial field, the communication field, etc. For another example, when the user wants to obtain a target compilation model for multi-language translation, for example, multi-language translation such as Chinese-English translation, Chinese-Russian translation, and English-Russian translation, when generating the training corpus of the target compilation model, it is necessary to generate source texts in different languages.
[0080] Specifically, the pre-trained language model supports text output in multiple languages, multiple application scenarios, and multiple text formats.
[0081] For example, a large-scale pre-trained language model (LLM) is used to generate source texts in multiple languages. Supported languages include but are not limited to Chinese, English, French, German, Russian, Italian, Japanese, Korean, Vietnamese, etc., covering mainstream languages and some low-resource languages. With multi-language support, the system can cope with multi-language needs in a global context, effectively cover multiple language scenarios, and support and process different language systems.
[0082] In addition, the source text generated using large-scale pre-trained language models can cover a variety of application scenarios, including but not limited to the following areas:
[0083] Economics: Produce texts related to finance, markets, investment, globalization, etc.
[0084] Science and Technology: covers artificial intelligence, information technology, quantum computing, aerospace technology and other scientific and technological fields;
[0085] Education: Produce texts on topics such as academic research, educational psychology, teaching methods, etc.;
[0086] History: texts dealing with historical events, cultural heritage, historical evolution, etc.;
[0087] Gastronomy: Generate texts on food culture, cooking techniques, catering industry, etc.
[0088] Culture: including art, drama, intangible cultural heritage, traditional festivals and other cultural contents;
[0089] Nature and Health: Generate content related to natural ecology, public health, mental health, and health preservation.
[0090] The ability to generate multiple application scenarios enables the pre-trained language model to provide customized high-quality translation materials for different professional fields and application scenarios. In addition, when generating source text, the style can be diversified and expanded according to the specific application scenario. The system supports the generation of multiple style forms, including but not limited to:
[0091] News: Generate objective and timely news report texts;
[0092] Narrative: Generate descriptive and narrative texts, such as memoirs, historical accounts, etc.
[0093] Dialogue: Generate dialogue text suitable for novels, dramas, interviews, etc.
[0094] Speech: Generate speech texts that express opinions or information to the public, such as academic reports;
[0095] Argumentative essays: Generate highly academic argumentative essays, suitable for reviews, editorials, etc.
[0096] Prose: Generates prose texts with free expression and elegant style;
[0097] Fiction: Generate short, novella or novel text suitable for literary creation;
[0098] Poetry: Generates rhythmic and beautiful poetry, such as free verse and metered verse.
[0099] Through multi-genre support, the system can adapt to different language styles and expressions, improving the diversity and wide application of translation materials.
[0100] Step 204: compile the source text by referring to the compilation model to obtain a first target text.
[0101] In implementation, the terminal obtains multiple reference compilation models, which may also be pre-trained language models, or other neural network models, machine learning models, etc. The embodiment of the present application does not limit the type and number of reference compilation models. Thus, after the terminal obtains the source text output by the pre-trained language model, multiple reference compilation models are called to compile multiple source texts to generate multiple versions of compiled translations (i.e., first target texts) with different structures and description styles, etc. For example, the multiple versions of compiled translations have different word orders, description words, etc.
[0102] Optionally, different reference compilation models may include different architectures and training methods, for example, different reference compilation models may be Llama, Qwen, etc., thus ensuring that each compiled translation has differences in semantic understanding, fluency, and style consistency. This multi-version compiled translation generation method can provide more diverse query options and enhance the flexibility and adaptability of the translation results.
[0103] Step 206: score the first target text using the scoring model to obtain a scoring result.
[0104] In implementation, after obtaining multiple versions of compiled translations (i.e., the first target text), the terminal scores the multiple versions of compiled translations by calling multiple scoring models to obtain scoring results corresponding to the multiple versions of compiled translations. For example, multiple scoring models may include, but are not limited to, traditional machine compilation models such as BLEU and ROUGE, and the embodiment of the present application does not limit the model type of the scoring model. In this way, multiple scoring models are used to score the compiled translations of each version from multiple dimensions such as semantic consistency, language fluency, stylistic adaptability, and style retention. The scoring range can be 1 to 10 points. Finally, the terminal comprehensively calculates the scores of each dimension and gives the initial scores of each version of the compiled translation, for example, by weighted average or weighted sum algorithm to obtain the initial scores of multiple versions of the compiled translation. Then. The scoring results of the compiled translations of multiple versions are calculated by the preset scoring rules to obtain the scoring results corresponding to the compiled translations of multiple versions.
[0105] Specifically, the scoring criteria for each dimension are as follows:
[0106] Semantic consistency: Check whether the translation retains the original semantics of the source text and is accurately communicated across multiple languages;
[0107] Language fluency: assessing the naturalness, fluency and grammatical correctness of the translation in the target language;
[0108] Stylistic adaptability: Based on the stylistic type of the source text, assess whether the translation is consistent with the source text in terms of format;
[0109] Style retention: Ensure that the translated text retains the emotional expression and tone of the source text after multilingual translation.
[0110] Optionally, a combination of multiple scoring models can be adjusted as needed to support different task objectives and evaluation focuses, thereby ensuring that the comprehensive quality of multi-version translations is effectively measured.
[0111] Step 208: Generate training corpus pairs for preference optimization training based on the scoring results.
[0112] Among them, the training corpus pairs are used to train the target compilation model.
[0113] In implementation, after obtaining the scoring results corresponding to the compiled translations of multiple versions, the terminal can compare and sort the compiled translations of multiple versions based on the scoring results, and further clarify the preferences of the compiled translations of each version through the sorting results, so as to add preference labels to the translations of multiple versions. Finally, a training corpus pair is generated based on the compiled translations of multiple versions carrying preference labels and the corresponding source texts. The training corpus pair is used to train the target compilation model so that the trained target compilation model can achieve the compilation target initially formulated (that is, the compilation requirements contained in the target text information input into the pre-trained language model).
[0114] Optionally, the compilation of the source text involved in the above data processing method can be to translate the source text, and then score the first target text obtained after translation, obtain the scoring results of each version of the translated text, and generate a training corpus pair for preference optimization training based on the scoring results, and the training corpus pair can be used to train the target translation model. The target translation model can achieve a model of a pre-determined translation target.
[0115] In the above data processing method, a source text is automatically generated with the help of a pre-trained language model, and the source text is compiled through multiple compilation models. The scoring model scores multiple versions of compiled translations to obtain scoring results for each compiled translation. Based on the scoring results corresponding to the compiled translations, the compiled translations are used to generate training corpus pairs, which contain compilation preferences accurately reflected by the scoring results. In this way, by adopting this method, high-quality training corpus can be quickly generated, thereby improving the efficiency of corpus generation.
[0116] Optionally, various types of models in the data processing method can also be integrated on the server, and then, when the user initiates a request to generate a training corpus pair, the generation of the training corpus pair is completed through the interaction between the terminal and the server. The specific implementation steps of this process are similar to the steps of performing data processing on the terminal side, and the embodiments of the present application will not be repeated. At the same time, since the data processing method of the present application can be applied to the terminal side, it can also be implemented based on the interaction between the terminal and the server. Therefore, the embodiments of the present application do not limit the specific application scenarios of the data processing method.
[0117] In an exemplary embodiment, Figure 3 As shown, after step 208, the data processing method further includes:
[0118] Step 302: Obtain an initial target compilation model.
[0119] In implementation, the terminal obtains an initial target compilation model, which is a target compilation model to be trained for preference optimization. The target compilation model may be a machine learning model such as a neural network model, and the specific type of the compilation model is not limited in the embodiment of the present application.
[0120] Step 304 , perform supervised fine-tuning on the initial target compilation model to obtain a supervised fine-tuned target compilation model.
[0121] During implementation, the terminal performs supervised fine-tuning (SFT) on the initial target compilation model to obtain the supervised fine-tuning target compilation model, and improves the compilation accuracy and other performance capabilities of the target compilation model through supervised fine-tuning.
[0122] Among them, the specific processing process of supervised fine-tuning includes the following steps:
[0123] Pre-training: First, the target compilation model is trained using an unsupervised dataset (an unlabeled dataset, which can be any text data in the corresponding language) to allow the compilation model to learn the basic structure and knowledge of the language.
[0124] Data collection and standards: Collect specific data related to the target task, for example, multiple versions of compiled translations obtained from multiple reference compilation models in this application, generate training language pairs, and select the best training language pair (i.e., the training language pair with the highest score) from the training language pairs as the annotated specific data for model supervision fine-tuning.
[0125] Supervised fine-tuning: The pre-trained target compilation model is further trained on a labeled dataset. Through these labeled data, the model can learn how to make predictions and inferences on specific tasks.
[0126] Evaluation and optimization: Evaluate the performance of the compiled model through the validation dataset, adjust the hyperparameters, and further optimize the output results of the compiled model so that the compiled model achieves the best performance on the target task.
[0127] Optionally, the supervised fine-tuning step for the target compilation model is an optional step, that is, in order to improve the performance of the compilation model, the target compilation model can be pre-supervised fine-tuned. Similarly, the target compilation model may not be fine-tuned, and preference optimization training may be performed directly on the target compilation model. Therefore, the embodiments of the present application do not limit the specific training process of the compilation model.
[0128] Step 306 , performing preference optimization training on the supervised fine-tuned target compilation model based on the training corpus to obtain a trained compilation model.
[0129] Among them, the trained target compilation model is used to implement the compilation task corresponding to the target text information.
[0130] In implementation, after obtaining the training corpus pair, the terminal uses the training corpus pair as a training sample to further train the compilation model after supervised fine-tuning, i.e., preference optimization training. The preference labels corresponding to each compiled translation in the training corpus pair are obtained based on the scoring results given by multiple scoring models. Furthermore, through the training corpus pair, an optimization method for directly adjusting the model generation results in combination with user preference feedback is implemented to improve the performance of the target compilation model. Finally, the trained target compilation model can achieve the compilation task corresponding to the target text information initially set.
[0131] Specifically, the specific process of performing preference optimization training DPO (Direct Preference Optimization) on the target compilation model includes the following steps:
[0132] Step 1. Collection of preference data: Usually, users or annotators rank multiple output results generated by the model and give preferences (i.e., which result is more in line with expectations).
[0133] These preference data become the core basis for model optimization. DPO directly uses human subjective preferences to better reflect the actual needs of users. In order to ensure the objectivity of the compiled translation preference results and avoid relying on manual annotation, the compiled translations are scored through multiple scoring models in this application, thereby obtaining a training corpus pair with preference optimization based on the scoring results.
[0134] Step 2. Sorting supervision: The optimization goal of DPO is to make the results generated by the compiled model more in line with preferences.
[0135] By comparing and ranking different candidate outputs generated by the model, the target compilation model can learn how to generate outputs that the user prefers, guiding the target compilation model to make preferred outputs.
[0136] Step 3. Contrastive learning: DPO is based on the idea of contrastive learning. During training, the target compilation model will learn the relative advantages and disadvantages between the two compiled translations in the training corpus pair.
[0137] This relative learning mechanism allows the model to better optimize the output quality in the high-dimensional generative space.
[0138] Step 4. Loss function design: DPO will design a special loss function to maximize the correctness of user preference ranking. For example, using a loss function based on contrastive learning, such as Pairwise Ranking Loss or Margin-based Loss, the model will adjust its generation distribution according to user preferences and reduce the probability of results that do not meet preferences.
[0139] Step 5. Model training and fine-tuning: DPO can be combined with large-scale pre-trained models (such as GPT, BERT, etc.) to fine-tune the model weights so that the output results of the target compilation model are more in line with the preference ranking.
[0140] During the process of supervised fine-tuning (SFT) or direct preference optimization, the target compilation model gradually learns and adapts to user preferences, improving the quality and consistency of the output results.
[0141] In this embodiment, during the DPO model training process, DPO training has the advantage of no reward function. Unlike the generative optimization method based on reinforcement learning, DPO does not require the pre-design of a complex reward function. Traditional reinforcement learning methods usually rely on reward signals to guide model generation, but the design of rewards may be very difficult and may not accurately reflect actual preferences. DPO directly uses preference data to avoid the challenges and instability brought about by reward function design.
[0142] In an optional embodiment, DPO can be combined with other optimization techniques (such as reinforcement learning, supervised fine-tuning, genetic algorithms, etc.) to form a hybrid optimization strategy. For example, DPO is first used to optimize the generated sorting, and then supervised fine-tuning (SFT) is used to further improve the generation quality of the model. Therefore, this application does not limit the training method of the target compilation model.
[0143] In this embodiment, based on a sufficient amount of high-quality training corpus pairs, the target compilation model is trained in a manner combining supervised fine-tuning and preference optimization to obtain a trained target compilation model, and the compilation performance capability of the target compilation model is improved.
[0144] In an exemplary embodiment, Figure 4 As shown, step 304 includes steps 401 to 402. Among them:
[0145] Step 401 : Filter the target compiled translation with the highest score among multiple versions of compiled translations.
[0146] In implementation, among multiple versions of compiled translations, multiple scoring models score each version of the compiled translations respectively, and obtain scoring results corresponding to each version of the compiled translations. Then, the terminal can sort the compiled translations based on the scoring results of the multiple versions of the compiled translations to achieve the screening of the target compiled translation. For example, the compiled translations of multiple versions of the same source text include: translation A, translation B and translation C, and the scoring results corresponding to these three versions of the translations are sorted from large to small as: translation A> translation B> translation C. Thus, the terminal selects the compiled translation with the highest scoring result as the target compiled translation, that is, translation A is the target compiled translation.
[0147] Step 402: Based on the target compiled text and the source text, an initial target compiled model is trained to obtain a supervised fine-tuned target compiled model.
[0148] In implementation, after determining the target compiled translation, the terminal uses the target compiled translation as the standard compiled translation of the source text, annotates the source text with the target compiled translation, implements supervised samples, and performs supervised learning training on the initial target compiled model, thereby obtaining a supervised fine-tuned target compiled model. In the supervised fine-tuning process, the hyperparameters of the initial target compiled model are adjusted through multiple iterations, the matching features between the target compiled translation and the source text are learned, and a corresponding loss function is set, which is used to calculate the loss difference between the output result of the compilation model and the target compiled translation. Based on the loss difference obtained by the loss function, it is judged whether the compilation model meets the training stop condition. For example, when the loss difference meets the preset loss threshold, it is determined that the training of the target compiled model is completed, and the supervised fine-tuned target compiled model is obtained.
[0149] In this embodiment, the compiled translation with the highest score among multiple versions of compiled translations is used as a reference compiled translation, and the initial target compiled model is trained by supervised learning-style supervised fine-tuning in combination with the source text, thereby improving the translation accuracy of the target compiled model.
[0150] In an exemplary embodiment, the process of generating a source text based on preset target text information is described in detail below. Figure 5 As shown, the specific processing process of step 202 includes steps 501 to 503, wherein:
[0151] Step 501: input the target text information into the pre-trained language model.
[0152] The target text information includes the text content information, stylistic form and language conversion type translated by the target compilation model.
[0153] In implementation, a pre-trained language model is pre-integrated in the terminal, and the pre-trained language model may be an LLM model, etc. The pre-trained language model provides an AI conversation page, in which the user can input the target text information, and the target text information includes the text content information, stylistic form, and language conversion type to be translated by the target compilation model, etc. For example, the target text information includes generating a "Chinese-English" economic news report, the content of which involves stock market analysis, and the generated Chinese news title may be "Stocks have risen sharply, and signs of economic recovery are obvious." The target text information can be used as a label for inputting a pre-trained language model, or as a corpus generation condition for a pre-trained language model, and input into the pre-trained language model, and displayed in the AI conversation page provided by the pre-trained language model to generate corpus results.
[0154] Step 502, the text content information, style form and language conversion type are processed by a pre-trained language model to obtain corpus material information that meets the requirements of the target text information.
[0155] In the implementation, the terminal searches for relevant corpus according to the input text content information, style form and language conversion type by running the pre-trained language model to obtain corpus material information that meets the requirements of the target text information. Among them, the information search scope of the pre-trained language model can be the database, knowledge base, etc. corresponding to the pre-trained language model itself, or various shared information contents disclosed on the entire network and platform. The embodiment of the present application does not limit the information source of the pre-trained speech model.
[0156] Step 503: Process the corpus material information according to the information synthesis algorithm in the pre-trained language model to obtain the source text.
[0157] In implementation, the terminal synthesizes the searched corpus material information according to the information synthesis algorithm in the pre-trained language model. For example, through information aggregation, information splicing, etc., a source text that meets the target text information requirements is generated. For example, if the target text information requires the generation of a source text in the form of a press release text, the article structure needs to include a title, introduction, body, conclusion, etc., and the content of the press release must ensure the objectivity and timeliness of the information.
[0158] In this embodiment, the source text is automatically generated through the pre-trained language model and the pre-set target text information, without relying on manual information search and information integration, thereby improving the efficiency of source text generation.
[0159] In an exemplary embodiment, the following is specifically described, such as Figure 6 As shown, the specific processing process of step 204 includes steps 601 to 602, wherein:
[0160] Step 601: Acquire multiple reference compilation models.
[0161] In implementation, the terminal obtains multiple reference compilation models. The multiple reference compilation models include different architectures and training methods. The multiple reference compilation models may include, but are not limited to, Llama, Qwen models, etc.
[0162] Step 602 : compile the source text using multiple reference compilation models to obtain compiled translations output by the compilation models for the same source text.
[0163] In implementation, the terminal compiles the source text through multiple reference compilation models to obtain the compiled translation output by each reference compilation model for the same source text. For example, for the economic scenario required in the above example, the source text involves stock market analysis, and the content of the source text includes the latest stock market data, signs of economic recovery, expert analysis, etc. The corresponding translation results given by different compilation models are as follows:
[0164] Translation A given by compilation model A: The stock market saw a significant rise, indicating clear signs of economic recovery.
[0165] Translation B given by compilation model B: There was a substantial increase in the stockmarket, signaling a strong economic recovery.
[0166] Compilation model C gives translation C: The stock market experienced a notable surge, suggesting clear indications of economic recovery.
[0167] In this embodiment, the terminal calls multiple compilation models to compile the source text, and generates translations with different styles and structures through the multiple compilation models to improve the compilation diversity of the compiled translations, and further enhance the flexibility and adaptability of the translation results.
[0168] In an exemplary embodiment, Figure 7As shown, the specific processing process of step 206 includes steps 701 to 703, wherein:
[0169] Step 701 : Scoring multiple versions of compiled translations using multiple scoring models to obtain initial scoring results of the compiled translations.
[0170] In implementation, multiple scoring models are pre-integrated in the terminal. The scoring model can use traditional machine translation evaluation models such as BLEU and ROUGE, or models such as TextRank and TF-IDF. In this way, the terminal calls multiple scoring models to implement multi-dimensional scoring of multiple versions of compiled translations, and obtains scores of different dimensions. Then, the terminal can perform weighted average or weighted sum calculations on the scores of different dimensions of the same version of the compiled translation to obtain the initial scoring results corresponding to the same version of the compiled translation. For example, for translation A, translation A is scored in four dimensions: semantic consistency, language fluency, problem adaptability, and style retention, and the scoring range for each dimension is 1-10 points. Finally, the scores of these four dimensions are weighted and summed, and the initial scoring result of translation A is 9 points. Among them, the scoring range of the initial scoring result is 1-10 points.
[0171] Step 702: sorting the compiled translations based on the initial scoring results obtained by the same scoring model to obtain the scoring and sorting results corresponding to the compiled translations.
[0172] In implementation, the terminal sorts the compiled translations of multiple versions based on the initial scoring results obtained by the same scoring model, and obtains multiple scoring sorting results corresponding to the compiled translations of multiple versions. Specifically, the scoring of multiple scoring models is as follows:
[0173] Scoring model 1: Translation A (10 points), Translation B (7 points), Translation C (8 points)
[0174] Scoring model 2: Translation A (8 points), Translation B (6 points), Translation C (5 points)
[0175] Scoring model 3: Translation A (8 points), Translation B (9 points), Translation C (6 points)
[0176] Among them, based on the initial scoring results of the compiled translations of multiple versions using the same scoring model, the compiled translations of multiple versions are sorted, and the initial sorting results of the compiled translations of multiple versions are obtained as follows:
[0177] Scoring model 1: Translation A> Translation C> Translation B
[0178] Scoring model 2: Translation A> Translation B> Translation C
[0179] Scoring model 3: Translation B>Translation A>Translation C.
[0180] Step 703: re-score each compiled translation based on the score ranking results of the same scoring model corresponding to each compiled translation, and obtain the score results of multiple scoring models for each compiled translation.
[0181] In implementation, the terminal re-assigns scores to the multiple versions of the compiled translations based on the score ranking results of the same scoring model corresponding to the multiple versions of the compiled translations. Specifically, since different scoring models have different scoring standards and scoring calculation rules for the compiled translations, the scoring standards of the multiple scoring models obtained in the end are not unified. Therefore, the terminal can re-assign scores to the multiple versions of the compiled translations according to the unified scoring rules based on the score ranking results of the multiple versions of the compiled translations given by each scoring model, so as to achieve the unification of the scoring standards.
[0182] For example, the unified scoring rule is to assign points according to the sorting order of each compiled translation in the scoring sorting result. The compiled translation ranked first gets 10 points, the compiled translation ranked second gets 9 points, the compiled translation ranked third gets 8 points, and so on, until all the compiled translations are scored.
[0183] Therefore, based on the sorting order given by the above initial sorting results, the scoring results after each scoring model re-assigns scores to the compiled translations of multiple versions are as follows:
[0184] Scoring model 1: Translation A (10 points), Translation C (9 points), Translation B (8 points)
[0185] Scoring model 2: Translation A (10 points), Translation B (9 points), Translation C (8 points)
[0186] Scoring model 3: Translation B (10 points), Translation A (9 points), Translation C (8 points)
[0187] After re-scoring, the terminal can calculate the scoring results of the multiple versions of the compiled translations according to the unified scoring rules and the weighted average algorithm, and after obtaining the scoring results of the multiple versions of the compiled translations, the final ranking results of the multiple versions of the compiled translations can also be obtained: that is, the final score of translation A is (10+10+9) / 3=9.66, the final score of translation B is (8+9+10) / 3=9, and the final score of translation C is (9+8+8) / 3=8.33. Therefore, the final ranking is translation A>translation B>translation C.
[0188] In this embodiment, the compiled translations are scored using multiple scoring models, and the scoring criteria of the scoring models are unified to obtain the scoring results of the compiled translations, thereby achieving optimization of the ranking of the compiled translations.
[0189] In an exemplary embodiment, Figure 8 As shown, the specific processing process of step 208 includes steps 801 to 803, wherein:
[0190] Step 801 : Pair the compiled translations of multiple versions in pairs to obtain compiled translation pairs.
[0191] In implementation, the terminal pairs the compiled translations of multiple versions in pairs to obtain compiled translation pairs. For example, taking the compiled translations of multiple versions as translation A, translation B, and translation C as an example, the multiple versions of translations are paired in pairs to obtain compiled translation pairs consisting of (translation A, translation B), (translation A, translation C), and (translation B, translation C).
[0192] Step 802 : compare the scoring results corresponding to the compiled translations in each compiled translation pair, and add a preference tag to each compiled translation in the compiled translation pair based on the comparison result to obtain an updated compiled translation pair.
[0193] In implementation, the terminal compares the compiled translations in each compiled translation pair based on the scoring results, and adds a preference tag to each compiled translation in the compiled translation pair based on the comparison results. Specifically, the preference tag includes an acceptance tag and a rejection tag, wherein the principle of adding the preference tag is to label the translation ranked first in the compiled translation pair with an acceptance tag, and to label the translation ranked later with a rejection tag. In this way, the terminal can obtain a compiled translation pair with a preference tag added, that is, an updated compiled translation pair. For example, for the sorting result of the final sorting result: translation A> translation B> translation C, a preference tag is added to the compiled translation pair composed of these multiple compiled translations, and the specific situation is as follows:
[0194] Compile translation pair: (Translation A, Translation B): Translation A (accepted), Translation B (rejected)
[0195] Compiled translation pair: (Translation A, Translation C): Translation A (accepted), Translation C (rejected)
[0196] Compile translation pair: (Translation B, Translation C): Translation B (accepted), Translation C (rejected).
[0197] Step 803: Add source text to the compiled translation in the updated compiled translation pair to generate a training corpus pair for preference optimization training.
[0198] In implementation, the terminal adds source text to the compiled translation in the updated compiled translation pair to generate a training corpus pair for preference optimization training. Specifically, after adding preference tags to the compiled translation in each compiled translation pair, the terminal adds source text to each compiled translation pair, thereby generating a final training corpus pair, which is used to implement preference optimization training for the compiled model.
[0199] For example, if the source text is the Chinese economics news release given in the above embodiment, the Chinese original text represents the source text, and then the generated training corpus pairs are represented as follows:
[0200] (Chinese original text + translation A, Chinese original text + translation B)
[0201] (Chinese original text + translation A, Chinese original text + translation C)
[0202] (Chinese original text + translation B, Chinese original text + translation C)
[0203] In this embodiment, the user feedback preference of the compiled corpus in each compiled corpus pair is determined by ranking the compiled translations of each version given by the scoring results, and then a training corpus pair carrying a preference label is generated for preference optimization training of the compilation model to improve the compilation performance capability of the compilation model.
[0204] In an optional embodiment, taking translation as an example, the present application provides several application scenarios of machine translation corpus generation and translation sorting optimization based on a pre-trained language model (also referred to as a large model) in a data processing method. Specifically, it is mainly applied to the following scenarios:
[0205] 1. Multilingual Corpus Construction
[0206] Academic research: Researchers need to build multilingual corpora for research in linguistics, machine learning, and natural language processing. This method can automatically generate high-quality corpora covering multiple languages, multiple scenarios, and multiple styles, providing rich data support for research.
[0207] Language resource development: Language resource development organizations need to build high-quality multilingual corpora to support language teaching, language model training, etc. This method can generate diverse translation corpora and improve the richness and quality of the corpus.
[0208] 2. Machine Compilation Model Training
[0209] Model training: Machine translation system developers need a large amount of high-quality translation data to train and optimize the model. This method can generate high-scoring translations and further improve the performance of the model and translation quality through direct preference optimization (DPO).
[0210] Data enhancement: During the training process, it is necessary to increase the diversity of training data through data enhancement technology. This method can generate multiple versions of translations, select the best version through scoring and sorting mechanisms, and enhance the diversity and quality of training data.
[0211] 3. Multilingual content generation
[0212] Content creation platform: Content creation platforms need to generate multilingual versions of content to cover a wider user base. This method can automatically generate high-quality multilingual content and improve the diversity and readability of the content.
[0213] Multilingual document generation: Enterprises or institutions need to generate multilingual versions of documents, such as product manuals, user guides, etc. This method can generate translation materials that meet professional requirements and improve the quality and consistency of documents.
[0214] 4. Translation Quality Assessment
[0215] Translation quality control: Translation companies or agencies need to evaluate and control translation quality. This method can score and sort translations through multiple models, provide objective evaluation results, and help improve translation quality.
[0216] Translation competitions: Translation competitions require scoring and ranking of participants’ translations. This method can automatically generate multiple versions of the translation and provide fair evaluation results through scoring and ranking mechanisms.
[0217] 5. Multilingual Education
[0218] Language teaching: Language teaching institutions need to generate diverse translation practice materials to help students improve their translation skills. This method can generate translation materials covering different scenarios and styles to enhance students' learning experience.
[0219] Online courses: Online education platforms need to generate multilingual versions of course materials to meet the needs of learners with different language backgrounds. This method can generate high-quality translation materials and improve the applicability and effectiveness of courses.
[0220] Specifically, for different scenarios, there are different translation requirements for the compilation model. Therefore, in different application scenarios, it is necessary to set different target text information according to the data processing method provided in the above embodiments of the present application, and then generate corresponding training corpus. The data processing method has been described in the above embodiments and will not be repeated here.
[0221] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0222] Based on the same inventive concept, the embodiment of the present application also provides a data processing device for implementing the data processing method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the one or more data processing device embodiments provided below can refer to the limitations on the data processing method above, and will not be repeated here.
[0223] In an exemplary embodiment, Fig. 9 As shown, a data processing device 900 is provided, comprising: a first generating module 901, a compiling module 902, a scoring module 903, and a second generating module 904, wherein:
[0224] A first generating module 901 is used to generate a source text based on a pre-trained language model and target text information;
[0225] A compiling module 902, configured to compile the source text by referring to the compiling model to obtain a first target text;
[0226] A scoring module 903 is used to score the first target text using a scoring model to obtain a scoring result;
[0227] The second generating module 904 is used to generate a training corpus pair for preference optimization training based on the scoring result, and the training corpus pair is used to train the target compilation model.
[0228] In one embodiment, the data processing device 900 further includes:
[0229] An acquisition module is used to obtain an initial target compilation model;
[0230] A supervised fine-tuning module is used to perform supervised fine-tuning on the initial compilation model to obtain a target compilation model after supervised fine-tuning;
[0231] The training module is used to perform preference optimization training on the supervised fine-tuned target compilation model based on the training corpus to obtain a trained target compilation model; wherein the trained target compilation model is used to implement the compilation task corresponding to the target text information.
[0232] In one embodiment, the supervised fine-tuning module is specifically used to select a target compiled translation with the highest score result from multiple versions of compiled translations;
[0233] Based on the target compiled translation and the source text, the initial target compilation model is trained to obtain the supervised fine-tuned target compilation model.
[0234] In one embodiment, the first generation module 901 is used to input the target text information into the pre-trained language model. The target text information includes the text content information, style form and language conversion type compiled by the target compilation model;
[0235] The text content information, style and language conversion type are processed through the pre-trained language model to obtain corpus material information that meets the requirements of the target text information;
[0236] The corpus material information is processed according to the information synthesis algorithm in the pre-trained language model to obtain the source text.
[0237] In one embodiment, the first target text includes compiled translations output by various compilation models for the same source text, and the compilation module 902 is specifically configured to obtain multiple reference compilation models;
[0238] The source text is compiled by using multiple reference compilation models to obtain compiled translations output by the reference compilation models for the same source text.
[0239] In one embodiment, the first target text includes compiled translations output by various compilation models for the same source text, and the scoring module 903 is specifically configured to score the compiled translations using multiple scoring models to obtain an initial scoring result of the compiled translations;
[0240] Based on the initial scoring results obtained by the same scoring model, the compiled translations are sorted to obtain the scoring sorting results corresponding to the compiled translations;
[0241] Based on the score ranking results of the same scoring model corresponding to each of the compiled translations, each compiled translation is re-scored to obtain the score results of the compiled translations by multiple scoring models.
[0242] In one embodiment, the second generating module 904 is specifically used to pair the compiled translations of multiple versions in pairs to obtain compiled translation pairs;
[0243] Comparing the scoring results corresponding to the compiled translations in each compiled translation pair, and adding a preference tag to each compiled translation in the compiled translation pair based on the comparison result, to obtain an updated compiled translation pair;
[0244] The source text is added to the compiled translation in the updated compiled translation pair to generate a training corpus pair for preference optimization training.
[0245] Each module in the above data processing device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in an electronic device in the form of hardware, or can be stored in a memory in an electronic device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0246] In an exemplary embodiment, an electronic device is provided. The electronic device may be a terminal, and its internal structure diagram may be as shown in FIG. Fig.10 As shown. The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and an external device. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a data processing method is implemented. The display unit of the electronic device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the electronic device casing, or an external keyboard, touchpad or mouse.
[0247] Those skilled in the art will understand that Fig.10 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.
[0248] In an exemplary embodiment, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0249] Generate source text based on the pre-trained language model and target text information;
[0250] Compile the source text by referring to the compilation model to obtain a first target text;
[0251] Scoring the first target text through the scoring model to obtain a scoring result;
[0252] Based on the scoring results, training corpus pairs for preference optimization training are generated, and the training corpus pairs are used to train the target compilation model.
[0253] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0254] Get the initial target compilation model;
[0255] Perform supervised fine-tuning on the initial target compilation model to obtain a supervised fine-tuned target compilation model;
[0256] Based on the training corpus, preference optimization training is performed on the supervised fine-tuned target compilation model to obtain a trained target compilation model; the trained target compilation model is used to implement the compilation task corresponding to the target text information.
[0257] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0258] Among multiple versions of compiled translations, select the target compiled translation with the highest score;
[0259] Based on the target compiled translation and the source text, the initial target compilation model is trained to obtain the supervised fine-tuned target compilation model.
[0260] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0261] Inputting target text information into the pre-trained language model, the target text information including text content information, stylistic form, and language conversion type compiled by the target compilation model;
[0262] The text content information, style and language conversion type are processed through the pre-trained language model to obtain corpus material information that meets the requirements of the target text information;
[0263] The corpus material information is processed according to the information synthesis algorithm in the pre-trained language model to obtain the source text.
[0264] In one embodiment, the first target text includes compiled translations output by various compilation models for the same source text, and the processor further implements the following steps when executing the computer program:
[0265] Obtain multiple reference compiled models;
[0266] The source text is compiled by using multiple reference compilation models to obtain compiled translations output by the reference compilation models for the same source text.
[0267] In one embodiment, the first target text includes compiled translations output by various compilation models for the same source text, and the processor further implements the following steps when executing the computer program:
[0268] Scoring multiple versions of compiled translations through multiple scoring models to obtain initial scoring results of the compiled translations;
[0269] Based on the initial scoring results obtained by the same scoring model, the compiled translations are sorted to obtain multiple scoring ranking results corresponding to the compiled translations;
[0270] Based on the score ranking results of the same score model corresponding to each compiled translation, each compiled translation is re-scored to obtain the score results of the multiple score models for each compiled translation.
[0271] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0272] Pair the compiled translations of multiple versions in pairs to obtain compiled translation pairs;
[0273] Comparing the scoring results corresponding to the compiled translations in each compiled translation pair, and adding a preference tag to each compiled translation in the compiled translation pair based on the comparison result, to obtain an updated compiled translation pair;
[0274] The source text is added to the compiled translation in the updated compiled translation pair to generate a training corpus pair for preference optimization training.
[0275] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0276] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0277] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0278] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0279] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0280] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A data processing method, characterized in that: The method comprises: Generate source text based on the pre-trained language model and target text information; Compile the source text by referring to the compilation model to obtain a first target text; Scoring the first target text by using a scoring model to obtain a scoring result; Based on the scoring result, a training corpus pair for preference optimization training is generated, and the training corpus pair is used to train a target compilation model.
2. The method according to claim 1, characterized in that: After generating a training corpus pair for preference optimization training based on the scoring result, the method further includes: Get the initial target compilation model; Performing supervised fine-tuning on the initial target compilation model to obtain a supervised fine-tuned target compilation model; Based on the training corpus, preference optimization training is performed on the supervised fine-tuned target compilation model to obtain a trained target compilation model; the trained target compilation model is used to implement the compilation task corresponding to the target text information.
3. The method according to claim 2, characterized in that The supervised fine-tuning of the initial target compilation model to obtain the supervised fine-tuned target compilation model includes: Among multiple versions of compiled translations, select the target compiled translation with the highest score; Based on the target compiled translation and the source text, the initial target compiled model is trained to obtain a supervised fine-tuned target compiled model.
4. The method according to claim 1, characterized in that: The generating of the source text based on the pre-trained language model and the target text information includes: Inputting target text information into a pre-trained language model, the target text information including text content information, stylistic form, and language conversion type compiled by the target compilation model; Processing the text content information, the stylistic form and the language conversion type through the pre-trained language model to obtain corpus material information that meets the requirements of the target text information; The corpus material information is processed according to the information synthesis algorithm in the pre-trained language model to obtain the source text.
5. The method according to claim 1, characterized in that The first target text includes compiled translations output by each compilation model for the same source text, and the source text is compiled by referring to the compilation model to obtain the first target text, including: Obtain multiple reference compiled models; The source text is compiled by using the multiple reference compilation models to obtain compiled translations output by the reference compilation models for the same source text.
6. The method according to claim 1, characterized in that The first target text includes compiled translations output by each compilation model for the same source text, and the scoring of the first target text by the scoring model to obtain a scoring result includes: Scoring the compiled translation using the plurality of scoring models to obtain an initial scoring result of the compiled translation; Based on the initial scoring results obtained by the same scoring model, the compiled translations are sorted to obtain the scoring sorting results corresponding to the compiled translations; Based on the score ranking results of the same score model corresponding to each of the compiled translations, each of the compiled translations is re-scored to obtain score results of the multiple score models for each of the compiled translations.
7. The method according to claim 1, characterized in that The step of generating a training corpus pair for preference optimization training based on the scoring result includes: Pair the compiled translations of multiple versions in pairs to obtain compiled translation pairs; Comparing the scoring results corresponding to the compiled translations in each compiled translation pair, and adding a preference tag to each compiled translation in the compiled translation pair based on the comparison result, to obtain an updated compiled translation pair; The source text is added to the compiled translation in the updated compiled translation pair to generate a training corpus pair for preference optimization training.
8. A data processing device, characterized in that: The device comprises: A first generation module, used to generate a source text based on a pre-trained language model and target text information; A compiling module, used for compiling the source text by referring to the compiling model to obtain a first target text; A scoring module, used to score the first target text through a scoring model to obtain a scoring result; The second generating module is used to generate a training corpus pair for preference optimization training based on the scoring result, and the training corpus pair is used to train the target compilation model.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.