Text data processing method and device, computer equipment, readable storage medium and program product
By updating the weight of the basic text model to match the target language type, the problem of poor accuracy in multilingual text processing is solved, and higher text processing accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510105544.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
Existing text processing techniques have poor accuracy in multilingual text processing, especially when dealing with language pairs with large language differences.
By obtaining the target language type of the target text and its corresponding model weights, update the weight of the trained basic text model, obtain the target text model, and process the target text through the model.
It improves the accuracy of multilingual text processing, enhances the generalization ability of the model to adapt to different languages, and can quickly generate high-quality text processing results in offline state.
Smart Images

Figure CN120012770A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of machine learning, and in particular to a text data processing method, apparatus, computer equipment, readable storage medium and program product. Background Art
[0002] With the deepening of globalization and the popularity of mobile devices such as smartphones, machine learning models are needed to assist in text processing, and machine learning models have become one of the indispensable applications in daily life. However, existing text processing technologies face some significant limitations. Most current text processing solutions rely on cloud services to complete text processing tasks. When dealing with multilingual text processing, cloud text processing solutions often use a common model to process all language pairs. However, due to the significant differences in grammar, vocabulary, and semantics between different languages, it is difficult for a single model to achieve uniform high-quality output in multilingual text processing tasks. The text processing effect in a multilingual environment is not ideal, especially when dealing with language pairs with large language differences, resulting in poor accuracy of text processing results. Summary of the invention
[0003] Based on this, it is necessary to provide a text data processing method, apparatus, computer device, computer-readable storage medium and computer program product that can improve the accuracy of multiple languages in response to the above technical problems.
[0004] In a first aspect, the present application provides a text data processing method, comprising:
[0005] Obtaining a target language type corresponding to the target text, and determining a model weight corresponding to the target language type;
[0006] The weight of the trained basic text model is updated by the model weight to obtain a target text model; the trained basic text model is obtained by training with a target data set and a target training method;
[0007] The target text is processed by the target text model to obtain a target text corresponding to the target language type.
[0008] In one embodiment, the method further comprises:
[0009] Acquire a first target corpus, where the first target corpus includes a first source text and a first target text corresponding to the first source text;
[0010] Combining a plurality of corpus style instructions with each of the first target corpora respectively to obtain a plurality of data groups, each of the data groups comprising a corpus style instruction, a first source text and a first target text, each of the corpus style instructions being obtained by diffusion processing and screening processing of the initially configured corpus style instruction;
[0011] For each data group, in a preset large language model, data processing is performed on the first source text based on the corpus style instruction to obtain a second target text; modification suggestion data is determined based on the second target text, and a target text corresponding to the first source text is obtained through the modification suggestion data, the first target text and the second target text;
[0012] A target data set is obtained based on the corpus style instructions, the first source text and the target text respectively corresponding to each of the data groups.
[0013] In one embodiment, the method further comprises:
[0014] Extracting an initial text processing corpus from a public data set, wherein the initial text processing corpus includes an initial source text and an initial processing text corresponding to the initial source text;
[0015] Each of the initial text processing corpora is screened from a target screening dimension to obtain a first target corpus, wherein the target screening dimension includes one or more of a text length dimension, a text structure dimension, and a grammatical correctness dimension.
[0016] In one embodiment, the modification suggestion data includes a first modification suggestion; and determining the modification suggestion data based on the second target text, and obtaining a target text corresponding to the first source text through the modification suggestion data, the first target text, and the second target text, includes:
[0017] Inputting the second target text into the large language model, determining whether the corpus style of the second target text matches the corpus style in the corpus style instruction through the large language model, obtaining corpus style difference data, and processing the first source text and the second target text through the large language model to obtain content modification data; obtaining a first modification suggestion based on the difference data and the content modification data;
[0018] The second target text is modified based on the first modification suggestion through the large language model to obtain a third target text; and a target text is obtained based on the third target text.
[0019] In one embodiment, the modification suggestion data further includes a second modification suggestion, and obtaining the target text based on the third target text includes:
[0020] Inputting a first target text and a third target text corresponding to the first source text into the large language model, evaluating the third target text based on the target evaluation dimension and the first target text, and obtaining an evaluation score and a second modification suggestion;
[0021] In the case where it is determined that the third target text does not meet the preset standard conditions, the third target text is improved based on the evaluation score and the second modification suggestion to obtain a modified third target text, and the evaluation of the third target text based on the target evaluation dimension and the first target text is re-executed to obtain an evaluation score and the second modification suggestion, until a third target text that meets the preset standard conditions is obtained; the target evaluation dimension includes one or more evaluation dimensions of fluency, accuracy, style consistency, and content completeness;
[0022] Based on the first target texts corresponding to the third target texts, the evaluation index values corresponding to the third target texts are calculated respectively, and the third target texts are screened based on the evaluation index values to obtain the target texts corresponding to the first source texts.
[0023] In one embodiment, the method further comprises:
[0024] Based on the corpus style instructions in the target data set, the first source text and the target text, the basic text model is pre-trained to obtain a character prediction probability corresponding to the first source text, wherein the character prediction probability represents a probability value of each character position being a character in a standard vocabulary;
[0025] A loss function is calculated based on the target text corresponding to the source text and the character prediction probability corresponding to the first source text, and a weight matrix of the basic text model is updated by the loss function until a preset training completion condition is met to obtain a pre-trained processing model;
[0026] The pre-trained processing model is trained by a target training method to obtain a trained basic text model, and the target training method includes one or more of supervised fine-tuning, distillation learning, reinforcement learning, and hyperparameter adjustment.
[0027] In one embodiment, the target training method includes any one of supervised fine-tuning, distillation learning, and reinforcement learning; the pre-trained processing model is trained by the target training method to obtain a trained basic text model, including:
[0028] Divide the target data set based on language type to obtain data subsets corresponding to each language type; for each language type, train the trained basic text model based on the data subset corresponding to the language type to obtain the model weight corresponding to the language type; process the model weight of the pre-trained basic text model based on the model weight corresponding to the language type to obtain the trained basic text model; or
[0029] The corpus style instructions, the first source text and the target text in the target data set are processed by the pre-trained processing model to perform text prediction processing, and the character prediction probability corresponding to the first source text is obtained, and the character prediction probability represents the probability value of each character position being a character in the target vocabulary; based on the target text corresponding to the source text and the character prediction probability corresponding to the first source text, a trained basic text model is obtained, and the target vocabulary is obtained after the preset large model processes the target data set; or
[0030] The pre-trained processing model is used to perform text prediction processing on the corpus style instructions, the first source text and the target text in the target data set to obtain the character prediction probability corresponding to the first source text, and the character prediction probability represents the probability value of each character position being a character in the standard vocabulary; based on the target text corresponding to the source text, the character prediction probability corresponding to the first source text and the environmental feedback data, a trained basic text model is obtained, and the environmental feedback data is the evaluation data of the external environment on the target text.
[0031] In one embodiment, the method further comprises:
[0032] Dividing the target data set based on language type to obtain data subsets corresponding to each language type;
[0033] The trained basic text model is processed based on the data subset corresponding to each of the language types to obtain the model weight corresponding to each of the language types.
[0034] In one embodiment, updating the weight of the trained basic text model by using the model weight to obtain the target text model includes:
[0035] The model weight corresponding to the target language type and the model weight of the trained basic text model are combined to obtain a target model weight;
[0036] The model weights in the trained basic text model are replaced by the target model weights to obtain a target text model.
[0037] A text data processing device, the device comprising:
[0038] A first acquisition module, used to acquire a target language type corresponding to a target text, and determine a model weight corresponding to the target language type;
[0039] A first updating module is used to update the weight of the trained basic text model by using the model weight to obtain a target text model; the trained basic text model is obtained by training with a target data set and a target training method;
[0040] The first translation module is used to process the target text through the target text model to obtain a target text corresponding to the target language type.
[0041] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0042] Obtaining a target language type corresponding to the target text, and determining a model weight corresponding to the target language type;
[0043] The weight of the trained basic text model is updated by the model weight to obtain a target text model; the trained basic text model is obtained by training with a target data set and a target training method;
[0044] The target text is processed by the target text model to obtain the target text corresponding to the target language type.
[0045] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0046] Obtaining a target language type corresponding to the target text, and determining a model weight corresponding to the target language type;
[0047] The weight of the trained basic text model is updated by the model weight to obtain a target text model; the trained basic text model is obtained by training with a target data set and a target training method;
[0048] The target text is processed by the target text model to obtain a target text corresponding to the target language type.
[0049] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:
[0050] Obtaining a target language type corresponding to the target text, and determining a model weight corresponding to the target language type;
[0051] The weight of the trained basic text model is updated by the model weight to obtain a target text model; the trained basic text model is obtained by training with a target data set and a target training method;
[0052] The target text is processed by the target text model to obtain a target text corresponding to the target language type.
[0053] The above-mentioned text data processing method, device, computer equipment, readable storage medium and program product, wherein the method may include: obtaining the target language type corresponding to the target text, and determining the model weight corresponding to the target language type; updating the weight of the trained basic text model by the model weight to obtain the target text model; the trained basic text model is obtained by training with the target data set and the target training method; processing the target text by the target text model to obtain the target text corresponding to the target language type. By adopting this method, the model can be trained by multiple training methods and high-quality data sets to ensure the accurate text processing effect of the model, improve the generalization ability of the model and the adaptability to different languages, and after determining the model weight corresponding to the target language type of the target text and merging it with the pre-trained processing model for data processing, high-quality text processing results can be quickly generated in an offline state to ensure the flexibility of multi-language text processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0055] Figure 1 A flowchart of a text data processing method in one embodiment;
[0056] Figure 2 A schematic diagram of a process for obtaining a target data set in one embodiment;
[0057] Figure 3 A schematic diagram of a process of obtaining a first target text in an embodiment;
[0058] Figure 4 A schematic diagram of a process for obtaining a target text in an embodiment;
[0059] Figure 5 A schematic diagram of a process for obtaining a target text in an embodiment;
[0060] Figure 6 A schematic diagram of a process for obtaining a basic text model in one embodiment;
[0061] Figure 7 A schematic diagram of a process for obtaining a target text model in one embodiment;
[0062] Figure 8 is a structural block diagram of a text data processing device in one embodiment;
[0063] Fig. 9 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0065] Machine Translation (MT) has become one of the indispensable applications in daily life. However, existing translation technologies face some significant limitations. Most current translation solutions rely on cloud services to complete translation tasks. When dealing with multilingual translation, cloud translation solutions often use a common model to handle all language pairs. However, due to the significant differences in grammar, vocabulary and semantics of different languages, it is difficult for a single model to achieve uniform high-quality output in multilingual translation tasks. The translation effect in a multilingual environment is not ideal, especially when dealing with language pairs with large language differences, resulting in poor accuracy of translation results. The translation method in traditional technology is generally cloud translation, but cloud translation is highly dependent on network connection. In the case of poor network conditions or inability to connect to the Internet, the speed and accuracy of translation are greatly reduced, which seriously affects the user experience. Secondly, cloud translation requires users to upload data to the server, which raises concerns about privacy and data security. For users who need to handle sensitive information, such privacy risks often limit the applicability of cloud translation.
[0066] Secondly, traditional technologies often use a universal model to process all language pairs. However, due to the significant differences in grammar, vocabulary and semantics of different languages, it is difficult for a single model to achieve uniform high-quality output in multilingual translation tasks, resulting in unsatisfactory translation results in multilingual environments, especially when dealing with language pairs with large language differences. The accuracy and naturalness of the translation results need to be improved.
[0067] The text data processing method provided in the embodiment of the present application can adjust the corpus style of the output translation text based on user needs, and can adapt to different scenarios and needs, using different corpus styles and tones. For example, in formal occasions, the translation text obtained by the text data processing method provided in the embodiment of the present application can be a rigorous expression, and the translation text output in daily conversation scenarios can be a colloquial expression. It can flexibly adapt to the needs of multiple scenarios and multiple corpus styles, avoid the limitations of translation applications, and expand the scope of application.
[0068] In summary, the text data processing method provided in the embodiment of the present application can avoid dependence on the network by training an offline translation model, and can also improve the timeliness of translation processing. Local data processing can also improve data security, avoid privacy risks and hardware performance limitations of mobile devices, and improve multi-language translation effects. Since the basic text model and the target text model in the embodiment of the present application are trained based on training data containing corpus style, the text data processing method provided in this embodiment can also flexibly adjust the corpus style of the output translation text according to different needs, achieve a balance between translation performance and security, and improve the flexibility and quality of multi-language translation.
[0069] In one embodiment, Figure 1 As shown, a text data processing method is provided. This embodiment uses the method applied to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The above-mentioned terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, projection devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. In this embodiment, the text data processing method includes the following steps:
[0070] Step 102: Obtain a target language type corresponding to the target text, and determine a model weight corresponding to the target language type.
[0071] Among them, the target text can be the text content that the user currently needs to translate, the target text can also be the text content that the user currently needs to extract text features or text encoding, and can also be the text content that the user currently needs to perform text segmentation processing; the following is a detailed description of the method in this embodiment applied to the text translation scenario. Correspondingly, the target language type corresponding to the target text can be the language type to be translated of the target text, and the language type of the target text itself is different from the target language type. For example, the language type of the target text itself can be the source language type of the target text, such as Chinese, English, French, etc. The model weight corresponding to the target language type can be the weight feature matrix corresponding to the target language type. The terminal can store model weights corresponding to multiple language types locally, and the terminal can be obtained after training based on the data subsets corresponding to the various language types in the target data set.
[0072] Specifically, the terminal obtains the target text and the target language type corresponding to the target text, that is, the language type into which the target text needs to be processed, and the target language type may be the language type to be processed selected by the user of the provided target text. In this way, the terminal can obtain the model weight (model weight matrix) corresponding to the target language type, for example, by querying in a weight database to obtain the model weight matching the target language type.
[0073] Optionally, the terminal obtains the target text and the target language type corresponding to the target text, that is, the language type to which the target text needs to be translated, and the target language type may be the language type to be translated selected by the user of the provided target text. In this way, the terminal can obtain the model weight (model weight matrix) corresponding to the target language type, for example, it may be queried in the weight database to obtain the model weight matching the target language type, and the weight database may store lightweight model weights corresponding to a plurality of language types, respectively, and the lightweight model weight may be a weight based on LoRA (Low-Rank Adaptation). The terminal may process the basic text model based on the target data set to obtain the model weights corresponding to each language type, respectively, and the target data set may be a data set obtained by screening and processing the translation corpus obtained based on the public data set, which is a relatively high-quality data set, and the process of obtaining the specific target data set will be described in detail in the subsequent embodiments, and will not be repeated here.
[0074] Optionally, in response to the user's trigger operation of the target language type, the terminal can determine the target language type corresponding to the trigger operation, that is, obtain the type to be translated input by the user to the terminal. At the same time, the terminal will query and load the model weight corresponding to the target language type in the weight database or the local storage space of the terminal. Ensure that the model weight corresponding to the target language type is obtained in a timely manner to ensure low latency for users during the translation process.
[0075] Step 104, updating the weight of the trained basic text model by using the model weight to obtain the target text model.
[0076] Among them, the trained basic text model is obtained after training with the target data set and the target training method.
[0077] Specifically, the terminal can update the weight of the pre-trained basic text model based on the model weight corresponding to the target language type to obtain the basic text model after the weight update, that is, the target text model. For example, after obtaining the model weight corresponding to the target language type, the terminal can merge the model weight with the weight of the trained basic text model to obtain the merged weight, and replace the model weight of the trained basic text model based on the merged weight to obtain the target text model, which is a text processing model that matches the target language type.
[0078] Optionally, the trained basic text model may be a basic translation model, and the target text model may be a target translation model, which may be a text translation model that matches the target language type. The terminal may update the weight of the pre-trained basic translation model based on the model weight corresponding to the target language type to obtain a basic translation model after the weight update, that is, to obtain a target translation model. For example, after obtaining the model weight corresponding to the target language type, the terminal may merge the model weight with the weight of the trained basic translation model to obtain a merged weight, and replace the model weight of the trained basic translation model based on the merged weight to obtain a target translation model.
[0079] Step 106: Process the target text using the target text model to obtain a target text corresponding to the target language type.
[0080] Specifically, the terminal can input the target text into the target text model, perform data processing on the target text through the target text model, and obtain a target text corresponding to the target language type; that is, the terminal can process the target text through the target text model, that is, process the language type of the target text, and obtain the target language type of the target text, that is, obtain the target text of the target language type.
[0081] Optionally, the target text may be a text to be translated. The terminal may input the text to be translated into a target translation model, and translate the text to be translated through the target translation model to obtain a target translation text corresponding to the text to be translated. The language type of the target translation text is the target language type, thereby ensuring the real-time output of the translation result.
[0082] Optionally, the target text may be a text to be encoded, and the target text model may be a target text encoding model; the terminal may input the text to be encoded into the target text model, encode the text to be encoded through the target text model, and obtain a target encoded text corresponding to the text to be encoded, and the language type of the target encoded text is the target language type.
[0083] In one example, the terminal can query the model weight corresponding to the target language type in the cloud in real time and download it to the terminal locally. In another example, the cloud server can be trained based on the target data set to obtain a trained basic text model, and the basic text model can be processed based on the data subsets corresponding to the various language types in the target data set to obtain the model weights corresponding to each language type; in this way, the terminal can download the model weights corresponding to each language type from the cloud server, and store the model weights corresponding to each language type in the local storage space; then the terminal can query the model weight corresponding to the target language type in the local storage space in an offline state or when not connected to the network, to achieve offline translation on the terminal side, and ensure the timeliness and stability of the translation processing. In other words, the terminal can realize lightweight model weight loading, store the trained text translation model separately, and the model weights corresponding to each language type respectively. After determining the target language type, the terminal can load the lightweight weight of the target language type in real time, and dynamically merge it with the trained text translation model, which can effectively reduce the model's memory usage and avoid the need to independently train and store large models for each language pair. Through hierarchical storage, model weight loading and merging can be achieved in milliseconds, ensuring low latency in user experience.
[0084] In another example, the MLC-LLM framework in the text data processing method provided in this embodiment uses ApacheTVM and microTVM as compilers to optimize the neural network reasoning process and adapt to mobile devices of various hardware architectures. The reasoning process is significantly accelerated through automated calculation graph optimization and hardware-specific optimization. For example, the MLC-LLM framework automatically selects appropriate parallel computing strategies, such as tensor splitting, memory management optimization, etc., according to the computing power, memory, and bandwidth of the device, thereby reducing the reasoning time, and the terminal can significantly reduce memory usage without significantly reducing the quality of translation. At the same time, the dynamic memory allocation and management mechanism in the reasoning process will release unnecessary memory usage in a timely manner according to the current task size, reduce redundant calculations, and the terminal can also adjust the frequency of the CPU and GPU according to the real-time reasoning load to ensure that power consumption is reduced while maintaining performance, and ensure that the end side completes long-term translation tasks.
[0085] In another example, the terminal can also complete the processing of user input data only through the local device until the translation result is obtained, avoiding the risk of sensitive information needing to be transmitted over the network. It is suitable for application scenarios with high privacy requirements, such as personal private communications or confidential document translation, etc.
[0086] In the above text data processing method, the target language type corresponding to the target text is obtained, and the model weight corresponding to the target language type is determined; the weight of the trained basic text model is updated by the model weight to obtain the target text model; the trained basic text model is obtained by training with the target data set and the target training method; the target text is processed by the target text model to obtain the target text corresponding to the target language type. By adopting this method, the model can be trained by multiple training methods and high-quality data sets to ensure the accurate text processing effect of the model, improve the generalization ability of the model and the adaptability to different languages, and after the model weight corresponding to the target language type of the target text is determined and merged with the pre-trained processing model for data processing, high-quality text processing results can be quickly generated in an offline state to ensure the flexibility of multi-language text processing.
[0087] That is to say, the above-mentioned text data method can obtain the target language type corresponding to the target text, and determine the model weight corresponding to the target language type; update the weight of the trained basic text model through the model weight to obtain the target text model; the trained basic text model is obtained after training through the target data set and the target training method; the target text is processed through the target text model to obtain the translation text of the target language type corresponding to the target text. By adopting this method, the model can be trained through a variety of training methods and high-quality data sets to ensure the accurate effect of model translation, improve the generalization ability of the model and the degree of adaptability to different languages, and after determining the model weight corresponding to the target language type of the target text and merging it with the pre-trained translation model for data processing, high-quality translation results can be quickly generated in an offline state to ensure the flexibility of multi-language translation.
[0088] In one embodiment, Figure 2 As shown, the text data processing method also includes:
[0089] Step 202: Obtain a first target corpus.
[0090] The first target corpus includes a first source text and a first target text corresponding to the first source text. The first target corpus may be determined based on a plurality of public data sets, the first source text may be a target text, and the first target text corresponding to the first source text is a translation text obtained by translating the first source text determined based on the public data sets. Each public data set may include data sets in a plurality of technical fields, such as news fields, social media fields, and technical document fields, etc.
[0091] Optionally, the first target corpus may be a first translation corpus, and the first target text contained in the first translation corpus may be a first translation text corresponding to the first source text; the first target corpus may be determined based on multiple public data sets, the first source text may be a text to be translated, and the first translation text corresponding to the first source text is a translation text obtained by translating the first source text determined based on the public data sets.
[0092] Specifically, the terminal may obtain multiple first target corpora based on public data sets corresponding to multiple fields, that is, obtain multiple first source texts and first target texts corresponding to the first source texts based on multiple public data.
[0093] Step 204 , combining the various corpus style instructions with the first target corpora respectively to obtain various data sets.
[0094] Among them, each data group includes a corpus style instruction, a first source text and a first target text. Each corpus style instruction is obtained after diffusion processing and screening processing of the initially configured corpus style instruction. The corpus style of the corpus in the public data set may be different. The corpus style instruction may represent the style represented by the translation text, that is, the instruction representing what kind of corpus style is required to be generated; for example, it may be a business corpus style, a spoken corpus style, a written corpus style, etc. The business corpus style represents translating the source text into a business style, and the translation text corresponding to the source text needs to be in a business style. The corpus style may be a translation style, and the corpus style instruction may be a translation style instruction.
[0095] Specifically, the terminal can obtain multiple corpus style instructions pre-configured manually, and the terminal can perform diffusion processing on the multiple corpus style instructions pre-configured manually based on the instruction diffusion model, and generate multiple diffused corpus style instructions by imitating the multiple corpus style instructions pre-configured through the instruction diffusion model. In this way, the terminal can screen the multiple corpus style instructions pre-configured manually and the multiple diffused corpus style instructions to obtain corpus style instructions that meet the style screening conditions. The instruction diffusion model can be a gpt-4o model, and the style screening conditions can be reasonable and commonly used corpus style instructions. Optionally, the specific screening process can be that the terminal can determine the frequency of use of each corpus style instruction based on statistical data, and determine the rationality of each corpus style instruction based on a semantic algorithm, and determine the corpus style instruction whose use frequency is greater than a preset use frequency threshold and whose rationality is greater than a preset rationality threshold as the corpus style instruction that meets the style screening conditions.
[0096] In this way, after obtaining multiple corpus style instructions, the terminal can randomly combine each corpus style instruction with each first target corpus to obtain each data group, and each data group can contain the first source text, the first target text corresponding to the first source text, and the corpus style instruction. In other words, the terminal can combine each corpus style instruction with each first target corpus to obtain multiple data pairs, that is, when the number of first corpus style instructions is the first number and the number of first target corpora is the second number, then the number of data pairs obtained by the final combination can be the product of the first number and the second number.
[0097] Step 206: for each data group, in a preset large language model, the first source text is processed based on the corpus style instruction to obtain a second target text. Modification suggestion data is determined based on the second target text, and a target text corresponding to the first source text is obtained by modifying the suggestion data, the first target text, and the second target text.
[0098] Among them, the preset large language models can be deep learning models with a large number of parameters. In the field of natural language processing (NLP), language patterns, grammar and semantics are learned by processing large amounts of text data to understand and generate human language.
[0099] Specifically, for each data group, the terminal can obtain a preset large language model, and use the large language model to process the first source text according to the corpus style instructions in the data group, that is, the first source text is processed into the target language type according to the corpus style instructions in the data group to obtain the second target text. The terminal can re-input the second target text into the large language model, verify the second target text through the large language model, and obtain the modification suggestion data corresponding to the second target text output by the large language model. In this way, the second target text can be modified and processed according to the modification suggestion data and the original text to be processed (first target text) through the large language model to obtain the modified text to be processed, that is, the target text corresponding to the first source text.
[0100] Optionally, for each data group, the terminal can obtain a preset large language model, and use the large language model to translate the first source text according to the translation style instruction in the data group, that is, translate the first source text into the target language type according to the translation style instruction in the data group to obtain a second translation text. The terminal can re-input the second translation text into the large language model, verify the second translation text through the large language model, and obtain the modification suggestion data corresponding to the second translation text output by the large language model. In this way, the second translation text can be modified by the large language model according to the modification suggestion data and the original translation text (first translation text) to obtain a modified translation text, that is, a target translation text corresponding to the first source text.
[0101] Step 208 , obtaining a target data set based on the corpus style instructions, the first source text, and the target text corresponding to each data set.
[0102] Specifically, the terminal can determine the corpus style instructions, the first source text and the target text respectively contained in each data group as the target data set; or, the terminal can determine the first source text and the target text corresponding to the first source text respectively contained in each data group as the target data set; or, the terminal can evaluate the target text corresponding to the first source text in each data group based on the dimension of the target evaluation index, obtain the evaluation index value corresponding to each target text, screen each target text based on each evaluation index value, and determine the screened target text, as well as the corpus style instructions and / or the first source text corresponding to the target text as the target data set.
[0103] Optionally, the process of obtaining the target data set, that is, the steps of step 202 to step 208, can also be completed on the cloud server side. The cloud server can extract the initial text processing corpus from the public data set, and screen each of the initial text processing corpus from the target screening dimension to obtain the first target corpus, and combine each corpus style instruction with each first target corpus to obtain the target data set. In this way, the cloud server can pre-train the basic text model based on the target data set to obtain a pre-trained processing model, and train the pre-trained processing model through a target training method to obtain a trained basic text model and model weights corresponding to each language type. The terminal can download the trained basic text model and the model weights of each language type from the cloud server in advance. In this way, after the terminal determines the target language type, the terminal can query the model weight of the target language type locally when it is offline or not connected to the network, and process the trained basic text model through the model weight of the target language type to obtain the target text model. Among them, the pre-trained processing model can be a pre-trained translation model, and the trained basic text model is a trained basic translation model.
[0104] In one example, the terminal can update the model weights of each language type stored locally. The specific process of updating can be: the public data set can be continuously updated, the cloud server can process the updated public data set to obtain the target data set, and the cloud server can update the model weights of each language type obtained in the previous training process based on the target data set. The terminal can connect to the cloud server after a preset time interval, download the latest model weights corresponding to each language type in the current cloud server, and save them locally, which can ensure the balance between the speed of model weight update and the accuracy of the currently stored model weights. The terminal can also download the updated model weights corresponding to each language type in real time after detecting that the model weights corresponding to each language type at the cloud server are updated, and save them to the terminal locally. The terminal can also obtain the data related to the user generated in the process of obtaining the translated text by performing text translation processing based on the user's target text, so that the terminal can adjust the model weights corresponding to each language type stored locally based on the data, obtain the updated model weights corresponding to each language type, and store them locally in the terminal. By updating the model weights corresponding to each language type stored in the terminal in a variety of ways, it is possible to ensure the timeliness of updating the model weights when the terminal is online or offline, thereby improving the accuracy of text translation and improving user experience.
[0105] Based on this, the text data processing method provided in this embodiment can achieve accurate text translation in an offline state through the MLC-LLM framework, make full use of the hardware resources of the terminal, maintain efficient reasoning performance, and ensure that heavy dynamic loading and hardware optimization make the reasoning process almost have no delay. Combined with the dynamic merging of lightweight weights and local optimization reasoning, the translation effect is close to the large cloud model. Intelligent memory management and energy optimization technology ensure efficient use of device resources in the reasoning process. All data processing is completed locally, avoiding the risk of privacy leakage. It can adapt to the offline translation needs of a variety of mobile devices and ensure that users can still obtain high-quality translation results close to real-time in an environment without a network connection.
[0106] In other words, the MLC-LLM framework supports parallel and real-time reasoning for multiple language pairs. When the user enters text in the translation application, the system will quickly load the applicable lightweight weights based on the input source language and the selected target language, and start the reasoning engine to generate the translation. Due to the adaptability of the framework to multiple languages, the system can efficiently handle translation tasks for different language pairs and ensure reasoning efficiency when switching between multiple language pairs. During the reasoning process, the MLC-LLM architecture ensures the real-time nature of the translation task through the parallel execution of multiple layers of transformers. Even in the case of longer text input, the translation results can be returned in a very short time to meet the user's usage needs.
[0107] In this embodiment, by pre-configuring multiple corpus style instructions and diffusing the corpus style instructions to generate more corpus style instructions, it is possible to ensure the stylization uniformity in the obtained target data set, avoid style differences, and provide reliable training data for subsequent models.
[0108] In one embodiment, Figure 3 As shown, the text data processing method also includes:
[0109] Step 302: extracting initial text processing corpus from the public data set.
[0110] The initial text processing corpus includes the initial source text and the initial processing text corresponding to the initial source text. The public data set may include the NLLB data set and the OPUS data set. Optionally, the initial text processing corpus may be an initial translation corpus, and the initial processing text corresponding to the initial source text may be a translation text corresponding to the initial source text obtained in the public data set.
[0111] Specifically, the terminal may extract multiple initial text processing corpora from multiple public data sets corresponding to multiple technical fields, that is, extract multiple initial source texts and extract initial processing texts corresponding to each initial source text.
[0112] Step 304 , screening each initial text processing corpus from the target screening dimension to obtain a first target corpus, where the first target corpus includes a first source text and a first target text corresponding to the first source text.
[0113] The target screening dimension includes one or more of the text length dimension, the text structure dimension, and the grammatical correctness dimension. The text length represents the number of characters of the source text and the translated text (processed text), the grammatical correctness dimension represents the degree of matching between the processed text and the grammatical norm, and the text structure dimension represents the integrity of the sentence structure of the source text and the processed text. The processed text may be the translated text after translation processing.
[0114] Specifically, the terminal can perform quality screening on each initial text processing corpus to obtain a first target corpus, that is, to obtain a first source text and a first target text corresponding to the first source text. Optionally, the terminal can first determine a target screening dimension, and can select one or more dimensions from a text length dimension, a text structure dimension, and a grammatical correctness dimension to determine as the target screening dimension. For example, the currently selected target screening dimension may include a text length dimension, a text structure dimension, and a grammatical correctness dimension.
[0115] In this way, for each public data set, the terminal can obtain the initial text processing corpus, and the terminal performs quality screening on the initial text processing corpus. The specific process of obtaining the first target corpus may include: the terminal can segment the initial source text through the tokenizer of the open source model (such as the Qwen model) to obtain multiple character tokens, obtain the number of characters of each initial source text, and obtain the character number screening range based on the preset minimum character number threshold and the preset maximum character number threshold. The terminal can eliminate the initial text processing corpus whose number of characters of the source text or the processed text corresponding to the source text does not meet the character number screening range, and obtain the initial text processing corpus that completes the text length dimension screening. Based on this, the terminal can use the preset perplexity quantification algorithm to calculate the perplexity of the initial processed text in each initial text processing corpus that completes the text length dimension screening, and eliminate the initial text processing corpus with a perplexity higher than the preset perplexity threshold to obtain the first target corpus that completes the quality screening. The perplexity of the translation corpus is negatively correlated with the completeness of the sentence structure of the translation corpus and the degree of matching with the preset grammatical norms. That is, the smaller the perplexity, the more complete the sentence structure representing the translation corpus, and the more the translated text in the translation corpus conforms to the preset grammatical norms.
[0116] In this embodiment, the quality of the initial text processing corpus is screened through the dimensions of text length, text structure, and grammatical correctness to improve the quality of the data set and provide a stable data foundation for subsequent model training.
[0117] In an exemplary embodiment, the modification suggestion data includes a first modification suggestion. Figure 4 As shown, the specific implementation process of the step of "determining the modification suggestion data based on the second target text, and obtaining the target text corresponding to the first source text by modifying the modification suggestion data, the first target text and the second target text" may include:
[0118] Step 402: Input the second target text into the large language model, determine whether the corpus style of the second target text matches the corpus style in the corpus style instruction through the large language model, obtain corpus style difference data, and process the first source text and the second target text through the large language model to obtain content modification data. Based on the difference data and the content modification data, obtain a first modification suggestion.
[0119] Specifically, the terminal can process the first source text according to the corpus style instruction in the data group through the large language model, that is, translate the first source text into the target language type according to the corpus style instruction in the data group to obtain the second target text. Based on this, the terminal can input the second target text into the large language model, determine whether the corpus style indicated by the corpus style instruction in the data group matches the second target text through the large language model, obtain the difference data of the corpus style, and determine the grammatical error data in the second target text through the large language model, and determine the completeness of the translation of the second target text for the first source text, the correctness of the interpretation, and the matching degree of the second target text and the first source text in content through the large language model and the first source text corresponding to the second target text, and obtain the first modification suggestion based on the above-mentioned corpus style difference data, grammatical error data, completeness, correctness of the interpretation, and the matching degree of the second target text and the first source text in content.
[0120] Among them, the second target text can be a second translated text obtained by translating the first source text based on the translation style instruction; the first target text can be the translated text corresponding to the first source text in the data set, that is, the first translated text; the corpus style instruction can be a translation style instruction.
[0121] Optionally, the terminal can translate the first source text according to the translation style instruction in the data group through the large language model, that is, translate the first source text into the target language type according to the translation style instruction in the data group to obtain the second translated text. Based on this, the terminal can input the second translated text into the large language model, determine whether the second translated text matches the translation style indicated by the translation style instruction in the data group through the large language model, obtain translation style difference data, and determine the grammatical error data in the second translated text through the large language model, and determine the completeness of the second translated text for the first source text translation, the correctness of the interpretation, and the content matching degree of the second translated text and the first source text through the large language model and the first source text corresponding to the second translated text, and obtain the first modification suggestion based on the above translation style difference data, grammatical error data, completeness, correctness of the interpretation, and the content matching degree of the second translated text and the first source text.
[0122] That is to say, the second translated text is checked through the large language model. For example, it can be checked whether the translation style of the second translated text is consistent with the translation style in the translation style instruction in the data group to obtain translation style difference data, and the large language model outputs grammatical errors (grammatical error data), inexpressiveness (the degree of content matching between the second translated text and the first source text), over-interpretation (the correctness of the interpretation) and omissions (the degree of completeness) in the second translated text, and a first modification suggestion is obtained based on the above data.
[0123] Step 404: Using the large language model, modify the second target text based on the first modification suggestion to obtain a third target text. Obtain the target text based on the third target text.
[0124] Specifically, after the terminal checks the second target text through the large language model and obtains the first modification suggestion, the terminal can modify the second target text according to the first modification suggestion through the large language model to obtain the third target text. In this way, the terminal can perform modification iterative processing and deep processing on the third target text through the large language model and the first target text in the data group to obtain the target texts corresponding to the first source texts in each data group, thereby ensuring the text accuracy and semantic consistency of the target text.
[0125] In this embodiment, the translation text output by the large language model is checked for corpus style through the large language model, and modification suggestions such as grammatical error data of the translation text are output, and the output translation text is modified based on the modification suggestions. This can obtain a higher quality translation text, further ensure the reliability of the data set used for model training, and thus improve the translation performance of the text translation model.
[0126] In an exemplary embodiment, the modification suggestion data also includes a second modification suggestion, such as Figure 5 As shown, the specific processing process of the step "obtaining the target text based on the third target text" includes:
[0127] Step 502: Input the first target text and the third target text corresponding to the first source text into the large language model, evaluate the third target text based on the target evaluation dimension and the first target text, and obtain an evaluation score and a second modification suggestion.
[0128] Among them, the first target text corresponding to the first source text can be the original text of the first source text extracted from the public data set after text processing, for example, it can be a translated text after translation processing, that is, a first translated text. The target evaluation dimensions include one or more of fluency, accuracy, style consistency and content integrity. Optionally, the target evaluation dimensions include fluency, accuracy, style consistency and content integrity. Fluency characterizes the natural fluency of the translated text, as well as the correctness of grammar and syntax. Accuracy characterizes whether the translated text accurately conveys the meaning of the source text, whether there are omissions or erroneous translations, and style consistency characterizes whether the translated text is consistent with the source text in terms of corpus style and tone, and whether it conforms to the expression habits of the target language type; content integrity characterizes whether the translated text completely contains important information in the source text.
[0129] Specifically, the terminal can input the first target text and the third target text of the first source text into the large language model, and obtain a comparison result after comparing the first target text and the third target text through the large language model. Based on the comparison result, the third target text is evaluated on the target evaluation dimension to obtain the evaluation scores corresponding to each target evaluation dimension, and obtain the second modification opinion, which may include the reason for the modification and the modification measures.
[0130] Step 504, when it is determined that the third target text does not meet the preset standard conditions, the third target text is improved based on the evaluation score and the second modification suggestion to obtain a modified third target text, and the third target text is evaluated again based on the target evaluation dimension and the first target text to obtain an evaluation score and the second modification suggestion, until the third target text that meets the preset standard conditions is obtained.
[0131] The target evaluation dimension includes one or more evaluation dimensions of fluency, accuracy, style consistency, and content completeness. The preset standard condition may be that a pre-configured evaluation score is greater than or equal to a preset evaluation score threshold.
[0132] Specifically, the terminal can determine whether the preset standard conditions are met based on the evaluation score of the target evaluation dimension of the third target text. If the evaluation score is less than the preset evaluation score threshold, it can be determined that the current third target text does not meet the preset standard conditions. In this way, the terminal can re-input the first source text, the third target text and the second modification suggestion into the large language model, and modify the third target text through the large language model, the first source text and the second modification suggestion to obtain the modified third target text. In this way, the terminal can re-execute the evaluation of the third target text based on the target evaluation dimension and the first target text in step 502 based on the modified third target text, obtain the evaluation score and the second modification suggestion, until the third target text that meets the preset standard conditions is obtained.
[0133] Step 506 , based on the first target texts corresponding to the third target texts, respectively calculate the evaluation index values corresponding to the third target texts, and screen the third target texts based on the evaluation index values to obtain the target texts corresponding to the first source texts.
[0134] Among them, the evaluation index value can be a numerical value corresponding to the evaluation index, and the evaluation index can be BLEU, COMET, or a large language model score; optionally, BLEU represents the degree of match between the third target text and the reference translation (the first target text), for example, the BLEU score can be calculated based on the degree of n-gram overlap, that is, the evaluation index value of the BLEU; COMET: the terminal can evaluate the semantic similarity and quality of the translated text and the reference translation through a pre-trained neural network model, and obtain an evaluation index value for the similarity; large language model score: the terminal can evaluate the fluency and naturalness of the third target text through a specific language model and score it to obtain an evaluation index value, and screening based on this dimension can ensure that the generated translated text conforms to the language habits of the target language.
[0135] Specifically, the terminal can calculate the evaluation index value of the third target text on each evaluation index based on the third target text and the first target text corresponding to the third target text, obtain a comprehensive evaluation index value based on the evaluation index value on each evaluation index, screen each comprehensive evaluation index value based on the comprehensive evaluation index value, and determine the third target text whose comprehensive evaluation index value is greater than or equal to a preset threshold as the target text. The preset threshold can be 90 points or 95 points. The embodiment of the present disclosure does not limit the specific value of the preset threshold, and those skilled in the art can determine it based on the actual application scenario.
[0136] In this embodiment, by scoring and screening each translated text on multiple automated evaluation dimensions, it can be ensured that the output translation corpus is higher than the preset standard, and it can be determined that the data set used for model training has sufficient quality to ensure the quality of text processing. It can also be ensured that the translated text generated when the model performs data processing conforms to the language habits of the target language type, thereby improving the translation quality and the data processing accuracy of the model.
[0137] In an exemplary embodiment, Figure 6 As shown, the text data processing method also includes:
[0138] Step 602: pre-train the basic text model based on the corpus style instructions in the target data set, the first source text, and the target text to obtain the character prediction probability corresponding to the first source text.
[0139] Among them, the character prediction probability represents the probability value of each character position being a character in the standard vocabulary. The character prediction probability may include the probability value of each character position being a character in the standard vocabulary, for example, it may be the highest probability value of a character in the standard vocabulary; for example, the standard vocabulary may include multiple words, and for the first character position, the terminal may obtain the probability of the first character position being each word in the standard vocabulary through a model, and determine the character prediction probability corresponding to the first character position with the highest probability (target probability value), representing that the character at the first character position is most likely the character corresponding to the highest probability, and the specific probability may be the target probability value. The character prediction probability at each character position may be determined based on the character prediction probabilities at each character position before the character position, for example, the character prediction probability at the third character position is determined based on the character prediction probability at the first character position and the character prediction probability at the second character position.
[0140] Specifically, the basic text model can be a general text processing base model, which has certain generalization ability and instruction-following ability. The terminal can use the target data set as training data for the text processing base model. The terminal can perform self-supervised learning training on the parameters of the attention layer, fully connected layer, input-output layer, embedding layer and normalization layer of the text processing base model to obtain a pre-trained processing model.
[0141] Optionally, the pre-trained processing model can be a pre-trained translation model; the basic text model can be a basic translation model, that is, a general translation base model, which has certain generalization ability and instruction-following ability. The terminal can use the target data set as training data for the translation base model, and the terminal can perform self-supervised learning training on the parameters of the attention layer, fully connected layer, input and output layer, embedding layer and normalization layer of the translation base model to obtain the pre-trained translation model.
[0142] For example, the terminal can concatenate the various translation corpora in the target data set to obtain a character string. For example, the terminal can use the various translation style instructions, the first source text, and the target text as a complete dialogue-style character string combined as data, and output multiple character strings to the basic translation model to be trained to obtain the character prediction probability corresponding to the first source text.
[0143] Step 604, calculate the loss function based on the target text corresponding to the source text and the character prediction probability corresponding to the first source text, and update the weight matrix of the basic text model through the loss function until the preset training completion conditions are met to obtain the pre-trained processing model.
[0144] The preset training completion condition may be that the number of training times reaches a preset threshold, or that the loss value corresponding to the loss function has converged.
[0145] Specifically, the terminal can determine the standard probability corresponding to each character position based on the target text corresponding to the source text, calculate the character prediction probability corresponding to the first source text based on the standard probability of each character position (for example, it can be 1), and obtain a loss function. Based on the loss function, the weight matrix of the basic text model is updated to obtain an updated basic text model. If the preset training completion condition is not met at present, step 602 is re-executed based on the updated basic text model until the preset training completion condition is met to obtain a pre-trained processing model.
[0146] Optionally, the terminal can obtain multiple test data sets based on the target data set, and pre-train multiple identical basic text models based on each test data set to obtain the character prediction probability corresponding to the first source text, and calculate the loss function based on the target text corresponding to the source text and the character prediction probability corresponding to the first source text. In this way, the terminal can screen based on each loss function and use the model with the smallest loss value as the pre-trained processing model that meets the preset training completion conditions, that is, obtain the pre-trained processing model.
[0147] Step 606, training the pre-trained processing model through a target training method to obtain a trained basic text model.
[0148] Among them, the target training method includes one or more of supervised fine-tuning, distillation learning, reinforcement learning, and hyperparameter adjustment.
[0149] Specifically, the terminal can train the pre-trained processing model through one or more of supervised fine-tuning, distillation learning, reinforcement learning, and hyperparameter adjustment to obtain a trained basic text model.
[0150] In this embodiment, model training is performed in a target training manner, which can ensure the lightweight, accuracy and efficiency of translation tasks in different language types.
[0151] In one embodiment, the target training method includes any one of supervised fine-tuning, distillation learning, and reinforcement learning, or they may be used in combination, which is not limited in the present disclosure; accordingly, the specific execution process of the step of "training the pre-trained processing model through the target training method to obtain a trained basic text model" includes one of the following situations:
[0152] Case 1: Divide the target data set based on language type to obtain data subsets corresponding to each language type. For each language type, train the trained basic text model based on the data subset corresponding to the language type to obtain the model weight corresponding to the language type. Process the model weight of the pre-trained basic text model based on the model weight corresponding to the language type to obtain the trained basic text model.
[0153] Specifically, the terminal can divide the target data set according to the language type of the target text corresponding to the source text, and obtain data subsets corresponding to each language type. For each data subset, the terminal can train a small number of newly added parameters in the trained basic text model based on the data subset and the LoRA or P-Tuning algorithm to obtain the model weight corresponding to the language type. The model weights of the pre-trained basic text model are processed based on the model weights corresponding to the language type. For example, the terminal can store the model weights corresponding to the language type, or replace the model weights of the pre-trained processing model based on the model weights of the language type to obtain a trained basic text model.
[0154] Optionally, the terminal can divide the target data set according to the language type of the translated text to obtain data subsets corresponding to each language type. For each data subset, the terminal can train a small number of newly added parameters in the trained basic translation model based on the data subset and the LoRA or P-Tuning algorithm to obtain the model weight corresponding to the language type. The model weight of the pre-trained basic translation model is processed based on the model weight corresponding to the language type. For example, the terminal can store the model weight corresponding to the language type, or replace the model weight of the pre-trained translation model based on the model weight of the language type to obtain a trained basic translation model.
[0155] Optionally, the terminal may train the pre-trained processing models corresponding to each data subset to obtain model weights corresponding to each language type, and the terminal may save the model weights corresponding to each language type in a local storage space.
[0156] Case 2: The corpus style instructions, the first source text, and the target text in the target data set are processed by the pre-trained processing model to perform text prediction processing, and the character prediction probability corresponding to the first source text is obtained. The character prediction probability represents the probability value of each character position being a character in the target vocabulary. Based on the target text corresponding to the source text and the character prediction probability corresponding to the first source text, a trained basic text model is obtained. The target vocabulary is obtained after the preset large model processes the target data set.
[0157] Among them, the target vocabulary can be a soft label output by a preset large model after processing the target data set, wherein the total probability value of each character in each target vocabulary can be 1 or other values less than 1 output by the preset large model.
[0158] Specifically, the terminal can use the target data set as the training data of the text processing base model, and the terminal can perform self-supervised learning training on the parameters of the attention layer, the fully connected layer, the input-output layer, the embedding layer and the normalization layer of the translation base model to obtain a pre-trained processing model. The terminal can determine the standard probability corresponding to each character position based on the target text corresponding to the source text and the target vocabulary, calculate the standard probability of each character position and the character prediction probability corresponding to the first source text to obtain a loss function, update the weight matrix of the basic text model based on the loss function, and obtain an updated basic text model. If the preset training completion condition is not met at present, the training step is re-executed based on the updated basic text model until the preset training completion condition is met to obtain a trained basic text model.
[0159] Case 3: The corpus style instructions, the first source text, and the target text in the target data set are processed by the pre-trained processing model to perform text prediction processing, and the character prediction probability corresponding to the first source text is obtained. The character prediction probability represents the probability value of each character position being a character in the standard vocabulary. Based on the target text corresponding to the source text, the character prediction probability corresponding to the first source text, and the environmental feedback data, a trained basic text model is obtained. The environmental feedback data is the evaluation data of the target text by the external environment.
[0160] Specifically, the terminal can use the target data set as the training data of the text processing base model, and the terminal can perform self-supervised learning training on the parameters of the attention layer, fully connected layer, input-output layer, embedding layer and normalization layer of the text processing base model to obtain a pre-trained processing model. The terminal can determine the standard probability corresponding to each character position based on the target text corresponding to the source text, calculate the character prediction probability corresponding to the standard probability of each character position and the first source text to obtain a loss function, update the weight matrix of the basic text model based on the loss function and the environmental feedback data to obtain an updated basic text model, and if the preset training completion condition is not met at present, re-execute the training steps based on the updated basic text model until the preset training completion condition is met to obtain a trained basic text model.
[0161] In this embodiment, model training is performed in a target training manner, which can ensure the lightweight, accuracy and efficiency of translation tasks in different language types.
[0162] In one embodiment, the method further comprises:
[0163] The target data set is divided based on language type to obtain data subsets corresponding to each language type; the trained basic text model is processed based on the data subsets corresponding to each language type to obtain model weights corresponding to each language type.
[0164] Specifically, the terminal can divide the target data set according to the language type of the translated text to obtain data subsets corresponding to each language type. For each data subset, the terminal can train a small number of newly added parameters in the trained basic text model based on the data subset and the LoRA or P-Tuning algorithm to obtain the model weight corresponding to the language type. For example, the terminal can store the model weight corresponding to the language type. The terminal can train the pre-trained processing model based on each data subset to obtain the model weight corresponding to each language type. The terminal can save the model weights corresponding to each language type in the local storage space.
[0165] Optionally, the cloud server can divide the target data set according to the language type of the translated text to obtain data subsets corresponding to each language type. For each data subset, the cloud server can train a small number of newly added parameters in the trained basic text model based on the data subset and the LoRA or P-Tuning algorithm to obtain the model weight corresponding to the language type. Then, the terminal can download the model weights corresponding to each language type in advance from the cloud server, for example, the model weights corresponding to each language type can be saved in the local storage space.
[0166] In this embodiment, lightweight training can be performed on the terminal or cloud server by dividing the data subsets contained in the target data set of the language type to obtain the model weights of each language type. This can effectively reduce the memory usage of the model and avoid the need for independent training and storage of large models for each language. It provides a data basis for achieving millisecond-level model weight loading and ensures low latency in user experience.
[0167] In one embodiment, Figure 7 As shown in FIG. 1 , the specific processing process of the step “updating the weight of the trained basic text model by the model weight to obtain the target text model” includes:
[0168] Step 702: merge the model weight corresponding to the target language type and the model weight of the trained basic text model to obtain the target model weight.
[0169] Step 704: Replace the model weights in the trained basic text model with the target model weights to obtain the target text model.
[0170] Specifically, the terminal can update the weight of the pre-trained basic text model based on the model weight corresponding to the target language type to obtain the basic text model after the weight update, that is, the target text model. For example, after obtaining the model weight corresponding to the target language type, the terminal can merge the model weight with the weight of the trained basic text model to obtain the merged weight, and replace the model weight of the trained basic text model based on the merged weight to obtain the target text model, which is a text processing model that matches the target language type.
[0171] In this embodiment, after determining the model weight corresponding to the target language type of the target text and merging it with the pre-trained translation model, data processing is performed, so that high-quality translation results can be quickly generated in an offline state, ensuring the flexibility of multi-language translation.
[0172] In the following, in conjunction with a specific embodiment, a detailed description of the specific implementation process of the above text translation method may include:
[0173] The text translation method provided in the embodiment of the present application can avoid the privacy issues of network dependence and data leakage through a localized and lightweight text translation model, and at the same time improve the accuracy of multilingual translation by optimizing the model structure, and support flexible style adjustment, thereby improving user experience and applicability. It includes a data preparation stage, a model training stage, and a reasoning stage. Each stage is optimized to achieve an efficient and accurate offline translation experience on a mobile device. In the data preparation stage, translations can be generated, self-reflected, translations can be modified, and model scores can be obtained to obtain a target data set; in the training model stage, training can be performed through a variety of training methods such as supervised fine-tuning, distillation learning, and reinforcement learning to obtain a trained text translation model; in the model reasoning stage, a language is selected, weights are loaded, and translations are generated, and text translation is performed through a text translation model that matches the selected target language type, thereby improving translation quality and the degree of adaptation to different language types.
[0174] That is to say, the text translation method provided in the embodiment of the present application provides an efficient local machine translation solution suitable for mobile devices (such as smart phones), especially for offline translation tasks. Its solution includes three main stages: data preparation, model training and end-side reasoning. In the data preparation stage, the initial translation corpus is first generated through the online large model, and the self-reflection mechanism is used to automatically evaluate and optimize the translation quality. Subsequently, the corpus higher than the predetermined quality standard is selected through model scoring for subsequent model training to ensure data quality. A variety of techniques are used in the model training stage, such as supervised fine-tuning (SFT), distillation learning (DL) and reinforcement learning (RL). The model is first pre-trained through all the corpus, and then fine-tuned for different language pairs to ensure more accurate translation effects. Through the combination of multiple learning methods, the generalization ability of the model and its performance in different translation tasks are improved. In the end-side reasoning stage, after the user selects the target language, the system loads the lightweight weights of the corresponding language pair (such as weights based on LoRA) and merges them with the pre-trained base model. The merged model can quickly generate high-quality translation results, making the translation performance close to the cloud level, while reducing computing resource consumption and latency, and improving translation efficiency. The advantages of this solution are that it enables offline translation through localized deployment, ensures data privacy, and adapts to scenarios with unstable networks; the lightweight weight loading mechanism reduces the computing resource consumption of the device; and the high-quality data preparation process ensures improved translation performance.
[0175] The first step is data preparation: The goal of the data preparation stage is to generate high-quality training data to ensure that the translation model has excellent performance in subsequent stages. Since the machine translation model relies on a large amount of parallel corpus (i.e., accurate translation of the same content in different languages), the quality of data preparation directly determines the effect of model training. This stage consists of the following key steps:
[0176] 1. Generate translation corpus. Specifically, the initial translation corpus comes from public datasets, such as NLLB datasets and OPUS datasets. These datasets provide a large number of parallel corpora in multiple languages including Chinese, English, French, German, etc., covering multiple fields. First, dataset selection, initial screening and style unification are carried out.
[0177] The specific process of dataset selection includes: obtaining large-scale initial translation data through public parallel corpus datasets. These datasets provide a large amount of corpus required for machine translation, but the quality and style are not uniform. Therefore, the initial corpus needs to be screened after acquisition. Public datasets such as NLLB and OPUS cover parallel corpora of multiple language pairs and are suitable for use in multilingual translation tasks. These datasets usually contain translations in multiple fields, such as news, social media, technical documents, etc., providing rich language alignment resources for subsequent training.
[0178] The specific process of initial screening includes: screening criteria include whether the length of the source text and the target text is reasonable, whether the sentence structure is complete, whether the translated sentence conforms to common grammatical norms, etc. Among them, the text length is calculated by using the tokenizer of an open source model (such as Qwen) to divide the original text into multiple tokens and then count the number of tokens, retaining only the corpus between the preset minimum and maximum lengths. The evaluation of whether the sentence structure is complete and whether it conforms to common grammatical norms can be quantified by perplexity, which is calculated as follows:
[0179]
[0180] Where N is the number of samples in the test set, q(x i ) indicates that the model is given a sample x i The lower the perplexity, the more complete the sentence structure is and the more it conforms to common grammatical norms. By calculating the perplexity of the original text and the translated text, corpus with obvious errors can be excluded.
[0181] The specific process of style unification includes: Since the translation styles of various parallel corpora in the public data set are very different, and may come from different translators or translators, there are large style differences, and direct use will affect the learning effect of the model. Therefore, it is necessary to ensure the consistency of the training data through style unification. First, a variety of instructions with different translation styles are manually designed, and then the GPT-4o model imitates these instructions to generate more instructions, and finally manually screens out reasonable and commonly used instructions (instructions of various translation styles). Then, the translation instructions and translation corpora are randomly combined to obtain multiple data groups. For each data group, an online large language model with excellent performance translates the first source text in the data group into the target language according to the instructions to obtain the second translation text corresponding to the first source text. This step requires ensuring that multiple instructions are fully combined with multiple language pairs, and each combination has sufficient translation corpus.
[0182] After generating the translation corpus according to the instruction, the generated translated text is passed to the big model again to obtain the first modification suggestion, which is obtained by the big language model after checking whether the translation style of the second translation text is consistent with the translation style in the instruction, and pointing out the grammatical errors, inexpressiveness, over-interpretation and omissions, etc. Finally, let the big model modify the second translation text according to the first modification suggestion in the previous step to generate the final translation text (obtaining the third target text).
[0183] 2. Self-reflection and translation revision, which includes self-reflection and feedback loops.
[0184] The specific process of self-reflection includes: the self-reflection mechanism is a self-correction process. The big model evaluates its own output and makes adjustments. The reflection mechanism is mainly used to allow the big model to refer to the original translation in the data set to further improve the generated translation. The specific steps are as follows: the first step is to pass the generated translation and the original translation into the big model, requiring it to compare the two translations (i.e., the first translation text and the third target text), and then score the generated translation from four aspects, namely fluency, accuracy, style consistency and content completeness, and put forward improvement suggestions, that is, to obtain the second modification suggestion; it should be noted that the style consistency here refers to the style between the original text and the generated translation text. Fluency represents whether the translation (third target text) is natural and fluent, whether the grammar and syntax are correct, and the corresponding scoring standards are: 1 (very unfluent) to 5 (very fluent); Accuracy represents whether the translation accurately conveys the meaning of the original text, without omissions or incorrect translations, and the corresponding scoring standards are: 1 (very inaccurate) to 5 (very accurate); Consistency of Style represents whether the translation is consistent with the original text in style and tone, and whether it conforms to the culture and expression habits of the target language, and the corresponding scoring standards are: 1 (completely inconsistent in style) to 5 (completely consistent in style); Completeness of Content represents whether the translation completely contains all the important information in the original text without omissions, and the corresponding scoring standards are: 1 (severely missing content) to 5 (completely complete content). In the second step, the original text (first source text), the generated translation (third target text) and the modification suggestions (second modification suggestions) are passed to the large language model together, and it is asked to modify the translation to obtain the modified third target text.
[0185] The specific process of the feedback loop includes: through multiple rounds of self-reflection and revision, the system will automatically optimize its results after completing the initial translation until the translated text reaches the high standards set by the system. Finally, the generated corpus will be further scored and screened to ensure its suitability for subsequent model training.
[0186] 3. Model scoring and final screening: In order to ensure that the corpus used for model training is of sufficient quality, the system will conduct strict scoring and screening on the generated translation corpus. The scoring process is based on automated indicators and combined with manual evaluation to ensure that the output corpus is higher than the predetermined standard.
[0187] Specifically, the terminal can screen the modified third target text through automated evaluation indicators to obtain the target text, and can also screen the target text based on manual evaluation to obtain the screening results; and the terminal can use the first source text and the target text as target data sets. Automated evaluation indicators may include: BLEU: used to evaluate the match between the translated text and the reference translation, and the score is calculated based on n-gram overlap. COMET: Through a pre-trained neural network model, the semantic similarity and quality of the translated text and the reference translation are evaluated. Large language model scoring: The system runs a special language model to evaluate and score the fluency and naturalness of the translation to ensure that the generated text conforms to the language habits of the target language.
[0188] The above training process can provide a large amount of high-quality and reliable translation corpus for subsequent model training through a series of rigorous self-reflection, scoring and screening mechanisms, ensuring the efficiency and accuracy of the offline translation system in the reasoning stage.
[0189] The second step is the model training phase: The model training phase is the most core step in the entire machine translation process, which directly determines the quality of the translation. In the model training phase, the main goal is to optimize the translation model through the most advanced training techniques, such as supervised fine-tuning (SFT), distillation learning (DL) or reinforcement learning (RL), to ensure that it is lightweight, accurate and efficient in translation tasks of different language pairs. In terms of specific implementation, these three training techniques can be used separately or in combination.
[0190] 1. Pre-training of base model:
[0191] Goal: Build a universal translation base model that has a certain level of generalization and instruction-following capabilities to handle translation tasks in multiple language pairs. The base model is not only used for general translation tasks, but also serves as the basis for subsequent training processes.
[0192] Data source: All translation corpora generated in the "data preparation phase" are used as training data.
[0193] Initial weight: In order to flexibly adjust the generated translation style, the model itself needs to have a certain ability to follow instructions, so the "Decode-Only" structure Transformer pre-trained model is selected, such as LLama, Qwen and other series models. In order to run smoothly on mobile terminals such as smartphones, the number of model parameters is controlled at around 2 billion (2B). Models that meet the requirements include: Qwen2-1.5B-Instruct, gemma-2-2B-it, etc. These models have been trained with corpora in multiple languages and have been fine-tuned with instructions. They have the ability to translate in multiple languages and are very suitable as base models.
[0194] Training strategy: After selecting the base model, open the parameters of all weight layers of the base model (including attention layer, fully connected layer, input and output layer, embedding layer and normalization layer) for self-supervised learning training. The translation instructions, original text and translated text of each corpus are uniformly combined into a complete dialogue string as input. The output is the probability of the next token of each token. The loss of the model is calculated using the cross entropy loss function. Finally, after optimizing the hyperparameters many times, the model with the smallest loss in the test set is selected as the translation base model, that is, the pre-trained translation model is obtained.
[0195] 2. Supervised fine-tuning (SFT): After pre-training, each target language pair is fine-tuned individually. The translation accuracy of the language pair is further improved by using specific corpus from the data subset corresponding to each language type in the target dataset.
[0196] Goal: Ensure the translation accuracy of the model on a specific language pair, and further optimize the grammatical characteristics and translation habits of certain language pairs. Finally, generate independent lightweight weights (similar to the LoRA structure) for each language pair for easy deployment.
[0197] Data source: bilingual parallel corpora for each language pair generated in the "data preparation phase". These data are regenerated and screened by the large model to ensure their high quality and representativeness, so that the model in the training process can effectively capture the linguistic characteristics of the target language pair.
[0198] Initial weights: Use the weights of the pre-trained translation base model. The translation base model is trained on a large-scale multilingual corpus and has good generalization ability.
[0199] Training strategy: Using LoRA or P-Tuning technology, the amount of parameters for fine-tuning specific tasks on the translation base model can be greatly reduced. With these two technologies, the model does not need to fully update all parameters, but only adjusts a small number of new parameters. These new parameters only account for about 1% to 5% of the base model parameters. Specifically, LoRA reduces the computational complexity of the matrix through low-rank decomposition, while P-Tuning uses parameter-efficient hint learning methods to obtain performance close to full fine-tuning with as few parameter changes as possible.
[0200] This supervised fine-tuning training method can save computing resources and storage space, and can also improve the fine-tuning efficiency of the model while ensuring translation accuracy. It is very suitable for environments with limited computing resources such as mobile devices. During fine-tuning, only the loss of the generated translation is calculated, and the input instructions and original text will be masked with a mask to ensure training efficiency and avoid additional consumption of computing resources.
[0201] 3. Distillation Learning (DL): The core of distillation learning is to transfer knowledge from a large model (teacher model) to a small model (student model) to ensure that the model running on resource-constrained mobile devices can also achieve translation quality close to that of the large model.
[0202] Objective: The main objective of distillation learning is the same as supervised fine-tuning, which is to improve the translation accuracy of the model on a specific language pair. However, distillation learning further optimizes the training process by introducing the guidance of the teacher model, so that the student model (pre-trained translation model) can more effectively learn the rich knowledge contained in the teacher model (pre-set large language model).
[0203] Data source: The data used for distillation learning comes from the bilingual parallel corpus of each language pair generated in the "data preparation phase".
[0204] Initial weights: Use the pre-trained translation base model weights. You can also combine the lightweight weights after supervised fine-tuning with the base model weights as the initial weights.
[0205] Training strategy: The training strategy of distillation learning is flexible and efficient. Only a small number of new parameters are adjusted using LoRA or P-Tuning technology. Unlike the one-hot encoded "hard labels" used in traditional supervised fine-tuning, distillation learning relies on the "soft labels" (target vocabulary) output by the teacher model. Soft labels contain probability distribution information and can therefore convey more knowledge. The student model (which can be a pre-trained translation model or a translation model trained through supervised fine-tuning) can have a deeper understanding of the target task and improve the translation quality of the model by learning these soft labels during the training process.
[0206] 4. Reinforcement Learning (RL): Reinforcement learning provides a self-optimization mechanism for model training. The model can adjust its translation strategy through environmental feedback to generate more natural and accurate translation results.
[0207] Goal: The translation effect of each target language pair needs to be optimized.
[0208] Data source: The training data in the reinforcement learning phase also comes from the bilingual parallel corpus generated in the "data preparation phase". However, reinforcement learning does not rely solely on these static training data, but introduces a dynamic feedback system generated by the interaction between the model and the environment. This environment can be a manually designed evaluation mechanism, or feedback based on user feedback and actual application scenarios. In this way, the model can gradually learn how to generate more appropriate translation results in a specific context.
[0209] Initial weights: Use the pre-trained translation base model weights. You can also combine the lightweight weights after supervised fine-tuning with the base model weights as the initial weights.
[0210] Reward mechanism: The core of reinforcement learning is to guide the learning process of the model through a reward mechanism. Every time the model generates a translation result, the environment will reward or punish the model's output based on multiple indicators such as the accuracy, naturalness, and coherence of the context. Common evaluation indicators can include BLEU scores (used to measure the closeness of the translation to the reference translation), fluency evaluation (to evaluate the naturalness of the translation), and user feedback (such as click-through rate and satisfaction through actual user feedback). By continuously adjusting the model's behavioral strategy, the model can optimize its translation strategy in the direction of obtaining higher rewards, thereby generating higher quality translations.
[0211] Training strategy: During the training process, the reinforcement learning model continuously explores different translation paths and adjusts them according to the feedback from the environment. In order to improve the training efficiency, reinforcement learning algorithms such as "strategy gradient" or "Q-learning" can be used. The model continuously interacts with the environment, that is, through the environmental feedback data output by the environment, and continuously corrects its own translation decisions in order to obtain greater rewards in the long-term goals (such as better translation performance). At the same time, reinforcement learning can also be combined with human feedback or expert evaluation to further refine and optimize the reward mechanism to ensure that the generated translation results are not only accurate in form, but also more in line with the needs of actual language use.
[0212] 5. Visualization and hyperparameter adjustment during training: Visualization of loss and score: In order to better monitor the performance of the model during training, tools such as TensorBoard are used to perform real-time visualization of the training loss and translation evaluation indicators of each language pair.
[0213] Loss curve: Monitor the loss changes during training and adjust hyperparameters such as learning rate in a timely manner to prevent the model from overfitting or underfitting.
[0214] Translation score curve: Evaluate the translation quality of the model on different language pairs through changes in indicators such as BLEU and COMET, ensuring that the model is continuously optimized during training.
[0215] Hyperparameter adjustment: Adjust hyperparameters such as learning rate, batch size, distillation temperature, etc. according to visual feedback to achieve the best training results.
[0216] The third step is the model inference stage: The end-side inference stage is the final link of the entire local machine translation solution. It adopts an optimized inference solution based on the MLC-LLM (Machine Learning Compiler for Large Language Models) framework to ensure efficient execution of the model on mobile devices. The goal of this stage is to achieve high-quality offline translation using a lightweight model architecture without relying on the cloud.
[0217] The text translation method provided in this embodiment can be applied to various scenarios where localized offline translation is required on mobile devices, especially in environments where the network connection is unstable or cloud resources cannot be accessed. The following are some typical usage scenarios:
[0218] No network environment: When users are in places without Internet access (such as remote areas, airplanes, or basements), the invention can provide users with efficient local machine translation services, avoiding reliance on online translation tools. In addition, users do not need to wait for network connection or worry about cloud service interruptions, thus ensuring the continuity of translation.
[0219] Privacy protection needs: In scenarios where sensitive or private information needs to be translated (such as passports, medical records, confidential documents, etc.), the invention avoids the risk of user data being transmitted over the network by performing all translation operations on the local device, effectively ensuring privacy security. Such scenarios include internal corporate document translation and personal privacy data processing.
[0220] Multi-language requirements: This product is suitable for users who need multi-language support, especially in environments where they need to switch between different languages quickly, such as international travel, multinational business meetings, or work scenarios in a multi-language environment. Users can dynamically load the weights of different language pairs according to their needs without having to download all models in advance, which greatly improves the flexibility and storage efficiency of the translation system.
[0221] Scenarios with limited device resources: For mobile devices with limited processing power and storage space (such as smartphones and tablets), this invention achieves efficient translation processing through lightweight weight loading and optimized inference engines, ensuring that the device can still run translation tasks smoothly even with limited resources. It is suitable for users who need efficient translation functions on business trips, travel or low-configuration devices.
[0222] Offline education and learning: In scenarios where foreign language learning or translation education is required, the invention can provide students and teachers with convenient offline translation support, especially in classrooms or outdoor learning environments, ensuring the availability of translation functions without a network, helping users to master multilingual content in real time.
[0223] Work environments with high security requirements: In some work scenarios with strict information security requirements, such as government agencies, the military, or security companies, the translation system can provide localized translation without the involvement of external servers, ensuring that all translation tasks are completed on local devices to avoid the leakage of sensitive information.
[0224] That is to say, the text translation method provided in this embodiment is particularly suitable for scenarios where translation needs to be performed without a network, with privacy protection, efficient resource utilization, multi-language support and a secure environment, providing users with a convenient, secure and efficient local machine translation solution, which can be specifically applied to smartphone translation applications, smart watches and wearable devices, portable translation devices, vehicle-mounted systems, smart home devices and enterprise-level translation software.
[0225] The following is an explanation of the technical terms involved in this embodiment:
[0226] Decode-Only structure, Decode-Only structure is an architecture used in natural language processing and generative models, mainly for text generation tasks. The characteristic of this structure is that it only contains a decoder part, but no encoder. The structural composition of the Decode-Only structure: In the Decode-Only model, only the decoder is used to generate the output. The Decode-Only model usually accepts a starting tag (such as "<start>") as input and gradually generates subsequent text based on this input. This means that it does not perform the encoding step when processing the input. During the generation process, the model will gradually predict the next word until a specific end tag (such as "<end>") is generated or the preset text length is reached. This process is called autoregressive generation.
[0227] LoRA (Low-Rank Adaptation) is a technique used to improve the performance of large models on specific tasks, especially when fine-tuning. LoRA reduces these requirements by introducing the concept of low-rank adaptation. The core idea of LoRA is to decompose the weights of the model into a low-rank matrix. This means that when updating the model, instead of adjusting all the parameters, they are only adapted in the low-rank space. This can significantly reduce the number of parameters that need to be updated.
[0228] Implementation: In LoRA, adaptation modules are usually added to certain layers of the model (such as linear layers), which learn task-specific features through low-rank matrices. During actual training, only these additional parameters are updated without changing the original pre-trained weights, which not only significantly reduces the computational cost, but also achieves good performance on smaller data sets and prevents overfitting. In addition, LoRA can be combined with other fine-tuning methods to provide a more flexible adaptation method. LoRA is widely used in various natural language processing and computer vision tasks, especially in scenarios that need to quickly adapt to new tasks, such as chatbots, text classification, etc.
[0229] MLC-LLM: is a compilation framework designed specifically for large language models. It aims to improve the performance and deployment efficiency of the model. It can solve the performance bottleneck of large language models on different hardware platforms, provide a unified compilation process, and optimize the inference speed and resource utilization. The framework supports a variety of hardware architectures, including CPUs, GPUs, and dedicated accelerators, and reduces the computational complexity of the model through automated optimization techniques. MLC-LLM uses a variety of strategies such as graph optimization and memory optimization to reduce memory usage and increase running speed while maintaining accuracy during model inference. It allows users to customize the compilation process according to specific tasks and hardware environments, supports various model architectures, and improves the flexibility and scalability of deployment. MLC-LLM is suitable for scenarios that require efficient reasoning, such as real-time dialogue systems, intelligent assistants, etc., and improves user experience by optimizing the reasoning performance of the model.
[0230] It should be noted that the term 'and / or' in this article is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, A and / or B can represent the three situations of: A exists alone, A and B exist at the same time, and B exists alone. In addition, 'at least one of the terms in this article' represents any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can represent any one or more elements selected from the set consisting of A, B, and C.
[0231] It should be noted that those skilled in the art can understand that in the above method of the specific implementation mode, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0232] It should be noted that the above description of the various embodiments tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. For the sake of brevity, they will not be repeated herein.
[0233] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0234] Based on the same inventive concept, the embodiment of the present application also provides a text data processing device for implementing the text data processing method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more text data processing device embodiments provided below can refer to the limitations of the text data processing method above, and will not be repeated here.
[0235] In an exemplary embodiment, Figure 8 As shown, a text data processing device 800 is provided, comprising:
[0236] A first acquisition module 802 is used to acquire a target language type corresponding to the target text and determine a model weight corresponding to the target language type;
[0237] The first updating module 804 is used to update the weight of the trained basic text model through the model weight to obtain the target text model; the trained basic text model is obtained by training with the target data set and the target training method;
[0238] The first text processing module 806 is used to process the target text through the target text model to obtain the target text corresponding to the target language type.
[0239] In one embodiment, the device further comprises:
[0240] A second acquisition module is used to acquire a first target corpus, where the first target corpus includes a first source text and a first target text corresponding to the first source text;
[0241] a combining module, used for combining a plurality of corpus style instructions with each of the first target corpora respectively, to obtain a plurality of data groups, each of the data groups comprising a corpus style instruction, a first source text and a first target text, each of the corpus style instructions being obtained after diffusion processing and screening processing of the initially configured corpus style instruction;
[0242] A second text processing module is used for, for each data group, performing data processing on the first source text based on the corpus style instruction in a preset large language model to obtain a second target text; determining modification suggestion data based on the second target text, and obtaining a target text corresponding to the first source text through the modification suggestion data, the first target text and the second target text;
[0243] The data set acquisition module is used to obtain a target data set based on the corpus style instructions, the first source text and the target text corresponding to each of the data groups.
[0244] In one embodiment, the device further comprises:
[0245] An extraction module is used to extract an initial text processing corpus from a public data set, wherein the initial text processing corpus includes an initial source text and an initial processing text corresponding to the initial source text;
[0246] The first screening module is used to screen each of the initial text processing corpora from a target screening dimension to obtain a first target corpus, wherein the first target corpus includes a first source text and a first target text corresponding to the first source text, and the target screening dimension includes one or more of a text length dimension, a text structure dimension, and a grammatical correctness dimension.
[0247] In one embodiment, the modification suggestion data includes a first modification suggestion; and the second text processing module is specifically configured to:
[0248] Inputting the second target text into the large language model, determining whether the corpus style of the second target text matches the corpus style in the corpus style instruction through the large language model, obtaining corpus style difference data, and processing the first source text and the second target text through the large language model to obtain content modification data; obtaining a first modification suggestion based on the difference data and the content modification data;
[0249] The second target text is modified based on the first modification suggestion through the large language model to obtain a third target text; and a target text is obtained based on the third target text.
[0250] In one embodiment, the modification suggestion data further includes a second modification suggestion, and the second text processing module is specifically configured to:
[0251] Inputting a first target text and a third target text corresponding to the first source text into the large language model, evaluating the third target text based on the target evaluation dimension and the first target text, and obtaining an evaluation score and a second modification suggestion;
[0252] In the case where it is determined that the third target text does not meet the preset standard conditions, the third target text is improved based on the evaluation score and the second modification suggestion to obtain a modified third target text, and the evaluation of the third target text based on the target evaluation dimension and the first target text is re-executed to obtain an evaluation score and the second modification suggestion, until a third target text that meets the preset standard conditions is obtained; the target evaluation dimension includes one or more evaluation dimensions of fluency, accuracy, style consistency, and content completeness;
[0253] Based on the first target texts corresponding to the third target texts, the evaluation index values corresponding to the third target texts are calculated respectively, and the third target texts are screened based on the evaluation index values to obtain the target texts corresponding to the first source texts.
[0254] In one embodiment, the device further comprises:
[0255] A pre-training module, configured to pre-train the basic text model based on the corpus style instructions in the target data set, the first source text, and the target text, to obtain a character prediction probability corresponding to the first source text, wherein the character prediction probability represents a probability value that each character position is a character in a standard vocabulary;
[0256] A first calculation module is used to calculate a loss function based on a target text corresponding to the source text and a character prediction probability corresponding to the first source text, and to update a weight matrix of the basic text model by the loss function until a preset training completion condition is met, thereby obtaining a pre-trained processing model;
[0257] The target training module is used to train the pre-trained processing model through a target training method to obtain a trained basic text model. The target training method includes one or more of supervised fine-tuning, distillation learning, reinforcement learning, and hyperparameter adjustment.
[0258] In one embodiment, the target training method includes any one of supervised fine-tuning, distillation learning, and reinforcement learning; the target training module is specifically used for:
[0259] Divide the target data set based on language type to obtain data subsets corresponding to each language type; for each language type, train the trained basic text model based on the data subset corresponding to the language type to obtain the model weight corresponding to the language type; process the model weight of the pre-trained basic text model based on the model weight corresponding to the language type to obtain the trained basic text model; or
[0260] The corpus style instructions, the first source text and the target text in the target data set are processed by the pre-trained processing model to perform text prediction processing, and the character prediction probability corresponding to the first source text is obtained, and the character prediction probability represents the probability value of each character position being a character in the target vocabulary; based on the target text corresponding to the source text and the character prediction probability corresponding to the first source text, a trained basic text model is obtained, and the target vocabulary is obtained after the preset large model processes the target data set; or
[0261] The pre-trained processing model is used to perform text prediction processing on the corpus style instructions, the first source text and the target text in the target data set to obtain the character prediction probability corresponding to the first source text, and the character prediction probability represents the probability value of each character position being a character in the standard vocabulary; based on the target text corresponding to the source text, the character prediction probability corresponding to the first source text and the environmental feedback data, a trained basic text model is obtained, and the environmental feedback data is the evaluation data of the external environment on the target text.
[0262] In one embodiment, the first update module is specifically used to:
[0263] The model weight corresponding to the target language type and the model weight of the trained basic text model are combined to obtain a target model weight;
[0264] The model weights in the trained basic text model are replaced by the target model weights to obtain a target text model.
[0265] Each module in the above-mentioned text data processing device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.
[0266] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig. 9As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to contrast images. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a text data processing method is implemented.
[0267] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0268] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps described in the above embodiment when executing the computer program.
[0269] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps described in the above embodiment are implemented.
[0270] In one embodiment, a computer program product is provided, including a computer program, which implements the steps described in the above embodiment when executed by a processor.
[0271] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0272] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0273] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0274] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0275] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A text data processing method, characterized in that: The method comprises: Obtaining a target language type corresponding to the target text, and determining a model weight corresponding to the target language type; The weight of the trained basic text model is updated by the model weight to obtain a target text model; the trained basic text model is obtained by training with a target data set and a target training method; The target text is processed by the target text model to obtain a target text corresponding to the target language type.
2. The method according to claim 1, characterized in that The method further comprises: Acquire a first target corpus, where the first target corpus includes a first source text and a first target text corresponding to the first source text; Combining a plurality of corpus style instructions with each of the first target corpora respectively to obtain a plurality of data groups, each of the data groups comprising a corpus style instruction, a first source text and a first target text, each of the corpus style instructions being obtained by diffusion processing and screening processing of the initially configured corpus style instruction; For each data group, in a preset large language model, data processing is performed on the first source text based on the corpus style instruction to obtain a second target text; Determining modification suggestion data based on the second target text, and obtaining a target text corresponding to the first source text through the modification suggestion data, the first target text, and the second target text; The target data set is obtained based on the corpus style instructions, the first source text and the target text respectively corresponding to each of the data groups.
3. The method according to claim 2, characterized in that The method further comprises: Extracting an initial text processing corpus from a public data set, wherein the initial text processing corpus includes an initial source text and an initial processing text corresponding to the initial source text; Each of the initial text processing corpora is screened from a target screening dimension to obtain the first target corpus, wherein the target screening dimension includes one or more of a text length dimension, a text structure dimension, and a grammatical correctness dimension.
4. The method according to claim 2, characterized in that: The modification suggestion data includes a first modification suggestion; the determining of the modification suggestion data based on the second target text, and obtaining a target text corresponding to the first source text through the modification suggestion data, the first target text, and the second target text, includes: Inputting the second target text into the large language model, determining whether the corpus style of the second target text matches the corpus style in the corpus style instruction through the large language model, obtaining corpus style difference data, and processing the first source text and the second target text through the large language model to obtain content modification data; obtaining a first modification suggestion based on the difference data and the content modification data; The second target text is modified based on the first modification suggestion through the large language model to obtain a third target text; and a target text is obtained based on the third target text.
5. The method according to claim 4, characterized in that The modification suggestion data also includes a second modification suggestion, and obtaining a target text based on the third target text includes: Inputting a first target text and a third target text corresponding to the first source text into the large language model, evaluating the third target text based on the target evaluation dimension and the first target text, and obtaining an evaluation score and a second modification suggestion; In the case where it is determined that the third target text does not meet the preset standard conditions, the third target text is improved based on the evaluation score and the second modification suggestion to obtain a modified third target text, and the evaluation of the third target text based on the target evaluation dimension and the first target text is re-executed to obtain an evaluation score and the second modification suggestion, until a third target text that meets the preset standard conditions is obtained; the target evaluation dimension includes one or more evaluation dimensions of fluency, accuracy, style consistency, and content completeness; Based on the first target texts corresponding to the third target texts, the evaluation index values corresponding to the third target texts are calculated respectively, and the third target texts are screened based on the evaluation index values to obtain the target texts corresponding to the first source texts.
6. The method according to claim 1, characterized in that The method further comprises: Pre-training the basic text model based on the corpus style instructions in the target data set, the first source text, and the target text to obtain character prediction probabilities corresponding to the first source text, wherein the character prediction probabilities represent probability values of each character position being a character in a standard vocabulary; A loss function is calculated based on the target text corresponding to the first source text and the character prediction probability corresponding to the first source text, and a weight matrix of the basic text model is updated by the loss function until a preset training completion condition is met to obtain a pre-trained processing model; The pre-trained processing model is trained by a target training method to obtain a trained basic text model, and the target training method includes one or more of supervised fine-tuning, distillation learning, reinforcement learning, and hyperparameter adjustment.
7. The method according to claim 6, characterized in that The target training method includes any one of supervised fine-tuning, distillation learning, and reinforcement learning; the pre-trained processing model is trained by the target training method to obtain a trained basic text model, including: Divide the target data set based on language type to obtain data subsets corresponding to each language type; for each language type, train the trained basic text model based on the data subset corresponding to the language type to obtain the model weight corresponding to the language type; process the model weight of the pre-trained basic text model based on the model weight corresponding to the language type to obtain the trained basic text model; or The corpus style instructions, the first source text and the target text in the target data set are processed by the pre-trained processing model to perform text prediction processing, and the character prediction probability corresponding to the first source text is obtained, and the character prediction probability represents the probability value of each character position being a character in the target vocabulary; based on the target text corresponding to the source text and the character prediction probability corresponding to the first source text, a trained basic text model is obtained, and the target vocabulary is obtained after the preset large model processes the target data set; or The pre-trained processing model is used to perform text prediction processing on the corpus style instructions, the first source text and the target text in the target data set to obtain the character prediction probability corresponding to the first source text, and the character prediction probability represents the probability value of each character position being a character in the standard vocabulary; based on the target text corresponding to the source text, the character prediction probability corresponding to the first source text and the environmental feedback data, a trained basic text model is obtained, and the environmental feedback data is the evaluation data of the external environment on the target text.
8. The method according to claim 6, characterized in that The method further comprises: Dividing the target data set based on language type to obtain data subsets corresponding to each language type; The trained basic text model is processed based on the data subset corresponding to each of the language types to obtain the model weight corresponding to each of the language types.
9. A text data processing device, characterized in that: The device comprises: A first acquisition module, used to acquire a target language type corresponding to a target text, and determine a model weight corresponding to the target language type; A first updating module is used to update the weight of the trained basic text model by using the model weight to obtain a target text model; the trained basic text model is obtained by training with a target data set and a target training method; The first translation module is used to process the target text through the target text model to obtain a target text corresponding to the target text.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Model inversion defense method and device based on compression fine tuning and medium
CN120234802A
A model inversion defense method and device based on compression fine-tuning and a medium
CN120234802B