Data Processing Method, Apparatus, Electronic Device, and Storage Medium

CN117034959BActive Publication Date: 2025-07-29BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310996726.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-08
Publication Date
2025-07-29
Estimated Expiration
2043-08-08

Smart Images

  • Figure CN117034959B_ABST
    Figure CN117034959B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a data processing method, apparatus, electronic device, and storage medium. Among them, the method includes: receiving a corpus to be processed; processing the corpus to be processed based on a target diffusion language model to obtain a target prediction result corresponding to the corpus to be processed; wherein the target diffusion language model is trained based on a plurality of corpus samples, and the masked corpus in the corpus samples corresponds to different masking rates; and presenting the target prediction result. The technical solution of the embodiments of the present disclosure realizes the effect that the target diffusion language model can process corpus data based on the principle of the diffusion model, and the obtained corpus data processing result can meet the requirements of the corpus processing task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of data processing, and in particular, to a data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, related research on processing corpus data using Artificial Intelligence (AI) has been gradually carried out. In this way, the processing requirements of users for corpus data are met.

[0003] Generally, an autoregressive language model can be used to process corpus data to obtain corresponding corpus processing results.

[0004] However, since the autoregressive language model is obtained after large-scale unsupervised pre-training and downstream fine-tuning, in the process of continuous iterative training, the autoregressive language model may naturally have problems such as error accumulation and lack of global vision. Therefore, the quality of the corpus processing results is poor and the accuracy is low, and the corpus processing results output by the model cannot meet the user's needs. Summary of the Invention

[0005] The present disclosure provides a data processing method, apparatus, electronic device, and storage medium to enable a target diffusion language model to process corpus data based on the principle of a diffusion model, and the obtained corpus data processing results can meet the requirements of corpus processing tasks.

[0006] In a first aspect, an embodiment of the present disclosure provides a data processing method, which includes:

[0007] Receiving the corpus to be processed;

[0008] Processing the corpus to be processed based on a target diffusion language model to obtain a target prediction result corresponding to the corpus to be processed; wherein the target diffusion language model is trained based on multiple corpus samples, and the masked corpus in the corpus samples corresponds to different masking rates;

[0009] Displaying the target prediction result.

[0010] In a second aspect, an embodiment of the present disclosure further provides a data processing apparatus, which includes:

[0011] A corpus receiving module, configured to receive the corpus to be processed;

[0012] A corpus processing module, configured to process the to-be-processed corpus based on a target diffusion language model to obtain a target prediction result corresponding to the to-be-processed corpus; wherein, the target diffusion language model is trained based on a plurality of corpus samples, and the masked corpus in the corpus samples corresponds to different masking rates;

[0013] A prediction result display module, configured to display the target prediction result.

[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, which includes:

[0015] One or more processors;

[0016] A storage device, configured to store one or more programs,

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method as described in any one of the embodiments of the present disclosure.

[0018] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the data processing method as described in any one of the embodiments of the present disclosure when executed by a computer processor.

[0019] The technical solution of the embodiment of the present disclosure realizes the effect of processing corpus data based on a diffusion model by receiving the to-be-processed corpus, and further processing the to-be-processed corpus based on the target diffusion language model to obtain a target prediction result corresponding to the to-be-processed corpus. Finally, the target prediction result is displayed, which solves the problems such as poor quality, low accuracy, and non-conformity with user requirements of the corpus processing result when using an autoregressive language model to process corpus data in the related art. It realizes that the target diffusion language model can process corpus data based on the principle of the diffusion model, and the obtained corpus data processing result can meet the requirements of the corpus processing task, improves the accuracy of the corpus data processing result, and enhances the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Combined with the drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale.

[0021] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of the present disclosure;

[0022] Figure 2Flow chart of a data processing method provided by an embodiment of the present disclosure;

[0023] Figure 3 Flow chart of a data processing method provided by an embodiment of the present disclosure;

[0024] Figure 4 Structural diagram of a data processing device provided by an embodiment of the present disclosure;

[0025] Figure 5 Structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0026] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0027] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0028] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0029] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0030] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0031] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0032] It is understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0033] For example, when responding to an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure based on the prompt message.

[0034] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0035] It is understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0036] It is understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0037] Before introducing this technical solution, an exemplary description of the application scenario can be given first. This technical solution can be applied to a scenario where an input corpus is processed according to a corpus processing task corresponding to a neural network model to obtain a corpus processing result matching the corpus processing task. The technical solution provided by the embodiments of the present disclosure adopts a target diffusion language model in the network model. The target diffusion language model can be a diffusion model for processing corpus data. This model can be a model constructed based on a pre-trained masked language, that is, the pre-trained Masked Language Model (MLM) is used as the diffusion language model. Furthermore, by training the masked language model in the training manner of the diffusion model, the trained diffusion language model can be used as the target diffusion language model. Thus, the effects of reducing the training cost, improving the training efficiency, and the model accuracy are achieved. Further, the target diffusion language model can be applied to any scenario corresponding to a corpus processing task. For example, it can be applied to a scenario of reviewing an input corpus. Among them, the review processing can be to summarize the input article, and the result of the summary is used as the review result of the article. In the embodiments of the present disclosure, the article to be reviewed can be used as the corpus to be processed. After obtaining the article to be reviewed, the article to be reviewed can be input into the target diffusion language model, and the summary result corresponding to the article to be reviewed can be obtained.

[0038] Before introducing the solution of the embodiments of the present disclosure, it should also be noted that the target diffusion language model constructed based on the embodiments of the present disclosure can be deployed on the server side or the client side. Among them, the server side can be a targeted service program that provides services and resources to the client side, and the device running the server side is the server. Correspondingly, the client side is a program corresponding to the server side that provides local services to users. At the same time, the client side and the server side can communicate based on various forms of text transfer protocols, such as the Hyper Text Transfer Protocol (HTTP). Exemplarily, the target diffusion language model in the embodiments of the present disclosure is integrated into an application software that supports multiple functions such as natural language processing or special effect image processing, and this software can be installed in an electronic device. Optionally, the electronic device can be a mobile terminal or a PC terminal, etc. The application software can be a type of software for processing data such as text, images, videos, or audios. The specific application software will not be elaborated here one by one, as long as it can realize the processing of data such as text, images, videos, or audios. It can also be a specially developed application program, integrated into the corresponding software, or integrated into the corresponding page, so that users can realize the processing of relevant data through the page integrated in the PC terminal.

[0039] Figure 1The flowchart of a data processing method provided by an embodiment of the present disclosure is applicable to the case where a corpus input into a model is processed based on a diffusion model to obtain a prediction result corresponding to the corpus input into the model, and the prediction result matches the corpus processing task corresponding to the corpus input into the model. This method can be executed by a data processing device, which can be implemented in the form of software and / or hardware. Optionally, it is implemented by an electronic device, which can be a mobile terminal, a PC, a server, etc.

[0040] As Figure 1 shown, the method of this embodiment may specifically include:

[0041] S110. Receive the corpus to be processed.

[0042] Among them, the corpus to be processed can be understood as the corpus to be processed for corpus processing. In this embodiment, the corpus to be processed can be a corpus in text form. In practical applications, the corpus to be processed can correspond to a corpus processing task. Exemplarily, if the corpus processing task is a text translation task, the corpus to be processed can be a corpus including the text to be translated; if the corpus processing task is an article review task, the corpus to be processed can be a corpus including the article to be reviewed; if the corpus processing task is an article abstract task, the corpus to be processed can be a corpus including the article to be abstracted.

[0043] In this embodiment, the corpus to be processed can be a corpus input in real time by a user through an input device (such as a keyboard, etc.) on a mobile terminal; or, it is a corpus uploaded to the corresponding server through an application software; or, it can also be a corpus pre-stored in the device storage space. Of course, in the actual application process, for an application software that provides natural language processing functions to users, each segment of the corpus data obtained after the server performs segmentation processing on the received corpus data can also be used as the corpus to be processed, and the embodiments of the present disclosure do not make specific limitations on this.

[0044] In practical applications, in order to identify the corpus processing task corresponding to the corpus to be processed, the corpus to be processed can carry an identifier representing the corresponding corpus processing task, and then, the corpus processing task corresponding to the corpus to be processed can be determined according to the corresponding identifier. Thus, the corpus processing result corresponding to the corpus processing task can be determined.

[0045] Based on this, before receiving the corpus to be processed, it further includes: editing the original corpus; marking a task identifier for the original corpus to obtain the corpus to be processed, so that the target diffusion language model processes the corpus to be processed based on the task identifier.

[0046] Among them, the original corpus can be the collected and unprocessed corpus. The language type of the original corpus can be any type. Optionally, the language type of the original corpus can be English, and correspondingly, the original corpus can be an English text. Exemplarily, the original corpus can be an English sentence, such as, "Diffusion language models can be so cool". It should be noted that the form of the original corpus can be any form. Optionally, the form of the original corpus can be in text form, audio form, video form, etc.

[0047] Among them, the original corpus can be the corpus input by the user in real time through the input device on the mobile terminal; or, it can be the corpus uploaded to the server through the application software; or, it can also be the corpus stored in the pre-constructed corpus. Of course, in the actual application process, for the application software that provides natural language processing functions to users, each segment of the corpus obtained after the server performs segmentation processing on the received corpus data can also be used as the original corpus, and the embodiments of the present disclosure do not make specific limitations in this regard. In actual applications, editing the original corpus can include various methods, and the editing methods of the original corpus will be described below.

[0048] The first method can be: displaying at least one corpus to be selected; and determining the original corpus in response to a selection trigger operation for the at least one corpus to be selected.

[0049] Among them, the corpus to be selected can be one or more. The corpus to be selected can be the corpus pre-stored in the corpus. In actual applications, at least one corpus to be selected can be displayed on the display interface, and the user can select from the displayed corpus to be selected through the trigger operation. When a selection trigger operation for any corpus to be selected is detected, the selection trigger operation is responded to, and the currently selected corpus to be selected is used as the original corpus.

[0050] The second method can be: responding to a corpus input operation to obtain the original corpus.

[0051] Among them, the corpus input operation can be understood as an operation of inputting the corpus based on the input device. In actual applications, the corpus can be input based on the input device, and then, when an input completion trigger operation is detected, the input corpus can be used as the original corpus.

[0052] It should be noted that since the corpus to be processed is in text form, the original corpus can be in any form. Therefore, after obtaining the original corpus, if the obtained original corpus is in a form other than text form, the form conversion process can be performed on the obtained original corpus to convert the obtained original corpus into a text-form corpus. Furthermore, the task identification marking process can be performed on the converted corpus. Thus, the corpus to be processed can be obtained.

[0053] Among them, the task identification can be understood as the identification representing the corpus processing task. The corpus processing task can be understood as the task of processing the input corpus to obtain the corresponding corpus processing result. Optionally, the corpus processing task can include translation tasks, summary tasks, review tasks, error identification tasks, etc. Optionally, the task identification can be a pre-set number corresponding to the corpus processing task of each task type. In practical applications, for each corpus processing task, a number corresponding to the corpus processing task of each task type can be pre-set. Furthermore, the mapping relationship between the task type and the number can be established. Thus, the task identification can be marked for the original corpus according to the pre-established mapping relationship. Or, the task identification can also be a pre-set keyword or key phrase corresponding to the corpus processing task of each task type. In practical applications, for each corpus processing task of each task type, the keyword or key phrase that can represent the task type of the corpus processing task can be pre-determined according to the processing method or other data of the corpus processing task, and the determined keyword or key phrase can be used as the task identification corresponding to the task type. Furthermore, the task identification can be marked for the original corpus according to the pre-determined keyword or key phrase. Exemplarily, the original corpus is "Diffusion language models can be so cool", the corpus processing task corresponding to the original corpus is a translation task, and the task identification corresponding to the translation task can be a keyword. For example, the keyword can be "Translate". Furthermore, the corpus to be processed obtained after marking the task identification for the original corpus can be "Translate 'Diffusion language models can be so cool' Answer in Chinese".

[0054] In practical applications, the original corpus can be obtained by editing the corpus. Further, the corpus processing task corresponding to the original corpus can be determined. Furthermore, the task identifier can be marked for the original corpus according to the corpus processing task corresponding to the original corpus. Furthermore, the original corpus marked with the task identifier can be used as the corpus to be processed, so that the target diffusion language model can process the corpus to be processed based on the language identifier. The advantage of such a setting is that: the target diffusion language model can process the corpus to be processed according to the task type corresponding to the corpus to be processed, thereby ensuring that the obtained target prediction result meets the corpus processing requirements.

[0055] S120. Process the corpus to be processed based on the target diffusion language model to obtain a target prediction result corresponding to the corpus to be processed.

[0056] Among them, the target diffusion language model can be understood as a diffusion model that takes corpus data as the input object to process the corpus data. Those skilled in the art can understand that the diffusion model (Diffusion Models) is a generative model based on iterative denoising. According to the type of data distribution it models, the diffusion model can be divided into a continuous diffusion model and a discrete diffusion model. For the generation process of images and audio, since both images and audio belong to continuous types of data, a continuous diffusion model can be used to process the input image or input audio to determine the generation result. The quality of its generation result is significantly higher than that of other generative models. For the generation process of text, since text belongs to discrete type of data, a discrete diffusion model can be used to process the input text to determine the generation result. However, the model construction cost of the discrete diffusion model is relatively high. At the same time, the training objective of the discrete diffusion model is consistent with that of the masked language model. Therefore, the pre-trained masked language model can be used as the diffusion language model, and the masked language model can be trained based on the corpus samples corresponding to the diffusion language model to obtain the target diffusion language model. It should be noted that using the trained masked language model as the diffusion language model reduces the model construction cost of the target diffusion language model.

[0057] Among them, the masked language model (Masked Language Model, MLM) can be a language model used to perform natural language processing tasks. The masked language model can process the corpus including the masked corpus to predict the corpus data masked in the corpus. Specifically, a certain masking rate is used to randomly select some words in the input corpus of the model and mask the selected words. Then, the processed corpus can be input into the trained masked language model for processing to obtain the prediction result, which can be the prediction result corresponding to the words masked in the corpus.

[0058] In this embodiment, the target diffusion language model can be trained based on multiple corpus samples. The masked corpus in the corpus samples corresponds to different masking rates. Among them, the corpus samples can include unmasked corpus and corresponding masked corpus. The masked corpus can be understood as the corpus obtained by masking at least part of the data in the corpus. The masking rate can be understood as the percentage between the amount of masked data in a corpus and the total amount of data in the corpus. Exemplarily, assume that a corpus includes 10 words. After masking any two words in the corpus, a masked corpus is obtained. The masking rate corresponding to this masked corpus can be 20%. It should be noted that the number of masking rates corresponding to the masked corpus in the corpus samples and the values corresponding to each masking rate are randomly determined. That is to say, the starting value of the masking rate and the step size between each masking rate are randomly sampled and determined in the interval of 0-1.

[0059] In practical applications, multiple unmasked training corpora can be obtained. Then, according to different masking rates, the obtained multiple training corpora can be processed to obtain the masked corpora corresponding to the training corpora. Furthermore, the training corpora and the corresponding masked corpora can be used as a corpus sample, and multiple corpus samples can be obtained. Further, the diffusion language model to be trained can be trained based on the constructed multiple corpus samples. Thus, the trained target diffusion language model can be obtained. It should be noted that training the masked language model with corpus samples of different masking rates can be to apply the training method of the diffusion model to the masked language model to activate the generation ability of the masked language model. Thus, the adaptability of the target diffusion language model to downstream application tasks is improved.

[0060] It should be noted that the target diffusion language model can be a single-task type processing model or a multi-task type processing model. Among them, the single-task type processing model can be a model only used to process a specific single corpus processing task. Correspondingly, the multi-task type processing model can be a model capable of processing multiple types of corpus processing tasks.

[0061] In the case where the target diffusion language model is a single-task type processing model, the corpus to be processed input into the target diffusion language model can correspond to the corpus processing task of the same task type, and this corpus processing task is consistent with the corpus processing task corresponding to the target diffusion language model. In the actual application process, after obtaining the corpus to be processed, the corpus to be processed can be input into the target diffusion language model, and the corpus to be processed can be processed based on the target diffusion language model. Thus, the corpus prediction result corresponding to the corpus to be processed can be output.

[0062] In the case where the target diffusion language model is a multi-task type model, after obtaining the corpus to be processed, the corpus to be processed can be input into the target diffusion language model. Further, the task type corresponding to the corpus to be processed can be determined based on the target diffusion language model. Furthermore, the corpus to be processed can be processed based on the target diffusion language model according to the corresponding task type. Thus, a corpus processing result corresponding to the task type can be obtained.

[0063] In practical applications, after obtaining the corpus to be processed, the corpus to be processed can be input into the target diffusion language model. Furthermore, the corpus to be processed can be processed based on the target diffusion language model. Thus, a target prediction result corresponding to the corpus to be processed can be obtained.

[0064] Among them, the target prediction result can be the prediction result obtained after the target diffusion language model processes the corpus to be processed according to the corpus processing task corresponding to the corpus to be processed. Optionally, the target prediction result can include any one of a translation result, an abstract result, a review result, and an error recognition result corresponding to the corpus to be processed. Exemplarily, if the corpus processing task corresponding to the corpus to be processed is a text translation task of translating an English text into a Chinese text, and the corpus to be processed is an English text to be translated, then the target prediction result can be the Chinese text corresponding to the English text after translation; if the corpus processing task corresponding to the corpus to be processed is a task of determining the abstract of the whole article, and the corpus to be processed is an article to be processed, then the target prediction result can be the article abstract corresponding to the article; if the corpus processing task corresponding to the corpus to be processed is a task of determining the review of the whole article, and the corpus to be processed is a text to be processed, then the target prediction result can be the text review corresponding to the text; if the corpus processing task corresponding to the corpus to be processed is a task of identifying error sentences, and the corpus to be processed is a text to be identified, then the target prediction result can be the error sentence identification result corresponding to the text.

[0065] S130. Display the target prediction result.

[0066] In practical applications, in the case of obtaining the target prediction result, the target prediction result can be displayed. Thus, the target prediction result corresponding to the corpus to be processed can be displayed on the target display interface. Among them, the target display interface can be a pre-determined display interface for displaying corpus prediction results.

[0067] In practical applications, for different application scenarios, the display position of the target prediction result can also change correspondingly. Exemplarily, for the corpus translation scenario integrated in a web page, the display interface can include two display areas at the same time. One of the display areas can be the receiving area for the corpus to be processed, and the other display area can be the display area for the target prediction result. In the case of detecting a corpus receiving operation for the receiving area of the corpus to be processed, the corpus to be processed can be received. Further, in the case of detecting a trigger operation for the corpus translation control, the trigger operation is responded to, and the translation result corresponding to the corpus to be processed is displayed in the display area of the target prediction result. Or, for the video playback scenario, the target video can be played in the display interface. When a trigger operation for the line translation control is detected, the translation result corresponding to the original video lines can be synchronously displayed in the display interface. At this time, the display area of the target prediction result is the video playback interface. Or, for the scenario of summarizing the meeting content, the terminal device pre-deployed with the target diffusion language model can receive the meeting audio data or the meeting text data. Further, in the case of detecting a trigger operation for the summarization processing control, the trigger operation can be responded to, and the corresponding meeting summary can be displayed in the display interface of the terminal device.

[0068] The technical solution of the embodiment of the present disclosure realizes the effect of processing corpus data based on the diffusion model by receiving the corpus to be processed, and further processing the corpus to be processed based on the target diffusion language model to obtain the target prediction result corresponding to the corpus to be processed. Finally, the target prediction result is displayed, which solves the problems such as poor quality, low accuracy, and non-conformity with user requirements of the corpus generation result when using the autoregressive language model to process corpus data in the related art. It realizes that the target diffusion language model can process corpus data based on the principle of the diffusion model, and the obtained corpus data processing result can meet the requirements of the corpus processing task, improves the accuracy of the corpus data processing result, and enhances the user experience.

[0069] Figure 2 It is a schematic flowchart of another data processing method provided by the embodiment of the present disclosure. Based on the above embodiment, the technical solution of this embodiment further refines how to process the corpus to be processed according to the target diffusion language model in the case that the target diffusion language model is a multi-task type processing model. The specific implementation can refer to the description of this embodiment. Among them, the same or similar technical features as those in the foregoing embodiment will not be repeated here.

[0070] As Figure 2 shown, the method of this embodiment may specifically include:

[0071] S210. Receive the corpus to be processed.

[0072] S220. Identify the task identifier carried by the corpus to be processed based on the target diffusion language model, and determine the task type corresponding to the corpus to be processed based on the task identifier.

[0073] Among them, the task type can be understood as the type of the corpus processing task to be processed. Optionally, the task type may include text translation tasks, article summary tasks, text review tasks, error text recognition tasks, etc. In the actual application process, multiple task types of corpus processing tasks can be determined in advance. Furthermore, for each task type, a task identifier corresponding to the task type can be set in advance, and thus the task identifier corresponding to each determined task type can be determined. After that, a mapping relationship between the task type and the task identifier can be established and deployed in the target diffusion language model.

[0074] In actual applications, when the target diffusion language model is a multi-task type processing model, after inputting the corpus to be processed into the target diffusion language model, the task identifier carried by the corpus to be processed can be identified based on the target diffusion language model to determine the task identifier carried by the corpus to be processed. Further, according to the mapping relationship pre-deployed in the target diffusion language model and the identified task identifier, the task type corresponding to the corpus to be processed can be determined. Thus, the corpus to be processed can be processed according to the corpus processing task corresponding to the task type.

[0075] S230. Process the corpus to be processed based on the target diffusion language model to obtain a target prediction result that matches the task type.

[0076] In this embodiment, after determining the task type corresponding to the corpus to be processed, the corpus to be processed can be processed based on the target diffusion language model according to the corpus processing task corresponding to the determined task type. Thus, a target prediction result that matches the task type can be obtained.

[0077] Optionally, if the task type corresponding to the corpus to be processed is a text translation type, the target prediction result may be the translation result corresponding to the corpus to be processed; if the task type corresponding to the corpus to be processed is an article summary type, the target prediction result may be the summary result corresponding to the corpus to be processed; if the task type corresponding to the corpus to be processed is a text review type, the target prediction result may be the review result corresponding to the corpus to be processed; if the task type corresponding to the corpus to be processed is a text error recognition type, the target prediction result may be the error recognition result corresponding to the corpus to be processed.

[0078] It should be noted that if the target diffusion language model is a single-task type processing model, the target diffusion language model is trained based on single-task type training samples. The target diffusion language model can be used only for performing corpus processing tasks of a single task type. In this case, after the corpus to be processed is input into the target diffusion language model, the corpus to be processed can be directly processed. Furthermore, the target prediction result corresponding to the corpus to be processed can be obtained.

[0079] S240. Display the target prediction result.

[0080] The technical solution of the embodiment of the present disclosure, by receiving the corpus to be processed, and then, based on the target diffusion language model, identifying the task identifier carried by the corpus to be processed, to determine the task type corresponding to the corpus to be processed based on the task identifier, and then, based on the target diffusion language model, processing the corpus to be processed to obtain the target prediction result matching the task type, and finally, displaying the target prediction result, realizes the effect of enabling the target diffusion language model to process the corpus according to the task type corresponding to the corpus, demonstrates the multi-task type processing ability and flexibility of the model, improves the adaptability of the model to multi-task types, and enhances the user experience.

[0081] Figure 3 It is a schematic flowchart of a data processing method provided by an embodiment of the present disclosure. Based on the technical solution of the above embodiment, before processing the corpus to be processed based on the target diffusion language model, the training corpus corresponding to different task types can be obtained, and the training corpus can be masked according to a preset different masking rate to obtain the masked corpus of the training corpus at different masking rates. Furthermore, a corpus sample can be constructed according to the training corpus and the corresponding masked corpus, so that the diffusion language model can be trained based on the corpus sample to obtain the target diffusion language model. The specific implementation manner can refer to the description of this embodiment. Among them, the technical features that are the same or similar to the foregoing embodiments will not be described in detail here.

[0082] As Figure 3 shown, the method of this embodiment may specifically include:

[0083] S310. Obtain the training corpus corresponding to at least one task type.

[0084] Among them, the training corpus can be understood as the sample corpus used to execute the training process. Similar to the corpus to be processed, the training corpus can be a sample corpus in text form. In this embodiment, the training corpus can be the corpus input by the input device; or, the corpus pre-stored in the corpus library; or, the corpus segmented by the corpus segmentation model, etc. It should be noted that for different task types, the corresponding training corpus can be the same or different, and the embodiments of the present disclosure do not make specific limitations on this. It should also be noted that the task type corresponding to the obtained training corpus can be one or more, and the type of the task type corresponding to the training corpus can match the type of the diffusion language model to be trained. When the diffusion language model to be trained is a single-task type processing model, the type of the task type corresponding to the obtained training corpus can be one, so that the trained target diffusion language model can execute the corpus processing task of a specific task type; when the diffusion language model to be trained is a multi-task type processing model, the type of the task type corresponding to the obtained training corpus can be multiple, so that the trained target diffusion language model can be applicable to the corpus processing tasks of multiple task types.

[0085] In practical applications, before training the diffusion language model to be trained, it is necessary to pre-construct multiple corpus samples to train the model based on the corpus samples. To improve the accuracy of the model, the corpus samples can be constructed as many and rich as possible. Specifically, at least one task type of corpus processing task to be executed by the target diffusion language model can be determined. Furthermore, the corresponding training corpus under the corresponding task type can be obtained according to the determined task type. Thus, the corpus samples can be constructed according to the obtained training corpus.

[0086] S320. Perform masking processing on the training corpus according to different pre-set masking rates to obtain the masked corpus corresponding to the training corpus.

[0087] Among them, the masking rate can be understood as the percentage of the amount of data masked in a corpus to the total amount of data in the corpus. In this embodiment, the different pre-set masking rates can be the masking rates randomly sampled from 0 to 1. That is to say, the number of masking rates and the value corresponding to each masking rate are randomly determined. The advantage of processing the training corpus with different masking rates to construct corpus samples is that the training target of the target diffusion language model can be determined as the full masking rate, which improves the processing ability of the target diffusion language model for corpora with different masking rates and enhances the scalability of the model.

[0088] It should be noted that the training corpus is masked according to different masking rates to construct corpus samples, which can be understood as the process of constructing training samples according to the training method of the diffusion model. Specifically, for the diffusion model performing the image generation task, its corresponding training process can be a process of iteratively adding image noise. Assuming the original image is a real image, image noise is added to the image through multiple accumulations. In this process, as the number of times increases, the obtained image gets closer and closer to pure noise. Furthermore, according to the images obtained in the above process, image samples for training the diffusion model can be constructed to complete the training of the diffusion model. For the diffusion model mentioned in the embodiments of the present disclosure for performing corpus processing tasks, the training corpus can be masked according to different masking rates to construct corpus samples for training the diffusion model. Thus, the diffusion model can be trained according to the constructed corpus samples to obtain a target diffusion language model that can perform corpus processing tasks.

[0089] In practical applications, after obtaining the training corpus corresponding to at least one task type, for the training corpus corresponding to each task type, the training corpus can be masked according to different preset masking rates. Furthermore, masked corpora with different masking rates corresponding to the training corpus can be obtained. It should be noted that the number of masked corpora corresponding to the training corpus is consistent with the number of masking rates for processing the training corpus. It should also be noted that for different training corpora, the number of masking rates corresponding thereto and the values corresponding to each masking rate can be the same or different, and the embodiments of the present disclosure do not make specific limitations thereon.

[0090] S330. Determine corpus samples based on the training corpus and the corresponding masked corpora.

[0091] Among them, the corpus sample is the training sample required for training the diffusion language model to be trained.

[0092] In practical applications, after obtaining the training corpus and the masked corpora corresponding to the training corpus, corpus samples can be constructed according to the training corpus and its corresponding masked corpora. Each corpus sample can include the training corpus and the masked corpora with different masking rates corresponding to the training corpus.

[0093] It should be noted that if the target diffusion language model is a multi-task type processing model, in order to enable the trained target diffusion language model to identify the task identifier corresponding to the model input corpus, the task identifier can be added to the corpus sample during the model training process to train the model based on the corpus sample with the added task identifier. Thus, the target diffusion language model can be enabled to identify the task identifier of the model input corpus.

[0094] Based on this, on the basis of the above technical solutions, it further includes: if the target diffusion language model is a multi-task type processing model, corresponding task identifiers are marked for the corpus samples to update the corpus samples.

[0095] Among them, the task identifier matches the task type.

[0096] In practical applications, if the model training objective is to obtain a target diffusion language model of a multi-task type processing model. After determining multiple corpus samples, the corresponding task identifiers can be determined according to the task types corresponding to the training corpora in the corpus samples, and the corresponding task identifiers are marked for the corpus samples. Thus, the corpus samples can be updated according to the marked corpus samples. The advantage of such a setting is that: it improves the adaptability of the target diffusion language model to multi-task type corpus processing tasks and enhances the scalability of the target diffusion language model.

[0097] S340. Train to obtain the target diffusion language model.

[0098] In this embodiment, after determining the corpus samples, the model to be trained can be trained based on the corpus samples. Thus, the target diffusion language model can be obtained.

[0099] In practical applications, among various types of models that can process corpus data, the training objective of the discrete diffusion model is consistent with that of the masked language model, both of which are to process the corpus data based on the trained language model to obtain the corpus processing result corresponding to the corpus data. At the same time, the masking rate used in the training process of the masked language model can correspond to the noise used in the training process of the diffusion model. Since in the actual application process, the training process of the discrete diffusion model is relatively cumbersome and the training cost is relatively high, and at the same time, the research and application in the field of model construction of the masked language model are relatively sufficient, that is, the trained masked language model is relatively easy to obtain. Therefore, in order to reduce the training cost of the diffusion model and improve the training efficiency of the diffusion model, the trained masked language model can be used as the diffusion language model to be trained. Furthermore, when training the diffusion language model to be trained, the model parameters of the masked language model can be adjusted according to the constructed corpus samples. Thus, the masked language model with adjusted parameters can be used as the trained target diffusion language model.

[0100] Optionally, training to obtain the target diffusion language model includes: obtaining a pre-trained masked language model; training and processing the masked language model based on the corpus samples to obtain the target diffusion language model.

[0101] Among them, the pre-trained masked language model can be understood as a trained masked language model. Those skilled in the art can understand that the masked language model (MLM) can be a language model for performing natural language processing tasks. The masked language model can process the corpus including the masked corpus to predict the corpus data masked in the corpus. Specifically, a part of words and phrases are randomly selected from the input corpus of the model at a certain masking rate, and the selected part of words and phrases are masked. Then, the processed corpus can be input into the trained masked language model for processing, and the prediction result can be obtained. The prediction result can be the prediction result corresponding to the words and phrases masked in the corpus.

[0102] It should be noted that the masked language model is obtained after training based on the training objective of a fixed masking rate. However, the training objective of the target diffusion language model is the full masking rate (random sampling from 0 to 1). The masked language model still lacks in language generation ability and is not sufficient to handle downstream application tasks. Therefore, after obtaining the masked language model, the masked language model can be trained based on the constructed corpus sample so that the trained target diffusion language model can be adapted to downstream application tasks.

[0103] In practical applications, the masked language model that has completed the pre-training stage can be obtained first. Further, the masked language model can be trained according to the constructed language sample to adjust the model parameters of the masked language model. Thus, the target diffusion language model can be obtained. The advantage of this setting is that: using the trained masked language model as the diffusion language model to be trained reduces the training cost of the target diffusion language model. Furthermore, after training the masked language model with the corpus sample, the adaptability of the target diffusion language model to downstream application tasks is improved.

[0104] Optionally, training the masked language model based on the corpus sample to obtain the target diffusion language model includes: inputting the training corpus in the corpus sample into the masked language model to obtain the actual output corpus; determining the loss value based on the actual output corpus and the corresponding masked corpus; correcting the model parameters in the masked language model based on the loss value, and taking the convergence of the loss function in the masked language model as the training objective to obtain the target diffusion language model.

[0105] It should be noted that for each corpus sample, the above method can be used to train it, so as to obtain the target diffusion language model.

[0106] Among them, the actual output corpus can be the corpus prediction result output after inputting the training corpus into the masked language model, and the corpus prediction result matches the task type of the training corpus. The loss value can be understood as the difference value between the actual output corpus and the corresponding masked corpus. The loss function can be determined based on the loss value and is used to characterize the difference degree between the actual output and the theoretical output.

[0107] In practical applications, after obtaining the masked language model, for each corpus sample, the training corpus in the corpus sample can be input into the masked language model to process the training corpus based on the masked language model, and the actual output corpus corresponding to the training corpus can be obtained. Further, the actual output corpus can be compared with the masked corpus in the current corpus sample to determine the loss value. Furthermore, the model parameters in the masked language model can be corrected according to the loss value. After that, the training error of the loss function in the masked language model, that is, the loss parameter, can be used as the condition for detecting whether the current loss function reaches convergence. For example, whether the training error is less than the preset error or whether the error change trend tends to be stable, or whether the current model iteration times is equal to the preset times, etc. If it is detected that the convergence condition is reached, such as the training error of the loss function reaches less than the preset error or the error change tends to be stable, it indicates that the current masked language model training is completed, and at this time, the iterative training can be stopped. If it is detected that the current does not reach the convergence condition, the current training sample can be further obtained to train the masked language model until the training error of the loss function is within the preset range. When the training error of the loss function reaches convergence, the currently trained masked language model can be used as the target diffusion language model. The advantage of such a setting is that: on the basis of improving the model accuracy and task matching degree, the training cost of the model is reduced, and the effect of activating the generation ability of the masked language model by using the diffusion model is achieved. Furthermore, the target diffusion language model can be a large-scale language model constructed based on the theoretical framework of the diffusion model.

[0108] S350. Receive the corpus to be processed.

[0109] S360. Process the corpus to be processed based on the target diffusion language model to obtain the target prediction result corresponding to the corpus to be processed.

[0110] S370. Display the target prediction result.

[0111] It should be noted that after obtaining the trained target diffusion language model, the input to-be-processed corpus can be processed based on the target diffusion language model. Furthermore, a corpus processing result corresponding to the to-be-processed corpus can be obtained, and this corpus processing result matches the task type corresponding to the to-be-processed corpus. The specific processing process of the target diffusion language model for processing the to-be-processed corpus can refer to the content described in steps S110 - S130.

[0112] In the technical solution of this embodiment of the present disclosure, by obtaining the training corpus corresponding to at least one task type, then, performing masking processing on the training corpus according to different pre-set masking rates to obtain the masked corpus corresponding to the training corpus. Furthermore, based on the training corpus and the corresponding masked corpus, corpus samples are determined, and a target diffusion language model is trained. Further, the to-be-processed corpus is received, and then, the to-be-processed corpus is processed based on the target diffusion language model to obtain the target prediction result corresponding to the to-be-processed corpus. Finally, the target prediction result is displayed, achieving the effect of training the masked language model with the model training objective of the full masking rate, so that the target diffusion language model adapts to downstream application tasks, improving the generalization, versatility, and practicality of the model. Furthermore, the prediction accuracy of the model for the input corpus and the task type matching degree of the prediction result are improved.

[0113] Figure 4 It is a schematic structural diagram of a data processing device provided by an embodiment of the present disclosure, as Figure 4 shown, the device includes: a corpus receiving module 410, a corpus processing module 420, and a prediction result display module 430.

[0114] Among them, the corpus receiving module 410 is used to receive the to-be-processed corpus; the corpus processing module 420 is used to process the to-be-processed corpus based on the target diffusion language model to obtain the target prediction result corresponding to the to-be-processed corpus; wherein, the target diffusion language model is trained based on multiple corpus samples, and the masked corpus in the corpus samples corresponds to different masking rates; the prediction result display module 430 is used to display the target prediction result.

[0115] Based on the above technical solutions, the device further includes: a corpus editing module and a corpus marking module.

[0116] The corpus editing module is used to edit the original corpus before receiving the to-be-processed corpus;

[0117] The corpus marking module is used to mark the task identifier for the original corpus to obtain the to-be-processed corpus, so that the target diffusion language model processes the to-be-processed corpus based on the task identifier.

[0118] Based on the above technical solutions, the target diffusion language model is a multi-task type processing model, and the corpus processing module 420 includes: an identification recognition unit and a corpus processing unit.

[0119] The identification recognition unit is used to identify the task identification carried by the corpus to be processed based on the target diffusion language model, so as to determine the task type corresponding to the corpus to be processed based on the task identification;

[0120] The corpus processing unit is used to process the corpus to be processed based on the target diffusion language model to obtain a target prediction result matching the task type.

[0121] Based on the above technical solutions, the device further includes: a training corpus acquisition module, a masked corpus determination module, and a corpus sample determination module.

[0122] The training corpus acquisition module is used to acquire the training corpus corresponding to at least one task type;

[0123] The masked corpus determination module is used to perform masked processing on the training corpus according to different pre-set masking rates to obtain a masked corpus corresponding to the training corpus;

[0124] The corpus sample determination module is used to determine the corpus sample based on the training corpus and the corresponding masked corpus.

[0125] Based on the above technical solutions, the device further includes: a corpus sample update module.

[0126] The corpus sample update module is used to mark the corresponding task identification for the corpus sample if the target diffusion language model is a multi-task type processing model, so as to update the corpus sample;

[0127] Wherein, the task identification matches the task type.

[0128] Based on the above technical solutions, the device further includes: a model training module.

[0129] The model training module is used to train and obtain the target diffusion language model after obtaining the corpus sample;

[0130] The model training module includes: a model acquisition unit and a model training unit.

[0131] The model acquisition unit is used to acquire a pre-trained masked language model;

[0132] The model training unit is used to train and process the masked language model based on the corpus sample to obtain the target diffusion language model.

[0133] Based on the above technical solutions, the model training unit includes: a corpus input subunit, a loss value determination subunit, and a model parameter correction subunit.

[0134] The corpus input subunit is configured to input the training corpus in the corpus sample into the masked language model to obtain an actual output corpus;

[0135] The loss value determination subunit is configured to determine a loss value based on the actual output corpus and the corresponding masked corpus;

[0136] The model parameter correction subunit is configured to correct the model parameters in the masked language model based on the loss value, and take the convergence of the loss function in the masked language model as a training target to obtain the target diffusion language model.

[0137] Based on the above technical solutions, the target prediction result includes any one of a translation result, an abstract result, a review result, and an error recognition result corresponding to the corpus to be processed.

[0138] The technical solution of the embodiment of the present disclosure receives the corpus to be processed, and further processes the corpus to be processed based on the target diffusion language model to obtain a target prediction result corresponding to the corpus to be processed, achieving the effect of processing corpus data based on the diffusion model. Finally, the target prediction result is displayed, solving the problems such as poor quality of corpus generation results, low accuracy, and corpus generation results not meeting user requirements when using the autoregressive language model to process corpus data in the related art. It realizes that the target diffusion language model can process corpus data based on the principle of the diffusion model, and the obtained corpus data processing result can meet the requirements of the corpus processing task, improving the accuracy of the corpus data processing result and enhancing the user experience.

[0139] The data processing device provided by the embodiment of the present disclosure can execute the data processing method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0140] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiment of the present disclosure.

[0141] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. The following refers to Figure 5 , which shows an electronic device suitable for implementing the embodiment of the present disclosure (such as Figure 5Schematic structural diagram of the terminal device or server) 500 therein. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0142] As Figure 5 shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage device 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The editing / output (I / O) interface 505 is also connected to the bus 504.

[0143] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0144] Particularly, according to the embodiments of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.

[0145] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0146] The electronic device provided in the embodiments of the present disclosure and the data processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0147] The embodiments of the present disclosure provide a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the data processing method provided in the above embodiments.

[0148] It should be noted that the computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0149] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0150] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.

[0151] The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: receive the corpus to be processed; process the corpus to be processed based on the target diffusion language model to obtain a target prediction result corresponding to the corpus to be processed; wherein the target diffusion language model is trained based on a plurality of corpus samples, and the masked corpus in the corpus samples corresponds to different masking rates; display the target prediction result.

[0152] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0154] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least two Internet protocol addresses".

[0155] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0156] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash memory), optical fibers, portable Compact Disc Read Only Memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0157] According to one or more embodiments of the present disclosure, [Example 1] provides a data processing method, the method comprising:

[0158] Receive the corpus to be processed;

[0159] Process the corpus to be processed based on the target diffusion language model to obtain a target prediction result corresponding to the corpus to be processed; wherein, the target diffusion language model is trained based on multiple corpus samples, and the masked corpus in the corpus samples corresponds to different masking rates;

[0160] Display the target prediction result.

[0161] According to one or more embodiments of the present disclosure, [Example 2] provides the method of Example 1, further including:

[0162] Optionally, before receiving the corpus to be processed, it further includes:

[0163] Edit the original corpus;

[0164] Mark a task identifier for the original corpus to obtain the corpus to be processed, so that the target diffusion language model processes the corpus to be processed based on the task identifier.

[0165] According to one or more embodiments of the present disclosure, [Example 3] provides the method of Example 1, further including:

[0166] Optionally, the target diffusion language model is a multi-task type processing model, and the processing the corpus to be processed based on the target diffusion language model to obtain a target prediction result corresponding to the corpus to be processed includes:

[0167] Identify the task identifier carried by the corpus to be processed based on the target diffusion language model, so as to determine the task type corresponding to the corpus to be processed based on the task identifier;

[0168] Process the corpus to be processed based on the target diffusion language model to obtain a target prediction result matching the task type.

[0169] According to one or more embodiments of the present disclosure, [Example 4] provides the method of Example 1, further including:

[0170] Optionally, it further includes: obtaining training corpus corresponding to at least one task type; performing masking processing on the training corpus according to preset different masking rates to obtain masked corpus corresponding to the training corpus; determining the corpus samples based on the training corpus and the corresponding masked corpus.

[0171] According to one or more embodiments of the present disclosure, [Example 5] provides the method of Example 4, further including:

[0172] Optionally, it further includes: if the target diffusion language model is a multi-task type processing model, corresponding task identifiers are marked for the corpus samples to update the corpus samples;

[0173] Among them, the task identifier matches the task type.

[0174] According to one or more embodiments of the present disclosure, [Example Six] provides the method of Example Four, and further includes:

[0175] Optionally, after obtaining the corpus samples, it further includes: training to obtain the target diffusion language model; the training to obtain the target diffusion language model includes: obtaining a pre-trained masked language model; training and processing the masked language model based on the corpus samples to obtain the target diffusion language model.

[0176] According to one or more embodiments of the present disclosure, [Example Seven] provides the method of Example Six, and further includes:

[0177] Optionally, the training and processing the masked language model based on the corpus samples to obtain the target diffusion language model includes:

[0178] Inputting the training corpus in the corpus samples into the masked language model to obtain the actual output corpus;

[0179] Determining a loss value based on the actual output corpus and the corresponding masked corpus;

[0180] Correcting the model parameters in the masked language model based on the loss value, and taking the convergence of the loss function in the masked language model as the training objective to obtain the target diffusion language model.

[0181] According to one or more embodiments of the present disclosure, [Example Eight] provides the method of Example One, and further includes:

[0182] Optionally, the target prediction result includes any one of a translation result, an abstract result, a review result, and an error recognition result corresponding to the corpus to be processed.

[0183] According to one or more embodiments of the present disclosure, [Example Nine] provides a data processing device, and the device includes:

[0184] A corpus receiving module, configured to receive the corpus to be processed;

[0185] A corpus processing module, configured to process the corpus to be processed based on the target diffusion language model to obtain a target prediction result corresponding to the corpus to be processed; wherein, the target diffusion language model is trained based on a plurality of corpus samples, and the masked corpus in the corpus samples corresponds to different masking rates;

[0186] A prediction result display module for displaying the target prediction result.

[0187] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0188] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0189] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A data processing method, characterized in that, Including: Receiving the corpus to be processed; Processing the corpus to be processed based on the target diffusion language model to obtain a target prediction result corresponding to the corpus to be processed; wherein, the target diffusion language model is trained based on multiple corpus samples, and the masked corpus in the corpus samples corresponds to different masking rates; the target prediction result includes any one of a translation result, an abstract result, a review result, and an error recognition result corresponding to the corpus to be processed; Displaying the target prediction result; Before receiving the corpus to be processed, it further includes: editing the original corpus; marking a task identifier for the original corpus to obtain the corpus to be processed, so that the target diffusion language model processes the corpus to be processed based on the task identifier; wherein, the task identifier is used to identify the corpus processing task corresponding to the corpus to be processed.

2. The method according to claim 1, wherein The target diffusion language model is a multi-task type processing model, and the processing the corpus to be processed based on the target diffusion language model to obtain a target prediction result corresponding to the corpus to be processed includes: Identifying the task identifier carried by the corpus to be processed based on the target diffusion language model to determine the task type corresponding to the corpus to be processed based on the task identifier; Processing the corpus to be processed based on the target diffusion language model to obtain a target prediction result matching the task type.

3. The method according to claim 1, wherein It further includes: Obtaining the training corpus corresponding to at least one task type; Performing masking processing on the training corpus according to different preset masking rates to obtain a masked corpus corresponding to the training corpus; Determining the corpus sample based on the training corpus and the corresponding masked corpus.

4. The method according to claim 3, characterized in that, It further includes: If the target diffusion language model is a multi-task type processing model, then marking a corresponding task identifier for the corpus sample to update the corpus sample; wherein, the task identifier matches the task type.

5. The method according to claim 3, characterized in that, After obtaining the corpus sample, it further includes: Training to obtain the target diffusion language model; The training to obtain the target diffusion language model includes: Obtaining a pre-trained masked language model; Training and processing the masked language model based on the corpus sample to obtain the target diffusion language model.

6. The method according to claim 5, wherein The training and processing the masked language model based on the corpus sample to obtain the target diffusion language model includes: Inputting the training corpus in the corpus sample into the masked language model to obtain an actual output corpus; Determining a loss value based on the actual output corpus and the corresponding masked corpus; Correcting the model parameters in the masked language model based on the loss value, and taking the convergence of the loss function in the masked language model as the training target to obtain the target diffusion language model.

7. A data processing device, characterized in that, Including: A corpus receiving module for receiving the corpus to be processed; A corpus processing module, configured to process the to-be-processed corpus based on a target diffusion language model to obtain a target prediction result corresponding to the to-be-processed corpus; wherein, the target diffusion language model is trained based on a plurality of corpus samples, and the masked corpus in the corpus samples corresponds to different masking rates; the target prediction result includes any one of a translation result, an abstract result, a review result, and an error recognition result corresponding to the to-be-processed corpus A prediction result display module, configured to display the target prediction result; The data processing device further includes: a corpus editing module and a corpus marking module; The corpus editing module is configured to edit the original corpus before receiving the to-be-processed corpus; The corpus marking module is configured to mark a task identifier for the original corpus to obtain the to-be-processed corpus, so that the target diffusion language model processes the to-be-processed corpus based on the task identifier; wherein, the task identifier is used to identify a corpus processing task.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method according to any one of claims 1-6.

9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to execute the data processing method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-task model training method, processing method, electronic equipment and storage medium

    CN114461366A