Text processing method and device, equipment, medium and program product
Through the target text model trained by text error correction tasks based on multiple association relationships, the text is corrected and error-corrected and error-correcting processing problem in the prior art is solved, and more efficient error correction capabilities are achieved and resource dependence is reduced.
Patent Information
- Application Number
- CN202510220459.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is not efficient in text error correction processing and relies on excessive vocabulary lists, knowledge bases or graphs, which increases the complexity and operating costs of the model.
A text processing method is provided, by obtaining the first text and performing error correction processing based on the target text model, which is obtained based on a variety of text error correction tasks with associated relationships. This method trains the model by designing multiple text error correction tasks with correlation, so that the model makes full use of the commonality and complementary information of the data during the training process, and enhances the model's recognition and recovery ability on error correction tasks.
It improves the efficiency of text error correction processing, greatly improves the recognition and resilience of the model in error correction tasks, reduces dependence on resources such as vocabulary lists, knowledge bases or graphs, and reduces the complexity and operating costs of the model.
Smart Images

Figure CN120146034A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of natural language processing, and particularly to a text processing method, apparatus, device, medium, and program product. Background Art
[0002] A large language model refers to a deep learning model trained using a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. Text error correction is one of the common tasks, used to identify and restore misspelled words in text, and is usually applied to short video products, novel products, document office products, etc., which can help check spelling errors and improve the user experience.
[0003] In related technologies, there are still limitations in the types of errors supported for splicing error correction. For example, it is mainly limited to equal-length variant types. At the same time, relying too heavily on resources such as word lists, knowledge bases, or graphs increases the complexity and operating cost of the model. Therefore, there is still a further need to optimize the processing method for the text error correction task. Summary of the Invention
[0004] In view of this, the present disclosure provides a text processing method, apparatus, device, medium, and program product to solve the problem of low efficiency in current text error correction processing.
[0005] In a first aspect, the present disclosure provides a text processing method, the method comprising:
[0006] Obtain a first text;
[0007] Perform text error correction processing on the first text based on a target text model, to obtain a second text, where the target text model is trained based on multiple text error correction tasks having an associated relationship.
[0008] In a second aspect, the present disclosure provides a text processing apparatus, the apparatus comprising:
[0009] A text acquisition module, configured to obtain a first text;
[0010] An error correction processing module, configured to perform text error correction processing on the first text based on a target text model, to obtain a second text, where the target text model is trained based on multiple text error correction tasks having an associated relationship.
[0011] In a third aspect, the present disclosure provides an electronic device, comprising: a memory and a processor, which are communicatively connected to each other, where the memory stores computer instructions, and the processor executes the computer instructions to execute the text processing method according to the first aspect or any corresponding implementation thereof.
[0012] Fourthly, the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the text processing method according to the first aspect or any corresponding embodiment thereof.
[0013] Fifthly, the present disclosure provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the text processing method according to the first aspect or any corresponding embodiment thereof.
[0014] The text processing method provided in this embodiment includes obtaining a first text; performing text error correction processing on the first text based on a target text model to obtain a second text, where the target text model is trained based on multiple text error correction tasks with associated relationships. This method can perform error correction processing on the first text in multiple application scenarios to obtain a second text. By designing multiple associated text error correction tasks to train the model, the model can make full use of the commonality and complementary information of the data during the training process, enhance the recognition and restoration ability of the model in error correction tasks, and improve the efficiency of text error correction processing. Description of the Drawings
[0015] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is a flowchart of the text processing method according to an embodiment of the present disclosure;
[0017] Figure 2 is a flowchart of the method for determining the target text model according to an embodiment of the present disclosure;
[0018] Figure 3 is a block diagram of the structure of the text processing device according to an embodiment of the present disclosure;
[0019] Figure 4 is a schematic diagram of the hardware structure of the electronic device according to an embodiment of the present disclosure. Detailed Embodiments
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0021] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the users and the users' authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0022] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested to be executed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that executes the operations of the technical solutions of the present disclosure according to the prompt message.
[0023] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0024] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0025] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations, and related provisions.
[0026] Spelling correction technology is used to identify and correct misspelled words in text, and can be used in various types of products, such as short video products, novel products, document office products, etc. Through the spelling correction function, the works published in the products can be reviewed, the incorrect content can be corrected or targeted processing can be carried out, so that the content of the works can be better spread, and at the same time, the reading experience of users can be improved. In related technologies, there are deficiencies in both the types of errors supported for spelling correction and the ability to identify and restore. Based on this, the embodiments of the present disclosure provide a text processing method, device, equipment, medium, and program product.
[0027] According to an embodiment of the present disclosure, an embodiment of a text processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0028] In this embodiment, a text processing method is provided, which can be used in terminals such as computers and tablets. Figure 1 It is a flowchart of the text processing method according to an embodiment of the present disclosure, as Figure 1 shown, and the process includes the following steps:
[0029] Step S101, obtain a first text.
[0030] The first text is an unprocessed original text string that may contain errors, representing the original text that may need error correction processing. The source of the first text depends on the application scenario of the embodiment of the present disclosure. The application scenario includes reviewing whether there are areas that need to be rectified in the comment content input by users in the comment area of a video software. Corresponding to this application scenario, the first text is the comment input or published by the user; the application scenario also includes reviewing whether there are areas that need to be rectified in the user comment area of a reading software and the reading content published by the author. Corresponding to this application scenario, the first text is the comment output or published by the user in the comment area and the reading content. The form of the first text is not limited, and it can be a sentence or a multiple-choice question for word discrimination, etc.
[0031] Step S102, perform text error correction processing on the first text based on a target text model to obtain a second text.
[0032] Take the first text as the input of the target text model. The target text model is a natural language processing model that has learned the ability to recognize and correct text errors through a large amount of training data. The target text model will analyze the input first text, identify the possible errors therein and correct the errors. The possible errors in the first text include typos, meaningless characters, etc. After the target text model performs text error correction processing on the first text, it outputs the modified text, that is, the second text. Compared with the first text, the errors in the second text have been corrected. The correction can be to correct the typos, or to change the special text in the first text to specific characters, or to identify the split characters in the first text. For example, if the user writes "good" as "woman", the target text model accurately identifies the illegal word and makes corresponding processing.
[0033] The target text model is trained based on multiple text error correction tasks with associated relationships. The target text model is obtained through multi-task model training. Multi-tasking improves the learning effect of each task by learning multiple related tasks, thereby enhancing the generalization ability and performance of the model, and obtaining the trained target text model. In the text error correction tasks with associated relationships, the input items of each text error correction task have associated relationships, and there are various types of text error correction tasks, such as: determining whether there are errors in the text, extracting the errors in the text, correcting the errors in the text, etc.
[0034] Specifically, during the training process, a large number of sample texts are collected. The sample texts include the original incorrect texts and the corresponding correct texts. The corresponding correct texts can be the complete texts after correcting the original incorrect texts, or the correct texts designed based on the tasks. Exemplarily, if the error correction task is to determine whether there are errors, then the correct text corresponding to the original incorrect text is "yes" or "no", and the correct text can be the label of the original incorrect text. That is, each text error correction task corrects a certain type of error in the text. By simultaneously learning these text error correction tasks in the large model, the parameters and structure of the large model base are optimized. Through ablation experiments, it is proved that the effect of jointly training multiple designed text error correction tasks with associated relationships is significant.
[0035] The text processing method provided in this embodiment includes obtaining a first text; performing text error correction processing on the first text based on the target text model to obtain a second text, where the target text model is trained based on multiple text error correction tasks with associated relationships. This method can perform error correction processing on the first text in multiple application scenarios to obtain a second text. By designing multiple associated text error correction tasks to train the model, the model can fully utilize the commonality and complementary information of the data during the training process, enhance the recognition and restoration ability of the model in error correction tasks, and improve the text error correction processing efficiency.
[0036] In this embodiment, a method for determining a target text model is provided, which can be used on terminals such as computers and tablets. Figure 2 It is a flowchart of the method for determining a target text model according to an embodiment of the present disclosure, as Figure 2 shown. The process includes the following steps:
[0037] Step S201, obtain a first text error correction task, and obtain multiple second text error correction tasks associated with the first text error correction task.
[0038] Among them, the first text error correction task includes the first sample text and the first label. The first text error correction task is the main error correction task, which includes the first sample text and the corresponding first label. The first sample text contains incorrect text, and the first label is the correct text after correcting the first sample text. According to the first text error correction task, multiple second text error correction tasks related to the first text error correction task are generated. The error correction types of the second text error correction tasks may be different from those of the first text error correction task.
[0039] For example, if the first sample text error correction task is to process the first sample text into the correct text, the second text error correction task can be to determine whether there are errors in the first sample text, identify and extract the errors in the first sample text, etc.
[0040] Step S202: Generate the second sample text and the second label corresponding to the second text error correction task based on the first sample text and the first label.
[0041] According to the second text error correction task, the first sample text, and the first label, generate the second sample text and the second label of the second text error correction task. Use the information in the first text error correction task and the type of the second text error correction task to generate the corresponding second sample text and second label. The second sample text can be directly generated from the first sample text, or through operations such as deformation processing or splicing processing of the first sample text and the first label. The second label is the processing result corresponding to the second sample text determined based on the second text error correction task and the second sample text. For example, if the second text error correction task is to determine whether there are errors in the second sample text, correspondingly, the second label is "yes" or "no".
[0042] Step S203: Train the initial text model based on the first sample text and the first label, and the second sample text and the second label to obtain the target text model.
[0043] The initial text model is an untrained text model. Input the first sample text and the first label of the first text error correction task, and the second sample text and the second label of the second text error correction task into the initial text model for joint training. The model learns multiple tasks and optimizes its performance on all tasks through shared representations. After training, an optimized text model, that is, the target text model, is obtained.
[0044] The method for determining the target text model provided in this embodiment improves its performance on a single task by simultaneously learning multiple related tasks. Combining multiple text error correction tasks helps the model better learn the shared features between different tasks, thereby improving the efficiency of text processing.
[0045] In some alternative embodiments, step S202 in the above embodiments includes:
[0046] Step a1: Based on the first text error correction task and the first sample text, determine the first label.
[0047] Wherein, the first label represents the correct text corresponding to the first sample text. The first text error correction task means performing error correction processing on the first sample text to obtain the correct text, and the correct text is the complete text after correcting the errors in the first sample text. The first label is the correct text.
[0048] Step a2: Based on the type of the second text error correction task, the first sample text, and the correct text, determine the second sample text and the second label.
[0049] There are various types of the second text error correction task. The second sample text is generated according to the type of the error correction task. The content of the second sample text is related to the first sample text. It can be the first sample text itself or the text after processing the first sample text and the correct text. After obtaining the second sample text, determine the processing result of the second sample text according to the second text error correction task, and the processing result is the corresponding second label. For different types of the second text error correction tasks, their processing methods and the second sample texts can be different or the same. The same type of error correction task can include multiple forms of the second sample text. Exemplarily, if the first sample text is "a pair of oil paintings", the first label is "a picture of oil paintings".
[0050] In some alternative embodiments, step a2 includes: If the type of the second text error correction task is error judgment, determine the second label based on the correct text; obtain the second sample text based on the first sample text, or the concatenation of the first sample text and the correct text.
[0051] Error judgment means judging whether there are errors in the text. The form of the second sample text includes the first sample text, or the concatenation of the first sample text and the correct text. In practical applications, the first sample text and the correct text can be concatenated through a delimiter. The second label represents "yes" or "no" output after the error judgment.
[0052] Exemplarily, if the first sample text is "a pair of oil paintings" and the first label is "a picture of oil paintings".
[0053] The first type of the second sample text is "a pair of oil paintings", and the second label is "yes";
[0054] The second type of the second sample text is "a pair of oil paintings <s>An oil painting", and the second label is "Yes".
[0055] Among them, <s>is a separator, which means that the preceding and following texts are concatenated together. This is only an example, and other separators can also be used in actual applications.
[0056] In some optional implementations, step a2 includes: if the type of the second text error correction task is error location, then based on the first sample text, or the splicing of the first sample text and the correct text, a second sample text is obtained; based on the correct text, the error of the second sample text is located to determine a second label. The second label represents the erroneous text in the second sample text.
[0057] Error location means finding errors in the text and displaying the errors. The second sample text includes the first sample text, or the first sample text and the correct text. In practical applications, the first sample text and the correct text can be connected by a separator. The second label is the error content output after error location.
[0058] For example, if the first sample text is "an oil painting", the first label is "an oil painting".
[0059] The second sample text of the first type is "a painting", and the second label is "a painting";
[0060] The second sample text is "a painting <s>An oil painting", and the second label is "deputy".
[0061] Among them, <s>The delimiter is used to splice the text before and after. This is just an example, and other delimiters can also be used in actual applications.
[0062] In some alternative embodiments, step a2 includes: if the type of the second text error correction task is error correction, obtaining a second sample text based on the first sample text, or the splicing of the first sample text and the correct text; correcting the error text in the second sample text based on the correct text to determine a second label. The second label represents the correct result corresponding to the error text in the second sample text.
[0063] Error correction means finding and correcting the errors in the text and displaying the text after the error text is corrected. The form of the second sample text includes the first sample text, or the splicing of the first sample text and the correct text. In actual applications, the first sample text and the correct text can be spliced through a delimiter. The second label is the correct result corresponding to the error text.
[0064] Exemplarily, if the first sample text is "a painting", the first label is "a picture".
[0065] The first type of second sample text is "a painting", and the second label is "picture";
[0066] The second type of second sample text is "a painting <s>An oil painting", and the second label is "piece".
[0067] Among them, <s>A delimiter is used to splice the text before and after. This is just an example, and other delimiters can also be used in actual applications.
[0068] In some alternative embodiments, step a2 includes: if the type of the second text error correction task is text completion, randomly occlude at least one character in the first sample text to obtain a third sample text; based on the splicing of the first sample text and the third sample text, obtain a second sample text.
[0069] Among them, the second label represents the correct text. For text completion, the original text needs to be partially occluded first, and the proportion of the occluded characters can be set according to actual needs. Randomly occlude one or more characters in the first sample text, and the occluded text is the third sample text. The occluded characters in the third sample text are replaced with characters representing occlusion. The model combines the first sample text to complete the sentence and obtains the second label.
[0070] Exemplarily, if the first sample text is "a painting", the first label is "a painting".
[0071] The third sample text is "a <m>Oil <m>”, the second sample text is "a painting <s>One <m>Oil <m>”, and the second label is "a painting".
[0072] Among them, <s>As a separator, it means to splice the text before and after. This is just an example, and other separators can also be used in actual applications. <m>Indicates masking.
[0073] In some alternative embodiments, step a2 includes: if the type of the second text error correction task is error type judgment, randomly masking at least one character in the first sample text to obtain a third sample text; determining a second sample text based on the first sample text, or a concatenation of the first sample text and the correct text, or a concatenation of the first sample text and the third sample text; judging the error type of the first sample text based on the correct text to determine a second label. The second label represents the error type of the first sample text.
[0074] There are various error types of texts. Taking the existence of typos as an example, the reasons for typos may be similar pronunciations, similar glyphs, etc. Randomly masking one or more characters in the first sample text, the text after masking is the third sample text, and the masked characters in the third sample text are replaced with characters representing masking.
[0075] The second sample text can be the first sample text, or a concatenation of the first sample text and the correct text, or a concatenation of the first sample text and the third sample text. The error type is determined through error type judgment, and the second label is the error type.
[0076] Exemplarily, if the first sample text is "a painting", and the first label is "a picture". The third sample text is "a <m>Oil <m>”.
[0077] The second sample text is "a painting <s>One <m>Oil <m>”, the second label is "a painting".
[0078] The first type of second sample text is "a painting", and the second label is "phonetically similar";
[0079] The second type of second sample text is "a painting" <s>An oil painting", and the second label is "phonetically similar".
[0080] The third second sample text is "a painting <s>One <m>Oil <m>”, and the second label is "phonetically similar".
[0081] Among them, <s>As a delimiter, it means to splice the text before and after. This is just an example, and other delimiters can also be used in actual applications. <m>Indicates masking.
[0082] In some alternative embodiments, step a2 includes: if it is determined that there is an error cause in the second text error correction task, randomly mask at least one character in the first sample text to obtain a third sample text; determine the second sample text based on the concatenation of the first sample text and the third sample text; analyze the error cause of the first sample text to obtain a second label. The second label represents the error cause.
[0083] The error cause judgment is used to judge the reason for the error in the text. One or more characters in the first sample text are randomly masked, and the text after being masked is the third sample text. The masked characters in the third sample text are replaced with characters representing masking. The second sample text is the concatenation of the first sample text and the third sample text. Based on the model, the error in the first sample text is analyzed. After analysis, a second label is assigned to the second sample text, and the second label represents the reason for the error.
[0084] Exemplarily, if the first sample text is "a painting", the first label is "a painting". The third sample text is "a <m>Oil <m>”.
[0085] The second sample text is "a painting <s>One <m>Oil <m>", the error type can be determined as phonetic similarity. Further analysis is conducted to determine the reason for phonetic similarity. Therefore, the second tag is the reason for phonetic similarity.
[0086] Among them, <s>As a delimiter, it means to splice the text before and after. This is just an example, and other delimiters can also be used in actual applications. <m>Indicates masking.
[0087] In some alternative embodiments, there can be various associated tasks used when training the model based on multitasking. The second text error correction tasks in the above embodiments are all examples. Actually, other text error correction tasks can also be designed specifically.
[0088] For example, select the associated option. Specifically:
[0089] Compare the following two options and do a multiple-choice question: A: A painting B: A pair of paintings. The corresponding correct answer is: A, and the correct answer is the second label.
[0090] The text processing method provided by the embodiments of the present disclosure constructs similar training tasks to jointly train the model for a certain type of task, and then can perform text error correction processing based on the trained model, enabling the model to repeatedly learn relevant knowledge points and improving the generalization ability of the model in the error correction task.
[0091] In this embodiment, a text processing device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0092] As a specific application scenario of the embodiments of the present disclosure, the comment input by the user in the comment area of the video software is "AA%A%A", where A represents Chinese or English characters, that is, the first text is "AA%A%A". Before this comment is published, based on the target text model, the first text is corrected for errors, and it is detected that the % symbol is a meaningless symbol, so the % symbol is removed to obtain the corrected content "AAAA", that is, the second text.
[0093] This embodiment provides a text processing device, as Figure 3 shown, including:
[0094] A text acquisition module 301, configured to acquire the first text;
[0095] An error correction processing module 302, configured to perform text error correction processing on the first text based on the target text model to obtain the second text, and the target text model is trained based on multiple text error correction tasks with associated relationships.
[0096] In some alternative embodiments, the device further includes:
[0097] An error correction task acquisition module, configured to acquire a first text error correction task, and obtain a plurality of second text error correction tasks associated with the first text error correction task, where the first text error correction task includes a first sample text and a first label;
[0098] A second text generation module, configured to generate a second sample text and a second label corresponding to the second text error correction task based on the first sample text and the first label;
[0099] A model generation module, configured to train an initial text model based on the first sample text and the first label, and the second sample text and the second label, to obtain the target text model.
[0100] In some alternative embodiments, the second text generation module includes:
[0101] A first label determination unit, configured to determine the first label based on the first text error correction task and the first sample text, where the first label represents the correct text corresponding to the first sample text;
[0102] A second generation unit, configured to determine the second sample text and the second label based on the type of the second text error correction task, the first sample text, and the correct text.
[0103] In some alternative embodiments, the second generation unit includes:
[0104] A first error judgment sub-unit, configured to, if the type of the second text error correction task is error judgment, determine the second label based on the correct text;
[0105] A second error judgment sub-unit, configured to obtain the second sample text based on the first sample text, or the concatenation of the first sample text and the correct text.
[0106] In some alternative embodiments, the second generation unit includes:
[0107] A first error location sub-unit, configured to, if the type of the second text error correction task is error location, obtain the second sample text based on the first sample text, or the concatenation of the first sample text and the correct text;
[0108] A second error location sub-unit, configured to perform error location on the second sample text based on the correct text to determine the second label, where the second label represents the error text in the second sample text.
[0109] In some alternative embodiments, the second generation unit includes:
[0110] The first error correction subunit is configured to, if the type of the second text error correction task is error correction, obtain a second sample text based on the first sample text or the concatenation of the first sample text and the correct text;
[0111] The second error correction subunit is configured to correct the error text in the second sample text based on the correct text to determine the second label, where the second label represents the correct result corresponding to the error text in the second sample text.
[0112] In some alternative embodiments, the second generation unit includes:
[0113] The first text completion subunit is configured to, if the type of the second text error correction task is text completion, randomly obscure at least one character in the first sample text to obtain a third sample text;
[0114] The second text completion subunit is configured to obtain the second sample text based on the concatenation of the first sample text and the third sample text, where the second label represents the correct text.
[0115] In some alternative embodiments, the second generation unit includes:
[0116] The first error type judgment subunit is configured to, if the type of the second text error correction task is error type judgment, randomly obscure at least one character in the first sample text to obtain a third sample text;
[0117] The second error type judgment subunit is configured to determine the second sample text based on the first sample text, or the concatenation of the first sample text and the correct text, or the concatenation of the first sample text and the third sample text;
[0118] The third error type judgment subunit is configured to judge the error type of the first sample text based on the correct text to determine the second label, where the second label represents the error type of the first sample text.
[0119] In some alternative embodiments, the second generation unit includes:
[0120] The first cause judgment subunit is configured to, if the second text error correction task represents error cause judgment, randomly obscure at least one character in the first sample text to obtain a third sample text;
[0121] The second cause judgment subunit is configured to determine the second sample text based on the concatenation of the first sample text and the third sample text;
[0122] A third cause determination subunit, configured to analyze the error cause of the first sample text to obtain a second tag, where the second tag represents the error cause.
[0123] The further function descriptions of the foregoing respective modules and units are the same as those in the corresponding foregoing embodiments, and will not be elaborated herein.
[0124] The text processing device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the foregoing functions.
[0125] This disclosure embodiment further provides an electronic device having the foregoing Figure 3 shown text processing device.
[0126] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an optional embodiment of this disclosure. As Figure 4 shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common main board or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional implementation manners, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 4 In
[0127] Processor 10 may be a central processor, a network processor, or a combination thereof. Among them, processor 10 may further include a hardware chip. The foregoing hardware chip may be an application specific integrated circuit, a programmable logic device, or a combination thereof. The foregoing programmable logic device may be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.
[0128] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the foregoing embodiment.
[0129] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device and the like. In addition, the memory 20 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories can be connected to the electronic device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0130] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memory.
[0131] The electronic device further includes a communication interface 30 for communicating the electronic device with other devices or a communication network.
[0132] Embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and to be stored in a local storage medium, so that the methods described herein can be stored in such software processed on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0133] A part of the present disclosure can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can call or provide the methods and / or technical solutions according to the present disclosure through the operations of the computer. Those skilled in the art should understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways for computer program instructions to be executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0134] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the present disclosure.< / m> < / s> < / m> < / m> < / s> < / m> < / m> < / m> < / s> < / m> < / m> < / s> < / s> < / m> < / m> < / s> < / m> < / m> < / m> < / s> < / m> < / m> < / s> < / m> < / m> < / s> < / s> < / s> < / s> < / s> < / s>
Claims
1. A text processing method, characterized in that: The method comprises: Get the first text; The first text is subjected to text error correction processing based on a target text model to obtain a second text, wherein the target text model is obtained by training based on a plurality of text error correction tasks having a correlation relationship.
2. The method according to claim 1, characterized in that The method for determining the target text model includes: Acquire a first text error correction task, and obtain a plurality of second text error correction tasks associated with the first text error correction task, wherein the first text error correction task includes a first sample text and a first label; Based on the first sample text and the first label, generating a second sample text and a second label corresponding to the second text error correction task; Based on the first sample text and the first label, and the second sample text and the second label, the initial text model is trained to obtain the target text model.
3. The method according to claim 2, characterized in that The generating, based on the first sample text and the first label, a second sample text and a second label corresponding to the second text error correction task includes: Determine the first label based on the first text error correction task and the first sample text, where the first label represents the correct text corresponding to the first sample text; Based on the type of the second text error correction task, the first sample text and the correct text, a second sample text and a second label are determined.
4. The method according to claim 3, characterized in that The determining the second sample text and the second label based on the type of the second text error correction task, the first sample text, and the correct text includes: If the type of the second text error correction task is error judgment, determining the second label based on the correct text; The second sample text is obtained based on the first sample text, or based on the concatenation of the first sample text and the correct text.
5. The method according to claim 3, characterized in that: The determining the second sample text and the second label based on the type of the second text error correction task, the first sample text, and the correct text includes: If the type of the second text error correction task is error location, the second sample text is obtained based on the first sample text, or the concatenation of the first sample text and the correct text; The errors of the second sample text are located based on the correct text to determine the second label, where the second label represents the erroneous text in the second sample text.
6. The method according to claim 3, characterized in that The determining the second sample text and the second label based on the type of the second text error correction task, the first sample text, and the correct text includes: If the type of the second text error correction task is error correction, a second sample text is obtained based on the first sample text, or a combination of the first sample text and the correct text; The erroneous text in the second sample text is corrected based on the correct text to determine the second label, where the second label represents a correct result corresponding to the erroneous text in the second sample text.
7. The method according to claim 3, characterized in that The determining the second sample text and the second label based on the type of the second text error correction task, the first sample text, and the correct text includes: If the type of the second text error correction task is text completion, randomly mask at least one character in the first sample text to obtain a third sample text; Based on the concatenation of the first sample text and the third sample text, the second sample text is obtained, and the second label represents the correct text.
8. The method according to claim 3, characterized in that The determining the second sample text and the second label based on the type of the second text error correction task, the first sample text, and the correct text includes: If the type of the second text error correction task is error type judgment, randomly masking at least one character in the first sample text to obtain a third sample text; Determine the second sample text based on the first sample text, or the splicing of the first sample text and the correct text, or the splicing of the first sample text and the third sample text; The error type of the first sample text is judged based on the correct text to determine a second label, where the second label represents the error type of the first sample text.
9. The method according to claim 3, characterized in that: The determining the second sample text and the second label based on the type of the second text error correction task, the first sample text, and the correct text includes: If the second text error correction task represents error cause judgment, randomly cover at least one character in the first sample text to obtain a third sample text; Determine the second sample text based on the concatenation of the first sample text and the third sample text; The error cause of the first sample text is analyzed to obtain a second label, where the second label represents the error cause.
10. A text processing device, characterized in that: The device comprises: A text acquisition module, used for acquiring a first text; The error correction processing module is used to perform text error correction processing on the first text based on a target text model to obtain a second text. The target text model is obtained by training based on multiple text error correction tasks that have a correlation relationship.
11. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the text processing method according to any one of claims 1 to 9 by executing the computer instructions.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the text processing method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the text processing method according to any one of claims 1 to 9.