A text error correction method, device, equipment and storage medium

By obtaining the candidate character set and associated score, the problem of incoherence correction of other words or words in the text is solved, and the high-accurate text error correction effect is achieved.

CN114254623BActive Publication Date: 2025-05-13HEBEI XUNFEI ARTIFICIAL INTELLIGENCE RES INST +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111527097.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-05-13
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

The prior art is difficult to effectively correct other words or words in text, resulting in incoherence of the text after correction.

Method used

The corrected text is determined by obtaining the candidate character set and association scores for each position in the text to be corrected. This method considers the correlation between characters at each position of the candidate text, and improves the corrected text accuracy.

Benefits of technology

It realizes accurate correction of other words or words in the text, and improves the consistency and accuracy of the text after correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114254623B_ABST
    Figure CN114254623B_ABST
Patent Text Reader

Abstract

The present application provides a text correction method, apparatus, device and storage medium, the method comprising: obtaining a text to be corrected; obtaining a candidate character set corresponding to a plurality of positions, the candidate character set corresponding to the positions including candidate characters having an association relationship with the characters located at the positions in the text to be corrected; obtaining association scores corresponding to a plurality of candidate texts, the character at each position of the candidate text being a candidate character in the candidate character set corresponding to the position; and determining a corrected text corresponding to the text to be corrected from the plurality of candidate texts according to the association scores corresponding to the plurality of candidate texts. Since the present application takes into account the association relationship between the candidate characters at each position of the candidate text, the association score of the candidate text can reflect the accuracy of the candidate text as a whole, and according to the association scores corresponding to the candidate texts, the corrected text corresponding to the text to be corrected can be accurately determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a text error correction method, device, equipment and storage medium. Background Art

[0002] With the rapid development of the Internet, more and more texts are published on the Internet, such as news, books, advertisements and other texts.

[0003] The text can be obtained by: inputting the text through a keyboard, or using OCR (Optical Character Recognition) technology to recognize the text contained in the image, or using speech-to-text technology to convert speech into text. In the process of obtaining the text, problems such as wrong characters and wrong words often occur.

[0004] Therefore, how to correct the wrong characters or words in the text is a difficult problem that technicians in this field need to solve urgently. Summary of the invention

[0005] In view of this, the present application provides a text error correction method, device, equipment and storage medium for correcting wrong characters or wrong words and wrong lines in a text. The technical solution is as follows:

[0006] According to a first aspect of an embodiment of the present disclosure, a text error correction method is provided, comprising:

[0007] Get the text to be corrected;

[0008] Acquire candidate character sets corresponding to a plurality of positions respectively, wherein the candidate character sets corresponding to the positions include candidate characters associated with the characters located at the positions in the text to be corrected;

[0009] Obtaining association scores corresponding to a plurality of candidate texts respectively, wherein the character at each position of the candidate text is a candidate character in the candidate character set corresponding to the position, and the association score corresponding to the candidate text represents the probability that each candidate character contained in the candidate text constitutes the candidate text;

[0010] According to the relevance scores respectively corresponding to the multiple candidate texts, a corrected text corresponding to the text to be corrected is determined from the multiple candidate texts.

[0011] In combination with the first aspect, in a first possible implementation manner, the step of obtaining a candidate character set corresponding to the position includes:

[0012] Acquire character attributes of the first character located at the position in the text to be corrected, wherein the character attributes include a glyph feature and / or a pronunciation feature;

[0013] Acquire a character attribute of a second character in the text to be corrected that is located at a first set position corresponding to the position, wherein the first set position corresponding to the position at least includes a position before the position and / or a position after the position;

[0014] Calculating the associated feature corresponding to the position based on the character attribute of the first character and the character attribute of the second character;

[0015] Based on the associated features and the preset character vector of the pending character, a candidate character belonging to the candidate character set corresponding to the position is obtained from the pending character.

[0016] In combination with the first aspect, in a second possible implementation manner, the step of calculating the associated feature corresponding to the position based on the character attribute of the first character and the character attribute of the second character includes:

[0017] Calculating a first associated feature based on the glyph feature of the first character and the glyph feature of the second character;

[0018] Calculating a second associated feature based on the phonetic feature of the first character and the phonetic feature of the second character;

[0019] Obtain a word embedding vector corresponding to the first character;

[0020] The association feature is calculated based on the first association feature, the second association feature and the word embedding vector of the first character.

[0021] In combination with the first aspect, in a third possible implementation manner, the step of calculating the association feature based on the first association feature, the second association feature, and the word embedding vector of the first character includes:

[0022] Based on the context information of the position in the text to be corrected, obtaining text features corresponding to the position, the text features including context semantic features and / or context syntactic features of the position;

[0023] The association feature is calculated based on the text feature, the first association feature, the second association feature and the word embedding vector of the first character.

[0024] In combination with the first aspect, in a fourth possible implementation manner, the step of obtaining a candidate character belonging to the candidate character set corresponding to the position from the undetermined character based on the associated feature and the preset character vector of the undetermined character includes:

[0025] Determine the product of the association feature and the preset character vectors of the multiple characters to be determined, as the association degree of the multiple characters to be determined;

[0026] A preset number of pending characters with higher relevance among the plurality of pending characters are divided into a candidate character set corresponding to the position, or, pending characters with a relevance greater than or equal to a first threshold are divided into a candidate character set corresponding to the position.

[0027] In combination with the first aspect, in a fifth possible implementation manner, obtaining a relevance score corresponding to the candidate text includes:

[0028] Obtaining connection scores corresponding to candidate characters included in the candidate text, wherein the connection scores corresponding to the candidate characters at least represent adjacent co-occurrence probabilities of each character in a character set, wherein the character set at least includes the candidate character and the next candidate character of the candidate character in the candidate text;

[0029] Obtaining the prediction probability corresponding to the candidate characters contained in the candidate text, the prediction probability corresponding to the candidate characters being the probability that the target position is the candidate character when the character before the target position is the character before the target position in the text to be corrected, the target position being the position of the candidate character in the candidate text;

[0030] Based on the connection scores corresponding to the candidate characters in the candidate text and the predicted probabilities, the association score of the candidate text is calculated.

[0031] In combination with the first aspect, in a sixth possible implementation manner, the step of obtaining a connection score corresponding to the candidate character includes:

[0032] Acquire a text feature corresponding to the target position, wherein the text feature corresponding to the target position is obtained based on context information of the target position in the text to be corrected, and the text feature includes a context semantic feature and / or a context syntactic feature of the target position;

[0033] Acquire a text feature corresponding to a second set position, where the text feature corresponding to the second set position is obtained based on context information of a second set position corresponding to the target position in the error text to be corrected, where the second set position corresponding to the target position includes a next position of the target position and / or a previous position of the target position;

[0034] Obtaining a first word embedding vector of the candidate character;

[0035] Obtaining a second word embedding vector of the candidate character located at the second set position in the candidate text;

[0036] Based on the text feature corresponding to the target position, the text feature corresponding to the second set position, the first character embedding vector and the second character embedding vector, a connection score corresponding to the candidate character is calculated.

[0037] In combination with the first aspect, in a seventh possible implementation manner, the step of calculating the relevance score of the candidate text based on the connection scores corresponding to the candidate characters in the candidate text and the predicted probabilities includes:

[0038] Calculate the relevance score of the candidate text based on the connection score corresponding to the candidate characters in the candidate text, the prediction probability corresponding to the candidate characters in the candidate text, and the relevance degree corresponding to the candidate characters in the candidate text;

[0039] Among them, the association degree corresponding to the candidate character is the product of the association feature corresponding to the position of the candidate character and the preset character vector of the candidate character, and the association feature corresponding to the position is obtained based on the character attributes of the first character located at the position in the text to be corrected, and the character attributes include glyph features and / or phonetic features.

[0040] In combination with the first aspect, in an eighth possible implementation manner, a pre-built text error correction model is used to obtain candidate character sets corresponding to multiple positions respectively;

[0041] Using the text error correction model to obtain relevance scores corresponding to multiple candidate texts;

[0042] The text error correction model is used to determine the candidate text with the highest relevance score among the multiple candidate texts as the corrected text corresponding to the text.

[0043] The text error correction model is obtained by training a machine learning model using sample text as input and taking the maximum correlation score of the annotated corrected text corresponding to the sample text output by the machine learning model as the training target.

[0044] According to a second aspect of an embodiment of the present disclosure, a text error correction device is provided, comprising:

[0045] A first acquisition module is used to acquire the text to be corrected;

[0046] A second acquisition module is used to acquire candidate character sets corresponding to a plurality of positions respectively, wherein the candidate character sets corresponding to the positions include candidate characters associated with the characters located at the positions in the text to be corrected;

[0047] A third acquisition module is used to obtain association scores corresponding to a plurality of candidate texts, wherein the character at each position of the candidate text is a candidate character in the candidate character set corresponding to the position, and the association score corresponding to the candidate text represents the probability that each candidate character contained in the candidate text constitutes the candidate text;

[0048] The determination module is used to determine the corrected text corresponding to the text to be corrected from the multiple candidate texts according to the relevance scores respectively corresponding to the multiple candidate texts.

[0049] In conjunction with the second aspect, in a first possible implementation manner, for each position, the second acquisition module includes:

[0050] A character attribute extraction model, used for obtaining the character attribute of the first character located at the position in the text to be corrected, wherein the character attribute includes a glyph feature and / or a phonetic feature;

[0051] The character attribute extraction model is used to obtain the character attribute of the second character in the text to be corrected that is located at a first set position corresponding to the position, wherein the first set position corresponding to the position at least includes a position before the position and / or a position after the position;

[0052] The character attribute extraction model is used to calculate the associated feature corresponding to the position based on the character attribute of the first character and the character attribute of the second character;

[0053] The candidate character scoring model is used to obtain candidate characters belonging to the candidate character set corresponding to the position from the pending characters based on the associated features and the preset character vector of the pending characters.

[0054] In combination with the second aspect, in a second possible implementation manner, for each candidate text, the third acquisition module includes:

[0055] A self-attention mechanism model is used to obtain connection scores corresponding to candidate characters contained in the candidate text, wherein the connection scores corresponding to the candidate characters at least represent the adjacent co-occurrence probability of each character in a character set, and the character set at least includes the candidate character and the next candidate character of the candidate character in the candidate text;

[0056] A text prediction model, used to obtain prediction probabilities corresponding to candidate characters contained in the candidate text, wherein the prediction probabilities corresponding to the candidate characters are probabilities that the target position is the candidate character when the character before the target position is the character before the target position in the text to be corrected, and the target position is the position of the candidate character in the candidate text;

[0057] The corrected text determination model is used to calculate the relevance score of the candidate text based on the connection scores corresponding to the candidate characters in the candidate text and the prediction probability.

[0058] According to a third aspect of an embodiment of the present disclosure, there is provided a text error correction device, including a memory and a processor;

[0059] The memory is used to store programs;

[0060] The processor is used to execute the program to implement various steps of the text error correction method.

[0061] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the various steps of the text error correction method shown in the first aspect are implemented.

[0062] Through the above technical solution, it can be known that the text correction method provided by the present application, after obtaining the text to be corrected, first obtains a set of candidate characters corresponding to multiple positions respectively, the candidate character set corresponding to the position includes candidate characters with an association relationship with the character located at the position in the text to be corrected, and then obtains the association scores corresponding to the multiple candidate texts respectively, the character at each position of the candidate text is a candidate character in the candidate character set corresponding to the position, and finally, according to the association scores corresponding to the multiple candidate texts respectively, the corrected text corresponding to the text to be corrected is determined from the multiple candidate texts. Since the present application considers the association relationship between the candidate characters at each position of the candidate text in the process of obtaining the association score of the candidate text, instead of considering the candidate characters at each position as independent characters, the association score of the candidate text can reflect the accuracy of the candidate text as a whole, which makes it possible to accurately determine the corrected text corresponding to the text to be corrected from the multiple candidate texts according to the association scores corresponding to the multiple candidate texts. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0064] Figure 1 A schematic diagram of the hardware architecture involved in the embodiments of the present application;

[0065] Figure 2 A schematic diagram of an implementation of a text error correction method provided in an embodiment of the present application;

[0066] Figure 3 A schematic diagram of a process for obtaining a corrected text by a text error correction model provided in an embodiment of the present application;

[0067] Figure 4 A schematic diagram of an implementation of a network structure of a character attribute feature extraction model provided in an embodiment of the present application;

[0068] Figure 5 A schematic diagram of an implementation method of the network structure of the self-attention mechanism model provided in an embodiment of the present application;

[0069] Figure 6 A network structure diagram of an implementation method of a text error correction model provided in an embodiment of the present application;

[0070] Figure 7 A schematic diagram of the structure of a text error correction device provided in an embodiment of the present application;

[0071] Figure 8 A schematic diagram of the structure of a text correction device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0072] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0073] The embodiments of the present application provide a text error correction method, apparatus, device and storage medium. Before introducing the technical solution provided by the embodiments of the present application, the relevant technologies and hardware architecture involved in the present application are first described.

[0074] First, the related art will be described.

[0075] In the related art, the characters to be corrected at various positions of the text to be corrected are input into a pre-constructed language model, and the language model outputs the corrected characters corresponding to each position. If the corrected character at position A output by the text correction model is different from the character to be corrected at position A, the character at position A in the text to be corrected is replaced with the corrected character to obtain the corrected text.

[0076] During the research process, the applicant found that the above idea has defects: Since the language model in the related technology belongs to a non-autoregressive model, which has an output independence assumption, that is, the corrected characters at each position of the output are independent and irrelevant to each other, this will result in the corrected text composed of the corrected characters corresponding to each position of the output being incoherent. That is, the language model in the related technology does not consider the relevance of the characters between positions in the corrected text, resulting in the corrected text composed of the corrected characters corresponding to each position of the output being incoherent. That is, the corrected text obtained through the language model in the related technology may still contain misspelled characters and words, that is, the language model in the related technology is inaccurate.

[0077] For example, if the text to be corrected is "我真户秃", the corrected character at position 1 output by the language model in the related technology is "我", the corrected character at position 2 is "真", the corrected character at position 3 is "尴", and the corrected character at position 4 is "涂", then the corrected text is "我真尴涂", but "我真尴涂" is still incoherent.

[0078] In view of the defects of the above idea, the inventor of this case continued to research. Through continuous research, a text error correction method with better effects was finally proposed. This text error correction method can be applied to any application scenario that requires text error correction. This text error correction method overcomes the defects of the above idea, that is, this text error correction method considers the relevance of the characters between positions in the corrected text and improves the accuracy of the obtained corrected text.

[0079] Secondly, the hardware architecture involved in the embodiments of the present application will be described.

[0080] As Figure 1 shown, it is a schematic diagram of the hardware architecture involved in the embodiments of the present application. This hardware architecture includes: an electronic device 11 and a server 12.

[0081] Exemplarily, the electronic device 11 can be any electronic product that can perform human-computer interaction with the user through one or more ways such as a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device. For example, a mobile phone, a laptop computer, a tablet computer, a handheld computer, a personal computer, a wearable device, a smart TV, a PAD, etc.

[0082] It should be noted that Figure 1 is just an example, and there can be various types of electronic devices, not limited to Figure 1 the laptop computer in.

[0083] Exemplarily, the server 12 may be a single server, or a server cluster consisting of multiple servers, or a cloud computing server center. The server 12 may include a processor, a memory, a network interface, and the like.

[0084] Exemplarily, the electronic device 11 may establish a connection and communicate with the server 12 via a wireless communication network; exemplary, the electronic device 11 may establish a connection and communicate with the server 12 via a wired network.

[0085] In an optional implementation, the user can upload the text to be corrected to the electronic device 11, and the electronic device 11 sends the text to be corrected to the server 12, so that the server 12 can correct the text to be corrected to obtain a corrected text.

[0086] In an optional implementation, the user can upload an image containing text to the electronic device 11, and the electronic device 11 can recognize the text from the image and send the text as text to be corrected to the server 12, so that the server 12 can correct the text to be corrected to obtain the corrected text.

[0087] In an optional implementation, the user can upload the voice to the electronic device 11, the electronic device 11 converts the voice into text, and sends the text as text to be corrected to the server 12, so that the server 12 can correct the text to be corrected to obtain the corrected text.

[0088] In an optional implementation, the electronic device 11 may send an image or voice to the server 12. The server 12 may recognize text from the image, or convert the voice into text, and then correct the text as the text to be corrected to obtain the corrected text.

[0089] Those skilled in the art should understand that the above-mentioned electronic devices and servers are only examples, and other existing or future electronic devices or servers that are applicable to the present disclosure should also be included in the protection scope of the present disclosure and are included here by reference.

[0090] First embodiment

[0091] The text error correction method provided in the embodiment of the present application is described below in combination with the above-mentioned related technologies and hardware architecture.

[0092] like Figure 2 FIG. 1 is a schematic diagram of an implementation of a text error correction method provided in an embodiment of the present application. The method can be applied to the server 12 , and the method involves the following steps S21 to S24 during implementation.

[0093] Step S21: Obtain the text to be corrected.

[0094] Exemplarily, the text to be corrected includes one or more of language characters from various countries such as Chinese characters, English characters, Korean characters, Japanese characters, etc.

[0095] The embodiments of the present application do not limit the number and type of characters included in the text to be corrected.

[0096] Exemplarily, in the embodiments of the present application, Chinese characters included in the text to be corrected are corrected.

[0097] Step S22: Obtain candidate character sets corresponding to multiple positions respectively, where the candidate character set corresponding to a position includes candidate characters having an association relationship with the character at the position in the text to be corrected.

[0098] Still taking the text to be corrected "我真户秃" as an example for illustration. Then in step S22, candidate character set 1 corresponding to position 1, candidate character set 2 corresponding to position 2, candidate character set 3 corresponding to position 3, and candidate character set 4 corresponding to position 4 can be obtained.

[0099] Among them, the candidate character set 1 corresponding to position 1 includes candidate characters having an association relationship with "我", for example, 我, 窝, 卧; the candidate character set 2 corresponding to position 2 includes candidate characters having an association relationship with "真", for example, 真, 针, 振; the candidate character set 3 corresponding to position 3 includes candidate characters having an association relationship with "户", for example, 户, 糊, 尴; the candidate character set 4 corresponding to position 4 includes candidate characters having an association relationship with "秃", for example, 秃, 涂, 图.

[0100] In summary, assume that the candidate character set 1 corresponding to position 1 is {我, 窝, 卧}; the candidate character set 2 corresponding to position 2 is {真, 针, 振}; the candidate character set 3 corresponding to position 3 is {户, 糊, 尴}; the candidate character set 4 corresponding to position 4 is {秃, 涂, 图}.

[0101] It can be understood that the candidate character set corresponding to a position includes the character at that position in the text to be corrected.

[0102] The above is only an example, and the number of candidate characters included in the candidate character sets corresponding to different positions may be the same or different.

[0103] It can be understood that the main reason for the appearance of miswritten characters or misused words is that miswritten characters are homophonic or similar in form to the correct characters. Exemplarily, candidate characters having an association relationship with a character refer to candidate characters having a relatively high degree of relevance to the glyph feature and / or phonetic feature of the character.

[0104] Next, the glyph feature and / or phonetic feature of a character will be described in combination with the type of the character.

[0105] If the character is a Chinese character, exemplarily, the glyph feature of the character is obtained based on the strokes of the character.

[0106] For example, if the character is "需", the strokes of 需 are:

[0107] If the character is a Chinese character, exemplarily, the phonetic feature of the character is obtained based on the pinyin of the character.

[0108] For example, if the character is "需", the pinyin of 需 is: xu.

[0109] If the character is a non-Chinese character, exemplarily, the character embedding vector of the character can be used as the glyph feature.

[0110] If the character is a non-Chinese character, exemplarily, the character embedding vector of the character can be used as the phonetic feature.

[0111] Step S23: Obtain the association scores corresponding to multiple candidate texts.

[0112] The character at each position of the candidate text is a candidate character in the candidate character set corresponding to the position. The association score corresponding to the candidate text represents the probability that the candidate characters included in the candidate text form the candidate text.

[0113] Among different candidate texts, there is at least one position where the candidate characters are different.

[0114] Still taking the text to be corrected as "我真户秃" as an example for illustration. If the candidate character set 1 corresponding to position 1 is {我, 窝, 卧}; the candidate character set 2 corresponding to position 2 is {真, 针, 振}; the candidate character set 3 corresponding to position 3 is {户, 糊, 尴}; the candidate character set 4 corresponding to position 4 is {秃, 涂, 图}. Then 81 candidate texts can be obtained.

[0115] The calculation formula for the number of candidate texts is as follows: the number of candidate characters included in candidate character set 1 * the number of candidate characters included in candidate character set 2 * the number of candidate characters included in candidate character set 3 * the number of candidate characters included in candidate character set 4 = 3 * 3 * 3 * 3 = 81.

[0116] The following lists 4 candidate texts: 我真糊涂, 我真户秃, 窝真户涂, 我真尴涂.

[0117] The association score of each candidate text can be obtained. For example, the association score of the candidate text "我真糊涂" refers to the probability that the candidate text is composed of "我" at position 1, "真" at position 2, "糊" at position 3, and "涂" at position 4.

[0118] In summary, in the process of obtaining the association score of the candidate text, the association relationship between the candidate characters at each position of the candidate text is considered, and the candidate characters at each position are no longer considered as independent characters. Therefore, the association score of the obtained candidate text can reflect the accuracy of the candidate text as a whole.

[0119] Step S24: Determine the corrected text corresponding to the text to be corrected from multiple candidate texts according to the association scores respectively corresponding to the multiple candidate texts.

[0120] Optionally, the candidate text with the highest association score among the multiple candidate texts can be determined as the corrected text corresponding to the text to be corrected. Still taking the text to be corrected as "我真户秃" as an example. If among the 81 candidate texts, the association score of the candidate text "我真糊涂" is the highest, it means that the probability of the candidate text "我真糊涂" as a whole is the highest, that is, it is more coherent. Then the candidate text "我真糊涂" is the corrected text.

[0121] The embodiment of the present application provides a text error correction method. After obtaining the text to be corrected, first obtain the candidate character sets respectively corresponding to multiple positions, then obtain the association scores respectively corresponding to multiple candidate texts, and finally determine the corrected text corresponding to the text to be corrected from multiple candidate texts according to the association scores respectively corresponding to the multiple candidate texts. Since in the process of obtaining the association score of the candidate text in the embodiment of the present application, the association relationship between the candidate characters at each position of the candidate text is considered, rather than considering the candidate characters at each position as independent characters, therefore, the association score of the candidate text can reflect the accuracy of the candidate text as a whole, which enables the corrected text corresponding to the text to be corrected to be accurately determined from multiple candidate texts according to the association scores respectively corresponding to the multiple candidate texts.

[0122] Second Embodiment

[0123] In an alternative implementation, the above steps S22 to S24 may be executed by a pre-constructed text error correction model. Then, as Figure 3 shown, it is a schematic diagram of the process of the text error correction model provided by the embodiment of the present application to obtain the corrected text.

[0124] The method for obtaining the corrected text by using the text error correction model includes a first step and a second step.

[0125] First step: Obtain the text to be corrected.

[0126] Second step: Input the text to be corrected into the pre-constructed text error correction model 31 to obtain the corrected text corresponding to the text to be corrected output by the text error correction model 31.

[0127] Among them, the text error correction model is obtained by training a machine learning model with the sample text as the input and the maximum correlation score between the sample text output by the machine learning model and the corresponding labeled corrected text as the training objective.

[0128] During the process of training the text error correction model, the text error correction model can output each candidate text and the corresponding correlation score for each candidate text. Each candidate text includes the labeled corrected text.

[0129] If the sample text is "我真户秃" (I am really bald), then the labeled corrected text is "我真糊涂" (I am really confused). That is, the labeled corrected text is the accurate corrected text corresponding to the sample text.

[0130] It can be understood that the specific implementation process of the second step includes the following steps A11 to A13.

[0131] Step A11: Use the pre-constructed text error correction model to obtain candidate character sets corresponding to multiple positions respectively.

[0132] In an optional implementation manner, the text error correction model includes: a candidate character set acquisition model 32 and a correlation relationship extraction model 33.

[0133] Exemplarily, the candidate character set acquisition model 32 can be used to obtain candidate character sets corresponding to multiple positions respectively.

[0134] Step A12: Use the text error correction model to obtain the correlation scores corresponding to multiple candidate texts respectively.

[0135] Exemplarily, the correlation relationship extraction model 33 can be used to obtain the correlation scores corresponding to multiple candidate texts respectively.

[0136] Step A13: Use the text error correction model to determine the corrected text corresponding to the text to be error-corrected from multiple candidate texts based on the correlation scores corresponding to multiple candidate texts respectively.

[0137] Exemplarily, the correlation relationship extraction model 33 can be used to determine that the candidate text with the highest correlation score among multiple candidate texts is the corrected text corresponding to the text to be error-corrected.

[0138] In an optional implementation manner, the above steps S22 to S24 are not all executed by the pre-constructed model.

[0139] Third Embodiment

[0140] Next, in combination with the above second embodiment, the network structure of the text error correction model will be described.

[0141] Exemplarily, the candidate character set acquisition model 32 includes: a character attribute extraction model 41 and a candidate character scoring model 42. The character attribute extraction model 41 is used to obtain the associated features corresponding to multiple positions respectively; the candidate character scoring model 42 is used to calculate the association degree of each pending character for each position based on the preset character vector of each pending character and the associated features corresponding to the position; based on the association degree of each pending character, determine the candidate character belonging to the candidate character set corresponding to the position from each pending character.

[0142] It is understandable that there are many ways to implement the network structure of the character attribute extraction model 41. The embodiments of the present application provide but are not limited to the following three.

[0143] The network structure of the first character attribute extraction model.

[0144] The character attribute extraction model 41 includes: a text feature extraction model 411 .

[0145] Exemplarily, the text feature extraction model 411 is a RoBerta model or a BERT model.

[0146] Among them, the RoBerta model is an improved model based on the BERT model. The RoBerta model uses a bidirectional transformer based on the self-attention mechanism as a feature extraction method, effectively utilizing the text features of the context.

[0147] Exemplarily, the text feature extraction model 411 can obtain text features corresponding to each position in the text to be corrected. For example, for position i, the text feature h i .

[0148] Exemplarily, the text features include contextual semantic features and / or contextual syntactic features. Exemplarily, the contextual semantic features at position i represent the meaning of the character at position i in the entire text to be corrected, and the contextual syntactic features at position i represent the part of speech of the character at position i in the entire text to be corrected.

[0149] In combination with the network structure of the first character attribute extraction model, the specific process of obtaining the associated features corresponding to multiple positions by the character attribute extraction model 41 is described. The process of obtaining the associated features corresponding to multiple positions includes the following steps B11 to B12.

[0150] Step B11: input the text to be corrected into the text feature extraction model 411.

[0151] Step B12: Acquire text features corresponding to multiple positions respectively through the text feature extraction model 411.

[0152] Exemplarily, for the text feature h at position i i , the text feature extraction model 411 can obtain the text feature h based on the characters before and after the character at position i in the text to be corrected i .

[0153] Step B13: Determine that the text features corresponding to multiple positions are the associated features corresponding to the multiple positions respectively

[0154] It can be understood that the reason for the existence of misspelled characters or words in the text to be corrected is that the misspelled characters or words are homophonic or similar in shape to the correct characters. The BERT model or the RoBerta model can effectively extract the text features corresponding to each position, but the BERT model or the RoBerta model does not consider the phonetic features and the shape features of the characters. The associated feature corresponding to each position output by the text feature extraction model 411 does not include the character attributes (character attributes include shape features and / or phonetic features) of the character at that position in the text to be corrected. Therefore, the candidate characters in the candidate character set corresponding to that position may not be related to the character attributes of the character at that position in the text to be corrected, that is, the obtained candidate character set corresponding to the position is inaccurate

[0155] In view of this, the embodiment of the present application provides a second character attribute extraction model 41

[0156] The network structure of the second character attribute extraction model 41 includes: a phonetic feature extraction model 412 and / or a shape feature extraction model 413

[0157] Next, in combination with the type of characters, the process of obtaining the shape features and / or phonetic features of characters will be described

[0158] If the character is a Chinese character, exemplarily, the phonetic feature extraction model 412 can obtain the pinyin of each character in the text to be corrected. Since Chinese characters may have multiple pronunciations, the phonetic feature extraction model 412 can obtain multiple pronunciations of Chinese characters

[0159] Exemplarily, the pinyin of a character is the initial and final sounds that make up the pinyin of the Chinese character. For example, the initial sound of the character "si" is "s", the final sound is "i", the tone is "4", and the pinyin of "si" is "si"

[0160] Exemplarily, the phonetic feature extraction model 412 can obtain similar pinyins based on the pinyin of each character. For example, the initial sound "s" is similar to the initial sound "sh"

[0161] Exemplarily, if the character is a Chinese character, the phonetic features of the character acquired by the phonetic feature extraction model 412 include: the pinyin of the character, or the tone and pinyin of the character, or the pinyin of the character and a pinyin with a high degree of similarity to the pinyin of the character, or the tone of the character, the pinyin of the character, and a pinyin with a high degree of similarity to the pinyin of the character.

[0162] If the character is a non-Chinese character, for example, English, the pronunciation feature extraction model 412 can obtain the pronunciation of the character, for example, the pronunciation of the character thank is

[0163] Exemplarily, the glyph feature extraction model may be a LSTM (Long Short-Term Memory) model.

[0164] If the character is a Chinese character, illustratively, the strokes of the character may be input into a glyph feature extraction model, thereby obtaining glyph features output by the glyph feature extraction model.

[0165] It is understandable that different Chinese characters have different numbers of strokes and specific forms of strokes. The glyph feature extraction model can be used to explore the internal relationship of the stroke combinations that make up the Chinese characters.

[0166] If the character is a non-Chinese character, illustratively, the word embedding vector of the character can be used as the glyph feature of the character.

[0167] Exemplarily, the text to be corrected is input into an embedding layer, and the embedding layer can map the characters in the input text to be corrected into corresponding word embedding vectors. The parameters of the embedding layer are updated during the model training process.

[0168] Assume that the phonetic feature in the character attribute corresponding to the character at position i in the text to be corrected is represented by p i Representation; for the glyph feature in the character attribute corresponding to the character at position i of the text to be corrected, q i Representation.

[0169] Combined with the network structure of the second character attribute extraction model, the specific process of obtaining the associated features corresponding to multiple positions by the character attribute extraction model 41 is described. The method for obtaining the associated features corresponding to each position provided in the embodiment of the present application includes the following steps B21 to B23.

[0170] Step B21: Obtain the phonetic feature of the first character at the position in the text through the phonetic feature extraction model 412.

[0171] Step B22: Obtaining the glyph feature of the first character located at the position in the text through the glyph feature extraction model 412.

[0172] Step B23: Determine that the phonetic feature and / or the glyph feature of the first character is the associated feature corresponding to the position.

[0173] Exemplarily, if the associated feature corresponding to the position includes the phonetic feature and the glyph feature of the first character, the phonetic feature and the glyph feature can be combined to obtain the associated feature. For example, the phonetic feature and the glyph feature can be input into a fully connected layer, and the phonetic feature and the glyph feature are simply fused through the fully connected layer to obtain the associated feature.

[0174] The associated features corresponding to each position obtained through the above Steps B21 to B23 take into account the glyph feature and / or the phonetic feature, so that the similarity between the candidate characters in the candidate character set corresponding to each position obtained and the phonetic feature and / or the glyph feature of the first character at that position in the text to be corrected is relatively high. The accuracy of obtaining the candidate character sets corresponding to each position is improved.

[0175] It can be understood that there is an association relationship between two or more adjacent characters in the text to be corrected. For example, if the text to be corrected is "The subject I like: Mathematics", and the corrected text is "The subject I like: Mathematics", the character before the character "wood" is "subject". It can be understood that "subject" can form a word, while "subwood" cannot form a word. Therefore, if the association relationship between two or more adjacent characters can be learned, the accuracy of the candidate character sets corresponding to each position can be further improved. And the number of candidate characters included in the candidate character set corresponding to each position is reduced.

[0176] For example, assume that the candidate character sets corresponding to two adjacent positions are candidate character set 1 and candidate character set 2 respectively. Then, if the association relationship between two or more adjacent characters is considered, the probability that any candidate character included in the obtained candidate character set 1 can form a word with a certain candidate character included in the candidate character set 2 is relatively high.

[0177] Based on this, under the network structure of the second character attribute extraction model, the specific process of the second character attribute extraction model 41 for obtaining the associated feature corresponding to any position is provided. This method includes the following Steps B31 to B33.

[0178] Step B31: Obtain the character attribute of the first character at the position in the text to be corrected, where the character attribute includes the glyph feature and / or the phonetic feature.

[0179] Exemplarily, the phonetic feature of the first character at the position in the text is obtained through the phonetic feature extraction model 412.

[0180] Exemplarily, the glyph feature of the first character at the position in the text is obtained by the glyph feature extraction model 412 .

[0181] Step B32: Obtain the character attribute of the second character in the text to be corrected that is located at a first set position corresponding to the position, wherein the first set position corresponding to the position at least includes a position before the position and / or a position after the position.

[0182] Exemplarily, the phonetic feature of the second character located at the first set position in the text is obtained by the phonetic feature extraction model 412.

[0183] Exemplarily, the glyph feature of the second character located at the first set position in the text is obtained by the glyph feature extraction model 412 .

[0184] Step B33: Based on the character attributes of the first character and the character attributes of the second character, calculate and obtain the associated features corresponding to the position.

[0185] Exemplarily, the number of the second characters may be one or more.

[0186] Assume that the position mentioned in step B31 to step B33 is position i. Exemplarily, if the first set position corresponding to position i includes position i-1 or position i+1, the second character includes: the character at position i-1 or the character at position i+1. Exemplarily, if the first set position corresponding to position i includes position i-1 and position i+1, the second character includes: the character at position i-1 and the character at position i+1. Exemplarily, if the first set position corresponding to position i includes: position i-2, position i-1, position i+1, position i+2. Then the second character includes: the character at position i-2, the character at position i-1, the character at position i+1, and the character at position i+2.

[0187] Assume that the character attributes of the first character (located at position i of the text to be corrected) include: the phonetic feature p i and / or character features q i The second character includes the character at position i-1 and the character at position i+1, then the character attributes of the second character include: the phonetic feature p i-1 and / or character features q i-1 , phonetic feature p i+1 and / or character features q i+1 .

[0188] In an optional implementation, there are multiple implementations of step B33, and the embodiments of the present application provide but are not limited to the following two.

[0189] The first implementation method of step B33 includes: inputting the attribute features of the first character and the character attributes of the second character into a fully connected layer, a normalized layer, or a convolutional layer; and calculating the associated features through the fully connected layer, the normalized layer, or the convolutional layer.

[0190] For example, taking the convolutional layer as an example, the associated feature = Conv(p i-1 ,p i ,p i+1 ,q i-1 ,q i ,q i+1 ).

[0191] The second implementation method of step B33 includes the following steps B331 to B333.

[0192] Step B331: Obtain a first associated feature based on the glyph feature of the first character and the glyph feature of the second character.

[0193] Exemplarily, the glyph features of the first character and the glyph features of the second character may be input into a fully connected layer, a normalized layer, or a convolutional layer to obtain a first associated feature. Exemplarily, taking the convolutional layer as an example, the first associated feature d i =Conv(q i-1 ,q i ,q i+1 ).

[0194] Step B332: Obtain a second associated feature based on the phonetic feature of the first character and the phonetic feature of the second character.

[0195] Exemplarily, the phonetic and shape features of the first character and the phonetic features of the second character can be input into a fully connected layer, a normalized layer, or a convolutional layer to obtain a second associated feature. Exemplarily, taking the convolutional layer as an example, the second associated feature.

[0196] Step B333: Based on the first association feature and the second association feature, calculate the association feature corresponding to the position.

[0197] Exemplarily, the first correlation feature and the second correlation feature may be input into a fully connected layer, a normalization layer, or a convolutional layer to obtain the correlation feature.

[0198] It can be understood that after the above-mentioned character attribute feature extraction model has been processed through multiple layers, the original features of the character located at the position in the text to be corrected are lost. In order to highlight the original features of the first character located at the position in the text to be corrected, the above-mentioned step S333 may include: obtaining the word embedding vector corresponding to the first character; based on the word embedding vector of the first character, the first associated feature and the second associated feature, calculating the associated feature corresponding to the position.

[0199] Exemplarily, the word embedding vector of the first character can be obtained through an embedding layer. The embedding layer can map the characters in the input text to be corrected into corresponding word embedding vectors, and the parameters of the embedding layer are updated during the model training process.

[0200] Exemplarily, the associated features corresponding to the positions can be calculated through a fully connected layer or a normalization layer or a convolutional layer.

[0201] Exemplarily, the word embedding vector of the first character can be obtained through one-hot encoding.

[0202] In the above steps B31 to B33, during the process of obtaining the associated features corresponding to each position, the association relationship between adjacent characters is considered, so the candidate character set obtained based on the associated features corresponding to the position is more accurate. However, the text features of the context in the entire text to be corrected are not considered, that is, the candidate characters in the candidate character set corresponding to the obtained position may not conform to the context. Based on this, the third character attribute feature extraction model is provided in the embodiments of the present application.

[0203] The network structure of the third character attribute feature extraction model includes: a text feature extraction model 411, a pronunciation feature extraction model 412, and a glyph feature extraction model 413.

[0204] As Figure 4 shown, it is a schematic diagram of an implementation manner of the network structure of the character attribute feature extraction model provided in the embodiments of the present application.

[0205] As Figure 4 shown, the candidate character set acquisition module further includes an input layer 40.

[0206] Exemplarily, the character attribute feature extraction model further includes: a convolutional layer 414, a convolutional layer 415, and a normalization layer 416.

[0207] Exemplarily, the convolutional layer 414 can be replaced by a fully connected layer or a normalization layer. Exemplarily, the convolutional layer 415 can be replaced by a fully connected layer or a normalization layer. Exemplarily, the normalization layer 416 can be replaced by a fully connected layer or a convolutional layer.

[0208] The text to be corrected (assumed to be "我真户秃") can be input into the input layer 40.

[0209] Exemplarily, the input layer includes an embedding layer, and the embedding layer can map the characters in the input text to be corrected into corresponding word embedding vectors, and the parameters of the embedding layer are updated during the model training process.

[0210] Exemplarily, before the text to be corrected is input into the input layer, the characters in the text to be corrected are first one-hot encoded to obtain a word embedding vector representation of each character in the text to be corrected.

[0211] Thus, the word embedding vector of each character in the text to be corrected is input into the normalization layer 416.

[0212] Exemplarily, the normalization layer 416 may be a BatchNorm layer, a LayerNorm layer, an InstanceNorm layer, a GroupNorm layer, or a SwitchableNorm layer. Figure 4 The description is made by taking the normalization layer 416 as the LayerNorm layer as an example.

[0213] For example, the normalization layer 46 may be replaced by a fully connected layer or a convolutional layer.

[0214] The following describes the specific process of the character attribute extraction model 41 executing the acquisition of associated features corresponding to a plurality of positions in combination with the third character attribute feature extraction model.

[0215] Under the third character attribute feature extraction model, the first process of obtaining the associated features corresponding to the position includes the following steps B41 and B43.

[0216] Step B41: Obtaining character attributes of the first character located at the position in the text to be corrected, wherein the character attributes include glyph features and / or pronunciation features.

[0217] Exemplarily, the phonetic feature of the first character at the position in the text is obtained by the phonetic feature extraction model 412 .

[0218] Exemplarily, the glyph feature of the first character at the position in the text is obtained by the glyph feature extraction model 412 .

[0219] Step B42: Based on the context information of the position in the text to be corrected, the text feature corresponding to the position is obtained through the text feature extraction model 411, and the text feature includes the semantic feature and / or syntactic feature of the character located at the position.

[0220] Step S43: Calculate the associated feature based on the text feature and the character attribute of the first character.

[0221] Exemplarily, the associated feature may be calculated by the normalization layer 416 based on the text feature and the character attribute of the first character.

[0222] It can be understood that after the above-mentioned character attribute feature extraction model has been processed through multiple layers, the original features of the character located at the position in the text to be corrected are lost. In order to highlight the original features of the first character at this position in the text to be corrected, the above-mentioned step S43 includes: obtaining the word embedding vector corresponding to the first character; based on the word embedding vector corresponding to the first character, the text features and the character attributes of the first character, calculating the associated features.

[0223] Under the third character attribute feature extraction model, the second process of obtaining the associated features corresponding to the position includes the following steps B51 and B56.

[0224] Step B51: Obtaining character attributes of the first character located at the position in the text to be corrected, wherein the character attributes include glyph features and / or pronunciation features.

[0225] Exemplarily, the phonetic feature of the first character at the position in the text is obtained by the phonetic feature extraction model 412 .

[0226] Exemplarily, the glyph feature of the first character at the position in the text is obtained by the glyph feature extraction model 412 .

[0227] Step B52: Obtain the character attribute of the second character in the text to be corrected that is located at a first set position corresponding to the position, wherein the first set position corresponding to the position at least includes a position before the position and / or a position after the position.

[0228] Exemplarily, the phonetic feature of the second character located at the first set position in the text is obtained by the phonetic feature extraction model 412.

[0229] Exemplarily, the glyph feature of the second character located at the first set position in the text is obtained by the glyph feature extraction model 412 .

[0230] Step B53: Obtain a first associated feature based on the glyph feature of the first character and the glyph feature of the second character.

[0231] Exemplarily, the first associated feature is obtained through the convolution layer 415 based on the glyph feature of the first character and the glyph feature of the second character.

[0232] Step B54: obtaining a second associated feature based on the phonetic feature of the first character and the phonetic feature of the second character.

[0233] Exemplarily, the second associated feature is obtained through the convolution layer 414 based on the phonetic feature of the first character and the phonetic feature of the second character.

[0234] Step B55: Based on the context information of the position in the text to be corrected, obtain text features corresponding to the position, wherein the text features include context semantic features and / or context syntactic features of the position.

[0235] Exemplarily, the text feature corresponding to the position is obtained based on the context information of the position in the text to be corrected through the text feature extraction model 411.

[0236] Step B56: Calculate the association feature based on the text feature, the first association feature, and the second association feature.

[0237] It can be understood that after the above-mentioned character attribute feature extraction model has been processed through multiple layers, the original features of the character located at the position in the text to be corrected are lost. In order to highlight the original features of the first character located at this position in the text to be corrected, the above-mentioned step B56 includes: obtaining the word embedding vector corresponding to the first character; based on the word embedding vector corresponding to the first character, the text feature, the first associated feature and the second associated feature, calculating the associated feature.

[0238] Exemplarily, the association feature may be calculated by the normalization layer 416 based on the text feature, the first association feature, and the second association feature.

[0239] Assuming that the first character is at position i, the text feature corresponding to position i is represented by h i Representation.

[0240] Figure 4 In the figure, the first set position corresponding to position i includes the previous position i-1 and the next position i+1 of the position i.

[0241] For position i, the word embedding vector w of the character at position i in the text to be corrected can be obtained through the input layer 40 i ; Obtain the phonetic feature p of the first character at position i in the text to be corrected through the phonetic feature extraction model 412 i , the phonetic feature p of the second character at position i-1 in the text to be corrected i-1 , the phonetic feature p of the second character at position i+1 in the text to be corrected i+1 ; Obtain the glyph feature q of the first character at position i in the text to be corrected through the glyph feature extraction model 413 i , the glyph feature q of the second character at position i-1 in the text to be corrected i-1 , the glyph feature q of the second character at position i+1 in the text to be corrected i+1 ; The phonetic feature p i-1、Phonetic features i 、Phonetic features i+1 Input to the convolution layer 414, the second correlation feature c can be obtained i =Conv(p i-1 ,p i ,p i+1 );The glyph feature q i-1 , font features q i , font features q i+1 Input to convolution 415 to obtain the first correlation feature d i =Conv(q i-1 ,q i ,q i+1 ). The text feature h corresponding to position i can be obtained by the text feature extraction module 411. i , the text feature h i , the second associated feature c i , the first associated feature d i and the word embedding vector w of the first character i Input to the normalization layer 416 to obtain the correlation feature o i =LayerNorm(c i +d i +w i +h i ).

[0242] In the process of calculating the associated features corresponding to each position in the third character attribute feature extraction model, due to the combination of the text features at that position in the text to be corrected, the character attributes of the second character at the first set position corresponding to the position, the character attributes of the first character at the position, and the word embedding vector of the first character, the obtained associated features corresponding to the position are more accurate, which can not only reflect the character features of the character at the position itself and the association relationship between the characters adjacent to the position, but also the contextual relationship of the position.

[0243] In the second embodiment, it is mentioned that not all of the above steps S22 to S24 are executed by a pre-built model.

[0244] Exemplarily, some of the above detailed steps of obtaining the associated features corresponding to the multiple positions are not implemented by the model. For example, the correspondence between characters and character attributes can be preset, and in the process of obtaining the character attributes of the first character or the second character, the correspondence between characters and character attributes can be searched without using the phonetic feature extraction model and the glyph feature extraction model.

[0245] For example, the code points corresponding to the strokes may be preset, and the glyph features corresponding to the characters are composed of the code points corresponding to the strokes constituting the characters.

[0246] Fourth embodiment

[0247] In combination with the second embodiment or the third embodiment described above, the process of "for each position, calculating the association degree of each pending character based on the association feature between the character vector of each pending character and the position" is described below.

[0248] The first implementation method of "for each position, calculating the association degree of each pending character based on the association features between the character vector of each pending character and the corresponding position" includes the following steps B61 to B62.

[0249] Step B61: searching for the preset character vector of each pending character from the correspondence between the preset characters and the character vectors.

[0250] Exemplarily, the character vector of a character is a character attribute; exemplary, the character vector of a character is a word embedding vector.

[0251] Step B62: multiplying the preset character vector of each undetermined character by the product of the associated feature corresponding to the position to determine the association degree of each undetermined character.

[0252] The first implementation described above does not pass the pre-built model.

[0253] The following describes the process of calculating the association degree of each pending character based on the association features of the preset character vector of each pending character and the corresponding position of the candidate character scoring model 42 for each position in combination with the third embodiment.

[0254] It is understandable that there are many ways to implement the candidate character scoring model 42, and the embodiments of the present application provide but are not limited to the following methods.

[0255] A first input terminal of the candidate character scoring model 42 is connected to an output terminal of the character attribute extraction model 41, and a second input terminal of the candidate character scoring model 42 is connected to a database.

[0256] Exemplarily, the database stores preset character vectors of each pending character. Exemplarily, the preset character vector of the pending character is a word embedding vector of the pending character, or the preset character vector of the pending character is a character attribute of the pending character, or the preset character vector of the pending character is obtained by training in the process of training a text error correction model.

[0257] Exemplarily, for each position, the candidate character scoring model 42 may calculate the degree of association between the character attribute corresponding to the position and each pending character based on the character attribute corresponding to the position and the preset character vector of each pending character.

[0258] For example, for position i, the correlation between the undetermined character (i, m) and position i is f(i, m)=o i *v m , where v m is a character vector containing the pointer to the pending character m at position i.

[0259] In an optional implementation, for each position, after obtaining the association degrees of multiple pending characters, a preset number of pending characters with higher association degrees among the multiple pending characters are divided into the candidate character set corresponding to the position, or, the pending characters with association degrees greater than or equal to a first threshold are divided into the candidate character set corresponding to the position.

[0260] Fifth embodiment

[0261] In an optional implementation, steps S23 to S24 may be performed through an association relationship extraction model.

[0262] Exemplarily, the step of obtaining the relevance score corresponding to each candidate text includes the following steps C01 to C03.

[0263] Step C01: obtaining semantic dependency information corresponding to each candidate character of each candidate text through an association relationship extraction model.

[0264] Exemplarily, the specific implementation method of step C01 includes the following steps C011 to step C013.

[0265] Step C011: obtaining connection scores corresponding to candidate characters contained in the candidate text, wherein the connection scores corresponding to the candidate characters at least represent the adjacent co-occurrence probability of each character in a character set, and the character set at least includes the candidate character and the next candidate character of the candidate character in the candidate text.

[0266] Step C012: Obtain the prediction probability corresponding to the candidate characters contained in the candidate text, the prediction probability corresponding to the candidate characters is the probability that the target position is the candidate character when the character before the target position is the character before the target position in the text to be corrected, and the target position is the position of the candidate character in the candidate text.

[0267] Step C013: Determine the connection scores and / or prediction probabilities corresponding to the candidate characters contained in the candidate text as the semantic dependency information corresponding to the candidate characters contained in the candidate text.

[0268] Step C02: Calculate the relevance score corresponding to each candidate text based on the semantic dependency information corresponding to each candidate character of each candidate text through the relevance extraction model.

[0269] Step C03: determining, through the association relationship extraction model, the candidate text with the highest association score among the plurality of candidate texts as the corrected text corresponding to the text to be corrected.

[0270] The network structure of the association relationship extraction model in the text error correction model is described below in conjunction with the second embodiment, the third embodiment, or the fourth embodiment.

[0271] It is understandable that there are many ways to implement the association relationship extraction model. The embodiments of the present application provide but are not limited to the following three.

[0272] The first association relationship extraction model includes: a self-attention mechanism model 43 and a corrected text determination model 45.

[0273] like Figure 5 , which is a schematic diagram of an implementation method of the network structure of the self-attention mechanism model provided in an embodiment of the present application.

[0274] Exemplarily, the self-attention mechanism model includes: Self-Attention Sub-Layer 51, normalization layer 52, normalization layer 53 and Feed-Forward Sub-Layer 54.

[0275] For example, the normalization layer 52 can be replaced by a fully connected layer or a convolutional layer. For example, the normalization layer 53 can be replaced by a fully connected layer or a convolutional layer.

[0276] Exemplarily, the self-attention mechanism model may further include an embedding layer, so that a word embedding vector of each candidate character may be obtained.

[0277] Exemplarily, the word embedding vector of each candidate character may be pre-set.

[0278] Exemplarily, the word embedding vector of the candidate character may be a character vector.

[0279] The output end of the text feature extraction module in the character attribute extraction model 41 (ie, the first output end of the character attribute extraction model) is connected to the input end of the Self-Attention Sub-Layer 51 .

[0280] Combine the following Figure 5The network structure of the self-attention mechanism model shown illustrates the process of the self-attention mechanism model obtaining the connection score corresponding to each candidate character of each candidate text. Assuming that the candidate character is located at the target position of the candidate text, the process includes the following steps C11 to C14.

[0281] Step C11: Acquire the text feature corresponding to the target position.

[0282] The text feature corresponding to the target position is obtained based on context information of the target position in the text to be corrected, and the text feature includes context semantic features and / or context syntactic features of the target position.

[0283] Exemplarily, the text to be corrected may be input into the text feature extraction model 411 , so that the text feature corresponding to the target position is obtained through the text feature extraction model 411 .

[0284] Step C12: Acquire the text feature corresponding to the second set position.

[0285] The text feature corresponding to the second set position is obtained based on the context information of the second set position corresponding to the target position in the erroneous text to be corrected, and the second set position corresponding to the target position includes the next position of the target position and / or the previous position of the target position.

[0286] Illustratively, the second set position may include one or more positions.

[0287] Exemplarily, the text to be corrected may be input into the text feature extraction model 411 , so that the text features corresponding to each position in the second set position are obtained through the text feature extraction model 411 .

[0288] Step C13: Obtain the first word embedding vector of the candidate character.

[0289] Exemplarily, the first word embedding vector of the candidate character can be obtained through the embedding layer.

[0290] Step C14: Calculate a first connection matrix based on the text feature corresponding to the target position, the text feature corresponding to the second set position, and the first character embedding vector.

[0291] Exemplarily, the first connection matrix can be calculated by Self-Attention Sub-Layer 51 based on the text features corresponding to the target position, the text features corresponding to the second set position, and the first character embedding vector.

[0292] For example, the target position is still taken as position i for illustration, then the first connection matrix p i,m=Attention(Q i,m *W Q ,K i *W K ,V i *W V ).

[0293] in, Q i,m =w i,m , where W Q , W K , W V is the mapping matrix that maps Q, K, and V to the same dimension, w i,m is the word embedding vector of the pointer to the pending character m at position i.

[0294] In summary, in the present embodiment, h i and h i+1 As a key-value pair (KV), the self-attention mechanism model considers the contextual relationship between position i and position i+1, and converts w i,m As the input sequence (Q), the attention mechanism model takes into account the character characteristics of the candidate character at position i.

[0295] Step C15: Based on the first connection matrix, calculate the connection score of the candidate character.

[0296] Exemplarily, the connection score may be obtained based on the first connection matrix by the Feed-Forward Sub-Layer 54 .

[0297] In an optional implementation, the calculation formula of the connection score g corresponding to the candidate character is as follows: S = FFN (concat (p i,m )); g( i,m,i+1,n ) = S*V. Wherein, V is a preset fixed vector, and the purpose of multiplying S and V is to make the matrix S become a fraction.

[0298] In an optional implementation, the above steps C11 to C15 do not consider the association between the candidate character at position i and the candidate character at position i+1. Based on this, step C15 may include the following steps C151 to C152.

[0299] Step C151: Obtain a second character embedding vector of the candidate character located at the second set position in the candidate text.

[0300] Exemplarily, a second word embedding vector of the candidate character at the second set position in the candidate text may be obtained through an embedding layer.

[0301] Exemplarily, if the second set position includes multiple positions, the second character embedding vectors corresponding to the candidate characters located at the multiple positions can be obtained.

[0302] Step C152: Calculate a second connection matrix based on the text features corresponding to the target position, the text features corresponding to the second set position, and the second character embedding vector.

[0303] Exemplarily, the second connection matrix is ​​calculated by Self-Attention Sub-Layer 51 based on the text features corresponding to the target position, the text features corresponding to the second set position, and the second character embedding vector.

[0304] Exemplarily, the second connection matrix q i,n =Attention(Q i+1n ,*W Q ,K i *W K ,V i *W V ), Q i+1,n =w i+1,n , w i+1,n is the word embedding vector of the pointer to the pending character n at position i+1.

[0305] Step C153: Obtain a connection score based on the first connection matrix and the second connection matrix.

[0306] Exemplarily, the connection score may be obtained by the Feed-Forward Sub-Layer 54 based on the first connection matrix and the second connection matrix.

[0307] In an optional implementation, the calculation formula of the connection score g corresponding to the candidate character is as follows: S = FFN (concat (p i,m ,q i,n )); g( i,m,i+1,n ) = S*V. Wherein, V is a preset fixed vector, and the purpose of multiplying S and V is to make the matrix S become a fraction.

[0308] In an optional implementation, in order to make the first connection matrix and the second connection matrix in a relatively fixed distribution so as to accelerate the convergence speed of the text correction model in the process of the text correction model, step C153 can specifically include the following method steps C1531 to C1533.

[0309] Step C1531: Based on the first connection matrix and the first word embedding vector of the candidate character located at the target position, a third connection matrix is ​​calculated.

[0310] Exemplarily, the third connection matrix may be calculated by the normalization layer 52 based on the first connection matrix and the first word embedding vector of the candidate character at the target position.

[0311] For example, Figure 5 Take the normalization layer as LayNorm layer as an example. Then the third connection matrix is ​​p' i,m =LayerNorm(p i,m +w i,m ).

[0312] Step C1532: Based on the second connection matrix and the word embedding vector of the candidate character located at the second set position, a fourth connection matrix is ​​calculated.

[0313] Exemplarily, the fourth connection matrix may be calculated by the normalization layer 53 based on the second connection matrix and the word embedding vector of the candidate character located at the second set position.

[0314] For example, Figure 5 Take the normalization layer as LayNorm layer as an example. Then the fourth connection matrix is ​​q' i,n =LayerNorm(q i,n +w i+1,n ).

[0315] Step C1533: Obtain a connection score based on the third connection matrix and the fourth connection matrix.

[0316] Exemplarily, the connection score is obtained by Feed-Forward Sub-Layer 54 based on the third connection matrix and the fourth connection matrix. Exemplarily, the formula for calculating the connection score is as follows: S = FFN (concat (p' i,m ,q' i,n )); g( i,m,i+1,n )=S*V.

[0317] In an optional implementation, the relevance score of the candidate text may be calculated by the corrected text determination model 45 based on the connection score of each candidate character in the candidate text.

[0318] Exemplarily, the relevance score of the candidate text=the sum of the connection scores of the candidate characters constituting the candidate text.

[0319] In an optional implementation, the relevance score of the candidate text may be calculated by the corrected text determination model 45 based on the connection score of each candidate character in the candidate text and the relevance of each candidate character.

[0320] Exemplarily, the relevance score of the candidate text=the sum of the connection scores of the candidate characters constituting the candidate text and the relevance of the candidate characters constituting the candidate text.

[0321] The first association relationship extraction model takes into account the association relationship between adjacent candidate characters in the candidate text, as well as the contextual text features between connected positions, so the connection score corresponding to each candidate character in the candidate text is relatively accurate.

[0322] The second association relationship extraction model includes a text prediction model 44 and a corrected text determination model 45 .

[0323] Exemplarily, the text prediction model may be a unidirectional text prediction model or a bidirectional text prediction model.

[0324] Exemplarily, the unidirectional text prediction model may be a GPT (Generative Pre-Training) model. The GPT model is a unidirectional Transformer language model that is pre-trained in advance. The probability of a candidate character in position i can be predicted based on the first i-1 characters of the text to be corrected. For example, the predicted probability of each candidate character being at position i is predicted based on [the character at position 1, the character at position 2, ..., the character at position i-1] in the text to be corrected.

[0325] Exemplarily, the bidirectional text prediction model may be a bidirectional LSTM model.

[0326] The bidirectional text prediction model can predict the probability of a candidate character at position i based on the first i-1 characters of the text to be corrected; and, based on the characters at position i+1 and later in the text to be corrected, predict the probability of each candidate character at position i, that is, based on [the character at position i+1, the character at position i+2, the character at position i+3, ...], predict the probability of each candidate character at position i.

[0327] Exemplarily, if the association relationship extraction model is a bidirectional text prediction model, the association score corresponding to the candidate character may be the average of the two prediction probabilities.

[0328] In combination with the network structure of the second association relationship extraction model, the step of obtaining the association score of the candidate text is described, and the step includes step C21 to step C22.

[0329] Step C21: obtaining the predicted probability of the candidate character being located at the position through a text prediction model.

[0330] Step C22: The corrected text determination model 45 calculates the relevance score corresponding to the candidate text based on the prediction probability corresponding to each candidate character in the candidate text.

[0331] Exemplarily, the relevance score of the candidate text=the sum of the predicted probabilities of the candidate characters constituting the candidate text.

[0332] Exemplarily, step C22 includes: calculating the association score corresponding to the candidate text based on the prediction probability corresponding to each candidate character in the candidate text and the association degree corresponding to each candidate character in the candidate text by the corrected text determination model 45.

[0333] Exemplarily, the relevance score of the candidate text=the sum of the predicted probabilities of the candidate characters constituting the candidate text and the relevance of the candidate characters constituting the candidate text.

[0334] Although the second association relationship extraction model takes the contextual relationship into consideration, it ignores the association relationship between two adjacent candidate characters in the candidate text. Based on this, the embodiment of the present application provides a network structure of a third association relationship extraction model.

[0335] The network structure of the third association relationship extraction model includes: a self-attention mechanism model 43, a text prediction model 44, and a model determined by correcting text 45.

[0336] The process of obtaining the relevance score of the candidate text is described in conjunction with the network structure of the third relevance extraction model. The method includes steps C31 to C33.

[0337] Step C31: Obtain the connection score corresponding to each candidate character contained in the candidate text through the self-attention mechanism model.

[0338] The process of obtaining the connection score corresponding to each candidate character contained in the candidate text through the self-attention mechanism model can be referred to the process shown in steps C11 to C15, which will not be repeated here.

[0339] Step C32: Obtain the prediction probability of each candidate character contained in the candidate text through the text prediction model.

[0340] The process of obtaining the prediction probability of each candidate character contained in the candidate text through the text prediction model can be referred to step C21, which will not be repeated here.

[0341] Step C33: The corrected text determination model 45 calculates the association score of the candidate text based on the connection score and the prediction probability corresponding to each candidate character in the candidate text.

[0342] Exemplary, the relevance score of the candidate text Among them, R(i,m) refers to the predicted probability of candidate character m at position i.

[0343] Exemplarily, step C33 includes: calculating the association score of the candidate text based on the connection score corresponding to each candidate character in the candidate text, the prediction probability corresponding to each candidate character in the candidate text, and the association degree corresponding to each candidate character in the candidate text by correcting the text determination model 45.

[0344] Exemplary, the relevance score of the candidate text

[0345] The third association relationship extraction model mentioned above takes into account the association relationship of the context, the association relationship between adjacent candidate texts in the candidate text, and the predicted probability of each candidate character being located at the corresponding position, thereby making the association score of the candidate text more accurate.

[0346] In an optional implementation, the process of obtaining the connection scores corresponding to the candidate characters contained in the candidate text may not pass through the model. The connection scores corresponding to the candidate characters may be adjacent co-occurrence probabilities.

[0347] Exemplarily, the process of obtaining connection scores corresponding to candidate characters included in the candidate text may include the following steps C41 to C43.

[0348] Step C41: for each candidate character in the candidate text, obtain a character set, wherein the character set at least includes the candidate character and the next candidate character of the candidate character in the candidate text.

[0349] Step C42: obtaining a target text including the character set from a plurality of texts, wherein two characters in the character set included in the target text are adjacent to each other.

[0350] Step C43: Based on the number of the target texts and the total number of the multiple texts, the adjacent co-occurrence probabilities corresponding to the candidate characters are calculated.

[0351] In order for those skilled in the art to better understand the process of obtaining the corrected text using the pre-built text error correction model in the embodiments of the present application, the process of obtaining the corrected text using the text error correction model is described below in combination with the second embodiment, the third embodiment, the fourth embodiment and the fifth embodiment.

[0352] like Figure 6 As shown, it is a network structure diagram of an implementation method of the text error correction model provided in an embodiment of the present application.

[0353] like Figure 6 As shown, the text error correction model includes: a character attribute extraction model 41, a candidate character scoring model 42, a self-attention mechanism model 43, a text prediction model 44 and a corrected text determination model 45.

[0354] Exemplarily, the character attribute extraction model 41, the candidate character scoring model 42, the self-attention mechanism model 43, the text prediction model 44, and the corrected text determination model 45 constitute a dynamic connection network (DCN, Dynamic Connected Networks).

[0355] Assume that the text to be corrected is "我真户秃" (I am really bald), and the candidate text is "我真尴涂" (I am really embarrassed). For the sake of clearly demonstrating the relationship between each model, the following takes the position i = position 3 as an example for explanation.

[0356] Through the character attribute extraction model 41, the text feature h3 corresponding to position 3 and the text feature h4 corresponding to position 4 in the text to be corrected can be obtained. The associated feature o3 of position 3 output by the character attribute extraction model 41, and the candidate character scoring model 42 is calculated based on the associated feature o3 and the character vector of "尴" located at position 3 in the candidate text, f(3,尴) = o3 * v 尴 , where v 尴 refers to the character vector of "尴".

[0357] The self-attention mechanism model 43 is based on the text feature h3 corresponding to position 3, the text feature h4 corresponding to position 4, the character embedding vector w 尴 of "尴", and the character embedding vector w 涂 of "涂" to calculate the connection score g(尴,涂).

[0358] Figure 6 Taking the unidirectional text prediction model as an example for the text prediction model, the content input to the text prediction model is "我真" (I am really).

[0359] If the text prediction model takes the bidirectional text prediction model as an example, the content input to the text prediction model is "我真" (I am really) and "秃" (bald).

[0360] The predicted probability R(3,尴) that the character at position 3 in the text to be corrected is "尴" is output through the text prediction model.

[0361] It can be understood that the above takes position 3 as an example for explanation. Through Figure 6 , the character attribute extraction model can also obtain f(1,我), f(2,真), f(4,涂) in the candidate text "我真尴涂".

[0362] Through Figure 6 , the self-attention mechanism model 43 and the character attribute extraction model 41 shown can also obtain g(我,真), g(真,尴), g(尴,涂).

[0363] Through Figure 6The text prediction model 44 shown can also obtain R(1, me), R(2, true), and R(4, smear).

[0364] Exemplarily, through the corrected text determination model 45, the following formula can be used: To obtain the correlation score of the candidate text "I'm really embarrassed and smeared". For the above example, N = 4, then:

[0365]

[0366] Through Figure 6 The corrected text determination model 45 shown can obtain the correlation scores corresponding to multiple candidate texts, so that the candidate text "I'm really confused" with the highest correlation score can be selected, and thus the corrected text is obtained.

[0367] Sixth Embodiment

[0368] Next, in combination with the above-mentioned second embodiment or third embodiment or fourth embodiment or fifth embodiment, the training process of the text error correction model mentioned in the embodiments of the present application will be described.

[0369] Exemplarily, the process of training the text error correction model includes the following steps D11 to D14.

[0370] Step D11: Input the sample text into the machine learning model.

[0371] Step D12: Obtain the correlation scores of the candidate texts corresponding to the sample text output by the machine learning model.

[0372] In an optional implementation manner, the network structure of the machine learning model can refer to the network structure of the text error correction model shown in the second embodiment or third embodiment or fourth embodiment or fifth embodiment.

[0373] In an optional implementation manner, at least one of the technologies such as artificial neural network, confidence network, reinforcement learning, transfer learning, inductive learning, and formula learning in machine learning is involved in the process of training the machine learning model.

[0374] Exemplarily, at least one of the technologies such as artificial neural network, confidence network, reinforcement learning, transfer learning, inductive learning, and formula learning in machine learning is involved in the process of training the machine learning model.

[0375] Exemplarily, the machine learning model can be any one of a neural network model, a logistic regression model, a linear regression model, a support vector machine (SVM), Adaboost, XGboost, and a Transformer-Encoder model.

[0376] Exemplarily, the neural network model can be any one of a recurrent neural network-based model, a convolutional neural network-based model, and a Transformer-encoder-based classification model.

[0377] Exemplarily, the machine learning model may be a deep hybrid model of a recurrent neural network-based model, a convolutional neural network-based model, and a Transformer-encoder-based classification model.

[0378] Exemplarily, the machine learning model can be any one of an attention-based deep model, a memory network-based deep model, and a short text classification model based on deep learning.

[0379] The short text classification model based on deep learning is a recurrent neural network (RNN) or a convolutional neural network (CNN) or a variant based on the recurrent neural network or the convolutional neural network.

[0380] For example, some simple domain adaptation modifications can be made on the pre-trained model to obtain a machine learning model.

[0381] Exemplarily, “simple domain adaptation” includes but is not limited to re-pre-training an already pre-trained model using large-scale unsupervised domain corpus, and / or compressing an already pre-trained model by model distillation.

[0382] Step D13: Calculate the loss function based on the relevance scores of the candidate texts corresponding to the sample text output by the machine learning model.

[0383] The greater the correlation score of the annotated corrected text corresponding to the sample text, the smaller the loss function, and the smaller the correlation score of the annotated corrected text corresponding to the sample text, the larger the loss function.

[0384] Exemplarily, the formula of the loss function can be as follows:

[0385]

[0386] Where M represents the total number of candidate texts. X represents the sample text, Y represents the annotated corrected text corresponding to the sample text, Score(X,Y) represents the correlation score of the annotated corrected text output by the machine learning model, and Y' j Represents the jth candidate text corresponding to the sample text, Score(X,Y j ') represents the relevance score of the jth candidate text output by the machine learning model.

[0387] Step D14: Update the parameters in the machine learning model with the goal of minimizing the loss function.

[0388] In an optional implementation, if the network structure of the machine learning model in the sixth embodiment is the network structure of the text error correction model shown in the second embodiment, the third embodiment, the fourth embodiment, or the fifth embodiment, then in order to quickly train the machine learning model, the text prediction model 44 and the text feature extraction module 411 may be pre-trained before training the machine learning model. Then the entire machine learning model is trained together.

[0389] The method is described in detail in the embodiments disclosed in the above-mentioned application. The method of the application can be implemented by various forms of devices. Therefore, the application also discloses a device, and a specific embodiment is given below for detailed description.

[0390] Seventh embodiment

[0391] like Figure 7 As shown, it is a structural diagram of a text error correction device provided in an embodiment of the present application, and the device includes: a first acquisition module 71, a second acquisition module 72, a third acquisition module 73 and a determination module 74, wherein:

[0392] The first acquisition module 71 is used to acquire the text to be corrected.

[0393] The second acquisition module 72 is used to acquire candidate character sets corresponding to multiple positions respectively, where the candidate character sets corresponding to the positions include candidate characters associated with the characters located at the positions in the text to be corrected.

[0394] The third acquisition module 73 is used to obtain the association scores corresponding to multiple candidate texts respectively, where the character at each position of the candidate text is a candidate character in the candidate character set corresponding to the position, and the association score corresponding to the candidate text represents the probability that each candidate character contained in the candidate text constitutes the candidate text.

[0395] The determination module 74 is used to determine the corrected text corresponding to the text to be corrected from the multiple candidate texts according to the relevance scores respectively corresponding to the multiple candidate texts.

[0396] In an optional implementation, the second acquisition module includes:

[0397] A first acquisition unit, configured to acquire, for each of the positions, character attributes of a first character located at the position in the text to be corrected, wherein the character attributes include a glyph feature and / or a phonetic feature;

[0398] A second acquisition unit is used to acquire, for each of the positions, a character attribute of a second character in the text to be corrected that is located at a first set position corresponding to the position, wherein the first set position corresponding to the position at least includes a position before the position and / or a position after the position;

[0399] A first calculation unit, configured to calculate an associated feature corresponding to the position based on a character attribute of the first character and a character attribute of the second character;

[0400] The third acquisition unit is used to obtain a candidate character belonging to the candidate character set corresponding to the position from the undetermined character based on the associated feature and a preset character vector of the undetermined character.

[0401] In an optional implementation, the first computing unit includes:

[0402] A first character calculation unit, used for calculating a first associated feature based on the glyph feature of the first character and the glyph feature of the second character;

[0403] A second character calculation unit, configured to calculate a second associated feature based on the phonetic feature of the first character and the phonetic feature of the second character;

[0404] A first acquisition subunit, configured to acquire a word embedding vector corresponding to the first character;

[0405] The third calculation subunit is used to calculate the association feature based on the first association feature, the second association feature and the word embedding vector of the first character.

[0406] In an optional implementation, the third computing subunit includes:

[0407] A first acquisition submodule, configured to acquire text features corresponding to the position based on context information of the position in the text to be corrected, wherein the text features include context semantic features and / or context syntactic features of the position;

[0408] The first calculation submodule is used to calculate the association feature based on the text feature, the first association feature, the second association feature and the word embedding vector of the first character.

[0409] In an optional implementation, the third obtaining unit includes:

[0410] A first determination subunit is used to determine the product of the association feature and the preset character vectors of the multiple characters to be determined, which is the association degree of the multiple characters to be determined;

[0411] The division subunit is used to divide a preset number of pending characters with higher association among the multiple pending characters into a candidate character set corresponding to the position, or to divide the pending characters with an association greater than or equal to a first threshold into a candidate character set corresponding to the position.

[0412] In an optional implementation, the third acquisition module includes:

[0413] a fourth acquisition unit, configured to acquire a connection score corresponding to a candidate character included in the candidate text, wherein the connection score corresponding to the candidate character at least represents a probability of adjacent co-occurrence of each character in a character set, wherein the character set at least includes the candidate character and the next candidate character of the candidate character in the candidate text;

[0414] a fifth obtaining unit, configured to obtain a prediction probability corresponding to a candidate character contained in the candidate text, wherein the prediction probability corresponding to the candidate character is a probability that the target position is the candidate character when the character before the target position is the character before the target position in the text to be corrected, and the target position is a position of the candidate character in the candidate text;

[0415] The second calculation unit is used to calculate the relevance score of the candidate text based on the connection scores corresponding to the candidate characters in the candidate text and the prediction probability.

[0416] In an optional implementation, the fourth obtaining unit includes:

[0417] A second acquisition subunit is used to acquire text features corresponding to the target position, wherein the text features corresponding to the target position are obtained based on context information of the target position in the text to be corrected, and the text features include context semantic features and / or context syntactic features of the target position;

[0418] a third acquisition subunit, configured to acquire a text feature corresponding to a second set position, wherein the text feature corresponding to the second set position is obtained based on context information of a second set position corresponding to the target position in the erroneous text to be corrected, wherein the second set position corresponding to the target position includes a next position of the target position and / or a previous position of the target position;

[0419] A fourth acquisition subunit, used to acquire a first word embedding vector of the candidate character;

[0420] a fifth acquisition subunit, configured to acquire a second character embedding vector of the candidate character located at the second set position in the candidate text;

[0421] The fourth calculation subunit is used to calculate the connection score corresponding to the candidate character based on the text feature corresponding to the target position, the text feature corresponding to the second set position, the first character embedding vector and the second character embedding vector.

[0422] In an optional implementation, the second computing unit includes:

[0423] a fifth calculation subunit, configured to calculate an association score of the candidate text based on a connection score corresponding to the candidate characters in the candidate text, a prediction probability corresponding to the candidate characters in the candidate text, and an association degree corresponding to the candidate characters in the candidate text;

[0424] Among them, the association degree corresponding to the candidate character is the product of the association feature corresponding to the position of the candidate character and the preset character vector of the candidate character, and the association feature corresponding to the position is obtained based on the character attributes of the first character located at the position in the text to be corrected, and the character attributes include glyph features and / or phonetic features.

[0425] In an optional implementation, a pre-built text error correction model is included.

[0426] A text error correction model is used to obtain a set of candidate characters corresponding to multiple positions respectively; obtain relevance scores corresponding to multiple candidate texts respectively; and determine the candidate text with the highest relevance score among the multiple candidate texts as the corrected text corresponding to the text.

[0427] The text error correction model is obtained by training a machine learning model using sample text as input and taking the maximum correlation score of the annotated corrected text corresponding to the sample text output by the machine learning model as the training target.

[0428] In an optional implementation, the loss function of the machine learning model is obtained based on the association scores of each candidate text corresponding to the sample text output by the machine learning model, and each candidate text includes the annotated corrected text, wherein the larger the association score of the annotated corrected text corresponding to the sample text, the smaller the loss function, and the smaller the association score of the annotated corrected text corresponding to the sample text, the larger the loss function.

[0429] In an optional implementation, the text error correction model includes: a character attribute extraction model and a candidate character scoring model, wherein the process of obtaining the candidate character set corresponding to each position is as follows:

[0430] The character attribute extraction model is used to obtain the character attribute of the first character located at the position in the text to be corrected, wherein the character attribute includes a glyph feature and / or a phonetic feature;

[0431] The character attribute extraction model is used to obtain the character attribute of the second character in the text to be corrected that is located at a first set position corresponding to the position, wherein the first set position corresponding to the position at least includes a position before the position and / or a position after the position;

[0432] The character attribute extraction model is used to calculate the associated feature corresponding to the position based on the character attribute of the first character and the character attribute of the second character;

[0433] The candidate character scoring model is used to obtain candidate characters belonging to the candidate character set corresponding to the position from the pending characters based on the associated features and the preset character vector of the pending characters.

[0434] In an optional implementation, the text error correction model includes: a self-attention mechanism model, a text prediction model, and a corrected text determination model, wherein the process of obtaining the relevance score corresponding to each candidate text includes:

[0435] The self-attention mechanism model is used to obtain connection scores corresponding to candidate characters contained in the candidate text, where the connection scores corresponding to the candidate characters at least represent the adjacent co-occurrence probability of each character in a character set, and the character set at least includes the candidate character and the next candidate character of the candidate character in the candidate text;

[0436] The text prediction model is used to obtain the prediction probability corresponding to the candidate characters contained in the candidate text, the prediction probability corresponding to the candidate characters being the probability that the target position is the candidate character when the character before the target position is the character before the target position in the text to be corrected, and the target position is the position of the candidate character in the candidate text;

[0437] The corrected text determination model is used to calculate the relevance score of the candidate text based on the connection scores corresponding to the candidate characters in the candidate text and the prediction probability.

[0438] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0439] Eighth embodiment

[0440] like Figure 8 As shown, it is a block diagram of the text correction device provided in an embodiment of the present application.

[0441] The text error correction device includes, but is not limited to: a processor 81 , a memory 82 , a network interface 83 , an I / O controller 84 , and a communication bus 85 .

[0442] It should be noted that those skilled in the art can understand that Figure 8 The structure of the text correction device shown in the figure does not constitute a limitation on the text correction device. The text correction device may include Figure 8 More or fewer components may be shown, or certain components may be combined, or the components may be arranged differently.

[0443] Combine the following Figure 8 A detailed introduction to the various components of the text correction device:

[0444] The processor 81 is the control center of the text correction device. It uses various interfaces and lines to connect the various parts of the entire text correction device. By running or executing software programs and / or modules stored in the memory 82, and calling the data stored in the memory 82, it performs various functions of the text correction device and processes data, thereby monitoring the text correction device as a whole. The processor 81 may include one or more processing units; illustratively, the processor 81 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 81.

[0445] The processor 81 may be a central processing unit (CPU), or an application specific integrated circuit ASIC (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;

[0446] The memory 82 may include a memory, such as a high-speed random access memory (RAM) 821 and a read-only memory (ROM) 822, and may also include a large-capacity storage device 823, such as at least one disk storage device, etc. Of course, the text error correction device may also include hardware required for other services.

[0447] The memory 82 is used to store instructions executable by the processor 81. The processor 81 has the following functions: obtaining the text to be corrected;

[0448] Acquire candidate character sets corresponding to a plurality of positions respectively, wherein the candidate character sets corresponding to the positions include candidate characters associated with the characters located at the positions in the text to be corrected;

[0449] Obtaining association scores corresponding to a plurality of candidate texts respectively, wherein the character at each position of the candidate text is a candidate character in the candidate character set corresponding to the position, and the association score corresponding to the candidate text represents the probability that each candidate character contained in the candidate text constitutes the candidate text;

[0450] According to the relevance scores respectively corresponding to the multiple candidate texts, a corrected text corresponding to the text to be corrected is determined from the multiple candidate texts.

[0451] Optionally, the detailed functions and extended functions of the program may refer to the above description.

[0452] A wired or wireless network interface 83 is configured to connect the text error correction device to a network.

[0453] The processor 81, the memory 82, the network interface 83 and the I / O controller 84 may be interconnected via a communication bus 85, which may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc.

[0454] In an exemplary embodiment, the text error correction device can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned text error correction method.

[0455] An embodiment of the present application also provides a computer-readable storage medium that can be directly loaded into the internal memory of a computer, such as the above-mentioned memory 82, and contains software code. After being loaded and executed by a computer, the computer program can implement the steps shown in any embodiment of the above-mentioned text error correction method.

[0456] It should be noted that the features described in the various embodiments in this specification can be replaced or combined with each other. For the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0457] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0458] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0459] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A text error correction method, characterized in that: include: Get the text to be corrected; Based on the association features corresponding to each position in the text to be corrected and the preset character vector of the undetermined character, a candidate character of a candidate character set corresponding to each position is obtained from the undetermined character to obtain candidate character sets corresponding to a plurality of positions, wherein the candidate character sets corresponding to the positions include candidate characters having an association relationship with the characters located at the positions in the text to be corrected; Obtaining association scores corresponding to a plurality of candidate texts respectively, wherein the character at each position of the candidate text is a candidate character in the candidate character set corresponding to the position, and the association score corresponding to the candidate text represents the probability that each candidate character contained in the candidate text constitutes the candidate text; Determining a corrected text corresponding to the text to be corrected from the multiple candidate texts according to the relevance scores respectively corresponding to the multiple candidate texts; Wherein, the step of obtaining a candidate character of a candidate character set corresponding to each position from the to-be-determined character based on the associated features corresponding to each position in the to-be-corrected text and the preset character vector of the to-be-determined character comprises: Determine the product of the association feature and the preset character vectors of the multiple characters to be determined, as the association degree of the multiple characters to be determined; A preset number of pending characters with higher relevance among the plurality of pending characters are divided into a candidate character set corresponding to the position, or, pending characters with a relevance greater than or equal to a first threshold are divided into a candidate character set corresponding to the position.

2. The text error correction method according to claim 1, characterized in that: Acquiring the associated features includes: Acquire character attributes of the first character located at the position in the text to be corrected, wherein the character attributes include a glyph feature and / or a pronunciation feature; Acquire a character attribute of a second character in the text to be corrected that is located at a first set position corresponding to the position, wherein the first set position corresponding to the position at least includes a position before the position and / or a position after the position; Based on the character attribute of the first character and the character attribute of the second character, the associated feature corresponding to the position is calculated.

3. The text error correction method according to claim 2, characterized in that: The calculating, based on the character attribute of the first character and the character attribute of the second character, the associated feature corresponding to the position includes: Calculating a first associated feature based on the glyph feature of the first character and the glyph feature of the second character; Calculating a second associated feature based on the phonetic feature of the first character and the phonetic feature of the second character; Obtain a word embedding vector corresponding to the first character; The association feature is calculated based on the first association feature, the second association feature and the word embedding vector of the first character.

4. The text error correction method according to claim 3, characterized in that: The calculating the association feature based on the first association feature, the second association feature and the word embedding vector of the first character includes: Based on the context information of the position in the text to be corrected, obtaining text features corresponding to the position, the text features including context semantic features and / or context syntactic features of the position; The association feature is calculated based on the text feature, the first association feature, the second association feature and the word embedding vector of the first character.

5. The text error correction method according to any one of claims 1 to 4, characterized in that: Obtaining the relevance score corresponding to the candidate text includes: Obtaining connection scores corresponding to candidate characters included in the candidate text, wherein the connection scores corresponding to the candidate characters at least represent adjacent co-occurrence probabilities of each character in a character set, wherein the character set at least includes the candidate character and the next candidate character of the candidate character in the candidate text; Obtaining the prediction probability corresponding to the candidate characters contained in the candidate text, the prediction probability corresponding to the candidate characters being the probability that the target position is the candidate character when the character before the target position is the character before the target position in the text to be corrected, the target position being the position of the candidate character in the candidate text; Based on the connection scores corresponding to the candidate characters in the candidate text and the predicted probabilities, the association score of the candidate text is calculated.

6. The text error correction method according to claim 5, characterized in that: Obtaining the connection score corresponding to the candidate character includes: Acquire a text feature corresponding to the target position, wherein the text feature corresponding to the target position is obtained based on context information of the target position in the text to be corrected, and the text feature includes a context semantic feature and / or a context syntactic feature of the target position; Acquire a text feature corresponding to a second set position, where the text feature corresponding to the second set position is obtained based on context information of a second set position corresponding to the target position in the text to be corrected, where the second set position corresponding to the target position includes a next position of the target position and / or a previous position of the target position; Obtaining a first word embedding vector of the candidate character; Obtaining a second word embedding vector of the candidate character located at the second set position in the candidate text; Based on the text feature corresponding to the target position, the text feature corresponding to the second set position, the first character embedding vector and the second character embedding vector, a connection score corresponding to the candidate character is calculated.

7. The text error correction method according to claim 5, characterized in that: The calculating the association score of the candidate text based on the connection scores corresponding to the candidate characters in the candidate text and the prediction probability includes: Calculate the relevance score of the candidate text based on the connection score corresponding to the candidate characters in the candidate text, the prediction probability corresponding to the candidate characters in the candidate text, and the relevance degree corresponding to the candidate characters in the candidate text; Among them, the association degree corresponding to the candidate character is the product of the association feature corresponding to the position of the candidate character and the preset character vector of the candidate character, and the association feature corresponding to the position is obtained based on the character attributes of the first character located at the position in the text to be corrected, and the character attributes include glyph features and / or phonetic features.

8. The text error correction method according to any one of claims 1, 2, 3, 4 or 6, characterized in that: The step of obtaining candidate character sets corresponding to the plurality of positions, obtaining association scores corresponding to the plurality of candidate texts, and determining a corrected text corresponding to the text to be corrected from the plurality of candidate texts according to the association scores corresponding to the plurality of candidate texts, comprises: Using a pre-built text error correction model, obtain candidate character sets corresponding to multiple positions; Using the text error correction model to obtain relevance scores corresponding to multiple candidate texts; Using the text error correction model, based on the correlation scores in the multiple candidate texts, determining a corrected text corresponding to the text to be corrected from the multiple candidate texts; The text error correction model is obtained by training a machine learning model using sample text as input and taking the maximum correlation score of the annotated corrected text corresponding to the sample text output by the machine learning model as the training target.

9. A text error correction device, characterized in that: include: A first acquisition module is used to acquire the text to be corrected; A second acquisition module is used to obtain candidate characters of a candidate character set corresponding to each position from the to-be-determined character using a third acquisition unit based on the associated features corresponding to each position in the to-be-determined text and a preset character vector of the to-be-determined character, so as to obtain candidate character sets corresponding to a plurality of positions respectively, wherein the candidate character sets corresponding to the positions include candidate characters having an associated relationship with the characters located at the positions in the to-be-determined text; A third acquisition module is used to obtain association scores corresponding to a plurality of candidate texts, wherein the character at each position of the candidate text is a candidate character in the candidate character set corresponding to the position, and the association score corresponding to the candidate text represents the probability that each candidate character contained in the candidate text constitutes the candidate text; A determination module, configured to determine a corrected text corresponding to the text to be corrected from the plurality of candidate texts according to the relevance scores respectively corresponding to the plurality of candidate texts; The third acquisition unit includes: A first determination subunit is used to determine the product of the association feature and the preset character vectors of the multiple characters to be determined, which is the association degree of the multiple characters to be determined; The division subunit is used to divide a preset number of pending characters with higher association among the multiple pending characters into a candidate character set corresponding to the position, or to divide the pending characters with an association greater than or equal to a first threshold into a candidate character set corresponding to the position.

10. The text error correction device according to claim 9, characterized in that: For each position, the second acquisition module includes: A character attribute extraction model, used for obtaining the character attribute of the first character located at the position in the text to be corrected, wherein the character attribute includes a glyph feature and / or a phonetic feature; The character attribute extraction model is used to obtain the character attribute of the second character in the text to be corrected that is located at a first set position corresponding to the position, wherein the first set position corresponding to the position at least includes a position before the position and / or a position after the position; The character attribute extraction model is used to calculate the associated feature corresponding to the position based on the character attribute of the first character and the character attribute of the second character; The candidate character scoring model is used to obtain candidate characters belonging to the candidate character set corresponding to the position from the pending characters based on the associated features and the preset character vector of the pending characters.

11. The text error correction device according to claim 9 or 10, characterized in that: For each candidate text, the third acquisition module includes: A self-attention mechanism model is used to obtain connection scores corresponding to candidate characters contained in the candidate text, wherein the connection scores corresponding to the candidate characters at least represent the adjacent co-occurrence probability of each character in a character set, and the character set at least includes the candidate character and the next candidate character of the candidate character in the candidate text; A text prediction model, used to obtain prediction probabilities corresponding to candidate characters contained in the candidate text, wherein the prediction probabilities corresponding to the candidate characters are probabilities that the target position is the candidate character when the character before the target position is the character before the target position in the text to be corrected, and the target position is the position of the candidate character in the candidate text; The corrected text determination model is used to calculate the relevance score of the candidate text based on the connection scores corresponding to the candidate characters in the candidate text and the prediction probability.

12. A text error correction device, characterized in that: including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the text error correction method as described in any one of claims 1 to 8.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the text error correction method as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Text error correction method and device

    CN111126045A