A method and device for correcting text
The method employs universal and domain-specific language models with a BK tree approach to efficiently correct text errors across multiple users in SaaS environments, overcoming the inefficiencies of user-specific training in existing query correction methods.
Patent Information
- Application Number
- CN202210351633.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-02
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-04-02
AI Technical Summary
In the prior art, the error correction process for corpus training language models for a single user is complex and inefficient, and cannot meet the different needs of multiple tenants in SaaS scenarios.
Provide a text error correction method and device, by obtaining query statements, performing general error correction and/or domain error correction, using common language model and domain language model to detect and correct text errors, and constructing a BK tree for domain error correction, to adapt to user needs in different fields.
It realizes efficient error correction of each user's query statement in the SaaS scenario, simplifies the error correction process, improves error correction efficiency and pertinence, and meets the query and error correction needs of users in different fields.
Smart Images

Figure CN114860870B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method and device for text error correction. Background Art
[0002] Software as a Service (SaaS) means that cloud application software is developed and maintained by a cloud provider, automatic software updates are provided, and the software is provided to users via the Internet on a pay-as-you-go basis. In the SaaS scenario, different tenants have different requirements for configurations and functions. That is to say, a system needs to meet the different needs of multiple tenants, and users under each tenant share a set of requirement configurations. Existing query error correction solutions mainly train a language model based on the user's corpus, which has high requirements for the user's own corpus. The user needs to provide a large amount of document data, and then error detection and correction are performed according to the language model, and finally the error correction result is obtained. For overlapping error correction scenarios between different users, error correction will be performed redundantly and cannot be deployed in a SaaS manner. The query error correction method in the prior art has the technical problems that the error correction process for training a language model based on the corpus of a single user is complex and inefficient. Summary of the Invention
[0003] The purpose of the embodiments of this application is to provide a method and device for text error correction, which solves the technical problems in the prior art that the error correction process for training a language model based on the corpus of a single user is complex and inefficient. The specific technical solutions are as follows:
[0004] In the first aspect of the implementation of this application, first, a method for text error correction is provided. The method includes: obtaining a query statement; wherein, the string in the query statement is used to represent the text to be error-corrected; performing general error correction and / or domain error correction on the text to be error-corrected carried by the query statement; wherein, general error correction refers to correcting general expression errors associated with the text, and domain error correction refers to correcting the mismatch between the text and keywords in the associated vertical domain; outputting the error correction result of the text to be error-corrected.
[0005] In the second aspect of the implementation of this application, a device for text error correction is further provided. The device includes: a first acquisition module, configured to obtain a query statement; wherein, the string in the query statement is used to represent the text to be error-corrected; an error correction module, configured to perform general error correction and / or domain error correction on the text to be error-corrected carried by the query statement; wherein, general error correction refers to correcting general expression errors associated with the text, and domain error correction refers to correcting the mismatch between the text and keywords in the associated vertical domain; an output module, configured to output the error correction result of the text to be error-corrected.
[0006] In a third aspect of the implementation of this application, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used to store a computer program; the processor is used to implement the method steps described in the first aspect when executing the program stored on the memory.
[0007] In a fourth aspect of the implementation of this application, a computer-readable storage medium is further provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, the computer is made to execute the text error correction method described in the first aspect above.
[0008] This application can be applied to the field of information retrieval technology for correcting query statements. The text error correction method and device provided in the embodiments of this application obtain a query statement; among them, the string in the query statement is used to represent the text to be corrected; perform general error correction and / or domain error correction on the text to be corrected carried by the query statement; where general error correction refers to correcting general expression errors associated with the text, and domain error correction refers to correcting the mismatch between the text and keywords in the associated vertical domain; output the error correction result of the text to be corrected; that is, perform general error correction and / or domain error correction on each user's query statement according to the configuration of the SaaS tenant, thereby solving the technical problems of complex and low-efficiency error correction process for training language models with single-user corpora in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.
[0010] Figure 1 It is one of the flowcharts of the text error correction method in the embodiments of this application;
[0011] Figure 2 It is another flowchart of the text error correction method in the embodiments of this application;
[0012] Figure 3 It is yet another flowchart of the text error correction method in the embodiments of this application;
[0013] Figure 4 It is still another flowchart of the text error correction method in the embodiments of this application;
[0014] Figure 5 It is yet another flowchart of the text error correction method in the embodiments of this application;
[0015] Figure 6 It is a schematic diagram of the BK tree of the text error correction method in the embodiments of this application;
[0016] Figure 7 This is the sixth flowchart of the text error correction method in the embodiments of the present application;
[0017] Figure 8 This is the seventh flowchart of the text error correction method in the embodiments of the present application;
[0018] Figure 9 This is the eighth flowchart of the text error correction method in the embodiments of the present application;
[0019] Figure 10 This is the ninth flowchart of the text error correction method in the embodiments of the present application;
[0020] Figure 11 This is the tenth flowchart of the text error correction method in the embodiments of the present application;
[0021] Figure 12 This is a schematic diagram of a demonstration BK tree for the text error correction method in the embodiments of the present application;
[0022] Figure 13 This is a schematic diagram of the structure of the text error correction device in the embodiments of the present application;
[0023] Figure 14 This is a schematic diagram of the structure of an electronic device in the embodiments of the present application. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0025] In subsequent descriptions, suffixes such as "module" and "unit" used to represent components are only for the convenience of description of the present application, and they have no specific meaning themselves. Therefore, "module" and "component" can be used interchangeably.
[0026] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The embodiments of the present application provide a text error correction method, as Figure 1 shown, including the following steps:
[0027] Step 102: Obtain a query statement; wherein, the string in the query statement is used to represent the text to be corrected;
[0028] It should be noted that, after obtaining the query statement, preprocess the query statement, such asFigure 2 As shown in the figure, it includes standardizing the pinyin in the query statement, unifying the full-width and half-width of the query statement, and removing punctuation marks from the query statement; the text to be corrected carried by the query statement after preprocessing is divided into three types: pinyin, Chinese characters, and combinations of pinyin and Chinese characters.
[0029] Step 104: Perform general error correction and / or domain error correction on the text to be corrected carried by the query statement; among them, general error correction refers to correcting general expression errors associated with the text, and domain error correction refers to correcting the mismatch between the text and keywords in the associated vertical domain.
[0030] It should be noted that performing general error correction and / or domain error correction on the text to be corrected can be only performing general error correction, only performing domain error correction, first performing general error correction and then performing domain error correction, or first performing domain error correction and then performing general error correction.
[0031] Step 106: Output the error correction result of the text to be corrected.
[0032] Through the above steps 102 to 106 of the embodiments of the present application, a query statement is obtained; among them, the string in the query statement is used to represent the text to be corrected; general error correction and / or domain error correction are performed on the text to be corrected carried by the query statement; among them, general error correction refers to correcting general expression errors associated with the text, and domain error correction refers to correcting the mismatch between the text and keywords in the associated vertical domain; the error correction result of the text to be corrected is output; that is to say, according to the configuration of the SaaS tenant, general error correction and / or domain error correction are performed on the query statements of each user, thereby solving the technical problems of complex and low-efficiency error correction in the prior art for training a language model with the corpus of a single user.
[0033] In an alternative embodiment of the embodiments of the present application, when the text to be corrected is Chinese characters, the general error correction involved in step 104 of the present application for the text to be corrected carried by the query statement is as follows Figure 3 As shown in the figure, it includes:
[0034] Step 302: Perform error detection on the string in the query statement.
[0035] Among them, it should be noted that when the text to be corrected is in Chinese characters, the error detection of the string in the query statement includes: obtaining the string fragment of the query statement in a sliding window manner, scoring the string fragment through a general language model trained with general corpus, and scoring the entire string in the query statement through the general language model. When the difference between the score of the string fragment and the score of the entire string is greater than a certain threshold, it is determined that the text corresponding to the string fragment has an error.
[0036] Step 304: Recall the corresponding candidate set from the preset confusion set for the text corresponding to the detected incorrect string; among them, the confusion set includes multiple texts and the candidate sets respectively corresponding to the multiple texts;
[0037] Among them, it should be noted that in the confusion set, all texts have corresponding candidate sets that can be used for replacement.
[0038] Step 306: Substitute the candidate texts in the candidate set into the text to be corrected in turn to obtain the candidate text to be corrected;
[0039] Step 308: Score the candidate text to be corrected based on the general language model; among them, the general language model is a language model trained based on general corpus and used to score general corpus;
[0040] Step 310: Sort the scored candidate text to be corrected and select the candidate text to be corrected with the highest score from the sorting result;
[0041] Step 312: When the difference between the score of the candidate text to be corrected with the highest score and the score of the text to be corrected is greater than the first preset threshold, determine the candidate text to be corrected with the highest score as the correction result.
[0042] Among them, it should be noted that when there is no candidate text to be corrected whose score difference from the score of the text to be corrected is greater than the first preset threshold, the text to be corrected is not corrected; in a demonstration example, the score of the candidate text to be corrected "laptop" is 95 points, the score of the text to be corrected is 80 points, the first preset threshold is 10 points, and the difference between the score of the candidate text to be corrected "laptop" and the score of the text to be corrected is 15 points. Since this difference is greater than the first preset threshold, the candidate text to be corrected "laptop" can be used as the correction result of the text to be corrected at this time.
[0043] It can be seen that the text correction method provided by the embodiments of the present application can perform error detection on the text to be corrected in Chinese characters, determine the candidate text to be corrected based on the general language model, and judge whether to correct the text to be corrected in Chinese characters.
[0044] In an alternative implementation of the embodiment of the present application, when the text to be corrected is pinyin, or the text to be corrected is a combination of pinyin and Chinese characters, the general correction of the text to be corrected carried in the query statement in step 104 of the present application, as Figure 4 shown, includes:
[0045] Step 402: When there are Chinese characters in the text to be corrected, convert the Chinese characters in the text to be corrected into corresponding pinyin;
[0046] It should be noted that when there are Chinese characters in the text to be corrected, that is, the text to be corrected is a combination of pinyin and Chinese characters. After converting the Chinese characters in the text to be corrected into corresponding pinyin and splicing the original pinyin, the text to be corrected changes from a combination of pinyin and Chinese characters to pinyin.
[0047] Step 404: Perform word segmentation on the pinyin corresponding to the text to be corrected to obtain a first word segmentation result;
[0048] It should be noted that the word segmentation process can adopt the forward maximum matching method and perform continuous long pinyin string word segmentation according to the pinyin dictionary; among them, the pinyin dictionary is obtained by exhaustively listing all pinyin combinations.
[0049] Step 406: Convert the pinyin word segmentation sequence in the first word segmentation result into a corresponding Chinese character sequence;
[0050] It should be noted that the conversion of the pinyin word segmentation sequence into a corresponding Chinese character sequence can be performed through a hidden Markov model trained by general corpus; among them, the hidden Markov model is a kind of general language model, which contains the conversion probability corresponding to the pinyin sequence and the Chinese character sequence.
[0051] Step 408: Based on the general language model, score the Chinese characters in the Chinese character sequence; among them, the general language model is a language model trained based on general corpus and used to score general corpus;
[0052] Step 410: Sort according to the scores of the Chinese character sequences to obtain a sorting result;
[0053] Step 412: Select the Chinese character sequence with the highest score from the Chinese character sequences with scores greater than the second preset threshold in the sorting result as the correction result of the text to be corrected.
[0054] Among them, it should be noted that in a demonstration example, the highest-scoring Chinese character sequence "knowledge graph" has a score of 90 points, and the second preset threshold is 80 points. The score of the highest-scoring Chinese character sequence "knowledge graph" is greater than the second preset threshold. At this time, the Chinese character sequence "knowledge graph" can be used as the error correction result of the text to be error-corrected. In another demonstration example, the highest-scoring Chinese character sequence "frosted glass" has a score of 70 points, and the second preset threshold is 75 points. The score of the highest-scoring Chinese character sequence "frosted glass" is less than the second preset threshold. At this time, the highest-scoring Chinese character sequence "frosted glass" will not be used as the error correction result of the text to be error-corrected, and the text to be error-corrected will not be error-corrected.
[0055] It can be seen that the text error correction method provided by the embodiments of the present application can process the text to be error-corrected of pure pinyin or a combination of pinyin and Chinese characters, and determine whether a Chinese character sequence can be used as the error correction result of the text to be error-corrected based on a general language model.
[0056] Before the domain error correction of the text to be error-corrected carried by the query statement involved in step 104 of the text error correction method provided by the embodiments of the present application, as Figure 5 shown, it includes:
[0057] Step 502: Obtain a domain dictionary, and store the Chinese characters corresponding to the domain words in the domain dictionary and the pinyin corresponding to the Chinese characters as key-value pairs in the target database; where the domain words are proper terms in multiple different domains, and the domain dictionary also includes the weights corresponding to the domain words.
[0058] Among them, it should be noted that the target database can be a Remote Dictionary Server (Redis for short) database; the domain dictionary is uploaded by the tenant in the SaaS scenario, and software users under the same tenant can use the same domain error correction configuration. For example, a tenant facing the medical field uploads a medical term domain dictionary, and the query statements input by the users under this tenant need to be error-corrected for the domain words in the medical field when querying and error-correcting.
[0059] Step 504: Construct a BK tree based on the pinyin corresponding to the Chinese characters in the domain words in the target database; where the BK tree is a data structure including the domain words as the root node and multiple child nodes, and the edit distance between the root node and the child nodes is used to represent how many times the pinyin corresponding to the root node needs to be processed to obtain the pinyin corresponding to the child node, and the edit distance between two child nodes is used to represent how many times the pinyin corresponding to the child node closer to the root node needs to be processed to obtain the pinyin corresponding to the child node farther from the root node.
[0060] Among them, it should be noted that the Chinese characters and their corresponding pinyin in the domain words of the target database are in a one-to-one correspondence relationship. Therefore, the Chinese characters in the domain words can be converted into the corresponding pinyin based on the target database; the BK (Burkhard Keller) tree uses the pinyin corresponding to a randomly selected Chinese character in the domain words as the root node; in a demonstration example, the root node has two child nodes, namely the first child node and the second child node, and the second child node has one child node, namely the third child node; among them, the pinyin corresponding to the root node is "doufu", the pinyin corresponding to the first child node is "dounai", the pinyin corresponding to the second child node is "shuofu", and the pinyin corresponding to the third child node is "shuoqi"; it takes 3 processes to process "doufu" into "dounai", so the edit distance between the root node and the first child node is 3; it takes 4 processes to process "doufu" into "shuofu", so the edit distance between the root node and the second child node is 4; it takes 2 processes to process "shuofu" into "shuoqi", so the edit distance between the second child node and the third child node is 2. Based on this, the schematic diagram of the constructed BK tree is as Figure 6 shown.
[0061] It can be seen that the text error correction method provided by the embodiment of the present application can construct a BK tree based on the Chinese characters and pinyin corresponding to the domain words in the domain dictionary uploaded by the tenant, and perform domain error correction on the query statements for the tenants in each domain.
[0062] In an optional implementation manner of the embodiment of the present application, when the text to be error-corrected is in Chinese characters, the domain error correction of the text to be error-corrected carried in the query statement involved in step 104 of the present application, as Figure 7 shown, includes:
[0063] Step 702: Convert the string used to represent Chinese characters into a string used to represent pinyin;
[0064] Step 704: Traverse the string used to represent pinyin through a sliding window to obtain the corresponding pinyin;
[0065] Step 706: Query the first candidate pinyin of the pinyin obtained through the sliding window based on the BK tree;
[0066] Step 708: Query the Chinese character corresponding to the first candidate pinyin from the target database to obtain the first candidate Chinese character;
[0067] Step 710: Replace the string used to represent the Chinese character to be error-corrected in the query statement with the string used to represent the first candidate Chinese character to obtain a candidate query statement;
[0068] Step 712: Score the candidate query statements based on the domain language model; wherein, the domain language model is a language model trained based on domain corpus and used to score the domain corpus.
[0069] Step 714: Sort the candidate query statements according to the scores of the candidate query statements.
[0070] Step 716: In the case that the difference between the score of the candidate query statement with the highest score and the score of the query statement is greater than the third preset threshold, determine the Chinese characters in the candidate query statement with the highest score as the error correction result.
[0071] It should be noted that, in an exemplary example, the score of the candidate query statement "solution" with the highest score is 85 points, the score of the query statement is 70 points, the third preset threshold is 10 points, and the difference between the score of the candidate query statement "solution" and the score of the query statement is 15 points. Since this difference is greater than the third preset threshold, the candidate query statement "solution" can be used as the error correction result of the query statement at this time.
[0072] It can be seen that the text error correction method provided by the embodiments of the present application is directed to a query statement whose text to be error corrected is Chinese characters. It can convert the text to pinyin and then use the domain language model to perform domain error correction on the query statement, improving the pertinence of query error correction and meeting the query error correction needs of users in different fields in the SaaS scenario.
[0073] In an alternative implementation manner of the embodiments of the present application, in the case that the text to be error corrected is pinyin, the domain error correction of the text to be error corrected carried by the query statement involved in step 104 of the present application, as Figure 8 shown, includes:
[0074] Step 802: Query the second candidate pinyin of the pinyin based on the BK tree.
[0075] Step 804: Query the Chinese characters corresponding to the candidate pinyin from the target database to obtain the second candidate Chinese characters.
[0076] Step 806: Query the weight corresponding to the second candidate Chinese character based on the domain dictionary, and determine the candidate Chinese character with the highest weight as the error correction result.
[0077] It can be seen that the text error correction method provided by the embodiments of the present application is directed to a query statement whose text to be error corrected is pinyin. It can perform domain error correction on the text to be error corrected according to the weights in the domain dictionary, improving the pertinence of query error correction and meeting the query error correction needs of users in different fields in the SaaS scenario.
[0078] In an alternative implementation of the embodiments of the present application, when the text to be corrected is a combination of pinyin and Chinese characters, the domain correction of the text to be corrected carried in the query statement involved in step 104 of the present application is as follows Figure 9 shown, including:
[0079] Step 902: Convert the Chinese characters in the text to be corrected into pinyin;
[0080] Step 904: Query the third candidate pinyin of the pinyin based on the BK tree;
[0081] Step 906: Query the Chinese characters corresponding to the candidate pinyin from the target database to obtain the third candidate Chinese characters;
[0082] Step 908: Query the weight corresponding to the third candidate Chinese character based on the domain dictionary, and determine the correction result as the candidate Chinese character with the highest weight;
[0083] Step 910: When the correction result indicates that correction is required, return the correction result;
[0084] Step 912: When the correction result indicates that correction is not required, convert the third candidate pinyin and the pinyin in the text to be corrected into Chinese characters;
[0085] Step 914: Convert the string representing the converted Chinese characters into a string representing pinyin;
[0086] Step 916: Traverse the string representing pinyin through a sliding window to obtain the corresponding pinyin;
[0087] Step 918: Query the fourth candidate pinyin of the pinyin obtained through the sliding window based on the BK tree;
[0088] Step 920: Query the Chinese characters corresponding to the candidate pinyin from the target database to obtain the fourth candidate Chinese characters;
[0089] Step 922: Replace the string representing the Chinese characters to be corrected in the query statement with the string representing the fourth candidate Chinese characters to obtain a candidate query statement;
[0090] Step 924: Score the candidate query statement based on the domain language model; wherein, the domain language model is a language model trained based on domain corpus and used to score domain corpus;
[0091] Step 926: Sort the candidate query statements according to the scores of the candidate query statements;
[0092] Step 928: When the difference between the score of the candidate query statement with the highest score and the score of the query statement is greater than the fourth preset threshold, determine the Chinese characters in the candidate query statement with the highest score as the candidate error correction result;
[0093] Step 930: When the candidate error correction result includes the pinyin in the text to be error corrected, determine the candidate error correction result as the error correction result;
[0094] Step 932: When the candidate error correction result does not include the pinyin in the text to be error corrected, prohibit error correction of the text to be error corrected.
[0095] It can be seen that the error correction method for text provided by the embodiments of the present application can convert a query statement with a combination of pinyin and Chinese characters into pinyin and then use a domain language model to perform domain error correction on the query statement, which improves the pertinence of query error correction and meets the query error correction needs of users in different domains in the SaaS scenario.
[0096] Combining the above steps 702 to 716, steps 802 to 816, and steps 902 to 932, the process of performing domain error correction on a query statement is as follows: Figure 10 As shown, when the text to be error corrected corresponding to the query statement is Chinese characters, convert the text to be error corrected into pinyin, divide the text to be error corrected into multiple parts using a sliding window, query the candidate pinyin of the pinyin based on the BK tree, and thus perform domain error correction, which is simply referred to as domain partial error correction; when the text to be error corrected corresponding to the query statement is pinyin, query the candidate pinyin of the pinyin based on the BK tree using the entire pinyin corresponding to the text to be error corrected, and thus perform domain error correction, which is simply referred to as overall error correction; when the text to be error corrected corresponding to the query statement is a combination of pinyin and Chinese characters, first convert the Chinese characters into pinyin, query the candidate pinyin of the pinyin based on the BK tree using the entire pinyin corresponding to the text to be error corrected, which is simply referred to as overall error correction. When the error correction result indicates that error correction is required, directly return the error correction result. When the error correction result indicates that error correction is not required, divide the text to be error corrected into multiple parts using a sliding window and query the candidate pinyin of the pinyin based on the BK tree to perform domain error correction, which is simply referred to as domain partial error correction, and error correction is only performed when the candidate error correction result includes the pinyin in the text to be error corrected, and the text to be error corrected is not error corrected when the candidate error correction result does not include the pinyin in the text to be error corrected.
[0097] In an alternative implementation manner of the embodiments of the present application, for the error correction method for text provided by the embodiments of the present application, the query of the candidate pinyin of the pinyin based on the BK tree involved in steps 706, 802, 904, and 918 is as follows: Figure 11 As shown, it includes:
[0098] Step 1102: Determine the edit distance range based on the edit distance between the pinyin and the root node in the BK tree and the fifth preset threshold;
[0099] It should be noted that the fifth preset threshold is preset according to the requirements of the SaaS tenant. For example, when the length of the pinyin string is greater than 10, the fifth preset threshold is set to 1, and when the length of the pinyin string is less than or equal to 10, the fifth preset threshold is set to 0; in a demonstration example, the edit distance D between the pinyin and the root node in the BK tree is 3, and the fifth preset threshold N is 1, then the edit distance range is [D - N, D + N], that is, [2, 4].
[0100] Step 1104: Determine candidate child nodes from the BK tree based on the edit distance range; among them, the pinyin corresponding to the candidate child nodes has an edit distance less than or equal to the fifth preset threshold;
[0101] It should be noted that to determine candidate child nodes from the BK tree based on the edit distance range, first, it is necessary to start from the root node to determine the child nodes with an edit distance within the edit distance range, and then find the child node with the smallest edit distance from the above-mentioned child nodes to the pinyin. When there is a child node with an edit distance less than the fifth preset threshold to the pinyin, this child node is determined as the candidate child node, and continue to search in the child nodes of this child node to see if there is a child node with an edit distance less than the fifth preset threshold to the pinyin. If there is, it is determined as the candidate child node; in a demonstration example, such as Figure 12As shown, when the fifth preset threshold N is 2, candidate pinyins for the pinyin "beifang" are queried based on the BK tree. The root node has a total of three child nodes, namely the first child node, the second child node, and the third child node. The first child node has a total of one child node, the fourth child node. The second child node has a total of two child nodes, the fifth child node and the sixth child node. The third child node has a total of one child node, the seventh child node. The pinyin corresponding to the root node is "dongfang", the pinyin corresponding to the first child node is "dongan", the pinyin corresponding to the second child node is "dongnan", the pinyin corresponding to the third child node is "xifang", the pinyin corresponding to the fourth child node is "xiuxi", the pinyin corresponding to the fifth child node is "chonglai", the pinyin corresponding to the sixth child node is "kunnan", and the pinyin corresponding to the seventh child node is "xibeifang". The edit distance D between the root node "dongfang" and the pinyin "beifang" is 4, N is 2, and the edit distance range is [D - N, D + N], that is, [2, 6]. The edit distance between the root node and the first child node is 2, the edit distance between the root node and the second child node is 2, the edit distance between the root node and the third child node is 4. The edit distance X between the pinyin corresponding to the third child node and the pinyin "beifang" is 2, X ≤ N. Therefore, the third child node is determined as the candidate child node. The edit distance X between the pinyin corresponding to the seventh child node and the pinyin "beifang" is 2, X ≤ N. Therefore, the seventh child node is determined as the candidate child node.
[0102] Step 1106: Determine the pinyin corresponding to the candidate child node as the candidate pinyin.
[0103] Among them, it should be noted that, as Figure 12 shown in the demonstration example, if the third child node and the seventh child node are determined as the candidate child nodes, then the pinyin "xifang" corresponding to the third child node and the pinyin "xibeifang" corresponding to the seventh child node are determined as the candidate pinyins.
[0104] It can be seen that the text error correction method provided by the embodiments of the present application can perform general error correction and / or domain error correction on the query statements of each user according to the configuration of the SaaS tenant, thereby solving the technical problems of complex and low - efficiency error correction in the prior art for training a language model with the corpus of a single user.
[0105] The embodiments of the present application provide a text error correction device. As Figure 13 shown, the device includes:
[0106] The first acquisition module 1302 is used to acquire a query statement; wherein, the string in the query statement is used to represent the text to be error - corrected.
[0107] An error correction module 1304 is used to perform general error correction and / or domain error correction on the text to be error-corrected carried in the query statement. Among them, general error correction refers to correcting the general expression errors associated with the text, and domain error correction refers to correcting the mismatch between the text and the keywords in the associated vertical domain.
[0108] An output module 1306 is used to output the error correction result of the text to be error-corrected.
[0109] Through the text error correction device provided by the embodiments of the present application, the query statement is obtained through the first acquisition module. Among them, the string in the query statement is used to represent the text to be error-corrected. The error correction module performs general error correction and / or domain error correction on the text to be error-corrected carried in the query statement. Among them, general error correction refers to correcting the general expression errors associated with the text, and domain error correction refers to correcting the mismatch between the text and the keywords in the associated vertical domain. The error correction result of the text to be error-corrected is output through the output module. That is to say, according to the configuration of the SaaS tenant, general error correction and / or domain error correction are performed on the query statements of each user, thus solving the technical problems of complex and low-efficiency error correction in the prior art for training a language model with the corpus of a single user.
[0110] In an alternative implementation manner of the embodiments of the present application, the error correction module 1304 provided by the embodiments of the present application may further include:
[0111] A detection unit is used to detect errors in the string in the query statement.
[0112] A recall unit is used to recall the corresponding candidate set from the preset confusion set for the text corresponding to the string detected with errors. Among them, the confusion set includes multiple texts and the candidate sets respectively corresponding to the multiple texts.
[0113] A first processing unit is used to sequentially substitute the candidate texts in the candidate set into the text to be error-corrected to obtain the candidate text to be error-corrected.
[0114] A first scoring unit is used to score the candidate text to be error-corrected based on a general language model. Among them, the general language model is a language model trained based on general corpus and used to score the general corpus.
[0115] A second processing unit is used to sort the scored candidate text to be error-corrected and select the candidate text to be error-corrected with the highest score from the sorting result.
[0116] A first determination unit is used to determine the candidate text to be error-corrected with the highest score as the error correction result when the difference between the score of the candidate text to be error-corrected with the highest score and the score of the text to be error-corrected is greater than the first preset threshold.
[0117] In an alternative implementation manner of the embodiment of the present application, when the text to be corrected is pinyin, or the text to be corrected is a combination of pinyin and Chinese characters, the error correction module 1304 provided by the embodiment of the present application may further include:
[0118] A first conversion unit, configured to convert Chinese characters in the text to be corrected into corresponding pinyin when there are Chinese characters in the text to be corrected;
[0119] A word segmentation unit, configured to perform word segmentation processing on the pinyin corresponding to the text to be corrected to obtain a first word segmentation result;
[0120] A second conversion unit, configured to convert the pinyin word segmentation sequence in the first word segmentation result into a corresponding Chinese character sequence;
[0121] A second scoring unit, configured to score Chinese characters in the Chinese character sequence based on a general language model; wherein, the general language model is a language model trained based on general corpus and used to score the general corpus;
[0122] A third processing unit, configured to sort according to the scores of the Chinese character sequence to obtain a sorting result;
[0123] A selection unit, configured to select the Chinese character sequence with the highest score from the Chinese character sequences with scores greater than a second preset threshold in the sorting result as the error correction result of the text to be corrected.
[0124] In an alternative implementation manner of the embodiment of the present application, the text error correction device provided by the embodiment of the present application may further include:
[0125] A second acquisition module, configured to acquire a domain dictionary and store the Chinese characters corresponding to the domain words in the domain dictionary and the corresponding pinyin as key-value pairs in a target database; wherein, the domain words are proper terms in multiple different domains, and the domain dictionary further includes the weights corresponding to the domain words;
[0126] A construction module, configured to construct a BK tree based on the pinyin corresponding to the Chinese characters in the domain words in the target database; wherein, the BK tree is a data structure including a root node and multiple child nodes with domain words, and the edit distance between the root node and the child nodes is used to represent how many times the pinyin corresponding to the root node needs to be processed to obtain the pinyin corresponding to the child node, and the edit distance between two child nodes is used to represent how many times the pinyin corresponding to the child node closer to the root node needs to be processed to obtain the pinyin corresponding to the child node farther from the root node.
[0127] In an alternative implementation manner of the embodiment of the present application, when the text to be corrected is Chinese characters, the error correction module 1304 provided by the embodiment of the present application may further include:
[0128] A third conversion unit, configured to convert a string for representing Chinese characters into a string for representing pinyin;
[0129] A first obtaining unit, configured to traverse the string for representing pinyin through a sliding window to obtain the corresponding pinyin;
[0130] A first query unit, configured to query a first candidate pinyin of the pinyin obtained through the sliding window based on a BK tree;
[0131] A fourth processing unit, configured to query Chinese characters corresponding to the first candidate pinyin from a target database to obtain first candidate Chinese characters;
[0132] A fifth processing unit, configured to replace the string for representing the Chinese character to be corrected in the query statement with the string for representing the first candidate Chinese character to obtain a candidate query statement;
[0133] A third scoring unit, configured to score the candidate query statement based on a domain language model; wherein, the domain language model is a language model trained based on domain corpus and used for scoring the domain corpus;
[0134] A sixth processing unit, configured to sort the candidate query statements according to the scores of the candidate query statements;
[0135] A second determination unit, configured to determine the Chinese character in the candidate query statement with the highest score as the error correction result when the difference between the score of the candidate query statement with the highest score and the score of the query statement is greater than a third preset threshold.
[0136] In an alternative implementation manner of the embodiment of the present application, when the text to be corrected is pinyin, the error correction module 1304 provided by the embodiment of the present application may further include:
[0137] A second query unit, configured to query a second candidate pinyin of the pinyin based on a BK tree;
[0138] A seventh processing unit, configured to query Chinese characters corresponding to the candidate pinyin from a target database to obtain second candidate Chinese characters;
[0139] A third determination unit, configured to query the weight corresponding to the second candidate Chinese character based on a domain dictionary, and determine the candidate Chinese character with the highest weight as the error correction result.
[0140] In an alternative implementation manner of the embodiment of the present application, when the text to be corrected is a combination of pinyin and Chinese characters, the error correction module 1304 provided by the embodiment of the present application may further include:
[0141] A fourth conversion unit, configured to convert the Chinese characters in the text to be corrected into pinyin;
[0142] A third query unit, configured to query a third candidate pinyin of the pinyin based on a BK tree;
[0143] An eighth processing unit, configured to query Chinese characters corresponding to the candidate pinyin from a target database to obtain third candidate Chinese characters;
[0144] A fourth determination unit, configured to query weights corresponding to the third candidate Chinese characters based on a domain dictionary, and determine an error correction result as the candidate Chinese character with the highest weight;
[0145] A return unit, configured to return an error correction result when the error correction result indicates that error correction is required;
[0146] A fifth conversion unit, configured to convert the third candidate pinyin and the pinyin in the text to be error-corrected into Chinese characters when the error correction result indicates that error correction is not required;
[0147] A sixth conversion unit, configured to convert a string representing the converted Chinese characters into a string representing pinyin;
[0148] A second acquisition unit, configured to traverse the string representing pinyin through a sliding window to obtain corresponding pinyin;
[0149] A fourth query unit, configured to query a fourth candidate pinyin of the pinyin obtained through the sliding window based on a BK tree;
[0150] A ninth processing unit, configured to query Chinese characters corresponding to the candidate pinyin from a target database to obtain fourth candidate Chinese characters;
[0151] A tenth processing unit, configured to replace the string representing the Chinese character to be error-corrected in the query statement with the string representing the fourth candidate Chinese character to obtain a candidate query statement;
[0152] A fourth scoring unit, configured to score the candidate query statement based on a domain language model; wherein, the domain language model is a language model trained based on domain corpus and used to score domain corpus;
[0153] An eleventh processing unit, configured to sort the candidate query statements according to the scores of the candidate query statements;
[0154] A fifth determination unit, configured to determine the Chinese characters in the candidate query statement with the highest score as candidate error correction results when the difference between the score of the candidate query statement with the highest score and the score of the query statement is greater than a fourth preset threshold;
[0155] A sixth determination unit, configured to determine the candidate error correction result as the error correction result when the candidate error correction result includes the pinyin in the text to be error-corrected;
[0156] A twelfth processing unit, configured to prohibit correcting the text to be corrected when the pinyin in the text to be corrected is not included in the candidate error correction results.
[0157] In an alternative implementation manner of the embodiment of the present application, for the text error correction device provided in the embodiment of the present application, the first query unit, the second query unit, the third query unit, and the fourth query unit respectively include:
[0158] A first determination subunit, configured to determine an edit distance range based on the edit distance between the pinyin and the root node in the BK tree and a fifth preset threshold;
[0159] A second determination subunit, configured to determine candidate child nodes from the BK tree based on the edit distance range; wherein, the pinyin corresponding to the candidate child nodes has a pinyin edit distance less than or equal to the fifth preset threshold;
[0160] A third determination subunit, configured to determine the pinyin corresponding to the candidate child nodes as candidate pinyins.
[0161] The embodiment of the present application further provides an electronic device, as Figure 14 shown, including a processor 1401, a communication interface 1402, a memory 1403, and a communication bus 1404. Among them, the processor 1401, the communication interface 1402, and the memory 1403 complete communication with each other through the communication bus 1404.
[0162] The memory 1403 is used to store a computer program;
[0163] When the processor 1401 is configured to execute the program stored on the memory 1403, it implements the Figure 1 method steps in, and the function it plays is the same as the Figure 1 method steps in, which will not be elaborated here.
[0164] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 14 only a thick line is used to represent it in, but it does not mean that there is only one bus or one type of bus.
[0165] The communication interface is used for communication between the above terminal and other devices.
[0166] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0167] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0168] In another embodiment provided by the present application, a computer-readable storage medium is further provided. Instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the text error correction method described in any one of the above embodiments.
[0169] In another embodiment provided by the present application, a computer program product containing instructions is further provided. When it runs on a computer, it causes the computer to execute the text error correction method described in any one of the above embodiments.
[0170] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0171] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0172] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0173] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A method for correcting errors in text, characterized in that, Including: Obtain a query statement; wherein, the string in the query statement is used to represent the text to be corrected. Perform only domain error correction on the text to be corrected carried by the query statement, or perform general error correction first and then domain error correction, or perform domain error correction first and then general error correction; wherein, the general error correction refers to correcting general expression errors associated with the text, and the domain error correction refers to correcting the mismatch between the text and keywords in the associated vertical domain. Output the error correction result of the text to be corrected. Wherein: When the text to be corrected is in Chinese characters, the general error correction of the text to be corrected carried by the query statement includes: detecting errors in the string in the query statement; recalling the corresponding candidate set from the preset confusion set for the text corresponding to the detected error string; wherein, the confusion set includes multiple texts and candidate sets respectively corresponding to the multiple texts; sequentially substituting the candidate texts in the candidate set into the text to be corrected to obtain a candidate text to be corrected; scoring the candidate text to be corrected based on a general language model; wherein, the general language model is a language model trained based on general corpus and used to score the general corpus; sorting the scored candidate texts to be corrected, and selecting the candidate text to be corrected with the highest score from the sorting result; when the difference between the score of the candidate text to be corrected with the highest score and the score of the text to be corrected is greater than a first preset threshold, determining the candidate text to be corrected with the highest score as the error correction result. When the text to be corrected is in pinyin, or the text to be corrected is a combination of pinyin and Chinese characters, the general error correction of the text to be corrected carried by the query statement includes: when there are Chinese characters in the text to be corrected, converting the Chinese characters in the text to be corrected into corresponding pinyin; performing word segmentation on the pinyin corresponding to the text to be corrected to obtain a first word segmentation result; converting the pinyin word segmentation sequence in the first word segmentation result into a corresponding Chinese character sequence; scoring the Chinese characters in the Chinese character sequence based on a general language model; wherein, the general language model is a language model trained based on general corpus and used to score the general corpus; sorting according to the scores of the Chinese character sequence to obtain a sorting result; selecting the Chinese character sequence with the highest score from the Chinese character sequences in the sorting result whose scores are greater than a second preset threshold as the error correction result of the text to be corrected. Before performing domain error correction on the text to be error-corrected carried in the query statement, it includes: obtaining a domain dictionary, and storing the Chinese characters corresponding to the domain words in the domain dictionary and the pinyin corresponding to the Chinese characters as key-value pairs in a target database; wherein, the domain words are proper terms in multiple different domains, and the domain dictionary further includes the weights corresponding to the domain words; constructing a BK tree based on the pinyin corresponding to the Chinese characters in the domain words in the target database; wherein, the BK tree is a data structure including the domain words as the root node and multiple child nodes, and the edit distance between the root node and the child nodes is used to represent how many times the pinyin corresponding to the root node needs to be processed to obtain the pinyin corresponding to the child node, and the edit distance between two child nodes is used to represent how many times the pinyin corresponding to the child node closer to the root node needs to be processed to obtain the pinyin corresponding to the child node farther from the root node; When the text to be error-corrected is in Chinese characters, performing domain error correction on the text to be error-corrected carried in the query statement includes: converting the string representing the Chinese characters into a string representing pinyin; traversing the string representing pinyin through a sliding window to obtain the corresponding pinyin; querying the first candidate pinyin of the pinyin obtained through the sliding window based on the BK tree; querying the Chinese characters corresponding to the first candidate pinyin from the target database to obtain the first candidate Chinese characters; replacing the string representing the Chinese characters to be error-corrected in the query statement with the string representing the first candidate Chinese characters to obtain a candidate query statement; scoring the candidate query statement based on a domain language model; wherein, the domain language model is a language model trained based on domain corpus and used to score the domain corpus; sorting the candidate query statements according to the scores of the candidate query statements; when the difference between the score of the candidate query statement with the highest score and the score of the query statement is greater than a third preset threshold, determining the Chinese characters in the candidate query statement with the highest score as the error correction result; When the text to be error-corrected is in pinyin, performing domain error correction on the text to be error-corrected carried in the query statement includes: querying the second candidate pinyin of the pinyin based on the BK tree; querying the Chinese characters corresponding to the candidate pinyin from the target database to obtain the second candidate Chinese characters; querying the weight corresponding to the second candidate Chinese characters based on the domain dictionary, and determining the candidate Chinese character with the highest weight as the error correction result; When the text to be corrected is a combination of pinyin and Chinese characters, domain correction of the text to be corrected carried in the query statement includes: converting the Chinese characters in the text to be corrected into pinyin; querying the third candidate pinyin of the pinyin based on the BK tree; querying the Chinese characters corresponding to the candidate pinyin from the target database to obtain the third candidate Chinese characters; querying the weight corresponding to the third candidate Chinese characters based on the domain dictionary, and determining the correction result as the candidate Chinese character with the highest weight; when the correction result indicates that correction is required, returning the correction result; when the correction result indicates that correction is not required, converting the third candidate pinyin and the pinyin in the text to be corrected into Chinese characters; converting the string representing the converted Chinese characters into a string representing pinyin; traversing the string representing pinyin through a sliding window to obtain the corresponding pinyin; querying the fourth candidate pinyin of the pinyin obtained through the sliding window based on the BK tree; querying the Chinese characters corresponding to the candidate pinyin from the target database to obtain the fourth candidate Chinese characters; replacing the string representing the Chinese characters to be corrected in the query statement with the string representing the fourth candidate Chinese characters to obtain a candidate query statement; scoring the candidate query statement based on a domain language model; wherein the domain language model is a language model trained based on domain corpus and used to score the domain corpus; sorting the candidate query statements according to the scores of the candidate query statements; when the difference between the score of the candidate query statement with the highest score and the score of the query statement is greater than a fourth preset threshold, determining the Chinese characters in the candidate query statement with the highest score as the candidate correction result; when the candidate correction result includes the pinyin in the text to be corrected, determining the candidate correction result as the correction result; when the candidate correction result does not include the pinyin in the text to be corrected, prohibiting correction of the text to be corrected.
2. The method according to claim 1, characterized in that Querying the candidate pinyin of the pinyin based on the BK tree includes: Determining an edit distance range based on the edit distance between the pinyin and the root node in the BK tree and a fifth preset threshold; Determining candidate child nodes from the BK tree based on the edit distance range; wherein the pinyin corresponding to the candidate child nodes has an edit distance less than or equal to the fifth preset threshold from the pinyin; Determining the pinyin corresponding to the candidate child nodes as the candidate pinyin.
3. An error correction device for a query statement, characterized in that, Including: A first acquisition module, configured to acquire a query statement; wherein the string in the query statement is used to represent the text to be corrected; An error correction module, configured to perform only domain correction on the text to be corrected carried in the query statement, or perform general error correction first and then domain correction, or perform domain correction first and then general error correction; wherein the general error correction refers to correcting the general expression errors associated with the text, and the domain error correction refers to correcting the mismatch between the text and the keywords in the associated vertical domain. An output module, configured to output an error correction result of the text to be error-corrected; Wherein: When the text to be error-corrected is Chinese, the error correction module is further configured to: detect errors in the string in the query statement; recall a corresponding candidate set from a preset confusion set for the text corresponding to the string detected with errors; wherein, the confusion set includes multiple texts and candidate sets respectively corresponding to the multiple texts; sequentially substitute the candidate texts in the candidate set into the text to be error-corrected to obtain a candidate text to be error-corrected; score the candidate text to be error-corrected based on a general language model; wherein, the general language model is a language model trained based on general corpus and used to score the general corpus; sort the scored candidate text to be error-corrected, and select the candidate text to be error-corrected with the highest score from the sorting result; when the difference between the score of the candidate text to be error-corrected with the highest score and the score of the text to be error-corrected is greater than a first preset threshold, determine the candidate text to be error-corrected with the highest score as the error correction result; When the text to be error-corrected is pinyin, or the text to be error-corrected is a combination of pinyin and Chinese characters, the error correction module is further configured to: when there are Chinese characters in the text to be error-corrected, convert the Chinese characters in the text to be error-corrected into corresponding pinyin; perform word segmentation processing on the pinyin corresponding to the text to be error-corrected to obtain a first word segmentation result; convert the pinyin word segmentation sequence in the first word segmentation result into a corresponding Chinese character sequence; score the Chinese characters in the Chinese character sequence based on a general language model; wherein, the general language model is a language model trained based on general corpus and used to score the general corpus; sort according to the scores of the Chinese character sequence to obtain a sorting result; select the Chinese character sequence with the highest score from the Chinese character sequences whose scores in the sorting result are greater than a second preset threshold as the error correction result of the text to be error-corrected; Before performing domain error correction on the text to be error-corrected carried in the query statement, obtain a domain dictionary, and store the Chinese characters corresponding to the domain words in the domain dictionary and the pinyin corresponding to the Chinese characters as key-value pairs in a target database; wherein, the domain words are proper terms in multiple different domains, and the domain dictionary further includes weights corresponding to the domain words; construct a BK tree based on the pinyin corresponding to the Chinese characters in the domain words in the target database; wherein, the BK tree is a data structure including a root node with the domain word and multiple child nodes, and the edit distance between the root node and the child nodes is used to represent how many times the pinyin corresponding to the root node needs to be processed to obtain the pinyin corresponding to the child node, and the edit distance between two child nodes is used to represent how many times the pinyin corresponding to the child node closer to the root node needs to be processed to obtain the pinyin corresponding to the child node farther from the root node; When the text to be corrected is in Chinese characters, the error correction module is further configured to: convert the string representing the Chinese characters into a string representing pinyin; traverse the string representing pinyin through a sliding window to obtain the corresponding pinyin; query the first candidate pinyin of the pinyin obtained through the sliding window based on the BK tree; query the Chinese characters corresponding to the first candidate pinyin from the target database to obtain the first candidate Chinese characters; replace the string representing the Chinese characters to be corrected in the query statement with the string representing the first candidate Chinese characters to obtain a candidate query statement; score the candidate query statement based on the domain language model, where the domain language model is a language model trained based on domain corpus and used to score the domain corpus; sort the candidate query statements according to the scores of the candidate query statements; when the difference between the score of the candidate query statement with the highest score and the score of the query statement is greater than a third preset threshold, determine the Chinese characters in the candidate query statement with the highest score as the error correction result; When the text to be corrected is in pinyin, the error correction module is further configured to: query the second candidate pinyin of the pinyin based on the BK tree; query the Chinese characters corresponding to the candidate pinyin from the target database to obtain the second candidate Chinese characters; query the weight corresponding to the second candidate Chinese characters based on the domain dictionary, and determine the candidate Chinese characters with the highest weight as the error correction result; When the text to be corrected is a combination of pinyin and Chinese characters, the error correction module is further configured to: convert the Chinese characters in the text to be corrected into pinyin; query the third candidate pinyin of the pinyin based on the BK tree; query the Chinese characters corresponding to the candidate pinyin from the target database to obtain the third candidate Chinese characters; query the weight corresponding to the third candidate Chinese characters based on the domain dictionary, and determine the error correction result as the candidate Chinese character with the highest weight; when the error correction result indicates that error correction is required, return the error correction result; when the error correction result indicates that no error correction is required, convert the third candidate pinyin and the pinyin in the text to be corrected into Chinese characters; convert the string representing the converted Chinese characters into a string representing pinyin; traverse the string representing pinyin through a sliding window to obtain the corresponding pinyin; query the fourth candidate pinyin of the pinyin obtained through the sliding window based on the BK tree; query the Chinese characters corresponding to the candidate pinyin from the target database to obtain the fourth candidate Chinese characters; replace the string representing the Chinese characters to be corrected in the query statement with the string representing the fourth candidate Chinese characters to obtain a candidate query statement; score the candidate query statement based on a domain language model, where the domain language model is a language model trained based on domain corpus and used to score the domain corpus; sort the candidate query statements according to the scores of the candidate query statements; when the difference between the score of the candidate query statement with the highest score and the score of the query statement is greater than a fourth preset threshold, determine the Chinese characters in the candidate query statement with the highest score as the candidate error correction result; when the candidate error correction result includes the pinyin in the text to be corrected, determine the candidate error correction result as the error correction result; when the candidate error correction result does not include the pinyin in the text to be corrected, prohibit error correction of the text to be corrected.
4. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is used to store a computer program; The processor is configured to implement the method according to any one of claims 1-2 when executing the program stored on the memory.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-2.
Citation Information
Patent Citations
Speech text error correction method, system and equipment based on vertical field and medium
CN110210029A
Automatic text error correction method and device, electronic equipment and storage medium
CN114154487A