An OCR typo correction method based on Bayes' theorem

Through the OCR typo correction method based on Bayesian theorem, the word frequency distribution model and dynamic programming optimization algorithm are used to solve the problem of high deep learning costs, and low-cost and highly interpretable typo correction is achieved, which is suitable for small businesses and specific fields.

CN119785368BActive Publication Date: 2025-07-29CHENGDU HARIT MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411838005.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-07-29
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

The existing OCR typo correction algorithm relies on deep learning, which leads to high hardware computing power and training and maintenance costs, making it difficult to popularize in small businesses or specific fields, and is not effective in small sample scenarios.

Method used

The OCR typo correction method based on Bayesian theorem is adopted to collect text data in professional fields, form a word frequency distribution model, use Bayesian formulas to correct typos, and use dynamic programming optimization algorithms to reduce training and maintenance costs.

Benefits of technology

It realizes low-cost OCR typo correction, no data annotation required for training, strong algorithm interpretation, reduces hardware requirements, and is suitable for small businesses and specific fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_14
    Figure SMS_14
  • Figure SMS_15
    Figure SMS_15
  • Figure SMS_16
    Figure SMS_16
Patent Text Reader

Abstract

The present invention discloses an OCR typo correction method based on Bayes' theorem, belonging to the technical field of typo correction. An OCR typo correction method based on Bayes' theorem includes the following steps: S1. Collect a large amount of text data in a professional field; form a word frequency distribution model, and perform conversion and training on the frequency of this model; S2. Correct typos through Bayes' formula and a correction algorithm; S3. Optimize the correction algorithm to enhance performance. Compared with mainstream methods, this method has extremely low training and maintenance costs, specifically manifested in: no data annotation is required for training, the algorithm has strong interpretability, it is not deep learning, has extremely low computing power costs, the entire algorithm is just a dynamic programming, does not rely on other additional tools, and is memory-deployable and maintainable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of typo correction, and more specifically, to an OCR typo correction method based on Bayes' theorem. Background Art

[0002] Currently, the mainstream typo correction algorithms mainly rely on deep learning and large models. With the excellent understanding ability of large models, they can achieve strong generalization.

[0003] However, the side effects brought by using large models, such as high hardware computing power costs and training and maintenance costs, make it difficult to popularize this technology in small enterprises or companies.

[0004] On the other hand, due to the high threshold of large models, it is difficult for them to play a role in specific fields (small samples). For example, in the medical field, "TJ" is an abbreviation for "physical health". Since the frequency of this word is very low, large models often regard these rare words as typos. Such problems often require specific tuning by senior large model algorithm engineers, which also greatly increases the usage cost of this technology. Summary of the Invention

[0005] 1. Technical problems to be solved:

[0006] Aiming at the problems existing in the prior art, based on Bayes' theory, the present invention designs and proposes an OCR typo correction technology with low cost in professional fields.

[0007] 2. Technical solutions:

[0008] To solve the above problems, the present invention adopts the following technical solutions.

[0009] An OCR typo correction method based on Bayes' theorem, comprising the following steps:

[0010] S1. Collect a large amount of text data in the professional field; form a word frequency distribution model, and perform conversion and training on the frequency of this model;

[0011] S2. Correct typos through Bayes' formula and correction algorithms;

[0012] S3. Optimize the correction algorithm to enhance performance.

[0013] Further improvement lies in that the specific content of S1 includes:

[0014] S11. Prepare a large amount of text data in the professional field, which may include relevant text inputs (admission and discharge records, outpatient records, test reports) in the medical scenario, and use the text data output by OCR as the basic data;

[0015] S12, training only needs to record the frequency of each word, and the frequency of other words when word a appears;

[0016] S13. The conditional probability of frequency conversion of the word frequency distribution model is:

[0017] ;

[0018] ;

[0019] Among them, P(t i+1 丨t i ) indicates that given t i Time t i+1 The probability of , and conditional probability;

[0020] t∈T represents any character, T is a set of characters; t i With t i+1 Represents two adjacent characters;

[0021] freq(t i+1 丨t i ) indicates that given t i Time t i+1 Frequency, Z is a fixed value of total word frequency, which is any value greater than max[freq(t i+1 丨t i )丨t∈T], therefore, Z=max[freq(t)丨t∈T].

[0022] A further improvement is that S2 specifically includes:

[0023] Bayes' formula is:

[0024] ;

[0025] P(A, B) is the joint probability of A and B;

[0026] P(A) is the prior probability or marginal probability of A, without considering any factors of B;

[0027] P(B|A) is the conditional probability of B given that A has occurred. The value obtained from A is called the posterior probability of B.

[0028] A further improvement is that the correction algorithm of S2 includes:

[0029] S21. Assume that a word is a typo, and set the joint distribution of the characters corresponding to the word to be lower than a probability threshold θ, and set the word to word=[t i ,t i+1 ,....t i+n ], if and only if:

[0030] θ≥P(t i ,t i+1 ,....t i+n ), the threshold θ is considered as an adjustable parameter;

[0031] S22, calculating the joint distribution based on the assumptions in step S21;

[0032] S23. Find and correct typos.

[0033] A further improvement is that the calculation formula for the joint distribution in S22 includes:

[0034] The joint probability of the front part of multiple words constructed according to the Bayesian formula is:

[0035] ;

[0036] The joint probability of the latter part of multiple words constructed according to the Bayesian formula is:

[0037] .

[0038] A further improvement is that the method for correcting typos in S23 includes:

[0039] S231, considering the joint distribution as an evaluation function, setting the conditions for typos, and setting the score value of the evaluation function to ≤θ;

[0040] S232, optimizing the evaluation function by maximizing its output;

[0041] S233. Use dynamic programming to solve.

[0042] A further improvement is that the transfer equation of the dynamic programming in S233 includes:

[0043] ;

[0044] dp[i] represents the maximum joint probability of the first i characters;

[0045] content[i-1] represents the previous character in the sequence;

[0046] t represents the candidate replacement character for the current character char[i];

[0047] P(t丨char[i-1]) represents the probability of t appearing under the condition of the previous character char[i-1].

[0048] A further improvement is that the optimization of the correction algorithm in S3 includes:

[0049] S31. Maximize the accuracy of identifying typos through bidirectional scanning and two-phase detection, with the condition that when z is a typo among a, b, z, d, e, both P(a, b, z) ≤ θ and P(e, d, z) ≤ θ must hold.

[0050] S32. When pruning the search space, delete the candidates with a conditional probability of 0 for sequence[i - 1].

[0051] 3. Beneficial effects:

[0052] Adopting the technical solution provided by the present invention, compared with the prior art, it has the following beneficial effects:

[0053] The present invention is reasonably designed. Compared with the mainstream methods, this method has extremely low training and maintenance costs, which are specifically manifested in:

[0054] (1) Training does not require data annotation, and the algorithm has strong interpretability;

[0055] (2) It is not deep learning, with extremely low computing power costs;

[0056] (3) The entire algorithm is just one dynamic programming, without relying on other additional tools for memory deployment and maintenance.

[0057] It should be noted that since the structures not introduced in the present invention do not involve the design key points and improvement directions of the present invention, they are the same as the prior art or can be implemented using the prior art, and will not be elaborated here. Detailed implementation manners

[0058] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0059] Embodiment:

[0060] To overcome the deficiencies of the prior art, based on the Bayesian theory, this patent designs and proposes an OCR typo correction technology with low cost in the professional field.

[0061] The entire algorithm is mainly divided into three major parts:

[0062] 1. Train the word frequency distribution model;

[0063] 2. Use Bayes' formula to correct typos;

[0064] 3. Performance enhancement.

[0065] The first part. Train the word frequency distribution model:

[0066] The main purpose of the word frequency distribution model is to produce a vocabulary that records the conditional probability distribution of characters and characters, as well as the probability of individual words.

[0067] For example, in the word "physical examination", we need to count the frequency of "physical examination" when given a body, and the probability of "physical examination" and "physical examination" appearing respectively.

[0068] The specific steps are as follows:

[0069] The first step is to prepare the data:

[0070] The present invention prepares a large amount of text data in professional fields. In medical scenarios, medical-related text inputs such as hospital admission and discharge records, outpatient records, test reports, etc. can be prepared.

[0071] Here we assume that typos are low-frequency data, so we can use the text data produced by OCR as the basic data.

[0072] Step 2: Training data:

[0073] Training only requires recording the frequency of each word and the frequency of other words when word a appears.

[0074] The third step is frequency conversion probability:

[0075] The conditional probability is and ;

[0076] Here: P(t i+1 丨t i ) indicates that given t i Time t i+1 The probability of , and conditional probability;

[0077] t∈T represents any character, T is a set of characters; t i With t i+1 Represents two adjacent characters;

[0078] freq(t i+1 丨t i ) indicates that given t i Time t i+1 Frequency, because Z is a fixed value of total word frequency, so any value greater than max[freq(t i+1 丨t i )丨t∈T], here let Z=max[freq(t)丨t∈T].

[0079] Part 2: Using the Bayesian formula to correct typos:

[0080] In this embodiment, assume that a word is a typo, and set the joint distribution of the characters corresponding to the word to be lower than a probability threshold θ. At the same time, set the word as word = [t i , t i+1 ,....t i+n , if and only if:

[0081] θ ≥ P(t i , t i+1 ,....t i+n ), this threshold θ is regarded as an adjustable parameter.

[0082] Based on the above assumptions, the correction algorithm can be divided into two steps:

[0083] 1. Calculate the joint distribution;

[0084] 2. Find and correct the typo.

[0085] The first step: Calculate the joint distribution:

[0086] The core of locating the typo is to calculate the joint distribution of the word. Therefore, the conditional probabilities between characters have been obtained in the first part. The joint probability of the first part of multiple characters constructed according to Bayes' formula is:

[0087] ;

[0088] The joint probability of the second part of multiple characters constructed according to Bayes' formula is:

[0089] ;

[0090] The second step: Find and correct the typo.

[0091] Set the joint distribution as the evaluation function, and regard this problem as an optimization problem:

[0092] Among them, the condition for the typo is: the score of the evaluation function

[0093] The optimization goal is: maximize the output of the evaluation function

[0094] For this, use dynamic programming to solve this problem:

[0095] Definition of dynamic programming:

[0096] dp[i] represents the maximum joint probability of the first i characters;

[0097] Transfer equation:

[0098] ;

[0099] content[i - 1] represents the previous character in the sequence;

[0100] t represents the candidate replacement character for the current character char[i];

[0101] P(t丨char[i-1]) represents the probability of t appearing under the condition of the previous character char[i-1].

[0102] The pseudo code is as follows:

[0103] pseudocode

[0104] function fix_typo(content, theta, char_space):

[0105] n = length(content)

[0106] dp = zero_array(size=n) / / Used to store the joint probability of the first i characters

[0107] sequence = copy(content) / / used to store the optimal character sequence

[0108] / / Initialize dp[0]

[0109] dp[0] = P(content[0])

[0110] for i = 1:n

[0111] dp[i] = dp[i-1]*P(content[i]|sequence[i-1])

[0112] if dp[i] < theta / / typo found

[0113] best_prob = 0

[0114] best_char = content[i]

[0115] for char from char_space

[0116] new_prob = dp[i-1]*P(char|sequence[i-1])

[0117] if new_prob > best_prob:

[0118] best_prob = new_prob

[0119] best_char = char

[0120] end for

[0121] / / Update dp

[0122] sequence[i] = best_char

[0123] dp[i] = best_prob

[0124] end if

[0125] end for

[0126] return sequence

[0127] end function

[0128] Part III, Performance Enhancement:

[0129] After the first part and the second part, it is basically possible to find and correct typos, but there are still two optimization points in the whole algorithm:

[0130] 1. Bidirectional scanning;

[0131] 2. Pruning the search space;

[0132] Bidirectional scanning:

[0133] The first part and the second part mainly introduce the calculation of single conditional probability. In order to maximize the accuracy of identifying typos, bidirectional detection is adopted. When z is a typo among a, b, z, d, e, it is necessary that both P(a, b, z) ≤ θ and P(e, d, z) ≤ θ hold.

[0134] When pruning the search space; when calculating , the space of char_space can vary based on different values of sequence[i - 1]. At this time, candidates with a conditional probability of 0 with respect to sequence[i - 1] are deleted to reduce the search space.

[0135] Example 1: To facilitate the description of the algorithm process, a specific example is given below:

[0136] The first step + the second step prepare data + record word frequency and:

[0137] Randomly select n files for recording, and record the word frequency, as well as 2-gram (the frequency of b when a exists);

[0138] Suppose the following two results are obtained:

[0139] Suppose there are 100 sentences of "high-density lipoprotein" data, among which 3 sentences have "lipoprotein" written as "lipoprotein self";

[0140] Then the resulting word frequency table is:

[0141] ;

[0142] The table of 2-gram is:

[0143]

[0144] The third step, conversion probability:

[0145] Let z be the highest frequency, which is 100 here,

[0146] Then the probability chart is:

[0147]

[0148] The table of 2-gram:

[0149]

[0150] The fourth step, find the misspelled words:

[0151] Suppose a sentence is "high-density lipoprotein self";

[0152] The idea of finding misspelled words is to find the minimum joint distribution;

[0153] This solution will calculate separately;

[0154] P(high, density) = 1 * 1 = 1

[0155] P(high, density, lipoprotein) = P(high, density) * P(lipoprotein | density) = 1 * 1 = 1

[0156] P(high, density, lipoprotein, lipid) = P(high, density, lipoprotein) * P(lipid | density) = 1 * 1 = 1

[0157] P(high, density, lipoprotein, lipid, protein) = P(high, density, lipoprotein, lipid) * P(protein | lipid) = 1 * 1 = 1

[0158] P(high, density, lipoprotein, lipid, protein, self) = P(high, density, lipoprotein, lipid, protein) * P(self | protein) = 1 * 0.3 = 0.3

[0159] It can be seen that the joint distribution of "high-density lipoprotein self" is 0.3, and it is also found that when traversing to "self", the probability drops suddenly.

[0160] If the suddenly dropped value is less than the originally defined value, it is considered that there is a misspelled word at this position;

[0161] Let θ = 0.5;

[0162] Then θ>0.3 can be concluded to be a typo.

[0163] Step 5 Correction:

[0164] The revised goal is to find a t that maximizes the probability of P(height, density, fat, protein, t);

[0165] Based on the Bayesian formula:

[0166] P(height, density, fat, egg, t) = P(height, density, fat, egg) * P(t|egg);

[0167] Maximizing P(height, density, fat, eggs, t) while keeping P(height, density, fat, eggs) constant is equivalent to maximizing P(t|eggs);

[0168] P(t|egg) contains:

[0169]

[0170] It can be seen that 0.97 is the maximum probability, so "自" will be corrected to "白".

[0171] "High-density lipoprotein" will be corrected to "high-density lipoprotein".

[0172] The above-mentioned embodiments only express a certain implementation method of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the patent scope of the present invention. It should be pointed out that for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent of the present invention shall be based on the attached claims.

Claims

1. An OCR typo correction method based on Bayes' theorem, characterized in that, The following steps are involved: S1. Collect a large amount of text data in professional fields; form a word frequency distribution model, and convert and train the frequency of the model; S2. Correct typos using the Bayesian formula and correction algorithm; S3. Optimize the correction algorithm to enhance performance; Specifically, S2 includes the Bayesian formula: ; P(A, B) is the joint probability of A and B; P(A) is the prior probability or marginal probability of A, without considering any factors of B; P(B|A) is the conditional probability of B given that A occurs. The value obtained from A is called the posterior probability of B. The correction algorithm of S2 includes: S21. Assume a word is a typo, and set the joint distribution of the characters corresponding to the word to be lower than a probability threshold θ. At the same time, set the word as word = [t i , t i+1 ,.... t i+n , if and only if: θ≥P(t i ,t i+1 ,....t i+n ), this threshold θ is regarded as an adjustable parameter; S22, calculating the joint distribution based on the assumptions in step S21; S23. Find and correct typos.

2. The OCR typo correction method based on Bayes' theorem according to claim 1, characterized in that, Said S1 specifically includes: S11. Prepare a large amount of text data in professional fields, including relevant text input in medical scenarios, and use the text data produced by OCR as basic data; S12, training only needs to record the frequency of each word, and the frequency of other words when word a appears; S13. The conditional probability of frequency conversion of the word frequency distribution model is: ; ; Among them, P(t i+1 丨t i ) indicates that given t i Time t i+1 The probability of , and conditional probability; t ∈ T represents any character, where T is a set of characters; t i and t i+1 represent two adjacent characters; freq(t i+1 丨t i ) indicates that given t i Time i+1 Frequency, Z is a fixed value of total word frequency, which is any value greater than max[freq(t i+1 丨t i )丨t∈T], therefore, Z=max[freq(t)丨t∈T].

3. The OCR typo correction method based on Bayes' theorem according to claim 1, characterized in that, The calculation formula of the joint distribution in S22 includes: The joint probability of the front part of multiple words constructed according to the Bayesian formula is: ; The joint probability of the latter part of multiple words constructed according to the Bayesian formula is: 。 4. The OCR typo correction method based on Bayes' theorem according to claim 1, wherein The method for correcting typos in S23 includes: S231, considering the joint distribution as an evaluation function, setting the conditions for typos, and setting the score value of the evaluation function to ≤θ; S232, optimizing the evaluation function by maximizing its output; S233. Use dynamic programming to solve.

5. A method for correcting OCR typos based on Bayes' theorem according to claim 4, characterized in that, The transfer equation of the S233 dynamic programming includes: ; dp[i] represents the maximum joint probability of the first i characters; content[i-1] represents the previous character in the sequence; t represents the candidate replacement character for the current character char[i]; P(t丨char[i-1]) represents the probability of t appearing under the condition of the previous character char[i-1].

6. The OCR typo correction method based on Bayes' theorem according to claim 1, characterized in that, The optimization of the correction algorithm in S3 includes: S31. Maximize the accuracy of identifying misspelled characters through bidirectional scanning and biphasic detection, satisfying the following conditions: when z is a misspelled character among a, b, z, d, e, P(a, b, z) ≤ θ and P(e, d, z) ≤ θ must both hold true. S32, when pruning the search space; by deleting the candidates whose conditional probability with sequence[i-1] is 0.

Citation Information

Patent Citations

  • Medical OCR (Optical Character Recognition) error correction method

    CN116306594A

  • Text error correction method and apparatus, and electronic device and storage medium

    WO2022174495A1