OCR (Optical Character Recognition) result correction method, device and equipment and computer storage medium
Through the OCR recognition result correction method based on a large language model, the problems of garbled characters and character confusion in complex documents and multilingual texts in OCR technology are solved, and efficient and accurate text correction is achieved with strong adaptability and applicable to various document formats and language scenarios.
Patent Information
- Application Number
- CN202510780596.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-10-10
AI Technical Summary
When processing texts with poor scan quality, complex document formats, or mixed languages, existing OCR technology often produces garbled characters, grammatical errors, and character confusion in the recognition results, resulting in low text usability and subsequent processing efficiency. In addition, existing methods have poor adaptability and high maintenance costs.
A large language model-based OCR recognition result correction method is adopted. By constructing a semantic correction instruction template, including task definition, rule definition and format constraint goals, the pre-trained language model is used to correct the text after OCR recognition. It can accurately correct garbled characters, character confusion and grammatical errors, and adapt to various document formats and multilingual scenarios.
It improves the accuracy and usability of text, adapts to multiple document formats and multi-language mixed text scenarios, reduces maintenance costs, shortens text correction cycles, and maintains the original format of the text, with a wide range of applications.
Smart Images

Figure CN120766090A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of text processing technology, and in particular to a method, apparatus, device, and computer storage medium for correcting OCR recognition results based on a large language model. Background Art
[0002] Optical character recognition (OCR) technology is widely used to convert PDF documents or images into editable text. However, in practice, OCR technology faces numerous challenges, particularly when processing text with poor scan quality, complex document formats, or mixed languages. Recognition results often contain garbled characters, grammatical errors, character confusion, and word segmentation. These issues severely impact the usability of the text and the efficiency of subsequent processing.
[0003] Traditional OCR post-processing methods typically rely on rule engines, regular expressions, or template matching techniques to extract and correct content. These OCR post-processing methods lack adaptability. Changes in document format or language environment require rewriting or adjusting rules, resulting in high maintenance costs and limited applicability. Furthermore, existing methods struggle to process mixed-language text and have limited understanding of complex contextual semantics. Summary of the Invention
[0004] In response to the deficiencies of the above-mentioned prior art, the present application provides a method for correcting OCR recognition results based on a large language model, comprising: obtaining a text to be corrected after OCR recognition, wherein the text to be corrected includes: garbled characters, character confusion errors, or grammatical errors caused by OCR recognition errors; constructing a semantic correction instruction template, wherein the semantic correction instruction template includes: a task definition, a rule definition, and a format constraint target; inputting the text to be corrected and the semantic correction instruction template into a large language model, and outputting a corrected text result through the large language model, wherein the large language model is a pre-trained language model.
[0005] Optionally, in an embodiment of the present application, the task definition includes: specifying the removal of garbled characters and semantic error correction of the OCR recognition results; the rule definition includes: defining semantic judgment rules including garbled characters and meaningless characters; the format constraint target includes: stipulating the format requirements for retaining the paragraph structure, punctuation marks and professional terms of the original text.
[0006] Optionally, in an embodiment of the present application, the garbled characters include: at least one of: irregular characters, #, @ or □; the character confusion errors include: misrecognition of numbers and letters.
[0007] Optionally, in an embodiment of the present application, the large language model is capable of correcting mixed texts including at least two languages.
[0008] Optionally, in the embodiment of the present application, the OCR recognition result correction method based on a large language model further includes: extracting structured field data from the text to be corrected.
[0009] In another aspect of the present application, an OCR recognition result correction device based on a large language model is also provided, comprising:
[0010] a text to be corrected acquisition unit configured to acquire an OCR-recognized text to be corrected, the text to be corrected including one of garbled characters, character confusion errors, or syntax errors caused by OCR recognition errors;
[0011] a semantic correction instruction template construction unit configured to construct a semantic correction instruction template, the semantic correction instruction template including task definition, rule definition, and format constraint targets;
[0012] a text correction unit configured to input the text to be corrected and the semantic correction instruction template into a large language model, and output a corrected text result through the large language model, the large language model being a pre-trained language model.
[0013] Optionally, in the embodiment of the present application, the semantic correction instruction template construction unit further includes:
[0014] a task definition unit configured to specify garbled character elimination and semantic error correction for the OCR recognition result;
[0015] a rule definition unit configured to define semantic judgment rules including garbled characters and meaningless characters;
[0016] a format constraint target unit configured to specify format requirements for retaining paragraph structures, punctuation marks, and professional terms of the original text.
[0017] Optionally, in the embodiment of the present application, the OCR recognition result correction device based on a large language model further includes a structured field extraction unit configured to extract structured field data from the text to be corrected.
[0018] In another aspect of the embodiment of the present application, an OCR recognition result correction device based on a large language model is also provided, comprising:
[0019] a memory configured to store a computer program;
[0020] a processor configured to call and execute the computer program to implement the steps of the OCR recognition result correction method based on a large language model according to any one of the above.
[0021] In another aspect of the embodiments of the present application, a storage medium is also provided, including a software program suitable for execution by a processor of the steps of the method for correcting OCR recognition results based on a large language model as described in any of the above.
[0022] The device for correcting OCR recognition results based on a large language model includes a computer program stored on a medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform the method described in the above aspects and achieve the same technical effects.
[0023] The embodiments described in the present application have the following beneficial effects:
[0024] The present application provides a method for correcting OCR recognition results based on a large language model, which can accurately correct garbled characters, character confusion errors or syntax errors caused by OCR recognition errors, effectively improve the accuracy and usability of the text, and provide high-quality basic data for subsequent text processing. At the same time, the generalization ability of the large language model can be used to adapt to various different format documents and multi-language mixed text scenarios, save maintenance costs, have a wide range of applications, improve correction efficiency, and shorten the text correction period. In addition, the correction method provided by the present application can maintain the original format of the text, including paragraphs, tables, symbols, and is friendly to subsequent editing and use. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the description of the embodiments will be briefly introduced. The drawings in the following description are only exemplary embodiments of the present application.
[0026] Figure 1 is a flowchart of the method for correcting OCR recognition results based on a large language model provided by the present application;
[0027] Figure 2A is a schematic diagram of the text to be corrected provided by the present application Figure 1 ;
[0028] Figure 2B is a schematic diagram of the text to be corrected provided by the present application Figure 1 ;
[0029] Figure 3A is a schematic diagram of the text to be corrected provided by the present application
[0030] Figure 3B is a schematic diagram of the text to be corrected provided by the present application
[0031] Figure 4 is a schematic diagram of the device for correcting OCR recognition results based on a large language model provided by the present application
[0032] Figure 5 It is a structural diagram of a device for correcting OCR recognition results based on a large language model provided in this application. DETAILED DESCRIPTION
[0033] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0034] Example 1
[0035] The following combination Figure 1 、 Figure 2A 、 Figure 2B 、 Figure 3A as well as Figure 3B The method 100 for correcting OCR recognition results based on a large language model of the present application is described in detail. Figure 1 This is a flowchart of a method for correcting OCR recognition results based on a large language model provided in this application. Figure 2A This is the text to be corrected provided by this application Figure 1 . Figure 2B This is the corrected text provided by this application Figure 1 . Figure 3A This is the second schematic diagram of the text to be corrected provided by this application. Figure 3B This is the second diagram of the corrected text provided by this application. Figure 1 As shown, a method 100 for correcting OCR recognition results based on a large language model includes: in step S101, obtaining a text to be corrected after OCR recognition, wherein the text to be corrected includes: garbled characters, character confusion errors, or grammatical errors caused by OCR recognition errors; in step S102, constructing a semantic correction instruction template, wherein the semantic correction instruction template includes: a task definition, a rule definition, and a format constraint target; in step S103, inputting the text to be corrected and the semantic correction instruction template into a large language model, and outputting a corrected text result through the large language model, wherein the large language model is a pre-trained language model.
[0036] In step S101 , a text to be corrected after OCR recognition is obtained, where the text to be corrected includes: garbled characters, character confusion errors, or grammatical errors caused by OCR recognition errors.
[0037] In an embodiment of the present application, during the application of optical character recognition (OCR) technology, the OCR recognition system will perform text extraction processing on the image or document to be recognized and generate corresponding text content. However, due to factors such as image quality, document complexity, and multilingual environment, there are often some errors in the text after OCR recognition. The present application processes the text to be corrected that has errors. The text to be corrected is the text content that may have one or more of garbled characters, character confusion errors, and grammatical errors caused by OCR recognition errors after OCR recognition. Garbled characters appear as unrecognizable strange symbols or texts in the text, character confusion errors are when OCR recognition mistakenly confuses and replaces certain similar characters, and grammatical errors are problems in the overall grammatical level of the text that do not conform to the corresponding language specifications. These erroneous texts all require further correction processing. Among them, the text to be corrected can be, for example, Figure 2A or Figure 3A The text shown.
[0038] Then, the process proceeds to step S102. In step S102, a semantic correction instruction template is constructed, wherein the semantic correction instruction template includes: a task definition, a rule definition, and a format constraint target.
[0039] In the embodiment of this application, constructing a semantic correction instruction template is one of the key steps. The semantic correction instruction template contains three important parts, namely task definition, rule definition, and format constraint target, as follows:
[0040] The task definition section clearly explains the task of the large language model, making it clear that this task is to correct text after OCR recognition, with the goal of removing garbled characters and correcting semantic errors. For example, correcting garbled characters, character confusion errors, or grammatical errors in the text.
[0041] The rule definition section sets a series of targeted rules based on language rules and text processing experience, providing detailed reference standards for corrections in the large language model. Examples of these rules include character matching, similar character differentiation, character confusion, and grammar rules, ensuring that the corrected text conforms to standards in terms of language structure and semantic expression.
[0042] The format constraint target part specifies the format requirements that the corrected text should meet to ensure that the final output text is neat and beautiful in format and conforms to the format specifications of common documents, making it convenient for users to further use and process it.
[0043] By building a comprehensive and detailed semantic correction instruction template, we can provide clear correction guidance for the large language model, making it more targeted and effective when processing OCR-recognized text, thereby improving the quality and effectiveness of text correction.
[0044] It can be understood that the above description of the semantic correction instruction template is only exemplary, and the construction of the semantic correction instruction template protected by this application is not limited to the contents listed above. Those skilled in the art can set and construct the semantic correction instruction template according to actual conditions, as long as the technical principles of this application can be implemented.
[0045] Furthermore, the task definition includes: specifying the removal of garbled characters and semantic error correction of OCR recognition results; the rule definition includes: defining semantic judgment rules including garbled characters and meaningless characters; the format constraint target includes: stipulating the format requirements of retaining the paragraph structure, punctuation marks and professional terms of the original text.
[0046] In the embodiment of the present application, the task definition, that is, clarifying the processing objectives of the large language model for the text after OCR recognition, can include two core tasks: garbled character removal and semantic error correction. By accurately locating and removing garbled characters in the text, and correcting character confusion and grammatical errors based on grammatical rules and contextual information, the text can be restored to accuracy and fluency at the semantic and grammatical levels, thereby effectively improving the readability and usability of the text.
[0047] Rule definition includes: defining semantic judgment rules for garbled and meaningless characters. For example, defining character matching rules to identify and correct garbled characters; defining similar character differentiation rules to resolve character confusion; and defining grammatical rules to correct grammatical errors in text. Through rule definition, garbled and meaningless characters in text can be accurately identified, thereby improving the accuracy and effectiveness of text correction and ensuring the quality of the corrected text.
[0048] Exemplarily, the garbled characters include: at least one of: unconventional characters, #, @, or □; the character confusion errors include: misrecognition of numbers and letters.
[0049] Garbled characters include at least one of the non-conventional characters, such as , #, @, or □. These characters usually appear in the text due to OCR recognition errors, and their presence seriously interferes with the readability and accuracy of the text.
[0050] Character confusion errors primarily occur when numbers and letters are misidentified during the OCR recognition process, such as misidentifying the number "6" as the letter "G" or the letter "O" as the number "0." These errors not only affect the semantic communication of the text but can also lead to serious misunderstandings or incorrect applications in professional fields. Therefore, identifying and correcting these garbled characters and character confusion errors is crucial to improving the accuracy and usability of OCR recognition results.
[0051] Format constraint objectives clarify the specific requirements for the format of the corrected text, which stipulates the format requirements for retaining the paragraph structure, punctuation marks and professional terminology of the original text. Specifically, the paragraph structure requirements ensure that the paragraph divisions of the text are consistent with the original text, and the logical relationship and hierarchical structure between the paragraphs are retained, so that the overall layout of the corrected text is consistent with the original text, which is convenient for readers to understand and read. The format requirements for punctuation marks emphasize the accurate restoration of various punctuation marks in the original text, including but not limited to periods, commas, question marks, exclamation marks, etc., to ensure that the semantic expression of the text is clear and in line with language habits. The format requirements for professional terminology are intended to ensure that the professional vocabulary, specific expressions and terms with special meanings involved in the text are consistent with the original text in format, such as specific fonts, capitalization, abbreviation rules, etc., to ensure that the accuracy and authority of professional terminology are maintained and meet the high requirements of professional fields for text accuracy.
[0052] Finally, the process proceeds to step S103 . In step S103 , the text to be corrected and the semantic correction instruction template are input into a large language model, and the large language model is used to output a corrected text result. The large language model is a pre-trained language model.
[0053] In an embodiment of the present application, the text to be corrected obtained in step S101 and the semantic correction instruction template constructed in step S102 are input into a large language model that has been deeply trained in advance. After being trained on a large-scale corpus, the large language model already has powerful language understanding and text generation capabilities. Its capabilities in cross-document structure understanding, entity recognition, and contextual reasoning can locate target fields and extract their corresponding values. The language model does not rely on fixed table structures, keyword positions and other rules, but completes field matching and extraction based on semantic understanding and contextual judgment. The language model outputs a corrected text result, which not only corrects problems such as garbled characters, character confusion errors, and grammatical errors in the original text, but also ensures the semantic coherence and logic of the text, while meeting the requirements of typesetting, punctuation usage, and paragraph structure specified in the format constraint target.
[0054] Furthermore, the large language model can correct mixed texts containing at least two languages.
[0055] In an embodiment of the present application, the large language model can effectively correct texts containing a mixture of at least two languages. The large language model is particularly critical when processing texts in complex multilingual environments because it can simultaneously identify and process garbled characters, character confusion errors, and grammatical errors in different languages without confusing the semantics and structures of different languages. Moreover, the large language model of the present application, through its deep learning algorithm and training on large-scale multilingual text data, can accurately distinguish and understand the different language parts in the text, and apply corresponding language rules and correction strategies to each part.
[0056] For example, Figure 2A This is the text to be corrected provided by this application Figure 1 . Figure 2B This is the corrected text provided by this application Figure 1 . Figure 2A The text to be corrected includes Chinese and English. The text results output by the large language model of this application are as follows Figure 2B As shown, according to Figure 2A and Figure 2B It can be concluded that the large language model of this application significantly improves the accuracy of text results.
[0057] Furthermore, the correction method 100 of the OCR recognition result based on the large language model of the present application may also include: extracting structured field data from the text to be corrected.
[0058] In the embodiment of the present application, in addition to correcting the text, structured field data can be further extracted from the text to be corrected. This process is to classify and organize the key information in the text so that it can be presented in a form that is easier to understand and process. Specifically, the process involves in-depth analysis of the text to identify and locate specific fields, such as name, date, address, phone number, email address, company name, etc. These fields usually have a clear format and location in the document, such as in a table, form or formal document.
[0059] To achieve this, the Big Language Model leverages its understanding of language patterns and semantics to identify these structured fields using predefined rules or pattern matching techniques. For example, it can use regular expressions to identify the specific formats of phone numbers and email addresses, or identify and extract information such as names and company names through contextual analysis. Furthermore, for multilingual text, the Big Language Model is able to identify structured fields in different languages and accurately extract the corresponding data.
[0060] For example, Figure 3A This is the second schematic diagram of the text to be corrected provided by this application. Figure 3B This is the second diagram of the corrected text provided by this application. Figure 3A The form is recognized by OCR, and the form text output by the large language model of this application is as follows Figure 3B As shown, according to Figure 3A and Figure 3B It can be concluded that the large language model of this application has significantly improved the recognition accuracy of forms.
[0061] It should be noted that Figure 2A and Figure 2B 、 Figure 3A and Figure 3B The examples shown are merely exemplary, and are used to demonstrate the capabilities and effects of the large language model of the present application in text correction and form recognition. The scope of protection of the present application covers all implementation methods that can achieve the effects of the present application.
[0062] In summary, the embodiments of the present application provide a method for correcting OCR recognition results based on a large language model, which can accurately correct garbled characters, character confusion errors, or grammatical errors caused by OCR recognition errors, effectively improve the accuracy and usability of the text, and provide high-quality basic data for subsequent text processing; at the same time, the present application utilizes the generalization ability of the large language model to adapt to a variety of different formats of documents and multi-language mixed text scenarios, saving maintenance costs and having a wide range of applications, while improving correction efficiency and shortening the text correction cycle. In addition, the correction method provided by the present application can maintain the original format of the text, including paragraphs, tables, and symbols, which is friendly to subsequent editing and use.
[0063] Example 2
[0064] Corresponding to the method embodiment, on the other side of the embodiment of the present application, a device for correcting OCR recognition results based on a large language model is also provided. Figure 4 A schematic diagram of the structure of the correction device for OCR recognition results based on a large language model provided in an embodiment of the present application is shown. The correction device for OCR recognition results based on a large language model is Figure 1 The device corresponding to the correction method of the OCR recognition result based on the large language model in the corresponding embodiment is implemented by a virtual device. Figure 1 In the corresponding embodiment, the method for correcting OCR recognition results based on a large language model, each virtual module constituting the device for correcting OCR recognition results based on a large language model can be executed by an electronic device, such as a network device, a terminal device, or a server. Specifically, the device for correcting OCR recognition results based on a large language model in the embodiment of the present application includes:
[0065] The text to be corrected acquiring unit 01 is used to acquire the text to be corrected after OCR recognition, wherein the text to be corrected includes: garbled characters, character confusion errors or grammatical errors caused by OCR recognition errors.
[0066] The semantic correction instruction template construction unit 02 is used to construct a semantic correction instruction template, wherein the semantic correction instruction template includes: a task definition, a rule definition, and a format constraint target.
[0067] Furthermore, the task definition includes: specifying the removal of garbled characters and semantic error correction of OCR recognition results; the rule definition includes: defining semantic judgment rules including garbled characters and meaningless characters; the format constraint target includes: stipulating the format requirements of retaining the paragraph structure, punctuation marks and professional terms of the original text.
[0068] Exemplarily, the garbled characters include: at least one of: unconventional characters, #, @, or □; the character confusion errors include: misrecognition of numbers and letters.
[0069] The text correction unit 03 is used to input the text to be corrected and the semantic correction instruction template into a large language model, and output a corrected text result through the large language model, where the large language model is a pre-trained language model.
[0070] Furthermore, the semantic correction instruction template construction unit also includes: a task definition unit, which is used to specify the removal of garbled characters and semantic error correction of OCR recognition results; a rule definition unit, which is used to define semantic judgment rules including garbled characters and meaningless characters; and a format constraint target unit, which is used to specify the format requirements for retaining the paragraph structure, punctuation marks and professional terms of the original text.
[0071] Furthermore, the large language model can correct mixed texts containing at least two languages.
[0072] Furthermore, the apparatus for correcting OCR recognition results based on a large language model further includes: a structured field extraction unit for extracting structured field data from the text to be corrected.
[0073] It should be noted that the specific implementation and technical effects of the correction device for OCR recognition results based on the large language model in the embodiment of the present application can be referred to Figure 1 The corresponding correction method of the OCR recognition result based on the large language model will not be described here in detail.
[0074] Example 3
[0075] Corresponding to the method embodiments, embodiments of the present application also provide a device, such as a terminal or server, for correcting OCR recognition results based on a large language model. The server may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal may be, but is not limited to, a smartphone, tablet computer, laptop computer, or desktop computer.
[0076] An example of a hardware structure block diagram of a correction device for OCR recognition results based on a large language model provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, this may include:
[0077] Processor 1, communication interface 2, memory 3 and communication bus 4;
[0078] The processor 1, the communication interface 2, and the memory 3 communicate with each other via the communication bus 4;
[0079] Optionally, the communication interface 2 may be an interface of a communication module, such as an interface of a GSM module;
[0080] The processor 1 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0081] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0082] The processor 1 is specifically configured to execute the computer program stored in the memory 3 to perform the following steps:
[0083] Obtaining a text to be corrected after OCR recognition, wherein the text to be corrected includes: garbled characters, character confusion errors, or grammatical errors caused by OCR recognition errors;
[0084] Constructing a semantic correction instruction template, wherein the semantic correction instruction template includes: a task definition, a rule definition, and a format constraint target;
[0085] The text to be corrected and the semantic correction instruction template are input into a large language model, and a corrected text result is output through the large language model, where the large language model is a pre-trained language model.
[0086] The above-mentioned product can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details that are not fully described in this embodiment, please refer to the correction method of OCR recognition results based on a large language model provided in the embodiment of the present application.
[0087] Example 4
[0088] In an embodiment of the present application, a storage medium is further provided. The storage medium may store a program suitable for execution by a processor, wherein the program is used to:
[0089] Obtaining a text to be corrected after OCR recognition, wherein the text to be corrected includes: garbled characters, character confusion errors, or grammatical errors caused by OCR recognition errors;
[0090] Constructing a semantic correction instruction template, wherein the semantic correction instruction template includes: a task definition, a rule definition, and a format constraint target;
[0091] The text to be corrected and the semantic correction instruction template are input into a large language model, and a corrected text result is output through the large language model, where the large language model is a pre-trained language model.
[0092] The above-mentioned product can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details that are not fully described in this embodiment, please refer to the correction method of OCR recognition results based on a large language model provided in the embodiment of the present application.
[0093] Optionally, the detailed functions and extended functions of the program may refer to the above description.
[0094] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0096] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0097] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0098] The above is merely an exemplary embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope described in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for correcting OCR recognition results based on a large language model, characterized in that: include: Obtaining a text to be corrected after OCR recognition, wherein the text to be corrected includes: garbled characters, character confusion errors, or grammatical errors caused by OCR recognition errors; Constructing a semantic correction instruction template, wherein the semantic correction instruction template includes: a task definition, a rule definition, and a format constraint target; The text to be corrected and the semantic correction instruction template are input into a large language model, and a corrected text result is output through the large language model, where the large language model is a pre-trained language model.
2. The method for correcting OCR recognition results based on a large language model according to claim 1, characterized in that: The task definition includes: specifying to remove garbled characters and perform semantic error correction on the OCR recognition results; The rule definition includes: defining semantic judgment rules including garbled characters and meaningless characters; The format constraint goal includes: preserving the format requirements of the paragraph structure, punctuation marks and professional terms of the original text.
3. The method for correcting OCR recognition results based on a large language model according to claim 2, characterized in that: The garbled characters include: at least one of an unconventional character, #, @ or □; The character confusion error includes: misidentification of numbers and letters.
4. The method for correcting OCR recognition results based on a large language model according to claim 3, characterized in that: The large language model is capable of correcting mixed texts comprising at least two languages.
5. The method for correcting OCR recognition results based on a large language model according to claim 1, wherein: Also includes: Extracting structured field data from the text to be corrected.
6. A device for correcting OCR recognition results based on a large language model, characterized in that: include: A text-to-be-corrected acquiring unit, configured to acquire the text to be corrected after OCR recognition, wherein the text to be corrected includes: garbled characters, character confusion errors, or grammatical errors caused by OCR recognition errors; A semantic correction instruction template construction unit, configured to construct a semantic correction instruction template, wherein the semantic correction instruction template includes: a task definition, a rule definition, and a format constraint target; The text correction unit is used to input the text to be corrected and the semantic correction instruction template into a large language model, and output a corrected text result through the large language model, where the large language model is a pre-trained language model.
7. The device for correcting OCR recognition results based on a large language model according to claim 6, characterized in that: The semantic correction instruction template construction unit further includes: A task definition unit, wherein the task definition unit is used to specify garbled code removal and semantic error correction for the OCR recognition result; A rule definition unit, the rule definition unit is used to define semantic judgment rules including garbled characters and meaningless characters; The format constraint target unit is used to specify format requirements for retaining the paragraph structure, punctuation marks, and professional terminology of the original text.
8. The device for correcting OCR recognition results based on a large language model according to claim 7, characterized in that: Also includes: The structured field extraction unit is used to extract structured field data from the text to be corrected.
9. A device for correcting OCR recognition results based on a large language model, characterized in that: include: memory for storing computer programs; A processor is configured to call and execute the computer program to implement the steps of the method for correcting OCR recognition results based on a large language model as described in any one of claims 1 to 5.
10. A storage medium, characterized in that: The method comprises a software program, wherein the software program is suitable for executing, by a processor, the steps of the method for correcting an OCR recognition result based on a large language model as claimed in any one of claims 1 to 5.
Citation Information
Cited By
Bill identification method, apparatus and device, and computer program product
CN122200714A