Text OCR and formula OCR fast fusion method based on mask mechanism

Through the method based on masking mechanism, the precise positioning and format adjustment of text and formulas is solved, and the problem of inaccurate positioning and complex and time-consuming fusion in the existing technology is achieved, and efficient and accurate text and formula OCR fusion is achieved, which is suitable for large-scale data processing.

CN120496088AInactive Publication Date: 2025-08-15BEIJING DIGITAL FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510388494.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing OCR fusion method of text and formulas has problems such as inaccurate positioning and complex and time-consuming fusion process, which may result in misalignment or mismatch, and is inefficient and cannot meet the needs of large-scale data processing.

Method used

Using a masking mechanism-based method, the text and formula areas are accurately positioned through deep learning and formula detection algorithms, coordinate mapping relationships are established, format adjustments and error corrections are carried out to achieve high-precision fusion of text and formulas.

Benefits of technology

It improves the positioning accuracy of text and formulas, shortens the fusion time, improves processing efficiency, and ensures the accuracy and aesthetics of the fusion results, reducing the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496088A_ABST
    Figure CN120496088A_ABST
Patent Text Reader

Abstract

The invention discloses a text OCR and formula OCR fast fusion method based on a mask mechanism, and belongs to the technical field of formula fusion, and the method comprises the following steps: S1, image acquisition; s2, image enhancement: carrying out graying, binaryzation, denoising and other preprocessing operations on the original image to improve the definition and quality of the image and facilitate subsequent processing; s3, text positioning and mask creation; s4, formula positioning and mask creation; s5, accurately positioning the text; s6, performing formula accurate positioning; s7, performing OCR processing on the text; s8, performing formula OCR processing; s9, establishing a coordinate mapping relation; s10, carrying out fusion processing; s11, format adjustment is carried out; and S12, error correction: the algorithm is checked manually or automatically. According to the method, the initial positioning and the accurate positioning of the text and the formula are combined, and optimization is carried out by virtue of a mask mechanism, so that the high-precision positioning of the text and the formula can be realized. Therefore, the problems of dislocation and mismatching are effectively avoided, and the fusion accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of formula fusion, and in particular to a fast fusion method of text OCR and formula OCR based on a mask mechanism. Background Art

[0002] In today's digital age, a large number of documents and materials contain mixed content such as text and mathematical formulas. To digitize this content, it is usually necessary to perform text OCR and formula OCR separately, and then fuse the two results. However, existing fusion methods have some shortcomings. On the one hand, the positioning of text and formulas is not accurate enough, resulting in possible misalignment or mismatch in the fused results. On the other hand, the fusion process is relatively complex and time-consuming, with low efficiency, and cannot meet the needs of large-scale data processing. Therefore, we propose a fast fusion method for text OCR and formula OCR based on a mask mechanism to solve this problem. Summary of the Invention

[0003] The purpose of the present invention is to provide a fast fusion method of text OCR and formula OCR based on a mask mechanism to solve the problems raised in the above background technology.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions:

[0005] A fast fusion method for text OCR and formula OCR based on a mask mechanism, including:

[0006] S1. Image acquisition: Acquire the original image containing text and formulas through a scanning device or image acquisition device, and convert the physical or virtual content containing text and formulas into a digital image format for subsequent processing and analysis;

[0007] S2, Image Enhancement: Perform pre-processing operations such as grayscale conversion, binarization, and denoising on the original image to improve the clarity and quality of the image and facilitate subsequent processing;

[0008] S3. Text localization and mask creation: Use a text detection algorithm, such as a deep learning-based text detection model, to locate the text area in the image. Based on the localization results, create a text mask with the same size as the original image. The pixel values of the text area are set to 1, and the pixel values of the non-text area are set to 0.

[0009] S4, formula location and mask creation: Use the formula detection algorithm to locate the formula area in the image, and generate a formula mask based on the location result. The pixel value of the formula area is set to 1, and the pixel value of other areas is set to 0;

[0010] S5. Precise text positioning: Based on the text mask, further analyze the text line, word and other information, and use projection method, connected domain analysis and other methods to accurately adjust the text position, optimize the text mask and make the text positioning more accurate;

[0011] S6. Formula Precision Positioning: Based on the formula mask, combined with the structural characteristics and symbolic features of the mathematical formula, each part of the formula is more carefully positioned and segmented, and the formula mask is corrected to ensure the positioning accuracy of the formula;

[0012] S7, text OCR processing: extract the precisely located text area from the original image and input it into the text OCR engine for character recognition to obtain the text recognition result;

[0013] S8. Formula OCR processing: For the precisely located formula area, use a professional formula OCR tool to identify it and obtain a structured representation of the formula;

[0014] S9, establishing a coordinate mapping relationship: establishing a coordinate mapping relationship between the text and the formula in the original image based on the text mask and the formula mask, recording the starting coordinates and ending coordinates of each text character and formula element and the relative position relationship between them;

[0015] S10, fusion processing: according to the coordinate mapping relationship, the text recognition results and the structured representation of the formula are fused. During the fusion process, the position and format of the text and the formula are automatically adjusted according to their relative positions to ensure that the fused result conforms to the content layout in the original image;

[0016] S11. Format adjustment: adjust the format of the fused result, including unified settings of font, font size, line spacing, etc., to make it more beautiful and standardized;

[0017] S12. Error correction: Correct possible errors in the fusion results through manual inspection or automatic proofreading algorithms to improve the accuracy of fusion.

[0018] Preferably, in said S2, the specific steps are:

[0019] S201, Grayscale: Convert the color information of each pixel in the original image into a grayscale value. Use the weighted average method to assign different weights to the R, G, and B channels according to the sensitivity of the human eye to different color channels. Calculate the grayscale value of each pixel to reduce the complexity of the image data, highlight the image outline and contrast, and facilitate subsequent processing.

[0020] S202, Binarization: Selecting an appropriate threshold value, dividing the pixels in the grayscale image into two categories based on the comparison result between their grayscale values and the threshold value, setting pixels greater than or equal to the threshold value to white, and pixels less than the threshold value to black, further simplifying the image data, making the text and formula areas contrast sharply with the background, and facilitating the location of the text and formulas;

[0021] S203, denoising: Use methods such as median filtering and Gaussian filtering to remove noise points in the image. Median filtering is achieved by replacing the value of each pixel with the median of the pixel values in its neighborhood. Gaussian filtering performs a convolution operation on the image through a convolution kernel to smooth the image, eliminate noise interference introduced during the image acquisition process, improve image quality, and avoid the impact of noise on subsequent positioning and recognition.

[0022] Preferably, the specific steps in S3 are as follows:

[0023] S301. Text Detection Model Selection and Application: Select appropriate deep learning-based text detection models and train them on large-scale text datasets to enable them to learn the characteristic patterns of text. Leveraging the powerful feature extraction capabilities of deep learning models, they can accurately detect text regions in images.

[0024] S302. Locating text regions: Applying the trained text detection model to the original image. The model outputs the probability that each pixel belongs to a text region. Based on a set threshold, the bounding box of the text region is determined, and the specific location of the text in the image is found, providing a basis for subsequent text mask creation.

[0025] S303. Create a text mask: Create a blank mask image of the same size as the original image. For the pixels in the text area determined by the text detection model, set their values to 1 to indicate that these pixels belong to text; set the pixel values of the non-text area to 0, and generate a mask for marking the text area so that the text can be accurately located and extracted in subsequent processing.

[0026] Preferably, the specific steps in S4 are as follows:

[0027] S401. Formula Detection Algorithm Selection and Application: Use detection algorithms specifically targeting mathematical formulas, such as rule-based and machine learning-based methods. These algorithms can be trained and optimized on publicly available formula datasets to accurately detect mathematical formula regions in images.

[0028] S402. Locating the formula area: Using the selected formula detection algorithm to process the original image, identify the bounding box of the formula. Formulas usually have complex structures and may contain multiple elements and nested relationships. Therefore, the algorithm needs to have high accuracy and robustness to determine the location of the formula in the image in preparation for the creation of the formula mask.

[0029] S403. Create a formula mask: Similar to creating a text mask, create a formula mask image with the same size as the original image. For pixels in the formula area, set their values to 1, and set the pixel values in the non-formula area to 0. Generate a mask for marking the formula area to facilitate subsequent precise positioning and processing of the formula.

[0030] Preferably, the specific steps in S5 are as follows:

[0031] S501. Analyze text line and character information: Further analyze the text area marked by the text mask. Through projection and connected component analysis, determine the number of text lines, the start and end positions of each line, and the approximate position of each character, and gain a deeper understanding of the structure and layout of the text, providing a basis for precise positioning.

[0032] S502, accurately adjust the text position: according to the text line and word information obtained by analysis, modify the text mask to optimize the positioning accuracy of the text and ensure that the text position is accurate.

[0033] Preferably, the specific steps in S6 are as follows:

[0034] S601. Analyze formula structure and symbol features: Analyze the formula area marked by the formula mask, such as identifying elements such as brackets, operators, and function names, and understanding the relationship between them. Analyze based on the characteristics of the symbols in the formula to understand the specific structure and symbol information of the formula, so as to perform more detailed positioning and segmentation.

[0035] S602. Detailed positioning and segmentation of formulas: Based on the structural characteristics and symbolic features of the formula, each part of the formula is more precisely positioned, and the formula mask is corrected to ensure the accurate position of each formula element, improve the positioning accuracy of the formula, and provide accurate data for subsequent formula recognition and fusion.

[0036] Preferably, the specific steps in S7 are as follows:

[0037] S701, text region extraction: extracting the text region from the original image based on the precisely located text mask, and performing operations such as image cropping to separate the text region so that it can be input into the OCR engine for recognition;

[0038] S702, OCR recognition: The extracted text area is input into the text OCR engine for character recognition. The OCR engine converts the characters in the image into text encoding that can be understood by the computer based on the pre-trained character model, obtains the text recognition result, and provides text data for subsequent fusion processing.

[0039] Preferably, the specific steps in S8 are as follows:

[0040] S801, formula area extraction: similarly, based on the precisely located formula mask, the formula area is extracted from the original image, and the formula area is extracted separately to prepare for processing by the formula OCR tool;

[0041] S802, Formula Recognition: Use professional formula OCR tools to recognize the extracted formula area. These tools usually use recognition algorithms specifically for mathematical formulas and can convert the formula into a structured representation to obtain the structured representation for integration with the text recognition results.

[0042] Preferably, the specific steps in S9 are as follows:

[0043] S901, coordinate recording: traverse each marked area in the text mask and formula mask, record its starting coordinates (x1, y1) and ending coordinates (x2, y2) in the original image, and record the relative position relationship between them, such as horizontal distance, vertical distance, etc., to provide accurate position information for subsequent fusion processing and ensure that the text and formula can maintain the correct relative position after fusion;

[0044] S902, mapping relationship storage: The established coordinate mapping relationship is stored in a suitable data structure to facilitate query and use during the fusion process, and to facilitate subsequent fusion of text and formula according to the coordinate mapping relationship.

[0045] Preferably, the specific steps in S10 are as follows:

[0046] S1001, read coordinate mapping relationship: read coordinate mapping relationship information from the stored data structure, obtain the position information of the text and formula in the original image, and provide a basis for fusion processing;

[0047] S1002. Fusion of text and formula: Based on the coordinate mapping relationship, the text recognition results and the structured representation of the formula are placed at corresponding positions in the original image. During the fusion process, the positions and formats of the text and formula are automatically adjusted according to their relative positions to achieve accurate fusion of the text and formula in the image, making it more visually natural and coherent.

[0048] Preferably, the specific steps in S11 are as follows:

[0049] S1101. Font uniformity setting: Adjust the fonts of the fused text and formulas according to the preset font requirements. For text, you can use the font setting function provided by the OCR engine; for formulas, you can use the formula editing tool to modify the font so that the entire fused image has a consistent font style and improve readability.

[0050] S1102. Font size adjustment: Adjust the size of text and formulas according to the preset font size standard. Similarly, the text can be adjusted using the relevant functions of the OCR engine, and the formula can be scaled using the formula editing tool to ensure that the text and formula are consistent in size and meet the standards.

[0051] S1103. Aesthetic settings: Further aesthetic settings can be made to the fused image as needed, such as adjusting line spacing, paragraph spacing, etc. For the text part, this can be set using the functions of the word processing software; for the formula part, adjustments can be made in the formula editing tool to enhance the overall aesthetics of the fused image, making it more comfortable and easy to read.

[0052] Preferably, the specific steps in S12 are as follows:

[0053] S1201, Manual Inspection: Professionals conduct a comprehensive inspection of the fused image, including text accuracy, formula correctness, and layout rationality. Using their extensive experience and expertise, they identify potential errors or inaccuracies. Through manual professional judgment, they identify potential errors in the fusion process, providing a reference for subsequent automated inspections.

[0054] S1202. Application of automatic inspection algorithms: Use automatic inspection algorithms to quickly inspect the fused images. These algorithms can inspect a large amount of content in a short period of time and detect some common errors, such as spelling errors and formula symbol errors, thereby improving inspection efficiency, reducing the workload of manual inspection, and ensuring the accuracy of inspection results.

[0055] The beneficial effects of the present invention are:

[0056] 1. The present invention describes a method for rapidly fusing text OCR and formula OCR based on a masking mechanism. By combining preliminary and precise positioning of text and formulas, and optimizing them with the aid of a masking mechanism, it is possible to achieve high-precision positioning of text and formulas. This effectively avoids problems of misalignment and mismatch, and improves the accuracy of fusion.

[0057] 2. The present invention discloses a method for rapidly fusing text OCR and formula OCR based on a mask mechanism. This method can rapidly integrate the results of text OCR and formula OCR. By establishing a coordinate mapping relationship, the relative positions of the text and formula in the original image are maintained, significantly shortening the fusion time and improving processing efficiency. This is particularly important for large-scale data processing and can significantly improve work efficiency.

[0058] 3. In the present invention, the method for rapidly fusing text OCR and formula OCR based on a masking mechanism, after post-processing operations such as format adjustment and error correction, produces a fused result that is not only accurate in content but also more standardized and aesthetically pleasing in appearance. The uniform font, size, and line spacing settings ensure that the final fused result has good readability and visual experience.

[0059] 4. The present invention describes a method for rapidly integrating text OCR and formula OCR based on a masking mechanism. The automated process design reduces the need for manual intervention throughout the integration process. Although some manual inspection may be required during the error correction phase, overall, this method significantly reduces labor costs and improves work efficiency.

[0060] 5. The mask-based rapid fusion method for text and formula OCR described in this invention is not only applicable to the fusion of common text and simple formulas, but can also process complex mathematical expressions and scientific formulas. Its strong adaptability makes it have broad application prospects in various fields such as academic research, engineering technology, and education. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is a flowchart of a fast fusion method for text OCR and formula OCR based on a mask mechanism proposed by the present invention. DETAILED DESCRIPTION

[0062] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0063] Reference Figure 1 , a fast fusion method for text OCR and formula OCR based on mask mechanism, including:

[0064] S1. Image acquisition: Acquire the original image containing text and formulas through a scanning device or image acquisition device, and convert the physical or virtual content containing text and formulas into a digital image format for subsequent processing and analysis;

[0065] S2, Image Enhancement: Perform pre-processing operations such as grayscale conversion, binarization, and denoising on the original image to improve the clarity and quality of the image and facilitate subsequent processing;

[0066] S3. Text localization and mask creation: Use a text detection algorithm, such as a deep learning-based text detection model, to locate the text area in the image. Based on the localization results, create a text mask with the same size as the original image. The pixel values of the text area are set to 1, and the pixel values of the non-text area are set to 0.

[0067] S4, formula location and mask creation: Use the formula detection algorithm to locate the formula area in the image, and generate a formula mask based on the location result. The pixel value of the formula area is set to 1, and the pixel value of other areas is set to 0;

[0068] S5. Precise text positioning: Based on the text mask, further analyze the text line, word and other information, and use projection method, connected domain analysis and other methods to accurately adjust the text position, optimize the text mask and make the text positioning more accurate;

[0069] S6. Formula Precision Positioning: Based on the formula mask, combined with the structural characteristics and symbolic features of the mathematical formula, each part of the formula is more carefully positioned and segmented, and the formula mask is corrected to ensure the positioning accuracy of the formula;

[0070] S7, text OCR processing: extract the precisely located text area from the original image and input it into the text OCR engine for character recognition to obtain the text recognition result;

[0071] S8. Formula OCR processing: For the precisely located formula area, use a professional formula OCR tool to identify it and obtain a structured representation of the formula;

[0072] S9, establishing a coordinate mapping relationship: establishing a coordinate mapping relationship between the text and the formula in the original image based on the text mask and the formula mask, recording the starting coordinates and ending coordinates of each text character and formula element and the relative position relationship between them;

[0073] S10, fusion processing: according to the coordinate mapping relationship, the text recognition results and the structured representation of the formula are fused. During the fusion process, the position and format of the text and the formula are automatically adjusted according to their relative positions to ensure that the fused result conforms to the content layout in the original image;

[0074] S11. Format adjustment: adjust the format of the fused result, including unified settings of font, font size, line spacing, etc., to make it more beautiful and standardized;

[0075] S12. Error correction: Correct possible errors in the fusion results through manual inspection or automatic proofreading algorithms to improve the accuracy of fusion.

[0076] In this embodiment, in S2, the specific steps are:

[0077] S201, Grayscale: Convert the color information of each pixel in the original image into a grayscale value. Use the weighted average method to assign different weights to the R, G, and B channels according to the sensitivity of the human eye to different color channels. Calculate the grayscale value of each pixel to reduce the complexity of the image data, highlight the image outline and contrast, and facilitate subsequent processing.

[0078] S202, Binarization: Selecting an appropriate threshold value, dividing the pixels in the grayscale image into two categories based on the comparison result between their grayscale values and the threshold value, setting pixels greater than or equal to the threshold value to white, and pixels less than the threshold value to black, further simplifying the image data, making the text and formula areas contrast sharply with the background, and facilitating the location of the text and formulas;

[0079] S203, denoising: Use methods such as median filtering and Gaussian filtering to remove noise points in the image. Median filtering is achieved by replacing the value of each pixel with the median of the pixel values in its neighborhood. Gaussian filtering performs a convolution operation on the image through a convolution kernel to smooth the image, eliminate noise interference introduced during the image acquisition process, improve image quality, and avoid the impact of noise on subsequent positioning and recognition.

[0080] In this embodiment, the specific steps in S3 are as follows:

[0081] S301. Text Detection Model Selection and Application: Select appropriate deep learning-based text detection models and train them on large-scale text datasets to enable them to learn the characteristic patterns of text. Leveraging the powerful feature extraction capabilities of deep learning models, they can accurately detect text regions in images.

[0082] S302. Locating text regions: Applying the trained text detection model to the original image. The model outputs the probability that each pixel belongs to a text region. Based on a set threshold, the bounding box of the text region is determined, and the specific location of the text in the image is found, providing a basis for subsequent text mask creation.

[0083] S303. Create a text mask: Create a blank mask image of the same size as the original image. For the pixels in the text area determined by the text detection model, set their values to 1 to indicate that these pixels belong to text; set the pixel values of the non-text area to 0, and generate a mask for marking the text area so that the text can be accurately located and extracted in subsequent processing.

[0084] In this embodiment, the specific steps in S4 are as follows:

[0085] S401. Formula Detection Algorithm Selection and Application: Use detection algorithms specifically targeting mathematical formulas, such as rule-based and machine learning-based methods. These algorithms can be trained and optimized on publicly available formula datasets to accurately detect mathematical formula regions in images.

[0086] S402. Locating the formula area: Using the selected formula detection algorithm to process the original image, identify the bounding box of the formula. Formulas usually have complex structures and may contain multiple elements and nested relationships. Therefore, the algorithm needs to have high accuracy and robustness to determine the location of the formula in the image in preparation for the creation of the formula mask.

[0087] S403. Create a formula mask: Similar to creating a text mask, create a formula mask image with the same size as the original image. For pixels in the formula area, set their values to 1, and set the pixel values in the non-formula area to 0. Generate a mask for marking the formula area to facilitate subsequent precise positioning and processing of the formula.

[0088] In this embodiment, the specific steps in S5 are as follows:

[0089] S501. Analyze text line and character information: Further analyze the text area marked by the text mask. Through projection and connected component analysis, determine the number of text lines, the start and end positions of each line, and the approximate position of each character, and gain a deeper understanding of the structure and layout of the text, providing a basis for precise positioning.

[0090] S502, accurately adjust the text position: according to the text line and word information obtained by analysis, modify the text mask to optimize the positioning accuracy of the text and ensure that the text position is accurate.

[0091] In this embodiment, the specific steps in S6 are as follows:

[0092] S601. Analyze formula structure and symbol features: Analyze the formula area marked by the formula mask, such as identifying elements such as brackets, operators, and function names, and understanding the relationship between them. Analyze based on the characteristics of the symbols in the formula to understand the specific structure and symbol information of the formula, so as to perform more detailed positioning and segmentation.

[0093] S602. Detailed positioning and segmentation of formulas: Based on the structural characteristics and symbolic features of the formula, each part of the formula is more precisely positioned, and the formula mask is corrected to ensure the accurate position of each formula element, improve the positioning accuracy of the formula, and provide accurate data for subsequent formula recognition and fusion.

[0094] In this embodiment, the specific steps in S7 are as follows:

[0095] S701, text region extraction: extracting the text region from the original image based on the precisely located text mask, and performing operations such as image cropping to separate the text region so that it can be input into the OCR engine for recognition;

[0096] S702, OCR recognition: The extracted text area is input into the text OCR engine for character recognition. The OCR engine converts the characters in the image into text encoding that can be understood by the computer based on the pre-trained character model, obtains the text recognition result, and provides text data for subsequent fusion processing.

[0097] In this embodiment, the specific steps in S8 are as follows:

[0098] S801, formula area extraction: similarly, based on the precisely located formula mask, the formula area is extracted from the original image, and the formula area is extracted separately to prepare for processing by the formula OCR tool;

[0099] S802, Formula Recognition: Use professional formula OCR tools to recognize the extracted formula area. These tools usually use recognition algorithms specifically for mathematical formulas and can convert the formula into a structured representation to obtain the structured representation for integration with the text recognition results.

[0100] In this embodiment, the specific steps in S9 are as follows:

[0101] S901, coordinate recording: traverse each marked area in the text mask and formula mask, record its starting coordinates (x1, y1) and ending coordinates (x2, y2) in the original image, and record the relative position relationship between them, such as horizontal distance, vertical distance, etc., to provide accurate position information for subsequent fusion processing and ensure that the text and formula can maintain the correct relative position after fusion;

[0102] S902, mapping relationship storage: The established coordinate mapping relationship is stored in a suitable data structure to facilitate query and use during the fusion process, and to facilitate subsequent fusion of text and formula according to the coordinate mapping relationship.

[0103] In this embodiment, the specific steps in S10 are as follows:

[0104] S1001, read coordinate mapping relationship: read coordinate mapping relationship information from the stored data structure, obtain the position information of the text and formula in the original image, and provide a basis for fusion processing;

[0105] S1002. Fusion of text and formula: Based on the coordinate mapping relationship, the text recognition results and the structured representation of the formula are placed at corresponding positions in the original image. During the fusion process, the positions and formats of the text and formula are automatically adjusted according to their relative positions to achieve accurate fusion of the text and formula in the image, making it more visually natural and coherent.

[0106] In this embodiment, the specific steps in S11 are as follows:

[0107] S1101. Font uniformity setting: Adjust the fonts of the fused text and formulas according to the preset font requirements. For text, you can use the font setting function provided by the OCR engine; for formulas, you can use the formula editing tool to modify the font so that the entire fused image has a consistent font style and improve readability.

[0108] S1102. Font size adjustment: Adjust the size of text and formulas according to the preset font size standard. Similarly, the text can be adjusted using the relevant functions of the OCR engine, and the formula can be scaled using the formula editing tool to ensure that the text and formula are consistent in size and meet the standards.

[0109] S1103. Aesthetic settings: Further aesthetic settings can be made to the fused image as needed, such as adjusting line spacing, paragraph spacing, etc. For the text part, this can be set using the functions of the word processing software; for the formula part, adjustments can be made in the formula editing tool to enhance the overall aesthetics of the fused image, making it more comfortable and easy to read.

[0110] In this embodiment, the specific steps in S12 are as follows:

[0111] S1201, Manual Inspection: Professionals conduct a comprehensive inspection of the fused image, including text accuracy, formula correctness, and layout rationality. Using their extensive experience and expertise, they identify potential errors or inaccuracies. Through manual professional judgment, they identify potential errors in the fusion process, providing a reference for subsequent automated inspections.

[0112] S1202. Application of automatic inspection algorithms: Use automatic inspection algorithms to quickly inspect the fused images. These algorithms can inspect a large amount of content in a short period of time and detect some common errors, such as spelling errors and formula symbol errors, thereby improving inspection efficiency, reducing the workload of manual inspection, and ensuring the accuracy of inspection results.

[0113] In this embodiment, when in use, a scanning device or image acquisition device is used to convert the original image containing text and formulas into a digital format, and pre-processing operations such as grayscale, binarization and denoising are performed on the original image to improve the clarity and quality of the image. A text detection model based on deep learning (such as CRAFT) and a detection algorithm specifically for mathematical formulas are used to locate the text area and formula area in the image respectively, and create corresponding text masks and formula masks. On the basis of preliminary positioning, the text line and word information as well as the structural characteristics and symbolic features of the formula are further analyzed to accurately adjust the position of the text and formula to optimize the quality. The text mask and formula mask are transformed, and the precisely located text area is extracted from the original image and input into a text OCR engine (such as Tesseract) for character recognition to obtain the text recognition result. Similarly, for the precisely located formula area, a professional formula OCR tool (such as Mathpix) is used for recognition to obtain the structured representation of the formula. According to the text mask and formula mask, the starting coordinates, ending coordinates and the relative position relationship between each text character and formula element are recorded. According to the coordinate mapping relationship, the text recognition result and the structured representation of the formula are fused. During the fusion process, the position and format of the text and formula are automatically adjusted according to their relative positions, and the format of the fused result is adjusted, including unified settings of font, font size, line spacing, etc., to make it more beautiful and standardized. Through manual inspection or automatic proofreading algorithm, possible errors in the fusion result are corrected to improve the accuracy of the fusion.

[0114] The above is a detailed introduction to the fast fusion method of text OCR and formula OCR based on a mask mechanism provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above examples is only used to help understand the method of the present invention and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A fast fusion method for text OCR and formula OCR based on mask mechanism, characterized by: include: S1. Image acquisition: acquiring the original image containing text and formulas through a scanning device or image acquisition device; S2, Image Enhancement: Perform pre-processing operations such as grayscale conversion, binarization, and denoising on the original image to improve the clarity and quality of the image and facilitate subsequent processing; S3. Text localization and mask creation: Use a text detection algorithm, such as a deep learning-based text detection model, to locate the text area in the image. Based on the localization results, create a text mask with the same size as the original image. The pixel values of the text area are set to 1, and the pixel values of the non-text area are set to 0. S4, formula location and mask creation: Use the formula detection algorithm to locate the formula area in the image, and generate a formula mask based on the location result. The pixel value of the formula area is set to 1, and the pixel value of other areas is set to 0; S5. Precise text positioning: Based on the text mask, we further analyze the text lines, characters, and other information. Through projection methods, connected domain analysis, and other methods, we precisely adjust the text position and optimize the text mask to make the text positioning more accurate. S6. Formula Precision Positioning: Based on the formula mask, combined with the structural characteristics and symbolic features of the mathematical formula, each part of the formula is more carefully positioned and segmented, and the formula mask is corrected to ensure the positioning accuracy of the formula; S7, text OCR processing: extract the precisely located text area from the original image and input it into the text OCR engine for character recognition to obtain the text recognition result; S8. Formula OCR processing: For the precisely located formula area, use a professional formula OCR tool to identify it and obtain a structured representation of the formula; S9, establishing a coordinate mapping relationship: establishing a coordinate mapping relationship between the text and the formula in the original image based on the text mask and the formula mask, recording the starting coordinates and ending coordinates of each text character and formula element, and the relative position relationship between them; S10, fusion processing: according to the coordinate mapping relationship, the text recognition results and the structured representation of the formula are fused. During the fusion process, the position and format of the text and formula are automatically adjusted according to their relative positions to ensure that the fused result conforms to the content layout in the original image; S11. Format adjustment: adjust the format of the fused result, including unified settings of font, font size, line spacing, etc., to make it more beautiful and standardized; S12. Error correction: Correct possible errors in the fusion results through manual inspection or automatic proofreading algorithms to improve the accuracy of fusion.

2. The fast fusion method of text OCR and formula OCR based on mask mechanism according to claim 1 is characterized in that: In said S2, the specific steps are: S201, grayscale conversion: convert the color information of each pixel in the original image into a grayscale value, and use a weighted average method to assign different weights to the R, G, and B channels according to the sensitivity of the human eye to different color channels, and calculate the grayscale value of each pixel; S202, Binarization: Select an appropriate threshold and classify the pixels in the grayscale image into two categories based on the comparison result between their grayscale values and the threshold. Pixels with values greater than or equal to the threshold are set to white, and pixels with values less than the threshold are set to black. S203, denoising: Use methods such as median filtering and Gaussian filtering to remove noise points in the image. Median filtering is achieved by replacing the value of each pixel with the median of the pixel values in its neighborhood. Gaussian filtering smoothes the image by convolving the convolution kernel with the image.

3. The method for fast fusion of text OCR and formula OCR based on mask mechanism according to claim 1, characterized in that: The specific steps in S3 are as follows: S301. Text Detection Model Selection and Application: Select appropriate deep learning-based text detection models and train them on large-scale text datasets to enable them to learn the characteristic patterns of text. S302, locating text areas: Applying the trained text detection model to the original image, the model outputs the probability that each pixel belongs to a text area, and determines the bounding box of the text area based on a set threshold; S303. Create a text mask: Create a blank mask image of the same size as the original image. For pixels in the text area determined by the text detection model, set their values to 1 to indicate that these pixels belong to text; set the pixel values of non-text areas to 0.

4. The method for fast fusion of text OCR and formula OCR based on mask mechanism according to claim 1, characterized in that: The specific steps in S4 are as follows: S401. Formula Detection Algorithm Selection and Application: Use detection algorithms specifically targeting mathematical formulas, such as rule-based and machine learning-based methods. These algorithms can be trained and optimized on publicly available formula datasets. S402, locating the formula area: using the selected formula detection algorithm to process the original image and identify the bounding box of the formula. Formulas usually have complex structures and may contain multiple elements and nested relationships, so the algorithm needs to have high accuracy and robustness; S403. Create a formula mask: Similar to creating a text mask, create a formula mask image with the same size as the original image. For pixels in the formula area, set their values to 1, and set the pixel values in the non-formula area to 0.

5. The fast fusion method of text OCR and formula OCR based on mask mechanism according to claim 1 is characterized in that: The specific steps in S5 are as follows: S501, analyzing text lines and characters: further analyzing the text area marked by the text mask, and determining the number of text lines, the start and end positions of each line, and the approximate position of each character through projection and connected component analysis; S502 , accurately adjusting the text position: modifying the text mask based on the text line and word information obtained through analysis.

6. The fast fusion method of text OCR and formula OCR based on mask mechanism according to claim 1, characterized in that: The specific steps in S6 are as follows: S601. Analyze formula structure and symbol features: Analyze the formula structure in the formula area marked by the formula mask, such as identifying elements such as brackets, operators, and function names, and understanding the relationship between them. Analyze based on the characteristics of the symbols in the formula. S602. Detailed positioning and segmentation of formulas: Based on the structural characteristics and symbolic features of the formula, each part of the formula is more precisely positioned, and the formula mask is corrected to ensure the accurate position of each formula element.

7. The fast fusion method of text OCR and formula OCR based on mask mechanism according to claim 1 is characterized in that: The specific steps in S7 are as follows: S701, text region extraction: extracting the text region from the original image based on the precisely located text mask, by image cropping and other operations; S702, OCR recognition: The extracted text area is input into a text OCR engine for character recognition. The OCR engine converts the characters in the image into text encoding that can be understood by the computer based on a pre-trained character model.

8. The method for fast fusion of text OCR and formula OCR based on mask mechanism according to claim 1, characterized in that: The specific steps in S8 are as follows: S801, formula area extraction: similarly, based on the precisely located formula mask, the formula area is extracted from the original image; S802. Formula recognition: Use professional formula OCR tools to recognize the extracted formula area. These tools usually use recognition algorithms specifically for mathematical formulas and can convert formulas into structured representations.

9. The fast fusion method of text OCR and formula OCR based on mask mechanism according to claim 1, characterized in that: The specific steps in S9 are as follows: S901, coordinate recording: traverse each marked area in the text mask and formula mask, record its starting coordinates (x1, y1) and ending coordinates (x2, y2) in the original image, and record the relative position relationship between them, such as horizontal distance, vertical distance, etc.; S902, mapping relationship storage: The established coordinate mapping relationship is stored in a suitable data structure to facilitate query and use during the fusion process, and to facilitate subsequent fusion of text and formula according to the coordinate mapping relationship.

10. The fast fusion method of text OCR and formula OCR based on mask mechanism according to claim 1, characterized in that: The specific steps in S10 are as follows: S1001, read coordinate mapping relationship: read coordinate mapping relationship information from a stored data structure; S1002, fusing text and formula: placing the text recognition result and the structured representation of the formula at corresponding positions in the original image according to the coordinate mapping relationship. During the fusion process, the position and format of the text and formula are automatically adjusted according to their relative positions; The specific steps in S11 are as follows: S1101. Unified font setting: Adjust the fonts of the merged text and formulas according to the preset font requirements. For text, the font setting function provided by the OCR engine can be used; for formulas, the font can be modified using the formula editing tool; S1102, font size adjustment: adjust the size of text and formulas according to the preset font size standard. Similarly, for text, the relevant functions of the OCR engine can be used, and for formulas, the formula editing tool can be used to scale; S1103, aesthetic settings: further aesthetic settings are performed on the fused image as needed, such as adjusting line spacing, paragraph spacing, etc. For text, these settings can be made using the functions of word processing software; for formulas, these settings can be made using the formula editing tool; The specific steps in S12 are as follows: S1201, Manual Inspection: Professionals conduct a comprehensive inspection of the fused image, including text accuracy, formula correctness, and layout rationality. They will identify possible errors or inaccuracies based on their extensive experience and expertise. S1202. Application of automatic inspection algorithms: Use automatic inspection algorithms to quickly inspect the fused images. These algorithms can inspect a large amount of content in a short period of time and detect some common errors, such as spelling errors and formula symbol errors.