An OCR recognition result processing method, device, equipment and storage medium

Through the LayoutLMv2 model and proximity algorithm, the problem of lack of structured OCR recognition results in the prior art is solved, efficient text classification and association matching are achieved, and structured display of recognition results is improved.

CN114511857BActive Publication Date: 2025-08-01SHANGHAI WEIWENJIA INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210089301.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-08-01
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

The existing OCR technology lacks structured information when identifying form image text results, resulting in the recognition results requiring manual screening or additional rules processing, which is poorly robust and inefficient.

Method used

The LayoutLMv2 pre-trained model is used to classify the OCR recognition results in text, calculate the coordinates of the center point of the text box through the proximity algorithm, and correlate and output the structured text content.

Benefits of technology

Efficient text classification and precise correlation matching of OCR recognition results are achieved, the structured display efficiency of recognition results is improved, and manual intervention is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511857B_ABST
    Figure CN114511857B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of OCR recognition processing, and particularly relates to an OCR recognition result processing method, device, equipment and storage medium. The present invention imports the OCR recognition result of a target image into a corresponding text recognition model for classification recognition, determines the first type of text and the second type of text, as well as the corresponding first type of text box and the second type of text box; by determining the first center point coordinates or the second center point coordinates of each text box, the nearest second center point coordinates corresponding to each first center point coordinate are calculated by using the proximity algorithm; by extracting the first type of text corresponding to each first center point coordinate and the second type of text corresponding to the second center point coordinate, the first type of text and the associated second type of text are paired and output for display; the OCR recognition result of the target image can be efficiently classified and extracted, and the classified texts are accurately associated and matched to structurally output and display the mutually associated text content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of OCR recognition processing, and particularly relates to an OCR recognition result processing method, device, equipment and storage medium. Background Art

[0002] OCR (Optical Character Recognition) technology mainly recognizes the text in an image as an editable string. Early OCR technology mainly recognized some simple document images. Due to the development of deep learning, current OCR technology has been widely applied to the text recognition of images in various complex scenarios.

[0003] Currently, the text recognition results of form images are all lines of text and corresponding coordinate positions. For these unstructured results, corresponding rules need to be provided additionally according to various application scenarios for editing and sorting before a structured display result can be obtained. The results recognized by existing OCR technology are only a string of editable strings and the text boxes and coordinate positions corresponding to the strings, without any structured information. For the recognition results, a series of rules often need to be established to screen and enter each item, or directly enter manually; the former has very poor robustness, and currently there is no completely effective set of rules for screening each item of information, and it is easy to have the situation of misaligned recognition results; the latter has low efficiency and requires a huge amount of human cost. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention provides an OCR recognition result processing method, device, equipment and storage medium. When applied, it can efficiently classify and extract the OCR recognition results, and perform accurate association matching on the classified text to structurally output and display the interrelated text content.

[0005] In a first aspect, the present invention provides an OCR recognition result processing method, including:

[0006] Obtain the OCR recognition result of the target image, where the OCR recognition result includes a plurality of text boxes and the text content in each text box, as well as the coordinate information of each text box;

[0007] Import the text content in each text box into the trained text recognition model to obtain the classification result of each text content;

[0008] Determine that each text content is the first type of text or the second type of text according to the classification result, and determine the text box corresponding to the first type of text as the first type of text box, and the text box corresponding to the second type of text as the second type of text box;

[0009] Calculate the center point coordinates of each text box based on the coordinate information of each text box. Among them, the center point coordinates of the first type of text box are the first center point coordinates, and the center point coordinates of the second type of text box are the second center point coordinates;

[0010] Calculate the second center point coordinates closest to each first center point coordinate through the proximity algorithm, and associate each first center point coordinate with its closest second center point coordinate one by one;

[0011] Associatively extract the first type of text in the first type of text box where each first center point coordinate is located, and the second type of text in the second type of text box where the second center point coordinate associated with each first center point coordinate is located;

[0012] Pair and output the associatively extracted first type of text and the corresponding second type of text for display.

[0013] Based on the above invention content, by importing the text content in the OCR recognition result of the target image into the corresponding text recognition model for classification recognition, it can be determined whether each text content is the first type of text or the second type of text, and it can be determined whether the corresponding text box is the first type of text box or the second type of text box; by calculating the coordinate information of each text box in the OCR recognition result of the target image, the center point coordinates of each text box can be obtained. The center point coordinates of the first type of text box are the first center point coordinates, and the center point coordinates of the second type of text box are the second center point coordinates; through the proximity algorithm, the second center point coordinates closest to each first center point coordinate can be calculated; by extracting the first type of text corresponding to each first center point coordinate and the second type of text corresponding to the second center point coordinates closest to each first center point coordinate, and then pairing and outputting the associatively extracted first type of text and the corresponding second type of text for display; it can efficiently classify and extract the OCR recognition result of the target image, and perform precise association matching on the classified text to structurally output and display the interrelated text content, which can be applied to various text association recognition and display scenarios of images.

[0014] In a possible design, the text recognition model uses the LayoutLMv2 pre-trained model, and its training process includes:

[0015] Obtain a text content sample set, which contains a number of first type of text samples with positive labels and a number of second type of text samples with negative labels;

[0016] Import the text content sample set into the LayoutLMv2 pre-trained model for classification training until the classification accuracy of the LayoutLMv2 pre-trained model for the first type of text samples and the second type of text samples reaches the set threshold.

[0017] In a possible design, the coordinate information of the text box includes the coordinates of a pair of diagonal points of the text box. Calculating the center point coordinates of the text box based on the coordinate information of the text box includes: calculating the center point coordinates of the text box based on the coordinates of the two diagonal points of the text box.

[0018] In a possible design, the coordinate information is planar coordinate information including an X-axis coordinate value and a Y-axis coordinate value. Calculating the center point coordinates of the text box based on the coordinates of the two diagonal points of the text box includes: adding the X-axis coordinate values and the Y-axis coordinate values of the two diagonal points respectively, and then dividing by 2 to obtain the center point coordinates of the text box.

[0019] In a possible design, calculating the second center point coordinates closest to each first center point coordinate through a proximity algorithm includes:

[0020] Calculating the Euclidean distance between each first center point coordinate and each second center point coordinate based on each first center point coordinate and each second center point coordinate;

[0021] Based on the Euclidean distance between each first center point coordinate and each second center point coordinate, comparing and determining the second center point coordinate closest to each first center point coordinate.

[0022] In a possible design, pairing and outputting and displaying the first type of text and the corresponding second type of text extracted by association includes:

[0023] Pairing each first type of text and the second type of text associated therewith one by one;

[0024] Outputting and displaying the paired first type of text and second type of text side by side in a single line with the first type of text in front and the second type of text behind.

[0025] In a second aspect, the present invention provides an OCR recognition result processing device, which includes an acquisition unit, a classification unit, a determination unit, a first calculation unit, a second calculation unit, an extraction unit, and a display unit, where:

[0026] The acquisition unit is used to acquire the OCR recognition result of the target image, and the OCR recognition result includes a plurality of text boxes, the text content in each text box, and the coordinate information of each text box;

[0027] The classification unit is used to import the text content in each text box into a trained text recognition model to obtain the classification result of each text content;

[0028] A determination unit, configured to determine each text content as a first - type text or a second - type text according to the classification result, and determine the text box corresponding to the first - type text as a first - type text box and the text box corresponding to the second - type text as a second - type text box;

[0029] A first calculation unit, configured to calculate the center - point coordinates of each text box according to the coordinate information of each text box, where the center - point coordinates of the first - type text box are the first center - point coordinates and the center - point coordinates of the second - type text box are the second center - point coordinates;

[0030] A second calculation unit, configured to calculate, through a proximity algorithm, the second center - point coordinates closest to each first center - point coordinate, and associate each first center - point coordinate with its closest second center - point coordinate one - by - one;

[0031] An extraction unit, configured to extract, in an associated manner, the first - type text in the first - type text box where each first center - point coordinate is located and the second - type text in the second - type text box where the second center - point coordinate associated with each first center - point coordinate is located;

[0032] A display unit, configured to pair - output and display the associated - extracted first - type text and the corresponding second - type text.

[0033] In a possible design, the device further includes a training unit, where the training unit is configured to obtain a text - content sample set, and the text - content sample set includes a number of first - type text samples with positive labels and a number of second - type text samples with negative labels; and import the text - content sample set into the LayoutLMv2 pre - trained model for classification training until the classification accuracy of the LayoutLMv2 pre - trained model for the first - type text samples and the second - type text samples reaches a set threshold.

[0034] In a third aspect, the present invention provides an OCR recognition result processing device, and the device includes:

[0035] A memory, configured to store instructions;

[0036] A processor, configured to read the instructions stored in the memory and execute any one of the methods in the first aspect according to the instructions.

[0037] In a fourth aspect, the present invention provides a computer - readable storage medium, where instructions are stored on the computer - readable storage medium, and when the instructions are run on a computer, the computer is enabled to execute any one of the methods in the first aspect.

[0038] In a fifth aspect, the present invention provides a computer program product including instructions, and when the instructions are run on a computer, the computer is enabled to execute any one of the methods in the first aspect.

[0039] The beneficial effects of the present invention are as follows:

[0040] By importing the text content in the OCR recognition result of the target image into the corresponding text recognition model for classification and recognition, the present invention can determine whether each text content is the first type of text or the second type of text, and determine whether the corresponding text box is the first type of text box or the second type of text box; by calculating the coordinate information of each text box in the OCR recognition result of the target image, the center point coordinates of each text box can be obtained. The center point coordinates of the first type of text box are the first center point coordinates, and the center point coordinates of the second type of text box are the second center point coordinates; through the proximity algorithm, the second center point coordinates closest to each first center point coordinate can be calculated; by extracting the first type of text corresponding to each first center point coordinate and the second type of text corresponding to the second center point coordinates closest to each first center point coordinate, and then pairing and outputting and displaying the associated first type of text and the corresponding second type of text; the OCR recognition result of the target image can be efficiently classified and extracted for text, and the classified texts can be accurately associated and matched to structurally output and display the mutually associated text content, which can be applied to various text association recognition and display scenarios of images. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 It is a schematic diagram of the method steps of the present invention;

[0043] Figure 2 It is a schematic diagram of the device structure of the present invention;

[0044] Figure 3 It is a schematic diagram of the device composition of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. It should be noted here that the description of these embodiment modes is used to help understand the present invention, but does not constitute a limitation to the present invention. The specific structural and functional details disclosed herein are only used to describe the exemplary embodiments of the present invention. However, the present invention can be embodied in many alternative forms and should not be construed as limited to the embodiments described herein.

[0046] It should be understood that the terms first, second, etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance. Although the terms first, second, etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, the first unit may be referred to as the second unit, and similarly, the second unit may be referred to as the first unit, without departing from the scope of the exemplary embodiments of the present invention.

[0047] Specific details are provided in the following description to facilitate a complete understanding of the exemplary embodiments. However, those of ordinary skill in the art should understand that the exemplary embodiments can be implemented without these specific details. For example, a system may be shown in a block diagram to avoid obscuring the example with unnecessary details. In other embodiments, well-known processes, structures, and techniques may not be shown with unnecessary details to avoid obscuring the exemplary embodiments.

[0048] Embodiment 1:

[0049] This embodiment provides an OCR recognition result processing method, as Figure 1 shown, including the following steps:

[0050] S101. Obtain the OCR recognition result of the target image, where the OCR recognition result includes a plurality of text boxes, the text content within each text box, and the coordinate information of each text box.

[0051] Specifically, when implemented, text recognition is performed on the target image through OCR recognition technology to obtain the OCR recognition result of the target image. The OCR recognition result includes a plurality of text boxes, the text content within each text box, and the coordinate information of each text box. The coordinate information is planar coordinate information including an X-axis coordinate value and a Y-axis coordinate value. The coordinate information of the text box includes the coordinates of a pair of diagonal points of the text box. The two diagonal points can be the diagonal points of the upper left corner and the lower right corner of the text box, or the diagonal points of the lower left corner and the upper right corner of the text box. The selection of the diagonal points is mainly for calculating the coordinates of the center point of the text box in the subsequent process. It is also possible to select the points of the four corners of the text box to complete the calculation of the coordinates of the center point of the text box in the subsequent process.

[0052] S102. Import the text content within each text box into the trained text recognition model to obtain the classification result of each text content.

[0053] In specific implementation, the text recognition model adopts the LayoutLMv2 pre-trained model, and its training process includes: obtaining a text content sample set, which contains a number of first-class text samples with positive labels and a number of second-class text samples with negative labels; then importing the text content sample set into the LayoutLMv2 pre-trained model for classification training until the classification accuracy of the LayoutLMv2 pre-trained model for the first-class text samples and the second-class text samples reaches a set threshold, and the threshold can be set according to actual requirements, such as set to 85%, 90%, 95%, 100%, etc.

[0054] Taking the identity information form image as an example here, the first-class text samples may include item texts such as "Name", "Gender", "Nationality", "Ethnic Group", "Date of Birth", "Identity Number", "Residence Address", etc. in the identity information form image, and the second-class text samples are the actual content texts corresponding to the items such as "Name", "Gender", "Nationality", "Ethnic Group", "Date of Birth", "Identity Number", "Residence Address", etc.

[0055] The trained text recognition model can be used to classify and recognize each text content to distinguish which are the corresponding first-class text samples and which are the corresponding second-class text samples, and then output the corresponding classification results.

[0056] S103. Determine whether each text content is a first-class text or a second-class text according to the classification result, and determine that the text box corresponding to the first-class text is the first-class text box, and the text box corresponding to the second-class text is the second-class text box.

[0057] In specific implementation, according to the classification result of the text recognition model, it can be effectively determined which text contents in the OCR recognition result of the target image are first-class texts and which text contents are second-class texts. If it is a text content of an unrecognizable type, it can be determined as other text and ignored. The text box corresponding to the first-class text can be determined as the first-class text box, and the text box corresponding to the second-class text can be determined as the second-class text box.

[0058] S104. Calculate the center point coordinates of each text box according to the coordinate information of each text box, where the center point coordinates of the first-class text box are the first center point coordinates, and the center point coordinates of the second-class text box are the second center point coordinates.

[0059] In specific implementation, the center point coordinates of each text box can be calculated by using the coordinate information of each text box in step S101. Taking the coordinates of the diagonal points of the upper left corner and the lower right corner of the text box as an example, the X-axis coordinate values and Y-axis coordinate values of the two diagonal points of the upper left corner and the lower right corner can be added respectively, and then divided by 2 to obtain the center point coordinates of the text box. After obtaining the center point coordinates of each text box, the center point coordinates of the first type of text box are set as the first center point coordinates, and the center point coordinates of the second type of text box are set as the second center point coordinates.

[0060] S105. Calculate the second center point coordinates closest to each first center point coordinate through the nearest neighbor algorithm, and associate each first center point coordinate with its closest second center point coordinate one by one.

[0061] In specific implementation, the process of calculating the second center point coordinates closest to each first center point coordinate through the nearest neighbor algorithm includes: calculating the Euclidean distance between each first center point coordinate and each second center point coordinate according to each first center point coordinate and the second center point coordinates; then, based on the Euclidean distances between each first center point coordinate and each second center point coordinate, compare and determine the second center point coordinates closest to each first center point coordinate. In the calculation process, if a second center point is a neighboring point of multiple first center points, then match this second center point to the first center point with the smallest Euclidean distance, and so on, so that each first center point uniquely corresponds to a second center point.

[0062] The nearest neighbor algorithm, that is, the K-Nearest Neighbor (KNN) classification algorithm, is one of the simplest methods in data mining classification technology. The nearest neighbor algorithm is a method of classifying each record in a data set. The core idea of the KNN algorithm is that if most of the K closest samples of a sample in the feature space belong to a certain category, then this sample also belongs to this category and has the characteristics of the samples in this category. This method only determines the category of the sample to be classified based on the categories of one or several closest samples when making a classification decision. When making a category decision, the KNN method is only related to a very small number of adjacent samples. Since the KNN method mainly relies on the limited surrounding neighboring samples rather than the method of discriminating the class domain to determine the category to which it belongs, the KNN method is more suitable for the sample set to be classified with more intersections or overlaps in the class domain than other methods.

[0063] The Euclidean distance, that is, the Euclidean distance or Euclidean metric, is a commonly used distance definition, referring to the actual distance between two points in an m-dimensional space, or the natural length of a vector (that is, the distance from this point to the origin). The Euclidean distance in two-dimensional and three-dimensional spaces is the actual distance between two points. The calculation formula on a two-dimensional plane can be expressed as:

[0064]

[0065] S106. Extract the first type of text in the first type of text boxes where each first center point coordinate is located, and the second type of text in the second type of text boxes where the second center point coordinates corresponding to each first center point coordinate are located, by associating them.

[0066] In specific implementation, after associating each first center point coordinate with its nearest second center point coordinate one by one, the first type of text and the second type of text in the text boxes where the associated first center point coordinate and the second center point coordinate are located can be extracted. Here, taking the identity information form image as an example again, if the first type of text extracted is "Name", then the second type of text associated with it is the actual content text "AA" corresponding to "Name"; if the first type of text extracted is "Gender", then the second type of text associated with it is the actual content text "BB" corresponding to "Gender"... and so on, associating each first type of text with the corresponding second type of text.

[0067] S107. Pair and output the extracted first type of text and the corresponding second type of text for display.

[0068] In specific implementation, the process of pairing and outputting each first type of text and the second type of text associated with it for display includes: pairing each first type of text with the second type of text associated with it one by one; and outputting and displaying the paired first type of text and the second type of text side by side in a single line with the first type of text in front and the second type of text behind. Here, taking the identity information form image as an example again, the final corresponding output display effect of the first type of text and the second type of text is as follows:

[0069]

[0070]

[0071] Through this embodiment, the OCR recognition results of the target image can be efficiently classified and extracted, and the classified texts can be accurately associated and matched to structurally output and display the mutually related text contents, which can be applied to various text association recognition and display scenarios of images.

[0072] Embodiment 2:

[0073] This embodiment provides an OCR recognition result processing device, as Figure 2 shown. The device includes an acquisition unit, a classification unit, a determination unit, a first calculation unit, a second calculation unit, an extraction unit, and a display unit, where:

[0074] An acquisition unit for acquiring the OCR recognition result of a target image, where the OCR recognition result includes a number of text boxes, the text content within each text box, and the coordinate information of each text box;

[0075] A classification unit for importing the text content within each text box into a trained text recognition model to obtain the classification result of each text content;

[0076] A determination unit for determining, based on the classification result, that each text content is a first type of text or a second type of text, and determining that the text box corresponding to the first type of text is a first type of text box, and the text box corresponding to the second type of text is a second type of text box;

[0077] A first calculation unit for calculating the center point coordinates of each text box based on the coordinate information of each text box, where the center point coordinates of the first type of text box are the first center point coordinates, and the center point coordinates of the second type of text box are the second center point coordinates;

[0078] A second calculation unit for calculating, through the proximity algorithm, the second center point coordinates closest to each first center point coordinate, and associating each first center point coordinate with its closest second center point coordinate one by one;

[0079] An extraction unit for associatively extracting the first type of text in the first type of text box where each first center point coordinate is located, and the second type of text in the second type of text box where the second center point coordinate associated with each first center point coordinate is located;

[0080] A display unit for pairing and outputting and displaying the associatively extracted first type of text and the corresponding second type of text.

[0081] Further, the device further includes a training unit for acquiring a text content sample set, where the text content sample set includes a number of first type of text samples with positive labels and a number of second type of text samples with negative labels; and importing the text content sample set into the LayoutLMv2 pre-training model for classification training until the classification accuracy of the LayoutLMv2 pre-training model for the first type of text samples and the second type of text samples reaches a set threshold.

[0082] Embodiment 3:

[0083] This embodiment provides an OCR recognition result processing device, as Figure 3 shown, at the hardware level, including:

[0084] A memory for storing instructions;

[0085] A processor for reading the instructions stored in the memory and executing the OCR recognition result processing method described in Embodiment 1 according to the instructions.

[0086] Optionally, the computer device further includes an internal bus and a communication interface. The processor, the memory, and the communication interface can be interconnected through the internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0087] The memory can but is not limited to include a random access memory (RAM), a read only memory (ROM), a flash memory, a first input first output (FIFO), and / or a first in last out (FILO), etc. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0088] Embodiment 4:

[0089] This embodiment provides a computer-readable storage medium, on which instructions are stored. When the instructions run on a computer, the computer is caused to execute the OCR recognition result processing method described in Embodiment 1. Among them, the computer-readable storage medium refers to a carrier for storing data, which can but is not limited to include a floppy disk, an optical disc, a hard disk, a flash memory, a USB flash drive, and / or a memory stick, etc. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0090] Embodiment 5:

[0091] This embodiment provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the OCR recognition result processing method described in Embodiment 1. Among them, the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0092] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An OCR recognition result processing method, characterized in that Including: Obtain the OCR recognition result of the target image, where the OCR recognition result includes a number of text boxes, the text content within each text box, and the coordinate information of each text box; Import the text content within each text box into the trained text recognition model to obtain the classification result of each text content; Determine whether each text content is a first type of text or a second type of text according to the classification result, and determine the text box corresponding to the first type of text as the first type of text box, and the text box corresponding to the second type of text as the second type of text box; Calculate the center point coordinates of each text box according to the coordinate information of each text box, where the center point coordinates of the first type of text box are the first center point coordinates, and the center point coordinates of the second type of text box are the second center point coordinates; Calculate the second center point coordinate closest to each first center point coordinate through the proximity algorithm, and associate each first center point coordinate with its closest second center point coordinate one by one; The process of calculating the second center point coordinate closest to each first center point coordinate through the proximity algorithm includes: calculating the Euclidean distance between each first center point coordinate and each second center point coordinate according to each first center point coordinate and the second center point coordinate, and comparing and determining the second center point coordinate closest to each first center point coordinate according to the Euclidean distance between each first center point coordinate and each second center point coordinate; Associatively extract the first type of text in the first type of text box where each first center point coordinate is located, and the second type of text in the second type of text box where each second center point coordinate associated with the first center point coordinate is located; Pair and output and display the associatively extracted first type of text and the corresponding second type of text.

2. The OCR recognition result processing method according to claim 1, wherein The text recognition model uses the LayoutLMv2 pre-trained model, and its training process includes: Obtain a text content sample set, where the text content sample set contains a number of first type of text samples with positive labels and a number of second type of text samples with negative labels; Import the text content sample set into the LayoutLMv2 pre-trained model for classification training until the classification accuracy of the LayoutLMv2 pre-trained model for the first type of text samples and the second type of text samples reaches the set threshold.

3. The OCR recognition result processing method according to claim 1, wherein The coordinate information of the text box includes the coordinates of a pair of diagonal points of the text box. The process of calculating the center point coordinates of the text box according to the coordinate information of the text box includes: calculating the center point coordinates of the text box according to the coordinates of the two diagonal points of the text box.

4. The OCR recognition result processing method according to claim 3, characterized in that The coordinate information is plane coordinate information including the X-axis coordinate value and the Y-axis coordinate value. The process of calculating the center point coordinates of the text box according to the coordinates of the two diagonal points of the text box includes: adding the X-axis coordinate values and the Y-axis coordinate values of the two diagonal points respectively, and then dividing by 2 to obtain the center point coordinates of the text box.

5. The OCR recognition result processing method according to claim 1, wherein The process of pairing and outputting and displaying the associatively extracted first type of text and the corresponding second type of text includes: Pair each first type of text with its associated second type of text one by one; Output and display the paired first type of text and second type of text in a single row side by side with the first type of text in the front and the second type of text in the back.

6. An OCR recognition result processing device, characterized in that, The device includes an acquisition unit, a classification unit, a determination unit, a first calculation unit, a second calculation unit, an extraction unit, and a display unit, where: The acquisition unit is configured to acquire the OCR recognition result of a target image, where the OCR recognition result includes a plurality of text boxes, the text content in each text box, and the coordinate information of each text box; The classification unit is configured to import the text content in each text box into a trained text recognition model to obtain the classification result of each text content; The determination unit is configured to determine, according to the classification result, that each text content is a first type of text or a second type of text, and determine that the text box corresponding to the first type of text is a first type of text box, and the text box corresponding to the second type of text is a second type of text box; The first calculation unit is configured to calculate the center point coordinates of each text box according to the coordinate information of each text box, where the center point coordinates of the first type of text box are the first center point coordinates, and the center point coordinates of the second type of text box are the second center point coordinates; The second calculation unit is configured to calculate, through a proximity algorithm, the second center point coordinate closest to each first center point coordinate, and associate each first center point coordinate with its closest second center point coordinate one by one; calculating, through the proximity algorithm, the second center point coordinate closest to each first center point coordinate includes: calculating the Euclidean distance between each first center point coordinate and each second center point coordinate according to each first center point coordinate and the second center point coordinate, and comparing and determining the second center point coordinate closest to each first center point coordinate according to the Euclidean distance between each first center point coordinate and each second center point coordinate; The extraction unit is configured to extract, in an associated manner, the first type of text in the first type of text box where each first center point coordinate is located, and the second type of text in the second type of text box where the second center point coordinate associated with each first center point coordinate is located; The display unit is configured to pair and output and display the extracted first type of text and the corresponding second type of text.

7. An OCR recognition result processing device according to claim 6, wherein The device further includes a training unit, where the training unit is configured to acquire a text content sample set, where the text content sample set includes a plurality of first type of text samples with positive labels and a plurality of second type of text samples with negative labels; and import the text content sample set into a LayoutLMv2 pre-trained model for classification training until the classification accuracy of the LayoutLMv2 pre-trained model for the first type of text samples and the second type of text samples reaches a set threshold.

8. An OCR recognition result processing device, characterized in that, The device includes: A memory for storing instructions; A processor for reading the instructions stored in the memory and executing the method according to any one of claims 1-5 according to the instructions.

9. A computer-readable storage medium, characterized in that, Instructions are stored on the computer-readable storage medium, and when the instructions run on the computer, the computer is caused to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Character recognition method and device, electronic equipment and storage medium

    CN113780098A

  • Table picture content extraction method, device and equipment based on artificial intelligence

    CN113887422A