Image recognition methods, systems, electronic devices and storage media

By using machine learning algorithms to correct image tilt and recognize structured table content, the problems of image tilt and noise interference in unstructured medical test reports have been solved, enabling fast and accurate data entry and table content recognition.

CN116778515BActive Publication Date: 2026-05-26BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV
Filing Date
2023-06-19
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In the existing digitization of unstructured medical test reports, image tilt and noise interference affect the accurate recognition of image content, especially the recognition of table content, which makes it difficult to accurately obtain the correspondence between data items.

Method used

Machine learning algorithms are used to perform image tilt correction and structured table content recognition, including image correction processing, optical character recognition, and data matching and extraction using preset algorithms. The correction angle is determined using pixel variance groups, and lost text information is extracted a second time using an OCR neural network.

Benefits of technology

It enables rapid data entry of images, improves the efficiency of digital work, and ensures accurate recognition of image content and acquisition of structured data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116778515B_ABST
    Figure CN116778515B_ABST
Patent Text Reader

Abstract

This application discloses an image recognition method, system, electronic device, and storage medium. The image recognition method includes: performing image correction processing on the image to be processed to obtain a forward image; performing image recognition on the forward image through optical character recognition to obtain image data information; and performing data matching and extraction on the image data information based on a preset algorithm to obtain document information with structured data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and more specifically, to an image recognition method, system, electronic device, and storage medium. Background Technology

[0002] With advancements in science and technology, the medical field has gradually evolved towards digital management. For example, in the current data collection of medical test reports, data integration with hospital information systems is a common data processing method.

[0003] However, in the current context of digitizing some unstructured medical test reports, most of these reports are in image, paper, or other document formats, making the digitization of these unstructured reports a challenging task.

[0004] Currently, the general approach is to convert medical test reports into image format, and then extract information from the images to obtain corresponding documents (such as tables) containing the extracted information.

[0005] However, the inventors of this application discovered that since images are generally obtained by photographing or scanning medical test reports, there may be a certain degree of tilt and noise interference. The tilt and noise interference of the image will affect the accurate recognition of the image content. In addition, since the table is a structured table, which is different from the recognition of general types of documents, the recognition of the table content also needs to consider the correspondence between data items.

[0006] Based on this, the technical solution that this application aims to solve is to provide an image recognition method for unstructured images (such as medical examination reports) to obtain corresponding tabular information. Summary of the Invention

[0007] This application discloses an image recognition method, system, electronic device, and storage medium. The image recognition method includes: performing image correction processing on the image to be processed to obtain a forward image; performing image recognition on the forward image through optical character recognition to obtain image data information; and performing data matching and extraction on the image data information based on a preset algorithm to obtain document information with structured data.

[0008] According to some embodiments of this application, image correction processing of the image to be processed to obtain a positive image includes: rotating the image to be processed multiple times within a preset angle range with a preset rotation angle step size; forming a pixel variance group based on the image pixel variance of the image to be processed after each rotation; and rotating the image to be processed by a correction angle to obtain a positive image, wherein the correction angle is the rotation angle corresponding to the maximum pixel variance in the pixel variance group.

[0009] According to some embodiments of this application, forming a pixel variance group based on the image pixel variance of the image to be processed after each rotation includes: determining the sum of row pixels or column pixels of the image to be processed after each rotation; calculating the variance of the sum of row pixels of all rows of the image to be processed to obtain the image pixel variance, or calculating the variance of the sum of column pixels of all columns of the image to be processed to obtain the image pixel variance; and forming a pixel variance group from all image pixel variances.

[0010] According to some embodiments of this application, the image data information includes text naming information and text position information. Based on a preset algorithm, the image data information is matched and extracted to obtain document information with structured data, including: determining multiple lines of text information of the document information based on the text naming information and text position information; determining the first line of text information of the multiple lines of text information based on a first preset matching rule; determining the line order information of the multiple lines of text information based on a second preset matching rule; and determining the document information based on the first line of text information and the line order information.

[0011] According to some embodiments of this application, determining the first line of text information among multiple lines of text information based on a first preset matching rule includes: determining the line of text information with the highest line matching degree with the preset first line of text information among the multiple lines of text information as the first line of text information; the formula for calculating the line matching degree is: P r = 2.0*M1 / T1, where P r M1 represents the line matching degree, where M1 is the number of characters in the intersection of the line text information and the string in the preset first line text information, and T1 is the sum of the lengths of the lines in the line text information and the string in the preset first line text information.

[0012] According to some embodiments of this application, determining the row order information of multiple row text information based on a second preset matching rule includes: sequentially determining the row text information with the highest column matching degree among the multiple row text information; the formula for calculating the column matching degree is: P c = 2.0 * M² / T², where P c For column matching degree, M2 is the number of characters in the intersection of the row text information and the string in the preset column text information, and T2 is the sum of the lengths of the strings in the row text information and the preset column text information.

[0013] According to some embodiments of this application, after extracting image data information by data matching based on a preset algorithm to obtain document information with structured data, the image recognition method further includes: determining that there is missing text information in the document information, then extracting the image of the adjacent field of the missing text information; and performing secondary extraction on the image of the adjacent field through optical character recognition to obtain document information including the missing text information.

[0014] Another aspect of this application provides an image recognition system. The image recognition system includes an image correction unit, a data recognition unit, and an image processing unit. The image correction unit performs image correction processing on the image to be processed to obtain a forward image; the data recognition unit performs image recognition on the forward image using optical character recognition to obtain image data information; and the image processing unit performs data matching and extraction on the image data information based on a preset algorithm to obtain document information with structured data.

[0015] According to another aspect of this application, an electronic device is provided. The electronic device includes one or more processors and a storage device for storing one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement the image recognition method described above.

[0016] According to another aspect of this application, a non-volatile computer-readable storage medium is provided. The storage medium stores a computer program that can implement the image recognition method described above.

[0017] This application obtains a frontal image by performing image correction processing on the image to be processed, and then performs image recognition on the frontal image through optical character recognition to obtain image data information. Based on a preset algorithm, data matching and extraction are performed on the image data information to obtain document information with structured data. This application can automatically perform image tilt correction and structured tabular data content recognition based on machine learning algorithms, enabling rapid image data entry and improving the efficiency of digital work. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A schematic diagram showing an example embodiment of the image recognition method of this application is illustrated.

[0020] Figure 2 A schematic diagram showing the image to be processed in an example embodiment of this application;

[0021] Figure 3 Another schematic diagram showing an example embodiment of the image recognition method of this application;

[0022] Figure 4 A schematic diagram showing the tilt angle versus image pixel variance in an example embodiment of this application;

[0023] Figure 5 A schematic diagram showing a frontal image of an example embodiment of this application;

[0024] Figure 6 Another schematic diagram showing an example embodiment of the image recognition method of this application;

[0025] Figure 7 A schematic diagram illustrating document information of an example embodiment of this application;

[0026] Figure 8 Another schematic diagram showing an example embodiment of the image recognition method of this application;

[0027] Figure 9 A schematic diagram of an image recognition system according to an example embodiment of this application is shown.

[0028] Explanation of reference numerals in the attached figures:

[0029] Image recognition system 1; image correction unit 10; data recognition unit 20; image processing unit 30. Detailed Implementation

[0030] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that this application will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.

[0031] The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of these specific details, or other methods, components, materials, devices, etc. In these cases, well-known structures, methods, devices, implementations, materials, or operations will not be shown or described in detail.

[0032] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.

[0033] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, rather than to describe a specific order.

[0034] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0035] According to one aspect of this application, an image recognition method is provided. This image recognition method can automatically perform image tilt correction and structured tabular data content recognition based on machine learning algorithms, which can quickly complete image data entry and improve the efficiency of digital work.

[0036] The present application will be described in detail below with reference to the accompanying drawings.

[0037] Figure 1 A schematic diagram illustrating an example embodiment of the image recognition method of this application is shown, such as... Figure 1 As shown, the image recognition method includes steps S100-S300. Exemplarily, the image recognition method can be executed by an image recognition system.

[0038] According to the example embodiment, in step S100, the image recognition system performs image correction processing on the image to be processed to obtain a positive image.

[0039] For example, the image to be processed can be a target image that requires digital processing. Exemplarily, the target image can be a medical test report image, which is obtained by photographing or scanning a medical test report. In the following description, this application uses a medical test report image as an example for illustration.

[0040] Figure 2 A schematic diagram showing an image to be processed according to an example embodiment of this application.

[0041] like Figure 2 As shown, the image to be processed is a medical test report. Due to operational errors during the photography or scanning process, the image to be processed has a certain tilt angle. In order to accurately obtain the correspondence between the contextual content and position in the target table during image recognition, it is necessary to correct the tilt of the image to be processed to obtain an upright image with a positive angle.

[0042] Figure 3 Another schematic diagram illustrating an image recognition method according to an example embodiment of this application is shown.

[0043] Optionally, such as Figure 3 As shown, step S100, in which the image recognition system performs image correction processing on the image to be processed, may include steps S110-S130.

[0044] In step S110, the image recognition system rotates the image to be processed multiple times within a preset angle range with a preset rotation angle step size.

[0045] For example, the preset angle range is the angle range set by the user according to actual needs, and the preset rotation angle step is also the rotation angle step set by the user according to actual needs. For example, the preset angle range can be -30° to 30°, and the preset rotation angle step can be 0.5°, but this application is not limited to this.

[0046] If an image recognition system rotates the image to be processed multiple times within an angle range of -30° to 30° with a rotation angle step of 0.5°, it can obtain 120 rotated images to be processed.

[0047] In step S120, the image recognition system forms a pixel variance group based on the image pixel variance of the image to be processed after each rotation.

[0048] For example, each rotated image has a pixel variance, and the pixel variances of the 120 rotated images are grouped into a pixel variance group.

[0049] In step S130, the image recognition system rotates and corrects the image to be processed by an angle to obtain a positive image, wherein the correction angle is the rotation angle corresponding to the maximum pixel variance in the pixel variance group.

[0050] Figure 4 This is a schematic diagram showing the tilt angle versus image pixel variance in an example embodiment of this application. Figure 4 The horizontal axis represents the image tilt angle, and the vertical axis represents the image pixel variance. According to... Figure 4 It can be seen that the image pixel variance is inversely proportional to the image tilt angle. Therefore, the rotation angle between the rotated image to be processed and the original image to be processed, corresponding to the largest pixel variance in the pixel variance group, is the correction angle that needs to be adjusted.

[0051] Figure 5 A schematic diagram showing a frontal image of an example embodiment of this application, such as... Figure 5 As shown, the image recognition system rotates and corrects the angle of the image to be processed to obtain the image as shown. Figure 5 The image shown is a positive image.

[0052] Optionally, in step S120, the image recognition system determines the sum of row pixels or column pixels of the image to be processed after each rotation, calculates the variance of the sum of row pixels of all rows of the image to be processed to obtain the image pixel variance, or calculates the variance of the sum of column pixels of all columns of the image to be processed to obtain the image pixel variance.

[0053] According to the example embodiment, the image pixel variance is obtained based on the row pixels or column pixels of the image to be processed after each rotation.

[0054] For example, taking row pixels as an example, the image recognition system sums the pixels of each row of the image to be processed after each rotation, and then calculates the variance of the sum of pixels of all rows, thereby obtaining the image pixel variance of the image to be processed.

[0055] According to the example embodiment, in step S200, the image recognition system performs image recognition on the front image through optical character recognition to obtain image data information.

[0056] For example, the image recognition system uses an OCR (Optical Character Recognition) neural network to perform image recognition on the forward image obtained in step S100, thereby obtaining image data information on the forward image.

[0057] Optionally, the image data information includes text naming information and text location information.

[0058] For example, an image recognition system can identify phrase information in a frontal image, as well as the location information of the phrase information in the frontal image.

[0059] According to the example embodiment, in step S300, image data information is extracted by data matching based on a preset algorithm to obtain document information with structured data.

[0060] For example, after identifying relevant information about phrases in a forward-facing image, the image recognition system extracts data from the phrases using a pre-defined algorithm and matches the extracted data to corresponding document information with structured data. For instance, matching the extracted data with corresponding positions in a table yields a table that includes image data, thus completing the digitization of the image.

[0061] Figure 6 Another schematic diagram illustrating an image recognition method according to an example embodiment of this application is shown.

[0062] Optionally, such as Figure 6 As shown, step S300 may also include steps S310-340.

[0063] In step S310, the image recognition system determines multiple lines of text information of the document information based on the text naming information and the text position information.

[0064] For example, an image recognition system can use the recognition results of an OCR (Optical Character Recognition) neural network to calculate the rightmost adjacent word and the leftmost adjacent word within a certain range for each word group, thereby determining the line text information of each row in the target table.

[0065] In step S320, the image recognition system determines the first line of text information of multiple lines of text information based on the first preset matching rule.

[0066] For example, after determining the row text information of each row in the target table, the first row text information (i.e., the table header information) is determined from the multiple rows of text information.

[0067] Optionally, the image recognition system determines the first line of text information as the line with the highest line matching degree among multiple lines of text information.

[0068] Users can set a preset first line of text information according to their actual needs. The image recognition system matches multiple lines of text information according to the preset code algorithm, and determines the line of text information with the highest matching degree with the preset first line of text information as the table header information.

[0069] The formula for calculating the row matching degree is as follows:

[0070] P r =2.0*M1 / T1

[0071] Among them, P r M1 represents the line matching degree, where M1 is the number of characters in the intersection of the line text information and the string in the preset first line text information, and T1 is the sum of the lengths of the lines in the line text information and the string in the preset first line text information.

[0072] In step S330, the image recognition system determines the line order information of multiple lines of text information based on the second preset matching rule.

[0073] Optionally, the image recognition system sequentially determines the row of text information that has the highest column matching degree with the preset column text information from multiple rows of text information.

[0074] For example, after determining the header information in the target table, the table item information is located based on the subfield of the header information. Preset column text information includes the table item information for each pre-defined row.

[0075] Users can set preset column text information, including the table entry information for each row, according to their actual needs. The image recognition system matches multiple rows of text information based on preset codes and determines the row text information with the highest column matching degree to each row's table entry information as the row text information for that row. Based on the same principle, the row order of each row's row text information can be determined.

[0076] The formula for calculating column matching degree is as follows:

[0077] P c =2.0*M2 / T2

[0078] Among them, P c For column matching degree, M2 is the number of characters in the intersection of the row text information and the string in the preset column text information, and T2 is the sum of the lengths of the strings in the row text information and the preset column text information.

[0079] In step S340, the image recognition system determines the document information based on the first line of text information and the line order information.

[0080] Figure 7 A schematic diagram illustrating document information of an example embodiment of this application.

[0081] For example, once an image recognition system determines the table header information and the table entries for each row, it can fully identify the table content. For each cell, by searching for its nearest header and entry information, the system can determine the row and column of the current cell, thus locating the table's position. Figure 7 The table information obtained after image recognition processing is shown.

[0082] Figure 8 Another schematic diagram illustrating an image recognition method according to an example embodiment of this application is shown.

[0083] Optionally, such as Figure 8 As shown, the image recognition method includes steps S100-S500. Steps S100-S300 have been described in detail above and will not be repeated here.

[0084] In step S400, if the image recognition system determines that there is missing text information in the document information, it will extract the image of the adjacent domain of the missing text information.

[0085] In step S500, the image recognition system performs secondary extraction on the images of adjacent fields using optical character recognition to obtain document information including missing text information.

[0086] For example, in table information, if there is missing cell information, its approximate location can be determined based on the adjacent positions of that cell. For instance, the images of the items above, below, to the left, and to the right (such as P1, P2, P3, P4) of the cell (P0) can be extracted, and the area enclosed by P1, P2, P3, P4 can be determined as the adjacent domain.

[0087] By using an OCR neural network to perform image recognition on the adjacent regions, the processing scope of the image can be narrowed, OCR processing noise can be reduced, and the completeness of image recognition can be improved. Through secondary extraction by the OCR neural network, missing text information can be identified.

[0088] Through the above example embodiments, this application can automatically perform image tilt correction and structured table content recognition based on machine learning algorithms, which can quickly complete image data entry and improve the efficiency of digital work.

[0089] According to another aspect of this application, an image recognition system is provided. Figure 9 A schematic diagram of an image recognition system according to an example embodiment of this application is shown. Figure 9 As shown, the image recognition system 1 includes an image correction unit 10, a data recognition unit 20, and an image processing unit 30. The image recognition system 1 is used to perform the image recognition method described above.

[0090] According to an example embodiment, the image correction unit 10 performs image correction processing on the image to be processed to obtain a positive image.

[0091] For example, the image to be processed can be a target image that requires digital processing. Exemplarily, the target image can be a medical test report image, which is obtained by photographing or scanning a medical test report. In the following description, this application uses a medical test report image as an example for illustration.

[0092] Due to operational errors during the photography or scanning process, the image to be processed may have a certain tilt angle. In order to accurately obtain the correspondence between the contextual content and position in the target table during image recognition, it is necessary to correct the tilt of the image to obtain an upright image with a positive angle.

[0093] Optionally, the image correction unit 10 performs multiple image rotations on the image to be processed within a preset angle range with a preset rotation angle step.

[0094] For example, the preset angle range is the angle range set by the user according to actual needs, and the preset rotation angle step is also the rotation angle step set by the user according to actual needs. For example, the preset angle range can be -30° to 30°, and the preset rotation angle step can be 0.5°, but this application is not limited to this.

[0095] If the image correction unit 10 rotates the image to be processed multiple times within the angle range of -30° to 30° with a rotation angle step of 0.5°, 120 rotated images to be processed can be obtained.

[0096] The image correction unit 10 forms a pixel variance group based on the image pixel variance of the image to be processed after each rotation.

[0097] For example, each rotated image has a pixel variance, and the pixel variances of the 120 rotated images are grouped into a pixel variance group.

[0098] The image correction unit 10 rotates and corrects the image to be processed by an angle to obtain a positive image, wherein the correction angle is the rotation angle corresponding to the maximum pixel variance in the pixel variance group.

[0099] Figure 4 This is a schematic diagram showing the tilt angle versus image pixel variance in an example embodiment of this application. Figure 4 The horizontal axis represents the image tilt angle, and the vertical axis represents the image pixel variance. According to... Figure 4 It can be seen that the image pixel variance is inversely proportional to the image tilt angle. Therefore, the rotation angle between the rotated image to be processed and the original image to be processed, corresponding to the largest pixel variance in the pixel variance group, is the correction angle that needs to be adjusted.

[0100] The image correction unit 10 determines the sum of row pixels or column pixels of the image to be processed after each rotation, calculates the variance of the sum of row pixels of all rows of the image to be processed to obtain the image pixel variance, or calculates the variance of the sum of column pixels of all columns of the image to be processed to obtain the image pixel variance.

[0101] According to the example embodiment, the image pixel variance is obtained based on the row pixels or column pixels of the image to be processed after each rotation.

[0102] For example, taking row pixels as an example, the image correction unit 10 sums the pixels of each row of the image to be processed after each rotation, and then calculates the variance of the sum of pixels of all rows, thereby obtaining the image pixel variance of the image to be processed.

[0103] According to the example embodiment, the data recognition unit 20 performs image recognition on the forward image through optical character recognition to obtain image data information.

[0104] For example, the data recognition unit 20 performs image recognition on the obtained positive image through an OCR (Optical Character Recognition) neural network to obtain image data information on the positive image.

[0105] Optionally, the image data information includes text naming information and text location information.

[0106] For example, an image recognition system can identify phrase information in a frontal image, as well as the location information of the phrase information in the frontal image.

[0107] According to the example embodiment, the image processing unit 30 performs data matching and extraction on image data information based on a preset algorithm to obtain document information with structured data.

[0108] For example, after determining the relevant information of phrase information in the forward image, the image processing unit 30 extracts the phrase information according to a preset code algorithm and matches the extracted data to the corresponding document information with structured data. For example, matching the extracted data with the corresponding positions in a table yields a table that includes image data information, thus completing the digital processing of the image to be processed.

[0109] Optionally, the image processing unit 30 determines multiple lines of text information of the document information based on the text naming information and the text position information.

[0110] For example, the image processing unit 30 calculates the rightmost adjacent word group closest to the end of each word group and the leftmost adjacent word group closest to the beginning of the word group within a certain upper and lower range based on the recognition results of the OCR (Optical Character Recognition) neural network. In this way, the line text information of each row in the target table can be determined.

[0111] The image processing unit 30 determines the first line of text information of multiple lines of text information based on the first preset matching rule.

[0112] For example, after determining the row text information of each row in the target table, the first row text information (i.e., the table header information) is determined from the multiple rows of text information.

[0113] Optionally, the image processing unit 30 determines the line of text information with the highest line matching degree among multiple lines of text information as the first line of text information.

[0114] If the user can set a preset first line text information according to actual needs, the image processing unit 30 will match multiple lines of text information according to the preset code algorithm, and determine the line text information with the highest matching degree with the preset first line text information as the table header information.

[0115] The formula for calculating the row matching degree is as follows:

[0116] P r =2.0*M1 / T1

[0117] Among them, P r M1 represents the line matching degree, where M1 is the number of characters in the intersection of the line text information and the string in the preset first line text information, and T1 is the sum of the lengths of the lines in the line text information and the string in the preset first line text information.

[0118] The image processing unit 30 determines the line order information of multiple lines of text information based on the second preset matching rule.

[0119] Optionally, the image processing unit 30 sequentially determines the row text information with the highest column matching degree among multiple row text information.

[0120] For example, after determining the header information in the target table, the table item information is located based on the subfield of the header information. Preset column text information includes the table item information for each pre-defined row.

[0121] Users can set preset column text information including table entry information for each row according to actual needs. The image processing unit 30 matches multiple rows of text information according to preset codes and determines the row text information with the highest column matching degree with the table entry information of each row as the row text information of that row. Based on the same principle, the row order of the row text information of each row can be determined.

[0122] The formula for calculating column matching degree is as follows:

[0123] P c =2.0*M2 / T2

[0124] Among them, P c For column matching degree, M2 is the number of characters in the intersection of the row text information and the string in the preset column text information, and T2 is the sum of the lengths of the strings in the row text information and the preset column text information.

[0125] The image processing unit 30 determines the document information based on the first line of text information and the line order information.

[0126] For example, once the image processing unit 30 determines the header information and the table item information for each row, it can fully identify the table content information. For the position of each cell, by searching for its nearest header and table item information, the row and column of the current cell can be determined, that is, the position of the table can be located.

[0127] Optionally, if the image processing unit 30 determines that there is missing text information in the document information, it extracts the image of the adjacent domain of the missing text information.

[0128] The image processing unit 30 performs secondary extraction on the image of adjacent fields using optical character recognition to obtain document information including missing text information.

[0129] For example, in table information, if there is missing cell information, its approximate location can be determined based on the adjacent positions of that cell. For instance, the images of the items above, below, to the left, and to the right (such as P1, P2, P3, P4) of the cell (P0) can be extracted, and the area enclosed by P1, P2, P3, P4 can be determined as the adjacent domain.

[0130] By using an OCR neural network to perform image recognition on the adjacent regions, the processing scope of the image can be narrowed, OCR processing noise can be reduced, and the completeness of image recognition can be improved. Through secondary extraction by the OCR neural network, missing text information can be identified.

[0131] Through the above example embodiments, this application can automatically perform image tilt correction and structured table content recognition based on machine learning algorithms, which can quickly complete image data entry and improve the efficiency of digital work.

[0132] According to another aspect of this application, an electronic device is provided. The electronic device includes one or more processors and a storage device for storing one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement the image recognition method described above.

[0133] According to another aspect of this application, a non-volatile computer-readable storage medium is provided. The storage medium stores a computer program that can implement the image recognition method described above.

[0134] Finally, it should be noted that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions of the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An image recognition method, characterized in that, include: Image correction processing is performed on the image to be processed to obtain a positive image; Image recognition is performed on the positive image using optical character recognition to obtain image data information, which includes text naming information and text location information. Based on a preset algorithm, data matching and extraction are performed on the image data to obtain document information with structured data, including: Determine multiple lines of text information of the document information based on the text naming information and the text position information, including: calculating the right adjacent word group closest to the end of each word group and the left adjacent word group closest to the beginning of each word group within a certain upper and lower range, so as to determine the line text information of each row in the target table; Determining the first line of text information of the plurality of lines of text information based on a first preset matching rule includes: determining the line of text information with the highest line matching degree with the preset first line of text information among the plurality of lines of text information as the first line of text information; Determining the row order information of the multiple rows of text information based on the second preset matching rule includes: sequentially determining the row of text information with the highest column matching degree with the preset column text information among the multiple rows of text information; The document information is determined based on the first line of text information and the line order information.

2. The image recognition method according to claim 1, characterized in that, The process of performing image correction on the image to be processed to obtain a positive image includes: The image to be processed is rotated multiple times within a preset angle range with a preset rotation angle step size. A pixel variance group is formed based on the image pixel variance of the image to be processed after each rotation; The image to be processed is rotated and corrected by an angle to obtain the positive image, wherein the correction angle is the rotation angle corresponding to the maximum pixel variance in the pixel variance group.

3. The image recognition method according to claim 2, characterized in that, The step of forming a pixel variance group based on the image pixel variance of the image to be processed after each rotation includes: Determine the sum of row or column pixels of the image to be processed after each rotation; The image pixel variance is obtained by calculating the variance of the row pixel sum of all rows of the image to be processed, or by calculating the variance of the column pixel sum of all columns of the image to be processed. The pixel variances of all images are grouped into the pixel variance group.

4. The image recognition method according to claim 1, characterized in that, The formula for calculating the row matching degree is: P r =2.0 M1 / T1 Among them, P r The line matching degree is M1, where M1 is the number of characters in the intersection of the line text information and the string in the preset first line text information, and T1 is the sum of the lengths of the line text information and the string in the preset first line text information.

5. The image recognition method according to claim 1, characterized in that, The formula for calculating the column matching degree is: P c =2.0 M2 / T2 Among them, P c M2 is the column matching degree, M2 is the number of characters in the intersection of the row text information and the string in the preset column text information, and T2 is the sum of the lengths of the row text information and the string in the preset column text information.

6. The image recognition method according to claim 1, characterized in that, After performing data matching and extraction on the image data information based on a preset algorithm to obtain document information with structured data, the image recognition method further includes: If it is determined that there is missing text information in the document information, then the image of the adjacent field of the missing text information is extracted; The adjacent domain images are extracted a second time using optical character recognition to obtain document information including the lost text information.

7. An image recognition system, characterized in that, The image recognition system is used to perform the image recognition method as described in any one of claims 1-6, and the image recognition system includes: The image correction unit performs image correction processing on the image to be processed to obtain a positive image; The data recognition unit performs image recognition on the frontal image through optical character recognition to obtain image data information, which includes text naming information and text location information. The image processing unit performs data matching and extraction on the image data information based on a preset algorithm to obtain document information with structured data, including: determining multiple lines of text information of the document information based on the text naming information and the text position information, including: calculating the right adjacent word group closest to the end of each word group and the left adjacent word group closest to the beginning of each word group within a certain upper and lower range to determine the line text information of each row in the target table; determining the first line of text information of the multiple lines of text information based on a first preset matching rule, including: determining the line text information with the highest line matching degree with the preset first line text information among the multiple lines of text information as the first line of text information; determining the row order information of the multiple lines of text information based on a second preset matching rule, including: sequentially determining the line text information with the highest column matching degree with the preset column text information among the multiple lines of text information; and determining the document information based on the first line of text information and the row order information.

8. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image recognition method as described in any one of claims 1-6.

9. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program implements the image recognition method as described in any one of claims 1-6.