Method and device for generating structured data
By performing text recognition and structural processing on bill image files to generate structured data, the problem of inaccurate recognition caused by bill information clarity and image stretching is solved, and data processing efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202210367829.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-04-08
AI Technical Summary
In the prior art, the automatic entry of bill information results in inaccurate recognition due to bill information clarity and image stretching issues, resulting in low data processing efficiency and prone to human errors.
By performing text recognition on image files, multiple recognition units are generated, and the data type is determined based on the text information and position information of the recognition unit. A suitable structured template is selected, and the recognition unit is structured, including merging, splitting, and deleting invalid units to generate structured data.
It improves data processing efficiency, reduces errors caused by manual processing, and ensures data accuracy and processing efficiency.
Smart Images

Figure CN114913537B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing technology, and in particular to a method and apparatus for generating structured data, a computer-readable storage medium, an electronic device, and a computer program product. Background Art
[0002] Currently, some information technology companies' daily data processing involves large volumes of orders, invoices, and other documents. To facilitate data processing and querying, companies need to assign staff to extract information from these large volumes of orders, invoices, and other documents and enter this information into a database. This data processing or importing method results in low data processing efficiency and can lead to data errors caused by manual processing.
[0003] To this end, existing technologies require automated entry of bill information, such as orders and invoices, through algorithms. However, the bill information analyzed by text recognition technology is often just scattered text box coordinate information. Furthermore, due to issues such as the bill's inherent clarity and image stretching, the recognition results are often inaccurate. Summary of the Invention
[0004] In view of this, the present invention proposes a method and device for generating structured data, a computer-readable storage medium, an electronic device, and a computer program product, aiming to solve the technical problem of inaccurate data when generating structured data from bill information due to the clarity of the bill information itself and image stretching.
[0005] According to one aspect of an embodiment of the present disclosure, a method for generating structured data is provided, comprising:
[0006] Performing text recognition on the image file to be processed to obtain a plurality of recognition units, wherein each recognition unit includes: position information and text information;
[0007] determining a data type of the image file to be processed based on the text information of the plurality of recognition units;
[0008] selecting a target structured template suitable for the image file to be processed from a plurality of structured templates according to the data type; and
[0009] The plurality of recognition units are subjected to structural processing based on one or more of the target structured template, the position information, and the text information, and structured data are generated according to the structurally processed recognition units.
[0010] Preferably, the performing structural processing on the plurality of recognition units based on one or more of the target structured template, position information, and text information includes:
[0011] Determining, from the plurality of recognition units, a recognition unit belonging to the format content according to the position information and / or text information of each format content in the target structured template; and
[0012] The identification units belonging to the format content among the multiple identification units are deleted.
[0013] Preferably, the performing structural processing on the plurality of recognition units based on one or more of the target structured template, position information, and text information includes:
[0014] determining the length and width of each identification unit based on the corresponding position information;
[0015] determining the number of characters in each recognition unit based on the corresponding text information;
[0016] Determining a recognition unit whose width is greater than its length and whose number of characters is greater than a preset threshold as a recognition unit belonging to the format content in the target structured template; and
[0017] The identified recognition units belonging to the format content are deleted.
[0018] Preferably, the performing structural processing on the plurality of recognition units based on one or more of the target structured template, position information, and text information includes:
[0019] determining the number of characters in each recognition unit based on the corresponding text information;
[0020] Delete the recognition units with zero characters.
[0021] Preferably, the performing structural processing on the plurality of recognition units based on one or more of the target structured template, position information, and text information includes:
[0022] Detect the distances between all adjacent characters in each recognition unit in sequence;
[0023] When it is detected that the distance between any adjacent characters is greater than the first character spacing threshold, the recognition units are split based on the adjacent characters whose distance is greater than the first character spacing threshold until the distance between all adjacent characters in each recognition unit is less than or equal to the first character spacing threshold.
[0024] Preferably, the performing structural processing on the plurality of recognition units based on one or more of the target structured template, position information, and text information includes:
[0025] Determine the distance between two adjacent recognition units;
[0026] When the distance between any two adjacent recognition units is less than the second character spacing threshold, the two adjacent recognition units are merged into one recognition unit.
[0027] Preferably, the performing structural processing on the plurality of recognition units based on one or more of the target structured template, position information, and text information includes:
[0028] selecting a reference position identification unit from the plurality of identification units;
[0029] Determine a first recognition unit adjacent to the reference position recognition unit in a transverse position and a second recognition unit adjacent to the reference position recognition unit in a longitudinal position;
[0030] determining a text inclination of the image file to be processed based on the reference position recognition unit, the first recognition unit, and the second recognition unit;
[0031] Based on the text inclination, a row and column position relationship of the plurality of recognition units in the image file to be processed is determined.
[0032] Preferably, the determining the row and column position relationship of the plurality of recognition units based on the text inclination includes:
[0033] generating a horizontal reference line and a vertical reference line based on the text inclination;
[0034] Determining an identification unit as a reference row among the plurality of identification units based on the horizontal reference line, and determining an average value of spacings between adjacent columns according to the identification unit as the reference row;
[0035] Determining recognition units as reference columns among the plurality of recognition units based on the vertical reference line, and determining an average value of spacings between adjacent rows according to the recognition units as reference columns; and
[0036] Based on the average distance between adjacent columns and the average distance between adjacent rows, the row and column positional relationship of the other recognition units except the recognition units in the reference row and the recognition units in the reference column is determined.
[0037] Preferably, the step of performing text recognition on the image file to be processed to obtain a plurality of recognition units includes:
[0038] Segmenting the text content in the image file to be processed to obtain a plurality of character lines;
[0039] Performing character segmentation on each character row to obtain a plurality of recognition units, and determining position information of the recognition units according to coordinates of the recognition units in the image file to be processed; and
[0040] Character recognition is performed on the characters in each recognition unit, so that the characters obtained through character recognition are used as text information of the recognition unit.
[0041] Preferably, the determining the data type of the image file to be processed based on the text information of the plurality of recognition units includes:
[0042] Searching the text information of the plurality of recognition units to determine keywords associated with the type;
[0043] The data type of the image file to be processed is determined based on the keyword associated with the type.
[0044] Preferably, the method of determining the data type of the image file to be processed based on the position information of the plurality of identification units further comprises:
[0045] Determine the number of recognition units;
[0046] Determining size information of each recognition unit based on the position information of the recognition unit;
[0047] The data type of the image file to be processed is determined based on the number of recognition units and size information of each recognition unit.
[0048] Preferably, the step of selecting a target structured template suitable for the image file to be processed from a plurality of structured templates according to the data type includes:
[0049] matching the data type with the data type of each structured template in a plurality of structured templates to determine a degree of match with each structured template; and
[0050] The structured module with the largest matching degree is selected as the target structured template suitable for the image file to be processed.
[0051] According to one aspect of an embodiment of the present disclosure, there is provided an apparatus for generating structured data, including:
[0052] A recognition unit, configured to perform text recognition on the image file to be processed to obtain a plurality of recognition units, wherein each recognition unit includes: position information and text information;
[0053] a determining unit, configured to determine a data type of the image file to be processed based on the text information of the plurality of recognition units;
[0054] a selecting unit, configured to select a target structured template suitable for the image file to be processed from a plurality of structured templates according to the data type; and
[0055] A processing unit is used to perform structural processing on the multiple recognition units based on one or more of the target structured template, position information, and text information, and generate structured data according to the recognition units that have undergone structural processing.
[0056] Preferably, the processing unit includes:
[0057] a first determining subunit, configured to determine, from the plurality of identification units, an identification unit belonging to the format content according to the position information and / or text information of each format content in the target structured template; and
[0058] The first deleting subunit is configured to delete the recognition units belonging to the format content from the plurality of recognition units.
[0059] Preferably, the processing unit further includes:
[0060] a second determining subunit, configured to determine the length and width of each identification unit based on the corresponding position information;
[0061] a third determining subunit, configured to determine the number of characters in each recognition unit based on corresponding text information;
[0062] a fourth determining subunit, configured to determine a recognition unit whose width is greater than its length and whose number of characters is greater than a preset threshold as a recognition unit belonging to the format content in the target structured template; and
[0063] The second deleting subunit is used to delete the identified identification unit belonging to the format content.
[0064] Preferably, the processing unit further includes:
[0065] a fifth determining subunit, configured to determine the number of characters in each recognition unit based on corresponding text information;
[0066] The third deleting subunit is used to delete the recognition unit whose number of characters is zero.
[0067] Preferably, the processing unit further includes:
[0068] A detection subunit, used to sequentially detect the distances between all adjacent characters in each recognition unit;
[0069] A processing subunit is used to split the recognition unit based on the adjacent characters whose distance is greater than the first character spacing threshold when it is detected that the distance between any adjacent characters is greater than the first character spacing threshold, until the distance between all adjacent characters in each recognition unit is less than or equal to the first character spacing threshold.
[0070] Preferably, the processing unit further includes:
[0071] a sixth determining subunit, configured to determine the distance between any two adjacent recognition units;
[0072] The merging subunit is configured to merge any two adjacent recognition units into one recognition unit when the distance between the two adjacent recognition units is less than a second character spacing threshold.
[0073] Preferably, the processing unit further includes:
[0074] a first selection subunit, configured to select a reference position recognition unit from the plurality of recognition units;
[0075] a seventh determining subunit, configured to determine a first identification unit adjacent to the reference position identification unit in a transverse position and a second identification unit adjacent to the reference position identification unit in a longitudinal position;
[0076] an eighth determining subunit, configured to determine a text inclination of the image file to be processed based on the reference position identifying unit, the first identifying unit, and the second identifying unit;
[0077] The ninth determining subunit is configured to determine, based on the text inclination, a row and column position relationship of the plurality of recognition units in the image file to be processed.
[0078] The ninth determining subunit is specifically configured to:
[0079] generating a horizontal reference line and a vertical reference line based on the text inclination;
[0080] Determining an identification unit as a reference row among the plurality of identification units based on the horizontal reference line, and determining an average value of spacings between adjacent columns according to the identification unit as the reference row;
[0081] Determining recognition units as reference columns among the plurality of recognition units based on the vertical reference line, and determining an average value of spacings between adjacent rows according to the recognition units as reference columns; and
[0082] Based on the average distance between adjacent columns and the average distance between adjacent rows, the row and column positional relationship of the other recognition units except the recognition units in the reference row and the recognition units in the reference column is determined.
[0083] Preferably, the identification unit includes:
[0084] a segmentation subunit, configured to segment the text content in the image file to be processed, thereby obtaining a plurality of character lines;
[0085] a tenth determining subunit, configured to segment each character row to obtain a plurality of recognition units, and determine position information of the recognition units according to coordinates of the recognition units in the image file to be processed; and
[0086] The recognition subunit is used to perform character recognition on the characters in each recognition unit, so as to use the characters obtained through character recognition as text information of the recognition unit.
[0087] Preferably, the determining unit is specifically configured to: search the text information of the plurality of recognition units to determine keywords associated with the type; and determine the data type of the image file to be processed based on the keywords associated with the type.
[0088] Preferably, the determination unit is further configured to: determine the number of recognition units; determine the size information of each recognition unit based on the position information of the recognition units; and determine the data type of the image file to be processed based on the number of recognition units and the size information of each recognition unit.
[0089] Preferably, the selection unit includes:
[0090] a matching subunit, configured to match the data type with the data type of each structured template in a plurality of structured templates to determine a degree of matching with each structured template; and
[0091] The second selection subunit is configured to select a structured module with the greatest matching degree as a target structured template suitable for the image file to be processed.
[0092] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute the method described in any one of the above embodiments.
[0093] According to another aspect of the embodiments of the present disclosure, an electronic device is provided, including:
[0094] processor;
[0095] a memory for storing instructions executable by the processor;
[0096] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method described in any one of the above embodiments.
[0097] According to another aspect of the embodiments of the present disclosure, a computer program product is provided, including computer-readable code. When the computer-readable code is run on a device, a processor in the device executes a method for implementing any of the above embodiments.
[0098] The method and apparatus for generating structured data, computer-readable storage medium, electronic device, and computer program product provided by the aforementioned embodiments of the present disclosure can, on the one hand, improve the efficiency of data processing or data import and avoid data errors caused by manual processing. On the other hand, they solve the technical problem of inaccurate data generated from bill information due to factors such as the clarity of the bill information itself and image stretching. By unifying the data structuring process for different bill templates, manual labor is efficiently replaced, costs are reduced, and data processing efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] The above and other objects, features, and advantages of the present invention will become more apparent through a more detailed description of the embodiments of the present invention in conjunction with the accompanying drawings. The accompanying drawings are provided to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and are not intended to limit the present invention. In the drawings, the same reference numerals generally represent the same components or steps.
[0100] Figure 1 is a flowchart of a method for generating structured data provided by an exemplary embodiment of the present disclosure;
[0101] Figure 2 is a schematic diagram of a bill example provided by an exemplary embodiment of the present disclosure;
[0102] Figure 3 is a schematic structural diagram of an apparatus for generating structured data provided by an exemplary embodiment of the present disclosure;
[0103] Figure 4 is a schematic diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0104] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0105] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0106] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meanings, nor do they indicate a necessary logical order between them.
[0107] It should also be understood that in the embodiments of the present disclosure, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.
[0108] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.
[0109] In addition, the term "and / or" in this disclosure is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.
[0110] It should also be understood that the description of the various embodiments in this disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.
[0111] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0112] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0113] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, they should be considered part of the specification.
[0114] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0115] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, among others.
[0116] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.
[0117] An embodiment of the present disclosure provides a method for generating structured data, including: performing text recognition on an image file to be processed to obtain multiple recognition units; determining the data type of the image file to be processed based on text information of the multiple recognition units; selecting a target structured template suitable for the image file to be processed from multiple structured templates according to the data type; and performing structured processing on the multiple recognition units based on one or more of the target structured template, position information, and text information, and generating structured data based on the structured recognition units.
[0118] Figure 1 FIG. 1 is a flow chart of a method for generating structured data provided by an exemplary embodiment of the present disclosure. Figure 1 As shown, the methods for generating structured data include:
[0119] Step 101 performs text recognition on the image file to be processed to obtain multiple recognition units, where each recognition unit includes: location information and text information. In practice, a large number of orders and invoices are obtained in the form of image files, for example, medical bills. Figure 2 FIG is a schematic diagram of a bill example provided by an exemplary embodiment of the present disclosure. Figure 2 As shown, a medical outpatient billing receipt in a certain city includes multiple data items, such as basic information, billing information, medical reimbursement information, and format information. Medical outpatient billing receipts are typically provided to information technology companies in the form of image files, and for various reasons, information technology companies are unable to obtain the structured data associated with these medical outpatient billing receipts.
[0120] In one embodiment, text recognition is performed on an image file to be processed to obtain a plurality of recognition units, including: segmenting the text content in the image file to be processed to obtain a plurality of character lines; segmenting each character line into characters to obtain a plurality of recognition units, determining the position information of the recognition unit according to the coordinates of the recognition unit in the image file to be processed; and performing character recognition on the characters in each recognition unit to use the characters obtained through character recognition as the text information of the recognition unit. Figure 2 As shown, for example, the text content of the medical outpatient billing data containing the bill code, electronic bill code, payee's unified social credit code, payee, item name, and western medicine fee is segmented into separate data lines to obtain multiple character lines. For example, all characters in the lines containing the bill code, electronic bill code, payee's unified social credit code, payee, item name, and western medicine fee can constitute a character line. Subsequently, each character line is segmented. For example, the character line containing the item name, quantity / unit, amount (yuan), remarks, item name, quantity / unit, amount (yuan), and remarks is segmented to obtain multiple recognition units, wherein each of the item name, quantity / unit, amount (yuan), remarks, item name, quantity / unit, amount (yuan), and remarks can be regarded as a recognition unit. Subsequently, the coordinates of each recognition unit among the item name, quantity / unit, amount (yuan), remarks, item name, quantity / unit, amount (yuan), and remarks in the image file of the medical outpatient billing data are determined, and the coordinates are used as the position information of the position recognition unit. Subsequently, character recognition is performed on the characters in each recognition unit, for example, character recognition is performed on the project name, quantity / unit, amount (yuan), and remarks, so that the characters obtained through character recognition are used as text information of the recognition unit.
[0121] Specifically, optical character recognition OCR (Optical Character Recognition) is used to analyze Figure 2 After the image file of the bill is processed through the OCR parsing interface, JSON data is obtained. The JSON data includes two fields: detect_box and text_info. The detect_box field stores the coordinate information of each text box identified in the bill. The text_info field stores the text information of each text box. Each text box can contain information such as top_k, position, text, score, bbox, and word_size.
[0122] Step 102 determines the data type of the image file to be processed based on the text information of the multiple recognition units. To improve the efficiency of structuring the data in the image file, different structuring processing templates are required for different bills. To this end, the data type of the image file to be processed must be determined, and a target structuring template suitable for the image file to be processed can be selected from multiple structuring templates based on the data type. The selected target structuring template improves the efficiency of structuring the data and effectively corrects errors during the recognition process.
[0123] In one embodiment, the data type of the image file to be processed is determined based on the text information of multiple recognition units, including: searching the text information of multiple recognition units to determine keywords associated with the type; and determining the data type of the image file to be processed based on the keywords associated with the type. Preferably, the keywords in the image file to be processed can reflect or indicate the specific category or data type of the ticket corresponding to the image file to be processed, for example, medical, parking, refueling, and other tickets. To this end, the present application performs a content search or keyword search on the text information from multiple recognition units of the image file to be processed, and determines the keywords associated with the type through the content search or keyword search. Keywords associated with the type usually have significant category characteristics, for example, medical, parking, refueling, etc. Therefore, the data type of the image file to be processed can be determined based on the keywords associated with the type, that is, when the keyword is medical, it can be determined that the data type of the image file to be processed is medical data or medical tickets.
[0124] In one embodiment, the data type of the image file to be processed is determined based on the position information of multiple recognition units: the number of recognition units is determined; the size information of each recognition unit is determined based on the position information of the recognition units; and the data type of the image file to be processed is determined based on the number of recognition units and the size information of each recognition unit. Preferably, different types of bills may have significantly different page layout, content layout, information item configuration, and other rules. Therefore, after performing text recognition on the image file to be processed, the recognition units obtained may have different position information, size information, and number information. To this end, the present application first determines the number of recognition units and the position information of each recognition unit, and then determines the size information of the recognition units based on the position information of the recognition units. Based on the size information of the recognition units, the length-to-width ratio of different text boxes can be determined. For example, because bills may use formatted data, some bills may have text boxes on the side with a width-to-length ratio greater than 5, while some bills may have text boxes on the bottom with a length-to-width ratio greater than 10. By identifying or matching the number of recognition units and the size information of each recognition unit, the data type of the image file to be processed can also be determined.
[0125] In actual applications, since there are a large number of bill templates, if matching calculations are performed for each template, the matching cost will be very high and the scalability will be very poor. Therefore, it is necessary to adopt a structured processing solution with strong versatility for multiple bill templates as much as possible. To this end, the embodiment of the present application first classifies the bills (image files to be identified) according to some features, and then summarizes and determines unified laws and rules for each category for unified processing. For example, the classification methods mainly include: classification based on keywords, and classification based on text box size and quantity. Among them, classification based on keywords is to divide categories based on the same keywords. For example, bills containing the keyword "medical" will be assigned to the same category or group (medical bill category). Classification based on text box size and quantity is to assign bills with similar information such as the number of text boxes and aspect ratio to the same category or group.
[0126] Step 103 selects a target structured template suitable for the image file to be processed from a plurality of structured templates based on the data type. After determining the data type of the image file to be processed based on the text information of the plurality of recognition units, a search or match is performed in the template library based on the data type, thereby selecting a target structured template suitable for the image file to be processed from the plurality of structured templates. For example, when the data type is medical data, a search or match is performed in the template library for structured templates for medical data, thereby selecting a structured template suitable for the image file to be processed from the plurality of structured templates. Furthermore, the data type can also be medical outpatient data, medical inpatient data, or a more detailed classification such as medical data from different regions.
[0127] Step 104: Structural processing is performed on the multiple recognition units based on one or more of the target structured template, location information, and text information, and structured data is generated based on the structured recognition units. After the preliminary structuring processing preparations have been completed, structural processing is performed on the multiple recognition units from the image file to be recognized, and structured data is generated based on the structured recognition units.
[0128] In one embodiment, a target structured template suitable for an image file to be processed is selected from a plurality of structured templates based on the data type, including: matching the data type with the data type of each structured template in the plurality of structured templates to determine the degree of matching with each structured template; and selecting the structured template with the greatest degree of matching as the target structured template suitable for the image file to be processed.
[0129] In one embodiment, during the process of structural processing of multiple recognition units, data in the text box (i.e., recognition unit) needs to be cleaned. Recognition technologies such as OCR will recognize all text information on the bill during the recognition process, so there will be some dirty data. For example: the vertical text on both sides of the bill (such as Figure 2 As shown, some useless information such as format information of the bill template, for example, printing by the Finance Bureau, receipt, etc., as well as some meaningless characters recognized due to unclear shooting or dirty bills. In addition, there may be no characters in the recognition unit recognized due to unclear shooting or dirty bills. The above contents will be recognized as text boxes by OCR. Therefore, this part of useless data needs to be filtered out. Specifically, the useless information on both sides of the bill is mainly filtered out through: 1) "Keywords", when keywords that often appear on both sides of the structured template appear in the text box, the text box will be deleted; 2) "Aspect ratio of text box", text boxes with a width greater than the length are judged as vertical boxes. However, simply judging the length and width size or the aspect ratio value may have misjudgment problems, so a restriction is added on the basis of "aspect ratio of text box", namely "number of text characters", for example, a text box with a width greater than the length and a number of characters greater than 2 will be deleted (in this case, it can be determined that the content in the text box is useless information such as format information). If the "number of text characters" limit is not increased, individual characters may be deleted, resulting in loss of information.
[0130] According to this method, multiple recognition units are structured based on one or more of the target structured template, position information, and text information, including: determining the recognition units belonging to the format content from the multiple recognition units based on the position information and / or text information of each format content in the target structured template; and deleting the recognition units belonging to the format content from the multiple recognition units. Additionally, multiple recognition units are structured based on one or more of the target structured template, position information, and text information, including: determining the length and width of each recognition unit based on the corresponding position information; determining the number of characters in each recognition unit based on the corresponding text information; determining the recognition units whose width is greater than the length and whose number of characters is greater than a preset threshold as recognition units belonging to the format content in the target structured template; and deleting the recognition units determined to belong to the format content. The text content or character content included in the recognition unit of the format content is information describing the bill type and information explaining precautions associated with the bill type.
[0131] In one embodiment, structural processing is performed on multiple recognition units based on one or more of a target structured template, location information, and text information, including: determining the number of characters in each recognition unit based on the corresponding text information; and deleting recognition units with zero character counts. As described above, due to unclear images or dirty bills, the recognition units may not contain any characters, so it is necessary to delete recognition units with zero character counts, thereby deleting recognition units with incorrect recognition.
[0132] In one embodiment, in the process of structurally processing multiple recognition units, it may be necessary to merge and split text boxes (i.e., recognition units). Recognition technologies such as OCR have two problems: 1) characters that should have existed in one text box are recognized as two text boxes, and 2) multiple characters that should have existed in two text boxes are recognized as existing in the same text box. The solution to these two problems is mainly to determine whether they belong to the same text box by the distance between the characters in the text box. For example, if the distance between adjacent characters in the same text box is greater than a first character spacing threshold, the text box should be split into two text boxes in the middle of the adjacent characters. Similarly, if the minimum distance between the character spacing between two adjacent text boxes is less than a second character spacing threshold, the two text boxes should be merged. According to the above method, all text boxes are traversed from beginning to end to complete the splitting and / or merging of text boxes.
[0133] In one embodiment, a plurality of recognition units are structured based on one or more of a target structured template, position information, and text information, including: sequentially detecting the distance between all adjacent characters in each recognition unit; and when it is detected that the distance between any adjacent characters is greater than a first character spacing threshold, splitting the recognition unit based on the adjacent characters whose distance is greater than the first character spacing threshold (i.e., when the distance between any adjacent characters in the recognition unit is greater than the first character spacing threshold, the recognition unit is split between the adjacent characters), until the distance between all adjacent characters in each recognition unit is less than or equal to the first character spacing threshold. For example, recognition unit A has 5 characters, and when the distance between the 3rd and 4th characters in recognition unit A is greater than the first character spacing threshold, recognition unit A is split based on the 3rd and 4th characters to obtain recognition units A1 and A2. Recognition unit A1 includes characters 1-3, and recognition unit A2 includes characters 4-5. According to another example, there are 10 characters in recognition unit B, and the distance between the 5th and 6th characters in recognition unit B is greater than the first character spacing threshold, and the distance between the 7th and 8th characters in recognition unit B is also greater than the first character spacing threshold. Recognition unit A is split based on the 5th and 6th characters to obtain recognition units B1 and B2, and recognition unit B2 is split based on the original 7th character (new 2nd character) and the original 8th character (new 3rd character) to obtain recognition units B2 and B3. Recognition unit B1 includes characters 1-5, recognition unit B2 includes characters 6-7, and recognition unit B3 includes characters 8-10. It can be seen that the above processing can be performed for each recognition unit in the order of the recognition units until the distance between all adjacent characters in each recognition unit is less than or equal to the first character spacing threshold (that is, there is no need to split the recognition unit).
[0134] In one embodiment, a plurality of recognition units are structured based on one or more of a target structured template, position information, and text information, including: determining the distance between any two adjacent recognition units; and merging the two adjacent recognition units into one recognition unit when the distance between any two adjacent recognition units is less than a second character spacing threshold. For example, recognition unit C and recognition unit D are two horizontally adjacent recognition units, and recognition unit C is to the left of recognition unit D. When the distance between the last character (the rightmost character) of recognition unit C and the first character (the leftmost character) of recognition unit D is less than the second character spacing threshold, recognition unit C and recognition unit D are merged into one recognition unit C.
[0135] In one embodiment, during the structuring process of multiple recognition units, it may be necessary to separate text boxes into rows and columns. This is because the bill information may be tilted during printing. To this end, the present application utilizes coordinate information to divide the multiple text boxes into rows and columns for structuring. The first text box in the upper left corner is identified and positioned as the first row and first column, thereby determining the base coordinates for the first row and first column. To address the tilt issue caused by the bill information during printing, a straight line is fitted using the adjacent row and column coordinates. Specifically, the text boxes are traversed, and the text boxes closest to the right and below the first text box are identified as the second text box in the first row and the third text box in the first column. After determining the three base text boxes, a horizontal straight line and a vertical straight line are fitted, respectively. The text boxes are traversed, and the remaining text boxes passing through these two straight lines are identified as the text boxes in the first row and the first column. After determining the text boxes in the first row and the first column, the row and column spacings are determined, and the average values of the adjacent row and column spacings are calculated. Based on the calculation of the average value of adjacent rows and the average value of column spacing, as well as the horizontal and vertical straight lines, the remaining text boxes are determined in corresponding positions, thereby completing the division of text boxes into rows and columns and completing the structural processing of data.
[0136] In one embodiment, a plurality of recognition units are subjected to structural processing based on one or more of a target structured template, position information, and text information, including: selecting a reference position recognition unit from the plurality of recognition units; determining a first recognition unit adjacent to the reference position recognition unit in a horizontal position and a second recognition unit adjacent to the reference position recognition unit in a vertical position; determining the text inclination of the image file to be processed based on the reference position recognition unit, the first recognition unit, and the second recognition unit; and determining the row and column position relationship of the plurality of recognition units in the image file to be processed based on the text inclination. Figure 2As shown, for example, the reference position identification unit is selected from multiple identification units as "Cefdinir Dispersible Tablets / 01.g*12t". It should be understood that this embodiment uses "Cefdinir Dispersible Tablets / 01.g*12t" as an example for explanation of the reference position identification unit for the purpose of explanation. As above, the first text box in the upper left corner can be selected. Determine the first identification unit "1 / box" adjacent to the reference position identification unit "Cefdinir Dispersible Tablets / 01.g*12t" in the horizontal position and the second identification unit "Lanqin Oral Liquid / 10ml*6 sticks" adjacent to the reference position identification unit "Cefdinir Dispersible Tablets / 01.g*12t" in the vertical position; based on the reference position identification unit "Cefdinir Dispersible Tablets / 01.g*12t", the first identification unit "1 / box" and the second identification unit "Lanqin Oral Liquid / 10ml*6 sticks", determine the text inclination of the image file to be processed. For example, when the inclination of the line connecting the reference position identification unit "Cefdinir Dispersible Tablets / 01.g*12t" and the first identification unit "1 / box" is 5 degrees, the horizontal inclination of the text inclination can be determined to be 5 degrees. For example, when the inclination of the line connecting the reference position identification unit "Cefdinir Dispersible Tablets / 01.g*12t" and the second identification unit "Lanqin Oral Liquid / 10ml*6 sticks" is 2 degrees, the horizontal inclination of the text inclination can be determined to be 2 degrees. Based on the text inclination, the row and column position relationship of multiple identification units in the image file to be processed is determined. That is, according to the text inclination, it can be determined that "Cefdinir Dispersible Tablets / 01.g*12t", "11.6300", "You Zifu", "Bai Rui Granules / 5g*4 bags", "5 / box" and "166.7000" and "You Zifu" are identification units in the same row.
[0137] In one embodiment, determining the row and column positional relationship of multiple recognition units based on text inclination includes: generating a horizontal reference line and a vertical reference line based on the text inclination; determining a recognition unit as a reference row among the multiple recognition units based on the horizontal reference line, and determining an average spacing between adjacent columns based on the recognition unit as the reference row; determining a recognition unit as a reference column among the multiple recognition units based on the vertical reference line, and determining an average spacing between adjacent rows based on the recognition unit as the reference column; and determining the row and column positional relationship of other recognition units other than the recognition unit in the reference row and the recognition unit in the reference column based on the average spacing between adjacent columns and the average spacing between adjacent rows. For example, when the horizontal inclination of the text inclination is determined to be 5 degrees and the horizontal inclination of the text inclination is determined to be 2 degrees, the horizontal reference line and the vertical reference line are generated based on the text inclination. Based on the horizontal baseline, the identification units as the reference rows among the multiple identification units are determined to be "Cefdinir Dispersible Tablets / 01.g*12t", "11.6300", "Self-paid", "Bai Rui Granules / 5g*4 bags", "5 / box", "166.7000" and "Self-paid", and based on the vertical baseline, the identification units as the reference columns among the multiple identification units are determined to be "Cefdinir Dispersible Tablets / 01.g*12t" and "Lanqin Oral Liquid / 10ml*6 bottles".
[0138] Subsequently, determine the spacing mean value of adjacent columns according to the recognition unit as the reference row, and determine the spacing mean value of adjacent rows according to the recognition unit as the reference row. Based on the spacing mean value of adjacent columns and the spacing mean value of adjacent rows, the row spacing and column spacing of data can be determined. For example, determine the row spacing between the row of "Cefdinir Dispersible Tablets / 01.g*12t" and the row of "Lanqin Oral Liquid / 10ml*6 sticks". Through the row spacing, on the basis of having determined the row of "Cefdinir Dispersible Tablets / 01.g*12t", determine the row position of the row of "Lanqin Oral Liquid / 10ml*6 sticks". In like manner, the column position of the character can be determined. Accordingly, based on the spacing mean value of adjacent columns and the spacing mean value of adjacent rows, determine the row and column positional relationship of other recognition units except the recognition unit of the reference row and the recognition unit of the reference column.
[0139] Figure 3 The apparatus for generating structured data provided by an exemplary embodiment of the present disclosure includes: an identification unit 301 , a determination unit 302 , a selection unit 303 , and a processing unit 304 .
[0140] The recognition unit 301 is used to perform text recognition on the image file to be processed to obtain a plurality of recognition units, wherein each recognition unit includes: position information and text information.
[0141] The determining unit 302 is configured to determine the data type of the image file to be processed based on the text information of the multiple recognition units.
[0142] The selection unit 303 is configured to select a target structured template suitable for the image file to be processed from a plurality of structured templates according to the data type.
[0143] The processing unit 304 is configured to perform structural processing on the plurality of recognition units based on one or more of the target structured template, the position information, and the text information, and generate structured data according to the structurally processed recognition units.
[0144] In one embodiment, the processing unit 304 includes:
[0145] a first determining subunit, configured to determine, from a plurality of identification units, an identification unit belonging to the format content according to the position information and / or text information of each format content in the target structured template; and
[0146] The first deleting subunit is used to delete the recognition units belonging to the format content from the multiple recognition units.
[0147] In one embodiment, the processing unit 304 further includes:
[0148] a second determining subunit, configured to determine the length and width of each identification unit based on the corresponding position information;
[0149] a third determining subunit, configured to determine the number of characters in each recognition unit based on corresponding text information;
[0150] a fourth determining subunit, configured to determine a recognition unit whose width is greater than its length and whose number of characters is greater than a preset threshold as a recognition unit belonging to the format content in the target structured template; and
[0151] The second deleting subunit is used to delete the identified identification unit belonging to the format content.
[0152] In one embodiment, the processing unit 304 further includes:
[0153] a fifth determining subunit, configured to determine the number of characters in each recognition unit based on corresponding text information;
[0154] The third deleting subunit is used to delete the recognition unit whose number of characters is zero.
[0155] In one embodiment, the processing unit 304 further includes:
[0156] A detection subunit, used to sequentially detect the distances between all adjacent characters in each recognition unit;
[0157] A processing subunit is used to split the recognition unit based on the adjacent characters whose distance is greater than the first character spacing threshold when it is detected that the distance between any adjacent characters is greater than the first character spacing threshold, until the distance between all adjacent characters in each recognition unit is less than or equal to the first character spacing threshold.
[0158] In one embodiment, the processing unit 304 further includes:
[0159] a sixth determining subunit, configured to determine the distance between any two adjacent recognition units;
[0160] The merging subunit is configured to merge any two adjacent recognition units into one recognition unit when the distance between the two adjacent recognition units is less than a second character spacing threshold.
[0161] In one embodiment, the processing unit 304 further includes:
[0162] a first selection subunit, configured to select a reference position recognition unit from a plurality of recognition units;
[0163] a seventh determining subunit, configured to determine a first identification unit adjacent to the reference position identification unit in a transverse position and a second identification unit adjacent to the reference position identification unit in a longitudinal position;
[0164] an eighth determining subunit, configured to determine a text inclination of the image file to be processed based on the reference position identifying unit, the first identifying unit, and the second identifying unit;
[0165] The ninth determining subunit is configured to determine, based on the text inclination, a row and column position relationship of the plurality of recognition units in the image file to be processed.
[0166] The ninth determination subunit is specifically used to: generate a horizontal baseline and a vertical baseline based on the inclination of the text; determine an identification unit as a reference row among multiple identification units based on the horizontal baseline, and determine the average spacing of adjacent columns based on the identification unit as the reference row; determine an identification unit as a reference column among multiple identification units based on the vertical baseline, and determine the average spacing of adjacent rows based on the identification unit as the reference column; and determine the row and column position relationship of other identification units except the identification unit of the reference row and the identification unit of the reference column based on the average spacing of adjacent columns and the average spacing of adjacent rows.
[0167] In one embodiment, the identification unit 301 includes:
[0168] A segmentation subunit is used to segment the text content in the image file to be processed, thereby obtaining multiple character lines;
[0169] a tenth determining subunit, configured to segment each character row to obtain a plurality of recognition units, and determine position information of the recognition units according to coordinates of the recognition units in the image file to be processed; and
[0170] The recognition subunit is used to perform character recognition on the characters in each recognition unit, so as to use the characters obtained through character recognition as text information of the recognition unit.
[0171] In one embodiment, the determining unit 302 is specifically configured to: search the text information of the multiple recognition units to determine keywords associated with the type; and determine the data type of the image file to be processed based on the keywords associated with the type.
[0172] In one embodiment, the determination unit 302 is further specifically used to: determine the number of recognition units; determine the size information of each recognition unit based on the position information of the recognition unit; and determine the data type of the image file to be processed based on the number of recognition units and the size information of each recognition unit.
[0173] In one embodiment, the selection unit 303 includes:
[0174] a matching subunit, configured to match the data type with the data type of each structured template in the plurality of structured templates to determine a degree of matching with each structured template; and
[0175] The second selection subunit is configured to select a structured module with the greatest matching degree as a target structured template suitable for the image file to be processed.
[0176] Figure 4 The electronic device according to an exemplary embodiment of the present disclosure may be a first device or a second device, or both of the first device and the second device, or a standalone device independent of the first device and the second device. The standalone device may communicate with the first device and the second device to receive collected input signals from the first device and the second device. Figure 4 FIG. 1 shows a block diagram of an electronic device according to an embodiment of the present disclosure. Figure 4 As shown, the electronic device includes one or more processors 410 and memory 420 .
[0177] The processor 410 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0178] The memory 420 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on a computer-readable storage medium, and the processor 410 may execute the program instructions to implement the method of generating structured data and / or other desired functions of the software program of each embodiment of the present disclosure above. In one example, the electronic device may further include: an input device 430 and an output device 440, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0179] In addition, the input device 430 may also include, for example, a keyboard, a mouse, and the like.
[0180] The output device 440 can output various information to the outside. The output device 440 can include, for example, a display, a speaker, a printer, a communication network and its connected remote output device, etc.
[0181] Of course, to simplify, Figure 4 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.
[0182] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions. When the computer program instructions are executed by a processor, the processor executes the steps in the method of generating structured data according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0183] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0184] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps in the method of generating structured data according to various embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.
[0185] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0186] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0187] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.
[0188] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0189] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.
[0190] It should also be noted that, in the apparatus, equipment and method of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present disclosure. The above description of the disclosed aspects is provided to enable any technician in this field to make or use the present disclosure. Various modifications to these aspects will be very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown here, but to the widest range consistent with the principles and novel features disclosed herein.
[0191] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A method for generating structured data, comprising: Performing text recognition on the image file to be processed to obtain a plurality of recognition units, wherein each recognition unit includes: position information and text information; determining a data type of the image file to be processed based on the text information of the plurality of recognition units; selecting a target structured template suitable for the image file to be processed from a plurality of structured templates according to the data type; and Performing structural processing on the plurality of recognition units based on one or more of the target structured template, the position information, and the text information, and generating structured data according to the structurally processed recognition units; The step of determining the data type of the image file to be processed based on the text information of the plurality of recognition units includes: Searching the text information of the plurality of recognition units to determine keywords associated with the type; determining the data type of the image file to be processed based on the keyword associated with the type; The performing structural processing on the multiple recognition units based on one or more of the target structured template, the position information, and the text information includes: selecting a reference position identification unit from the plurality of identification units; Determine a first recognition unit adjacent to the reference position recognition unit in a transverse position and a second recognition unit adjacent to the reference position recognition unit in a longitudinal position; determining a text inclination of the image file to be processed based on the reference position recognition unit, the first recognition unit, and the second recognition unit; Determining, based on the text inclination, a row and column position relationship of the plurality of recognition units in the image file to be processed; The step of determining the row and column position relationship of the plurality of recognition units based on the text inclination includes: generating a horizontal reference line and a vertical reference line based on the text inclination; Determining an identification unit as a reference row among the plurality of identification units based on the horizontal reference line, and determining an average value of spacings between adjacent columns according to the identification unit as the reference row; Determining recognition units as reference columns among the plurality of recognition units based on the vertical reference line, and determining an average value of spacings between adjacent rows according to the recognition units as reference columns; and Based on the average distance between adjacent columns and the average distance between adjacent rows, the row and column positional relationship of the other recognition units except the recognition units in the reference row and the recognition units in the reference column is determined.
2. The method according to claim 1, wherein The performing structural processing on the multiple recognition units based on one or more of the target structured template, the position information, and the text information includes: Determining, from the plurality of recognition units, a recognition unit belonging to the format content according to the position information and / or text information of each format content in the target structured template; and The identification units belonging to the format content among the multiple identification units are deleted.
3. The method according to claim 1, wherein The performing structural processing on the multiple recognition units based on one or more of the target structured template, the position information, and the text information includes: determining the length and width of each identification unit based on the corresponding position information; determining the number of characters in each recognition unit based on the corresponding text information; Determining a recognition unit whose width is greater than its length and whose number of characters is greater than a preset threshold as a recognition unit belonging to the format content in the target structured template; and The identified recognition units belonging to the format content are deleted.
4. The method according to claim 1, wherein The performing structural processing on the multiple recognition units based on one or more of the target structured template, the position information, and the text information includes: determining the number of characters in each recognition unit based on the corresponding text information; Delete the recognition units with zero characters.
5. The method according to claim 1, wherein The performing structural processing on the multiple recognition units based on one or more of the target structured template, the position information, and the text information includes: Detect the distances between all adjacent characters in each recognition unit in sequence; When it is detected that the distance between any adjacent characters is greater than the first character spacing threshold, the recognition units are split based on the adjacent characters whose distance is greater than the first character spacing threshold until the distance between all adjacent characters in each recognition unit is less than or equal to the first character spacing threshold.
6. The method according to claim 1, wherein the structural processing of the plurality of recognition units based on one or more of the target structured template, position information, and text information comprises: Determine the distance between two adjacent recognition units; When the distance between any two adjacent recognition units is less than the second character spacing threshold, the two adjacent recognition units are merged into one recognition unit.
7. The method according to claim 1, wherein The text recognition is performed on the image file to be processed to obtain a plurality of recognition units, including: Segmenting the text content in the image file to be processed to obtain a plurality of character lines; Performing character segmentation on each character row to obtain a plurality of recognition units, and determining position information of the recognition units according to coordinates of the recognition units in the image file to be processed; and Character recognition is performed on the characters in each recognition unit, so that the characters obtained through character recognition are used as text information of the recognition unit.
8. The method according to claim 1 , further comprising determining a data type of the image file to be processed based on position information of the plurality of recognition units: Determine the number of recognition units; Determining size information of each recognition unit based on the position information of the recognition unit; The data type of the image file to be processed is determined based on the number of recognition units and size information of each recognition unit.
9. The method according to claim 1 or 8, wherein The selecting, according to the data type, a target structured template suitable for the image file to be processed from a plurality of structured templates includes: matching the data type with the data type of each structured template in a plurality of structured templates to determine a degree of match with each structured template; and The structured module with the largest matching degree is selected as the target structured template suitable for the image file to be processed.
10. A device for generating structured data, comprising: A recognition unit, configured to perform text recognition on the image file to be processed to obtain a plurality of recognition units, wherein each recognition unit includes: position information and text information; a determining unit, configured to determine a data type of the image file to be processed based on the text information of the plurality of recognition units; a selecting unit, configured to select a target structured template suitable for the image file to be processed from a plurality of structured templates according to the data type; and a processing unit, configured to perform structural processing on the plurality of recognition units based on one or more of the target structured template, the position information, and the text information, and generate structured data according to the structurally processed recognition units; The determining unit is further configured to: Searching the text information of the plurality of recognition units to determine keywords associated with the type; determining the data type of the image file to be processed based on the keyword associated with the type; The performing structural processing on the multiple recognition units based on one or more of the target structured template, the position information, and the text information includes: selecting a reference position identification unit from the plurality of identification units; Determine a first recognition unit adjacent to the reference position recognition unit in a transverse position and a second recognition unit adjacent to the reference position recognition unit in a longitudinal position; determining a text inclination of the image file to be processed based on the reference position recognition unit, the first recognition unit, and the second recognition unit; Determining, based on the text inclination, a row and column position relationship of the plurality of recognition units in the image file to be processed; The step of determining the row and column position relationship of the plurality of recognition units based on the text inclination includes: generating a horizontal reference line and a vertical reference line based on the text inclination; Determining an identification unit as a reference row among the plurality of identification units based on the horizontal reference line, and determining an average value of spacings between adjacent columns according to the identification unit as the reference row; Determining recognition units as reference columns among the plurality of recognition units based on the vertical reference line, and determining an average value of spacings between adjacent rows according to the recognition units as reference columns; and Based on the average distance between adjacent columns and the average distance between adjacent rows, the row and column positional relationship of the other recognition units except the recognition units in the reference row and the recognition units in the reference column is determined.
11. A computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 9.
12. An electronic device comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Physical examination report information structuring method and device, readable storage medium and terminal
CN112686258A
Text structured recognition method and device, electronic equipment and computer readable medium
CN113723158A
Text information acquisition method and system
CN113779935A