Method, apparatus, and computer-readable storage medium for recognizing characters in digital documents
By classifying digital document segments as handwritten or printed text zones, the method optimizes character recognition by applying OCR or ICR algorithms efficiently, addressing inefficiencies in current methods and improving processing speed and accuracy.
Patent Information
- Application Number
- JP2023574705
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-01
- Filing Date
- 2021-06-02
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2041-06-02
AI Technical Summary
Current character recognition methods for handwritten and printed characters are inefficient and burdensome, especially when processing large documents, as they often require applying both optical character recognition (OCR) and intelligent character recognition (ICR) to the entire document without considering the content type.
A method and apparatus that classify digital document segments as either handwritten or printed text zones using parameter values and threshold values based on distribution profiles, allowing for efficient application of OCR or ICR algorithms based on the zone classification.
This approach reduces processing time and improves accuracy by applying the appropriate recognition algorithm to specific text zones, avoiding redundant processing and enhancing overall efficiency.
Smart Images

Figure 0007758758000002 
Figure 0007758758000003 
Figure 0007758758000004
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 17 / 335,547, filed June 1, 2021, which is incorporated by reference herein in its entirety for all purposes.
[0002] This disclosure relates generally to the field of character recognition, and more particularly to recognizing handwritten and printed characters. [Background technology]
[0003] Typically, character recognition is performed using various character recognition algorithms. Current character recognition methods utilized to recognize handwritten and printed characters are burdensome, slow, and inefficient, especially when a document contains a large number of characters to be processed.
[0004] The foregoing "Background" discussion is intended to generally present the context for the present disclosure. The inventor's work, to the extent described in this "Background" section, is not expressly or impliedly admitted as prior art to the present disclosure, as are aspects of the description that may not be admitted as prior art at the time of filing. Summary of the Invention
[0005] The present disclosure relates to recognizing characters in digital documents.
[0006] According to an embodiment, the present disclosure further relates to a method for recognizing characters in a digital document, the method including: classifying, by a processing circuit, a segment of the digital document as containing text; calculating, by the processing circuit, at least one parameter value associated with the classified segment of the digital document; determining, by the processing circuit, a zone parameter value based on the calculated at least one parameter value; classifying, by the processing circuit, the segment of the digital document as a handwritten text zone or a printed text zone based on the determined zone parameter value and a threshold value, wherein the threshold value is based on selecting an intersection of a handwritten text distribution profile and a printed text distribution profile, each of the handwritten text distribution profile and the printed text distribution profile being associated with a zone parameter that corresponds to the determined zone parameter value; and generating, by the processing circuit, a modified version of the digital document based on the classification.
[0007] According to an embodiment, the present disclosure further relates to a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for recognizing characters in a digital document, the method including: classifying a segment of the digital document as containing text; calculating at least one parameter value associated with the classified segment of the digital document; determining a zone parameter value based on the calculated at least one parameter value; classifying the segment of the digital document as a handwritten text zone or a printed text zone based on the determined zone parameter value and a threshold value, wherein the threshold value is based on selecting an intersection of a handwritten text distribution profile and a printed text distribution profile, each of the handwritten text distribution profile and the printed text distribution profile being associated with a zone parameter that corresponds to the determined zone parameter value; and generating a modified version of the digital document based on the classification.
[0008] According to an embodiment, the present disclosure further relates to an apparatus for recognizing characters in a digital document, the apparatus comprising a processing circuit configured to: classify a segment of the digital document as containing text; calculate at least one parameter value associated with the classified segment of the digital document; determine a zone parameter value based on the calculated at least one parameter value; classify the segment of the digital document as a handwritten text zone or a printed text zone based on the determined zone parameter value and a threshold value, wherein the threshold value is based on selecting an intersection of a handwritten text distribution profile and a printed text distribution profile, each of the handwritten text distribution profile and the printed text distribution profile being associated with a zone parameter that corresponds to the determined zone parameter value; and generate a modified version of the digital document based on the segment classification.
[0009] The preceding paragraphs have been provided by way of general introduction and are not intended to limit the scope of the claims that follow. The described embodiments, together with further advantages, will be best understood by reference to the following detailed description taken in conjunction with the accompanying drawings, in which:
[0010] A more complete understanding of the present disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a block diagram of a character recognition system according to one embodiment of the present disclosure. [Figure 2A] 1 is a flowchart illustrating a character recognition method according to an embodiment of the present disclosure. [Figure 2B] 1 is a flowchart illustrating a character recognition method according to an embodiment of the present disclosure. [Figure 2C]1 is a flowchart illustrating a sub-process of a character recognition method according to an embodiment of the present disclosure. [Figure 3] FIG. 10 illustrates a distribution function curve according to one embodiment of the present disclosure. [Figure 4] 1 is a flowchart illustrating generating selectable versions of a document based on a character recognition process according to one embodiment of the present disclosure. [Figure 5] 1 illustrates an exemplary image of an object according to one embodiment of the present disclosure. [Figure 6] 1 is a flowchart illustrating a character recognition method using a trained machine learning algorithm, according to one embodiment of the present disclosure. [Figure 7] FIG. 2 is a detailed block diagram illustrating an exemplary classifier server in accordance with certain embodiments of the present disclosure. [Figure 8] FIG. 2 is a detailed block diagram illustrating an exemplary user device in accordance with certain embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] As used herein, the terms "a" or "an" are defined as one or more than one. As used herein, the term "plurality" is defined as two or more than two. As used herein, the term "another" is defined as at least a second or more. As used herein, the terms "including" and / or "having" are defined as comprising (i.e., open language). References throughout this specification to "one embodiment," "particular embodiment," "embodiment," "implementation," "example," or similar terms mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of such phrases or in various places throughout the specification are not necessarily all referring to the same embodiment. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments, without limitation.
[0013] Additionally, terms such as "approximately," "proximate," "minor," and similar terms generally refer to ranges that include the identified value within a margin of 20%, 10%, or preferably 5%, in certain embodiments, and any value therebetween.
[0014] The term "user" and other related terms are used interchangeably to refer to a person using a character recognition circuit, a character recognition system, or a system that sends input to a character recognition circuit / system.
[0015] 1 is a block diagram of a character recognition system 100 according to one embodiment of the present disclosure. The character recognition system 100 includes user devices 102(1)-102(n), a digital document 104, a network 106, a classifier server device 108, and a database storage device 110. The user devices 102(1)-102(n) may also be referred to as a pool of user devices 102(1)-102(n). Components of the system 100 may include computing devices (e.g., computers, servers, etc.) having memory that stores data and / or software instructions (e.g., server code, client code, databases, etc.). In some embodiments, one or more computing devices may be configured to execute software instructions stored on one or more memory devices to perform one or more operations consistent with disclosed embodiments.
[0016] The user devices 102(1)-102(n) may be tablet computing devices, mobile phones, laptop computing devices, and / or personal computing devices, but may also include any other user communication devices. In particular embodiments, the user devices 102(1)-102(n) may be smartphones. However, those skilled in the art will understand that the features described herein may be adapted to be implemented in other devices (e.g., servers, electronic readers, cameras, navigation devices, etc.).
[0017] Digital document 104 may be a digital version of any document that can be stored or displayed on user device 102(1)-102(n). Digital document 104 may be in any format, such as a Portable Document Format (PDF), an image file such as the Joint Photographic Experts Group (JPEG) file format, a word processing document such as those commonly used in Microsoft Word (.doc, .docx), or other digital formats known to those skilled in the art. Additionally, digital document 104 may be a digital version of a physical document that can be captured and converted into a digitized document by user device 102(1)-102(n).
[0018] User devices 102(1)-102(n) may be scanners, fax machines, cameras, or other similar devices used to convert physical documents to generate digital documents 104. Those skilled in the art will appreciate that the present disclosure is not limited to any particular device used to generate digital documents 104. Digital documents 104 may be any artifact having textual content. Digital documents 104 may include printed text, handwritten text, graphics, barcodes, QR codes, lines, images, shapes, colors, structures, formats, layouts, or other identifiers. Digital documents 104 may be forms with input fields that can be filled out by a user, personal identification documents that may include text, images or other identifiers, characters, notes, etc.
[0019] The network 106 may comprise one or more types of computer networking structures configured to provide communication, exchange data, or both, between components of the system 100. For example, the network 106 may include any type of network (including infrastructure) that provides communication, exchanges information, and / or facilitates the exchange of information, such as the Internet, a private data network, a virtual private network using a public network, a LAN or WAN network, a Wi-Fi™ network, and / or other suitable connections that may enable information exchange between various components of the system 100. The network 106 may also include a public switched telephone network ("PSTN") and / or a wireless cellular network. The network 106 may be a secured network or an unsecured network. In some embodiments, one or more components of the system 100 may communicate directly via dedicated communication links. The user devices 102(1)-102(n), the classifier server device 108, and the database storage device 110 may be configured to communicate with each other via the network 106.
[0020] The classifier server device 108 may be one or more network-accessible computing devices configured to perform one or more operations consistent with the disclosed embodiments, as discussed more fully below. As discussed below, the classifier server device 108 may be a network device that stores instructions for classifying zones within the digital document 104 as handwritten text zones or printed text zones.
[0021] The database storage device 110 may be communicatively coupled, directly or indirectly, to the classifier server device 108 and the user devices 102(1)-102(n) via the network 106. The database storage device 110 may also be part of the classifier server device 108 (i.e., not a separate device). The database storage device 110 may include one or more memory devices that store information and are accessed and / or managed by one or more components of the system 100. By way of example, the database storage device 110 may include an Oracle™ database, a Sybase™ database, or other relational or non-relational databases, such as Hadoop sequence files, HBase, or Cassandra. The database storage device 110 may include computing components (e.g., a database operating system, a network interface, etc.) configured to receive and process requests for data stored in the memory device of the database storage device 110 and to provide data from the database storage device 110. The database storage device 110 may be configured to store instructions for classifying zones within the digital document 104 as handwritten text zones or printed text zones.
[0022] In one embodiment, this technique is utilized to recognize handwritten and / or printed characters within a digital document 104. Briefly summarized, in one embodiment, the method of the present disclosure first includes segmenting the digital document. The segments of the digital document may then be classified into text zones and non-text zones. Each of the text zones may then be classified as containing handwritten text or printed text based on one or more parameters associated with the text zone. A character recognition process may then be applied to the text zones based on their classification as either handwritten text zones or printed text zones. For example, an intelligent character recognition (ICR) process may be applied to handwritten text zones, and an optical character recognition (OCR) process may be applied to printed text zones.
[0023] By classifying different text zones into their respective types of text, each zone of a digital document does not need to be processed twice (i.e., by both ICR and OCR). This approach contrasts with conventional approaches that apply OCR to entire regions of a document without knowing or considering the document's content. For example, OCR may be acceptable for certain machine-generated text portions (i.e., printed text portions), but is inefficient and / or inaccurate in processing handwritten text portions. Thus, applying OCR to the entire document fails to provide speed or accuracy when processing documents containing text.
[0024] According to one embodiment, the method of the present disclosure may employ a single recognition module that performs each of the OCR and ICR processes.
[0025] In one embodiment, the OCR process includes an algorithm that recognizes printed text within a digital document. The printed text recognized by the OCR algorithm may include alphabetic, numeric, and / or special characters, but may also include any other digital characters.
[0026] In one embodiment, the ICR process includes an algorithm that recognizes handwritten text within a digital document. Handwritten text that can be recognized by the ICR process includes alphabetic, numeric, and / or special characters, but can also include any other handwritten characters.
[0027] Referring again to the figures, FIGS. 2A, 2B, and 2C are flow diagrams illustrating a method of character recognition by character recognition system 100, according to an exemplary embodiment of the present disclosure. The method may begin when a digital document 104 is received by classifier server device 108 from user device 102(1) of multiple user devices 102(1)-102(n). User devices 102(1)-102(n) may be scanners, fax machines, cameras, or other similar devices. In one embodiment, digital document 104 may be generated by user device 102(1) by converting a physical document (e.g., a page containing handwritten text, printed text, and / or graphics) into digital document 104, or by creating digital document 104 in any format, such as PDF, image, Word, or other digital format known to those skilled in the art.
[0028] In one embodiment, the digital document 104 may be received from a user device 102(1)-102(n). Additionally, the digital document 104 may be generated by a user device 102(1)-102(n).
[0029] At step 202 of method 200, the classifier server device 108 receives a request for character recognition of the digital document 104 from the user device 102(1). In one embodiment, the user may use a user interface displayed on the user device 102(1) to upload the digital document 104 to the classifier server device 108 and may further select a button on the user interface of the user device 102(1) to send a request to the classifier server device 108 and initiate the character recognition process of the digital document 104. The classifier server device 108 receives the digital document 104 along with instructions to initiate the character recognition process of the digital document 104. Thus, upon receiving the digital document 104, the classifier server device 108 initiates the character recognition process by performing segmentation on the digital document 104, as described below in step 204 of method 200.
[0030] In one embodiment, the classifier server device 108 receives requests from any of the user devices 102(1)-102(n) for character recognition to convert the digital document 104 into a selectable document.
[0031] In one embodiment, digital document 104 may be a document that includes text zones and non-text zones. The text zones may include handwritten and / or printed text, but may also include any other type of text data. The non-text zones may include graphic zones, magnetic ink character recognition (MICR), machine-readable zones (MRZ), optical mark recognition (OMR), images, diagrams, drawings, and / or tables, but may also include any other type of non-text data.
[0032] In one embodiment, digital document 104 may be a document containing handwritten text. In one embodiment, digital document 104 may be a document containing printed text. In one embodiment, digital document 104 may be a document containing handwritten and printed text.
[0033] In one embodiment, and as a simple example, this digital document 104 may be a single page containing the contents of a business letter. Thus, the single page includes a logo, a greeting, a body, and a signature.
[0034] In step 204 of method 200, the classifier server device 108 segments the digital document 104 into a plurality of zones by applying one or more segmentation algorithms to the digital document 104. The one or more segmentation algorithms may include classical computer vision-based approaches or artificial intelligence-based approaches. The one or more segmentation algorithms may deploy semantic segmentation or instance segmentation. Furthermore, the one or more segmentation algorithms may include threshold segmentation, edge detection segmentation, k-means clustering, and region-based convolutional neural networks. It will be appreciated that method 200 may similarly be implemented such that the digital document 104 received in step 202 is pre-segmented by another software program, device, or method.
[0035] In one embodiment, the one or more segmentation algorithms may include a region-based segmentation algorithm, an edge-based segmentation algorithm, a threshold- and color-based segmentation algorithm, and / or an object analysis-based segmentation algorithm, but may also include any other segmentation algorithm. Region-based segmentation algorithms may include, but are not limited to, watershed and region-growing segmentation algorithms. Edge-based segmentation algorithms may include, but are not limited to, edge detection algorithms and active contour algorithms. Threshold- and color-based segmentation algorithms may include, but are not limited to, Otsu, multi-level color, fixed threshold, adaptive threshold, dynamic threshold, and / or auto-threshold algorithms. Object analysis-based segmentation algorithms are further described below.
[0036] In one embodiment, the classifier server device 108 utilizes an object analysis-based segmentation algorithm to extract all objects from the digital document 104. Objects (e.g., horizontal lines, vertical lines, text objects, graphic objects, etc.) are first detected in the digital document 104 by the classifier server device 108. In one embodiment, the classifier server device 108 classifies the detected objects as non-text objects and text objects. Text objects may include characters (e.g., handwritten characters, printed characters), numbers (e.g., handwritten numbers, printed numbers), letters (e.g., handwritten letters, printed letters), etc. Non-text objects may include horizontal lines, vertical lines, graphic objects, etc.
[0037] To this end, the text object may be divided into multiple blocks by the classifier server device 108. A block may be, for example, any grouping of characters (e.g., a single character of a letter or number, a word, a sentence, a paragraph, etc.). A block may be distinguished from another block based on whether the distance between them may be greater than a predetermined distance. For example, the distance between two paragraphs may be much greater than the distance between two characters in a word. Thus, such distance may be used in determining the size of a block. In one embodiment, each block may be a different size. In one embodiment, some blocks may be the same size, while other blocks may be different sizes. In one embodiment, the minimum distance between blocks may be twice the size of a character, and the distance between characters may be less than or equal to 0.2 times the size of a character. The distances defined above provide relationships between characters that may be implemented in various metrics, such as pixels. In one embodiment, the text object is divided into blocks by the classifier server device 108 based on the height and stroke width corresponding to the text object. For example, characters with the same height and stroke width are grouped together to form a text group. By way of example, a text group may include a word, a sentence, a paragraph, etc. Accordingly, each text group containing characters with the same height and stroke width is referred to as a unique text zone. The non-text objects may then be processed by the classifier server device 108 in a similar manner to text objects. By way of example, the classifier server device 108 utilizes an object analysis-based segmentation algorithm to extract all non-text objects from the digital document 104. Non-text objects (e.g., horizontal lines, vertical lines, graphic objects, etc.) are detected within the digital document 104. In one embodiment, the classifier server device 108 groups each of the non-text objects into a non-text group. By way of example, the classifier server device 108 determines whether the intersection of a horizontal line and a vertical line forms a table and identifies the table as a unique non-text zone.Furthermore, if it is determined that horizontal and vertical lines do not form a table, each horizontal and vertical line is determined as a unique non-text zone by the classifier server device 108. The classifier server device 108 utilizes an object analysis-based segmentation algorithm to identify graphic objects, such as, but not limited to, images, drawings, etc., as each unique non-text zone. Each graphic grouping may be determined as a unique non-text zone by the classifier server device 108. In one embodiment, text objects may include characters (e.g., handwritten characters, printed characters), numbers (e.g., handwritten numbers, printed numbers), characters (e.g., handwritten characters, printed characters), etc. For example, a text block may be any grouping of characters (e.g., a single character such as a letter or number, a group of words, a group of sentences, a group of paragraphs, etc.). Furthermore, groupings of characters having the same height and stroke width (e.g., a single character such as a letter or number, a group of words containing characters of the same height and stroke width, a group of sentences containing words with the same height and stroke width, a group of paragraphs containing sentences with the same height and stroke width, etc.) together form a text group. Thus, each text group can be determined by the classifier server device 108 as a unique text zone.
[0038] In addition to the above, the segmentation algorithm may include a machine learning algorithm trained to segment documents, which may include a neural network algorithm trained to segment documents, a convolutional neural network algorithm trained to segment documents, a deep learning algorithm trained to segment documents, and a reinforcement learning algorithm trained to segment documents, but may also include any other neural network algorithm trained to segment documents.
[0039] Subsequently, in step 206 of method 200, the classifier server device 108 classifies each of the plurality of zones as a text zone or a non-text zone. Classification by the classifier server device 108 may be performed by known image classification techniques, including, but not limited to, artificial neural networks (e.g., convolutional neural networks), linear regression, logistic regression, decision trees, support vector machines, random forests, naive Bayes, and k-nearest neighbors.
[0040] In an example where the digital document 104 is a business letter, the digital document 104 may be segmented into at least four zones, and these segmented zones are then classified. The first zone may be a text zone containing the greeting of the business letter. The second zone may be a text zone containing the body of the business letter. The third zone may be a text zone containing the signature of the business letter. The fourth zone may be a non-text zone containing the logo of the business letter. The fifth zone may be a non-text zone forming the remaining white space of the business letter.
[0041] At step 208 of method 200, the classifier server device 108 calculates at least one parameter value for each text zone. In one embodiment, step 208 of method 200 may further include identifying text objects within each text zone and calculating at least one parameter value for each identified text object. In one embodiment, the at least one parameter value may be based on, among other things, the extracted shape of the text object, the extracted dimensions and perimeter of the text object, the stroke width of the text object, the assigned number of text objects, and pixel information related to color and intensity. The at least one parameter value may also include, for each text zone, curvature, center of mass, image characteristics, and attributes of neighboring zones.
[0042] In one embodiment, the at least one parameter value may include, among others, a dimension and perimeter value (i.e., an object dimension value and an object perimeter value) associated with each of the text objects, a color associated with each of the text objects, a contour associated with each of the text objects, an edge position associated with each of the text objects, an area associated with each of the text objects, a shape associated with each of the text objects, an object curvature value associated with each of the text objects, a center of mass value associated with each of the text objects, an object font stroke width associated with each of the text objects, an object height associated with each of the text objects, an object density value associated with each of the text objects, object pixel color information (e.g., color and intensity), a number of black pixels associated with each of the text objects, and / or an object pixel intensity associated with each of the text objects, but may also include any other type of parameter value. The at least one parameter value may also be referred to as one or more parameters. In one example, the object density associated with each of the text objects may be the ratio of the number of text objects in a text zone normalized by the zone width corresponding to the text zone associated with each of the text objects multiplied by the zone height corresponding to the text zone associated with each of the text objects.
[0043] In one embodiment, when it relates to a given text object within a text zone, the at least one parameter value may include a height of the text object within the text zone and a font stroke width of the text object within the text zone. Further, the at least one parameter value may be a ratio between the height of the text object within the text zone and the font stroke width of the text object within the text zone.
[0044] 5 provides an exemplary illustration of at least one parameter value that may be derived from a character in a text zone. FIG. 5 illustrates the character "C," which may be a text object 500A. Additionally, the at least one parameter value may be an object font stroke width 502 of the text object "C," an object height 504 of the text object "C," and / or an object width 506 of the text object "C." In one embodiment, the font stroke width 502 of the text object "C" is
number
[0045] 2A-2C , in step 210 of method 200, the classifier server device 108 assigns a ranking value to each calculated parameter value based on all of the parameter values calculated in step 208 of method 200 for a given text zone. In one embodiment, the ranking value may be an object ranking value for each text object in a given zone based on each other object ranking value of each other text object in the given zone (described in more detail below). For example, the object ranking value of each of the text objects in a first zone (i.e., the greeting) of the digital document 104 may be assigned a value between 0 and 15 based on the object value associated with each of the text objects in the first zone. In other words, the text object in the first zone having the lowest object value may be assigned an object ranking value of 0, while the text object in the first zone having the highest object value may be assigned an object ranking value of 15. Multiple text objects may have the same object ranking value. Additionally, any other value in the range may be included.
[0046] In the example described herein, the first zone may include a letter greeting. The letter greeting may be, for example, "Hello World." An object ranking value may be determined for each text object in the greeting. In other words, each of "H," "e," "l," "l," "o," "W," "o," "r," "l," and "d" may be assigned an object ranking value. Based on the object value determination method described above, the text object "H" may be assigned an object ranking value of 11, and the text object "o" may be assigned an object ranking value of 4. It should be understood that such object ranking values are merely exemplary and are not intended to reflect runtime evaluation.
[0047] At step 212 of method 200, the classifier server device 108 determines at least one zone parameter value for each text zone. In one embodiment, the at least one zone parameter value may correspond to an average value of the at least one parameter value calculated at step 208 of method 200. In one embodiment, for a given text zone, the at least one zone parameter value may be determined as a calculation (e.g., an average) based on zone pixel information such as, among other things, the density of text objects within the text zone, the average dimension ratio of the text objects within the text zone, context uniformity, and the color, area, contours, edges, and text objects contained therein.
[0048] In one embodiment, when it relates to a single text zone, at least one zone parameter value of the text zone may be an average of the dimension ratios of each text object within the text zone, for example, each dimension ratio may be a comparison of the height of the text object to the stroke width of the text object.
[0049] In subprocess 212 of method 200, and assuming that only one zone parameter value is considered for each text zone of a digital document, a relationship between the zone parameter value and a threshold value can be identified to assist in classifying the text zone as a handwritten text zone or a printed text zone, as described with reference to FIG. 2C.
[0050] In step 216 of sub-process 214 of Figure 2C, average ranking values may be assigned to the zone parameter values determined in step 212 of method 200. The average ranking values may be assigned in a similar manner to the ranking values of step 210 of method 200.
[0051] In step 218 of sub-process 214, a threshold value can be identified by probing the distribution profiles associated with each of the handwritten and printed text. In other words, the threshold value is identified from the reference data as a value, such as a ranking value, above which a given text zone is most likely to be a handwritten text zone and below which a given text zone is most likely to be a printed text zone. The phrase "most likely" reflects the confidence based on how the distribution profiles relate to each other, i.e., the confidence that allows a given text zone to be subsequently classified as either a handwritten text zone or a printed text zone.
[0052] In one embodiment, the thresholds may be obtained from a reference table, such as a look-up table, that includes thresholds based on the type of zone parameter value. To this end, the thresholds in the reference table may be based on distribution profiles generated for each type of zone parameter value, may be calculated in real time, pre-calculated, or calculated by a third party, and may be based on reference data associated with "labeled" text zones, including those previously identified as either handwritten or printed text.
[0053] In one embodiment, each of the handwritten text distribution profile and the printed text distribution profile may be generated by the classifier server device 108 or retrieved from a local or remote storage device and may be based on reference data corresponding to the handwritten text and the printed text, respectively. Furthermore, each distribution profile may be based on the particular zone parameter value being evaluated. Thus, for example, if the zone parameter value is the aspect ratio of an object, each of the handwritten text distribution profile and the printed text distribution profile used to determine the associated threshold may be based on the aspect ratio of the object.
[0054] In general, the handwritten text distribution profile may be a distribution curve based on a histogram of a sample of reference data and associated ranking values. The reference data may be, for example, sample text zones classified as handwritten text. In one embodiment, the reference data may be based on zone parameter values with associated ranking values. For example, the zone parameter value may be object density. In one embodiment, the reference data may be based on parameter values of each object in the sample text zone, each of the sample text zones including objects with calculated parameter values and subsequently assigned ranking values, etc. For example, the calculated parameter value may be object stroke width. Thus, if object parameter values are used, each pair of calculated parameter value and assigned ranking value from each of the sample text zones may be used to populate the histogram. If zone parameter values are used, each pair of zone parameter value and assigned ranking value from each of the sample text zones may be used to populate the histogram. In other words, the histogram may be populated by, for each ranking value, counting the number of handwritten text sample text zones with the same ranking value. A distribution curve, referred to herein as a handwritten text distribution profile, can then be generated from the histogram by normalizing the histogram values by the total number of text zones in the handwritten text sample.
[0055] Similarly, the printed text distribution profile may be a distribution curve based on a histogram of a sample of reference data and associated ranking values. The reference data may be, for example, sample text zones classified as printed text. In one embodiment, the reference data may be based on zone parameter values with associated ranking values. For example, the zone parameter value may be object density. In one embodiment, the reference data may be based on parameter values of each object within the sample text zone, each of the sample text zones including objects with calculated parameter values and subsequently assigned ranking values. For example, the calculated parameter value may be object stroke width. Thus, if object parameter values are used, each pair of calculated parameter value and assigned ranking value from each of the sample text zones may be used to populate the histogram. If zone parameter values are used, each pair of zone parameter value and assigned ranking value from each of the sample text zones may be used to populate the histogram. In other words, the histogram may be populated by, for each ranking value, counting the number of printed text sample text zones with the same ranking value. A distribution curve, referred to herein as a printed text distribution profile, can then be generated from the histogram by normalizing the histogram values by the total number of text zones in the printed text sample.
[0056] In one embodiment, the printed text distribution profile and the handwritten text distribution profile may be as shown in FIG. 3. Shown are handwritten text distribution profile 308 and printed text distribution profile 316. Horizontal axis 312 defines the ranking value of the object, while vertical axis 310 defines the probability of observing a given ranking value within a handwritten text zone or a printed text zone. Handwritten text distribution profile 308 and printed text distribution profile 316 may each be generated for and correspond to a parameter value of interest. Intersection 320 may be identified as a threshold, similar to step 218 of subprocess 214. In other words, intersection 320 defines the ranking value as a threshold.
[0057] In one embodiment, the intersection 320 between the handwritten text distribution profile 308 and the printed text distribution profile 316 may be selected as the intersection of the distribution profiles that maximizes the area under the curves. Theoretically, as it relates to FIG. 3 , the prediction X follows the probability density function FH (e.g., handwritten text distribution profile 308) if the prediction actually belongs to the class “positive / handwritten,” and follows FM (e.g., printed text distribution profile 316) otherwise. Therefore, the true positive rate is given by the area under the handwritten text distribution profile 308 above a threshold T defined by the intersection 320. In other words, the true positive rate is the sum of the values of the handwritten text distribution profile 308 that are greater than the intersection 320 multiplied by the distance between them. Meanwhile, the true negative rate is given by the area under the printed text distribution curve, which is the sum of the printed text distribution curve values that are less than the threshold T multiplied by the distance between them. The threshold T of the present disclosure may be the intersection that maximizes the true positive rate and the true negative rate.
[0058] In reality, the area under the curve of each distribution profile is located on both sides of the intersection. In fact, as shown in Figure 3, the maximized area under the curve of each distribution profile is achieved at the intersection 320 defined by the ranking value "9." In relation to the ranking value "9," the area under the curve of the handwritten text distribution profile 308 to the right of the "9" is maximized, while the area under the curve of the printed text distribution profile 316 to the left of the "9" is maximized. Other intersection points between the two distribution profiles, such as about 3, are not selected because they do not result in the maximized area under the curve.
[0059] Returning now to Figure 2C, in step 220 of sub-process 214, the assigned average ranking value may be compared to the threshold value obtained in step 218 of sub-process 214. Referring briefly to Figure 3, this comparison may be a binary classification type problem, where the average ranking value may be evaluated to determine whether it is above or below a threshold. Generally, in binary classification, the class prediction for each instance is often made based on a variable X, which is a "score" calculated for that instance. Given a threshold T, if X>T, the instance is classified as "positive," and otherwise as "negative."
[0060] Thus, based on the comparison, and referring again to FIG. 2B, the classifier server device 108 classifies the text zone as either one of a handwritten text zone or a printed text zone at step 222 of method 200 or step 224 of method 200, respectively.
[0061] If the comparison in step 220 of sub-process 214 indicates that the ranking value of the text zone is greater than the threshold, the text zone is classified as a handwritten text zone in step 222 of method 200 and assigned an identifier indicating that the zone is a handwritten text zone. The identifier may, for example, be part of software code inserted into the text zone to indicate that the text zone is a handwritten text zone.
[0062] If the comparison in step 220 of sub-process 214 indicates that the ranking value of the text zone is less than the threshold, the text zone is classified as a printed text zone in step 222 of method 200 and assigned an identifier indicating that the zone is a printed text zone. The identifier may, for example, be part of software code inserted into the text zone to indicate that the text zone is a printed text zone.
[0063] Thus, in step 226 of method 200, the classifier server device 108 generates a modified version of the digital document 104, the modified version of the digital document 104 including the identifiers assigned to its text zones.
[0064] After modifying digital document 104 according to the method of Figures 2A-2C, the modified digital document may be processed according to method 400 of Figure 4. For purposes of illustration, it may be assumed that the modified digital document includes four text zones, two of which are classified as handwritten text zones and assigned corresponding first identifiers indicating the same, and two of which are classified as printed text zones and assigned corresponding second identifiers indicating the same.
[0065] 4 is a flowchart illustrating a method 400 for generating selectable versions of a document based on a character recognition process, according to one embodiment of the present disclosure. The method begins by accessing a modified version of the digital document 104 stored in the database storage device 110. In one embodiment, the modified version of the digital document 104 may be pre-stored by the classifier server device 108 in step 226 of the method 200 of FIG. 2B. In one embodiment, the modified version of the digital document 104 may be received by the classifier server device 108 from the user device 102(1) and / or may be stored locally on the classifier server device 108. In one embodiment, the modified version of the digital document 104 may be generated by the user device 102(1) and stored in the database storage device 110.
[0066] In one embodiment, modified versions of the digital document 104 may be received from the user devices 102(1)-102(n) by the classifier server device 108 and / or may be stored locally on the classifier server device 108. In one embodiment, modified versions of the digital document 104 may be generated by the user devices 102(1)-102(n) and stored in the database storage device 110.
[0067] 4, the classifier server device 108 identifies a first identifier assigned to a corresponding text zone. In step 404 of the method 400, the classifier server device 108 determines that the first identifier is an identifier of a handwritten text zone and therefore identifies that an ICR algorithm should be applied to recognize the text in the handwritten text zone. Accordingly, the classifier server device 108 applies the ICR algorithm to the handwritten text zone in the modified version of the digital document 104. Similarly, in step 406 of the method 400, the classifier server device 108 identifies a second identifier assigned to the corresponding text zone. In step 408 of the method 400, the classifier server device 108 determines that the second identifier is an identifier of a printed text zone and therefore identifies that an OCR algorithm should be applied to recognize the text in the printed text zone. Accordingly, the classifier server device 108 applies the OCR algorithm to the printed text zone in the modified version of the digital document 104.
[0068] The classifier server device 108 then generates a selectable version of the digital document 104 at step 410 of method 400 based on the identified and recognized text zones. The selectable version of the digital document 104 facilitates selection of its text objects by the user devices 102(1)-102(n). The selectable version of the digital document 104 includes handwritten and printed characters that are selectable by the user devices 102(1)-102(n). The selectable version of the digital document 104 may then be optionally transmitted to the user devices at step 412 of method 400. For example, the selectable version of the digital document 104 may be transmitted to any of the user devices 102(1)-102(n). The user devices 102(1)-102(n) may apply a handwriting recognition algorithm, pre-stored on the user devices 102(1)-102(n), to all zones in the modified version of the digital document 104. In this embodiment, the handwriting recognition algorithm identifies zones corresponding to the first identifier inserted into the modified version of the document 104. Upon identifying the zones corresponding to the first and second identifiers, the handwriting recognition algorithm recognizes characters corresponding to those zones assigned the first identifier, and the printed character recognition algorithm recognizes characters corresponding to those zones assigned the second identifier. Thus, the handwritten and printed characters are then converted to be selectable by the user devices 102(1)-102(n) via the user interfaces corresponding to the user devices 102(1)-102(n).
[0069] An advantage of classifying text zones as either handwritten or printed text zones is that the classifier server device 108 may utilize an ICR algorithm to recognize characters within the handwritten text zones of the digital document 104, and separately utilize an OCR algorithm to recognize characters within the printed text zones of the digital document 104. Thus, each recognition algorithm can be efficiently applied based on the requirements of a given text zone, as indicated by its assigned identifier.
[0070] Understanding that neither a handwritten character recognition algorithm (e.g., the ICR described above) nor an OCR algorithm for printed text objects is applied to every zone throughout the digital document 104, the present disclosure provides a technical advantage of efficiently processing a document by applying an ICR algorithm to only specific zones of the document. This technique recognizes handwritten and printed characters without applying both an ICR algorithm and an OCR algorithm to every zone throughout the digital document 104, and because every character throughout the document is not processed by both an ICR algorithm and an OCR algorithm, this technique significantly reduces the time required to analyze every character in the digital document. Thus, this technique provides the advantage of recognizing characters in a document only once, providing a highly efficient, fast process of character recognition, especially in environments where a large number of characters are processed.
[0071] Thus, the present disclosure not only achieves more accurate and robust results, but also saves computational resources / processing power, since the entire document (all zones) does not need to be processed by both handwritten character recognition algorithms (e.g., ICR) and printed character recognition algorithms (e.g., OCR).
[0072] In one embodiment, the classifier server device 108 may perform the operations associated with FIGS. 2A-2C sequentially, randomly, or in combination with each other based on the following conditions before reaching step 226 of method 200:
[0073] In a first situation, according to one embodiment, at least one parameter value may be a ratio of object height to object font stroke width, and the classifier server device 108 may determine the zone parameter value as an average of the ratios of object height to object font stroke width. When the zone parameter value exceeds a threshold, the classifier server device 108 determines that all objects in the zone are handwritten text objects and performs step 222 of method 200 to classify the text zone as a handwritten text zone. Conversely, when the zone parameter value does not exceed the threshold, the classifier server device 108 determines that all objects in the zone are printed text objects and performs step 224 of method 200 to classify the text zone as a printed text zone.
[0074] In a second situation, when the zone parameter value may be object density, the classifier server device 108 determines an average object density as the zone parameter value based on the ratio of the number of objects in the text zone to the multiplication of the text zone width and the text zone height. When the zone parameter value exceeds a threshold, the classifier server device 108 determines that all objects in the text zone are handwritten text objects and performs step 222 of method 200 to classify the text zone as a handwritten text zone. Conversely, when the zone parameter value does not exceed the threshold, the classifier server device 108 determines that all objects in the text zone are printed text objects and performs step 224 of method 200 to classify the text zone as a printed text zone.
[0075] In a third situation, the classifier server device 108 identifies a text zone that has been classified as a handwritten text zone during the first and second situations described above and labels the text zone as a reliable handwritten text zone.
[0076] In a fourth situation, the classifier server device 108 identifies a text zone that has been classified as a printed text zone during the first and second situations described above and labels the text zone as a reliable printed text zone.
[0077] In a fifth situation, the classifier server device 108 identifies unreliable text zones taking into account the third and fourth situations described above. The unreliable text zones are text zones that are not classified as handwritten text zones during either the first situation or the second situation, and text zones that are not classified as printed text zones during either the first situation or the second situation. Therefore, upon identifying unreliable zones, the classifier server device 108 identifies adjacent zones surrounding each unreliable zone. Upon identifying adjacent zones surrounding an unreliable zone, the classifier server device 108 evaluates whether a majority of the adjacent zones are reliable handwritten text zones. If the classifier server device 108 determines that a majority of the adjacent zones are reliable handwritten text zones, it labels the unreliable zone as a reliable handwritten text zone. However, if it determines that a majority of the adjacent zones are not reliable handwritten text zones, it labels the unreliable zone as a reliable printed text zone.
[0078] 6 is a flowchart illustrating a method for character recognition by character recognition system 100 using a trained machine learning algorithm, according to one embodiment of the present disclosure. The method begins when classifier server device 108 receives a digital document 104 from user device 102(1). Digital document 104 may be generated by user device 102(1) by converting a physical document (e.g., a page containing text and / or graphics) into digital document 104, or by creating digital document 104 in any format, such as pdf, image, word, or other digital format known to those skilled in the art.
[0079] In one embodiment, digital document 104 may be received from any of user devices 102(1)-102(n). Additionally, digital document 104 may be generated by any of user devices 102(1)-102(n).
[0080] In one embodiment, the classifier server device 108 trains the machine learning algorithm by utilizing training test data and training result data pre-stored on the classifier server device 108 and / or by accessing training test data and training result data stored in the database storage device 110. The training test data includes a set of digital documents, each of which includes a text zone. Further, the training result data includes classification results associated with each of the documents in the set of digital documents included in the training test data. The training result data stores classification results corresponding to the actual classification of zones (classified as either handwritten text zones or printed text zones) in the documents of the set of digital documents.
[0081] The classifier server device 108 performs training of the machine learning algorithm by classifying the training test data based on the operations described in Figures 2A-4. After the classifier server device 108 classifies the training test data, the classifier server device 108 utilizes the machine learning algorithm to verify whether the classification performed on the training test data is accurate. The verification of the classification on the training test data is performed by comparing the classification results of the zones included in the set of documents in the training test data with the classification results of the corresponding zones included in the training result data.
[0082] If, based on the comparison, the classifier server device 108 determines that the result is inaccurate, the classifier server device 108 identifies the inaccuracy and therefore trains the machine learning algorithm to make accurate classification predictions in the future. Further, if, based on the comparison, the classifier server device 108 determines that the result is accurate, the classifier server device 108 identifies the accuracy and trains the machine learning algorithm to make similar accurate classification predictions in the future.
[0083] At step 602 of method 600, the classifier server device 108 receives a request for character recognition of the digital document 104 from the user device 102(1). In one embodiment, the user may use a user interface displayed on the user device 102(1) to upload the digital document 104 to the classifier server device 108 and may further select a button on the user interface of the user device 102(1) to send a request to the classifier server device 108 and initiate the character recognition process for the digital document 104. The classifier server device 108 receives the digital document 104 along with instructions to initiate the character recognition process for the digital document 104. Thus, upon receiving the digital document 104, the classifier server device 108 initiates the character recognition process by performing segmentation on the digital document 104, as described in step 604 below.
[0084] In one embodiment, the classifier server device 108 receives a request for character recognition of a digital document 104 from any of the user devices 102(1)-102(n).
[0085] In one embodiment, digital document 104 may be a document that includes text zones and non-text zones. The text zones may include handwritten and / or printed text, but may also include any other type of text data. The non-text zones may include graphic zones, magnetic ink character recognition (MICR), machine-readable zones (MRZ), optical mark recognition (OMR), images, diagrams, drawings, and / or tables, but may also include any other type of non-text data.
[0086] In one embodiment, digital document 104 may be a document containing handwritten text. In one embodiment, digital document 104 may be a document containing printed text. In one embodiment, digital document 104 may be a document containing handwritten text as well as printed text.
[0087] In step 604 of method 600, the classifier server device 108 segments the digital document 104 into multiple zones by applying one or more segmentation algorithms to the digital document 104. This operation of segmenting the digital document 104 into multiple zones by applying one or more segmentation algorithms may be similar to step 204 of FIG. 2A, described in detail above.
[0088] At step 606 of method 600, the classifier server device 108 classifies each of the plurality of zones as a text zone or a non-text zone. Classification by the classifier server device 108 may be performed by known image classification techniques, including, but not limited to, artificial neural networks (e.g., convolutional neural networks), linear regression, logistic regression, decision trees, support vector machines, random forests, naive Bayes, and k-nearest neighbors.
[0089] In step 608 of method 600, the classifier server device 108 analyzes the image pixels associated with each of the characters in the text zone and applies a trained machine learning algorithm that may have been pre-trained to classify the text zone as a handwritten text zone or a machine-printed text zone. Machine learning algorithms include neural network algorithms trained to segment documents, convolutional neural network algorithms trained to segment documents, deep learning algorithms trained to segment documents, and reinforcement learning algorithms trained to segment documents, but may also include any other neural network algorithm trained to classify a zone as a handwritten text zone or a machine-printed text zone.
[0090] The classifier server device 108 utilizes a machine learning algorithm to analyze each image pixel corresponding to each character in the text zone as input and determine whether each character can be a handwritten character or a machine-printed character. If the classifier server device 108 determines that a character is a handwritten character, the classifier server device 108 increments a first output pin value corresponding to the handwritten character. Furthermore, if the classifier server device 108 determines that a character is a machine-printed character, the classifier server device 108 increments a second output pin value corresponding to the machine-printed character.
[0091] Upon analyzing each image pixel corresponding to each character in the text zone, the classifier server device 108 compares the first output pin value corresponding to the handwritten character with the second output pin value corresponding to the machine-printed character. If the first output pin value corresponding to the handwritten character has a value greater than the second output pin value corresponding to the machine-printed character, the classifier server device 108 classifies the text zone as a handwritten text zone, and method 600 proceeds to step 610. If the first output pin value corresponding to the handwritten character has a value less than the second output pin value corresponding to the machine-printed character, the text zone may be classified as a machine-printed text zone, and method 600 proceeds to step 612.
[0092] At step 610 of method 600, the classifier server device 108 classifies the text zone as a handwritten text zone and assigns an identifier to the text zone, as described above with reference to Figures 2A-2C. The identifier may be a label indicating that the text zone is a handwritten text zone. The identifier may be part of software code inserted into each text zone of the document 104 to indicate that the text zone may be a handwritten text zone.
[0093] At step 612 of method 600, the classifier server device 108 classifies the text zone as a printed text zone and assigns an identifier to the text zone, as described above with reference to Figures 2A-2C. The identifier may be a label that indicates that the text zone is a printed text zone. The identifier may be part of software code inserted into each text zone of the document 104 to indicate that the text zone may be a printed text zone.
[0094] In step 614 of method 600, once the classifier server device 108 has performed the operation of classifying each of the identified text zones in the digital document 104 as either a handwritten text zone or a printed text zone, the classifier server device 108 generates a modified version of the digital document 104 and stores the modified version of the digital document 104 including the identifiers assigned to the text zones in the database storage device 110 and / or stores the modified version of the digital document 104 locally on the classifier server device 108. This modified version of the digital document 104 including the identifiers may be further utilized as described above in method 400 of FIG.
[0095] In one embodiment, as described with reference to step 608 of method 600, the classifier server device 108 receives as input the outlines of objects or characters corresponding to each of the text zones. The classifier server device 108 utilizes a machine learning algorithm to analyze the outlines of each character in the text zone and determine whether each character may be a handwritten character or a machine-printed character. The classifier server device 108 determines whether each character is a handwritten character and, for each character determined as a handwritten character, increments a first output pin value corresponding to the handwritten character. Furthermore, the classifier server device 108 determines whether each character is a printed character and, for each character determined as a machine-printed character, increments a second output pin value corresponding to the machine-printed character.
[0096] Upon analyzing the contours corresponding to each character in the text zone, the classifier server device 108 compares the first output pin value corresponding to the handwritten character with the second output pin value corresponding to the machine-printed character. If the classifier server device 108 determines that the first output pin value corresponding to the handwritten character has a value that may be greater than the second output pin value corresponding to the machine-printed character, the classifier server device 108 classifies the zone as a handwritten text zone, and method 600 proceeds to step 610, as described above. If the classifier server device 108 determines that the first output pin value corresponding to the handwritten character has a value that may be less than the second output pin value corresponding to the machine-printed character, the classifier server device 108 classifies the zone as a machine-printed text zone, and method 600 proceeds to step 612, as described above.
[0097] In one embodiment, the operations of Figures 2A, 2B, 2C, 4, and 6 may be implemented in classifier server device 108. In one embodiment, the operations of Figures 2A, 2B, 2C, 4, and 6 may be implemented by installing a software application ("app") on user devices 102(1)-102(n), such that the operations of Figures 2A, 2B, 2C, and 4 are implemented by the software application installed on user devices 102(1)-102(n) without communicating with any external network devices.
[0098] In one embodiment, the operations of FIGS. 2A, 2B, 2C, 4, and 6 may be implemented by installing software applications on user devices 102(1)-102(n) such that the operations of FIGS. 2A, 2B, 2C, 4, and 6 are implemented by the software applications installed on user devices 102(1)-102(n) while communicating with external network devices.
[0099] An advantage of classifying a zone as a handwritten text zone is that classifying a zone as a handwritten text zone causes the classifier server device 108 to utilize a handwritten character recognition algorithm that only recognizes the characters associated with those zones classified as handwritten text zones. This provides an efficient method of recognizing characters because the handwritten character recognition algorithm only recognizes the characters associated with those zones classified as handwritten text zones and does not process all the characters in the entire document. Furthermore, classifying a zone as a printed text zone causes the classifier server device 108 to apply a printed character recognition algorithm that only recognizes the characters associated with those zones classified as printed text zones. This provides an efficient method of recognizing characters because the printed character recognition algorithm only recognizes the characters associated with those zones classified as printed text zones and does not process all the characters in the entire document.
[0100] Because handwriting recognition algorithms are not applied to an entire document to identify handwritten characters, and because printed character recognition algorithms are not applied to an entire document to identify printed characters, the present disclosure provides a technical advantage of efficiently processing a document by applying a handwriting recognition algorithm or a printed character recognition algorithm only to specific zones of the document, instead of applying both to the entire document to recognize handwritten characters and printed characters, respectively. Such a technique not only achieves more accurate and robust results, but also saves computational resources / processing power, because the entire document (all zones) does not need to be processed by both a handwriting recognition algorithm (e.g., ICR) and a printed character recognition algorithm (e.g., OCR).
[0101] Therefore, this technique significantly reduces the time required to analyze every character in a digital document because every character in the entire document is not processed by both the ICR and OCR algorithms. Therefore, this technique offers the advantage of recognizing characters in a document only once, providing a very efficient, fast process for character recognition, especially in environments where a large amount of characters are processed.
[0102] Each of the functions of the described embodiments may be implemented by one or more processing circuits (also referred to as controllers). A processing circuit includes a programmed processor (e.g., CPU 700 of FIG. 7), since a processor includes circuitry. A processing circuit may also include devices such as application specific integrated circuits (ASICs) and circuit components arranged to perform the recited functions. The processing circuit may be part of the classifier server device 108, as discussed in more detail with respect to FIG. 7.
[0103] 7 is a detailed block diagram illustrating an exemplary classifier server device 108 according to certain embodiments of the present disclosure. In FIG. 7, the classifier server device 108 includes a CPU 700 and a query manager application 750. In one embodiment, the classifier server device 108 includes a database storage device 110 coupled to a storage controller 724. In one embodiment, the database storage device 110 may be a separate, individual (external) device and may be accessed by the classifier server device 108 via a network 720 (network 106 of FIG. 1 ).
[0104] The CPU 700 performs the processes described in this disclosure. Process data and instructions may be stored in memory 702. These processes and instructions (discussed with respect to FIGS. 2A-6) may also be stored on a storage medium disk 704, such as a hard drive (HDD) or portable storage medium, or may be stored remotely.
[0105] Furthermore, the discussed features of the present disclosure may be provided as utility applications, background daemons, or operating system components, or combinations thereof, running in combination with CPU 700 and an operating system such as Microsoft Windows or other versions, UNIX, Solaris, LINUX, Apple MAC-OS, and other systems known to those skilled in the art.
[0106] The hardware elements for implementing the operation of the classifier server device 108 may be implemented by a variety of circuit elements known to those skilled in the art. For example, the CPU 700 may be an Intel Xenon or Core processor, or an AMD Opteron processor, or other processor types as would be recognized by one skilled in the art.
[0107] The classifier server device 108 of FIG. 7 also includes a network controller 706, such as an Intel Ethernet PRO network interface card from Intel Corporation, for interfacing with a network 720. As can be appreciated, the network 720 can be a public network, such as the Internet, or a private network, such as a LAN or WAN network, or any combination thereof, and can also include PSTN or ISDN subnetworks. The network 720 can also be wired, such as an Ethernet network, or wireless, such as a cellular network, including EDGE, 3G, and 4G wireless cellular systems. The wireless network can also be WiFi, Bluetooth, or any other known form of wireless communication. The classifier server device 108 can communicate with external devices, such as the database storage device 110 and the pool of user devices 102(1)-102(n), via the network 720.
[0108] The classifier server device 108 further includes a display controller 708, such as an NVIDIA GeForce GTX or Quadro graphics adapter from NVIDIA, USA, for interfacing with a display 770. An I / O interface 712 interfaces with a keyboard 714 and / or a mouse and a touchscreen panel 716, and / or interfaces separately from the display 770. Additionally, the classifier server device 108 may be connected to a pool of user devices 102(1)-102(n) via the I / O interface 712 or through a network 720. The pool of user devices 102(1)-102(n) may send requests from the database storage device 110 via the storage controller 724 as queries that are processed by the query manager application 750, which may include retrieving data from memory 702, or triggering the execution of the processes discussed in FIGS. 2A-6.
[0109] The storage controller 724 connects to the storage media with a communications bus 726, which may be ISA, EISA, VESA, PCI, or the like, for interconnecting all of the components of the classifier server device 108. A description of the display 770, keyboard and / or mouse 714, and the general features and functionality of the display controller 708, storage controller 724, network controller 706, and I / O interface 712 will be omitted herein for the sake of brevity, as these features are known.
[0110] 7 may transmit or receive digital documents 104 to or from the pool of user devices 102(1)-102(n) over a network 720. When the classifier server device 108 receives a digital document 104 as part of a request to initiate a character recognition process, the classifier server device 108 stores the received digital document 104 in its memory 702. For example, the pool of user devices 102(1)-102(n) may transmit a modified version of the digital document 104 from the classifier server device 108 over the network 720, or the pool of user devices 102(1)-102(n) may receive a selectable version of the modified document from the classifier server device 108 over the network 720, or a camera 809 of the pool of user devices 102(1)-102(n) may capture an image of a physical document and transmit the image to the classifier server device 108. The pool of user devices 102(1)-102(n) may also implement one or more functions of the classifier server device 108 on the hardware of one of the pool of user devices 102(1)-102(n), for example, user device 102(1), as further shown in FIG. 8.
[0111] FIG. 8 is a detailed block diagram 800 illustrating an exemplary user device from the pool of user devices 102(1)-102(n). By way of example, FIG. 8 illustrates user device 102(1) in accordance with certain embodiments of the present disclosure. In certain embodiments, user device 102(1) may be a smartphone. However, those skilled in the art will understand that the features described herein may be adapted to be implemented in other devices (e.g., laptops, tablets, servers, electronic readers, cameras, navigation devices, etc.). The exemplary user device 102(1) includes a controller 810 and wireless communication processing circuitry 802 connected to an antenna 801. A speaker 804 and a microphone 805 are connected to audio processing circuitry 803.
[0112] The controller 810 may include one or more central processing units (CPUs) and may control elements within the user device 102(1) to perform functions related to communication control, audio signal processing, control for audio signal processing, still and video image processing and control, and other types of signal processing. The controller 810 may perform these functions by executing instructions stored in memory 850. For example, the processes shown in FIGS. 2A, 2B, 3, 4, and 5 may be stored in memory 850. As an alternative to or in addition to local storage in memory 850, functions may be performed using instructions stored on an external device accessed over a network or on a non-transitory computer-readable medium.
[0113] The user device 102(1) includes control lines CL and data lines DL as internal communication bus lines. Control data to / from the controller 810 can be transmitted over the control lines CL. The data lines DL can be used to transmit audio data, display data, etc.
[0114] The antenna 801 transmits and receives electromagnetic wave signals between base stations to conduct wireless-based communications, such as various forms of cellular telephone communications. The wireless communication processing circuitry 802 controls communications conducted between the user device 102(1) and other external devices, such as the classifier server device 108, via the antenna 801. The wireless communication processing circuitry 802 may control communications between base stations for cellular telephone communications.
[0115] The speaker 804 emits an audio signal corresponding to the audio data provided by the audio processing circuit 803. The microphone 805 detects ambient audio and converts the detected audio into an audio signal. The audio signal may then be output to the audio processing circuit 803 for further processing. The audio processing circuit 803 demodulates and / or decodes audio data read from the memory 850 or audio data received by the wireless communication processing circuit 802 and / or the short-range wireless communication processing circuit 807. In addition, the audio processing circuit 803 may decode the audio signal acquired by the microphone 805.
[0116] The exemplary user device 102(1) may also include a display 811, a touch panel 830, operation keys 840, and short-range communication processing circuitry 807 connected to the antenna 806. The display 811 may be a liquid crystal display (LCD), an organic electroluminescent display panel, or another display screen technology.
[0117] Touch panel 830 may include a physical touch panel display screen and a touch panel driver. Touch panel 830 may include one or more touch sensors for detecting input actions on the operating surface of the touch panel display screen.
[0118] For simplicity, this disclosure assumes that the touch panel 830 is a capacitive touch panel technology. However, it should be understood that aspects of this disclosure can be readily applied to other touch panel types (e.g., resistive touch panels) having alternative structures. In particular aspects of this disclosure, the touch panel 830 can include transparent electrode touch sensors arranged in the XY direction on the surface of a transparent sensor glass.
[0119] The operation keys 840 may include one or more buttons or similar external control elements that may generate operation signals based on detected input by a user. In addition to output from the touch panel 830, these operation signals may be provided to the controller 810 for performing associated processing and control. In certain aspects of the present disclosure, processing and / or functions associated with external buttons, etc. may be performed by the controller 810 in response to input operations on the display screen of the touch panel 830 rather than external buttons, keys, etc. In this manner, external buttons on the user device 800 may be eliminated in lieu of performing input via touch operations, thereby improving watertightness.
[0120] The antenna 806 may transmit and receive electromagnetic signals to / from other external devices, and the short-range wireless communication processing circuit 807 may control wireless communication between other external devices. Bluetooth, IEEE 802.11, and Near Field Communication (NFC) are non-limiting examples of wireless communication protocols that may be used for device-to-device communication via the short-range wireless communication processing circuit 807.
[0121] The user device 102(1) may include a camera 809 that includes a lens and a shutter for capturing a photograph of the surroundings around the user device 102(1). In one embodiment, the camera 809 captures the surroundings on the opposite side of the user device 102(1) from the user. The captured photographic image may be displayed on the display panel 811. A memory circuit stores the captured photograph. The memory circuit may be within the camera 809 or may be part of the memory 850. The camera 809 may be a separate feature attached to the user device 102(1) or may be an integrated camera feature.
[0122] The user device 102(1) may include an application that requests data processing from the classifier server device 108 over the network 720.
[0123] In the above description, any process, illustration, or block in the flowchart should be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or step in the process; alternative implementations are included within the scope of the exemplary embodiments of the invention in which functions may be performed in a different order than shown or discussed, including substantially simultaneously or in reverse order, depending on the functionality involved, as will be understood by those skilled in the art. The various elements, features, and processes described herein may be used independently of one another or may be combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure.
[0124] While specific embodiments have been described, these embodiments are presented by way of example only and are not intended to limit the scope of the present disclosure.
[0125] Indeed, the novel methods, apparatus, and systems described herein may be embodied in a variety of other forms, and various omissions, substitutions, and changes in the forms of the methods, apparatus, and systems described herein may be made without departing from the spirit of the present disclosure. The accompanying claims and their equivalents are intended to cover such forms or modifications that would fall within the scope and spirit of the present disclosure. For example, the technology may be configured for cloud computing, in which a single function is shared and processed collaboratively among multiple devices over a network.
[0126] The devices, methods, and computer-readable media discussed herein are examples. Various embodiments may omit, substitute, or add various procedures or components, as appropriate. For example, features described with respect to particular embodiments may be combined in various other embodiments. Different aspects and elements of the embodiments may be combined in a similar manner. Various components of the diagrams provided herein may be embodied in hardware and / or software. Also, technology evolves, and thus many of the elements are examples that do not limit the scope of the disclosure to those specific examples.
[0127] The methods, apparatus, and devices discussed herein are examples. Various embodiments may omit, substitute, or add various procedures or components, as appropriate. For example, features described with respect to particular embodiments may be combined in various other embodiments. Different aspects and elements of the embodiments may be combined in a similar manner. Various components of the diagrams provided herein may be embodied in hardware and / or software. Also, technology evolves, and thus many of the elements are examples that do not limit the scope of the disclosure to those specific examples.
[0128] Obviously, many modifications and variations are possible in light of the above teachings. It is therefore to be understood that, within the scope of the appended claims, the present disclosure may be practiced otherwise than as specifically described herein.
[0129] Embodiments of the present disclosure may also be as described in the insert below.
[0130] (1) A method for recognizing characters in a digital document, the method including: classifying, by a processing circuit, a segment of the digital document as containing text; calculating, by the processing circuit, at least one parameter value associated with the classified segment of the digital document; determining, by the processing circuit, a zone parameter value based on the calculated at least one parameter value; classifying, by the processing circuit, the segment of the digital document as a handwritten text zone or a printed text zone based on the determined zone parameter value and a threshold value, wherein the threshold value is based on selecting an intersection of a handwritten text distribution profile and a printed text distribution profile, each of the handwritten text distribution profile and the printed text distribution profile being associated with a zone parameter that corresponds to the determined zone parameter value; and generating, by the processing circuit, a modified version of the digital document based on the classification.
[0131] (2) The method of (1), further comprising: performing handwriting character recognition on the segment to recognize handwritten text when the segment is classified by the processing circuitry as a handwritten text zone; and performing printed character recognition on the segment to recognize printed text when the segment is classified by the processing circuitry as a printed text zone.
[0132] (3) A method according to (1) or (2), wherein the determined zone parameter value is an object density value calculated by the processing circuit by multiplying the segment ratio by the segment height, the segment ratio being the ratio of the number of objects in the segment to the segment width.
[0133] (4) A method according to any one of (1) to (3), wherein generating a modified version of the digital document includes assigning a handwritten text identifier to the segment by the processing circuit when the segment is classified as a handwritten text zone, and assigning a printed text identifier to the segment by the processing circuit when the segment is classified as a printed text zone.
[0134] (5) A method according to any one of (1) to (4), further comprising generating, by the processing circuitry, a first selectable version of the digital document as a modified version of the digital document, wherein the first selectable version of the digital document includes the recognized handwritten text in the segment.
[0135] (6) The method of any one of (1) to (5), further comprising generating, by the processing circuitry, a second selectable version of the digital document as a modified version of the digital document, wherein the second selectable version of the digital document includes the recognized printed text in the segment.
[0136] (7) The method of any one of (1) to (6), further comprising: retrieving, by the processing circuit, threshold values from a lookup table containing threshold values based on the type of zone parameter value; each of the threshold values being based on a corresponding handwritten text distribution profile and a corresponding printed text distribution profile from the database; the corresponding handwritten text distribution profile being based on a histogram associated with the text zone labeled as handwritten text; and the corresponding printed text distribution profile being based on a histogram associated with the text zone labeled as printed text.
[0137] (8) A method according to any one of (1) to (7), wherein each of the histograms is based on a comparison of parameter values of objects in the labeled text zone with corresponding rankings of the objects, and the corresponding rankings are based on the parameter values of the object with the parameter values of each other object in the labeled text zone.
[0138] (9) A method according to any one of (1) to (8), wherein selecting the intersection of the handwritten text distribution profile and the printed text distribution profile includes maximizing the area under the curve for each of the handwritten text distribution profile and the printed text distribution profile by a processing circuit, and wherein a threshold is a value above which the segment is likely to be handwritten text.
[0139] (10) A method according to any one of (1) to (9), wherein the classifying includes: calculating, by a processing circuit, a zone ranking value based on the determined zone parameter values and a ranking value associated with each of the calculated at least one parameter value; and comparing, by the processing circuit, the zone ranking value with a threshold value, wherein the segment of the digital document is classified as a handwritten text zone when the zone ranking value satisfies the threshold value.
[0140] (11) A non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for recognizing characters in a digital document, the method including: classifying a segment of the digital document as containing text; calculating at least one parameter value associated with the classified segment of the digital document; determining a zone parameter value based on the calculated at least one parameter value; classifying the segment of the digital document as a handwritten text zone or a printed text zone based on the determined zone parameter value and a threshold value, wherein the threshold value is based on selecting an intersection of a handwritten text distribution profile and a printed text distribution profile, each of the handwritten text distribution profile and the printed text distribution profile being associated with a zone parameter that corresponds to the determined zone parameter value; and generating a modified version of the digital document based on the classification.
[0141] (12) The non-transitory computer-readable medium of (11), further comprising: performing handwriting character recognition on the segment to recognize handwritten text when the segment is classified as a handwritten text zone; and performing printed character recognition on the segment to recognize printed text when the segment is classified as a printed text zone.
[0142] (13) A non-transitory computer-readable medium as described in (11) or (12), wherein the determined zone parameter value is an object density value calculated by multiplying the segment ratio by the segment height, and the segment ratio is the ratio of the number of objects in the segment to the segment width.
[0143] (14) A non-transitory computer-readable medium described in any one of (11) to (13), wherein generating a modified version of the digital document includes assigning a handwritten text identifier to the segment when the segment is classified as a handwritten text zone, and assigning a printed text identifier to the segment when the segment is classified as a printed text zone.
[0144] (15) A non-transitory computer-readable medium described in any one of (11) to (14), further comprising generating a first selectable version of the digital document as a modified version of the digital document, the first selectable version of the digital document including the recognized handwritten text in the segments and the recognized printed text in the segments.
[0145] (16) A non-transitory computer-readable medium according to any one of (11) to (15), further comprising obtaining threshold values from a lookup table containing threshold values based on a type of zone parameter value, each of the threshold values being based on a corresponding handwritten text distribution profile and a corresponding printed text distribution profile from a database, the corresponding handwritten text distribution profile being based on a histogram associated with a text zone labeled as handwritten text, and the corresponding printed text distribution profile being based on a histogram associated with a text zone labeled as printed text.
[0146] (17) A non-transitory computer-readable medium described in any one of (11) to (16), wherein each of the histograms is based on a comparison of parameter values of objects in a labeled text zone with corresponding rankings of the objects, and the corresponding rankings are based on the parameter values of the object with the parameter values of each other object in the labeled text zone.
[0147] (18) A non-transitory computer-readable medium according to any one of (11) to (17), wherein selecting the intersection of the handwritten text distribution profile and the printed text distribution profile includes maximizing the area under the curve for each of the handwritten text distribution profile and the printed text distribution profile, and wherein a threshold is a value above which the segment is likely to be handwritten text.
[0148] (19) A non-transitory computer-readable medium according to any one of (11) to (18), wherein the classifying includes calculating a zone ranking value based on the determined zone parameter value and a ranking value associated with each of the calculated at least one parameter value, and comparing the zone ranking value with a threshold value, wherein the segment of the digital document is classified as a handwritten text zone when the zone ranking value satisfies the threshold value.
[0149] (20) An apparatus for recognizing characters in a digital document, the apparatus comprising a processing circuit configured to: classify a segment of the digital document as containing text; calculate at least one parameter value associated with the classified segment of the digital document; determine a zone parameter value based on the calculated at least one parameter value; classify the segment of the digital document as a handwritten text zone or a printed text zone based on the determined zone parameter value and a threshold value, wherein the threshold value is based on selecting an intersection of a handwritten text distribution profile and a printed text distribution profile, each of the handwritten text distribution profile and the printed text distribution profile being associated with a zone parameter that corresponds to the determined zone parameter value; and generate a modified version of the digital document based on the segment classification.
[0150] Thus, the foregoing discussion discloses and describes merely exemplary embodiments of the present disclosure. As will be understood by those skilled in the art, the present disclosure may be embodied in other specific forms without departing from its spirit. Accordingly, the disclosure of the present disclosure is intended to be illustrative, but not limiting, of the scope of the disclosure and other claims. This disclosure, including any readily discernible variations of the teachings herein, defines, in part, the scope of the foregoing claim terms so that the subject matter of the invention is not dedicated to the public.
Claims
1. 1. A method for recognizing characters in a digital document, said method comprising: classifying, by a processing circuit, a segment of the digital document as containing text; calculating, by the processing circuitry, at least one parameter value associated with the classified segment of the digital document; determining, by the processing circuitry, a zone parameter value based on the calculated at least one parameter value; classifying, by the processing circuitry, the segment of the digital document as a handwritten text zone or a printed text zone based on the determined zone parameter value and a threshold value, the threshold value being based on a selection of an intersection of a handwritten text distribution profile and a printed text distribution profile, each of the handwritten text distribution profile and the printed text distribution profile being associated with a zone parameter corresponding to the determined zone parameter value; generating, by the processing circuitry, a modified version of the digital document based on the classifying; The method wherein the determined zone parameter value is an object density value calculated by the processing circuitry by multiplying a segment ratio by a segment height, the segment ratio being the ratio of the number of objects in the segment to the segment width.
2. performing handwriting recognition on the segment to recognize handwritten text when the segment is classified by the processing circuit as a handwritten text zone; 2. The method of claim 1, further comprising: when the processing circuit classifies the segment as a printed text zone, performing printed character recognition on the segment to recognize printed text.
3. said generating said modified version of said digital document comprises: assigning a handwritten text identifier to the segment when the processing circuit classifies the segment as a handwritten text zone; and assigning, by the processing circuitry, a printed text identifier to the segment when the segment is classified as a printed text zone.
4. 3. The method of claim 2, further comprising generating, by the processing circuitry, a first selectable version of the digital document as the modified version of the digital document, the first selectable version of the digital document including the recognized handwritten text in the segment.
5. 3. The method of claim 2, further comprising generating, by the processing circuitry, a second selectable version of the digital document as the modified version of the digital document, the second selectable version of the digital document including the recognized printed text in the segment.
6. further comprising: retrieving, by the processing circuitry, the thresholds from a lookup table containing thresholds based on the type of the zone parameter value, each of the thresholds being based on a corresponding handwritten text distribution profile and a corresponding printed text distribution profile from a database; the corresponding handwritten text distribution profile is based on a histogram associated with a text zone labeled as handwritten text; The method of claim 1 , wherein the corresponding printed text distribution profile is based on a histogram associated with a text zone labeled as printed text.
7. 7. The method of claim 6, wherein each of the histograms is based on a comparison of parameter values of objects in a labeled text zone with a corresponding ranking of the objects, the corresponding ranking being based on the parameter values of the object and parameter values of each other object in the labeled text zone.
8. The selection of the intersection of the handwritten text distribution profile and the printed text distribution profile comprises:
2. The method of claim 1, further comprising maximizing, by the processing circuitry, an area under a curve for each of the handwritten text distribution profile and the printed text distribution profile, the threshold being a value above which the segment is likely to be handwritten text.
9. The classification is calculating, by the processing circuitry, a zone ranking value based on the determined zone parameter values and a ranking value associated with each of the calculated at least one parameter value; and comparing, by the processing circuitry, the zone ranking value to a threshold value, wherein when the zone ranking value satisfies the threshold value, the segment of the digital document is classified as a handwritten text zone.
10. 1. A non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for recognizing characters in a digital document, the method comprising: classifying a segment of the digital document as containing text; calculating at least one parameter value associated with the classified segments of the digital document; determining a zone parameter value based on the calculated at least one parameter value; classifying the segment of the digital document as a handwritten text zone or a printed text zone based on the determined zone parameter value and a threshold value, the threshold value being based on selecting an intersection of a handwritten text distribution profile and a printed text distribution profile, each of the handwritten text distribution profile and the printed text distribution profile being associated with a zone parameter corresponding to the determined zone parameter value; generating a modified version of the digital document based on the classifying; a non-transitory computer-readable medium, wherein the determined zone parameter value is an object density value calculated by multiplying a segment ratio by a segment height, the segment ratio being the ratio of the number of objects in the segment to the segment width.
11. when the segment is classified as a handwritten text zone, performing handwriting recognition on the segment to recognize handwritten text; 11. The non-transitory computer-readable medium of claim 10, further comprising: when the segment is classified as a printed text zone, performing printed character recognition on the segment to recognize printed text.
12. said generating said modified version of said digital document comprises: assigning a handwritten text identifier to the segment when the segment is classified as a handwritten text zone; and assigning a printed text identifier to the segment when the segment is classified as a printed text zone.
13. 12. The non-transitory computer-readable medium of claim 11, further comprising generating a first selectable version of the digital document as the modified version of the digital document, the first selectable version of the digital document including the recognized handwritten text in the segment and the recognized printed text in the segment.
14. obtaining the thresholds from a lookup table containing thresholds based on the type of the zone parameter value, each of the thresholds being based on a corresponding handwritten text distribution profile and a corresponding printed text distribution profile from a database; the corresponding handwritten text distribution profile is based on a histogram associated with a text zone labeled as handwritten text; The non-transitory computer-readable medium of claim 10 , wherein the corresponding printed text distribution profile is based on a histogram associated with a text zone labeled as printed text.
15. 15. The non-transitory computer-readable medium of claim 14, wherein each of the histograms is based on a comparison of parameter values of objects in a labeled text zone to a corresponding ranking of the objects, the corresponding ranking being based on the parameter values of the object and parameter values of each other object in the labeled text zone.
16. The selection of the intersection of the handwritten text distribution profile and the printed text distribution profile comprises:
11. The non-transitory computer-readable medium of claim 10, further comprising maximizing an area under a curve for each of the handwritten text distribution profile and the printed text distribution profile, the threshold being a value above which the segment is likely to be handwritten text.
17. The classification is calculating a zone ranking value based on the determined zone parameter values and a ranking value associated with each of the calculated at least one parameter value; and comparing the zone ranking value to a threshold value, wherein when the zone ranking value meets the threshold value, the segment of the digital document is classified as a handwritten text zone.
18. 1. An apparatus for recognizing characters in a digital document, said apparatus comprising: a processing circuit, the processing circuit classifying a segment of the digital document as containing text; calculating at least one parameter value associated with the classified segments of the digital document; determining a zone parameter value based on the calculated at least one parameter value; classifying the segment of the digital document as a handwritten text zone or a printed text zone based on the determined zone parameter value and a threshold value, the threshold value being based on selecting an intersection of a handwritten text distribution profile and a printed text distribution profile, each of the handwritten text distribution profile and the printed text distribution profile being associated with a zone parameter corresponding to the determined zone parameter value; generating a modified version of the digital document based on the segment classification; the determined zone parameter value is an object density value calculated by the processing circuitry by multiplying a segment ratio by a segment height, the segment ratio being the ratio of the number of objects in the segment to the segment width.
Citation Information
Patent Citations
Image decision method, image processing unit, and image output unit
JP2007087196A
Method and apparatus for classifying machine printed text and handwritten text
US20150339543A1
Character recognition method, character recognition device, and character recognition program
WO2011074067A1