Optical Character Recognition Segmentation and Processing Method, Computer Program, Hardware Device (Optical Character Recognition Segmentation)

The method enhances OCR technology by segmenting documents into regions based on text types, removing noise, and applying specific OCR codes, effectively addressing the challenges of complex document layouts and improving text recognition accuracy.

JP7691178B2Active Publication Date: 2025-06-11INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021188156
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-24
Filing Date
2021-11-18
Publication Date
2025-06-11
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

Existing optical character recognition (OCR) technologies struggle to effectively segment documents containing multiple types of text data, such as printed text, handwritten text, and images, due to complex document layouts and the inability to trace the source of recognized text.

Method used

A method and system for improving OCR technology by detecting different types of text data in a document, dividing the document into multiple text regions, removing optical noise, and selecting and executing appropriate OCR software codes for each region to extract computer-readable text.

Benefits of technology

The solution enables accurate segmentation and recognition of diverse text types within complex documents, improving the overall efficiency and effectiveness of OCR processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691178000001
    Figure 0007691178000001
  • Figure 0007691178000002
    Figure 0007691178000002
  • Figure 0007691178000003
    Figure 0007691178000003
Patent Text Reader

Abstract

To provide a method, system, and computer program product for segmenting and processing documents for optical character recognition.SOLUTION: The method includes receiving a document and detecting different types of text data. The document is divided into a plurality of text regions associated with the different types of the text data. Optical noise is removed from each text region and differing optical character recognition software code is selected for application to each text region. The differing optical character recognition software code is executed with respect to each text region resulting in extractable computer readable text located within each text region.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to a method for segmenting a document for optical character recognition, and more particularly to a method for improving optical character recognition software technology associated with detecting, evaluating, and dividing an electronic document structure into a plurality of regions for extracting computer-readable text within the region of the electronic document structure, and related systems.

Summary of the Invention

Problems to be Solved by the Invention

[0002] The present invention generally relates to a method for segmenting a document for optical character recognition, and more particularly to a method for improving optical character recognition software technology associated with detecting, evaluating, and dividing an electronic document structure into a plurality of regions for extracting computer-readable text within the region of the electronic document structure, and related systems.

Means for Solving the Problems

[0003] A first aspect of the present invention provides an optical character recognition segmentation and processing method, the method comprising receiving, by a processor of a hardware device, a document to be processed; detecting, by the processor, different types of text data of the document; dividing, by the processor, the document into a plurality of text regions associated with different types of text data, wherein each text region of the plurality of text regions has a single type of text data; removing, by the processor, optical noise from each text region; selecting, by the processor, different optical character recognition software codes to be applied to each text region; and executing, by the processor, in response to the selection, different optical character recognition software codes for each text region to obtain extractable computer-readable text within each text region.

[0004] A second aspect of the present invention provides a computer program product comprising a computer-readable hardware storage device storing computer-readable program code, the computer-readable program code having an algorithm for implementing an optical character recognition segmentation and processing method, the method comprising receiving, by a processor, a document to be processed; detecting, by the processor, different types of text data in the document; dividing, by the processor, the document into a plurality of text regions associated with the different types of text data, each text region of the plurality of text regions having a single type of text data; removing, by the processor, optical noise from each text region; selecting, by the processor, different optical character recognition software codes to apply to each text region; and obtaining, by the processor, extractable computer-readable text within each text region by executing different optical character recognition software codes in response to the selection for each text region.

[0005] A third aspect of the present invention provides a hardware device comprising a processor coupled to a computer-readable memory unit, the memory unit having instructions which, when executed by the processor, implement an optical character recognition segmentation and processing method, the method comprising receiving, by the processor, a document to be processed; detecting, by the processor, different types of text data in the document; dividing, by the processor, the document into a plurality of text regions associated with different types of text data, each text region of the plurality of text regions having a single type of text data; removing, by the processor, optical noise from each text region; selecting, by the processor, different optical character recognition software codes to apply to each text region; and obtaining, by the processor, extractable computer-readable text within each text region by executing different optical character recognition software codes in response to the selection for each text region.

[0006] The present invention advantageously provides a simple method and related system that can accurately segment documents for optical character recognition.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

[0008]

Figure 3

[0009]

Figure 4

[0010]

Figure 5

[0011]

Figure 6

[0012]

Figure 7

[0013]

Figure 8

[0014]

Figure 9

[0015]

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0016] Figure 1 shows a system 100 for improving optical character recognition software technology associated with detecting, evaluating, and dividing an electronic document structure into multiple regions to extract computer-readable text within the region of the electronic document structure according to an embodiment of the present invention. A general optical character recognition (OCR) process may face complex document layout data / text types that appear simultaneously within the same document, particularly printed text, text of a curved seal, tilted text, tables, handwriting, etc. Therefore, since the relevant models and external interfaces suitable for the data / text types are also different, a single OCR recognition model may not have the ability to effectively recognize and process multiple data / text types. Thus, system 100 enables a process of detecting different data / text types to be divided, executing them via different models, and combining them to generate a final result. In addition, a general OCR algorithm may not functionally support the traceability of the recognized text (e.g., whether the recognized text is due to a form, text, handwriting, or seal). Therefore, system 100 is configured to identify multiple types of text from complex layout data and trace the source of the recognized text.

[0017] System 100 enables the following modules to be executed via a hardware device 139. 1. A cognitive OCR module that detects different types of layout data / text based on image semantic segmentation technology. 2. An OCR recognition algorithm dispatcher module that distributes different types of data / text to an optimal OCR recognition algorithm. 3. A noise removal module that removes noise based on a self-attention mechanism that can effectively remove the noise of the segmentation unit.

[0018] System 100 enables an advantageous process that enables an OCR component for complex data / text layouts. Similarly, System 100 enables a noise removal process based on a self-attention mechanism that effectively removes any noise in the segmentation unit. The text recognition process (enabled by System 100) is configured to recommend an optimal recognition algorithm to adapt to the recognition situation of complex fonts.

[0019] The system 100 of FIG. 1 includes a hardware device 139 (i.e., dedicated hardware), OCR software code 138 (e.g., within a data storage device), a database 151, and a network interface controller 153 interconnected through a network 117. The hardware device 139 includes a dedicated circuit 127 (which may include dedicated software), a sensor 112, and a machine learning software code / hardware structure 121 (i.e., including machine learning software code). The interface controller 153 may include any type of device or apparatus that securely interfaces hardware and software to the network. The OCR software code 138 comprises different types of software that perform different OCR processes with respect to different types of data and / or text or both of an electronic document. The sensor 112 may include any type of internal or external sensor, including in particular an ultrasonic three-dimensional sensor module, a temperature sensor, an ultrasonic sensor, an optical sensor, an image search device, an audio search device, a humidity sensor, a voltage sensor, a pressure sensor, etc. The hardware device 139 may comprise an embedded device. An embedded device is defined herein as a dedicated device or computer comprising a combination of computer hardware and software (fixed-capability or programmable) specifically designed to perform a dedicated function. A programmable embedded computer or device may comprise a dedicated programming interface. In one embodiment, the hardware device 139 may comprise a dedicated hardware device having dedicated (non-general-purpose) hardware and circuitry (i.e., dedicated discrete non-general-purpose analog, digital, and logic-based circuitry) that executes the processes described with respect to FIGS. 1 - 10 (either independently or in combination).Dedicated discrete non - general - purpose analog, digital, and logic - based circuits are designed only for implementing an automated process that improves optical character recognition software technology, which is associated with detecting, evaluating, and dividing an electronic document structure into multiple regions in order to extract computer - readable text within the region of the electronic document structure. The dedicated specially - designed components (e.g., dedicated integrated circuits, such as application - specific integrated circuits (ASICs), etc.) may be included. Network 117 may include any type of network, particularly including 5G telecommunication networks, local area networks (LANs), wide area networks (WANs), the Internet, wireless networks, etc. Alternatively, network 117 may include an application programming interface (API).

[0020] System 100 is enabled to execute an OCR recognition method based on pre - semantic segmentation including the following steps. 1. The cognitive OCR mechanism enables detecting different types of layout data / text based on image semantic segmentation technology. 2. The OCR recognition algorithm dispatcher enables distributing different types of layout data / text to the optimal OCR recognition algorithm. 3. Execute noise removal technology based on the self - attention mechanism used to effectively remove the noise of the segmentation unit. 4. Utilize the image semantic segmentation algorithm for different text regions and utilize the image noise removal algorithm based on the self - attention mechanism. 5. After the image semantic segmentation algorithm is executed, a classification operation is performed so that the classification category of the algorithm index of different recognition methods and each algorithm index are associated with a specific algorithm. Subsequently, the optimal recognition algorithm is selected for the relevant region. 6. To recommend classification types for different independent semantic units, a classification algorithm is utilized.

[0021] FIG. 2 shows an algorithm that details a process flow enabled by the system 100 of FIG. 1 to improve optical character recognition software technology related to detecting, evaluating, and dividing an electronic document structure into multiple regions to extract computer-readable text within the region of the electronic document structure according to an embodiment of the present invention. The steps in the algorithm of FIG. 2 may each be enabled and executed in any order by a computer processor that executes computer code. Additionally, the steps in the algorithm of FIG. 2 may each be enabled and executed in combination by the hardware device 139 and the OCR software code 138. At step 200, a document is received for processing (e.g., from the database 151 of FIG. 1). At step 202, different types of text data of the document are detected. The different types of text data may be constituted by formats such as, in particular, table format, watermark format, handwritten format, and rotated text format.

[0022] At step 204, the document is divided into a plurality of text regions associated with the different types of the above text data. Each text region contains a single type of text data. At step 208, for each text region, optical noise (e.g., unwanted background text and images of the document) is removed. The noise removal process may be executed via a self-attention mechanism executed via natural language processing. Removing optical noise from each text region may include the following. 1. Encoding each text region as a sum meaning vector of each text region. 2. Enabling a 3×3 window to divide each text region into a plurality of sub-regions such that each sub-region is encoded as a 1×300 vector. 3. Generate a dot product for each sub-region (based on the sum meaning vector). Each dot product has a specific score representing the importance level of each sub-region. 4. Compare each specific score with a score threshold. 5. Eliminate all pixels of each sub-region that exceed the score threshold (based on the result of the comparison) to perform noise removal.

[0023] In step 210, different optical character recognition software codes are selected and applied to each text region. Additionally, each text region may be classified with respect to a single type of text data such that the selection is performed based on the result of the classification. Each instance of the different optical character recognition software codes may include self-learning software codes stored in a dedicated database (e.g., database 151 in FIG. 1).

[0024] In step 212, different optical character recognition software codes are executed for each text region to obtain extractable computer-readable text within each text region. The extractable computer-readable text within each text region may be configured to be used with respect to cut or copy and paste functions.

[0025] Figure 3 shows an internal structure diagram of the machine learning software / hardware structure 121 (and / or circuit 127) of FIG. 1 according to an embodiment of the present invention. The machine learning software / hardware structure 121 includes a detection module 304, a document segmentation module 310, a noise removal module 308, a selection execution module 314, and a communication controller 302. The detection module 304 includes dedicated hardware and software that controls all functions related to the text data type detection stage of FIGS. 1 and 2. The document segmentation module 310 includes dedicated hardware and software that controls all functionality related to the text area segmentation functionality that implements the process described with respect to the algorithm of FIG. 2. The noise removal module 308 includes dedicated hardware and software that controls all functions related to the noise removal stage of FIG. 2. The selection execution module 314 includes dedicated hardware and software that controls all functions related to the selection and execution of OCR code. The communication controller 302 is enabled to control all communication between the detection module 304, the document segmentation module 310, the noise removal module 308, and the selection execution module 314.

[0026] Figure 4 shows an OCR segmentation process 400 according to an embodiment of the present invention. The OCR segmentation process 400 is executed to train an image semantic segmentation algorithm resource pool 407 for execution on different text regions 405a... 405d (including text portions 414a... 414d) from an original text page 402. A noise removal process is executed on the different (semantic) text regions 405a... 405d. Similarly, a classification operation is performed to obtain classification categories for algorithm indexes of different recognition methods. Each algorithm index is associated with a specific algorithm of algorithms 407a... 407d such that an optimal recognition algorithm is selected for each text region.

[0027] FIG. 5 shows a system 500 that executes an image semantic segmentation algorithm on different text regions 502a... 502d of a document 502 according to an embodiment of the present invention. The image semantic segmentation algorithm is executed via a position attention module 510. In addition, the system 500 is configured to execute an image noise removal algorithm based on the self-attention mechanism of a channel attention module 512. The image semantic segmentation algorithm uses a dual attention network (DANET) for scene segmentation, which is an open-source algorithm, to classify each pixel in the original document 502 (image). As a result of the final result, the original document 502 can be divided into different semantic regions 514a... 514d and coordinated for template area cutting. All target areas are cut into independent semantic units.

[0028] FIG. 6 shows a noise removal process according to an embodiment of the present invention. Before noise removal from the original picture 602, in many cases, additional elements of other semantic units are included within the semantic unit to be cut. The additional elements become noise (i.e., noise background 604) that interferes with the text recognition process. Therefore, the noise background 604 must be removed before the text recognition process is executed. The self-attention mechanism in natural language processing (NLP) may enable noise extraction through the following steps. 1. An encoder module is executed that encodes the original picture 602 as a 1×300 vector with the total semantic vector of the original picture 602. 2. A 3×3 window may be enabled to divide the original picture 602 into a plurality of sub-pictures. Each sub-picture is encoded as a 1×300 vector. 3. The sum meaning vector of the original picture 602 can be made to generate a dot product representing a specific score in combination with each 1×300 vector for each sub-picture. The score represents the importance of the sub-picture in the original picture 602. The above process enables information entropy as a filter that obtains a score threshold, eliminates all pixels greater than the score to obtain the noise background 604, performs a pixel subtraction process using the original picture 602 to remove the noise background 604, and generates a picture 608 after the noise background 604 is removed.

[0029] Figure 7 shows a classification algorithm 700 according to an embodiment of the present invention. With respect to the meaning region 702 generated from the segmentation process, a classification operation 704 is performed so that the algorithm index of the recognition process with different classification categories is associated. Each algorithm index is associated with a specific algorithm 708 (for example, an algorithm for detecting and recognizing curved text). The specific algorithm may enable a process of dividing each meaning region and selecting an optimal recognition algorithm for the related region. After removal and noise removal, an independent semantic classification operation 704 is enabled that recommends classification types for different independent semantic units in order to find the optimal recognition technology for different independent semantic units. The above recognition algorithm is pre-trained and placed in a model pool and waits to be acquired for operation. When each semantic unit finds its optimal recognition technology, it becomes possible to extract text 710 regarding its source.

[0030] Figure 8 shows a computer system 90 (for example, the hardware device 139 of FIG. 1) used by or composed of the system of FIG. 1 that improves optical character recognition software technology associated with detecting, evaluating, and dividing an electronic document structure into a plurality of regions in order to extract computer-readable text within the region of the electronic document structure according to an embodiment of the present invention.

[0031] Aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects that are generally referred to herein as "circuits," "modules," or "systems."

[0032] The present invention may be a system, a method, or a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions for causing a processor to implement aspects of the present invention.

[0033] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exclusive list of more specific examples of the computer-readable storage medium includes a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. The computer-readable storage medium is not to be construed as being a transitory signal per se, such as a high-frequency or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0034] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions to be stored in a computer-readable storage medium within each respective computing / processing device.

[0035] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk®, C++, and conventional procedural programming languages such as the “C” programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuit for performing aspects of the present invention.

[0036] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0037] Computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, mobile device, smartwatch, or other programmable data processing device to create a machine, such that the instructions executed via the processor of the computer or other programmable data processing device create means for implementing the functions / acts specified in one or more blocks of the flowchart, block diagram, or both. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing device, or other device to function in a particular manner, such that the computer-readable storage medium containing the instructions comprises a manufacture including instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart, block diagram, or both.

[0038] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing device, or other device to cause a series of operational steps to be performed on the computer, other programmable device, or other device to create a computer-implemented process, such that the instructions which execute on the computer, other programmable device, or other device implement the functions / acts specified in one or more blocks of the flowchart, block diagram, or both.

[0039] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that comprise one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed as one step, simultaneously, substantially simultaneously, in a partially or wholly temporally overlapping manner, depending on the functionality involved, or the blocks may sometimes be executed in the reverse order. It is also noted that each block of the block diagrams, flowchart diagrams, or both, and combinations of blocks in the block diagrams, flowchart diagrams, or both, can be implemented by a special-purpose hardware-based system that performs the specified function or operation, or by a combination of special-purpose hardware and computer instructions.

[0040] The computer system 90 shown in FIG. 8 includes a processor 91, an input device 92 coupled to the processor 91, an output device 93 coupled to the processor 91, and memory devices 94 and 95 respectively coupled to the processor 91. The input device 92 may be, in particular, a keyboard, a mouse, a camera, a touch screen, etc. The output device 93 may be, in particular, a printer, a plotter, a computer screen, a magnetic tape, a removable hard disk, a floppy disk, etc. The memory devices 94 and 95 may be, in particular, a hard disk, a floppy disk, a magnetic tape, an optical storage such as a compact (CD) or digital video disk (DVD), a dynamic random access memory (DRAM), a read-only memory (ROM), etc. The memory device 95 includes computer code 97. The computer code 97 includes an algorithm (e.g., the algorithm of FIG. 2) associated with detecting, evaluating, and dividing an electronic document structure into a plurality of regions in order to extract computer-readable text within the region of the electronic document structure, which improves optical character recognition software technology. The processor 91 executes the computer code 97. The memory device 94 includes input data 96. The input data 96 includes the input required by the computer code 97. The output device 93 displays the output from the computer code 97. Either or both of the memory devices 94 and 95 (or one or more additional memory devices such as a read-only memory (ROM) device or firmware 85) may include an algorithm (e.g., the algorithm of FIG. 2), have computer-readable program code embedded therein, or have other data stored therein, or both, and may be used as a computer-usable medium (or computer-readable medium or program storage device), and the computer-readable program code includes the computer code 97. Generally, the computer program product (or alternatively, the product) of the computer system 90 may include a computer-usable medium (or program storage device).

[0041] In some embodiments, rather than being stored on and accessed from a hard drive, optical disk, or other writable, rewritable, or removable hardware memory device 95, the stored computer program code 84 (e.g., including an algorithm) may be stored on a static, non-removable, read-only storage medium such as a ROM device or firmware 85, or may be accessed directly by a processor 91 from such a static, non-removable, read-only medium. Similarly, in some embodiments, the stored computer program code 97 may be stored as a ROM device or firmware 85, or may be accessed directly by a processor 91 from such a ROM device or firmware 85 rather than from a more dynamic or removable hardware data storage device 95 such as a hard drive or optical disk.

[0042] Furthermore, any of the components of the present invention can be created, integrated, hosted, maintained, deployed, managed, serviced, etc. by a service provider that provides improvements to optical character recognition software technology associated with detecting, evaluating, and dividing an electronic document structure into multiple regions in order to extract computer-readable text within the region of the electronic document structure. Accordingly, the present invention discloses a process for the deployment, creation, integration, hosting, or maintenance of a computing infrastructure, or a combination thereof, including integrating computer-readable code into a computer system 90, and can execute a method that enables a process for improving optical character recognition software technology associated with detecting, evaluating, and dividing an electronic document structure into multiple regions in order to extract computer-readable text within the region of the electronic document structure by combining the code with the computer system 90. In another embodiment, the present invention provides a business method that executes the process steps of the present invention for a subscription fee, advertising, or fee-based system, or a combination thereof. That is, a service provider, such as a solution integrator, can provide the ability to enable a process for improving optical character recognition software technology associated with detecting, evaluating, and dividing an electronic document structure into multiple regions in order to extract computer-readable text within the region of the electronic document structure. In this case, the service provider can create, maintain, support, etc. a computing infrastructure that executes the process steps of the present invention for one or more customers. In return, the service provider can receive payment from the customer under an agreement for a subscription fee and / or a fee, and / or the service provider can receive payment from the sale of advertising content to one or more third parties.

[0043] FIG. 8 shows a computer system 90 as a hardware and software configuration, and any configuration of hardware and software, which would be known to those skilled in the art, may be utilized in conjunction with the computer system 90 of FIG. 8 for the purposes described above. For example, memory devices 94 and 95 may be part of a single memory device rather than separate memory devices. Cloud computing environment

[0044] This disclosure includes a detailed description of cloud computing, but it should be understood that the implementation of the teachings referred to herein is not limited to a cloud computing environment. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0045] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0046] The characteristics are as follows:

[0047] On-demand self-service: Cloud consumers can unilaterally provision computing capabilities, such as server time and network storage, automatically as needed without the need for human interaction with the service provider.

[0048] Broad network access: The capabilities are network-accessible and are accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (such as mobile phones, laptops, and PDAs).

[0049] Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, and different physical and virtual resources are dynamically assigned and reallocated according to demand. Consumers generally have no control over or knowledge of the exact location of the resources provided, but there is location independence in the sense that they may be able to specify the location at a higher level of abstraction (such as a country, state, or data center).

[0050] Rapid elasticity: Capabilities can be rapidly and elastically provisioned, in some cases automatically, scaling out quickly and then rapidly releasing and scaling in. To the consumer, the capabilities available for provisioning often appear limitless and can be purchased in any quantity at any time.

[0051] Measured service: The cloud system automatically controls and optimizes resource use by leveraging metering capabilities at some level of abstraction appropriate to the type of service (such as storage, processing, bandwidth, and active user accounts). It can monitor, control, and report resource use to provide transparency to both the provider and the consumer of the services being utilized.

[0052] The service model is as follows:

[0053] Software as a Service (SaaS): The ability provided to the consumer is to use the provider's applications running on cloud infrastructure. The applications are accessible from various client devices through a client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, which includes the network, servers, operating systems, storage, or even the individual application capabilities, except in some cases for limited user-specific application configurations.

[0054] Platform as a Service (PaaS): The ability provided to the consumer is to deploy consumer-created or acquired applications created using programming languages and tools supported by the provider on cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, which includes the network, servers, operating systems, or storage, but has control over the deployed applications and, in some cases, the environmental configuration in which the applications are hosted.

[0055] Infrastructure as a Service (IaaS): The ability provided to the consumer is to provision other basic computing resources, such as processing, storage, network, and any software that the consumer can include operating systems and applications, to deploy and run. The consumer does not manage or control the underlying cloud infrastructure, but has limited control over the operating system, storage, control over the deployed applications, and, in some cases, the selection of networking components (e.g., host firewall).

[0056] The deployment models are as follows:

[0057] Private Cloud: The cloud infrastructure is operated solely for an organization. It may be managed by the organization or a third party and may exist on - premise or off - premise.

[0058] Community Cloud: The cloud infrastructure is shared by several organizations and supports a specific community that shares concerns (e.g., mission, security requirements, policies, and compliance considerations). It may be managed by the organization or a third party and may exist on - premise or off - premise.

[0059] Public Cloud: The cloud infrastructure is made available to the general public or a large industry group and is owned by an organization that sells cloud services.

[0060] Hybrid Cloud: The cloud infrastructure is a combination of two or more clouds (private, community, or public) that retain their own entities but are linked together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load - balancing between clouds).

[0061] Cloud computing environments are service - oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0062] Next, referring to FIG. 9, an exemplary cloud computing environment 50 is shown. As illustrated, cloud computing environment 50 includes one or more cloud computing nodes 10, which may be used to communicate with local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or cellular phone 54A, desktop computer 54B, laptop computer 54C, or in-vehicle computer system 54N, or combinations thereof. Nodes 10 may communicate with each other. They may be physically or virtually grouped (not shown) in one or more networks, such as private, community, public, or hybrid clouds as described above, or combinations thereof. This enables cloud computing environment 50 to provide infrastructure, platform, or software, or combinations thereof, as a service where cloud consumers do not need to maintain resources on local computing devices. The types of computing devices 54A, 54B, 54C, and 54N shown in FIG. 9 are intended to be merely exemplary, and it is understood that cloud computing nodes 10 and cloud computing environment 50 can communicate with any type of computerized device through any type of network and / or network addressable connection (e.g., using a web browser).

[0063] Next, referring to FIG. 10, a set of function abstraction layers provided by cloud computing environment 50 (see FIG. 9) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 10 are intended to be merely exemplary and that embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0064] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and networking components 66. In some embodiments, the software components include network application server software 67 and database software 68.

[0065] The virtualization layer 70 provides an abstraction layer from which examples of virtual entities may be provided, such as virtual server 71, virtual storage 72, virtual network 73 including a virtual private network, virtual applications and operating systems 74, and virtual client 75.

[0066] In one example, the management layer 80 may provide the functions described below. Resource provisioning 81 provides for the dynamic procurement of computing resources and other resources utilized to execute tasks within a cloud computing environment. Measurement and pricing 82 provides for cost tracking when resources are utilized within a cloud computing environment and for billing or invoicing for the consumption of those resources. In one example, these resources may include application software licenses. Security provides for verification of identification information for cloud consumers and tasks and for protection of data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 87 provides for the allocation and management of cloud computing resources such that the required service levels are met. Service level agreement (SLA) planning and fulfillment 88 provides for advance commitments and procurement of cloud computing resources for which future needs are predicted in accordance with the SLA.

[0067] The workload layer 101 provides examples of functionality that may be utilized in a cloud computing environment. Examples of workloads and functions that may be provided from this layer include mapping and navigation 102, software development and lifecycle management 103, provision of education in virtual classrooms 133, data analysis processing 134, transaction processing 106, and improvements in optical character recognition software technology associated with detecting, evaluating, and dividing an electronic document structure into multiple regions for extracting computer-readable text within the region of the electronic document structure 107.

[0068] Although embodiments of the present invention have been described herein for purposes of illustration, many modifications and changes will be apparent to those skilled in the art. Accordingly, the appended claims are intended to cover all such modifications and changes as falling within the true spirit and scope of the invention.

Claims

1. Receiving, by a processor of a hardware device, a document to be processed; Detecting, by the processor, different types of text data in the document; Dividing, by the processor, the document into a plurality of text regions associated with the different types of the text data, wherein each text region of the plurality of text regions has a single type of the text data; Removing, by the processor, optical noise from each of the text regions; Selecting, by the processor, different optical character recognition software codes to be applied to each of the text regions; Executing, by the processor, in response to the selecting, the different optical character recognition software codes for each of the text regions to obtain extractable computer-readable text within each of the text regions; Comprising; The removing of the optical noise from each of the text regions is Encoding, by the processor, each of the text regions as a sum meaning vector of each of the text regions; Enabling, by the processor, a window of a predetermined size to divide each of the text regions into the plurality of sub-regions such that each sub-region of the plurality of sub-regions is encoded as a vector of a predetermined number of elements; Generating, by the processor, a dot product for each of the sub-regions based on the sum meaning vector, wherein each dot product has a specific score representing the importance level of each of the sub-regions; An optical character recognition segmentation and processing method having.

2. The removing of the optical noise from each of the text regions is Comparing, by the processor, each of the specific scores with a score threshold; Executing, by the processor, the removing by erasing all pixels of each sub-region exceeding the score threshold based on the result of the comparing; The optical character recognition segmentation and processing method according to claim 1, further comprising.

3. Receiving, by a processor of a hardware device, a document to be processed; Detecting, by the processor, different types of text data in the document; Dividing, by the processor, the document into a plurality of text regions associated with the different types of the text data, wherein each text region of the plurality of text regions has a single type of the text data; Removing, by the processor, optical noise from each of the text regions; Selecting, by the processor, different optical character recognition software codes to be applied to each of the text regions; Executing, by the processor, in response to the selecting, the different optical character recognition software codes for each of the text regions to obtain extractable computer-readable text within each of the text regions; comprising; The method for optical character recognition segmentation and processing, wherein the removing is performed via a self-attention mechanism executed via natural language processing.

4. The method further comprising classifying, by the processor, each of the text regions with respect to the single type of the text data, and the selecting is performed based on a result of the classifying. The method for optical character recognition segmentation and processing according to any one of claims 1 to 3.

5. The method for optical character recognition segmentation and processing according to any one of claims 1 to 4, wherein the optical noise comprises unwanted background text and images of the document.

6. The method for optical character recognition segmentation and processing according to any one of claims 1 to 5, wherein each of the different optical character recognition software codes has self-learning software codes stored in a dedicated database.

7. The method for optical character recognition segmentation and processing according to any one of claims 1 to 6, wherein the extractable computer-readable text within each of the text regions is configured to be used with respect to cut or copy and paste functions.

8. The optical character recognition segmentation and processing method according to any one of claims 1 to 7, wherein the different types of text data are constituted by a format selected from the group consisting of a table format, a watermark format, a handwritten format, and a rotated text format.

9. Further comprising providing at least one support service for at least one of creating, integrating, hosting, maintaining, and deploying computer-readable code in the hardware device, wherein the computer-readable code is executed by a computer processor to implement the receiving, detecting, splitting, removing, selecting, and executing. The optical character recognition segmentation and processing method according to any one of claims 1 to 8.

10. A computer program comprising computer-readable program code, wherein when the computer-readable program code is executed by a processor of a hardware device, it has an algorithm for implementing an optical character recognition segmentation and processing method, and the optical character recognition segmentation and processing method comprises: Receiving, by the processor, a document to be processed; Detecting, by the processor, different types of text data in the document; Splitting, by the processor, the document into a plurality of text regions associated with the different types of text data, wherein each text region of the plurality of text regions has a single type of the text data; Removing, by the processor, optical noise from each of the text regions; Selecting, by the processor, different optical character recognition software codes to be applied to each of the text regions; Executing, by the processor, in response to the selecting, the different optical character recognition software codes for each of the text regions to obtain extractable computer-readable text within each of the text regions; Including The removing of the optical noise from each of the text regions is encoding each of the text regions by the processor as a sum meaning vector of each of the text regions; enabling the processor to divide each of the text regions into the plurality of sub-regions such that each of the sub-regions of the plurality of sub-regions is encoded as a vector having a predetermined number of elements; generating, by the processor, a dot product for each of the sub-regions based on the sum meaning vector, each of the dot products having a specific score representing the importance level of each of the sub-regions; A computer program having the above.

11. removing the optical noise from each of the text regions, comparing, by the processor, each of the specific scores with a score threshold; executing the removing by the processor by erasing all pixels of each sub-region exceeding the score threshold based on the result of the comparing; The computer program according to claim 10, further comprising the above.

12. A computer program comprising computer-readable program code, wherein when the computer-readable program code is executed by a processor of a hardware device, it has an algorithm for implementing an optical character recognition segmentation and processing method, and the optical character recognition segmentation and processing method includes: receiving, by the processor, a document to be processed; detecting, by the processor, different types of text data of the document; dividing, by the processor, the document into a plurality of text regions associated with the different types of text data, each of the text regions of the plurality of text regions having a single type of the text data; removing optical noise from each of the text regions by the processor; selecting, by the processor, different optical character recognition software codes to be applied to each of the text regions; In response to the selecting, the processor executes the different optical character recognition software codes for the respective text regions to obtain extractable computer-readable text within each of the text regions. comprising The removing is performed via a self-attention mechanism executed via natural language processing, a computer program. **Claim 13** The optical character recognition segmentation and processing method further comprises classifying, by the processor, the respective text regions for the single type of text data, and the selecting is performed based on a result of the classifying. The computer program according to any one of claims 10 to 12. **Claim 14** The computer program according to any one of claims 10 to 13, wherein the optical noise comprises unwanted background text and images of the document. **Claim 15** The computer program according to any one of claims 10 to 14, wherein each of the different optical character recognition software codes has self-learning software codes stored in a dedicated database. **Claim 16** The computer program according to any one of claims 10 to 15, wherein the extractable computer-readable text within each of the text regions is configured to be used with respect to cut or copy and paste functions. **Claim 17** The computer program according to any one of claims 10 to 16, wherein the different types of text data are constituted by a format selected from the group consisting of a table format, a watermark format, a handwritten format, and a rotated text format. **Claim 18** A hardware device comprising a processor coupled to a computer-readable memory unit, wherein the computer-readable memory unit has instructions that, when executed by the processor, implement an optical character recognition segmentation and processing method, and the optical character recognition segmentation and processing method receiving, by the processor, a document to be processed; detecting, by the processor, different types of text data of the document; The processor divides the document into a plurality of text regions associated with the different types of the text data, wherein each text region of the plurality of text regions has a single type of the text data; The processor removes optical noise from each of the text regions; The processor selects different optical character recognition software codes to be applied to each of the text regions; The processor, in response to the selecting, executes the different optical character recognition software codes on each of the text regions to obtain extractable computer-readable text within each of the text regions; comprising: The removing of the optical noise from each of the text regions is The processor encodes each of the text regions as a sum meaning vector of each of the text regions; The processor enables a window of a predetermined size to divide each of the text regions into the plurality of sub-regions such that each sub-region of the plurality of sub-regions is encoded as a vector of a predetermined number of elements; The processor generates a dot product for each of the sub-regions based on the sum meaning vector, wherein each dot product has a specific score representing the importance level of each of the sub-regions; A hardware device having the above.

19. The removing of the optical noise from each of the text regions is The processor compares each of the specific scores with a score threshold; The processor, based on the result of the comparing, erases all pixels of each sub-region exceeding the score threshold to perform the removing; The hardware device according to claim 18, further comprising the above.

20. A hardware device comprising a processor coupled to a computer-readable memory unit, wherein the computer-readable memory unit has instructions that, when executed by the processor, implement an optical character recognition segmentation and processing method, and the optical character recognition segmentation and processing method is receiving, by the processor, a document to be processed; detecting, by the processor, different types of text data in the document; dividing, by the processor, the document into a plurality of text regions associated with the different types of the text data, wherein each text region of the plurality of text regions has a single type of the text data; removing, by the processor, optical noise from each of the text regions; selecting, by the processor, different optical character recognition software codes to be applied to each of the text regions; executing, by the processor, in response to the selecting, the different optical character recognition software codes on each of the text regions to obtain extractable computer-readable text within each of the text regions; comprising; the hardware device, wherein the removing is performed via a self-attention mechanism executed via natural language processing. Claims 21 The optical character recognition segmentation and processing method, further comprising classifying, by the processor, each of the text regions with respect to the single type of the text data, and the selecting is performed based on a result of the classifying; The hardware device according to any one of claims 18 to 20.

Citation Information

Patent Citations

  • REFLECTION FILM AND TRANSLUCENT REFLECTION FILM FOR OPTICAL RECORDING MEDIUM AND Ag ALLOY SPUTTERING TARGET FOR FORMING THESE FILMS

    JP2006092646A