Text detection algorithm for separating words detected as a text bounding box

CN116524528BActive Publication Date: 2026-08-14INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-30
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

这可导致丢失的单词(例如,出现在文本行中的单词被文本检测算法忽略)

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524528B_ABST
    Figure CN116524528B_ABST
Patent Text Reader

Abstract

This invention relates to a text detection algorithm for separating words detected as a text bounding box. A method, computer system, and computer program product for text detection are provided. The invention may include training a text detection model. The invention may include performing text detection on an input image using the trained text detection model. The invention may include determining whether at least one bounding box among a plurality of bounding boxes generated using the input image has an aspect ratio higher than a threshold. The invention may include, based on determining that at least one bounding box among a plurality of bounding boxes generated using the input image has an aspect ratio higher than a threshold, enlarging any text within that at least one bounding box, and performing text detection on the new image using the trained text detection model. The invention may include outputting an output image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates generally to the field of computing, and more specifically, to text detection algorithms. Background Technology

[0002] Lines of text drawn in a small font size may be detected by text detection algorithms using only a single text bounding box, where the text detection model cannot separate one line of text from another. This can lead to missed words (e.g., words appearing in a line of text are ignored by the text detection algorithm). Summary of the Invention

[0003] Embodiments of the present invention disclose a method, computer system, and computer program product for text detection. The present invention may include training a text detection model. The present invention may include performing text detection on an input image using the trained text detection model. The present invention may include determining whether at least one bounding box among a plurality of bounding boxes generated using the input image has an aspect ratio higher than a threshold. The present invention may include, based on the determination that at least one bounding box among the plurality of bounding boxes generated using the input image has an aspect ratio higher than the threshold, enlarging any text within the at least one bounding box, and performing text detection on the new image using the trained text detection model. The present invention may include outputting an output image. Attached Figure Description

[0004] These and other objects, features, and advantages of the invention will become apparent from the following detailed description of exemplary embodiments of the invention, which will be read in conjunction with the accompanying drawings. Various features in the drawings are not to scale, as they are illustrated for clarity to enable those skilled in the art to understand the invention in conjunction with specific embodiments. In the drawings:

[0005] Figure 1 A networked computer environment according to at least one embodiment is shown;

[0006] Figure 2 This is an operational flowchart illustrating a process for text detection according to at least one embodiment;

[0007] Figure 3A and 3B These are exemplary illustrations of the original image and the new image, respectively, according to at least one embodiment;

[0008] Figure 4 According to at least one embodiment Figure 1 A block diagram depicting the internal and external components of a computer and server;

[0009] Figure 5 Includes embodiments according to this disclosure Figure 1A block diagram of an exemplary cloud computing environment for a computer system depicted in the diagram; and

[0010] Figure 6 According to embodiments of this disclosure Figure 5 A block diagram of the functional layers of an exemplary cloud computing environment. Detailed Implementation

[0011] Detailed embodiments of the claimed structures and methods are disclosed herein; however, it should be understood that the disclosed embodiments are merely illustrative of the claimed structures and methods, which may be embodied in different forms. The invention can be embodied in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided to make this disclosure thorough and complete, and to fully convey the scope of the invention to those skilled in the art. Details of well-known features and techniques may be omitted in the specification to avoid unnecessarily obscuring the presented embodiments.

[0012] This invention can be a system, method, and / or computer program product at any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the invention.

[0013] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or recessed structures with instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0014] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or downloaded via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network) to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.

[0015] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​or similar programming languages ​​such as the "C" programming language. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform aspects of this invention, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions to personalize the electronic circuits by utilizing state information from the computer-readable program instructions.

[0016] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0017] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.

[0018] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0019] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a non-consecutive order. For example, depending on the functions involved, two consecutively shown blocks may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.

[0020] The exemplary embodiments described below provide a system, method, and program product for text detection. Therefore, the embodiments can improve the field of optical character recognition by performing secondary detection on text that may not have been correctly detected by first magnifying misrepresented or missing text and then running the text recognition software again. More specifically, the invention may include training a text detection model. The invention may include performing text detection on an input image using the trained text detection model. The invention may include determining whether at least one bounding box among a plurality of bounding boxes generated using the input image has an aspect ratio higher than a threshold. The invention may include, based on determining that at least one bounding box among a plurality of bounding boxes generated using the input image has an aspect ratio higher than a threshold, magnifying any text within that at least one bounding box and performing text detection on the new image using the trained text detection model. The invention may include outputting an output image.

[0021] As previously mentioned, text lines drawn in small font sizes may be detected by text detection algorithms using only a single text bounding box, where the text detection model cannot separate one line of text from another. This can lead to missing words (e.g., words appearing in a text line are ignored by the text detection algorithm).

[0022] Therefore, among other operations, it may be advantageous to magnify the text that has been misrepresented or is missing and run the text recognition software a second time on that part of the input image.

[0023] According to at least one embodiment, the present invention can improve the ability to detect small-sized text in an image by creating bounding boxes around each individual word rather than the entire line of text, thus enabling words to be detected without loss.

[0024] According to at least one embodiment, the present invention can overcome the inaccuracies of text detection algorithms, such as those that occur when using small font sizes. This algorithm can... (Python and all Python-based trademarks are trademarks or registered trademarks of the Python Software Foundation) or any other high-level general-purpose programming language, because the algorithm's functionality does not depend on The use of language.

[0025] According to at least one embodiment, optical character recognition (OCR) can be used to recognize text within digital images (e.g., scanned documents and / or images). The OCR function developed herein works by marking pixels in the digital image, which indicate the presence of text within those pixels. Bounding boxes can then be placed around the pixels marked with text to convert these portions of the digital image into text. If the system determines, based on analysis of the bounding box size, that there may be more than one word within the generated bounding boxes, the text can be enlarged (e.g., made larger), and the analysis is repeated until it is determined that each generated bounding box contains only one word. Once this is complete, the image can be converted into text, and the text can be returned.

[0026] Reference Figure 1 This describes an exemplary networked computer environment 100 according to one embodiment. The networked computer environment 100 may include a computer 102 having a processor 104 and a data storage device 106, the computer 102 being capable of running software program 108 and text detection program 110a. The networked computer environment 100 may also include a server 112 and a communication network 116, the server 112 being capable of running a text detection program 110b that can interact with a database 114. The networked computer environment 100 may include multiple computers 102 and servers 112, only one of which is shown. The communication network 116 may include different types of communication networks, such as wide area networks (WANs), local area networks (LANs), telecommunications networks, wireless networks, public switched networks, and / or satellite networks. It should be understood that... Figure 1 This illustration is merely one example of an implementation and does not imply any limitation regarding the environment in which different embodiments may be implemented. Many modifications can be made to the depicted environment based on design and implementation requirements.

[0027] Client computer 102 can communicate with server computer 112 via communication network 116. Communication network 116 may include connections such as wired or wireless communication links, or fiber optic cables. (See reference...) Figure 4As discussed, server computer 112 may include internal component 902a and external component 904a, and client computer 102 may include internal component 902b and external component 904b. Server computer 112 may also operate in a cloud computing service model (such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS)). Server 112 may also reside in a cloud computing deployment model, such as a private cloud, community cloud, public cloud, or hybrid cloud. Client computer 102 may be, for example, a mobile device, telephone, personal digital assistant, netbook, laptop computer, tablet computer, desktop computer, or any type of computing device capable of running programs, accessing a network, and accessing database 114. According to various embodiments of this example, text detection programs 110a, 110b may interact with database 114, which may be embedded in different storage devices, such as, but not limited to, computer / mobile device 102, networked server 112, or cloud storage service.

[0028] According to this embodiment, users of client computer 102 or server computer 112 can use text detection programs 110a and 110b respectively to solve the problem that the text detection algorithm cannot accurately detect all words in the input image by: enlarging the incorrectly represented or missing text, and performing a secondary detection on the parts of the text that may not have been correctly detected by using the text detection algorithm again. See below for further details. Figure 2 , Figure 3A and Figure 3B The text detection method will be explained in more detail.

[0029] Now for reference Figure 2 The diagram depicts an operational flowchart of an exemplary text detection process 200 used by text detection programs 110a and 110b according to at least one embodiment.

[0030] In section 202, a text detection algorithm is trained. This algorithm can utilize a neural network (i.e., a model, a deep learning model) to predict words or lines of text by displaying an image of text (where each pixel can be labeled as text or not) to the neural network, and the model can be trained accordingly to predict whether each pixel of an image contains text for future images. Deep learning text detection models may include, but are not limited to, WordUNet (which performs text detection), etc. Some comprehensive systems may use a version of WordUNet to perform at least part of the text detection, but may also include additional steps such as receiving the image via OCR.

[0031] Training a text detection algorithm can be accomplished by generating text using a generator. The generator is designed to produce random text that is variable in size, font, and / or background, and this random text can be injected with additional noise. Training can also involve creating bounding boxes around each word and training a neural network, for example, outputting each pixel in the image that is within the bounding box as having a label of 1, and each pixel in the image that is not within the bounding box as having a label of 0.

[0032] Bounding boxes can be generated using the four corners; however, on-the-fly labeling is accomplished by essentially replacing the pixels in the image where text detection procedures 110a, 110b determine the presence of a word with a specific color (e.g., white). To determine the presence of letters and words, each character can be surrounded by a Gaussian distribution (e.g., a normal distribution with a bell-shaped curve). For this model, a Gaussian distribution can depict the center of a letter as more important than spaces between letters (e.g., the center of a letter produces a more noticeable distribution, while spaces between letters, words, or lines produce no distribution). This can guide the model to learn the actual location of characters within words and does not assign equal weight to every point within the bounding box, thus contributing to the overall performance of the model.

[0033] As previously described, text detection procedures 110a and 110b can apply a Gaussian distribution around the center of each character (e.g., letters in a word). The neural network can provide the resulting Gaussian distribution (e.g., a Gaussian distribution including each letter and the spaces between those letters) as output. Using the resulting Gaussian distribution, text detection procedures 110a and 110b can be trained to determine whether each Gaussian distribution indicates the presence of a letter.

[0034] Once the neural network is trained, whenever an image is input, it can generate an image based on the calculated probability of whether each pixel in the image contains text. This image includes the label 'white' in all locations where the model believes text is present and the label 'black' in all locations where the model believes text is absent. Essentially, the model predicts the presence of text for each pixel in the image. A second algorithm, called connected components, can then be used to combine white regions into bounding boxes. The connected component algorithm can be used to place bounding boxes around the detected text in the image. The bounding boxes can start from a single white pixel and can expand until they reach a black pixel (e.g., the connected component algorithm can use the boundaries of the white portions of the image to place the bounding boxes).

[0035] At step 204, text detection is performed. The text used by text detection programs 110a and 110b can be uploaded by the user. The text can be a part of a document or the entire document. It doesn't matter whether a complete document or a part of a document is uploaded, because text detection programs 110a and 110b can detect text in any uploaded image.

[0036] As previously described regarding step 202 above, once the model is trained, text detection can be performed using WordUNet, meaning that an image can be taken as input and the probability of each pixel in the image containing text can be output. A threshold can be applied to generate the resulting binary image with black and white pixels, representing 'yes' or 'no,' indicating the presence or absence of text in those pixels, respectively. A connected component algorithm can then be used to place bounding boxes around each word of the detected text in the image.

[0037] At 206, text bounding boxes are identified that have an aspect ratio or a box height (i.e., bounding box height) divided by a box width (i.e., bounding box width) higher than a predetermined threshold (e.g., the Otsu threshold). For example, if a very wide bounding box exists, text detection procedures 110a, 110b may attempt to determine whether there is more than one word within that text bounding box. Similarly, if a line is detected instead of a word, the aspect ratio may be greater than that of a single word. The image can always have a constant resolution (e.g., 600 / 800) regardless of how the image can be resized on the screen, and bounding boxes can also be created at the same resolution. Image resolution can help determine whether a bounding box is higher than the predetermined threshold.

[0038] This determination can be based on a ratio of box height to box width (e.g., a threshold). For example, a text line may be very wide (e.g., box width), but the box height may always be the same (e.g., the height of the text). For example, the threshold could be a ratio of box height to box width calculated for the longest word in the dictionary and multiplied by two. If detected text exists with a ratio higher than this threshold, the detected text is unlikely to be just a single word. According to at least one alternative embodiment, the threshold can be modified as needed.

[0039] At position 208, the identified text bounding box is enlarged and copied to the new image. If, at position 206, the detected text is determined to have an aspect ratio or box height divided by box width above a predetermined threshold, that portion of the image can be enlarged (e.g., made larger). For portions of text determined to be above the threshold, a larger text size improves the algorithm's ability to detect individual words.

[0040] Enlargement can be achieved by using standard image resizing algorithms (e.g., bilinear interpolation) to enlarge the image to a range where text detection works better. As mentioned earlier, this may be relevant in cases where the text detection algorithm might not detect small-sized text. The enlarged portion of the original image can then be copied to a new image (e.g., created in the same file format as the original input image) for another run by the text detection algorithm.

[0041] At 210, text detection is performed on the new image. Once the text is enlarged, text detection is performed again using WordUNet. The text detection performed here can be the same as described above with respect to step 204, except that the detected text portion is only the enlarged text portion, as previously described with respect to step 208.

[0042] In step 212, text detection procedures 110a and 110b determine which text bounding boxes to use. If, after enlarging a portion of the image, the text detection algorithm generates more bounding boxes than those generated for the original image, then text detection procedures 110a and 110b can determine that the bounding boxes generated from the enlarged image are correct bounding boxes and can include these bounding boxes in the final output image. The basic principle here is that after enlarging small bounding boxes, there may be bounding boxes that may no longer be recognized as text and can be filtered out at this stage, thereby improving the speed and / or efficiency of the computational system by avoiding attempts to transform incorrectly represented portions of the image.

[0043] Finally, text detection programs 110a and 110b can combine the bounding boxes generated from the original image and the bounding boxes generated from the magnified portion of the image in the final output image.

[0044] Now for reference Figure 3A and Figure 3B Exemplary illustrations of an original image and a new image are depicted according to at least one embodiment. The text lines depicted in the original image 3A and the new image 3B have been run using different text detection algorithms, with the text detection algorithm utilized in the original image 3A most closely mirroring prior art prior to this invention. As can be seen in the original image 3A, each individual word is not enclosed by a bounding box, but rather the entire text line is detected within a large bounding box. Conversely, the new image 3B depicts the illustrative output of this invention, where most words are detected individually and enclosed by individual bounding boxes. This increases the efficiency of the text detection algorithm and improves the quality of the text detection results.

[0045] Understandable. Figure 2 , 3A3B provides only an illustration of one embodiment and does not imply any limitation on how different embodiments may be implemented. Many modifications may be made to the depicted embodiments based on design and implementation requirements.

[0046] Figure 4 This is an illustrative embodiment of the present invention. Figure 1 Block diagram 900 depicts the internal and external components of a computer. It should be understood that... Figure 4 The illustrations provided are merely illustrative of one embodiment and do not imply any limitation regarding the environment in which different embodiments may be implemented. Many modifications may be made to the depicted environment based on design and implementation requirements.

[0047] Data processing systems 902 and 904 represent any electronic device capable of executing machine-readable program instructions. Data processing systems 902 and 904 may represent smartphones, computer systems, PDAs, or other electronic devices. Examples of computing systems, environments, and / or configurations that data processing systems 902 and 904 may represent include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, network PCs, minicomputer systems, and distributed cloud computing environments that include any of the above systems or devices.

[0048] The user client computer 102 and the network server 112 may include Figure 4 The internal components 902a, 902b and external components 904a, 904b shown are illustrated. Each group of internal components 902a, 902b includes one or more processors 906, one or more computer-readable RAMs 908 and one or more computer-readable ROMs 910 on one or more buses 912, as well as one or more operating systems 914 and one or more computer-readable tangible storage devices 916. One or more operating systems 914, software programs 108 and text detection programs 110a in client computer 102 and text detection programs 110b in network server 112 may be stored on one or more computer-readable tangible storage devices 916 for execution by one or more processors 906 via one or more RAMs 908 (which typically include cache memory). Figure 4 In the illustrated embodiment, each computer-readable tangible storage device 916 is a disk storage device of an internal hard disk drive. Alternatively, each computer-readable tangible storage device 916 is a semiconductor storage device, such as ROM 910, EPROM, flash memory, or any other computer-readable tangible storage device capable of storing computer programs and digital information.

[0049] Each set of internal components 902a, 902b also includes an R / W drive or interface 918 for reading from and writing to one or more portable computer-readable tangible storage devices 920, such as CD-ROM, DVD, Memory Stick, magnetic tape, disk, optical disc, or semiconductor storage devices. Software programs (such as software program 108 and text detection programs 110a and 110b) may be stored on one or more corresponding portable computer-readable tangible storage devices 920, read from and loaded into corresponding hard disk drives 916 via the corresponding R / W drive or interface 918.

[0050] Each set of internal components 902a, 902b may also include a network adapter (or switch port card) or interface 922, such as a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G or 4G wireless interface card, or other wired or wireless communication links. Software program 108 and text detection program 110a in client computer 102 and text detection program 110b in network server computer 112 may be downloaded from an external computer (e.g., a server) via a network (e.g., the Internet, a local area network, or another wide area network) and the corresponding network adapter or interface 922. From the network adapter (or switch port adapter) or interface 922, software program 108 and text detection program 110a in client computer 102 and text detection program 110b in network server computer 112 are loaded into the corresponding hard disk drive 916. The network may include copper wire, fiber optic, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.

[0051] Each set of external components 904a, 904b may include a computer display monitor 924, a keyboard 926, and a computer mouse 928. External components 904a, 904b may also include a touchscreen, a virtual keyboard, a touchpad, a pointing device, and other human-computer interface devices. Each set of internal components 902a, 902b also includes a device driver 930 connected to the computer display monitor 924, keyboard 926, and computer mouse 928. The device driver 930, R / W driver or interface 918, and network adapter or interface 922 include hardware and software (stored in storage device 916 and / or ROM 910).

[0052] It should be understood beforehand that although this disclosure includes a detailed description of cloud computing, the implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed hereafter.

[0053] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.

[0054] The features are as follows:

[0055] On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power, such as server time and network storage, as needed, without requiring manual interaction with the service provider.

[0056] Wide Area Network (WAN) Access: Capabilities are available on the network and accessed through standard mechanisms that facilitate the use of heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0057] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. Location independence has significance because consumers typically do not control or know the exact location of the resources provided, but can specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0058] Rapid Flexibility: In some cases, the ability to scale outwards and inwards quickly and flexibly can be provided. For consumers, the available capacity often appears unlimited and can be purchased in any quantity at any time.

[0059] Measurement services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency for both service providers and consumers.

[0060] The service model is as follows:

[0061] Software as a Service (SaaS): The capability offered to consumers is the ability to use the provider's applications running on cloud infrastructure. Applications can be accessed from various client devices through thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application capabilities, with possible exceptions such as limited user-specific application configuration settings.

[0062] Platform as a Service (PaaS): This provides consumers with the ability to deploy consumer-created or acquired applications onto cloud infrastructure using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and the configuration of any application hosting environments.

[0063] Infrastructure as a Service (IaaS): This provides consumers with the capability to deliver processing, storage, networking, and other basic computing resources that enable them to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do have control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0064] The deployment model is as follows:

[0065] Private cloud: Cloud infrastructure operated solely by an organization. It can be managed by the organization or a third party and can exist inside or outside a building.

[0066] Community cloud: Cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.

[0067] Public cloud: Cloud infrastructure available to the general public or large industrial groups and owned by organizations that sell cloud services.

[0068] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and applications to be ported together (e.g., cloud bursting for load balancing between clouds).

[0069] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure of a network of interconnected nodes.

[0070] Now for reference Figure 5The diagram illustrates an illustrative cloud computing environment 1000. As shown, the cloud computing environment 1000 includes one or more cloud computing nodes 100, whose local computing devices used by cloud consumers (such as, for example, personal digital assistants (PDAs) or cellular phones 1000A, desktop computers 1000B, laptop computers 1000C, and / or automotive computer systems 1000N) can communicate with the cloud computing nodes 100. The nodes 100 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment 1000 to provide Infrastructure as a Service, Platform as a Service, and / or Software as a Service without requiring cloud consumers to maintain resources on their local computing devices for these services. It is to be understood that... Figure 5 The types of computing devices 1000A-N shown are for illustrative purposes only, and computing node 100 and cloud computing environment 1000 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).

[0071] Now for reference Figure 6 This illustrates a set of functional abstraction layers 1100 provided by the cloud computing environment 1000. It should be understood beforehand that... Figure 6 The components, layers, and functions shown are intended to be illustrative only, and embodiments of the invention are not limited thereto. As described, the following layers and corresponding functions are provided:

[0072] The hardware and software layer 1102 includes hardware and software components. Examples of hardware components include: a mainframe 1104; a server 1106 based on a RISC (Reduced Instruction Set Computer) architecture; a server 1108; a blade server 1110; a storage device 1112; and a network and networking component 1114. In some embodiments, the software components include network application server software 1116 and database software 1118.

[0073] The virtualization layer 1120 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 1122; virtual storage 1124; virtual network 1126, including virtual private network; virtual application and operating system 1128; and virtual client 1130.

[0074] In one example, management layer 1132 may provide the functionality described below. Resource Provisioning 1134 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and Pricing 1136 provides cost tracking when resources are used in the cloud computing environment and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User Portal 1138 provides access to the cloud computing environment for consumers and system administrators. Service Level Management 1140 provides cloud resource allocation and management to ensure that required service levels are met. Service Level Agreement (SLA) Planning and Fulfillment 1142 provides pre-scheduling and procurement of cloud resources for anticipated future needs according to the SLA.

[0075] Workload layer 1144 provides examples of functionalities that can leverage a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 1146; software development and lifecycle management 1148; virtual classroom instruction provision 1150; data analysis and processing 1152; transaction processing 1154; and text detection 1156. Text detection programs 110a and 110b provide a way to address the problem that text detection algorithms cannot accurately detect all words in an input image by magnifying the misrepresented or missing text and performing a secondary detection on the potentially incorrectly detected portions of the text again using the text detection algorithm.

[0076] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or technical improvements to technologies found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for text detection, the method comprising: Train the text detection model; Use a trained text detection model to perform text detection on the input image; Determine whether at least one bounding box among a plurality of bounding boxes generated using the input image has an aspect ratio higher than a threshold, the aspect ratio higher than the threshold indicating that the at least one bounding box contains text that is incorrectly represented or missing. Based on the determination that at least one bounding box among the plurality of bounding boxes generated using the input image has an aspect ratio higher than the threshold, any text within the at least one bounding box is magnified; Copy the enlarged image within the at least one bounding box, which has an aspect ratio higher than the threshold, to a new image file with the same file format as the input image; The new image is subjected to text detection using a trained text detection model; At least one bounding box generated using the input image and at least one bounding box generated using the new image are combined to create an output image, wherein, when it is determined that at least one bounding box among the plurality of bounding boxes generated using the input image contains more than one word, any corresponding portion of the input image is replaced with text detected by the at least one bounding box generated iteratively using the new image, which is created around each individual word rather than the entire line of text; as well as The output image is then output.

2. The method according to claim 1, wherein, The trained text detection model is a neural network that predicts words or lines of text by labeling each pixel in the input image as text or not as text and placing one of the multiple bounding boxes generated using the input image around a set of concurrent pixels labeled as text.

3. The method according to claim 2, wherein, A Gaussian distribution is calculated based on the set of concurrent pixels marked as text within the bounding box.

4. The method according to claim 1, wherein, The threshold is calculated by dividing the height of the bounding box for the longest word in the dictionary by twice the width of the bounding box.

5. The method according to claim 1, wherein, Enlarging the text within the at least one bounding box further includes: The enlarged text is copied to the new image, which has a constant resolution.

6. The method according to claim 1, wherein, Performing text detection on the input image also includes: Based on the presence or absence of a word, each pixel of the input image is replaced with a specific color.

7. The method according to claim 6, further comprising: The connected component algorithm is used to combine each pixel with a specific color into multiple bounding boxes.

8. A computer system for text detection, comprising: One or more processors, one or more computer-readable memories, one or more computer-readable storage media, and program instructions stored on at least one of the one or more storage media for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing the method according to any one of claims 1-7.

9. A computer program product for text detection, comprising program instructions executable by a processor to cause the processor to perform the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Word bounding box detection

    US10127673B1

  • Compositional model for text recognition

    US20200226400A1

  • Computer Vision Systems and Methods for Information Extraction from Text Images Using Evidence Grounding Techniques

    US20210192201A1