Deep learning-based optical character recognition method and system

The deep learning-based OCR method addresses scene text recognition challenges by using reference points for text regions, ensuring accurate extraction and efficient training, despite detection errors, and improving model performance.

JP7732704B2Active Publication Date: 2025-09-02NAVER CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024034245
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-03-08
Filing Date
2024-03-06
Publication Date
2025-09-02
Estimated Expiration
2044-03-06

Smart Images

  • Figure 0007732704000005
    Figure 0007732704000005
  • Figure 0007732704000006
    Figure 0007732704000006
  • Figure 0007732704000007
    Figure 0007732704000007
Patent Text Reader

Abstract

To provide a deep learning-based optical character recognition method implemented with at least one processor.SOLUTION: A deep learning-based optical character recognition method includes the steps of detecting at least one text region from an image, generating at least one reference point associated with the at least one text region, and extracting at least one text from the image based on the image and the at least one reference point.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to optical character recognition methods and systems, and more particularly to deep learning-based optical character recognition methods and systems that extract text from an image based on the image and reference points by generating reference points to reference text regions contained in the image. [Background technology]

[0002] Optical character recognition (OCR) is a technology that detects text from an image and identifies the type of text it detects. Conventional OCR technology has been used to recognize characters in documents. However, recognizing text found in everyday life, such as roadside signs, is technically challenging due to the wide variety of possible cases. For example, text contained in a diagonal or curved image can be difficult to recognize because text regions of various shapes, rather than rectangular, may exist. Text displayed in the surrounding environment or on objects is called scene text.

[0003] To recognize such scene text, conventional technologies consist of a first stage in which a detector is used to detect text regions from an image, and the detected regions are edited by adjusting their size and transmitted to a recognizer, followed by a second stage in which the recognizer recognizes the text. However, in such conventional technologies, if the text regions cannot be detected, both the text recognition and the text may fail, and the content of the text may be lost when adjusting the size of the detected regions. Furthermore, the artificial neural network executed in each stage must be trained repeatedly, which may reduce the efficiency of computational resources. Furthermore, when updating the detector, the recognizer must also be retrained, which may be disadvantageous in terms of management and maintenance of related services. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Korean Patent Publication No. 10-2015-0125376 Summary of the Invention [Problem to be solved by the invention]

[0005] The present disclosure provides a deep learning-based optical character recognition method and system (apparatus) for solving the above-mentioned problems. [Means for solving the problem]

[0006] The present disclosure can be implemented in numerous ways, including as a method, an apparatus (system), or a computer readable computer program product.

[0007] According to one embodiment of the present disclosure, a deep learning-based optical character recognition method may include detecting at least one text region from an image, generating at least one reference point associated with the at least one text region, and extracting at least one text from the image based on the image and the at least one reference point.

[0008] A computer readable computer program may be provided for executing a method according to an embodiment of the present disclosure on a computer.

[0009] An optical character recognition system according to one embodiment of the present disclosure may include a detector that detects at least one text region from an image and generates at least one reference point associated with the at least one text region, and a recognizer that extracts at least one text from the image based on the image and the at least one reference point. [Effects of the Invention]

[0010] According to various embodiments of the present disclosure, even if an error occurs in the process of detecting a text region in an image, text can be correctly extracted from the image by using the entire image and reference points. That is, even if an error occurs in detecting a text region, loss of text content can be prevented. Furthermore, since the final text recognition result does not solely depend on the detected text region, it can also be advantageous in extracting rotated text in an image.

[0011] According to various embodiments of the present disclosure, the detector and recognizer are trained in one step instead of two steps, which may improve training efficiency. Furthermore, by using images and reference points without adjusting the size of detected text regions, it is possible to prevent recognition errors caused by distorted or truncated long characters. Furthermore, by sharing multi-scale features extracted from the backbone with the detector and recognizer, it is possible to improve model performance and inference speed.

[0012] According to various embodiments of the present disclosure, text in an image can be recognized and decoded by utilizing reference points and overall image features without pooling or masking regions of interest in the image, which allows for smooth text extraction even if regions of interest are not detected correctly.

[0013] According to various embodiments of the present disclosure, complex scene text, such as crossed text, text within text, and various fonts and sizes, can be extracted more accurately. Furthermore, since the final text recognition result does not depend heavily on the detection of text regions within the image, it is possible to be resilient to errors in detecting text regions. Furthermore, even if the text within the image is rotated, the text can be extracted correctly.

[0014] According to various embodiments of the present disclosure, by matching which text regions contain specific word units of text, the detected text may be linked more organically, thereby providing a translation service of higher quality after extracting text from a document using optical character recognition.

[0015] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary knowledge in the technical field to which the present disclosure belongs (referred to as an "ordinary engineer") from the description in the claims.

[0016] Non-limiting embodiments of the present disclosure will be described with reference to the accompanying drawings, as described below, in which like reference numerals indicate like elements and in which: [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 1 illustrates an example of an optical character recognition method according to one embodiment of the present disclosure. [Figure 2] FIG. 2 is a schematic diagram illustrating an example configuration in which an information processing system is communicatively coupled to multiple user terminals for optical character recognition according to one embodiment of the present disclosure. [Figure 3] FIG. 3 is a block diagram illustrating the internal configuration of a user terminal and an information processing system according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a schematic diagram illustrating an optical character recognition system according to one embodiment of the present disclosure. [Figure 5] FIG. 5 is a diagram illustrating an example of extracting text from an image according to one embodiment of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating an example process of an optical character recognition system according to an embodiment of the present disclosure. [Figure 7] FIG. 7 is a diagram illustrating an example of an optical character recognition result according to an embodiment of the present disclosure. [Figure 8] FIG. 8 is a diagram illustrating an example in which text is detected character by character according to an embodiment of the present disclosure. [Figure 9] FIG. 9 is a diagram illustrating an example of detecting line-unit and paragraph-unit text regions in a document according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a flowchart illustrating a method according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0018] Hereinafter, specific implementations of the present disclosure will be described in detail with reference to the accompanying drawings. However, in the following description, detailed descriptions of well-known functions and configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.

[0019] In the accompanying drawings, identical or corresponding components are denoted by the same reference numerals. In addition, in the following description of the embodiments, duplicated descriptions of identical or corresponding components may be omitted. However, the omission of a description of a component does not mean that such a component is not included in a certain embodiment.

[0020] The advantages and features of the disclosed embodiments, and methods for achieving them, will become apparent from the following detailed description of the embodiments taken in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below, and may be embodied in various different forms. The embodiments are provided solely so that this disclosure will be thorough and complete, and will adequately convey the scope of the invention to those skilled in the art.

[0021] The terms used in this specification will be briefly explained, and the disclosed embodiments will be specifically described. The terms used in this specification are generally used and widely, taking into consideration the function of the present disclosure. However, these terms may change depending on the intentions or precedents of engineers in the relevant field, the emergence of new technologies, etc. In addition, in certain cases, the applicant may arbitrarily select terms, and in such cases, their meanings will be described in detail in the description of the relevant invention. Therefore, the terms used in this disclosure should be defined based on the meanings of the terms and the overall content of this disclosure, rather than simply by their names.

[0022] In this specification, the singular expression includes the plural expression unless the context clearly dictates otherwise. Furthermore, the plural expression includes the singular expression unless the context clearly dictates otherwise. When a part of the entire specification is said to include a certain element, it does not mean that other elements are excluded, but that other elements may also be included, unless otherwise specified.

[0023] Furthermore, the terms "module" or "unit" as used herein refer to a software or hardware component, and the "module" or "unit" performs some function. However, the term "module" or "unit" is not limited to software or hardware. A "module" or "unit" may be configured to reside on an addressable storage medium or to execute one or more processors. Thus, by way of example, a "module" or "unit" may include components such as software components, object-oriented software components, class components, and task components, as well as at least one of a process, a function, an attribute, a procedure, a subroutine, a program code segment, a driver, firmware, microcode, a circuit, data, a database, a data structure, a table, an array, or a variable. The components and "modules" or "units" may indicate that the functionality provided therein may be combined into fewer components and "modules" or "units" or further separated into additional components and "modules" or "units."

[0024] According to one embodiment of the present disclosure, a "module" or "unit" may be implemented with a processor and memory. "Processor" should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, "processor" may refer to an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. "Processor" may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such combination. Additionally, "memory" should be broadly interpreted to include any electronic component capable of storing electronic information. "Memory" may refer to various types of processor-readable media such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage devices, registers, etc. Memory is said to be in electronic communication with a processor if the processor can read information from and / or write information to the memory. Memory that is integrated into a processor is in electronic communication with the processor.

[0025] In this disclosure, a "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, a system may be configured with one or more server devices. As another example, a system may be configured with one or more cloud devices. As yet another example, a system may be configured and operated by a server device and a cloud device together.

[0026] In this disclosure, "display" may refer to any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided by a computing device.

[0027] In the present disclosure, "each of a plurality of A's" or "each of a plurality of A's" may refer to each of all components included in the plurality of A's, or may refer to each of some components included in the plurality of A's.

[0028] End-to-end scene text spotting techniques that recognize arbitrary text instances have made significant improvements. Common methods for scene text detection are region-of-interest pooling or segmentation masking to limit the features of a single text instance. However, if the detection is inaccurate (e.g., one or more characters are cut off), the recognizer in such methods may have difficulty decoding or generating the correct character image sequence.

[0029] This disclosure provides an optical character recognition method including a detection-agnostic end-to-end recognizer, taking into account the difficulty of accurately determining word boundaries using a detector alone, for example, in scene text spotting problems. The disclosed method may reduce the tight dependency between the detector and the recognizer by linking the detector and the recognizer using a single reference point for each piece of text instead of using detected text regions. Furthermore, the disclosed method may allow the recognizer to recognize text displayed as a reference point along with features of the entire image. In other words, since only a single point is required to recognize text, text can be extracted from an image without any type of detector or bounding polygon. Additionally, the disclosed method may provide an optical character recognition method and system that is competitive in regular and irregular text spotting benchmarks and robust against text detection errors.

[0030] End-to-end scene text spotting techniques have been applied in various fields, including information extraction, image retrieval, and visual query answering. In general, an end-to-end scene text spotting pipeline consists of a detector and a recognizer. The detector localizes text instances in an image in the form of boxes or polygons, and the recognizer receives each localized text region as input and decodes the characters in each patch of the image.

[0031] Traditional scene text spotting pipelines use a somewhat tightly coupled framework between the detector and recognizer. In particular, if the detector encounters a recognition error for a text region, the recognizer may be fed an image containing only the text itself, with the target text clipped. In this case, the recognition performance of the pipeline can be heavily influenced by the performance of the image patches output by the detector and recognizer. Recently, end-to-end scene text spotting methods have adopted a more loosely coupled framework by extracting features using region-of-interest (ROI) pooling or masking and restricting the recognizer's input region to a single word. Using localized features in the recognizer can reduce the recognizer's reliance on clipped regions in the detector, but detector errors can still accumulate and lead to recognition failures. Furthermore, feature pooling and masking require data with bounding boxes or polygons to train end-to-end scene text spotting models, even if precise boundary information is not required for the final application.

[0032] This disclosure provides a novel end-to-end recognizer that significantly reduces dependency on the accuracy of detection results. Instead of relying on a detector to extract precise text regions, the detector can generate reference points for each text instance. The recognizer can then comprehensively recognize the text surrounding the reference points. Specifically, given the reference points, the recognizer can be trained to pay attention to specific text instance regions while decoding text sequences. Furthermore, by requiring the detector to only require a single reference point, a greater variety of detection algorithms and annotations can be applied. Furthermore, rotated or curved text instances can be handled naturally without pooling operations and polygon-type annotations.

[0033] In this disclosure, a "text region" may refer to a region in an image that contains text. Here, a text region may be represented in the form of a rectangular bounding box, but is not limited thereto, and may also be represented in the form of a polygonal bounding polygon. Furthermore, a text region may be configured with vertex coordinates. For example, if a text region is in the form of a bounding box, the text region may be configured with top-left, top-right, bottom-right, and bottom-left coordinates.

[0034] In this disclosure, a "recognizer" may refer to a module in an optical character recognition system that recognizes text regions in an image. The recognizer may also include a text decoder. That is, the recognizer may recognize the text regions to generate a character image sequence and convert the character image sequence into text using the text decoder.

[0035] FIG. 1 illustrates an optical character recognition method according to an embodiment of the present disclosure. A first example 110 is an example of a conventional optical character recognition method. The conventional optical character recognition method includes a first step in which a detector detects a text region 112 from an image, adjusts the size of the detected text region 112, and transmits the size to a recognizer. A second step in which the recognizer recognizes text in the detected text region 112. However, in this conventional method, if the detector erroneously detects the text region 112, both steps of text recognition may fail. Furthermore, when the size of the detected text region 112 is adjusted, the text content may be lost. Furthermore, since the artificial neural network executed in each step must be trained repeatedly, the efficiency of computing resources may decrease. Furthermore, when the detector is updated, the recognizer must also be retrained, which may increase the time and cost required for management or maintenance of related services.

[0036] A second example 120 is an example of an optical character recognition method of the present disclosure. In one embodiment, at least one text region may be detected from an image. Specifically, features related to text in units of words may be extracted from the image. Position information for the at least one text region may be generated based on the features. The position information for the text region may be the vertex coordinates of a bounding box or a bounding polygon of the text region.

[0037] In one embodiment, a reference point 122 associated with a text region may be generated. The reference point 122 may include the center point of the text region. In this case, the center point of the text region may be determined based on the position information of the text region. As shown in the figure, if the bounding box of the text region includes the character "sesame," the reference point 122 may be generated at the center of the character "sesame." An example of how the reference point is displayed will be described in detail below with reference to FIG. 7.

[0038] In one embodiment, text may be recognized or extracted from an image based on the image and the reference point 122. Specifically, multiple offset points 124_1-124_4 adjacent to the reference point 122 may be generated. The multiple offset points 124_1-124_4 may guide a recognizer to focus on areas adjacent to the multiple offset points 124_1-124_4 and extract text. At least one piece of text may be extracted from the image based on the image and the multiple offset points 124_1-124_4. For example, the recognizer may extract "sesame" from the image by focusing on the multiple offset points 124_1-124_4 in the image. While FIG. 1 illustrates four offset points, this is not limiting.

[0039] In one embodiment, multiple text extractions may be performed at least partially in parallel. For example, a region containing "sesame" may be detected as a first text region, and a region containing "stick" may be detected as a second text region. In this case, a first reference point may be generated in the first text region, and a second reference point may be generated in the second text region. Furthermore, "sesame" may be extracted from the first text region based on the image and the first reference point. Furthermore, "stick" may be extracted from the second text region based on the image and the second reference point, at least partially in parallel with the extraction of "sesame."

[0040] With this configuration, even if there is an error in the process of detecting the text region of the image, the text can be correctly extracted from the image by using the entire image and reference points. In other words, even if an error occurs in detecting the text region, the loss of the text content can be prevented. Furthermore, since the final text recognition result does not solely depend on the detected text region, it is also advantageous for extracting rotated text in the image.

[0041] 2 is a schematic diagram illustrating a configuration in which an information processing system 230 is communicatively coupled to a plurality of user terminals 210_1, 210_2, and 210_3 for optical character recognition according to one embodiment of the present disclosure. As shown in the figure, the plurality of user terminals 210_1, 210_2, and 210_3 may be coupled to the information processing system 230, which may provide optical character recognition services, via a network 220. Note that the plurality of user terminals 210_1, 210_2, and 210_3 may include terminals of users who will receive the optical character recognition service.

[0042] In one embodiment, information processing system 230 may include one or more server devices and / or databases, or one or more cloud computing service-based distributed computing devices and / or distributed databases, that may store, provide, and execute computer-executable programs (e.g., downloadable applications) and data related to providing optical character recognition services, etc.

[0043] The optical character recognition service provided by the information processing system 230 may be provided to users through a web browser or a web browser extension program of an optical character recognition service application installed in each of the user terminals 210_1, 210_2, and 210_3. For example, the information processing system 230 may provide information or respond to a request for extracting text from an image received from the user terminals 210_1, 210_2, and 210_3 through the optical character recognition service application.

[0044] A plurality of user terminals 210_1, 210_2, and 210_3 may communicate with the information processing system 230 via a network 220. The network 220 may be configured to enable communication between the plurality of user terminals 210_1, 210_2, and 210_3 and the information processing system 230. Depending on the installation environment, the network 220 may be configured as a wired network such as Ethernet (registered trademark), a wired home network (Power Line Communication), a telephone line communication device, and RS serial communication, a mobile communication network, a wireless network such as WLAN (Wireless LAN), Wi-Fi (registered trademark), Bluetooth (registered trademark), and ZigBee (registered trademark), or a combination thereof. The communication method is not limited and may include not only a communication method utilizing a communication network that the network 220 may include (for example, a mobile communication network, a wired Internet, a wireless Internet, a broadcast communication network, a satellite communication network, etc.), but also short-range wireless communication between the user terminals 210_1, 210_2, and 210_3.

[0045] 2, a mobile phone terminal 210_1, a tablet terminal 210_2, and a PC terminal 210_3 are shown as examples of user terminals, but are not limited thereto, and the user terminals 210_1, 210_2, and 210_3 may be any computing devices capable of wired and / or wireless communication and on which an optical character recognition service application or a web browser, etc., can be installed and executed. For example, the user terminals may include AI speakers, smartphones, mobile phones, navigation systems, computers, notebooks, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, game consoles, wearable devices, IoT (Internet of Things) devices, VR (Virtual Reality) devices, AR (Augmented Reality) devices, set-top boxes, etc. Also, although FIG. 2 shows three user terminals 210_1, 210_2, and 210_3 communicating with the information processing system 230 via the network 220, this is not limited thereto, and a different number of user terminals may be configured to communicate with the information processing system 230 via the network 220.

[0046] 2 illustrates user terminals 210_1, 210_2, and 210_3 receiving optical character recognition services from information processing system 230, but is not limited to this. For example, optical character recognition services may be provided by optical character recognition programs / applications installed in user terminals 210_1, 210_2, and 210_3 without communication with information processing system 230. Also, information processing system 230 is illustrated as a single device, but is not limited to this, and information processing system 230 may be comprised of multiple devices.

[0047] FIG. 3 is a block diagram illustrating the internal configuration of a user terminal 210 and an information processing system 230 according to an embodiment of the present disclosure. The user terminal 210 refers to any computing device capable of executing applications, a web browser, and the like, and capable of wired / wireless communication, and may include, for example, the mobile phone terminal 210_1, the tablet terminal 210_2, and the PC terminal 210_3 of FIG. 2 . As shown in the figure, the user terminal 210 may include a memory 312, a processor 314, a communication module 316, and an input / output interface 318. Similarly, the information processing system 230 may include a memory 332, a processor 334, a communication module 336, and an input / output interface 338. As shown in FIG. 3 , the user terminal 210 and the information processing system 230 may be configured to communicate information and / or data over the network 220 using their respective communication modules 316 and 336. Additionally, the input / output device 320 may be configured to input information and / or data to or output information and / or data generated from the user terminal 210 via the input / output interface 318 .

[0048] The memory 312, 332 may include any non-transitory computer-readable recording medium. According to one embodiment, the memory 312, 332 may include a permanent mass storage device such as a read only memory (ROM), a disk drive, a solid state drive (SSD), a flash memory, etc. As another example, a non-transitory mass storage device such as a ROM, an SSD, a flash memory, a disk drive, etc. may be included in the user terminal 210 or the information processing system 230 as a separate permanent storage device distinct from the memory. The memory 312, 332 may also store an operating system and at least one program code.

[0049] Such software components may be loaded from a computer-readable recording medium separate from the memories 312, 332. Such separate computer-readable recording medium may include a recording medium directly connectable to the user terminal 210 and the information processing system 230, but may also include a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, or a memory card. As another example, the software components may be loaded into the memories 312, 332 via the communication modules 316, 336 rather than from a computer-readable recording medium. For example, at least one program may be loaded into the memories 312, 332 based on a computer program installed by a file provided over the network 220 by a developer or a file distribution system that distributes application installation files.

[0050] The processors 314, 334 may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 314, 334 by the memory 312, 332 or the communication modules 316, 336. For example, the processors 314, 334 may be configured to execute instructions received from program code stored in a storage device, such as the memory 312, 332.

[0051] The communication modules 316 and 336 may provide a configuration or function for the user terminal 210 and the information processing system 230 to communicate with each other via the network 220, and may provide a configuration or function for the user terminal 210 and / or the information processing system 230 to communicate with other user terminals or other systems (e.g., a separate cloud system). For example, a request or data (e.g., a request to extract text from an image) generated by the processor 314 of the user terminal 210 using program code stored in a storage device such as the memory 312 may be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, a control signal or command provided under the control of the processor 334 of the information processing system 230 may be received by the user terminal 210 via the communication module 316 of the user terminal 210 via the communication module 336 and the network 220.

[0052] The input / output interface 318 may be a means for interfacing with the input / output device 320. As an example, the input device may include a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, a mouse, etc., and the output device may include a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface 318 may be a means for interfacing with a device in which the configuration or functions for performing input and output are integrated into one, such as a touch screen. For example, when the processor 314 of the user terminal 210 processes instructions of a computer program loaded into the memory 312, a service screen configured using information and / or data provided by the information processing system 230 or another user terminal may be displayed on the display via the input / output interface 318. Although FIG. 3 illustrates the input / output device 320 as not being included in the user terminal 210, this is not limiting and the input / output device 320 and the user terminal 210 may be configured as a single device. Furthermore, the input / output interface 338 of the information processing system 230 may be a means for interfacing with an input or output device (not shown) that may be coupled to or included in the information processing system 230. In Fig. 3, the input / output interfaces 318, 338 are shown as elements configured separately from the processors 314, 334, but are not limited thereto, and the input / output interfaces 318, 338 may also be configured to be included in the processors 314, 334.

[0053] The user terminal 210 and the information processing system 230 may include more components than those shown in FIG. 3 . However, it is not necessary to explicitly show most of the conventional components. In one embodiment, the user terminal 210 may be implemented to include at least some of the input / output devices 320 described above. The user terminal 210 may also include other components such as a transceiver, a global positioning system (GPS) module, a camera, various sensors, a database, etc. For example, if the user terminal 210 is a smartphone, it may include components typically included in smartphones. For example, the user terminal 210 may be implemented to further include various components such as an acceleration sensor, a gyro sensor, a microphone module, a camera module, various physical buttons, buttons using a touch panel, an input / output port, and a vibrator for vibration.

[0054] While a program for an optical character recognition service application or the like is running, the processor 314 can receive text, images, videos, sounds, and / or actions, etc., entered or selected via input devices such as a touch screen, keyboard, camera including an audio sensor and / or image sensor, microphone, etc. connected to the input / output interface 318, and can store the received text, images, videos, sounds, and / or actions, etc. in the memory 312 or provide them to the information processing system 230 via the communication module 316 and the network 220.

[0055] The processor 314 of the user terminal 210 may be configured to manage, process, and / or store information and / or data received from the input / output device 320, other user terminals, the information processing system 230, and / or multiple external systems. The information and / or data processed by the processor 314 may be provided to the information processing system 230 via the communication module 316 and the network 220. The processor 314 of the user terminal 210 may transfer and output information and / or data to the input / output device 320 via the input / output interface 318. For example, the processor 314 may display the received information and / or data on the screen of the user terminal 210.

[0056] The processor 334 of the information processing system 230 may be configured to manage, process, and / or store information and / or data received from the plurality of user terminals 210 and / or the plurality of external systems. The information and / or data processed by the processor 334 may be provided to the user terminal 210 via the communication module 336 and the network 220.

[0057] 4 is a schematic diagram illustrating an example of an optical character recognition system according to one embodiment of the present disclosure. In one embodiment, the optical character recognition system may include a backbone (not shown), a transformer encoder 420, a detector (or location head) 430, and a recognizer 440. Note that the recognizer 440 may include a text decoder.

[0058] In one embodiment, the transformer encoder 420 may combine the multi-scale feature maps generated by the backbone, and the detector 430 may set reference points 432 for text instances and bounding boxes. The recognizer 440 may generate a character image sequence by recognizing text regions based on the multi-scale features and reference points 432 of the input image 410, and convert the generated character image sequence into text using a text decoder.

[0059] In one embodiment, an input image 410 is provided to a backbone, whereby feature maps (e.g., C2, C3, C4, and C5) can be extracted from the input image 410. The resolution of the extracted feature maps can be 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the resolution of the input image 410, respectively. The feature maps can then be projected onto multiple channels (e.g., 256 channels) by applying a fully-connected (FC) layer and group normalization. The projected feature maps can then be merged, flattened, and concatenated into feature tokens of size (L2+L3+L4+L5)×256. Here, L i is C i (or H / 2 i ×W / 2 i ) as input. In this case, the Transformer Encoder 420 may take it as input and output refined features. The refined features, along with the reference points 432 via the detector 430, may be used by the recognizer 440 to autoregressively generate text sequences within the text instances.

[0060] In one embodiment, the Transformer Encoder 420 may use deformable attention, which scales linearly with input length. Because the cost of traditional self-attention training or inference increases quadratically with input length, using a Transformer for multi-scale feature concatenation may be inefficient. In contrast, deformable attention may provide the encoder and decoder with more efficient and accurate localization results. The deformable attention may be calculated as follows:

[0061]

number

[0062] In one embodiment, to recognize a text object using position information, a reference point (or center position of a text instance) 432 can be predicted using a detector 430 that detects text from multi-scale features. A bounding polygon of the text instance can be extracted using a segmentation map. For example, an L2-sized feature token can be extracted from the feature corresponding to feature map C2 and reconstructed to a size of (H / 4, W / 4). A segmentation head consisting of transposed convolution, group normalization, and ReLU (Rectified Linear Unit) can be used to obtain binary and threshold maps. Additionally, the center coordinate of the text instance can be determined as the reference point 432 in the detected bounding polygon. While the text detection method described here is based on a segmentation-based detection method, various detection methods can be applied.

[0063] In one embodiment, a recognizer 440 including a text decoder, e.g., a transformer decoder, may autoregressively predict text sequences from text instances associated with text regions, using deformable attention 444 to reference the image 410 and reference points 432. In this case, a query Q for the text decoder may be generated based on the text embedding, positional embedding, and reference points q ref Also, the key K and value V of the text decoder may be feature tokens of the transformer encoder 420. The query Q may be propagated through self-attention 442, deformable attention 444, and feed-forward layer 446.

[0064] In one embodiment, during the training phase of the recognizer 440, the bounding boxes of the text regions are sampled from the image, and the center coordinates of the bounding boxes can be used as the reference points 432 for the text decoder. This allows the detector 430 and the recognizer 440 to be trained independently. Also, to reduce the coordinate difference between the model's predictions and the actual values, the center coordinates can be randomly varied during the training phase using Equation 2 below.

[0065]

number

[0066] In one embodiment, the loss function L used for training can be expressed as Equation 3 below.

[0067]

number

[0068] In one embodiment, the inference stage may use the probability map of the detector 430. In this case, the probability map is binarized to a specified threshold, and connected components may be extracted as regions in the binary map. Note that the size of the extracted regions may be smaller than the actual text regions. Therefore, the extracted regions may be expanded using an offset D in the Vatti clipping algorithm, as shown in Equation 4 below.

[0069]

number

[0070] This configuration allows the detector and recognizer to be trained in one step instead of two, improving training efficiency. Furthermore, by using images and reference points without adjusting the size of the detected text region, it is possible to prevent recognition errors caused by distorted or truncated long characters. Furthermore, by sharing multi-scale features extracted from the backbone with the detector and recognizer, model performance and inference speed can be improved.

[0071] FIG. 5 illustrates an example of extracting text from an image 510 according to an embodiment of the present disclosure. In one embodiment, text can be recognized from the image 510 using a reference point 512 and offset points 514_1 to 514_4. Specifically, a recognizer 520 including deformable attention can sequentially and autoregressively decode each character by referring to the features of the image 510 and the reference point 512. Note that the deformable attention may not perform attention on the entire features of the image 510, but may perform attention by sampling only a few points. That is, attention may be performed using the reference point 512 and offset points 514_1 to 514_4.

[0072] In one embodiment, offset points 514_1 to 514_4 may be generated at positions adjacent to reference point 512. Specifically, the position of each of the offset points 514_1 to 514_4 may be calculated by adding a model-predicted offset from the coordinates of reference point 512. In this case, the recognizer 520 may determine the position where attention is performed based on the keys and values ​​of offset points 514_1 to 514_4. As a result, the recognizer 520 may decode text by generating characters autoregressively, with attention being focused on the reference point 512 and the surrounding offset points 514_1 to 514_4. For example, the recognizer 520 may sequentially and autoregressively decode each character of "sesame" by referring to offset points 514_1 to 514_4. Although FIG. 5 illustrates four offset points 514_1 to 514_4, this is not limiting.

[0073] This configuration allows text in an image to be recognized and decoded by utilizing reference points and global image features without pooling or masking regions of interest in the image, which allows for smooth text extraction even if regions of interest are not detected correctly.

[0074] 6 illustrates an example process of an optical character recognition system according to one embodiment of the present disclosure. In one embodiment, a backbone 620 may generate image features (or multi-scale feature maps) from an image 610. A transformer encoder 630 may also tune and encode the image features.

[0075] In one embodiment, the word detector 640 may detect the location of text by word from features encoded in the image 610. Specifically, the word detector 640 may extract features related to the text by word. The word detector 640 may also generate location information for at least one text region based on the features. The word detector 640 may also generate a reference point based on the location information for the text region. The reference point may be the center point of the detected text region.

[0076] In one embodiment, based on the reference points generated by the word detector 640 and the encoded features of the image 610, recognizers (recognizers 1 to N) 650_1 to 650_N may recognize and extract text from the image 610. Specifically, the recognizers 650_1 to 650_N may generate a plurality of offset points adjacent to the reference points. The recognizers 650_1 to 650_N may recognize the text from the image 610 by paying attention to the reference points and the plurality of offset points, and generate a character image sequence. The recognizers 650_1 to 650_N may then convert the generated character image sequence into text by decoding it.

[0077] In one embodiment, word detector 640 may detect multiple text regions from feature encoded image 610 at least partially in parallel, where reference points may be generated for each text region, allowing recognizers 650_1 through 650_N to extract text in parallel for each of the N text regions.

[0078] In one embodiment, a line detector 660 may detect line-by-line text locations in the feature encoded image 610. Specifically, the line detector 660 may extract features related to line-by-line text. The line detector 660 may also generate location information for at least one text region based on the features. An example of detecting line-by-line text locations will be described in detail below with reference to FIG. 9.

[0079] In one embodiment, a paragraph detector 670 may detect paragraph-based text positions in the feature-encoded image 610. Specifically, the paragraph detector 670 may extract features related to the text in paragraph units. The paragraph detector 670 may also generate position information for at least one text region based on the features. An example of detecting paragraph-based text positions will be described in detail below with reference to FIG. 9.

[0080] In one embodiment, the word detector 640, the line detector 660, and the paragraph detector 670 may be coupled in parallel to the transformer encoder 630. That is, the word detector 640, the line detector 660, and the paragraph detector 670 may detect word, line, and paragraph-based text regions at each level at least partially in parallel.

[0081] In one embodiment, the detected text regions at each level may be correlated with each other through a post-processing step 680. Specifically, the positions of words detected by the word detector 640, the positions of lines detected by the line detector 660, and the positions of paragraphs detected by the paragraph detector 670 may be matched with each other. As a result, position information of line-based text regions and / or paragraph-based text regions containing text extracted by the recognition units 650_1 to 650_N may be detected.

[0082] FIG. 7 illustrates an example of an optical character recognition result according to an embodiment of the present disclosure. A first example 710 illustrates a conventional optical character recognition result. Conventional optical character recognition methods extract text from an image by pooling or masking regions of interest. In this case, extraction of complex, dense text or text with geometric shapes may fail. For example, in the ambiguous word "HOME" containing a geometric shape, a conventional optical character recognition method may determine only "ME" as the region of interest due to the geometric shape of the "O." Such a failure to detect text may also result in failure to recognize text in the image.

[0083] A second example 720 is an example of an optical character recognition result of the present disclosure. The optical character recognition method of the present disclosure can detect text regions and generate reference points (denoted by the "+" shape in FIG. 7) without pooling or masking regions of interest. In this case, the recognizer does not need to rely heavily on the detected text regions. In other words, even if the text region detection result is inaccurate, the recognizer can correctly extract the text in the image by considering the entire extracted image and concentrating on the reference points.

[0084] This configuration allows for more accurate extraction of complex scene text, such as crossed text, text within text, and texts of various fonts and sizes within an image. Furthermore, since the final text recognition result does not depend heavily on the detection of text regions within an image, it is robust against misdetection of text regions. Furthermore, even if the text within an image is rotated, the text can be extracted correctly.

[0085] FIG. 8 illustrates an example of character-based text detection according to an embodiment of the present disclosure. In one embodiment, after a recognizer 810 recognizes word-based text, character-based text may be detected. Specifically, an auto-regressive decoder 820 may be added to the recognizer 810 to prevent the optical character recognition model from becoming heavy. The auto-regressive decoder 820 may auto-regressively recognize each character included in the extracted text using a classification score for the extracted text.

[0086] In one embodiment, the autoregressive decoder 820 can simultaneously recognize and detect each character by predicting the location and angle of each recognized character. The autoregressive decoder 820 can also be trained to correctly recognize and detect characters for foreign languages ​​and all character orientations. As shown in Figure 8, the autoregressive decoder 820 can detect each of "c, h, o, c, o...".

[0087] In one embodiment, pseudo-labeling may be applied to train the autoregressive decoder 820. Specifically, a teacher model may train labeled data. In this case, the trained teacher model may generate pseudo-labeled data for weakly labeled data. Here, the weakly labeled data may include data in which word-based text regions are detected in an image. The pseudo-labeled data may also include data in which the teacher model detects characters in the weakly labeled data. The autoregressive decoder 820 may then train both the pseudo-labeled data and the labeled data as a student model.

[0088] 9 is a diagram illustrating an example of detecting line-based and paragraph-based text regions in a document according to an embodiment of the present disclosure. A first example 910 illustrates an example of detecting word-based text in a document using the optical character recognition method of the present disclosure. A second example 920 illustrates an example of detecting line-based text in a document using the optical character recognition method of the present disclosure. A third example 930 illustrates an example of detecting paragraph-based text in a document using the optical character recognition method of the present disclosure.

[0089] In one embodiment, features extracted from the Transformer Encoder can be used to extract text-related features from a document. Based on the features, a word detector, a line detector, and a paragraph detector can perform optical character recognition (OCR) on a word, line, and paragraph basis, respectively. In this case, the word detector, the line detector, and the paragraph detector can detect text at each level at least partially in parallel. In a post-processing stage, matching can be performed to determine which line-based text regions and which paragraph-based text regions contain a particular word-based text.

[0090] This configuration allows for more organic linking of detected text by matching which text region a specific word unit is included in. This allows for better quality translation services to be provided after extracting text from a document using optical character recognition.

[0091] 10 is a flowchart illustrating an example of a method 1000 according to an embodiment of the present disclosure. In one embodiment, the method 1000 may be performed by at least one processor. The method 1000 may begin with the processor detecting at least one text region from an image (S1010). Specifically, the processor may extract features related to text on a word-by-word basis from the image. The processor may also generate position information for the at least one text region based on the features.

[0092] The processor may then generate at least one reference point associated with the at least one text region (S1020), where each of the at least one reference point may include a center point of each of the at least one text region.

[0093] Then, the processor may extract at least one text from the image based on the image and the at least one reference point (S1030). Specifically, the processor may generate a plurality of offset points adjacent to the at least one reference point. The plurality of offset points may guide the recognizer to focus on an area adjacent to the plurality of offset points and extract the text. The processor may also extract at least one text from the image based on the image and the plurality of offset points.

[0094] In one embodiment, the at least one text region may include a first text region and a second text region. The at least one reference point may include a first reference point associated with the first text region and a second reference point associated with the second text region. In this case, the processor may extract the first text from the first text region based on the image and the first reference point. Furthermore, the processor may extract the second text from the second text region based on the image and the second reference point, in parallel with at least a portion of the step of extracting the first text.

[0095] In one embodiment, the extracted text may be word-by-word text, in which case the processor may autoregressively detect each of at least one character contained in the extracted text using the classification score for the extracted text, and may predict the position and angle of each of the at least one character within the image.

[0096] In one embodiment, the processor may extract line-by-line text-related features from the image, and based on the features, the processor may generate position information for at least one text region.

[0097] In one embodiment, the processor may extract features associated with the text in paragraph units from the image, and based on the features, the processor may generate position information for at least one text region.

[0098] In one embodiment, the processor may detect at least one word-based text region from the image. The processor may also detect at least one line-based text region from the image. The processor may then detect at least one paragraph-based text region from the image. In this case, detecting the word-based text region, detecting the line-based text region, and detecting the paragraph-based text region may be performed at least partially in parallel. Furthermore, the processor may detect position information of at least one of the line-based text region or the paragraph-based text region containing the extracted text.

[0099] The above method may be provided as a computer program stored on a computer-readable recording medium for execution by a computer. The medium may permanently store the computer-executable program or may temporarily store it for execution or download. The medium may also be various recording or storage means in the form of a single or multiple pieces of hardware combined together, and is not limited to media directly connected to any computer system but may also be distributed over a network. Examples of media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and ROM, RAM, flash memory, and other media configured to store program instructions. Other examples of media include recording or storage media managed by app stores that distribute applications, or by sites or servers that provide or distribute various other software.

[0100] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, such techniques may be implemented by hardware, firmware, software, or a combination thereof. Those of ordinary skill in the art will understand that the various illustrative logic blocks, modules, circuits, and algorithm steps described in connection with the present disclosure may also be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design requirements imposed on the overall system. Those of ordinary skill in the art may implement the described functionality in various ways for each particular application, but such implementations should not be interpreted as departing from the scope of the present disclosure.

[0101] In a hardware implementation, the processing units used to perform the techniques may be implemented within one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in this disclosure, computers, or combinations thereof.

[0102] Thus, the various illustrative logic blocks, modules, and circuits described in connection with this disclosure may be implemented or performed by a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but the processor may also be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other configuration.

[0103] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), nonvolatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage device, etc. The instructions may also be executable by one or more processors and cause the processors to perform certain aspects of the functions described in this disclosure.

[0104] If implemented in software, the techniques may be stored on or transmitted via a computer-readable medium as one or more instructions or code. Computer-readable media includes both computer storage media and communication media, including any medium that facilitates transfer of a computer program from one place to another. Storage media may be any available medium that can be accessed by a computer. By way of non-limiting example, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to transport or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Furthermore, any connection is properly termed a computer-readable medium.

[0105] For example, if the software is transferred from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted wire, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted wire, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of medium. As used herein, disk and disc include CDs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0106] A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in a user terminal.

[0107] Although the embodiments described above are described as utilizing aspects of the presently disclosed subject matter in one or more stand-alone computer systems, the present disclosure is not limited thereto and may be implemented in connection with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the subject matter in the present disclosure may be implemented in multiple processing chips or devices, and storage may be affected similarly across multiple devices. Such devices may include PCs, network servers, and even portable devices.

[0108] Although the present disclosure has been described herein with reference to certain embodiments, various modifications and changes may be made thereto without departing from the scope of the present disclosure as would be understood by one of ordinary skill in the art to which the presently disclosed invention pertains, and such modifications and changes should be considered to fall within the scope of the claims appended hereto. [Explanation of symbols]

[0109] 110: First example 112: Text area 120: Second example 122: Reference Point 124_1~124_4: Multiple offset points

Claims

1. 1. A deep learning-based Optical Character Recognition (OCR) method executed by at least one processor, comprising: Detecting at least one text region from the image based on model data obtained by deep learning; inferring and generating at least one reference point associated with the at least one detected text region so as to reduce a difference between the model data and the at least one detected text region; predicting and extracting at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, Detecting at least one text region from the image comprises: extracting text-related features from the image on a word-by-word basis; and generating position information for the at least one text region based on the features.

2. The deep learning based optical character recognition method of claim 1 , wherein each of the at least one reference point comprises a center point of each of the at least one text region.

3. A deep learning-based optical character recognition (OCR) method executed by at least one processor, comprising: Detecting at least one text region from the image based on model data obtained by deep learning; inferring and generating at least one reference point associated with the at least one detected text region so as to reduce a difference between the model data and the at least one detected text region; predicting and extracting at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, the at least one text area includes a first text area and a second text area; the at least one reference point includes a first reference point associated with the first text region and a second reference point associated with the second text region; The step of extracting at least one text from the image comprises: extracting a first text from the first text region based on the image and the first reference points; extracting second text from the second text region based on the image and the second reference points, at least in part in parallel with extracting the first text; Deep learning-based optical character recognition methods, including:

4. A deep learning-based optical character recognition (OCR) method executed by at least one processor, comprising: Detecting at least one text region from the image based on model data obtained by deep learning; inferring and generating at least one reference point associated with the at least one detected text region so as to reduce a difference between the model data and the at least one detected text region; predicting and extracting at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, The step of extracting at least one text from the image comprises: generating a plurality of offset points adjacent to the at least one reference point; extracting at least one text from the image based on the image and the plurality of offset points; Deep learning-based optical character recognition methods, including:

5. The deep learning-based optical character recognition method of claim 4 , wherein the offset points guide text extraction by attention to regions adjacent to the offset points.

6. A deep learning-based optical character recognition (OCR) method executed by at least one processor, comprising: Detecting at least one text region from the image based on model data obtained by deep learning; inferring and generating at least one reference point associated with the at least one detected text region so as to reduce a difference between the model data and the at least one detected text region; predicting and extracting at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, the extracted text is word-by-word text, The method comprises: The deep learning-based optical character recognition method further includes a step of autoregressively detecting at least one character included in the extracted text using a classification score for the extracted text.

7. The deep learning based optical character recognition method of claim 6 , further comprising predicting a position and an angle of each of the at least one character within the image.

8. A deep learning-based optical character recognition (OCR) method executed by at least one processor, comprising: Detecting at least one text region from the image based on model data obtained by deep learning; inferring and generating at least one reference point associated with the at least one detected text region so as to reduce a difference between the model data and the at least one detected text region; predicting and extracting at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, Detecting at least one text region from the image comprises: extracting line-by-line text-related features from the image; generating position information for at least one text region based on the features; Deep learning-based optical character recognition methods, including:

9. A deep learning-based optical character recognition (OCR) method executed by at least one processor, comprising: Detecting at least one text region from the image based on model data obtained by deep learning; inferring and generating at least one reference point associated with the at least one detected text region so as to reduce a difference between the model data and the at least one detected text region; predicting and extracting at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, Detecting at least one text region from the image comprises: extracting text-related features from the image in paragraph units; generating position information for at least one text region based on the features; Deep learning-based optical character recognition methods, including:

10. A deep learning-based optical character recognition (OCR) method executed by at least one processor, comprising: Detecting at least one text region from the image based on model data obtained by deep learning; inferring and generating at least one reference point associated with the at least one detected text region so as to reduce a difference between the model data and the at least one detected text region; predicting and extracting at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, Detecting at least one text region from the image comprises: detecting at least one word-based text region from the image; detecting at least one line-by-line text region from the image; detecting at least one paragraph-based text region from the image; The deep learning-based optical character recognition method, wherein the steps of detecting word-based text regions, detecting line-based text regions, and detecting paragraph-based text regions are performed at least partially in parallel.

11. The deep learning-based optical character recognition method of claim 10 , further comprising detecting position information of at least one of a line-based text region or a paragraph-based text region including the extracted text.

12. A computer readable computer program for executing the method according to any one of claims 1 to 11 on a computer.

13. A deep learning-based Optical Character Recognition (OCR) system executed by at least one processor, comprising: a detector that detects at least one text region from an image based on model data obtained by deep learning, and infers and generates at least one reference point associated with the at least one text region so as to reduce a difference between the model data and the detected at least one text region; a recognizer that predicts and extracts at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, a backbone for extracting at least one text-related feature from the image; a transformer encoder for encoding the features; an optical character recognition system further comprising:

14. A deep learning-based Optical Character Recognition (OCR) system executed by at least one processor, comprising: a detector that detects at least one text region from an image based on model data obtained by deep learning, and infers and generates at least one reference point associated with the at least one text region so as to reduce a difference between the model data and the detected at least one text region; a recognizer that predicts and extracts at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, The detector comprises: extracting bounding polygons of text regions from the image using a segmentation map; extracting the center coordinates of the text region from the bounding polygon; An optical character recognition system that determines the center coordinate as the reference point.

15. A deep learning-based Optical Character Recognition (OCR) system executed by at least one processor, comprising: a detector that detects at least one text region from an image based on model data obtained by deep learning, and infers and generates at least one reference point associated with the at least one text region so as to reduce a difference between the model data and the detected at least one text region; a recognizer that predicts and extracts at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, The recognizer autoregressively predicts at least one text from a text instance associated with the at least one text region using deformable attention while referring to the image and the at least one reference point.

16. A deep learning-based Optical Character Recognition (OCR) system executed by at least one processor, comprising: a detector that detects at least one text region from an image based on model data obtained by deep learning, and infers and generates at least one reference point associated with the at least one text region so as to reduce a difference between the model data and the detected at least one text region; a recognizer that predicts and extracts at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, The recognizer a first recognizer for extracting first text from the image based on first reference points associated with a first text region and the image; a second recognizer for extracting second text from the image based on second reference points associated with second text regions and the image; An optical character recognition system, wherein the first recognizer and the second recognizer perform text extraction at least partially in parallel.

17. A deep learning-based Optical Character Recognition (OCR) system executed by at least one processor, comprising: a detector that detects at least one text region from an image based on model data obtained by deep learning, and infers and generates at least one reference point associated with the at least one text region so as to reduce a difference between the model data and the detected at least one text region; a recognizer that predicts and extracts at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, The recognizer an optical character recognition system that uses a classification score for the extracted text to autoregressively detect at least one character contained in the extracted text and predict a position and angle of each of the at least one character within the image.

18. A deep learning-based Optical Character Recognition (OCR) system executed by at least one processor, comprising: a detector that detects at least one text region from an image based on model data obtained by deep learning, and infers and generates at least one reference point associated with the at least one text region so as to reduce a difference between the model data and the detected at least one text region; a recognizer that predicts and extracts at least one text from the image based on the image and the at least one reference point so as to minimize a loss function; Including, The detector comprises: a first detector for detecting word-based text regions; a second detector for detecting line-by-line text regions; a third detector for detecting text regions in paragraph units; The optical character recognition system, wherein the first detector, the second detector, and the third detector perform text region detection at least partially in parallel.

Citation Information

Patent Citations

  • Text region determination method and device, equipment and readable storage medium

    CN113076814A

  • Apparatus for image optical character recoginition and method thereof

    KR1020150125376A

  • Text line detection in images

    US20160026899A1