Method and apparatus for detecting text in an image
The method and apparatus address the limitations of conventional OCR by allowing flexible text detection based on user-defined conditions, enabling efficient recognition of text in images according to specified criteria.
Patent Information
- Application Number
- JP2025542164
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-14
- Filing Date
- 2024-02-01
- Publication Date
- 2026-02-10
AI Technical Summary
Conventional OCR methods are limited in recognizing characters only in predetermined formats and lack the ability to select appropriate detection formats based on user requirements, leading to inefficient text detection.
A method and apparatus for detecting text in an image that includes receiving an image and a text detection condition, using a text detection model to generate a sequence indicating detection results that fit the specified criteria, allowing for flexible text detection based on user-defined conditions.
Enables the detection of text according to various types and locations desired by the user, overcoming limitations of conventional OCR by efficiently recognizing text regardless of image size or content length.
Smart Images

Figure 2026504946000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a method and apparatus for detecting text in an image, and more particularly to a method and apparatus for detecting text from an image based on an instruction including a text detection condition. [Background technology]
[0002] Optical Character Recognition (OCR) generally refers to the technology of detecting or recognizing characters from images containing printed or handwritten characters, such as those printed on a printer. OCR technology is used to detect characters from images obtained by scanning or photographing documents containing characters, but it is also used to recognize and translate characters in real time from images of characters printed on objects, signs, etc.
[0003] However, conventional OCR methods have limitations in that they can only recognize the characters themselves in an image and can only detect characters in a predetermined format. Furthermore, when performing OCR, there is a problem in that it is not possible to select an appropriate detection format depending on the user's requirements and the performance of the character detection method for the image to be text-detected. Summary of the Invention [Problem to be solved by the invention]
[0004] To solve the above-mentioned problems, the present disclosure provides a method for detecting text in an image, a computer-readable non-transitory recording medium having instructions recorded thereon, and an apparatus (system). [Means for solving the problem]
[0005] The present disclosure may be implemented in numerous ways, including as a method, a system (apparatus), or a computer program stored on a readable recording medium.
[0006] According to one embodiment of the present disclosure, a method for detecting text in an image includes receiving an image containing text, receiving an instruction word containing a text detection condition, and inputting the image and the instruction word into a text detection model to generate a sequence indicating detection results of text instances contained in the image that fit the text detection condition.
[0007] A non-transitory computer-readable recording medium is provided that stores instructions for executing a method for detecting text in an image according to an embodiment of the present disclosure on a computer.
[0008] According to one embodiment of the present disclosure, an information processing system includes a communication module, a memory, and at least one processor coupled to the memory and configured to execute at least one computer-readable program contained in the memory, wherein the communication module receives an image containing text and receives instructions including text detection criteria, the at least one program including instructions for inputting the image and the instructions into a text detection model and generating a sequence indicating detection results of text instances contained in the image that meet the text detection criteria. [Effects of the Invention]
[0009] According to some embodiments of the present disclosure, text in an image can be detected according to various types of text detection conditions, allowing the user to visualize or output the detected text from the image according to the detection location, area, or form desired by the user.
[0010] According to some embodiments of the present disclosure, by detecting text in an image according to a detection condition selected from a plurality of different text detection conditions, it is possible to efficiently detect or recognize text that is suitable for the selected detection condition.
[0011] According to some embodiments of the present disclosure, it is possible to solve the problem of limited information volume of the text detection result output by a conventional sequence generation-based image text detection model, and thus, it is possible to detect text from a position or area in an image desired by a user, regardless of the size of the image input to the text detection model or the length of the text contained in the image.
[0012] The effects of the present disclosure should not be limited to those described above, and other effects not described will be clearly understood by a person having ordinary knowledge in the technical field to which the present disclosure pertains (hereinafter referred to as "a person skilled in the art") from the description in the claims. [Brief explanation of the drawings]
[0013] Embodiments of the present disclosure will now be described with reference to the accompanying drawings, in which like reference numerals indicate like elements, but are not limited to the drawings. [Figure 1] 10A and 10B are diagrams illustrating an example of a detection order and results of text instances according to an embodiment of the present disclosure. [Figure 2] 1 is a schematic diagram illustrating a configuration in which an information processing system is communicatively connected to a plurality of user terminals to provide a text-in-image detection service according to an embodiment of the present disclosure. [Figure 3] 1 is a block diagram showing the internal configuration of a user terminal and an information processing system according to an embodiment of the present disclosure. [Figure 4] FIG. 2 is a block diagram showing the internal structure and input / output data of a text detection model according to an embodiment of the present disclosure. [Figure 5] FIG. 1 illustrates an example sequence generated by a text detection model according to an embodiment of the present disclosure. [Figure 6] FIG. 10 is a diagram illustrating an image in which the results of detecting text instances based on the generated sequences are visualized according to one embodiment of the present disclosure. [Figure 7]FIG. 1 illustrates a process by which multiple text instances are detected in sequence in accordance with one embodiment of the present disclosure. [Figure 8] FIG. 10 is a diagram illustrating a text detection result according to an embodiment of the present disclosure and a text detection result according to the prior art. [Figure 9] 1 is a flowchart illustrating a method for detecting text in an image, according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the accompanying drawings. However, in the following description, detailed descriptions of well-known functions and configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.
[0015] In the accompanying drawings, the same or corresponding components are denoted by the same reference numerals. Furthermore, in the following description of the embodiments, duplicated descriptions of the same or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such a component is not included in any of the embodiments.
[0016] The advantages and features of the disclosed embodiments, as well as methods for achieving them, will become apparent from the following detailed description of the embodiments, taken in conjunction with the accompanying drawings. However, the present disclosure should not be construed as being limited to the following embodiments, and may be realized in various different forms. These embodiments are merely provided so that the present disclosure will be complete and will enable those skilled in the art to fully appreciate the scope of the invention.
[0017] The terms used in this specification will be briefly explained, and the disclosed embodiments will be described in detail. The terms used in this specification have been selected as widely used and common terms as much as possible, taking into consideration the functionality of the present disclosure. However, these terms may vary depending on the intentions of engineers in the relevant field, legal precedents, the emergence of new technologies, etc. In addition, terms arbitrarily selected by the applicant may also be used in specific cases. In such cases, the meanings of the terms will be described in detail in the relevant sections of the description of the invention. Therefore, the terms used in this disclosure should be defined based on the meaning of the terms and the overall content of the present disclosure, rather than simply by the name of the terms.
[0018] In this specification, the singular form includes the plural form unless the context clearly dictates otherwise. Furthermore, the plural form includes the singular form unless the context clearly dictates otherwise. Throughout the specification, when a part is described as including a certain element, it does not mean that other elements are excluded, but that other elements may also be included, unless otherwise specified.
[0019] Additionally, the terms "module" or "unit" as used herein refer to a software or hardware component, and the "module" or "unit" performs a certain function. However, this does not imply that the "module" or "unit" is limited to software or hardware. A "module" or "unit" may be configured to reside on an addressable storage medium or to execute one or more processors. Thus, as an example, a "module" or "unit" may include at least one of the following components: a software component, an object-oriented software component, a class component, and a task component; a process; a function; an attribute; a procedure; a subroutine; a program code segment; a driver; firmware; microcode; a circuit; data; a database; a data structure; a table; an array; or a variable. The functionality provided within a component and a "module" or "unit" may be combined into a "module" or "unit" with fewer components or further separated into additional components and "modules" or "units."
[0020] According to one embodiment of the present disclosure, a "module" or "unit" may be implemented with a processor and memory. "Processor" should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, "processor" may refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), etc. "Processor" may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such configuration. Additionally, "memory" should be broadly interpreted to include any electronic component capable of storing electronic information. "Memory" may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), nonvolatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage devices, and registers. A memory and a processor are said to be in electronic communication if the processor can read information from or store information in the memory. Memory that is integrated into a processor is in electronic communication with the processor.
[0021] In this disclosure, a "text instance" may refer to a portion of a single piece of text, including letters, numbers, symbols, etc., that a text recognition model is to detect or recognize. For example, if a text recognition model is trained to recognize addresses from images, each address used as training data or each word that makes up an address may correspond to a text instance.
[0022] 1 is a diagram illustrating an example of a text instance detection order and results according to an embodiment of the present disclosure. As shown in FIG. 1, an image 110 containing text and a text detection condition 120 may be input to a text detection model 130, thereby generating a detection result 140 of text instances contained in the image that are suitable for the text detection condition. For example, the image 110 may include various types of text, such as numbers, symbols, and characters composed of any language. The image 110 and the text detection condition 120 may be received from a user.
[0023] In one embodiment, the text detection condition 120 may be input to the text detection model 130 in the form of an instruction. For example, the text detection condition 120 may be input in the form of a sequence in which one or more tokens are arranged in a certain order. Alternatively, the text detection condition 120 may be input in the form of natural language. In this case, a natural language processing model that converts natural language into a sequence may be additionally used, or a natural language conversion function may be additionally included in the text detection model 130, so that a natural language sentence indicating the text detection condition 120 may be converted into a sequence.
[0024] The text detection condition 120 may include information related to the detection form of a text instance. In one embodiment, the detection form may include at least one of the center point of the text instance, the bounding box of the text instance, or a polygon including the text instance. In this case, the detection form may be displayed in a predetermined number of coordinate forms for the detection form. For example, the detection form of a center point, a bounding box, a rectangle, or a polygon other than a rectangle may be displayed in one or more coordinate forms, such as one coordinate indicating the center point, two coordinates indicating the upper left vertex and the lower right vertex of the bounding box, the coordinates of the four vertices of the rectangle, or the coordinates of all vertices of the polygon.
[0025] For example, if the text detection criteria are input such that text in image 110 is detected relative to the center point of the text instance, the location of each text instance in image 110 may be specified as the coordinates of the center point of each text instance (e.g., (x, y) = (100, 100)).
[0026] In another example, if text in image 110 is detected relative to the bounding box of the text instance, the location of each text instance may be identified as the coordinates of the upper left vertex and the lower right vertex of a rectangle that encloses each text instance (e.g., (x, y)=(10, 10), (x, y)=(30, 40)).
[0027] The text detection condition may include at least one of a detection start position or a detection area of a text instance in an image. In one embodiment, a text instance may be detected in a predetermined direction (e.g., left-to-right or top-to-bottom) from a preset detection start position in the image. In this case, if a detection start position is not set, the top left corner of the image may be specified as the detection start position of the text instance. In another embodiment, a text instance existing within a preset detection area in the image may be detected. The detection area may be specified by various shapes and their positions.
[0028] Additionally, the text detection condition may include other types of conditions, such as, but not limited to, a detection language (e.g., detecting only text instances written in a set detection language, outputting the results of translating text instances into a set detection language, or outputting the results of translating only text instances written in a set language into another set language), a type of text instance (e.g., detecting only specific symbols or numbers from text, or detecting only letters), a start text of a text instance (e.g., detecting text instances beginning with “A”), a text context (e.g., detecting title text, detecting address or destination text), a detection order within an image (e.g., detecting the 13th to 20th instances of text instances included in an image), and a specific emotional state (e.g., detecting text instances expressing sadness).
[0029] The text detection model 130 may correspond to a transformer artificial neural network model including an encoder and a decoder. In this case, when the image 110 is input to the encoder of the text detection model 130 and the image features extracted by the encoder and the text detection condition 120 are input to the decoder, the text detection model 130 may detect text instances included in the image 110. In one embodiment, during the training phase of the text detection model 130, the detection start position is set to the top left corner of the image with a probability of 0.5, and an arbitrary point within the image with the remaining probability, so that the text detection model 130 may be trained to detect text instances from any detection start position. Specific configurations of the text detection model 130 and input / output data of each configuration will be described in detail with reference to FIG. 4.
[0030] The detection result 140 may be generated by inputting the image 110 and the text detection conditions 120 into the text detection model 130. For example, in the first detection result 141, the positions of the center points of each of the text instances included in the image 110 may be displayed on the image. In the second detection result 142 and the fifth detection result 145, the bounding boxes of each of the text instances included in the image 110 may be displayed on the image. In the third detection result 143 and the fourth detection result 144, a rectangle and a polygon including each of the text instances included in the image 110 may be displayed on the image. Meanwhile, in the fifth detection result 145, a detection start position 145_1 of the text instance is identified, and it can be seen that the detection results of the text instance are displayed on the image from left to right and top to bottom from the detection start position 145_1.
[0031] The detection results 140 may be identified or visualized by generating a sequence showing the detection of text instances contained in the image 110 that match the text detection criteria 120. The specific configuration of the sequence will be described in more detail with reference to FIG.
[0032] 2 is a schematic diagram illustrating a configuration in which an information processing system 230 is communicatively connected to multiple user terminals 210_1, 210_2, and 210_3 to provide a text-in-image detection service according to one embodiment of the present disclosure. The information processing system 230 may include a system(s) capable of providing the text-in-image detection service. In one embodiment, the information processing system 230 may include one or more server devices and / or databases, or one or more distributed computing devices and / or distributed databases based on cloud computing services, capable of storing, providing, and executing computer-executable programs (e.g., downloadable applications) and data related to the text-in-image detection service. For example, the information processing system 230 may include a separate system (e.g., a server) for the text-in-image detection service.
[0033] The in-image text detection service provided by the information processing system 230 may be provided to users through a text detection application installed on each of the multiple user terminals 210_1, 210_2, and 210_3.
[0034] A plurality of user terminals 210_1, 210_2, and 210_3 may communicate with the information processing system 230 via a network 220. The network 220 may be configured to enable communication between the plurality of user terminals 210_1, 210_2, and 210_3 and the information processing system 230. Depending on the installation environment, the network 220 may be configured as a wired network such as Ethernet, a wired home network (Power Line Communication), a telephone line communication device, and RS-serial communication, a mobile communication network, a wireless network such as WLAN (Wireless LAN), Wi-Fi, Bluetooth, and ZigBee, or a combination thereof. The communication method is not limited, and may include not only a communication method utilizing a communication network that can be included in the network 220 (for example, a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, a satellite network, etc.), but also a short-range wireless communication between the user terminals 210_1, 210_2, and 210_3.
[0035] For example, the plurality of user terminals 210_1, 210_2, and 210_3 may transmit an image including text and a command including a text detection condition to the information processing system 230 via the network 220, and the information processing system 230 may receive the command.
[0036] 2 illustrates a mobile phone terminal 210_1, a tablet terminal 210_2, and a PC terminal 210_3 as examples of user terminals, but the user terminals 210_1, 210_2, and 210_3 may be any computing devices capable of wired and / or wireless communication and capable of installing and executing an in-image text detection application, etc. For example, the user terminals may include smartphones, mobile phones, navigation systems, PCs, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), tablet PCs, game consoles, wearable devices, internet of things (IoT) devices, virtual reality (VR) devices, and augmented reality (AR) devices. Also, while FIG. 2 shows three user terminals 210_1, 210_2, and 210_3 communicating with the information processing system 230 via the network 220, this is not limited thereto, and a different number of user terminals may be configured to communicate with the information processing system 230 via the network 220.
[0037] FIG. 3 is a block diagram showing the internal configuration of a user terminal 210 and an information processing system 230 according to an embodiment of the present disclosure. The user terminal 210 may refer to any computing device capable of executing a text-in-image detection application and capable of wired / wireless communication, and may include, for example, the mobile phone terminal 210_1, tablet terminal 210_2, or PC terminal 210_3 shown in FIG. 2 . As shown in the figure, the user terminal 210 may include a memory 312, a processor 314, a communication module 316, and an input / output interface 318. Similarly, the information processing system 230 may include a memory 332, a processor 334, a communication module 336, and an input / output interface 338. As shown in FIG. 3 , the user terminal 210 and the information processing system 230 may be configured to communicate information and / or data via the network 220 using their respective communication modules 316 and 336. Additionally, the input / output device 320 may be configured to input information and / or data to the user terminal 210 via the input / output interface 318, and to output information and / or data generated from the user terminal 210.
[0038] The memory 312, 332 may include any non-transitory computer-readable storage medium. According to one embodiment, the memory 312, 332 may include a permanent mass storage device such as a read only memory (ROM), a disk drive, a solid state drive (SSD), or a flash memory. As another example, a permanent mass storage device such as a ROM, an SSD, a flash memory, or a disk drive may be included in the user terminal 210 or the information processing system 230 as a separate permanent storage device distinct from the memory. The memory 312, 332 may also store an operating system and at least one program code (e.g., code for an application associated with the in-image text detection service).
[0039] Such software components may be loaded from a computer-readable recording medium separate from the memory 312, 332. Such separate computer-readable recording medium may include a recording medium directly connectable to the user terminal 210 and the information processing system 230, such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, or a memory card. As another example, the software components may be loaded into the memory 312, 332 through the communication modules 316, 336, which are not computer-readable recording media. For example, at least one program may be loaded into the memory 312, 332 based on a computer program (e.g., an application associated with an in-image text detection service) to be installed by a file provided over the network 220 by a developer or a file distribution system that distributes application installation files.
[0040] The processors 314, 334 may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 314, 334 by the memories 312, 332 or the communications modules 316, 336. For example, the processors 314, 334 may be configured to execute instructions received according to program code stored in a storage device, such as the memories 312, 332.
[0041] The communication modules 316, 336 may provide configurations or functions for the user terminal 210 and the information processing system 230 to communicate with each other via the network 220, and may provide configurations or functions for the user terminal 210 and / or the information processing system 230 to communicate with other user terminals or other systems (such as, for example, a separate cloud system) via the network 220. As an example, a request or data (e.g., a text detection request) generated by the processor 314 of the user terminal 210 in accordance with program code stored in a storage device such as the memory 312 may be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, a control signal or command provided under the control of the processor 334 of the information processing system 230 may be received by the user terminal 210 via the communication module 316 of the user terminal 210 via the communication module 336 and the network 220.
[0042] The input / output interface 318 may be a means for interfacing with the input / output device 320. As an example, the input device may include a device such as a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, or a mouse, while the output device may include a device such as a display, a speaker, or a haptic feedback device. As another example, the input / output interface 318 may be a means for interfacing with a device in which the configuration or functions for performing input and output are integrated into one, such as a touchscreen. Although FIG. 3 illustrates a device in which the input / output device 320 is not included in the user terminal 210, this is not limiting and the input / output device 320 and the user terminal 210 may be configured as a single device. Furthermore, the input / output interface 338 of the information processing system 230 may be a means for interfacing with an input or output device (not shown) connected to or included in the information processing system 230. Although FIG. 3 shows the input / output interfaces 318, 338 as elements configured separately from the processors 314, 334, this is not limiting and the input / output interfaces 318, 338 may be configured to be included in the processors 314, 334.
[0043] The user terminal 210 and the information processing system 230 may include more components than those shown in FIG. 3 . However, it is not necessary to clearly illustrate most of the conventional components. In one embodiment, the user terminal 210 may be implemented to include at least some of the input / output devices 320 described above. The user terminal 210 may also include other components such as a transceiver, a global positioning system (GPS) module, a camera, various sensors, and a database. For example, if the user terminal 210 is a smartphone, the user terminal 210 may include components typically included in a smartphone, such as an acceleration sensor, a gyro sensor, a microphone module, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration.
[0044] According to one embodiment, the processor 314 of the user terminal 210 may be configured to run a text-in-image detection application or a web browser application that provides text-in-image detection services. Program code associated with the application may be loaded into the memory 312 of the user terminal 210. During application operation, the processor 314 of the user terminal 210 may receive information and / or data provided by the input / output device 320 via the input / output interface 318 or from the information processing system 230 via the communication module 316, process the received information and / or data, and store it in the memory 312. Furthermore, such information and / or data may be provided to the information processing system 230 via the communication module 316.
[0045] During operation of the text-in-image detection application, the processor 314 may receive voice data, text, images, videos, etc. input or selected via input devices such as a touchscreen, keyboard, camera including an audio sensor and / or image sensor, microphone, etc. connected to the input / output interface 318, and may store the received voice data, text, images, videos, etc. in the memory 312 or provide them to the information processing system 230 via the communication module 316 and the network 220. In one embodiment, the processor 314 may receive user input to select a graphical object displayed on the display via the input device, and provide data / requests corresponding to the received user input to the information processing system 230 via the network 220 and the communication module 316.
[0046] The processor 314 of the user terminal 210 may transmit and output information and / or data to the input / output device 320 via the input / output interface 318. For example, the processor 314 of the user terminal 210 may output the processed information and / or data through the output device 320, such as a display-capable device (e.g., a touch screen, a display, etc.), an audio-capable device (e.g., a speaker), or the like.
[0047] The processor 334 of the information processing system 230 may be configured to manage, process, and / or store information and / or data received from the plurality of user terminals 210 and / or the plurality of external systems. The information and / or data processed by the processor 334 may be provided to the user terminal 210 via the communication module 336 and the network 220.
[0048] 4 is a block diagram showing the internal configuration and input / output data of a text detection model 430 according to an embodiment of the present disclosure. Fig. 4 shows an example of the configuration of the text detection model 430 executed by an information processing system or a user terminal. The text detection model 430 may correspond to a transformer model including an encoder 432 and a decoder 434, or a modified model thereof. For example, the text detection model 430 may correspond to a multi-way transformer model.
[0049] An image 410 containing text received from a user may be input to an encoder 432. In one embodiment, the encoder 432 may extract image features from the image 410. The image features extracted by the encoder 432 may then be input to a decoder 434.
[0050] The text detection condition 420, which is received from a user or preset, may be input to the decoder 434. If the text detection condition 420 is not configured in the form of an instruction word that can be input to the decoder 434 (e.g., a natural language), a process of converting the text detection condition 420 into the form of an instruction word may be separately performed before the text detection condition 420 is input to the decoder 434.
[0051] The decoder 434 may generate a sequence 440 associated with the text instances included in the image 410 from the image features and the text detection criteria 420. In one embodiment, the sequence 440 may include one or more sequences. For example, the sequence 440 may include a start sequence (first sequence) indicating the detection results of a predetermined number of text instances from the text detection start position, and one or more intermediate or end sequences (second sequences) indicating the detection results of a predetermined number of text instances from the position of the last detection result in the start sequence. This configuration allows all text instances included in the image to be detected without limiting the length or amount of information of the detectable text instances. In this case, each of the one or more sequences may be generated by an auto-regressive decoder, as described in more detail with reference to FIG. 7.
[0052] In one embodiment, sequence 440 may indicate at least one of the location or content of each text instance in image 410. Specifically, sequence 440 may include multiple tokens that indicate the detection result or the start and end of the detection result of the text instance. That is, each of the multiple tokens included in sequence 440 may indicate the location or content of each corresponding text instance in image 410. This will be described in more detail with reference to FIG. 5.
[0053] 5 is a diagram illustrating an example of a sequence 500 generated by a text detection model according to an embodiment of the present disclosure. The sequence 500 may indicate at least one of the location or content of a text instance within an image. Specifically, the sequence 500 may include a start position sequence 510 associated with a text detection start position, a start token 520 indicating the start of the detection result, a detection result sequence 530 indicating the detection result of the text instance, and an end token 540 indicating the end of the detection result.
[0054] The detection result sequence 530 may indicate the detection of a predetermined number of text instances from the text detection start position. Because the length of the sequence 500 is limited, the detection result sequence 530 may include sequences 531, 532, 533, 534, and 535 associated with each of the predetermined number (e.g., 20) of text instances. Each of the sequences 531, 532, 533, 534, and 535 may indicate the detection of one or more text instances in a predetermined direction (e.g., from left to right and from top to bottom) from the text detection start position in the image.
[0055] Specifically, the sequence 531 associated with the first detected text instance may include a detected form token 531_1, a coordinate sequence 531_2 associated with the detected form, a content sequence 531_3 associated with the content of the text instance, and a padding sequence 531_4, although not limited thereto, and the sequence 531 may further include tokens or sequences associated with other text detection conditions.
[0056] The detection form token 531_1 may indicate the detection form of the text instance. The detection form token 531_1 may be determined based on the detection conditions input by the user. For example, if the detection forms of the text instance input by the user are a center point, a bounding box, a rectangle, and a polygon, the detection form token 531_1 may be " <point> 」、「 <bbox> 」、「 <quad> 」、「 <poly>" may be displayed.
[0057] In one embodiment, the length of the image coordinate sequence 531_2 may be determined depending on the type of the detected form token 531_1. For example, if the detected form tokens 531_1 are each <point> 」、「 <bbox> 」、「 <quad> 」、「 <poly>", the image coordinate sequence 531_2 may be composed of a plurality of tokens each including 1, 2, 4, or 16 pieces of coordinate information. In this case, the coordinate information included in the image coordinate sequence 531_2 may correspond to coordinate information normalized based on the size of each image.
[0058] The content sequence 531_3 may indicate the content of the text instance. In one embodiment, the content sequence 531_3 may be generated after the coordinate sequence 531_2. In this case, the coordinate sequence 531_2 indicating the coordinate detection result and the content sequence 531_3 indicating the content detection result may be distinguished by individual time stamps.
[0059] In one embodiment, content sequence 531_3 may include one or more tokens, one for each character included in the text instance, and the tokens included in content sequence 531_3 may then be concatenated to determine the content of the text instance.
[0060] The padding sequence 531_4 may be added next to the content sequence 531_3 so that the total number of tokens included in the content sequence 531_3 and the padding sequence 531_4 remains constant. For example, if the total number of tokens included in the content sequence 531_3 and the padding sequence 531_4 is predetermined to be 25 and the number of tokens in the content sequence 531_3 is 6, the padding sequence 531_4 may contain 19 " <pad>Alternatively, for example, if the number of characters included in the text instance exceeds 25, tokens corresponding to the excess characters may be omitted so that the number of tokens included in the content sequence 531_3 is 25. In this case, the padding sequence 531_4 may not be included in the sequence 531.
[0061] What has been described for the construction of the sequence 531 associated with the first detected text instance is equally applicable to the other sequences 532, 533, 534, 535.
[0062] The termination token 540 may be inserted in response to all text instances being found in the image, for example, if a predetermined number of text instances or fewer have been found up to the bottom right corner of the image, the termination token 540 may be inserted at the end of the sequence 500.
[0063] Alternatively, if a predetermined number of text instances have already been detected before all text instances have been detected up to the bottom right corner of the image, the end token 540 may not be inserted. In this case, the position of the last detected text instance may be updated as the detection start position, and a predetermined number of additional text instances may be detected from the updated detection start position, from left to right and from top to bottom. This process may be repeated until all text instances in the image have been detected. The end token 540 may be inserted after all iterations have been completed, and the sequence 500 may further include one or more start position sequences, start tokens, and detection result sequences corresponding to each iteration before the end token 540 is inserted.
[0064] Additionally, each text instance based on sequence 500 may be translated into a particular language and then displayed on the image replacing an existing text instance, or a summary of the image or text instance may be provided (e.g., information to the effect that the image is likely to be an image of a hotel sign), or the text instance may be provided as an audio output.
[0065] Fig. 6 is a diagram illustrating an image 600 in which the detection result of a text instance based on a sequence generated by one embodiment of the present disclosure is visualized. Fig. 6 illustrates that the position and content of each text instance included in the example sequence described with reference to Fig. 5 are displayed in the image 600 based on the position and content information within the image of the respective text instances. In this case, the position of each text instance within the image may be displayed on the image in the corresponding detection form based on each detection form token in Fig. 5.
[0066] For example, in Figure 5, <point>The location of the "OLD" text instance detected by the detected morphological token may be displayed on the image as the location of the center point of the "OLD" text instance. <bbox>The location of the "MILL" text instance detected according to the detected feature token may be displayed in the form of a box on the image. As shown in FIG. 6, the text instance may be displayed near the location of the detected text instance. In this case, each of the detected features displayed on the image may be identified based on a predetermined number of coordinates on the image for the detected feature.
[0067] 7 is a diagram illustrating a process for sequentially detecting multiple text instances according to one embodiment of the present disclosure. For ease of explanation, the first through fourth decoders 742, 744, 746, and 748 are shown as separate decoders in FIG. 7, but they may also be configured as a single decoder.
[0068] After the input image 710 is input to the encoder 720, the features of the input image extracted by the encoder 720 may be input to a first decoder 742. At this time, a first detection start position 732 (e.g., x=0, y=0) may also be input to the first decoder 742. The first detection start position 732 may be received from a user terminal or may correspond to the top left corner of the input image 710.
[0069] The first decoder 742 may generate a first sequence associated with the text instances included in the input image based on the features of the input image and the first detection start position 732. The first sequence may indicate the detection results of a predetermined number of text instances in a predetermined direction (e.g., from left to right and from top to bottom) from the text detection start position in the image. Additionally, a first output image 752 may be generated based on the first sequence, in which the detection results of the text instances are visualized on the image.
[0070] The detection start position may then be updated to a second detection start position 734 (e.g., x=0, y=325), which is the position of the last detection result in the first sequence. The second decoder 744 may generate a second sequence indicating the detection results of a predetermined number of text instances from left to right and top to bottom from the second detection start position 734. Additionally, a second output image 754 may be generated based on the second sequence, in which the detection results included in the second sequence are visualized on the image.
[0071] Similarly, this process may be repeated until all text instances in the input image 710 have been detected. For example, the third decoder 746 may generate a third sequence indicating the detection results of a predetermined number of text instances from left to right and top to bottom starting from a third detection start position 736 (e.g., x=750, y=450), which is the position of the last detection result in the second sequence, and the fourth decoder 748 may generate a fourth sequence indicating the detection results of a predetermined number of text instances from left to right and top to bottom starting from a fourth detection start position 738 (e.g., x=500, y=675), which is the position of the last detection result in the third sequence. Additionally, a third output image 756 and a fourth output image 758 may be generated in which the detection results included in the third and fourth sequences are visualized on the image, respectively.
[0072] In response to detecting all text instances in the input image 710, a sequence including the detection results of the entire text instances may be generated. For example, a sequence including the detection results of the entire text instances may be generated based on the first to fourth sequences. Then, an output image 760 may be generated in which the detection results of the entire text instances are visualized on the input image 710.
[0073] FIG. 8 is a diagram showing a text detection result according to an embodiment of the present disclosure and a text detection result according to the prior art.
[0074] In the text detection results 812 and 814 using conventional sequence generation-based OCR technology, only some text instances of the entire text are displayed due to the length limitation of the output data in the conventional text detection model. In contrast, in the text detection results 822 and 824 for the same image according to an embodiment of the present disclosure, it can be seen that all text instances included in the image are detected and displayed in the manner described with reference to FIG.
[0075] 9 is a flowchart illustrating a method 900 for detecting text in an image according to one embodiment of the present disclosure. The method 900 may be performed by at least one processor of an information processing system or a user terminal. The method 900 may begin by the processor receiving an image containing text (S910).
[0076] The processor may then receive an instruction including a text detection condition (S920). In one embodiment, the text detection condition may include information associated with a detection form of the text instance. For example, the detection form may include at least one of a center point of the text instance, a bounding box of the text instance, or a polygon containing the text instance, and may be displayed in the form of a predetermined number of coordinates for the detection form.
[0077] In one embodiment, the text detection criteria may include at least one of a detection start position or a detection area of a text instance within an image.
[0078] In one embodiment, the text detection criteria may include the detection language of the text instance.
[0079] The processor may then input the image and the instruction words into a text detection model and generate a sequence indicating a detection result of a text instance included in the image that meets the text detection conditions (S930). In this case, the text detection model may correspond to a transformer model including an encoder and a decoder. Specifically, the text detection model may extract image features from the image using an encoder, and generate a sequence associated with the text instance included in the image from the image features and the instruction words using a decoder. The sequence generated by the processor may indicate at least one of the location or content of the text instance within the image, and may include a plurality of tokens indicating the detection result of the text instance or the start and end of the detection result.
[0080] In one embodiment, the sequences generated by the processor may include one or more sequences indicating detection of one or more text instances in a predetermined direction from a text detection start location within the image, where the one or more sequences may include a first sequence indicating detection of a predetermined number of text instances from the text detection start location and a second sequence indicating detection of a predetermined number of text instances from a location of the last detection result in the first sequence.
[0081] In one embodiment, if the text detection condition includes a detection language for the text instance, the processor may input the image and instruction words into a text detection model and generate a sequence indicating the detection results for the text instance configured in the detection language.
[0082] The processor may then visualize the detection results of the text instances on the image based on the sequence, for example, the processor may display the position and content of each text instance on the image.
[0083] 9 and the above description are merely examples and may be implemented differently in some embodiments, such as by omitting one or more steps, changing the order of the steps, overlapping one or more steps, or repeating one or more steps multiple times.
[0084] The above-described method may be provided as a computer program recorded on a computer-readable recording medium for execution by a computer. The medium may continuously record a computer-executable program or may temporarily record the program for execution or download. The medium may be a variety of recording or storage means in the form of a single piece of hardware or a combination of multiple pieces of hardware, and may be a medium directly connected to a computer system or a medium distributed over a network. Examples of media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to record program instructions, such as ROM, RAM, and flash memory. Other examples of media include recording media or storage media managed by app stores, websites, and servers that distribute and distribute applications and other various software.
[0085] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, such techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will appreciate that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the present disclosure may be implemented in electronic hardware, computer software, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design requirements imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementation should not be interpreted as departing from the scope of the present disclosure.
[0086] In a hardware implementation, the processing units utilized to perform the techniques may be implemented with one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in this disclosure, computers, or combinations thereof.
[0087] Accordingly, the various illustrative logic blocks, modules, and circuits described in connection with this disclosure may be implemented with or performed by a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other configuration.
[0088] In a firmware and / or software implementation, the techniques may be implemented with instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processor(s) to perform certain aspects of the functions described in this disclosure.
[0089] While the above-described embodiments have been described as utilizing aspects of the presently disclosed subject matter on one or more stand-alone computer systems, the present disclosure should not be limited thereto and may be implemented in connection with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the presently disclosed subject matter may be implemented on multiple processing chips or devices, and storage may be shared across multiple devices. These devices may include PCs, network servers, and portable devices.
[0090] Although the present disclosure has been described herein with reference to certain embodiments, various modifications and changes may be made thereto without departing from the scope of the present disclosure, which would be understood by those skilled in the art to which the present disclosure pertains. Furthermore, such modifications and changes must be considered to fall within the scope of the appended claims.< / bbox> < / point> < / pad> < / poly> < / quad> < / bbox> < / point> < / poly> < / quad> < / bbox> < / point>
Claims
1. 1. A method for detecting text in an image, performed by at least one processor, comprising: receiving an image containing text; receiving a command including a text detection condition; and a step of inputting the image and the instruction words into a text detection model and generating a sequence indicating detection results of text instances included in the image that meet the text detection criteria.
2. The method of claim 1 , wherein the text detection model is a transformer model that includes an encoder and a decoder.
3. The step of generating the sequence comprises: extracting, by the encoder, image features from the image; and The method of claim 2 , further comprising generating, by the decoder, a sequence associated with a text instance contained in the image from the image features and the instruction words.
4. The method of claim 1 , wherein the text detection condition includes information associated with a detection manner of the text instance.
5. The method of claim 4 , wherein the detection features include at least one of a center point of the text instance, a bounding box of the text instance, or a polygon containing the text instance.
6. The method of claim 4 , wherein the detected feature is displayed in the form of a predetermined number of coordinates for the detected feature.
7. The method of claim 1 , wherein the text detection conditions include at least one of a detection start position or a detection area of the text instance in the image.
8. The method of claim 1 , wherein the sequences include one or more sequences indicating detection results of one or more text instances in a predetermined direction from a text detection start position in the image.
9. The one or more sequences a sequence of start positions indicating detection results for a predetermined number of text instances from the text detection start positions; and The method of claim 8 , further comprising a detection result sequence indicating detection results for the predetermined number of text instances from the position of the last detection result in the start position sequence.
10. the text detection condition includes a detection language of the text instance; The step of generating the sequence comprises: The method of claim 1 , further comprising inputting the image and the instruction words into a text detection model and generating a sequence indicating detection results of text instances composed of the detection language.
11. The method of claim 1 , wherein the sequence indicates the location or content of the text instance within the image.
12. The method of claim 1 , wherein the sequence includes a plurality of tokens indicating the detection of the text instance or the beginning and end of the detection.
13. The method of claim 1 , further comprising the step of visualizing the detection results of the text instances based on the sequence on the image.
14. A program for executing the method according to claim 1 on a computer.
15. An information processing system, communication module, memory, and at least one processor coupled to the memory and configured to execute at least one computer-readable program contained in the memory; Including, The communication module includes: Receive an image containing text, receiving a command including a text detection condition; The at least one program an instruction for inputting the image and the instruction into a text detection model and generating a sequence indicating detection results of text instances included in the image that meet the text detection criteria;
Citation Information
Patent Citations
Character translation method and device
JP2019537103A
A spatial attention model for image caption generation
JP2019537147A
Processing visual input
JP2020534590A
Information providing method and system based on pointing
JP2022167734A
Text extraction method, text extraction model training method, device and equipment
JP2022172381A