Method and device for detecting text in image
By receiving instructions for image and text detection conditions and using a text detection model to generate a sequence of detection results, the problems of limitations in recognizing characters within images and the limited amount of output information in existing technologies are solved, and efficient text detection and visualization based on user requirements are achieved.
Patent Information
- Application Number
- CN202480012011.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-14
- Filing Date
- 2024-02-01
- Publication Date
- 2025-09-19
AI Technical Summary
Existing optical character recognition technology has limitations in recognizing characters within images, is unable to select appropriate detection formats based on user requirements, and has a limited amount of output information.
By receiving an image including text and instructions for text detection conditions, a text detection model is used to generate a sequence representing text detection results, thereby realizing detection and visualization of text instances in the image.
The detected text is visualized or output according to user preferences, and the text that meets the selected detection conditions is effectively detected or identified, solving the problem of limited output information volume in the prior art.
Smart Images

Figure CN120677510A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for detecting text in an image, and in particular, to a method and device for detecting text from an image based on an instruction including text detection conditions. Background Art
[0002] Optical character recognition (OCR) is a technology used to detect or recognize characters from images, including characters written by hand or printed by a printer. This technology is used not only to detect characters from images obtained by scanning or photographing documents containing characters, but also to recognize or translate characters in real time from images of characters printed on objects, plaques, and the like.
[0003] However, existing optical character recognition methods are limited to recognizing characters within images or detecting characters only in a predetermined format. Furthermore, during optical character recognition, there is a problem with selecting an appropriate detection format based on user requirements or performance differences among character detection methods for the target image. Summary of the Invention
[0004] Technical issues
[0005] In order to solve the above-mentioned problems, an object of the present invention is to provide a method for detecting text in an image, a non-transitory computer-readable recording medium having instructions recorded thereon, and a device (system).
[0006] Technical Solution
[0007] The present invention can be implemented in various ways, including methods, systems (devices) and / or computer programs stored in readable storage media.
[0008] According to one embodiment of the present invention, a method for detecting text in an image includes the following steps: receiving an image including text; receiving an instruction including text detection conditions; and inputting the image and the instruction into a text detection model to generate a sequence of detection results representing detection of text instances included in the image based on the text detection conditions.
[0009] An embodiment of the present invention provides a non-transitory computer-readable recording medium having recorded thereon instructions for executing a method for detecting text in an image on a computer.
[0010] An information processing system according to an embodiment of the present invention includes: a communication module; a memory; and at least one processor, the at least one processor being connected to the memory and configured to execute at least one computer-readable program contained in the memory. The communication module is configured to receive an image containing text and an instruction containing text detection conditions, wherein the at least one program includes instructions for inputting the image and the instruction into a text detection model to generate a sequence of detection results representing detection results of text instances contained in the image based on the text detection conditions.
[0011] Effects of the Invention
[0012] According to some embodiments of the present invention, as text in an image is detected based on various types of text detection conditions, the text detected from the image may be visualized or output according to a user-preferred detection position, region, or form.
[0013] According to some embodiments of the present invention, text in an image may be detected according to a detection condition selected from a plurality of different text detection conditions to effectively detect or recognize text that meets the selected detection condition.
[0014] Some embodiments of the present invention can address the limitations of existing sequence generation-based text detection models in outputting information in text detection results. Thus, text can be detected from the desired location or region within an image, regardless of the image size or length of the text input to the text detection model.
[0015] The effects of the present invention are not limited to the effects mentioned above. Ordinary technicians in the technical field to which the present invention belongs (hereinafter referred to as "ordinary technicians") can clearly understand other effects not mentioned through the description of the scope of the invention patent application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Hereinafter, multiple embodiments of the present invention will be described with reference to the accompanying drawings. In the process, like reference numerals represent like structures, but the present invention is not limited thereto.
[0017] Figure 1 2 is an illustrative diagram of the steps and structure for detecting text instances according to an embodiment of the present invention.
[0018] Figure 2 This is a schematic diagram of a connection structure for an information processing system according to an embodiment of the present invention to communicate with multiple user terminals in order to provide a text detection service within an image.
[0019] Figure 3 It is a block diagram showing the internal structure of a user terminal and an information processing system according to an embodiment of the present invention.
[0020] Figure 4 A block diagram illustrating the internal structure and input and output data of a text detection model according to an embodiment of the present invention.
[0021] Figure 5 An example diagram of a sequence generated by a text detection model according to an embodiment of the present invention.
[0022] Figure 6 FIG. 1 is a diagram showing an image of a detection result of a text instance based on a generated sequence visualization according to an embodiment of the present invention.
[0023] Figure 7 A diagram illustrating a process of sequentially detecting multiple text instances according to an embodiment of the present invention.
[0024] Figure 8 1 is a diagram showing text detection results according to an embodiment of the present invention and text detection results according to the prior art.
[0025] Figure 9 FIG. 1 is a flowchart illustrating a method for detecting text in an image according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] Hereinafter, the specific contents for implementing the present invention will be described in detail with reference to the accompanying drawings. However, in the following description, when there is a risk of unnecessarily obscuring the main purpose of the present invention, the specific description of the related known functions or structures will be omitted.
[0027] In the accompanying drawings, identical or corresponding structural elements are given the same reference numerals. Furthermore, in describing the following embodiments, repeated descriptions of identical or corresponding structural elements will be omitted. However, even if the description of a structural element is omitted, it does not mean that such structural element is not included in any embodiment.
[0028] The advantages, features and implementation methods of the disclosed embodiments can be found in the attached Figure 1 However, the present invention is not limited to the embodiments disclosed below and can be implemented through different implementation methods. Moreover, these embodiments are used to ensure the completeness of this disclosure. This disclosure is only provided to enable those skilled in the art to fully understand the scope of the present invention.
[0029] In this specification, the terms used will be briefly described and the disclosed embodiments will be described in detail. In this specification, although the common terms currently in widespread use have been selected as much as possible in consideration of the functions of the present invention, the terms used herein may change according to the intentions or conventions of ordinary technicians in the relevant fields, the emergence of new technologies, etc. In addition, in certain cases, there are terms arbitrarily selected by the applicant, in which case their meanings are described in detail in the corresponding description of the present invention. Therefore, in this specification, the terms used should be defined based on the meanings of the terms and the full content of this specification, and are not limited to the names of the terms.
[0030] In this specification, unless the context clearly indicates the singular, expressions in the singular include expressions in the plural. Furthermore, unless the context clearly indicates the plural, expressions in the plural include expressions in the singular. Throughout this specification, when a portion is indicated as including a certain structural element, unless there is a specific description to the contrary, it means that other structural elements may also be included, and does not exclude other structural elements.
[0031] Furthermore, in this specification, the term "module" or "section" used refers to a software structural element or a hardware structural element, and the "module" or "section" performs a certain function. However, a "module" or "section" is not limited to software or hardware. A "module" or "section" may exist in an accessible storage medium, or may also regenerate one or more processors. Therefore, as an example, a "module" or "section" may include at least one of a software structural element, an object-oriented software structural element, a class structural element, and a task structural element, a process, a function, a property, a program, a subroutine, a program code segment, a driver, firmware, microcode, a circuit, data, a database, a data structure, a list, an array, or a variable. The functions provided internally by the structural elements and the "module" or "section" can be combined by a smaller number of structural elements and "modules" or "sections" or further separated into additional structural elements and "modules" or "sections".
[0032] According to one embodiment of the present invention, a "module" or "unit" may be implemented by a processor and memory. "Processor" is broadly interpreted to include general-purpose processors, central processing units (CPUs), microprocessors, digital signal processors (DSPs), controllers, microcontrollers, state machines, and the like. In various contexts, "processor" may also refer to application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), and the like. For example, "processor" may also refer to a combination of processing devices, such as a digital signal processor and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors combined with a digital signal processing chip, or any other combination of such structures. Furthermore, "memory" is broadly interpreted to include any electronic component capable of storing electronic information. "Memory" may also refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage devices, and cache. When a processor reads information from or writes information to a memory, the memory and the processor are in electronic communication. Memory integrated with the processor is in electronic communication with the processor.
[0033] In the present invention, a "text instance" refers to a single text segment containing characters, numbers, symbols, etc., that is detected or recognized by a text recognition model. For example, if a text recognition model is trained to identify distance addresses from an image, the individual distance addresses used as learning data or detected, or the words that make up the distance addresses, can each be equivalent to a text instance.
[0034] Figure 1 FIG1 is an illustration of the steps and structure of detecting text instances according to an embodiment of the present invention. Figure 1 As shown, as an image 110 including text and text detection conditions 120 are input to a text detection model 130, a detection result 140 is generated, which detects text instances included in the image based on the text detection conditions. For example, image 110 may include various types of text, such as numbers, symbols, and characters in any language. Furthermore, image 110 and text detection conditions 120 may be received from a user.
[0035] In one embodiment, the text detection condition 120 can be input to the text detection model 130 in the form of instructions. For example, the text detection condition 120 can be input in the form of a sequence of one or more tokens arranged in a specified order. Alternatively, the text detection condition 120 can be input in the form of natural language. In this case, a natural language processing model that converts natural language into a sequence can be additionally used, or the text detection model 130 can additionally include a natural language conversion function that converts the natural language sentence representing the text detection condition 120 into a sequence.
[0036] The text detection conditions 120 may include information related to the detection form of the text instance. In one embodiment, the detection form may include at least one of the center point of the text instance, the bounding box of the text instance, or the polygon including the text instance. In this case, the detection form may be displayed as a predetermined number of coordinate forms for the detection form. For example, the detection form of the center point, bounding box, quadrilateral, and polygons other than quadrilaterals may be displayed as one or more coordinate forms, such as one coordinate representing each center point, two coordinates representing the upper left intersection point and the lower right intersection point of the bounding box, four intersection coordinates of the quadrilateral, and all intersection coordinates of the polygon.
[0037] For example, if the text detection conditions are input in a manner such that text within image 110 is detected based on the center point of a text instance, each position of the text instance within image 110 can be specified as the coordinates of each center point of the corresponding text instance (e.g., x, y = 100, 100).
[0038] In another example, if text within image 110 is detected based on the bounding box of the text instance, the respective positions of the text instance can be specified as the upper left intersection coordinates and the lower right intersection coordinates of the rectangle enclosing each text instance (e.g., x, y = 10, 10; x, y = 30, 40).
[0039] The text detection conditions may include at least one of a detection starting position or a detection area of a text instance within an image. In one embodiment, the text instance may be detected starting from a preset detection starting position within the image along a predetermined direction (e.g., from left to right and from top to bottom). In this case, if the detection starting position is not set, the upper left end point of the image may be designated as the detection starting position of the text instance. In yet another embodiment, the text instance existing within a preset detection area within the image may be detected. The detection area may be designated by various forms of graphics and their positions.
[0040] Additionally, the text detection conditions may include multiple types of conditions. For example, the text detection conditions may include detecting a language (e.g., detecting only text instances composed of a set detection language, outputting a result of translating a text instance composed of a set detection language, outputting a result of translating a text instance composed of a set language into another set language, etc.), a type of text instance (e.g., detecting only specific symbols or numbers or only characters in a text), a starting text of a text instance (e.g., detecting a text instance starting with "A"), a context of a text (e.g., detecting a title text, detecting an address or destination text), a detection order within an image (e.g., detecting the 13th to 20th instances of text instances included in an image), a specific emotional state (e.g., detecting a text instance expressing a sad emotion), etc., but are not limited thereto.
[0041] The text detection model 130 may be equivalent to a Transformer artificial neural network model including an encoder and a decoder. In this case, if the image 110 is input to the encoder of the text detection model 130 and the image features extracted by the encoder and the text detection conditions 120 are input to the decoder, the text detection model 130 can detect the text instance included in the image 110. According to one embodiment, in the learning step of the text detection model 130, the detection starting position is set to the upper left end point of the image with a probability of 0.5 and is set to any point in the image with the remaining probability, so that the text detection model 130 can learn to detect text instances starting from any detection starting position. Later, refer to Figure 4 The specific structure of the text detection model 130 and the input and output data of each structure are described in detail.
[0042] Detection results 140 can be generated by inputting image 110 and text detection conditions 120 into text detection model 130. For example, in first detection result 141, the positions of the center points of the text instances included in image 110 can be displayed on the image. Furthermore, in second detection result 142 and fifth detection result 145, the bounding boxes of the text instances included in image 110 can be displayed on the image. Furthermore, in third detection result 143 and fourth detection result 144, quadrilaterals and polygons including the text instances included in image 110 can be displayed on the image. On the other hand, in fifth detection result 145, with the detection starting position 145_1 of the specified text instance being specified, it can be confirmed that the detection results of the text instances are displayed on the image from the left to the right and from the top to the bottom starting from the corresponding detection starting position 145_1.
[0043] The detection result 140 may be specified or visualized by generating a sequence representing the detection results of the text instance included in the image 110 based on the text detection condition 120. Figure 5Describe in detail the exact structure of the sequence.
[0044] Figure 2 This is a schematic diagram of the connection structure for an information processing system 230 according to one embodiment of the present invention to communicate with multiple user terminals 210_1, 210_2, and 210_3 in order to provide an intra-image text detection service. The information processing system 230 may include a system (multiple systems) capable of providing an intra-image text detection service. In one embodiment, the information processing system 230 may include a computer executable program (e.g., a downloadable application) related to the intra-image text detection service and one or more server devices and / or databases capable of storing, providing, and executing data, or one or more distributed computing devices and / or distributed databases based on cloud computing services. For example, the information processing system 230 may include an additional system (e.g., a server) for the intra-image text detection service.
[0045] The text detection service within an image provided by the information processing system 230 may be provided to users through text detection applications installed in the plurality of user terminals 210_1 , 210_2 , and 210_3 .
[0046] Multiple user terminals 210_1, 210_2, and 210_3 can communicate with information processing system 230 via network 220. Network 220 enables communication between multiple user terminals 210_1, 210_2, and 210_3 and information processing system 230. For example, network 220 can be implemented by wired networks such as Ethernet, power line communication, telephone line communication, and RS-serial communication, mobile communication networks, wireless local area networks (WLANs), mobile hotspots (Wi-Fi), Bluetooth, and ZigBee, or a combination thereof, depending on the installation environment. The communication method is not limited and, in addition to communication methods using communication networks that can be included in network 220 (e.g., mobile communication networks, wired Internet, wireless Internet, broadcast networks, satellite networks, etc.), can also include short-range wireless communication between user terminals 210_1, 210_2, and 210_3.
[0047] For example, multiple user terminals 210_1, 210_2, and 210_3 may transmit images including text and instructions including text detection conditions to the information processing system 230 via the network 220, and the information processing system 230 may receive them.
[0048] exist Figure 2In the example of user terminals, although a mobile phone terminal 210_1, a tablet terminal 210_2 and a personal computer (PC) terminal 210_3 are shown, they are not limited thereto. The user terminals 210_1, 210_2 and 210_3 can be any computing device that can perform wired communication and / or wireless communication and can install and execute an application for detecting text in an image. For example, user terminals may include smart phones, mobile phones, navigators, computers, notebook computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), tablet computers, game consoles, wearable devices, Internet of Things (IoT) devices, virtual reality (VR) devices, augmented reality (AR) devices, etc. In addition, in Figure 2 In the figure, although three user terminals 210_1, 210_2, and 210_3 are shown communicating with the information processing system 230 via the network 220, the present invention is not limited thereto, and the number of user terminals communicating with the information processing system 230 via the network 220 may also be different.
[0049] Figure 3 The block diagram shows the internal structure of the user terminal 210 and the information processing system 230 according to one embodiment of the present invention. The user terminal 210 can refer to any computing device that can execute an application program for detecting text in an image and can perform wired / wireless communication, for example, it can include Figure 2 As shown in the figure, the user terminal 210 may include a memory 312, a processor 314, a communication module 316 and an input / output interface 318. Similarly, the information processing system 230 may include a memory 332, a processor 334, a communication module 336 and an input / output interface 338. Figure 3 As shown, the user terminal 210 and the information processing system 230 can use the communication modules 316 and 336 to send and receive information and / or data via the network 220. In addition, the input / output device 320 can input information and / or data to the user terminal 210 through the input / output interface 318, or can output information and / or data generated by the user terminal 210.
[0050] The memories 312 and 332 may include any non-transitory computer-readable recording medium. According to one embodiment, the memories 312 and 332 may include a non-volatile mass storage device such as a read-only memory (ROM), a disk drive, a solid-state drive (SSD), or a flash memory. As another example, a non-volatile mass storage device such as a read-only memory, a solid-state drive, a flash memory, or a disk drive may be included in the user terminal 210 or the information processing system 230 as an additional permanent storage device other than the memory. Furthermore, the memories 312 and 332 may store an operating system and at least one program code (e.g., code for an application program related to the text detection service in an image).
[0051] Such software structural elements may be loaded from an additional computer-readable recording medium different from the memories 312 and 332. Such additional computer-readable recording medium may include a recording medium that can be directly connected to the user terminal 210 and the information processing system 230, for example, a computer-readable recording medium such as a floppy disk drive, a disk, a tape, a DVD / CD-ROM drive, and / or a memory card. As another example, the software structural elements may also be loaded into the memories 312 and 332 via the communication modules 316 and 336 instead of the computer-readable recording medium. For example, at least one program may be loaded into the memories 312 and 332 based on a computer program (e.g., an application related to a text detection service in an image) installed from a file provided by a developer or a file distribution system that distributes application installation files via the network 220.
[0052] The processors 314 and 334 can process instructions of a computer program by performing basic arithmetic, logical, and input / output operations. Instructions can be provided to the processors 314 and 334 via the memory 312 and 332 or the communication modules 316 and 336. For example, the processors 314 and 334 can execute instructions received based on program code stored in a recording device such as the memory 312 and 332.
[0053] The communication modules 316 and 336 may provide a structure or function that enables the user terminal 210 and the information processing system 230 to communicate with each other via the network 220, and may provide a structure or function that enables the user terminal 210 and / or the information processing system 230 to communicate with other user terminals or other systems (e.g., additional cloud systems, etc.). As an example, a request or data (e.g., a text detection request, etc.) generated by the processor 314 of the user terminal 210 based on program code stored in a recording device such as the memory 312 may be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, the user terminal 210 may receive control information or instructions provided by the processor 334 of the information processing system 230 via the communication module 336 and the network 220 through the communication module 316 of the user terminal 210.
[0054] The input / output interface 318 refers to a device for connecting to the input / output device 320. As an example, the input device may include a camera with an audio sensor and / or an image sensor, a keyboard, a microphone, a mouse, and the like. Moreover, the output device may include a display, a speaker, a haptic feedback device, and the like. As another example, the input / output interface 318 may be an interface device that integrates the structure or function of a device for performing input and output functions into one, such as a touch screen. Figure 3 In the embodiment, although the input / output device 320 is not included in the user terminal 210, it is not limited to this and may constitute a single device together with the user terminal 210. Furthermore, the input / output interface 338 of the information processing system 230 may be connected to the information processing system 230, or may be an interface device for a device (not shown) for input or output that the information processing system 230 may include. Figure 3 In the embodiment, the input / output interfaces 318 and 338 are separated from the processors 314 and 334 as separate components, but the present invention is not limited thereto. The input / output interfaces 318 and 338 may be included in the processors 314 and 334 .
[0055] Compared to Figure 3The user terminal 210 and the information processing system 230 may include more structural elements. However, it is not necessary to explicitly show most of the existing technical structures. In one embodiment, the user terminal 210 may include at least a portion of the input and output devices 320. In addition, the user terminal 210 may also include other structural elements such as a radio transceiver, a global positioning system (GPS) module, a camera, various sensors, and a database. For example, if the user terminal 210 is a smart phone, it may include structural elements that a smart phone generally includes. For example, the user terminal 210 may also include an acceleration sensor, a gyroscope sensor, a microphone module, a camera module, various physical buttons, buttons using a touchpad, input and output ports, a vibrator for vibration, and other structural elements.
[0056] According to one embodiment, the processor 314 of the user terminal 210 may operate an in-image text detection application or a web browser application that provides an in-image text detection service. In this case, the program code associated with the corresponding application may be loaded into the memory 312 of the user terminal 210. While the application is operating, the processor 314 of the user terminal 210 may receive information and / or data provided by the input / output device 320 via the input / output interface 318, or may receive information and / or data from the information processing system 230 via the communication module 316, and may process the received information and / or data and store it in the memory 312. Furthermore, such information and / or data may be provided to the information processing system 230 via the communication module 316.
[0057] During the operation of the text-in-image detection application, the processor 314 may input or receive selected voice data, text, images, videos, etc. through an input device such as a touch screen, a keyboard, a camera with an audio sensor and / or an image sensor, a microphone, etc. connected to the input / output interface 318. The received voice data, text, images, and / or videos may be stored in the memory 312 or provided to the information processing system 230 through the communication module 316 and the network 220. In one embodiment, the processor 314 receives user input for selecting a graphical object displayed on the display through an input device, and may provide data / request corresponding to the received user input to the information processing system 230 through the network 220 and the communication module 316.
[0058] The processor 314 of the user terminal 210 may transmit and output information and / or data to the input / output device 320 via the input / output interface 318. For example, the processor 314 of the user terminal 210 may output the processed information and / or data via the output device 320, such as a display output device (e.g., a touch screen, a display, etc.) or a voice output device (e.g., a speaker).
[0059] The processor 334 of the information processing system 230 can manage, process, and / or store information and / or data received from multiple user terminals 210 and / or multiple external systems. The information and / or data processed by the processor 334 can be provided to the user terminal 210 via the communication module 336 and the network 220.
[0060] Figure 4 FIG. 4 is a block diagram illustrating the internal structure and input and output data of a text detection model 430 according to an embodiment of the present invention. Figure 4 The following figure shows an example of the structure of a text detection model 430 executed by an information processing system or a user terminal. The text detection model 430 may be equivalent to a Transformer model including an encoder 432 and a decoder 434, and its variants. For example, the text detection model 430 may be equivalent to a multi-way Transformer model.
[0061] The image 410 including text received from the user may be input to the encoder 432. In one embodiment, the encoder 432 may extract features of the image from the image 410. Next, the features of the image extracted by the encoder 432 may be input to the decoder 434.
[0062] The text detection conditions 420 received from the user or preset may be input to the decoder 434. If the text detection conditions 420 are not in a command form (e.g., natural language) that can be input to the decoder 434, a process may be additionally performed to convert the text detection conditions 420 into a command form before inputting to the decoder 434.
[0063] The decoder 434 can generate a sequence 440 related to the text instance included in the image 410 based on the features of the image and the text detection conditions 420. In one embodiment, the sequence 440 may include more than one sequence. For example, the sequence 440 may include: a starting sequence (first sequence), which represents the detection results of a predetermined number of text instances detected starting from the text detection starting position; and one or more intermediate sequences or ending sequences (second sequences), which represent the detection results of a predetermined number of text instances detected starting from the position of the last detection result of the starting position sequence. Through this structure, the length or amount of information of the detectable text instance is not limited, and all text instances included in the image can be detected. In this case, more than one sequence can be generated by an auto-regressive decoder respectively. In this regard, refer to later Figure 7 Provide detailed explanation.
[0064] In one embodiment, sequence 440 may represent at least one of the position or content of each text instance in image 410. Specifically, sequence 440 may include multiple markers representing the detection result of the text instance or the beginning and end of the detection result. That is, the multiple markers included in sequence 440 may respectively represent the position or content of each corresponding text instance in image 410. Figure 5 Provide detailed explanation.
[0065] Figure 5 This figure illustrates an example of a sequence 500 generated by a text detection model according to an embodiment of the present invention. Sequence 500 may represent at least one of the location or content of a text instance within an image. Specifically, sequence 500 may include a start position sequence 510 associated with the start position of text detection, a start marker 520 indicating the beginning of the detection results, a detection result sequence 530 indicating the detection results of the text instance, and an end marker 540 indicating the end of the detection results.
[0066] Detection result sequence 530 may represent detection results for a predetermined number of text instances starting from the text detection starting position. As the length of sequence 500 is limited, detection result sequence 530 may include sequences 531, 532, 533, 534, and 535, respectively, associated with a predetermined number (e.g., 20) of text instances. Sequences 531, 532, 533, 534, and 535 may respectively represent detection results for detecting one or more text instances along a predetermined direction (e.g., from the left to the right and from the top to the bottom) starting from the text detection starting position within the image.
[0067] Specifically, sequence 531 associated with the first detected text instance may include a detection form marker 531_1, a coordinate sequence 531_2 associated with the detected form, a content sequence 531_3 associated with the content of the text instance, and a padding sequence 531_4. However, this is not limiting and sequence 531 may also include markers and sequences associated with other text detection conditions.
[0068] The detection shape mark 531_1 may indicate the detection shape of the text instance. The detection shape mark 531_1 may be determined based on the detection condition input by the user. For example, if the detection shapes of the text instance input by the user are respectively a center point, a bounding box, a quadrilateral, and a polygon, the detection shape mark 531_1 may be respectively indicated as “ <point> ”、" <bbox> ”、" <quad> ”、" <poly>”.
[0069] In one embodiment, the length of the image coordinate sequence 531_2 may be determined according to the type of the detection form mark 531_1. For example, if the detection form marks 531_1 are respectively " <point> ”、" <bbox> ”、" <quad> ”、" <poly>", the image coordinate sequence 531_2 may be composed of multiple tags including 1 coordinate information, 2 coordinate information, 4 coordinate information, and 16 coordinate information, respectively. In this case, the coordinate information included in the image coordinate sequence 531_2 may correspond to a coordinate phenomenon normalized based on the size of each image.
[0070] Content sequence 531_3 may represent the content of a text instance. In one embodiment, content sequence 531_3 may be generated after coordinate sequence 531_2. In this case, coordinate sequence 531_2 representing the coordinate detection result and content sequence 531_3 representing the content detection result may be distinguished by separate timestamps.
[0071] In one embodiment, the content sequence 531_3 may include more than one token having each character included in the text instance. Subsequently, the content of the text instance may be determined by concatenating the tokens included in the content sequence 531_3.
[0072] The padding sequence 531_4 is appended to the next sequence of the content sequence 531_3 and can be used to constantly maintain the sum of the number of tokens included in the content sequence 531_3 and the padding sequence 531_4. For example, if the sum of the number of tokens included in the content sequence 531_3 and the padding sequence 531_4 is 25, and if the number of tokens in the content sequence 531_3 is 6, then the padding sequence 531_4 can be composed of 19 tokens. <pad>" tag. Differently, for example, if the number of characters included in the text instance is greater than 25, the tags corresponding to the different characters can be omitted so that the number of tags included in the content sequence 531_3 becomes 25. In this case, the padding sequence 531_4 may not be included in the sequence 531.
[0073] The description of the structure of the sequence 531 related to the first detected text instance is also applicable to the other sequences 532, 533, 534, 535.
[0074] In response to all text instances being detected within the image, an end marker 540 may be inserted. For example, if fewer than a predetermined number of text instances have been detected as of the bottom right end of the image, an end marker 540 may be inserted at the end of the sequence 500.
[0075] Alternatively, if a predetermined number of text instances have been detected before all text instances are detected at the lower right end of the image, end marker 540 may not be inserted. In this case, the position of the last detected text instance is updated as the detection starting position, and a predetermined number of other text instances may be detected from the updated detection starting position in the left-to-right and top-to-bottom directions. This process may be repeated until all text instances in the image are detected. End marker 540 may be inserted after all iterative processes are completed. Before inserting end marker 540, sequence 500 may further include one or more starting position sequences, starting markers, and detection result sequences corresponding to each iterative process.
[0076] Additionally, after each text instance is translated into a specific language based on sequence 500, it can be displayed instead of the existing text instance on the image, or summary information of the image and text instance can be provided (for example, information indicating that the corresponding image is presumed to be an image of a plaque of an accommodation store, etc.), and the text instance can also be provided through voice output.
[0077] Figure 6 FIG6 is a diagram illustrating an image 600 of a detection result of a text instance based on a generated sequence visualization according to an embodiment of the present invention. Figure 6 Shown based on reference Figure 5 The position and content information of each text instance included in the sequence example are displayed on the image 600. In this case, the position of each text instance in the image can also be based on Figure 5 Each detection form mark is displayed on the image in the corresponding detection form.
[0078] For example, in Figure 5 In, based on <point>"Detection of the position of the "OLD" text instance detected by the morphological marker can be displayed as the center point position of the "OLD" text instance on the image. As another example, based on the " <bbox>The position of the "MILL" text instance detected by the detection form mark can be displayed as a box form on the image. Figure 6 As shown, the corresponding text instance can be displayed around the detected text instance position. In this case, each detection form displayed on the image can specify the detection form based on a predetermined number of coordinate forms on the image.
[0079] Figure 7 1 is a diagram showing a process of sequentially detecting multiple text instances according to an embodiment of the present invention. Figure 7 A first decoder 742, a second decoder 744, a third decoder 746, and a fourth decoder 748 are shown separately, but they may constitute one decoder.
[0080] After the input image 710 is input to the encoder 720, the features of the input image extracted by the encoder 720 may be input to the first decoder 742. In this case, a first detection starting position (e.g., x=0, y=0) 732 may also be input to the first decoder 742. The first detection starting position 732 is received from the user terminal and may correspond to the upper leftmost endpoint of the input image 710.
[0081] The first decoder 742 may generate a first sequence related to text instances included in the input image based on features of the input image and the first detection start position 732. The first sequence may represent detection results of a predetermined number of text instances detected along a predetermined direction (e.g., left-to-right and top-to-bottom directions) starting from the text detection start position within the image. Additionally, a first output image 752 may be generated based on the first sequence, visualizing the detection results of the text instances on the image.
[0082] Next, the detection starting position may be updated to a second detection starting position (e.g., x=0, y=325) 734 as the position of the last detection result of the first sequence. The second decoder 744 may generate a second sequence representing the detection results of detecting a predetermined number of text instances along the left-to-right and top-to-bottom directions starting from the second detection starting position 734. Additionally, a second output image 754 may be generated based on the second sequence, in which the detection results included in the second sequence are visualized on the image.
[0083] As described above, this process can be repeatedly performed until all text instances in the input image 710 are detected. For example, the third decoder 746 can generate a third sequence representing detection results of detecting a predetermined number of text instances along the left-to-right and top-to-bottom directions starting from a third detection starting position (e.g., x=750, y=450) 736, which is the position of the last detection result of the second sequence. The fourth decoder 748 can generate a fourth sequence representing detection results of detecting a predetermined number of text instances along the left-to-right and top-to-bottom directions starting from a fourth detection starting position (e.g., x=500, y=675) 738, which is the position of the last detection result of the third sequence. Additionally, a third output image 756 and a fourth output image 758 can be generated, in which the detection results included in the third and fourth sequences are respectively visualized on the image.
[0084] In response to detecting all text instances within input image 710, a sequence including detection results for all text instances may be generated. For example, a sequence including detection results for all text instances may be generated based on the first to fourth sequences. Then, an output image 760 may be generated that visualizes the detection results for all text instances on input image 710.
[0085] Figure 8 1 is a diagram showing text detection results according to an embodiment of the present invention and text detection results according to the prior art.
[0086] In the text detection results 812 and 814 of the existing sequence generation-based optical character recognition technology, the existing text detection model may only display a portion of the text instances in the entire text due to the length limitation of the output data. In contrast, for the same image, in the text detection results 822 and 824 of an embodiment of the present invention, it can be confirmed that the text instances are based on the reference text. Figure 7 The described method detects and displays all instances of text contained in an image.
[0087] Figure 9 1 is a flowchart illustrating a method 900 for detecting text within an image according to an embodiment of the present invention. The method 900 may be executed by at least one processor of an information processing system or a user terminal. The method 900 may begin by the processor receiving an image including text (step S910).
[0088] Next, the processor may receive an instruction including text detection conditions (step S920). In one embodiment, the text detection conditions may include information related to the detection form of the text instance. For example, the detection form may include at least one of the center point of the text instance, the bounding box of the text instance, or the polygon including the text instance. The detection form may be displayed as a predetermined number of coordinate forms.
[0089] In one embodiment, the text detection condition may include at least one of a detection start position or a detection area of a text instance within an image.
[0090] In one embodiment, the text detection condition may include the detected language of the text instance.
[0091] Subsequently, the processor may input an image and instructions to the text detection model to generate a sequence representing a detection result of detecting a text instance included in the image based on a text detection condition (step S930). In this case, the text detection model may be equivalent to a Transformer model including an encoder and a decoder. Specifically, the text detection model may extract image features from an image through an encoder, and may generate a sequence related to a text instance included in the image based on the image features and instructions through a decoder. The sequence generated by the processor may represent at least one of the position or content of the text instance within the image, and may include multiple markers representing the detection result of the text instance or the beginning and end of the detection result.
[0092] In one embodiment, the sequence generated by the processor may include one or more sequences representing detection results of detecting one or more text instances along a predetermined direction starting from a text detection starting position within an image. In this case, the one or more sequences may include: a first sequence representing detection results of detecting a predetermined number of text instances starting from the text detection starting position; and a second sequence representing detection results of detecting a predetermined number of text instances starting from the position of the last detection result in the first sequence.
[0093] In one embodiment, if the text detection condition includes a detection language of a text instance, the processor may input an image and instructions to a text detection model to generate a sequence representing detection results of a text instance composed of the detection language.
[0094] The processor may then visualize the detection results of the text instances on the image based on the sequence. For example, the processor may display the location and content of each text instance on the image.
[0095] Figure 9 The flowchart shown and the content described are merely examples, and in some embodiments, different implementations may be used. For example, one or more steps may be omitted, or the order of the steps may be changed, or one or more steps may be performed simultaneously, or one or more steps may be performed repeatedly.
[0096] In order to enable the computer to execute, the method can be provided by a computer program stored in a computer-readable recording medium. The medium can also be a temporary storage for continuously storing, executing or downloading computer executable programs. In addition, the medium can be a variety of recording units or storage units formed by a single or multiple hardware combinations, and is not limited to media that directly access any computer system, and can also be distributed on a network. As an example, the medium includes magnetic media such as hard disks, floppy disks and tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floppy disks (floptical disks), and structures for storing program instructions such as ROM, RAM, and flash memory. In addition, as another example, the medium can be a recording medium or storage medium managed by an application store that sells application programs or a website or server that provides and sells a variety of other software.
[0097] The methods, work or techniques of the present invention may be implemented by a variety of units. For example, the techniques may also be implemented by hardware, firmware, software or a combination thereof. In connection with the disclosure of the present invention, it should be understood by those skilled in the art to which the present invention pertains that the various exemplary logic blocks, modules, circuits and algorithm steps described may also be implemented by electronic hardware, computer software or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, the above briefly describes various exemplary structural elements, blocks, modules, circuits and steps based on a functional perspective. However, whether such functionality is implemented by hardware or software depends on the design requirements assigned to the specific application and the overall system. Those skilled in the art to which the present invention pertains may also implement the described functionality in a variety of ways for each specific application, and therefore, such implementation should not be interpreted as departing from the scope of the present invention.
[0098] In a hardware instance, the processing unit used to perform the techniques may also be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions of the present invention, computers, or a combination thereof.
[0099] Therefore, the various exemplary logic blocks, modules, and circuits described in conjunction with the present invention may also be implemented or executed by a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete logic gates or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions of the present invention. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any existing processor, controller, microcontroller, or state machine. The processor may also be implemented by a combination of computing devices, for example, a digital signal processor, a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a digital signal processing chip, or a combination of any other structures.
[0100] In the firmware and / or software examples, the techniques may also be implemented by instructions stored on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, compact discs (CD), magnetic or optical data storage devices, etc. The instructions may be executed by one or more processors and may cause the processor(s) to perform the functions described herein according to a particular embodiment.
[0101] In the embodiments described above, the presently disclosed subject matter embodiments can be implemented using one or more standalone computer systems. However, the present invention is not limited thereto and can be implemented in any computing environment, such as a network or distributed computing environment. Furthermore, the present invention can be implemented using multiple processing chips or devices, and storage can be similarly implemented across multiple devices. Such devices can also include personal computers, network servers, and portable devices.
[0102] Although this specification describes some embodiments of the present disclosure, a person skilled in the art may make various modifications and variations without departing from the scope of the present invention. Such modifications and variations are also within the scope of the invention patent applications attached to this specification.< / bbox> < / point> < / pad> < / poly> < / quad> < / bbox> < / point> < / poly> < / quad> < / bbox> < / point>
Claims
1. A method for detecting text in an image, executed by at least one processor, characterized in that: The steps include: receiving an image including text; receiving an instruction including a text detection condition; and The image and the instruction are input to a text detection model to generate a sequence representing detection results of text instances included in the image based on the text detection conditions.
2. The method for detecting text in an image according to claim 1, wherein: The text detection model is a Transformer model including an encoder and a decoder.
3. The method for detecting text in an image according to claim 2, wherein: The steps of generating the sequence include the following steps: extracting features of the image from the image by the encoder; and A sequence related to a text instance included in the image is generated by the decoder from features of the image and the instructions.
4. The method for detecting text in an image according to claim 1, wherein: The text detection condition includes information related to the detection form of the text instance.
5. The method for detecting text in an image according to claim 4, wherein: The detection form includes at least one of a center point of the text instance, a bounding box of the text instance, or a polygon including the text instance.
6. The method for detecting text in an image according to claim 4, wherein: The detection form is displayed as a coordinate form of a number predetermined for the detection form.
7. The method for detecting text in an image according to claim 1, wherein: The document detection condition includes at least one of a detection start position or a detection area of the text instance in the image.
8. The method for detecting text in an image according to claim 1, wherein: The sequence includes one or more sequences representing detection results of detecting one or more text instances along a predetermined direction starting from a text detection start position within the image.
9. The method for detecting text in an image according to claim 8, wherein: The one or more sequences include: a starting position sequence representing detection results of a predetermined number of text instances starting from the text detection starting position; and The detection result sequence represents the detection results of detecting the predetermined number of text instances starting from the position of the last detection result of the starting position sequence.
10. The method for detecting text in an image according to claim 1, wherein: The text detection conditions include the detection language of the text instance, The step of generating the sequence includes the following steps: inputting the image and the instruction into a text detection model to generate a sequence representing detection results of the text instance composed of the detection language.
11. The method for detecting text in an image according to claim 1, wherein: The sequence represents the location or content of the text instance within the image.
12. The method for detecting text in an image according to claim 1, wherein: The sequence includes a plurality of tokens representing the detection result of the text instance or the beginning and the end of the detection result.
13. The method for detecting text in an image according to claim 1, wherein: The method further includes the following step: visualizing the detection result of the text instance on the image based on the sequence.
14. A non-transitory computer-readable recording medium, characterized in that Instructions for executing the method according to claim 1 on a computer are recorded.
15. An information processing system, characterized in that: include: Communication module; Memory; and at least one processor, connected to the memory, configured to execute at least one computer-readable program contained in the memory, The communication module is used to receive an image including text and receive an instruction including text detection conditions, The at least one program includes instructions for executing the following steps: inputting the image and the instructions into a text detection model to generate a sequence representing detection results of text instances included in the image based on the text detection conditions.