Text recognition method, device, medium and electronic device
By decoding the position of the target text line in the text recognition process and performing parallel decoding, the problem of separation of detection modules and identification modules in the prior art is solved, and more efficient and accurate text recognition is achieved.
Patent Information
- Application Number
- CN202210411604.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-04-19
AI Technical Summary
The existing end-to-end text recognition network cannot achieve the deep fusion of detection modules and identification modules, resulting in text detection and recognition being divided into two parts, affecting the recognition accuracy.
The text content is synchronized by decoding the position of the target text line in the image to be identified in the decoding stage, and then decoding all target text lines in parallel, thereby reducing the complexity of the decoding cycle.
On the premise of ensuring recognition accuracy, the text recognition process is greatly accelerated, the impact of text detection accuracy on recognition accuracy is reduced, and the recognition accuracy is improved.
Smart Images

Figure CN114758342B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of text recognition, and in particular, to a text recognition method, device, medium and electronic device. Background Art
[0002] The current mainstream text recognition framework requires text detection as a pre-task, and then recognizes the detected text area. Such a framework has a complex process, a long model calculation time, and requires more parameters to be stored. It also requires additional manual joint debugging. There are also many end-to-end text recognition networks, but the existing end-to-end text recognition networks still cannot achieve a deep integration of the detection module and the recognition module. Text detection and text recognition are still divided into two parts. Therefore, the accuracy of text recognition still needs to rely on the accuracy of text detection. Inaccurate text detection will greatly affect the accuracy of text recognition. Summary of the invention
[0003] This summary is provided to introduce concepts in a brief form that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, the present disclosure provides a text recognition method, which includes: acquiring an image to be recognized; extracting image features in the image to be recognized; determining the coordinate position of a target text line according to the image features; and decoding the target text line in parallel according to the coordinate position to synchronously obtain the text in all the target text lines in the image to be recognized.
[0005] In a second aspect, the present disclosure provides a text recognition device, which includes: an acquisition module for acquiring an image to be recognized; a feature extraction module for extracting image features in the image to be recognized; a coordinate determination module for determining the coordinate position of a target text line according to the image features; and a text recognition module for decoding the target text line in parallel according to the coordinate position, and synchronously obtaining the text in all the target text lines in the image to be recognized.
[0006] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when the program is executed by a processing device.
[0007] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device on which a computer program is stored; and a processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0008] Through the above technical scheme, when performing end-to-end text recognition on the image to be recognized, the position of the target text line in the image to be recognized can be decoded first in the decoding stage, and then all the target text lines in the image to be recognized can be decoded in parallel, and the text in each target text line in the image to be recognized can be obtained by synchronous decoding and recognition, thereby greatly reducing the decoding cycle complexity in the text recognition process, and greatly accelerating the text recognition while ensuring the model recognition accuracy. In addition, since there is no need to accurately locate the position of the text line in the image to be recognized, it is only necessary to determine the coordinates of any point in the text line to realize text recognition of the text line, thereby greatly reducing the influence of text detection accuracy on text recognition accuracy, and improving the accuracy of text recognition.
[0009] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale. In the drawings:
[0011] Figure 1 The figure is a flowchart of a text recognition method according to an exemplary embodiment of the present disclosure.
[0012] Figure 2 is a flowchart of a text recognition method according to another exemplary embodiment of the present disclosure.
[0013] Figure 3 It is a schematic diagram of a decoding sequence of a target text line in a text recognition method according to yet another exemplary embodiment of the present disclosure.
[0014] Figure 4 is a flowchart of a text recognition method according to another exemplary embodiment of the present disclosure.
[0015] Figure 5 It is a structural block diagram of a text recognition device according to an exemplary embodiment of the present disclosure.
[0016] Figure 6 A schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0018] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0019] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0020] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0021] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0024] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0025] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0026] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0027] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0028] Figure 1 FIG. 1 is a flowchart of a text recognition method according to an exemplary embodiment of the present disclosure. Figure 1 As shown, the method includes steps 101 to 104.
[0029] In step 101, an image to be identified is obtained. The image to be identified may be an image of a preset size including any image content.
[0030] In step 102, image features are extracted from the image to be identified.
[0031] In this embodiment, the method for extracting the image features can be to extract the image features through any feature extraction network. For example, the feature map can be extracted from the image to be identified only through a convolutional neural network (CNN) as the image feature, or the feature map can be first extracted from the image to be identified through the convolutional neural network (CNN), and then the feature map is further encoded through an encoder to obtain the encoded image features. Specifically, the method for extracting the image features of the image to be identified can be set according to the needs of actual applications, and the method for extracting image features is not limited in the present disclosure.
[0032] In step 103, the coordinate position of the target text line is determined according to the image features.
[0033] In step 104, the target text lines are decoded in parallel according to the coordinate positions, and the text in all the target text lines in the image to be recognized is obtained synchronously.
[0034] The target text line is a text line that may be predicted to exist in the image to be recognized. The coordinate position of the target text line can be the coordinates (X, Y) of any point that can ensure the position of the target text line, where X can be the horizontal coordinate corresponding to the width of the image to be recognized, and Y can be the vertical coordinate corresponding to the height of the image to be recognized.
[0035] That is, after acquiring the image features in the image to be identified, the coordinate positions of the text lines that may exist in the image to be identified are first predicted, and then the text is further identified based on the predicted coordinate text of the target text line, and in the text recognition step, all target text lines can be decoded in parallel, which greatly reduces the complexity of the decoding process.
[0036] Take an image to be recognized including 60 lines of text as an example, assuming that the maximum length of the text in each text line is 25, then the total length of each text line including the coordinate position (X, Y) is 27 (the length of the coordinate position is 2). When recognizing text by conventional autoregressive decoding method, a total of 60*27=1620 loop decodings are required; while by the text decoding method given in step 103 and step 104 of this embodiment, only 60*2 loop decodings (the length of the coordinate position is 2) are required to determine the coordinate position of the target text line in step 103. Since the 60 target text lines can be decoded in parallel in step 104, only the number of decodings corresponding to the length of the text in one target text line (25 times) is required to realize the text recognition of all 60 target text lines, and only 60*2+25=145 loop decodings are required in total, which reduces the number of loop decodings by 11 times compared with conventional text recognition, and reduces the decoding cycle complexity from O(n*T) to O(n*2+(T-2)).
[0037] Through the above technical scheme, when performing end-to-end text recognition on the image to be recognized, the position of the target text line in the image to be recognized can be decoded first in the decoding stage, and then all the target text lines in the image to be recognized can be decoded in parallel, and the text in each target text line in the image to be recognized can be obtained by synchronous decoding and recognition, thereby greatly reducing the decoding cycle complexity in the text recognition process, and greatly accelerating the text recognition while ensuring the model recognition accuracy. In addition, since there is no need to accurately locate the position of the text line in the image to be recognized, it is only necessary to determine the coordinates of any point in the text line to realize text recognition of the text line, thereby greatly reducing the influence of text detection accuracy on text recognition accuracy, and improving the accuracy of text recognition.
[0038] Figure 2 FIG. 1 is a flowchart of a text recognition method according to another exemplary embodiment of the present disclosure. Figure 2 As shown, the method includes step 201.
[0039] In step 201, the coordinate position of each target text line is used as a start symbol to decode the target text lines in parallel in the same decoding sequence, and to synchronously obtain the text in all the target text lines.
[0040] The decoding sequence for parallel decoding of the target text line can be as follows Figure 3 As shown in . Wherein, the target text line 1 is any target text line in the image to be recognized, including the coordinate position 2 (coord: x, y) as the start symbol, the recognized text content (transcription: HELLO), and the completion character 4 (<PAD>: Φ); the other target text lines in the image to be recognized are arranged after the target text line 1 in a random order, and the decoding length corresponding to each target text line can be a fixed value. If the text content in the target text line is not long enough, the completion character 4 can be used to align and complete the text of insufficient length; the sentence start identifier 5 ( <sos>:<) and sentence end identifier 6 ( <eos>:>) is the start mark and end mark of the entire decoding sequence. In this way, when all target text lines in the image to be recognized are decoded in parallel, the coordinate position corresponding to each target text line can be used as the start mark to decode all target text lines synchronously in the decoding sequence, so as to synchronously obtain the text in all target text lines in the image to be recognized.
[0041] Figure 4 FIG. 1 is a flowchart of a text recognition method according to another exemplary embodiment of the present disclosure. Figure 4 As shown, the method includes step 401.
[0042] In step 401, the category to which the target text line actually lies in the image to be identified is predicted based on the image features, and the category is determined as the coordinate position of the target text line; wherein the category includes a horizontal coordinate category and a vertical coordinate category.
[0043] That is, when determining the coordinate position corresponding to each target text line in the image to be identified, it is done through a classification prediction method. The correspondence between the category and the real coordinate position in the image to be identified can be pre-set. For example, the horizontal coordinate (image width) and vertical coordinate (image height) of the image to be identified can be divided into N areas and M areas respectively. The range of coordinates included in each area corresponds to a category. For example, the first category of the horizontal coordinate can correspond to the coordinate position in the range of 0-100 on the horizontal coordinate (X axis), and the first category of the vertical coordinate can correspond to the coordinate position in the range of 0-90 on the vertical coordinate (Y axis). When predicting the coordinate position, if the category of the coordinate position of the target text line is predicted to be the first category of the horizontal coordinate and the first category of the vertical coordinate, then the coordinate position of the target text line is represented to be within the interval ((0-100), (0-90)).
[0044] In conventional end-to-end text recognition methods, for the detection of text positions, a regression method is usually used to detect specific text positions. Due to the size of the image, the text position coordinates finally predicted are often of huge values, which leads to problems such as a long detection process and inaccurate position detection results. However, the method in this embodiment no longer uses a regression method to predict the specific coordinates of the text in the image to be recognized, but directly determines the coordinate position of each target text line in the image to be recognized by classification prediction through the mapping relationship between the pre-constructed coordinate information and the category information. This can speed up the prediction speed of the coordinate position, and since the number of categories can be controlled within a certain number, it can further accelerate the decoding and recognition of the target text line. In addition, since the accuracy of the category prediction of the text position is higher than that of directly predicting the specific coordinate position by regression, it is less likely to produce errors, so it can also reduce the influence of the text detection accuracy on the text recognition accuracy to a certain extent, and further improve the accuracy of text recognition.
[0045] In a possible implementation, the extracting of image features in the image to be identified includes: extracting image features in the image to be identified by an encoder based on sliding window encoding. That is, in the feature extraction process, when using an encoder to determine the image features, the encoder based on sliding window encoding can be used. In this way, the time loss of the encoder's self-attention mechanism can be reduced, and the encoding speed can be further improved. At the same time, due to the small number of sliding windows, the encoding accuracy can also be improved to a certain extent. Among them, the encoder based on sliding window encoding can be, for example, an encoder in Swin-transformer.
[0046] In a possible implementation, the extracting of image features from the image to be recognized, the determining of the coordinate position of the target text line according to the image features, and the decoding of the target text line in parallel according to the coordinate position to synchronously obtain the text in all the target text lines in the image to be recognized are all implemented by a pre-trained end-to-end text recognition model. That is, the pre-trained end-to-end text recognition model is used in the present disclosure to perform the following steps: Figure 1 The text recognition steps from step 102 to step 104 are shown in FIG.
[0047] Among them, the text recognition method in the present disclosure also includes a training method for training the end-to-end text recognition model, specifically including: training the end-to-end text recognition model through target training data, wherein the target training data is obtained through a target data augmentation method, and the target data augmentation method includes at least one of the following: rotation, cropping, random change of image size, random change of image attributes, and point enhancement.
[0048] In a possible implementation, the annotation data in the target training data is the coordinates of a single annotation point representing the image position of each text line in the target training data, and the annotation point corresponds to the text line one by one. In this way, when annotating the training data to obtain the target training data, only a single point annotation is required at any position in each text line in the training data, which reduces a lot of annotation costs and is more conducive to the training of the end-to-end text recognition model.
[0049] Figure 5 FIG. 1 is a structural block diagram of a text recognition device according to an exemplary embodiment of the present disclosure. Figure 5 As shown, the device includes: an acquisition module 10, used to acquire the image to be identified; a feature extraction module 20, used to extract image features in the image to be identified; a coordinate determination module 30, used to determine the coordinate position of the target text line according to the image features; a text recognition module 40, used to decode the target text line in parallel according to the coordinate position, and synchronously obtain the text in all the target text lines in the image to be identified.
[0050] Through the above technical scheme, when performing end-to-end text recognition on the image to be recognized, the position of the target text line in the image to be recognized can be decoded first in the decoding stage, and then all the target text lines in the image to be recognized can be decoded in parallel, and the text in each target text line in the image to be recognized can be obtained by synchronous decoding and recognition, thereby greatly reducing the decoding cycle complexity in the text recognition process, and greatly accelerating the text recognition while ensuring the model recognition accuracy. In addition, since there is no need to accurately locate the position of the text line in the image to be recognized, it is only necessary to determine the coordinates of any point in the text line to realize text recognition of the text line, thereby greatly reducing the influence of text detection accuracy on text recognition accuracy, and improving the accuracy of text recognition.
[0051] In a possible implementation, the text recognition module 40 is further used to: use the coordinate position of each target text line as a start symbol to decode the target text lines in parallel in the same decoding sequence to synchronously obtain the text in all the target text lines.
[0052] In a possible implementation, the coordinate determination module 30 is also used to: predict the category to which the target text line actually lies in the image to be identified based on the image features, and determine the category as the coordinate position of the target text line; wherein the category includes a horizontal coordinate category and a vertical coordinate category.
[0053] In a possible implementation manner, the feature extraction module 20 is further configured to extract image features in the image to be identified by using an encoder that performs encoding based on a sliding window.
[0054] In a possible implementation, the feature extraction module 20, the coordinate determination module 30, and the text recognition module 40 are all modules in a pre-trained end-to-end text recognition model.
[0055] In a possible implementation, the device further includes: a model training module (not shown), configured to train the end-to-end text recognition model using target training data, wherein the target training data is obtained using a target data augmentation method, and the target data augmentation method includes at least one of the following: rotation, cropping, random change of image size, random change of image attributes, and point enhancement.
[0056] In a possible implementation, the annotation data in the target training data are coordinates of a single annotation point in the target training data that represents an image position where each text line is located, and the annotation point corresponds to the text line one by one.
[0057] Reference below Figure 6 , which shows a schematic diagram of the structure of an electronic device 600 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0058] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0059] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0060] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0061] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0062] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0063] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0064] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: obtains the image to be recognized; extracts image features in the image to be recognized; determines the coordinate position of the target text line according to the image features; decodes the target text line in parallel according to the coordinate position, and synchronously obtains the text in all the target text lines in the image to be recognized.
[0065] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0066] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0067] The modules involved in the embodiments described in the present disclosure may be implemented by software or hardware. The name of a module does not limit the module itself in some cases. For example, an acquisition module may also be described as a "module for acquiring an image to be recognized".
[0068] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0069] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0070] According to one or more embodiments of the present disclosure, Example 1 provides a text recognition method, including: acquiring an image to be recognized; extracting image features in the image to be recognized; determining the coordinate position of a target text line based on the image features; decoding the target text line in parallel based on the coordinate position, and synchronously obtaining the text in all the target text lines in the image to be recognized.
[0071] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein the target text lines are decoded in parallel according to the coordinate positions, and the text in all the target text lines is synchronously obtained, including: using the coordinate positions of each of the target text lines as a starting symbol to decode the target text lines in parallel in the same decoding sequence, and the text in all the target text lines is synchronously obtained.
[0072] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1, wherein determining the coordinate position of the target text line based on the image features includes: predicting the category to which the target text line belongs at its actual position in the image to be identified based on the image features, and determining the category as the coordinate position of the target text line; wherein the category includes a horizontal axis category and a vertical axis category.
[0073] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 1, wherein extracting image features in the image to be identified includes: extracting image features in the image to be identified by an encoder that performs encoding based on a sliding window.
[0074] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1, wherein the extracting of image features in the image to be recognized, the determining of the coordinate positions of the target text lines based on the image features, and the parallel decoding of the target text lines based on the coordinate positions to synchronously obtain the text in all the target text lines in the image to be recognized are all implemented by a pre-trained end-to-end text recognition model.
[0075] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 5, which further includes: training the end-to-end text recognition model through target training data, wherein the target training data is obtained through a target data augmentation method, and the target data augmentation method includes at least one of the following: rotation, cropping, random change of image size, random change of image attributes, and point enhancement.
[0076] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 6, wherein the annotation data in the target training data is the coordinates of a single annotation point in the target training data that represents the image position of each text line, and the annotation point corresponds one-to-one to the text line.
[0077] According to one or more embodiments of the present disclosure, Example 8 provides a text recognition device, including: an acquisition module for acquiring an image to be recognized; a feature extraction module for extracting image features in the image to be recognized; a coordinate determination module for determining the coordinate position of a target text line according to the image features; and a text recognition module for decoding the target text line in parallel according to the coordinate position, and synchronously obtaining the text in all the target text lines in the image to be recognized.
[0078] According to one or more embodiments of the present disclosure, Example 9 provides a computer-readable medium having a computer program stored thereon, which implements the steps of any of the methods described in Examples 1-7 when executed by a processing device.
[0079] According to one or more embodiments of the present disclosure, Example 10 provides an electronic device, comprising: a storage device on which a computer program is stored; and a processing device for executing the computer program in the storage device to implement the steps of any one of the methods described in Examples 1-7.
[0080] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0081] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0082] Although the subject matter has been described in language specific to structural features and / or method logic actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims. Regarding the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment related to the method, and will not be elaborated here.< / eos> < / sos>
Claims
1. A text recognition method, It is characterized in that The method comprises: Obtain an image to be recognized; Extracting image features from the image to be identified; Determine the coordinate position of the target text line according to the image features; By using an end-to-end text recognition model, the target text lines are decoded in parallel according to the coordinate positions, and the text in all the target text lines in the image to be recognized is synchronously obtained; The decoding of the target text lines in parallel according to the coordinate positions to synchronously obtain the text in all the target text lines comprises: The coordinate position of each target text line is used as a start symbol to decode the target text lines in parallel in the same decoding sequence, and to synchronously obtain the text in all the target text lines.
2. The method according to claim 1, It is characterized in that Determining the coordinate position of the target text line according to the image features includes: Predicting the category to which the target text line actually lies in the image to be recognized based on the image features, and determining the category as the coordinate position of the target text line; The categories include a horizontal axis category and a vertical axis category.
3. The method according to claim 1, It is characterized in that The extracting of image features from the image to be identified comprises: The image features in the image to be recognized are extracted by an encoder that performs encoding based on a sliding window.
4. The method according to claim 1, It is characterized in that The extracting of image features in the image to be recognized, the determining of the coordinate positions of target text lines according to the image features, and the parallel decoding of the target text lines according to the coordinate positions to synchronously obtain the text in all the target text lines in the image to be recognized are all implemented by a pre-trained end-to-end text recognition model.
5. The method according to claim 4, It is characterized in that The method further comprises: The end-to-end text recognition model is trained by target training data, wherein the target training data is obtained by a target data augmentation method, and the target data augmentation method includes at least one of the following: rotation, cropping, random change of image size, random change of image attributes, and point enhancement.
6. The method according to claim 5, It is characterized in that The annotation data in the target training data are the coordinates of a single annotation point in the target training data that represents the image position of each text line, and the annotation point corresponds to the text line one by one.
7. A text recognition device, It is characterized in that The device comprises: An acquisition module, used for acquiring an image to be recognized; A feature extraction module, used to extract image features from the image to be identified; A coordinate determination module, used to determine the coordinate position of the target text line according to the image features; A text recognition module, configured to decode the target text lines in parallel according to the coordinate positions through an end-to-end text recognition model, and synchronously obtain the text in all the target text lines in the image to be recognized; The text recognition module is further used to: use the coordinate position of each target text line as a start symbol to decode the target text lines in parallel in the same decoding sequence, and synchronously obtain the text in all the target text lines.
8. A computer readable medium having a computer program stored thereon, It is characterized in that When the program is executed by a processing device, the steps of the method described in any one of claims 1 to 6 are implemented.
9. An electronic device, It is characterized in that include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Automobile charging pile interface detection intelligent interaction system based on computer vision
CN113569849A