Image text recognition method and device
The first character position of complex or special arrangement sequential text is determined by the startup unit in the image text recognition model, which solves the problem of low accuracy in special arrangement sequential text recognition in existing OCR technology, and realizes accurate recognition of complex arrangement sequential text.
Patent Information
- Application Number
- CN202111676256.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The existing OCR technology is effective in identifying text arranged from left to right at a level, but the accuracy of text recognition in special arrangement orders is low, especially for texts with complex or special arrangement orders, which is difficult to determine the first text character, which affects the decoding process.
The image text recognition model is adopted, which is composed of an encoding unit, a startup unit and a decoding unit. The startup unit determines the area where the first text character is located in the image to be identified in the preset arrangement order, and realizes accurate recognition of the text in the complex arrangement order by learning the relevant information of the first text character.
It improves the recognition accuracy of texts with complex or special arrangement orders, and can obtain the character sequence of recognized texts more accurately.
Smart Images

Figure CN114332842B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and particularly to an image text recognition method and device. Background Art
[0002] OCR (Optical Character Recognition) is a technology for recognizing and analyzing images including text to obtain the text in the images. Using OCR technology, the text in the images can be extracted and recognized.
[0003] The characters in the image text are arranged in a certain order. Usually, the text in the image is arranged in the order from left to right horizontally. In addition, there is also some text in the images that is arranged in a relatively special order, such as staggered arrangement, vertical arrangement, etc.
[0004] Currently, OCR technology can recognize text arranged in the order from left to right horizontally, while the recognition accuracy for text with a special arrangement order is relatively low. Summary of the Invention
[0005] In view of this, embodiments of this application provide an image text recognition method and device, which can accurately recognize text with a relatively special and complex arrangement order in the image.
[0006] To solve the above problems, the technical solutions provided by the embodiments of this application are as follows:
[0007] In a first aspect, embodiments of this application provide an image text recognition method, and the method includes:
[0008] Input the image to be recognized into an image text recognition model. The image to be recognized includes at least two to-be-recognized text characters arranged in a preset arrangement order. The image text recognition model is composed of an encoding unit, a starting unit, and a decoding unit. The starting unit is used to determine the area where the first to-be-recognized text character in the to-be-recognized image is located according to the preset arrangement order;
[0009] Obtain the recognized text character sequence output by the image text recognition model.
[0010] In a second aspect, embodiments of this application provide an image text recognition device, and the device includes:
[0011] An input unit for inputting an image to be recognized into an image text recognition model, where the image to be recognized includes at least two text characters to be recognized arranged in a preset arrangement order, and the image text recognition model is composed of an encoding unit, a starting unit, and a decoding unit. The starting unit is used to determine the area where the first text character to be recognized in the preset arrangement order in the image to be recognized is located;
[0012] An acquisition unit for acquiring the recognized text character sequence output by the image text recognition model.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0014] One or more processors;
[0015] A storage device on which one or more programs are stored,
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any embodiment of the first aspect.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable medium, characterized in that a computer program is stored thereon, and when the program is executed by a processor, the method described in any embodiment of the first aspect is implemented.
[0018] Thus, the embodiments of the present application have the following beneficial effects:
[0019] An embodiment of the present application provides an image text recognition method and device, which input an image to be recognized into an image text recognition model and acquire the recognized text character sequence output by the image text recognition model. Among them, the image text recognition model is composed of an encoding unit, a starting unit, and a decoding unit. The starting unit can determine the area where the first character to be recognized in the preset arrangement order in the image to be recognized is located. The image text recognition model can learn the relevant information of the first text character in the preset arrangement order in the image to be recognized, and then can recognize the text characters to be recognized in the image to be recognized based on the first text character, and can accurately recognize the text characters in the image in the preset arrangement order to obtain a relatively accurate recognized text character sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a framework schematic diagram of an exemplary application scenario provided by an embodiment of the present application;
[0021] Figure 2 It is a flowchart of an image text recognition method provided by an embodiment of the present application;
[0022] Figure 3Flowchart of a method for training an image text recognition model provided by an embodiment of the present application;
[0023] Figure 4 A training image provided by an embodiment of the present application;
[0024] Figure 5 Schematic structural diagram of an encoding unit provided by an embodiment of the present application;
[0025] Figure 6 Schematic structural diagram of a decoding unit provided by an embodiment of the present application;
[0026] Figure 7 Schematic structural diagram of a starting unit provided by an embodiment of the present application;
[0027] Figure 8 Schematic structural diagram of an image text recognition device provided by an embodiment of the present application;
[0028] Figure 9 Schematic diagram of the basic structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0029] To make the above objects, features, and advantages of the present application more obvious and understandable, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0030] To facilitate understanding of the technical solution provided by the present application, the background technology related to the present application will be described first below.
[0031] Currently, the OCR technology based on deep learning is mainly divided into the CTC (Connectionist Temporal Classification) method represented by CRNN (Convolutional Recurrent Neural Network) and the Attention method represented by Transformer. The above two methods have strong recognition capabilities for horizontally arranged text, but the accuracy of the recognition results for text with a more complex or special arrangement order is not high. After studying the image text recognition process, it is found that for text characters with a complex or special arrangement order, it is difficult for the image text recognition model to determine the first text character, which affects the decoding process and thus affects the recognition results of the image text recognition model.
[0032] Based on this, an embodiment of the present application provides an image text recognition method and apparatus. The image to be recognized is input into an image text recognition model, and a recognized text character sequence output by the image text recognition model is obtained. Among them, the image to be recognized includes at least two to-be-recognized text characters arranged in a preset arrangement order. The image text recognition model is composed of an encoding unit, a starting unit, and a decoding unit. The starting unit can determine the area where the first to-be-recognized text character arranged in the preset arrangement order in the image to be recognized is located. The image text recognition model can learn the relevant information of the first text character in the preset arrangement order in the image to be recognized, and then can recognize the to-be-recognized text characters in the image to be recognized based on the first text character, and can accurately recognize the text characters arranged in the preset arrangement order in the image, so as to obtain a relatively accurate recognized text character sequence.
[0033] To facilitate understanding of the image text recognition method provided by the embodiment of the present application, the following is described in conjunction with Figure 1 the following scenario example shown. Refer to Figure 1 This figure is a framework schematic diagram of an exemplary application scenario provided by an embodiment of the present application.
[0034] In practical applications, the image 101 to be recognized includes five to-be-recognized text characters. The to-be-recognized text characters are "Roast", "Lacquer", "Training", "Class", and "Course". The to-be-recognized text characters are arranged in a relatively complex arrangement order. The image 101 to be recognized is input into the trained image text recognition model 102, and "Roast Lacquer Training Class" 103 output by the image text recognition model is obtained, that is, the recognized text character sequence.
[0035] Those skilled in the art can understand that Figure 1 the framework schematic diagram shown is only an example in which the implementation manner of the present application can be realized. The scope of application of the implementation manner of the present application is not limited by any aspect of this framework.
[0036] Based on the above description, the following will describe in detail the training method of the image text recognition model provided by the present application with reference to the accompanying drawings.
[0037] An embodiment of the present application provides an image text recognition method. Refer to Figure 2 shown in the figure, this figure is a flowchart of an image text recognition method provided by an embodiment of the present application. The image text recognition method includes S201 - S202:
[0038] S201: Input the image to be recognized into the image text recognition model.
[0039] The image to be recognized is an image for which text characters need to be recognized. The image to be recognized includes at least two text characters to be recognized arranged in a preset arrangement order. Among them, the preset arrangement order is determined based on the reading order of the text characters to be recognized. The preset arrangement order can be a relatively complex and special arrangement order such as vertical arrangement, diagonal arrangement, multi-line arrangement, etc. For example, see Figure 3 As shown, this figure is an image to be recognized provided by an embodiment of the present application. Among them, "burn", "cured meat", "training", and "class" are the text characters to be recognized included in the image to be recognized. The text characters to be recognized are arranged in a preset arrangement order where the first row is "burn" and "cured meat", the second row is "training", and finally "class".
[0040] Input the image to be recognized into the image text recognition model. The image text recognition model is composed of an encoding unit, a starting unit, and a decoding unit. The image text recognition model is trained by the training method of the image text recognition model.
[0041] In a possible implementation manner, an embodiment of the present application provides a specific implementation manner of the training method of the image text recognition model. Please refer to the following text.
[0042] Among them, the starting unit in the image text recognition model is used to determine the attention area of the first character. The attention area of the first character is the area of the first text character to be recognized arranged in the preset arrangement order in the image to be recognized. Taking Figure 3 as an example, the attention area of the first character is the area where "burn" is located.
[0043] S202: Obtain the recognized text character sequence output by the image text recognition model.
[0044] The recognized text character sequence is the recognition result obtained by the image text recognition model for recognizing the text characters to be recognized in the image to be recognized. The sorting sequence of the recognized text characters in the recognized text character sequence is determined by the image text recognition model.
[0045] Based on the relevant content of the above S201 - S202, it can be seen that the image text recognition model is composed of an encoding unit, a starting unit, and a decoding unit. The starting unit can determine the area where the first text character to be recognized in the preset arrangement order in the image to be recognized is located. The image text recognition model can learn the relevant information of the first text character in the preset arrangement order in the image to be recognized, and then can recognize the text characters to be recognized in the image to be recognized based on the first text character, so as to accurately recognize the text characters arranged in a relatively complex arrangement order and obtain a relatively accurate recognized text character sequence.
[0046] In a possible implementation manner, an embodiment of the present application provides a specific implementation manner of the training method of the image text recognition model.
[0047] See Figure 4 , which is a flowchart of a method for training an image text recognition model provided by an embodiment of the present application. As Figure 4 shown, the method may include S401 - S402:
[0048] S401: Input the training image into the image text recognition model to be trained, and obtain the first - character attention region generated by the start unit.
[0049] The training image is an image for training the image text recognition model to be trained. The training image includes at least two training text characters arranged in a preset arrangement order.
[0050] The image text recognition model to be trained may adopt a seq2seq model structure. The image text recognition model to be trained is composed of an encoding unit, a start unit, and a decoding unit. Inputting the training image into the image text recognition model to be trained can obtain the first - character attention region output by the start unit.
[0051] Among them, in the image text recognition model to be trained, the encoding unit is used to encode the input training image to obtain an encoded image feature vector.
[0052] See Figure 5 shown, which is a schematic structural diagram of an encoding unit provided by an embodiment of the present application. The encoding unit is composed of an encoder. Each encoder is composed of an embedding (Input Embedding) layer, a positional encoding (PositionalEncoding) layer, and n encoding layers. Among them, n is a positive integer. The encoding layer includes a multi - head attention (Multi - HeadAttention) layer, a residual connection and layer normalization (Residual Connection and Layer Normalization, Add&Norm) layer, and a feedforward layer.
[0053] Input the training image into the encoding unit to obtain the image feature vector Froi1 output by the encoding unit.
[0054] The start unit is used to determine the first - character attention region based on the image feature vector input by the encoding unit. The first - character attention region is calculated by the start unit. The first - character attention region is used to indicate the region of the first training text character arranged in the preset arrangement order in the training image determined by the start unit. The start unit determines the first training text character, which can enable the decoding unit to learn the relevant information of the first training text character in the preset arrangement order.
[0055] In a possible implementation manner, the embodiments of the present application provide a specific structure of a startup unit and a method for determining the first character attention area. Please refer to the following.
[0056] The decoding unit is used to decode based on the first feature vector output by the startup unit to obtain a text character sequence. The decoding unit may be composed of a decoder. Among them, the decoding unit may be composed of at least two decoders. Refer to Figure 6 As shown, this figure is a schematic structural diagram of a decoding unit provided by the embodiments of the present application. Each decoder is composed of an embedding (OutputEmbedding) layer, a positional encoding (Positional Encoding) layer, n decoders, a linear layer, and a classification (softmax) layer. Among them, n is a positive integer. The decoder includes a masked multi-head attention (Masked Multi-HeadAttention) layer, a multi-head attention (Multi-Head Attention) layer, a residual connection and layer normalization (ResidualConnection and Layer Normalization, Add&Norm) layer, and a feedforward layer.
[0057] Among them, y0 is the first feature vector output by the startup unit. The decoding unit uses y0 output by the startup unit as the input at the starting time step, and based on the image feature vectors y1,..., y t-1 , y t , y t+1 and so on to decode in sequence to obtain a text character sequence. Among them, t is a positive integer.
[0058] S402: Use the first character attention area and the first character text area identifier of the training image to train the text recognition model of the image to be trained, and obtain the text recognition model of the image.
[0059] The first character text area identifier of the training image is used to identify the text area where the first training text character arranged in a preset arrangement order is located among the training text characters. The first character text area identifier is used to indicate the text area where the more accurate first training text character is located.
[0060] Using the first character text area identifier of the training image as a supervision signal can make the area of the first training text character arranged in a preset arrangement order determined by the startup unit in the training image more accurate. In this way, the information related to the first training text character arranged in a preset arrangement order included in the feature vector output by the startup unit can be more accurate.
[0061] In a possible implementation manner, an embodiment of the present application provides a specific implementation manner for training the image text recognition model to be trained by using the first character attention region and the first character text region identifier of the training image. For details, please refer to the following text.
[0062] Based on the relevant content of S401 - S402 above, adding a start unit can enable the image text recognition model to be trained to pay more attention to the first training text character arranged in the preset arrangement order. By using the first character attention region and the first character text region identifier of the training image, the gap between the first training text character determined by the start unit and arranged in the preset arrangement order, and the first training text character arranged in the preset arrangement order in the training image can be measured. The image text recognition model trained in this way can learn the relevant information of the first training text character arranged in the preset arrangement order, and a more accurate text recognition result can be obtained based on the first training text character arranged in the preset arrangement order.
[0063] In a possible implementation manner, the start unit is used to determine the first character attention region and generate a first feature vector based on the input start symbol and the image feature vector.
[0064] Among them, the start symbol is a special symbol used to indicate the start of the text. The start symbol can specifically be an SOS (start of string) identification symbol. The image feature vector is generated by the encoding unit based on the training image. The first feature vector is used to be input into the decoding unit. The decoding unit can decode the first feature vector to obtain a text character sequence.
[0065] Furthermore, an embodiment of the present application provides a start unit. The start unit is composed of a decoder. Refer to Figure 7 As shown, this figure is a schematic structural diagram of a start unit provided by an embodiment of the present application. Among them, the start unit includes a Starter Embedding layer, a Positional Encoding layer, a decoder, a linear layer, and a softmax layer. Wherein, n is a positive integer. The decoder includes a Masked Multi - HeadAttention layer, a Multi - Head Attention layer, a Residual Connection and Layer Normalization (Add&Norm) layer, and a Feedforward layer.
[0066] The start embedding layer is used to encode the start symbol of the input to obtain the second feature vector e_s. The second feature vector e_s is consistent with the channel dimension of the image feature vector.
[0067] Input e_s and the image feature vector into the multi-head attention layer. In a possible implementation, the F output by the encoding unit can be flattened in advance to obtain F roi1 . Then input e_s and F roi2 into the multi-head attention layer. The multi-head attention layer can use e_s as the query, and F roi2 as the key and value. Utilize the first-character attention region to affect the attention part of the query and key, so that the start module can better learn the relevant information of the first training text character arranged in the preset arrangement order. roi2
[0068] In a possible implementation, the first-character attention region can be determined based on the attention heat map. The attention heat map can be calculated by the attention mechanism. The first-character text region identifier of the training image can be represented by the first-character text region distribution map. Specifically, the first-character text region distribution map can be determined based on the boundary coordinates of the text region where the first text character arranged in the preset arrangement order is located in the training image. The first-character text region distribution map can be, for example, a Gaussian distribution in the boundary coordinates.
[0069] Using the difference between the first-character attention region and the first-character text region identifier of the training image can measure the accuracy of the first text character determined by the start unit arranged in the preset arrangement order. In a possible implementation, the embodiments of the present application provide a specific implementation method for training the to-be-trained image text recognition model using the first-character attention region and the first-character text region identifier of the training image, including:
[0070] Calculate the loss function using the first-character attention region and the first-character text region identifier of the training image;
[0071] Adjust the model parameters of the start model using the loss function.
[0072] Specifically, when the first-character attention region is determined based on the attention heat map and the first-character text region identifier of the training image can be represented by the first-character text region distribution map, the corresponding loss function can be a loss function for measuring the distribution distance such as the Sliced Wasserstein distance.
[0073] Based on the image text recognition method provided by the above method embodiment, the embodiments of the present application also provide an image text recognition device, which will be described below with reference to the accompanying drawings.
[0074] See Figure 8 as shown, this figure is a schematic structural diagram of an image text recognition device provided by an embodiment of the present application. As Figure 8 shown, the image text recognition device includes:
[0075] An input unit 801, configured to input an image to be recognized into an image text recognition model. The image to be recognized includes at least two to-be-recognized text characters arranged in a preset arrangement order. The image text recognition model is composed of an encoding unit, a start unit, and a decoding unit. The start unit is configured to determine a region where the first to-be-recognized text character in the preset arrangement order is located in the image to be recognized;
[0076] An acquisition unit 802, configured to acquire a recognized text character sequence output by the image text recognition model.
[0077] In a possible implementation manner, the image text recognition model is trained by the following method:
[0078] Input a training image into a to-be-trained image text recognition model to obtain a first character attention region determined by the start unit. The training image includes at least two training text characters arranged in a preset arrangement order. The to-be-trained image text recognition model is composed of an encoding unit, the start unit, and a decoding unit. The first character attention region is the region of the first training text character in the preset arrangement order determined by the start unit in the training image;
[0079] Use the first character attention region and the first character text region identifier of the training image to train the to-be-trained image text recognition model to obtain an image text recognition model. The first character text region is used to identify the text region where the first training text character arranged in the preset arrangement order is located among the training text characters.
[0080] In a possible implementation manner, the start unit is configured to determine a first character attention region and generate a first feature vector based on an input start symbol and an image feature vector. The image feature vector is generated by the encoding unit based on the training image. The first feature vector is used to be input into the decoding unit for decoding.
[0081] In a possible implementation manner, the start unit is composed of a decoder. The decoder includes a multi-head attention layer. The start embedding layer is configured to encode the input start symbol to obtain a second feature vector. The multi-head attention layer is configured to determine a first character attention region based on the input second feature vector and the image feature vector.
[0082] In a possible implementation, the first-character attention region is determined based on an attention heat map, and the first-character text region identifier of the training image is represented by a first-character text region distribution map.
[0083] In a possible implementation, training the text recognition model of the image to be trained by using the first-character attention region and the first-character text region identifier of the training image includes:
[0084] Calculating a loss function by using the first-character attention region and the first-character text region identifier of the training image;
[0085] Adjusting the model parameters of the starting model by using the loss function.
[0086] In a possible implementation, the loss function is the Sliced Wasserstein distance.
[0087] Based on the image text recognition method provided in the above method embodiments, the present application further provides an electronic device, including: one or more processors; a storage device, on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the image text recognition method as described in any of the above embodiments.
[0088] Reference is made below to Figure 9 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing the embodiments of the present application. The terminal device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (portable android devices, tablet computers), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs (televisions), desktop computers, etc. Figure 9 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0089] As Figure 9As shown, the electronic device 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0090] Generally, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 9 an electronic device 900 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively.
[0091] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above functions defined in the method of the embodiment of the present application are executed.
[0092] The electronic device provided by the embodiment of the present application and the image text recognition method provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment may be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0093] Based on an image text recognition method provided by the above method embodiment, an embodiment of the present application provides a computer storage medium, on which a computer program is stored, wherein the program, when executed by a processor, implements the image text recognition method as described in any of the above embodiments.
[0094] It should be noted that the computer-readable medium described above in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0095] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0096] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0097] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to execute the above image text recognition method.
[0098] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0100] The units involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the unit / module does not constitute a limitation on the unit itself in some cases. For example, the voice data acquisition module can also be described as the "data acquisition module".
[0101] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0102] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0103] According to one or more embodiments of the present application, [Example 1] provides an image text recognition method, the method comprising:
[0104] Inputting an image to be recognized into an image text recognition model, the image to be recognized including at least two to-be-recognized text characters arranged in a preset arrangement order, the image text recognition model being composed of an encoding unit, a starting unit, and a decoding unit, the starting unit being configured to determine a region where a first to-be-recognized text character in the preset arrangement order is located in the image to be recognized;
[0105] Obtaining a recognized text character sequence output by the image text recognition model.
[0106] According to one or more embodiments of the present application, [Example 2] provides an image text recognition method, and the image text recognition model is trained by the following method:
[0107] Inputting a training image into a to-be-trained image text recognition model to obtain a first character attention region determined by the starting unit, the training image including at least two training text characters arranged in the preset arrangement order, the to-be-trained image text recognition model being composed of the encoding unit, the starting unit, and the decoding unit, the first character attention region being a region in the training image where a first training text character in the preset arrangement order determined by the starting unit is located;
[0108] Training the to-be-trained image text recognition model by using the first character attention region and a first character text region identifier of the training image, to obtain an image text recognition model, the first character text region being used to identify a text region where a first training text character arranged in the preset arrangement order among the training text characters is located.
[0109] According to one or more embodiments of the present application, [Example Three] provides an image text recognition method. The startup unit is used to determine the first character attention region and generate a first feature vector based on the input start symbol and the image feature vector. The image feature vector is generated by the encoding unit based on the training image, and the first feature vector is used to be input into the decoding unit for decoding.
[0110] According to one or more embodiments of the present application, [Example Four] provides an image text recognition method. The startup unit consists of a decoder, and the decoder includes a multi-head attention layer. The startup embedding layer is used to encode the input start symbol to obtain a second feature vector, and the multi-head attention layer is used to determine the first character attention region based on the input second feature vector and the image feature vector.
[0111] According to one or more embodiments of the present application, [Example Five] provides an image text recognition method. The first character attention region is determined based on an attention heat map, and the first character text region identification of the training image is represented by a first character text region distribution map.
[0112] According to one or more embodiments of the present application, [Example Six] provides an image text recognition method. Training the to-be-trained image text recognition model by using the first character attention region and the first character text region identification of the training image includes:
[0113] Calculating a loss function by using the first character attention region and the first character text region identification of the training image;
[0114] Adjusting the model parameters of the startup model by using the loss function.
[0115] According to one or more embodiments of the present application, [Example Seven] provides an image text recognition method. The loss function is the Sliced Wasserstein distance.
[0116] According to one or more embodiments of the present application, [Example Eight] provides an image text recognition device. The device includes:
[0117] An input unit, configured to input an image to be recognized into an image text recognition model. The image to be recognized includes at least two to-be-recognized text characters arranged in a preset arrangement order. The image text recognition model consists of an encoding unit, a startup unit, and a decoding unit. The startup unit is used to determine the region where the first to-be-recognized text character in the image to be recognized is located according to the preset arrangement order;
[0118] An acquisition unit, configured to acquire the recognized text character sequence output by the image text recognition model.
[0119] According to one or more embodiments of the present application, [Example Nine] provides an image text recognition device, and the image text recognition model is trained by the following method:
[0120] Input the training image into the image text recognition model to be trained, and obtain the first character attention region determined by the starting unit. The training image includes at least two training text characters arranged in a preset arrangement order. The image text recognition model to be trained is composed of an encoding unit, the starting unit, and a decoding unit. The first character attention region is the region of the first training text character in the preset arrangement order determined by the starting unit in the training image;
[0121] Use the first character attention region and the first character text region identifier of the training image to train the image text recognition model to be trained, and obtain the image text recognition model. The first character text region is used to identify the text region where the first training text character arranged in the preset arrangement order is located.
[0122] According to one or more embodiments of the present application, [Example Ten] provides an image text recognition device. The starting unit is used to determine the first character attention region and generate a first feature vector based on the input starting symbol and image feature vector. The image feature vector is generated by the encoding unit based on the training image, and the first feature vector is used to input the decoding unit for decoding.
[0123] According to one or more embodiments of the present application, [Example Eleven] provides an image text recognition device. The starting unit is composed of a decoder. The decoder includes a multi-head attention layer. The starting embedding layer is used to encode the input starting symbol to obtain a second feature vector, and the multi-head attention layer is used to determine the first character attention region based on the input second feature vector and the image feature vector.
[0124] According to one or more embodiments of the present application, [Example Twelve] provides an image text recognition device. The first character attention region is determined based on an attention heat map, and the first character text region identifier of the training image is represented by a first character text region distribution map.
[0125] According to one or more embodiments of the present application, [Example Thirteen] provides an image text recognition device. The step of using the first character attention region and the first character text region identifier of the training image to train the image text recognition model to be trained includes:
[0126] Calculate a loss function using the first character attention region and the first character text region identifier of the training image;
[0127] Adjust the model parameters of the startup model by using the loss function.
[0128] According to one or more embodiments of the present application, [Example XIV] provides an image text recognition device, and the loss function is the sliced Wasserstein distance.
[0129] According to one or more embodiments of the present application, [Example XV] provides an electronic device, including:
[0130] One or more processors;
[0131] A storage device storing one or more programs thereon,
[0132] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of [Example I] to [Example VII].
[0133] According to one or more embodiments of the present application, [Example XVI] provides a computer-readable medium storing a computer program, wherein when the program is executed by a processor, the method described in any one of [Example I] to [Example VII] is implemented.
[0134] It should be noted that the various embodiments in this specification are described in a progressive manner. The key point of each embodiment is the difference from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0135] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0136] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0137] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in a software module executed by a processor, or in a combination thereof. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium well-known in the art.
[0138] The foregoing description of the disclosed embodiments enables those skilled in the art to make or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image text recognition method, characterized in that, The method includes: Inputting an image to be recognized into an image text recognition model, where the image to be recognized includes at least two text characters to be recognized arranged in a preset arrangement order. The image text recognition model is composed of an encoding unit, a starting unit, and a decoding unit. The starting unit is used to determine the region where the first text character to be recognized in the preset arrangement order in the image to be recognized is located. The starting unit includes a starting embedding layer, a position encoding layer, a decoder, a linear layer, and a classification layer. The decoder includes a masked multi-head attention layer, a multi-head attention layer, a residual connection, a layer normalization layer, and a feed-forward layer; Obtaining a recognized text character sequence output by the image text recognition model; During the training process of the image text recognition model, the starting unit is used to determine a first character attention region and generate a first feature vector based on an input starting symbol and an image feature vector. The image feature vector is generated by the encoding unit based on a training image. The first feature vector is used to be input into the decoding unit for decoding. The training image is used to train the image text recognition model to be trained to obtain the image text recognition model.
2. The method according to claim 1, wherein The image text recognition model is obtained by training with the following method: Inputting a training image into the image text recognition model to be trained to obtain the first character attention region determined by the starting unit. The training image includes at least two training text characters arranged in the preset arrangement order. The image text recognition model to be trained is composed of the encoding unit, the starting unit, and the decoding unit. The first character attention region is the region of the first training text character in the preset arrangement order determined by the starting unit in the training image; Using the first character attention region and the first character text region identifier of the training image to train the image text recognition model to be trained to obtain an image text recognition model. The first character text region is used to identify the text region where the first training text character arranged in the preset arrangement order in the training text characters is located.
3. The method according to claim 1, wherein The starting embedding layer is used to encode the input starting symbol to obtain a second feature vector. The multi-head attention layer is used to determine the first character attention region based on the input second feature vector and the image feature vector.
4. The method according to claim 3, characterized in that, The first character attention region is determined based on an attention heat map. The first character text region identifier of the training image is represented by a first character text region distribution map.
5. The method according to claim 2, wherein The using the first character attention region and the first character text region identifier of the training image to train the image text recognition model to be trained includes: Calculating a loss function using the first character attention region and the first character text region identifier of the training image; Adjusting the model parameters of the starting unit using the loss function.
6. The method according to claim 5, wherein The loss function is the Sliced Wasserstein distance.
7. An image text recognition device, characterized in that, The device includes: An input unit for inputting an image to be recognized into an image text recognition model, where the image to be recognized includes at least two text characters to be recognized arranged in a preset arrangement order, and the image text recognition model is composed of an encoding unit, a starting unit, and a decoding unit. The starting unit is used to determine the region where the first text character to be recognized arranged in the preset arrangement order is located in the image to be recognized. The starting unit includes a starting embedding layer, a position encoding layer, a decoder, a linear layer, and a classification layer. The decoder includes a masked multi-head attention layer, a multi-head attention layer, a residual connection, a layer normalization layer, and a feed-forward layer; An acquisition unit for acquiring the recognized text character sequence output by the image text recognition model; During the training process of the image text recognition model, the starting unit is used to determine the first character attention region and generate a first feature vector based on the input starting symbol and image feature vector. The image feature vector is generated by the encoding unit based on the training image. The first feature vector is used to input the decoding unit for decoding. The training image is used to train the image text recognition model to be trained to obtain the image text recognition model.
8. An electronic device, characterized in that, Comprising: One or more processors; A storage device on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable medium, characterized in that, On which a computer program is stored, wherein the program, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Text recognition model training method and device and text recognition method and device
CN113111871A
Character recognition network model training method, character recognition method, apparatuses, terminal, and computer storage medium therefor
WO2021115159A1