Method, device, readable medium and electronic device for text recognition

By fusing semantic and sequence features into the text recognition model, the problem of low accuracy in recognizing complex text lines in existing models is solved, achieving higher recognition accuracy.

CN114495081BActive Publication Date: 2026-03-27BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-12
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing text recognition models based on CTC and sequence decoding cannot accurately recognize some characters when processing complex text lines of images, resulting in low recognition accuracy.

Method used

By fusing the semantic and sequence features of text line images, and using a pre-trained text recognition model to extract and fuse features, more complete semantic fusion features are obtained to improve recognition accuracy.

Benefits of technology

By fusing semantic and sequence features, the recognition accuracy of text line images was improved, the features were made more complete, and the recognition effect was enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495081B_ABST
    Figure CN114495081B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a text recognition method, device, readable medium and electronic equipment, and relates to the technical field of computers, comprising: obtaining a text line picture to be recognized; inputting the text line picture into a pre-trained text recognition model to obtain text in the text line picture output by the text recognition model; wherein the text recognition model is obtained by training a preset training model according to a first target character, the first target character is a character obtained by character conversion on a first semantic fusion feature, the first semantic fusion feature is a feature obtained by fusion processing of a first target sequence feature in a sample picture used for training and a first semantic feature of the sample picture, and the first semantic feature is a feature obtained according to the first target sequence feature. In this way, the text line picture feature extraction can be more complete, thereby improving the accuracy of text image recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to a method, apparatus, readable medium, and electronic device for text recognition. Background Technology

[0002] With the widespread application of text recognition technology, the requirements for the accuracy of text line image recognition are becoming increasingly higher, demanding the ability to accurately identify each character in the text line image. Related technologies utilize text recognition models based on CTC (Connectionist Temporal Classification) or sequence decoding to identify the text content in text line images.

[0003] However, both of the above methods only utilize the features of the text line image itself. For some more complex text line images, when the above text recognition models are used to recognize the text line image, some characters may not be recognized correctly, resulting in a low accuracy rate for text line image recognition. Summary of the Invention

[0004] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] In a first aspect, this disclosure provides a text recognition method, comprising: acquiring a text line image to be recognized; inputting the text line image into a pre-trained text recognition model to obtain the text in the text line image output by the text recognition model; wherein the text recognition model is obtained by training a preset training model based on a first target character, the first target character being a character obtained by character conversion of a first semantic fusion feature, the first semantic fusion feature being a feature obtained by fusing a first target sequence feature in a training sample image with a first semantic feature of the sample image, and the first semantic feature being a feature obtained based on the first target sequence feature.

[0006] Secondly, this disclosure provides a text recognition apparatus, comprising: an acquisition module for acquiring a text line image to be recognized; and a recognition module for inputting the text line image into a pre-trained text recognition model to obtain the text in the text line image output by the text recognition model; wherein the text recognition model is obtained by training a preset training model based on a first target character, the first target character being a character obtained by converting a first semantic fusion feature into a character, the first semantic fusion feature being a feature obtained by fusing a first target sequence feature in a training sample image with a first semantic feature of the sample image, and the first semantic feature being a feature obtained based on the first target sequence feature.

[0007] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the text recognition method described in the first aspect.

[0008] Fourthly, this disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; and a processing device for executing the computer program in the storage device to implement the steps of the text recognition method described in the first aspect.

[0009] The above technical solution involves acquiring an image of a text line to be recognized; inputting this image into a pre-trained text recognition model to obtain the text output by the model; wherein the text recognition model is trained on a pre-trained model based on a first target character, which is obtained by converting a first semantic fusion feature into a character; the first semantic fusion feature is obtained by fusing the first target sequence feature in the training sample image with the first semantic feature of the sample image; and the first semantic feature is obtained based on the first target sequence feature. In other words, when recognizing a text line image, this disclosure fuses the semantic and sequence features of the text line image using a text recognition model to obtain the semantic fusion feature of the text line image, and then obtains the text content of the text line image based on the semantic fusion feature. This makes the features of the text line image more complete, thereby improving the accuracy of text image recognition.

[0010] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:

[0012] Figure 1 This is a flowchart of a text recognition method provided according to an exemplary embodiment;

[0013] Figure 2 This is a structural diagram of a text recognition model provided according to an exemplary embodiment;

[0014] Figure 3 This is a flowchart of a training method for a text recognition model according to an exemplary embodiment;

[0015] Figure 4 This is a flowchart of another method for training a text recognition model according to an exemplary embodiment;

[0016] Figure 5 This is a structural diagram of a text recognition device according to an exemplary embodiment;

[0017] Figure 6 This is a block diagram of an electronic device provided according to an exemplary embodiment. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0020] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0021] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0022] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0023] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0024] First, the application scenarios of this disclosure are explained. This disclosure can be applied to the scenario of recognizing text lines in images. With the widespread application of text recognition technology, people have increasingly higher requirements for the accuracy of text line image recognition, needing to be able to accurately identify each character in the text line image. In related technologies, OCR (Optical Character Recognition) models are often used to recognize text in text line images. However, whether it is a CTC-based or sequence decoding-based OCR model, it only uses the features of the text line image itself. For some more complex text line images, such as advertisements and movie posters, the feature information in the text line image is relatively rich. When recognizing such text line images using the above-mentioned text recognition models, some characters may not be recognized correctly, resulting in a lower accuracy rate for text line image recognition.

[0025] To address the aforementioned technical problems, this disclosure provides a method, apparatus, readable medium, and electronic device for text recognition. When recognizing a text line image, a text recognition model fuses the semantic and sequence features of the text line image to obtain its semantic fusion features. The text content of the text line image is then obtained based on these semantic fusion features. This makes the features of the text line image more complete, thereby improving the accuracy of text image recognition.

[0026] The present disclosure will now be described in conjunction with specific embodiments.

[0027] Figure 1 This is a flowchart of a text recognition method according to an exemplary embodiment, such as... Figure 1 As shown, the method may include the following steps:

[0028] In step S101, the image of the text line to be identified is obtained.

[0029] In step S102, the text line image is input into a pre-trained text recognition model to obtain the text in the text line image output by the text recognition model.

[0030] The text recognition model is obtained by training a preset training model based on a first target character. The first target character is a character obtained by converting the first semantic fusion feature into a character. The first semantic fusion feature is a feature obtained by fusing the first target sequence feature in the sample image used for training with the first semantic feature of the sample image. The first semantic feature is a feature obtained based on the first target sequence feature.

[0031] In some embodiments, such as Figure 2 As shown, the text recognition model may include a feature extraction model, at least one semantic fusion model, a decoding model, and a first fully connected layer. In the case where there are multiple semantic fusion models, the multiple semantic fusion models are coupled in series in sequence.

[0032] In one possible implementation, the text recognition model includes a semantic fusion model. In this case, the output of the feature extraction model can be used as the input of the semantic fusion model, and the output of the semantic fusion model can be used as the input of the decoding model.

[0033] In another possible implementation, the text recognition model comprises multiple semantic fusion models, which are sequentially coupled in series. The output of the feature extraction model serves as the input to the first semantic fusion model, and so on, until the last semantic fusion model. The output of the last semantic fusion model is then used as the input to the decoding model. By sequentially extracting features from the text line image through multiple semantic fusion modules, the semantic information of the characters in the text line image can be extracted more accurately.

[0034] This feature extraction model is used to obtain features of a second sequence to be encoded from an input image of a text line.

[0035] The feature extraction model can be, for example, based on the CNN (Convolutional Neural Networks) framework and trained using existing model training methods, which will not be elaborated here. The feature extraction model can also be, for example, the image2vector model.

[0036] For example, the feature extraction model obtains a second sequence feature to be encoded from the input text line image. In order to improve the efficiency of data processing, the second sequence feature to be encoded can be a fixed-dimensional sequence feature, such as 512.

[0037] The semantic fusion model is used to encode the second target sequence features output by the feature extraction model to obtain the second target sequence features, and to obtain the second semantic features of the text line image based on the second target sequence features. The second target sequence features and the second semantic features are then fused to obtain the second semantic fusion feature.

[0038] For example, such as Figure 2 As shown, the second target sequence feature can be added to the second semantic feature to obtain the second semantic fusion feature.

[0039] In some embodiments, the semantic fusion model may include: an encoding sub-model, a second fully connected layer, and a language sub-model;

[0040] This encoding sub-model is used to encode the second target sequence features output by the feature extraction model to obtain the second target sequence features.

[0041] The encoding sub-model can be any feasible encoder algorithm under the encoder model architecture, and this disclosure does not impose any restrictions on it. The encoding sub-model is used to encode the second feature to be encoded into a vector feature (second target sequence feature), for example, using the attention mechanism to extract a 512-dimensional vector sequence feature.

[0042] For example, the second target sequence feature output by the feature extraction model is encoded through this encoding sub-model, that is, the second target sequence feature is transformed into a fixed-dimensional vector feature (second target sequence feature).

[0043] The second fully connected layer is used to convert the second target sequence features output by the encoding submodel into the second character to be processed.

[0044] The language sub-model is used to extract the second semantic feature from the second character to be processed output by the second fully connected layer.

[0045] This language sub-model can be, for example, based on the BERT (Bidirectional Encoder Representations from Transformers) framework and trained using existing model training methods, which will not be elaborated here. Using this language sub-model, semantic features are extracted from the second character to be processed, resulting in the second semantic feature.

[0046] Furthermore, considering that the second character to be processed may contain some useless information, which will cause interference and affect the efficiency of model recognition, in order to further improve the accuracy of feature extraction and the efficiency of data processing, the confidence level of the second character to be processed can be obtained. Second characters with a confidence level greater than or equal to a preset confidence threshold are input into the language sub-model. In other words, second characters with a confidence level less than the preset confidence threshold are removed, thus removing useless information (i.e., noise) from the second character. This effectively improves both the accuracy of feature extraction and the efficiency of data processing.

[0047] A decoding model is used to decode the second semantic fusion feature output by the semantic fusion model;

[0048] The first fully connected layer is used to convert the decoded second semantic fusion feature output by the decoding model into the second target character.

[0049] Using the above method, when recognizing text line images, the semantic and sequence features of the text line image are fused through a text recognition model to obtain the semantic fusion features of the text line image. The text content of the text line image is then obtained based on these semantic fusion features. This makes the features of the text line image more complete, thereby improving the accuracy of text image recognition.

[0050] The training method of the above text recognition model will be explained below, such as... Figure 3 As shown, this text recognition model can be trained using the following steps:

[0051] In step S301, multiple sample images for training are acquired.

[0052] It should be noted that each of the multiple sample images obtained is an image containing text.

[0053] In step S302, the first target sequence features are obtained from the sample image.

[0054] Understandably, the first sequence feature to be encoded can be extracted from the sample image, and the first sequence feature to be encoded can be encoded to obtain the first target sequence feature.

[0055] For example, the image2vector model described above can be used to extract the first sequence feature to be encoded from a sample image with a fixed dimension, such as 512. This first sequence feature can then be encoded using an encoding sub-model to obtain the first target sequence feature. For instance, the attention mechanism of the encoder model can be used to extract the 512-dimensional vector sequence feature from the first sequence feature to be encoded, thus obtaining the first target sequence feature.

[0056] In step S303, the first semantic feature of the sample image is obtained based on the first target sequence feature.

[0057] For example, the first semantic features of a sample image can be obtained through the semantic fusion model described above.

[0058] In some embodiments, such as Figure 4 As shown, obtaining the first semantic feature of the sample image based on the first target sequence feature may include the following steps:

[0059] In step S3031, the first target sequence feature is converted into the first character to be processed.

[0060] For example, the first target sequence features can be converted into the first character to be processed through the second fully connected layer mentioned above, so that the language sub-model can extract the semantic features of the sample image.

[0061] In step S3032, the first semantic feature is extracted from the first character to be processed.

[0062] For example, the first semantic feature can be extracted from the first character to be processed using the BERT model described above, so as to obtain semantic features containing stronger semantics in the sample image.

[0063] In step S304, the first target sequence feature is fused with the first semantic feature to obtain the first semantic fusion feature.

[0064] In this step, the first target sequence feature is fused with the first semantic feature to obtain the semantic fusion feature that best reflects the semantic information in the sample image, thereby improving the accuracy of the model recognition results.

[0065] For example, the first target sequence feature and the first semantic feature can be added together to obtain the first semantic fusion feature.

[0066] In one possible implementation, the text recognition model includes a semantic fusion module that uses the first semantic fusion feature obtained by the semantic fusion model as input to the decoding model.

[0067] In another possible implementation, the text recognition model includes multiple semantic fusion models, where the first semantic fusion feature can be used as input to the next semantic fusion model. This process continues until the last semantic fusion model, where the output of the last semantic fusion model is used as input to the decoding model. By sequentially extracting features from the text image using multiple semantic fusion modules, the semantic information of the characters in the text image can be extracted more accurately.

[0068] In step S305, the first semantic fusion feature is converted into the first target character.

[0069] Understandably, the first semantic fusion feature can be decoded and converted into the corresponding first target character.

[0070] For example, the first semantic fusion feature can be decoded using a decoder model, and the decoded first semantic fusion feature can be converted into the corresponding first target character through the first fully connected layer.

[0071] In step S306, the preset training model is trained based on the first target character to obtain the text recognition model.

[0072] The first semantic fusion feature is obtained by fusing the first target sequence feature with the first semantic feature, and the first semantic fusion feature is converted into the first target character. The preset training model is then trained based on the first target character to obtain the text recognition model.

[0073] Furthermore, considering that the first character to be processed may contain some useless information, which will interfere with the model's recognition, in order to further improve the accuracy of feature extraction and the efficiency of data processing, the confidence level of the first character to be processed can be obtained. Characters with a confidence level greater than or equal to a preset confidence threshold are input into the language sub-model. In other words, characters with a confidence level less than the preset confidence threshold are removed, thus removing useless information (i.e., noise) from the first character to be processed. This effectively improves the accuracy of feature extraction and the efficiency of data processing. The language sub-model then extracts the first semantic feature from the first character to be processed with a confidence level greater than or equal to the preset confidence threshold.

[0074] Using the above method, when recognizing text line images, the semantic and sequence features of the text line image are fused through a text recognition model to obtain the semantic fusion features of the text line image. The text content of the text line image is then obtained based on these semantic fusion features. This makes the features of the text line image more complete, thereby improving the accuracy of text image recognition.

[0075] Figure 5 This is a structural diagram of a text recognition device according to an exemplary embodiment, such as... Figure 5 As shown, the device 500 includes:

[0076] The acquisition module 501 is used to acquire images of the text lines to be recognized;

[0077] The recognition module 502 is used to input the text line image into a pre-trained text recognition model to obtain the text in the text line image output by the text recognition model;

[0078] The text recognition model is obtained by training a preset training model based on a first target character. The first target character is a character obtained by converting the first semantic fusion feature into a character. The first semantic fusion feature is a feature obtained by fusing the first target sequence feature in the sample image used for training with the first semantic feature of the sample image. The first semantic feature is a feature obtained based on the first target sequence feature.

[0079] Optionally, the text recognition model is trained in the following way:

[0080] Obtain multiple sample images for training;

[0081] Extract the first target sequence features from the sample image;

[0082] The first semantic feature of the sample image is obtained based on the first target sequence feature;

[0083] The first target sequence feature is fused with the first semantic feature to obtain the first semantic fusion feature;

[0084] The first semantic fusion feature is converted into the first target character;

[0085] The text recognition model is obtained by training the preset training model based on the first target character.

[0086] Optionally, obtaining the first semantic feature of the sample image based on the first target sequence feature includes:

[0087] Convert the features of the first target sequence into the first character to be processed;

[0088] Extract the first semantic feature from the first character to be processed.

[0089] Optionally, the method further includes:

[0090] Obtain the confidence level of the first character to be processed;

[0091] The extraction of the semantic feature from the first character to be processed includes:

[0092] The first semantic feature is extracted from the first character to be processed whose confidence level is greater than or equal to the preset confidence threshold.

[0093] Optionally, the first target sequence feature is fused with the first semantic feature to obtain the first semantic fusion feature, which includes:

[0094] The first target sequence feature and the first semantic feature are added together to obtain the first semantic fusion feature.

[0095] Optionally, obtaining the first target sequence features from the sample image includes:

[0096] The first target sequence feature is obtained by extracting the first sequence feature to be encoded from the sample image and encoding the first sequence feature.

[0097] Optionally, converting the first semantic fusion feature into the first target character includes:

[0098] The first semantic fusion feature is decoded, and the decoded first semantic fusion feature is converted into the corresponding first target character.

[0099] Optionally, the text recognition model includes a feature extraction model, at least one semantic fusion model, a decoding model, and a first fully connected layer, wherein, when there are multiple semantic fusion models, the multiple semantic fusion models are coupled in series sequentially.

[0100] This feature extraction model is used to obtain the features of the second sequence to be encoded from the input image of the text line;

[0101] The semantic fusion model is used to encode the second sequence features output by the feature extraction model to obtain the second target sequence features, and to obtain the second semantic features of the text line image based on the second target sequence features. The second target sequence features and the second semantic features are then fused to obtain the second semantic fusion feature.

[0102] A decoding model is used to decode the second semantic fusion feature output by the semantic fusion model;

[0103] The first fully connected layer is used to convert the decoded second semantic fusion feature output by the decoding model into the second target character.

[0104] Optionally, the semantic fusion model includes: an encoding sub-model, a second fully connected layer, and a language sub-model;

[0105] This encoding sub-model is used to encode the second target sequence feature output by the feature extraction model to obtain the second target sequence feature.

[0106] The second fully connected layer is used to convert the second target sequence features output by the encoding submodel into the second character to be processed;

[0107] The language sub-model is used to extract the second semantic feature from the second character to be processed output by the second fully connected layer.

[0108] Using the aforementioned device, when recognizing text line images, the semantic and sequence features of the text line image are fused through a text recognition model to obtain the semantic fusion features of the text line image. The text content of the text line image is then obtained based on these semantic fusion features. This makes the features of the text line image more complete, thereby improving the accuracy of text image recognition.

[0109] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0110] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 600 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0111] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0112] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0113] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0114] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0115] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0116] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0117] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire a text line image to be recognized; input the text line image into a pre-trained text recognition model to obtain the text in the text line image output by the text recognition model; wherein the text recognition model is obtained by training a preset training model based on a first target character, the first target character being a character obtained by character conversion of a first semantic fusion feature, the first semantic fusion feature being a feature obtained by fusing a first target sequence feature in a training sample image with a first semantic feature of the sample image, and the first semantic feature being a feature obtained based on the first target sequence feature.

[0118] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0120] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a module does not necessarily limit the module itself; for example, an acquisition module can also be described as "a module for acquiring images of text lines to be recognized".

[0121] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0122] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0123] According to one or more embodiments of this disclosure, Example 1 provides a text recognition method, the method comprising: acquiring a text line image to be recognized; inputting the text line image into a pre-trained text recognition model to obtain text in the text line image output by the text recognition model; wherein the text recognition model is obtained by training a preset training model based on a first target character, the first target character being a character obtained by character conversion of a first semantic fusion feature, the first semantic fusion feature being a feature obtained by fusing a first target sequence feature in a sample image used for training with a first semantic feature of the sample image, and the first semantic feature being a feature obtained based on the first target sequence feature.

[0124] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the text recognition model is trained by: acquiring a plurality of sample images for training; acquiring a first target sequence feature from the sample images; acquiring a first semantic feature of the sample images based on the first target sequence feature; fusing the first target sequence feature with the first semantic feature to obtain a first semantic fusion feature; converting the first semantic fusion feature into a first target character; and training a preset training model based on the first target character to obtain the text recognition model.

[0125] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein obtaining the first semantic feature of the sample image based on the first target sequence feature includes: converting the first target sequence feature into a first character to be processed; and extracting the first semantic feature from the first character to be processed.

[0126] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 3, the method further comprising: obtaining the confidence level of the first character to be processed; the step of extracting the semantic feature from the first character to be processed comprises: extracting the first semantic feature from the first characters to be processed whose confidence level is greater than or equal to a preset confidence threshold.

[0127] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 2, wherein fusing the first target sequence feature with the first semantic feature to obtain the first semantic fusion feature includes: adding the first target sequence feature and the first semantic feature to obtain the first semantic fusion feature.

[0128] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 2, wherein obtaining the first target sequence feature from the sample image includes: extracting the first sequence feature to be encoded from the sample image and encoding the first sequence feature to be encoded to obtain the first target sequence feature.

[0129] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 6, wherein converting the first semantic fusion feature into the first target character includes: decoding the first semantic fusion feature and converting the decoded first semantic fusion feature into the corresponding first target character.

[0130] According to one or more embodiments of this disclosure, Example 8 provides a method of any one of Examples 1 to 7, wherein the text recognition model includes a feature extraction model, at least one semantic fusion model, a decoding model, and a first fully connected layer, wherein, when there are multiple semantic fusion models, the multiple semantic fusion models are coupled in series sequentially; the feature extraction model is used to obtain a second sequence feature to be encoded from the input text line image; the semantic fusion model is used to encode the second sequence feature to be encoded output by the feature extraction model to obtain a second target sequence feature, and to obtain a second semantic feature of the text line image based on the second target sequence feature, and to fuse the second target sequence feature with the second semantic feature to obtain a second semantic fusion feature; the decoding model is used to decode the second semantic fusion feature output by the semantic fusion model; the first fully connected layer is used to convert the decoded second semantic fusion feature output by the decoding model into a second target character.

[0131] According to one or more embodiments of this disclosure, Example 9 provides the method of Example 8, wherein the semantic fusion model includes: an encoding sub-model, a second fully connected layer, and a language sub-model; the encoding sub-model is used to encode the second target sequence features output by the feature extraction model to obtain second target sequence features; the second fully connected layer is used to convert the second target sequence features output by the encoding sub-model into second characters to be processed; and the language sub-model is used to extract the second semantic features from the second characters to be processed output by the second fully connected layer.

[0132] According to one or more embodiments of this disclosure, Example 10 provides a text recognition apparatus, the apparatus comprising: an acquisition module for acquiring a text line image to be recognized; and a recognition model for inputting the text line image into a pre-trained text recognition model to obtain text in the text line image output by the text recognition model; wherein the text recognition model is obtained by training a preset training model based on a first target character, the first target character being a character obtained by character conversion of a first semantic fusion feature, the first semantic fusion feature being a feature obtained by fusing a first target sequence feature in a training sample image with a first semantic feature of the sample image, and the first semantic feature being a feature obtained based on the first target sequence feature.

[0133] According to one or more embodiments of the present disclosure, Example 11 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1 to 9.

[0134] According to one or more embodiments of this disclosure, Example 12 provides an electronic device including: a storage device having a computer program stored thereon; and a processing device for executing the computer program in the storage device to implement the steps of the method described in any one of Examples 1 to 9.

[0135] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0136] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0137] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A text recognition method, characterized in that, The method includes: Obtain the image of the text line to be recognized; The text line image is input into a pre-trained text recognition model to obtain the text in the text line image output by the text recognition model; The text recognition model is obtained by training a preset training model based on a first target character. The first target character is a character obtained by converting a first semantic fusion feature into a character. The first semantic fusion feature is a feature obtained by fusing a first target sequence feature in a sample image used for training with a first semantic feature of the sample image. The first semantic feature is a feature obtained based on the first target sequence feature. The sample image is an image used to train the text recognition model. The first semantic feature is determined by: converting the first target sequence feature into a first character to be processed; obtaining the confidence level of the first character to be processed; extracting the first semantic feature from the first characters to be processed whose confidence level is greater than or equal to a preset confidence threshold; the first target sequence feature is obtained by extracting a first sequence feature to be encoded from the sample image and encoding the first sequence feature to be encoded.

2. The method according to claim 1, characterized in that, The text recognition model was trained in the following way: Obtain multiple sample images for training; Obtain the first target sequence features from the sample images; The first semantic feature of the sample image is obtained based on the first target sequence feature; The first target sequence feature is fused with the first semantic feature to obtain the first semantic fusion feature; Convert the first semantic fusion feature into the first target character; The text recognition model is obtained by training a preset training model based on the first target character.

3. The method according to claim 2, characterized in that, The step of fusing the first target sequence feature with the first semantic feature to obtain the first semantic fusion feature includes: The first target sequence feature and the first semantic feature are added together to obtain the first semantic fusion feature.

4. The method according to claim 2, characterized in that, The step of converting the first semantic fusion feature into the first target character includes: The first semantic fusion feature is decoded, and the decoded first semantic fusion feature is converted into the corresponding first target character.

5. The method according to any one of claims 1 to 4, characterized in that, The text recognition model includes a feature extraction model, at least one semantic fusion model, a decoding model, and a first fully connected layer. In the case where there are multiple semantic fusion models, the multiple semantic fusion models are coupled in series sequentially. The feature extraction model is used to obtain the second sequence features to be encoded from the input text line image; The semantic fusion model is used to encode the second sequence features output by the feature extraction model to obtain the second target sequence features, and to obtain the second semantic features of the text line image based on the second target sequence features. The second target sequence features and the second semantic features are then fused to obtain the second semantic fusion features. The decoding model is used to decode the second semantic fusion feature output by the semantic fusion model; The first fully connected layer is used to convert the decoded second semantic fusion feature output by the decoding model into a second target character.

6. The method according to claim 5, characterized in that, The semantic fusion model includes: an encoding sub-model, a second fully connected layer, and a language sub-model; The encoding sub-model is used to encode the second target sequence features output by the feature extraction model to obtain the second target sequence features. The second fully connected layer is used to convert the second target sequence features output by the encoding sub-model into a second character to be processed; The language sub-model is used to extract the second semantic features from the second character to be processed output by the second fully connected layer.

7. A text recognition device, characterized in that, The device includes: The acquisition module is used to acquire images of the text lines to be recognized; The recognition module is used to input the text line image into a pre-trained text recognition model to obtain the text in the text line image output by the text recognition model; The text recognition model is obtained by training a preset training model based on a first target character. The first target character is a character obtained by converting a first semantic fusion feature into a character. The first semantic fusion feature is a feature obtained by fusing a first target sequence feature in a sample image used for training with a first semantic feature of the sample image. The first semantic feature is a feature obtained based on the first target sequence feature. The sample image is an image used to train the text recognition model. The first semantic feature is determined by: converting the first target sequence feature into a first character to be processed; obtaining the confidence level of the first character to be processed; extracting the first semantic feature from the first characters to be processed whose confidence level is greater than or equal to a preset confidence threshold; the first target sequence feature is obtained by extracting a first sequence feature to be encoded from the sample image and encoding the first sequence feature to be encoded.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method according to any one of claims 1 to 6.

9. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Character recognition model training method and device, and character recognition method and device

    CN113657399A