Text Recognition in Images

By using a single character model to train text data in multiple languages ​​and different orientations, the problem of excessive consumption of model training complexity and computing resources in the prior art is solved, and efficient and simple text recognition is achieved.

CN113269009BActive Publication Date: 2025-05-30MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010093899.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-02-14
Publication Date
2025-05-30
Estimated Expiration
2040-02-14

AI Technical Summary

Technical Problem

When existing text recognition technologies process text in different orientations and multiple languages ​​in images, they need to design multiple dedicated models, resulting in increased model training complexity and excessive computing resource consumption.

Method used

Using a single character model, by training multiple training text line areas containing different directions and multiple language texts, the probability distribution information of the character model units in the target text line area is determined without determining the direction and language of the text.

Benefits of technology

It realizes that there is no need to determine text orientation and language during the text recognition process, improves recognition efficiency and simplicity, and reduces the need for storage space and computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113269009B_ABST
    Figure CN113269009B_ABST
Patent Text Reader

Abstract

According to an implementation of the present disclosure, a solution for text recognition in images is proposed. In this solution, a target text line region expected to have text to be recognized is determined from an image. Using a single-character model, probability distribution information of at least one character model unit presented in the target text line region is determined. The single-character model is trained based on: multiple training text line regions and corresponding ground truth texts in the multiple training text line regions. The texts in the multiple training text line regions are organized in different orientations, and / or the ground truth texts include texts related to multiple languages (e.g., texts related to Latin languages and oriental languages). Based on the determined probability distribution information, the text in the target text line region can be determined. The application of the single-character model makes the text recognition process more efficient and simpler.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Text recognition refers to recognizing text in an image. The text form in the image text can be printed text, handwritten text, digital ink text, etc. The image containing text can be a digital image taken by an electronic device, a scanned version of a document, a text image rendered by digital ink, and any other image containing text. Text recognition in images has many uses. For example, it can be used to digitize handwritten characters, to recognize expected information such as license plate numbers and document information from the captured images, to digitize scanned documents, to implement image-based information retrieval, to digital ink recognition systems, and so on. Many text recognition technologies have been proposed. However, due to the diverse forms of text presented in images, it is desirable to provide a more optimized solution for text recognition. Summary of the Invention

[0002] According to an implementation of the present disclosure, a text recognition solution for images is proposed. In this solution, a target text line region expected to have text to be recognized is determined from the image. Using a single character model, probability distribution information of at least one character model unit presented in the target text line region is determined. The single character model is trained based on: multiple training text line regions and the corresponding ground truth text in the multiple training text line regions. The text in the multiple training text line regions is organized in different orientations, and / or the ground truth text includes text related to multiple languages (e.g., text related to Latin languages and oriental languages). Based on the determined probability distribution information, the text in the target text line region can be determined. Through this solution, it is not necessary to train multiple character models for different text orientations and / or different languages. The application of the single character model makes it unnecessary to determine the text orientation and / or language in the target text line region during the text recognition process, thus being more efficient and convenient.

[0003] The Summary of the Invention section is provided to introduce a selection of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify the key features or main features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Brief Description of the Drawings

[0004] Figure 1 A block diagram showing a computing environment capable of implementing multiple implementations of the present disclosure;

[0005] Figure 2 A block diagram showing a text recognition module according to some implementations of the present disclosure;

[0006] Figure 3 An example showing a target text line region determined from an image according to some implementations of the present disclosure;

[0007] Figures 4A to 4C Some examples of text line preprocessing according to some implementations of the present disclosure are shown;

[0008] Figure 5 A block diagram of a text recognition module according to some other implementations of the present disclosure is shown; and

[0009] Figure 6A and Figure 6B A flowchart of a process for text recognition according to some implementations of the present disclosure is shown.

[0010] In these figures, the same or similar reference signs are used to denote the same or similar elements. Detailed Description

[0011] The present disclosure will now be described with reference to several example implementations. It should be understood that these implementations are described only to enable those of ordinary skill in the art to better understand and thus implement the present disclosure, and do not imply any limitation on the scope of the present disclosure.

[0012] As used herein, the term "comprising" and its variants are to be construed as open-ended terms meaning "including but not limited to". The term "based on" is to be construed as "at least partially based on". The terms "one implementation" and "an implementation" are to be construed as "at least one implementation". The term "another implementation" is to be construed as "at least one other implementation". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0013] In this document, a "machine learning model" may also be referred to as a "learning model", "learning network", "network model", or "model". A "neural network" or "neural network model" is a type of deep machine learning model. The parameter set of a machine learning model is determined through training. A machine learning model uses the trained parameter set to map the received input to the corresponding output. Therefore, the training process of a machine learning model can be considered as learning the mapping or association relationship from the input to the output from the training data.

[0014] Figure 1 A block diagram of a computing device 100 capable of implementing multiple implementations of the present disclosure is shown. It should be understood that Figure 1 the computing device 100 shown is merely exemplary and should not impose any limitation on the functions and scope of the implementations described in the present disclosure. As Figure 1As shown, computing device 100 includes a computing device 100 in the form of a general-purpose computing device. The components of computing device 100 may include, but are not limited to, one or more processors or processing units 110, a memory 120, a storage device 130, one or more communication units 140, one or more input devices 150, and one or more output devices 160.

[0015] In some implementations, computing device 100 may be implemented as various user terminals or service terminals. A service terminal may be a server, a large computing device, etc. provided by various service providers. A user terminal may be any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that computing device 100 is capable of supporting any type of user interface (such as a "wearable" circuit, etc.).

[0016] Processing unit 110 may be an actual or virtual processor and is capable of performing various processes according to programs stored in memory 120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of computing device 100. Processing unit 110 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0017] Computing device 100 generally includes multiple computer storage media. Such media may be any accessible media available to computing device 100, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 120 may be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 130 may be removable or non-removable media and may include machine-readable media, such as a memory, a flash drive, a disk, or any other media that can be used to store information and / or data and can be accessed within computing device 100.

[0018] Computing device 100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 1As shown, a disk drive for reading from and writing to a removable, non-volatile disk and an optical disk drive for reading from and writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data medium interfaces.

[0019] The communication unit 140 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device 100 can be implemented in a single computing cluster or multiple computer machines that are capable of communicating via a communication connection. Thus, the computing device 100 can operate in a networked environment using a logical connection with one or more other servers, personal computers (PCs), or another general network node.

[0020] The input device 150 can be one or more various input devices such as a mouse, keyboard, trackball, voice input device, etc. The output device 160 can be one or more output devices such as a display, speaker, printer, etc. The computing device 100 can also communicate with one or more external devices (not shown) as needed via the communication unit 140, external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the computing device 100, or communicate with any device that enables the computing device 100 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0021] In some implementations, in addition to being integrated on a single device, some or all of the various components of the computing device 100 can also be provided in the form of a cloud computing architecture. In a cloud computing architecture, these components can be remotely located and can work together to implement the functions described in this disclosure. In some implementations, cloud computing provides computing, software, data access, and storage services that do not require the end user to be aware of the physical location or configuration of the system or hardware providing these services. In various implementations, cloud computing uses appropriate protocols to provide services over a wide area network such as the Internet. For example, a cloud computing provider provides applications over a wide area network, and they can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated at a remote data center location or they can be dispersed. The cloud computing infrastructure can provide services through a shared data center, even though they appear as a single access point for the user. Thus, the components and functions described herein can be provided from a service provider at a remote location using a cloud computing architecture. Alternatively, they can be provided from a conventional server, or they can be directly or otherwise installed on a client device.

[0022] The computing device 100 can be used to implement text recognition in multiple implementations of the present disclosure. The memory 120 may include one or more modules having one or more program instructions that can be accessed and run by the processing unit 110 to implement the functions of the various implementations described herein. For example, the memory 120 may include a text recognition module 122 for performing text recognition in an image.

[0023] During text recognition, the computing device 100 is capable of receiving an image 170 to be processed through the input device 150. The image 170 may be input by the user or specified by the user, or obtained or received from other sources. The image 170 presents text, such as Chinese text "Single-point vegetables", English text "Vegetables", Chinese and English mixed text "Sandy potatoes SandPotato", and so on. The text recognition module 122 is configured to perform text recognition processing on the image 170 and use the text recognized from the image 170 as the output 180. The output 180 may optionally be output via the output device 160, such as being presented to the user or provided to an external device. In some implementations, the output 180 may be stored for subsequent use and / or used as an input for subsequent processing of the image 170 (e.g., information retrieval based on the image 170). The embodiments of the present disclosure are not limited in this regard.

[0024] It should be understood that Figure 1 The components and arrangements of the illustrated computing device are merely examples, and a computing device suitable for implementing the example implementations described in the present disclosure may include one or more different components, other components, and / or different arrangements. Figure 1 The input images and the outputs of text recognition shown are also merely examples. The computing device may be suitable for processing any other image according to the example implementations of the present disclosure to recognize or attempt to recognize the text in the image. The text in the image may be printed text, handwritten text, and digital ink text. The image used for performing text recognition may be any type of image, such as a digital image taken by an electronic device, a scanned version of a document, an image containing handwritten text or digital ink text, and any other type of image.

[0025] As mentioned above, there may be many variations in the layout of the text presented in the image. For example, in the image, a group of characters may be arranged in a vertical direction (e.g., Figure 1 the Chinese text "Single-point vegetables" in the image 170), while another group of characters may be arranged in a horizontal direction (e.g., Figure 1the English text “Vegetables” in the image 170 and other texts). In addition, in many practical applications, such as in business documents, store signs, restaurant menus, etc., mixed-language text is very common (e.g., Figure 1 “Sand Potato” in the image 170). In this article, “mixed-language text” refers to the coexistence of characters in two or more languages in a sentence or a part of the text. Correspondingly, “single-language text” refers to the text in which only characters of a single language exist. Of course, in mixed-language text and single-language text, in addition to language-specific characters, there may also be some common characters, such as numbers, punctuation marks, or other symbols.

[0026] Most text recognition schemes usually locate an image area where text may exist in the image, and then perform processing on this image area to recognize the text that may exist therein. Considering the differences in character arrangement directions and characters of different languages, in conventional schemes, it is required to design multiple dedicated text recognition models for each orientation and multiple dedicated text recognition models for different languages to process the corresponding image areas. This requires that after locating a specific image area, it is first necessary to determine the orientation or language of the text in the image area, and then input the image area into the text recognition model trained for the determined orientation or language for further processing.

[0027] This approach not only increases the complexity of model training, but also requires more storage space for model storage and more computing resources for implementing complex text recognition processes during the actual text recognition process. If the image contains not only text with different orientations but also text in different languages, it is necessary to design corresponding text recognition models for a certain orientation of each language, which further increases the complexity of model training and use. In addition, the text recognition result of a specific image area is highly dependent on the accuracy of the orientation and language recognition of the text in the image area, which may be difficult to guarantee in some complex situations (e.g., image distortion, frequent alternation of characters in different languages, etc.).

[0028] According to the implementation of the present disclosure, a solution for text recognition in images is proposed. This solution uses a single model to recognize text organized in various orientations and / or text in different languages in an image, without involving multiple separate dedicated models to achieve text recognition specific to a language or text orientation. The single model used in this solution can be a single-character model, which is used to determine the probability distribution information of at least one character model unit presented in the target text line region. The single-character model is trained based on multiple training text line regions and the corresponding ground truth text in the multiple training text line regions. The text in the training text line regions can include text organized in different orientations and / or text related to multiple languages, so that the single-character model can determine the character occurrence probabilities of text in different orientations and / or different languages.

[0029] Using such a single-character model, any target text line region determined from the image is directly provided as the model input, without determining the orientation and / or language of the text in the target text line region. The use of a single model makes the text recognition process faster and simpler, and is particularly suitable for text recognition in complex applications (e.g., images with multiple orientations and / or mixed-language text). The application of a single model reduces the requirements for storage space and computational resource consumption. Moreover, the requirements of a single model can also simplify the model training process, without consuming processing resources and time to train different models for specific orientations or specific languages.

[0030] The following will continue to describe in detail an example implementation of text recognition from an image with reference to the accompanying drawings.

[0031] Figure 2 An example structure of a text recognition module according to some implementations of the present disclosure is shown. The text recognition module can be, for example, Figure 1 the text recognition module 122 in the computing device 100. For the convenience of discussion, the text recognition module will be described with reference to Figure 1 As shown in the figure, the text recognition module 122 includes a text line detector 210, a single-character model 220, and a text decoder 230 to implement the recognition of some or all of the text in the image 170.

[0032] The text line detector 210 is configured to determine a target text line region 212 in the image 170. The target text line region refers to the region in the image 170 that is expected to have text to be recognized, such as a text region with a specific arrangement. The specific arrangement can include, for example, arranging each character in the text in a horizontal straight line, vertical straight line, curve, or inclined line, etc. The text line detector 210 can detect one or more target text line regions 212 from the image 170.

[0033] The text line detector 210 may be configured to detect or locate a target text line region 212 from the image 170 using various text line detection methods. In some implementations, the text line detector 210 may utilize an automatic text line detection algorithm to process the image 170 to determine one or more target text line regions 212 from the image 170. For example, the text line detector 210 may utilize a machine learning model or a neural network (e.g., a machine learning model based on a relational network) to automatically detect one or more target text line regions 212 from the image 170. When determining the target text line region 212, the text line detector 210 does not require the ability to recognize specific characters or text in the image 170, but rather determines whether a set of pixels or larger image units in the image 170 are likely to present text. Alternatively or additionally, the text line detector 210 may also determine or assist in determining one or more target text line regions 212 in the image 170 by means of manual calibration or the like. It should be understood that the text line detector 210 may employ any text line detection technology currently in use or to be developed in the future to detect the target text line region 212 in the image 170.

[0034] Figure 3 An example of the target text line region 212 determined from the image 170 is shown. As shown, the text line detector 210 may determine multiple target text line regions 212-1, 212-2,..., 212-10 (collectively or individually referred to as the target text line region 212) in the image 170, and each target text line region 212 is considered to have text to be recognized. It should be understood that Figure 3 The division of the target text line regions shown is merely an example. In other embodiments, the text line detector 210 may determine one or more target text line regions in the image 170 to be of other sizes, dimensions, and / or orientations. For example, if a line of text in the image 170 is arranged otherwise (e.g., multiple characters are arranged in a line with a certain curvature rather than positioned vertically or horizontally), the text line detector 210 may also detect the target text line region 212 that bounds such a line of text.

[0035] In some cases, the text in the target text line region 212 of the image 170 may be organized in any orientation. For example, in Figure 3In the example, the text in the target text line region 212-1 is organized in a vertical orientation, while the text in the target text lines 212-2 to 212-10 is organized in a horizontal orientation. In this document, "vertically oriented" text means that multiple characters of the text are written (arranged) vertically, so the bottom of one character is closer to the top of an adjacent character; "horizontally oriented" text means that multiple characters of the text are written (arranged) horizontally, so the side of one character is closer to the side of an adjacent character. For reading, a reader can generally read "vertically oriented" text in a roughly top-to-bottom order, and read "horizontally oriented" text in a roughly left-to-right order. However, it should be understood that in an image, "vertically oriented" text does not necessarily appear in the image in a vertical direction, but can have a larger or smaller angular offset relative to the vertical axis. Similarly, "horizontally oriented" text may not necessarily appear in the image in a horizontal direction, but can have a larger or smaller angular offset relative to the horizontal axis. The specific offset situation depends on the specific design of the text in the image, the process of taking or acquiring the image, and so on.

[0036] In addition to or as an alternative to potentially having different orientations, the text in the target text line region 212 of the image 170 may also have characters in different languages. In some cases, the text in a target text line region 212 may be one or more characters in a single language. For example, Figure 3 the target text line region 212-1 in [description] includes multiple Chinese characters in only Chinese, the target text line region 212-2 includes multiple letters in only English, and Figure 3 the target text line regions 212-3, 212-5, 212-7, and 212-9 in [description] include mixed-language text with Chinese characters and English letters. In some embodiments, in addition to characters in various languages, one or more target text line regions 212 may also include other common symbols, such as numbers, punctuation marks, currency symbols, etc. For example, Figure 3 the target text line regions 212-4, 212-6, 212-8, and 212-10 in [description].

[0037] According to the implementation of the present disclosure, it is desirable to be able to use a single model to identify text with any orientation and / or text in different languages in the target text line region 212 without specifically distinguishing the orientation and / or language of the text. Specifically, the text decoder 230 can use a single character model 220 to support such single identification.

[0038] In some implementations, before being provided to the text decoder 230, some preprocessing may be performed on the target text line region 212 determined from the image 170. The text recognition module 122 may include one or more sub-modules (not shown) for performing corresponding preprocessing operations on the target text line region 212. In some implementations, the purpose of the preprocessing is to normalize different target text line regions recognized from the image to a single shape and / or size, to facilitate the analysis and processing by the single-character model 220 and the text decoder 230.

[0039] In some examples, the preprocessing operation may include a size normalization operation to enlarge or reduce the target text line region to a predetermined size for subsequent processing. For example, Figure 4A and Figure 4B shows scaling the target text line regions 212-1 and 212-2 in the image 170 to a predetermined size.

[0040] Alternatively or additionally, the preprocessing operation may further include an orientation classification and a rotation operation based on the orientation classification, so that the characters of the text in a line are distributed in a predetermined direction. For example, the target text line region 212 having text oriented longitudinally may be rotated so that multiple characters in the text are distributed in the horizontal direction. As Figure 4C shown, the target text line region 212-1 having text oriented longitudinally may be rotated counterclockwise by 90 degrees so that multiple characters are sequentially distributed in the horizontal direction. Text tilted at other angles relative to the vertical axis may also be rotated so that the characters are sequentially distributed in the horizontal direction, while the target text line region having text oriented horizontally and distributed in the horizontal direction may remain unchanged. If the text of the target text line region oriented horizontally is tilted at a certain angle relative to the horizontal axis, the target text line region may also be rotated to the horizontal direction. In another example, the target text line regions having different text orientations and different tilt angles may also be rotated to be sequentially distributed in the vertical direction, or may also be rotated to a certain angle (such as tilted 45 degrees) relative to the vertical axis or the horizontal axis.

[0041] In still other examples, if the target text line region 212 has curved lines (i.e., a line of text is defined by curved lines), the target text line region 212 may be corrected to generate a text line region having a predetermined shape (e.g., a rectangular shape). For example, for the target text line region 212 in which multiple characters are arranged in a line with a certain curvature, the target text line region may be corrected to a predetermined shape through various correction operations such as stretching and shrinking. Alternatively or additionally, if the text in the target text line region 212 is mirror text, mirror processing may also be performed on the target text line region 212.

[0042] It should be understood that the above are only some examples of preprocessing operations applicable to the target text line region 212. Depending on the actual application, other preprocessing operations can also be additionally or alternatively applied. The scope of implementation of the present disclosure is not limited in this regard.

[0043] Return reference Figure 2 , after the target text line region 212 (after the predetermined processing) is provided to the text decoder 230. The text decoder 230 uses the trained single-character model 220 to determine the probability distribution information of one or more character model units presented by the target text line region 212. As used herein, a character model unit is the basic unit for the single-character model 220 to perform probability determination, and each character model unit includes one or more characters or symbols in a predetermined character set. The probability distribution information indicates the conditional probability of each possible character model unit based on the target text line region 212. In some implementations, the text decoder 230 uses the single-character model 220 to determine the sequence of character model units with the highest occurrence probability in the target text line region 212.

[0044] The predetermined character set can include multiple characters used in at least one predetermined language. A character can be the basic unit used in a language. In languages such as Latin-based languages, a character includes letters used to form words, and in Eastern languages, a character can include a single Chinese character. The specific characters included in the predetermined character set depend on the design of the single-character model 220, which will be discussed in more detail below. In some implementations, in addition to the characters in a certain language, the predetermined character set can also include one or more general symbols, such as numbers, punctuation marks adopted in the language, currency symbols, or other symbols.

[0045] In some implementations, the single-character model 220 can be configured to be "uniform" for different orientations of the text, that is, the single-character model 220 has the ability to process the target text line region with text organized in any orientation. In some implementations, the single-character model 220 can be configured to be "uniform" for texts in different languages, that is, the single-character model 220 has the ability to process the target text line region with texts in different languages (for example, single-language texts in different languages and mixed-language texts). In still other implementations, the single-character model 220 can also be configured to be "uniform" for different orientations of the text and for texts in multiple languages, that is, a single model is used to achieve text recognition for any orientation and multiple languages. Therefore, the use of the single-character model 220 eliminates the need to perform the determination of the orientation and / or language of the text in the target text line region.

[0046] In an implementation of the present disclosure, in order to obtain the ability to recognize text in different orientations and / or multiple languages, the single-character model 220 is configured as a machine learning model and learns the corresponding ability from training data through machine learning. The training data for training the single-character model 220 includes a plurality of training text line regions and the text annotated in these training text line regions (also referred to as "ground truth text" or "known text"). In some implementations, in order for the single-character model 220 to recognize text organized in different orientations in various target text line regions, during the training phase, the plurality of training text line regions include ground truth text in different orientations. For example, the ground truth text in some training text line regions is arranged in a vertical orientation, while the ground truth text in other training text line regions is arranged in a horizontal orientation. To make the model training more accurate, the training text line regions and the ground truth text presented therein may have different angular variations in the vertical or horizontal orientation, such as being offset by a certain angle relative to the vertical axis or the horizontal axis. In some implementations, the training text line regions may also be used for model training after some preprocessing (such as size normalization, rotation, correction, mirror operation, etc.).

[0047] In some implementations, in order for the single-character model 220 to recognize the character model units of multiple languages that may appear in the target text line regions, during the training phase, the ground truth text in the plurality of training text line regions in the training data may include multiple texts related to these languages. The text for training may include single-language text for each language and, in some cases, may also include mixed-language text. Here, the mixed-language text may include characters of any two or more predetermined languages. In some cases, the mixed-language text is not necessary, and by using the single-language text of multiple languages as training data, the single-character model 220 can also learn the character features in different languages, so as to be able to recognize the single-language text and mixed-language text of these languages.

[0048] In some implementations, the ground truth text in the training data can be text in Latin languages and text in Oriental languages, including one or more characters of at least one Latin language and one or more characters of at least one Oriental language. Latin languages include, for example but not limited to, languages such as English, French, German, Dutch, Italian, Spanish, Portuguese, etc. and their variant languages. Oriental languages are sometimes also referred to as Asian languages, and examples include but are not limited to languages such as Chinese (including Simplified Chinese and Traditional Chinese), Japanese, Korean, etc. and their variant languages. Chinese, Japanese, and Korean are also referred to as CJK languages. For example, the ground truth text can include a single Chinese character in one or more training text line regions, English in one or more training text line regions, and may also include a mixed text of Chinese and English that appears in one or more training text line regions. The ground truth text for training the single-character model 220 can also include text in three or more languages, such as including Chinese, Japanese, and English, etc. It should be understood that training text line regions and ground truth text corresponding to one or more other languages other than Latin languages and Oriental languages can be used to train the single-character model 220.

[0049] In some implementations, in addition to or as an alternative to the mixture of different language families, the ground truth text in the training data can also include text in different languages of the same language family or relatively similar language families. For example, the ground truth text for training can include text in different Latin languages (such as English and French), text in different Oriental languages (such as Chinese and Japanese), and / or their mixed language text. Generally speaking, if it is desired that the single-character model 220 can recognize text in multiple languages, the ground truth text with these languages can be used as training data for model training.

[0050] In some implementations, if it is desired that the single-character model 220 can perform the determination of the probability distribution of the character model unit for text with arbitrary orientations and multiple languages, the training data can be selected to have variations in both text orientation and language, that is, the text in some training text line regions is organized in multiple different orientations, and the text in some training text line regions includes characters of multiple predetermined languages. It can be understood that if it is only necessary to train the single-character model 220 to have the ability to recognize text organized in different orientations, the ground truth text in the training text line regions can include characters of a single language. Similarly, if it is only necessary to train the single-character model 220 to have the ability to recognize text in multiple languages without a single requirement for text orientation, all the ground truth text in the training text line regions can be arranged only in a single orientation (in a vertical or horizontal arrangement).

[0051] As mentioned above, the single character model 220 is trained to determine the probability distribution information of the character model units presented in the target text line area. Depending on the design of the single character model 220, the predetermined character set includes characters that may appear in the language targeted by the single character model 220. For example, if the single character model 220 is only trained to recognize text arranged in different orientations in a certain language, the predetermined character set includes the characters of the language and one or more common symbols that may appear in the application of the language. If the single character model 220 is trained to recognize text in multiple languages, the predetermined character set may include the characters of these languages, and one or more common symbols that may appear in the application of these languages.

[0052] Depending on the model selection and specific configuration, the single character model 220 can adopt different frameworks, different model structures, and be trained with different objective functions. In some implementations, the machine learning algorithms that can be used by the single character model 220 may include algorithms suitable for image processing and natural language processing, examples of which include but are not limited to convolutional neural networks (CNN), long short-term memory (LSTM) models, deep bidirectional LSTM (DB LSTM) models, recurrent neural networks (RNN), Transformers, feedforward sequential storage networks (FSMN), encoding and decoding models based on attention mechanisms, decision tree-based models (e.g., random forest models), support vector machines (SVM), or a combination of the foregoing multiple models / networks, etc. The specific working principles of these models / networks are well known to those skilled in the art and will not be described in detail here. For example, the single character model 220 can be a combination of CNN and DBLSTM models. It should also be understood that currently developed or future improvements to machine learning models can be applied accordingly to the example implementations of the present disclosure.

[0053] The training process of the single character model 220 enables the model to learn and focus on the specific features of characters (and universal symbols) of different orientations and / or different languages ​​from the training data. The single character model 220 can be trained using different objective functions based on the selected model architecture. For example, connection temporal classification (CTC), minimum cross entropy estimation criterion (CE), maximum mutual information estimation (MMIE), minimum classification error criterion (MCE), minimum word / phoneme error criterion (MWE / MPE) and other discriminative training criteria. The training of the single character model 220 can also be completed alternatively or additionally using methods such as stochastic gradient descent and forward error correction. In some implementations, the training of the single character model 220 can be completed by a device other than the computing device 100 to perform text recognition, such as a device with strong computing power. Of course, in some implementations, it is also feasible for the computing device 100 (alone or in combination with other computing devices) to perform model training.

[0054] The above discussion has described how the single-character model 220 achieves unity in text orientation and multiple languages. The single single-character model 220 obtained through training can thus implement a composite text recognition task. Accordingly, without the need to concern about the orientation and / or the mixture of languages of the text in the target text line region 212, the target text line region 212 can be directly input into the text decoder 230 to utilize the single single-character model 220, and the optimal sequence of character model units can be decoded.

[0055] In some implementations, in addition to the single-character model 220, the text decoder 230 can also utilize a language model and a predefined dictionary to obtain better text recognition results. The predefined dictionary includes text units in various languages. As used herein, a text unit refers to a text segment in a language that has a specific meaning, such as a word, a phrase, a sentence, etc. A text unit includes one or more character model units. In some implementations, the predefined dictionary also indicates the mapping between text units and character model units, that is, it indicates which character model units constitute each text unit. The language model is used to constrain the grammatical relationships between text units in various languages.

[0056] For any possible sequence of character model units, the text decoder 230 combines the probability distribution information of the character model units calculated from the target text line region 212 by the single-character model 220, the text units in the dictionary 520, and the language constraint scores at the text unit level calculated from the single language model 510, and selects the sequence of character model units with the optimal comprehensive score as the recognition result for output.

[0057] The language model can, for example, be an n-gram language model used in natural language processing. Some other examples of language models include a maximum entropy model, a Hidden Markov Model (HMM), a Conditional Random Field (CRF) model, a Recurrent Neural Network (RNN), a Long Short-Term Memory (LSTM) model, a Gated Recurrent Unit model (GRU), a Transformer, and other Neural Network Language Models (NNLMs). The specific working principles of these models / networks are well-known to those skilled in the art and will not be elaborated in detail herein. Through the training process, the language model can all be trained to be able to measure whether a sequence of character model units conforms to the constraints of one or more specific languages (e.g., grammatical constraints). By means of the language model, the finally determined text can conform to the constraints of one or more predefined languages, and meaningful text that may actually occur can be obtained.

[0058] In some implementations, the language model can be a single language model that can apply the constraints of multiple predefined languages to determine the text recognition of the target text line region 212. Figure 5An example of the text recognition module 122 according to some implementations of the present disclosure is shown, in which a single language model 510 and a predetermined dictionary 520 are added. The dictionary 520 may be stored in an internal storage device or an external storage device of the computing device 100 and is accessible to the text recognition module 122. The multiple predetermined languages may include any different languages, such as may include one or more Latin languages, one or more Oriental languages, and / or any other languages. The "single" of the single language model 510 lies in the ability to use a single model to uniformly learn and apply the constraints of multiple different languages (e.g., grammar constraints) to recognize the text of each language or the mixed text of these languages.

[0059] Among the multiple predetermined languages, there are relatively large differences in the number of common text units of different languages and the length of the characters they contain. For example, words in Latin languages and characters in Oriental languages. Therefore, in order to balance the size of the text unit sets in different languages and support the recognition of out-of-vocabulary words that do not appear in the recognition dictionary 520, in some implementations, sub-words can be used as the basic units (also known as "text sub-units") of the language model and the predetermined dictionary in certain languages. The text sub-units may include a part of a text unit in a Latin language. New vocabulary often may appear in Latin languages because as the language is used and developed, different characters may be used to form new text units (e.g., new words). The text sub-units can be included in the dictionary 520, and the combination of these text sub-units can thus be used to cover these possible new text units.

[0060] One or more text sub-units can be obtained by different methods, such as methods of word morphology analysis like stem-suffix splitting, sub-word learning methods based on large-scale corpora (such as Morfessor, G1G, and byte pair encoding (BPE), etc.). If the BPE method is adopted, the text sub-units are learned from the existing corpus with words as text units or the pre-computed word frequency statistics using the BPE algorithm. In some implementations, after obtaining one or more text sub-units, the corresponding text units in the corpus used to train the language model will all be converted into corresponding text sub-unit sequences through the BPE algorithm, and then various types of language models are trained.

[0061] In some implementations, the multiple predetermined languages targeted by the single language model 510 may correspond to the text in the multiple predetermined languages that the single character model 220 can process. Specifically, if the single character model 220 is configured to determine the probability distribution information of characters in multiple predetermined languages (e.g., Latin languages and Oriental languages) in the target text line region 212, the single language model 510 is also configured to determine and apply language constraints for these predetermined languages to identify the corresponding text in the target text line region 212. Alternatively, if the single character model 220 is configured to process text in any orientation of a specific language, the single language model 510 may be configured to determine and apply language constraints for the specific language and one or more other languages (e.g., if there are one or more other single character models configured to process text in any orientation of other languages).

[0062] To enable the single language model 510 to determine and apply constraints for multiple languages, the single language model 510 is configured as a machine learning model and learns the corresponding capabilities from training data through machine learning. The training data for training the single language model 510 may include corpora based on multiple predetermined languages. The single language model can score the grammatical expressions of sentences composed of any one or more language text units within a specific language. Thus, the single language model 510 can judge at a coarser granularity than the character level (e.g., text unit granularity) whether the occurrence of a certain combination of characters in the target text line region conforms to the constraints of a specific language.

[0063] The corpora include language materials that actually appear in the actual use of the multiple languages targeted by the single language model 510, such as language materials from various sources such as novels, web pages, news, newspapers and magazines, papers, blogs, etc. The language materials obtained from each source can be digitized and, after certain analysis and processing, can be stored in a corpus for use when performing model training. In some implementations, the corpora may include single language texts in multiple languages or may include mixed language texts in multiple languages.

[0064] Depending on the specific language model adopted, a corresponding training algorithm can be used to train the single language model 510. The implementation of the present disclosure is not limited in terms of the specific training algorithm of the single language model 510. In some implementations, the training of the single character model 220 can be completed by a device outside the computing device 100 that is to perform text recognition, such as a device with strong computing capabilities. Of course, in some implementations, the model training can be performed by the computing device 100 (alone or in combination with other computing devices).

[0065] It should be understood that although a single language model is discussed above, in some implementations, language-specific language models can be applied to impose constraints. For example, if the single-character model 220 deterministically determines the probability distribution information of the character model units presented in the target text line region for multiple predetermined languages, multiple language models can be applied to take into account the constraints of these predetermined languages.

[0066] When determining the text in the target text line region 212, the text decoder 230 uses the probability distribution information of the single-character model 220, the text units of the dictionary 520, and the constraints applied by the single language model 510 to identify the text in the target text line region 212 as the output 180. In some implementations, the text decoder 230 can utilize a weighted finite state transducer (WFST)-based decoding model (or decoding network) to determine the text in the target text line region 212. For an input sequence (such as a sequence of character model units), the WFST will determine whether to accept this sequence. If accepted, it will output its corresponding output sequence (such as a sequence of words) and its score.

[0067] The process of recognizing the text in the target text line region 212 can be considered as using an efficient search algorithm to optimize the combination of the dictionary and the language model into the WFST network, and combining the probability scores of the character model units provided by the character model to quickly find a globally optimal path. The output result corresponding to this path is the recognized text. It should be understood that in addition to the WFST-based decoding model, other static or dynamic search decoding algorithms that combine the character model, the language model, and the dictionary can also be applied to the text decoder 230.

[0068] In some implementations, the text recognition module 122 may further include additional sub-modules (not shown) for further measuring whether the text determined by the text decoder 230 for the target text line region 212 is correct or usable, so as to further improve the accuracy of text recognition. For example, the text recognition module 122 may further include an accept / reject sub-module for determining whether the text output by the text decoder 230 is text that may appear in the actual application of the language, such as text with actual meaning. This can avoid misidentifying non-text patterns that appear in the image as text. For another example, the text recognition module 122 may further include a confidence sub-module for determining the reliability of the text output by the text decoder 230. It should be understood that the text recognition module 122 may additionally or alternatively include one or more other sub-modules for implementing other desired functions. The implementations of the present disclosure are not limited in this regard.

[0069] In the text recognition module 122, one or more sub-modules employ machine learning or deep learning models / networks to implement corresponding functions, such as the text line detector 210, the single character model 220, the single language model 510, the text decoder 230, etc. In these implementations, for the functions to be implemented by each sub-module, the corresponding machine learning or deep learning models / networks can be trained separately based on the corresponding training data. It is also possible to train the multiple machine learning or deep learning models / networks included in the text recognition module 122 in an end-to-end manner either after separate training or at the beginning, so as to achieve the goal of recognizing text from the input image. In some implementations, the text line detector 210 can be trained separately independent of other sub-modules. Of course, other implementations are not limited to this.

[0070] Figure 6A FIG. 4 shows a flowchart of a process 600 according to some implementations of the present disclosure. The process 600 can be implemented by the computing device 100, for example, it can be implemented at the text recognition module 122 of the computing device 100.

[0071] At block 610, the computing device 100 determines a target text line region in the image, where text to be recognized is expected to be present. At block 620, the computing device 100 uses the single character model to determine probability distribution information of at least one character model unit presented in the target text line region. The single character model is trained based on: multiple training text line regions in which text is organized in different orientations, and the corresponding ground truth text of the multiple training text line regions. Each character model unit includes at least one character or symbol. At block 630, the computing device 100 determines the text in the target text line region based on the determined probability distribution information.

[0072] In some implementations, the corresponding ground truth text in the multiple training text line regions includes multiple texts related to multiple predetermined languages, and each text in the multiple texts includes single language text or mixed language text.

[0073] In some implementations, the multiple predetermined languages include at least one of the following: at least one Latin language and at least one Oriental language.

[0074] In some implementations, determining the text in the target text line region includes: generating the text in the target text line region based on the determined probability distribution information and with the aid of a single language model and a predetermined dictionary, where the predetermined dictionary includes at least text units of multiple predetermined languages, each text unit includes at least one character model unit, and the single language model constrains the syntactic relationships between the text units of multiple languages.

[0075] In some implementations, in the case of multiple predetermined languages including Latin languages, the predetermined dictionary further includes at least one text subunit determined from the corpus of the Latin language, and each text subunit includes a part of a text unit of the Latin language. In some implementations, the multiple predetermined languages include at least one Oriental language.

[0076] In some implementations, the determination of at least one text subunit includes performing byte pair encoding (BPE) on the corpus of the Latin language.

[0077] In some implementations, the single language model includes an n-gram language model.

[0078] In some implementations, determining the text in the target text line region includes: using a decoding model based on a weighted finite state transducer (WFST) to determine the text in the target text line region.

[0079] Figure 6B A flowchart of a process 602 according to some other implementations of the present disclosure is shown. The process 602 can be implemented by the computing device 100, for example, it can be implemented at the text recognition module 122 of the computing device 100.

[0080] In block 640, the computing device 100 determines a target text line region in the image, and text to be recognized is expected to be in the target text line region. In block 650, the computing device 100 uses a single character model to determine the probability distribution information of at least one character model unit presented in the target text line region without determining the orientation and language of the text in the target text line region. The single character model is trained based on: multiple training text line regions in which the text is organized in different orientations and the corresponding ground truth text in the multiple training text line regions, and the ground truth text includes at least text related to Latin languages and Oriental languages. Each character model unit includes at least one character or characters. In block 660, the computing device 100 determines the text in the target text line region based on the determined probability distribution information.

[0081] In some implementations, determining the text in the target text line region includes: generating the text in the target text line region based on the determined probability distribution information and with the aid of a single language model and a predetermined dictionary, where the predetermined dictionary includes at least text units of multiple predetermined languages, each text unit includes at least one character model unit, and the single language model constrains the syntactic relationships between text units of multiple languages.

[0082] In some implementations, in the case of multiple predetermined languages including Latin languages, the predetermined dictionary further includes at least one text subunit determined from the corpus of the Latin language, and each text subunit includes a part of a text unit of the Latin language. In some implementations, the multiple predetermined languages include at least one Oriental language.

[0083] In some implementations, the determination of at least one text subunit includes performing byte pair encoding (BPE) on a corpus of Latin languages.

[0084] In some implementations, a single language model includes an n-gram language model.

[0085] In some implementations, determining the text in the target text line region includes: using a decoding model based on a weighted finite state transducer (WFST) to determine the text in the target text line region.

[0086] Some example implementations of the present disclosure are listed below.

[0087] In a first aspect, the present disclosure provides a computer-implemented method. The method includes: determining a target text line region in an image, where the target text line region is expected to have text to be recognized; using a single character model to determine probability distribution information of at least one character model unit presented in the target text line region, each character model unit including at least one character or symbol, and the single character model being trained based on: multiple training text line regions in which text is organized in different orientations, and the corresponding ground truth text of the multiple training text line regions; and determining the text in the target text line region based on the determined probability distribution information.

[0088] In some implementations, the corresponding ground truth text in the multiple training text line regions includes multiple texts related to multiple predetermined languages, and each text in the multiple texts includes a single language text or a mixed language text.

[0089] In some implementations, the multiple predetermined languages include at least one of the following: at least one Latin language and at least one Oriental language.

[0090] In some implementations, determining the text in the target text line region includes: generating the text in the target text line region based on the determined probability distribution information and with the aid of a single language model and a predetermined dictionary, where the predetermined dictionary includes at least text units of multiple predetermined languages, each text unit including at least one character model unit, and the single language model constrains the syntactic relationships between the text units of multiple languages.

[0091] In some implementations, when the multiple predetermined languages include a Latin language, the predetermined dictionary further includes at least one text subunit determined from a corpus of the Latin language, each text subunit including a part of a text unit of the Latin language. In some implementations, the multiple predetermined languages include at least one Oriental language.

[0092] In some implementations, the determination of at least one text subunit includes performing byte pair encoding (BPE) on a corpus of Latin languages.

[0093] In some implementations, a single language model includes an n-gram language model.

[0094] In some implementations, determining the text in the target text line region includes: using a decoding model based on a weighted finite state transducer (WFST) to determine the text in the target text line region.

[0095] In a second aspect, the present disclosure provides an electronic device. The electronic device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, the instructions, when executed by the processing unit, cause the device to perform the following actions: determining a target text line region in an image, in which text to be recognized is expected; using a single character model to determine probability distribution information of at least one character model unit presented in the target text line region, each character model unit including at least one character or symbol, the single character model being trained based on: multiple training text line regions in which text is organized in different orientations, and corresponding ground truth texts of the multiple training text line regions; and determining the text in the target text line region based on the determined probability distribution information.

[0096] In some implementations, the corresponding ground truth texts in the multiple training text line regions include multiple texts related to multiple predetermined languages, each text in the multiple texts including a single language text or a mixed language text.

[0097] In some implementations, the multiple predetermined languages include at least one of the following: at least one Latin language and at least one Oriental language.

[0098] In some implementations, determining the text in the target text line region includes: generating the text in the target text line region based on the determined probability distribution information and with the aid of a single language model and a predetermined dictionary, where the predetermined dictionary includes at least text units of multiple predetermined languages, each text unit including at least one character model unit, and the single language model constrains the syntactic relationships between the text units of multiple languages.

[0099] In some implementations, when the multiple predetermined languages include a Latin language, the predetermined dictionary further includes at least one text subunit determined from the corpus of the Latin language, each text subunit including a part of a text unit of the Latin language. In some implementations, the multiple predetermined languages include at least one Oriental language.

[0100] In some implementations, the determination of at least one text subunit includes performing byte pair encoding (BPE) on the corpus of the Latin language.

[0101] In some implementations, a single language model includes an n-gram language model.

[0102] In some implementations, determining the text in the target text line region includes: using a decoding model based on a weighted finite state transducer (WFST) to determine the text in the target text line region.

[0103] In a third aspect, the present disclosure provides a computer program product, which is tangibly stored in a non-transitory computer storage medium and includes machine-executable instructions that, when executed by a device, cause the device to execute one or more implementations of the above method.

[0104] In a fourth aspect, the present disclosure provides a computer-readable medium having machine-executable instructions stored thereon that, when executed by a device, cause the device to execute one or more implementations of the method of the first aspect above.

[0105] In a fifth aspect, the present disclosure provides a computer-implemented method. The method includes: determining a target text line region in an image, where the target text line region is expected to have text to be recognized; using a single-character model to determine probability distribution information of at least one character model unit presented in the target text line region without determining the orientation and language of the text in the target text line region, each character model unit including at least one character or symbol, and the single-character model being trained based on: multiple training text line regions in which the text is organized in different orientations and the corresponding ground truth text in the multiple training text line regions, the ground truth text including at least text related to Latin languages and oriental languages; and determining the text in the target text line region based on the determined probability distribution information.

[0106] In some implementations, determining the text in the target text line region includes: generating the text in the target text line region based on the determined probability distribution information and with the aid of a single language model and a predetermined dictionary, where the predetermined dictionary includes at least text units of multiple predetermined languages, each text unit including at least one character model unit, and the single language model constrains the syntactic relationships between the text units of multiple languages.

[0107] In some implementations, when the multiple predetermined languages include Latin languages, the predetermined dictionary further includes at least one text subunit determined from the corpus of the Latin language, each text subunit including a part of a text unit of the Latin language. In some implementations, the multiple predetermined languages include at least one oriental language.

[0108] In some implementations, at least one text subunit is determined by performing byte pair encoding (BPE) on the corpus of the Latin language.

[0109] In some implementations, the single language model includes an n-gram language model.

[0110] In some implementations, determining the text in the target text line region includes: using a decoding model based on a weighted finite state transducer (WFST) to determine the text in the target text line region.

[0111] In a sixth aspect, the present disclosure provides an electronic device. The electronic device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, the instructions, when executed by the processing unit, cause the device to perform the following actions: determining a target text line region in an image, in which text to be recognized is expected; using a single-character model to determine probability distribution information of at least one character model unit presented in the target text line region, without determining the orientation and language of the text in the target text line region, each character model unit includes at least one character, and the single-character model is trained based on: training text in the training image in multiple training text line regions organized in different orientations and corresponding ground truth text in the multiple training text line regions, the ground truth text includes at least text related to Latin languages and oriental languages; and determining the text in the target text line region based on the determined probability distribution information.

[0112] In some implementations, determining the text in the target text line region includes: generating the text in the target text line region based on the determined probability distribution information and with the aid of a single language model and a predetermined dictionary, where the predetermined dictionary includes at least text units of multiple predetermined languages, each text unit includes at least one character model unit, and the single language model constrains the syntactic relationships between text units of multiple languages.

[0113] In some implementations, when the multiple predetermined languages include Latin languages, the predetermined dictionary further includes at least one text subunit determined from the corpus of the Latin language, each text subunit includes a part of a text unit of the Latin language. In some implementations, the multiple predetermined languages include at least one oriental language.

[0114] In some implementations, determining at least one text subunit includes performing byte pair encoding (BPE) on the corpus of the Latin language.

[0115] In some implementations, the single language model includes an n-gram language model.

[0116] In some implementations, determining the text in the target text line region includes: using a decoding model based on a weighted finite state transducer (WFST) to determine the text in the target text line region.

[0117] In a seventh aspect, the present disclosure provides a computer program product tangibly stored in a non-transitory computer storage medium and including machine-executable instructions that, when executed by a device, cause the device to perform one or more implementations of the above-described method.

[0118] In an eighth aspect, the present disclosure provides a computer-readable medium having stored thereon machine-executable instructions that, when executed by a device, cause the device to perform one or more implementations of the method of the above fifth aspect.

[0119] The functions described above herein can be performed at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0120] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0121] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0122] In addition, although the operations are depicted in a particular order, this should be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation may also be implemented separately or in any suitable sub-combination in multiple implementations.

[0123] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An electronic device, comprising: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform the following actions: determine a target text line region organized in an arbitrary orientation in an image, the target text line region expected to have text to be recognized; using a single-character model, determine probability distribution information of at least one character model unit presented in the target text line region organized in an arbitrary orientation, each character model unit including at least one character or symbol, the single-character model being trained based on: multiple training text line regions in which text is organized in different orientations, and corresponding ground truth texts of the multiple training text line regions having different angular variations in a vertical orientation or a horizontal orientation; and based on the determined probability distribution information, determine the text in the target text line region.

2. The device according to claim 1, wherein the corresponding ground truth texts in the multiple training text line regions include multiple texts related to multiple predetermined languages, each text of the multiple texts including a single-language text or a mixed-language text.

3. The device according to claim 2, wherein the multiple predetermined languages include at least one of the following: at least one Latin language and at least one Oriental language.

4. The device according to claim 1, wherein determining the text in the target text line region comprises: generating the text in the target text line region based on the determined probability distribution information and with the aid of a single-language model and a predetermined dictionary, wherein the predetermined dictionary includes at least text units of multiple predetermined languages, each text unit including at least one character model unit, and the single-language model constrains the syntactic relationships between the text units of the multiple predetermined languages.

5. The device according to claim 4, wherein when the multiple predetermined languages include a Latin language, the predetermined dictionary further includes at least one text subunit determined from a corpus of the Latin language, each text subunit including a part of a text unit of the Latin language.

6. The device according to claim 5, wherein the determination of the at least one text subunit includes performing byte pair encoding (BPE) on the corpus of the Latin language.

7. The device according to claim 4, wherein the single-language model includes an n-gram language model.

8. The device according to claim 1, wherein determining the text in the target text line region comprises: using a decoding model based on a weighted finite state transducer (WFST) to determine the text in the target text line region.

9. A computer-implemented method, comprising: determine a target text line region organized in an arbitrary orientation in an image, the target text line region expected to have text to be recognized; Using a single-character model, determine probability distribution information of at least one character model unit presented in the target text line region organized in an arbitrary orientation, each character model unit including at least one character or symbol, and the single-character model is trained based on: multiple training text line regions in which the text is organized in different orientations, and corresponding ground truth texts of the multiple training text line regions having different angular variations in the vertical or horizontal orientation; And Based on the determined probability distribution information, determine the text in the target text line region.

10. The method according to claim 9, wherein the corresponding ground truth texts in the multiple training text line regions include multiple texts related to multiple predetermined languages, and each text in the multiple texts includes a single-language text or a mixed-language text.

11. The method according to claim 10, wherein the multiple predetermined languages include at least one of the following: at least one Latin language and at least one Oriental language.

12. The method according to claim 9, wherein determining the text in the target text line region Includes: Based on the determined probability distribution information and by means of a single-language model and a predetermined dictionary, generate the text in the target text line region, Wherein the predetermined dictionary includes at least text units of multiple predetermined languages, each text unit including at least one character model unit, and the single-language model constrains the syntactic relationships between the text units of the multiple predetermined languages.

13. The method according to claim 12, wherein in the case where the multiple predetermined languages include a Latin language, the predetermined dictionary further includes at least one text subunit determined from the corpus of the Latin language, and each text subunit includes a part of a text unit of the Latin language.

14. The method according to claim 13, wherein the determination of the at least one text subunit includes performing byte pair encoding (BPE) on the corpus of the Latin language.

15. The method according to claim 12, wherein the single-language model includes an n-gram language model.

16. The method according to claim 9, wherein determining the text in the target text line region Includes: Use a decoding model based on a weighted finite state transducer (WFST) to determine the text in the target text line region.

17. A computer program product, the computer program product being tangibly stored in a non-transitory computer storage medium and including machine-executable instructions that, when executed by a device, cause the device to: Determine a target text line region organized in an arbitrary orientation in an image, and text to be recognized is expected to be present in the target text line region; Using a single-character model, determine probability distribution information of at least one character model unit presented in the target text line region organized in an arbitrary orientation, each character model unit including at least one character or symbol, and the single-character model being trained based on: a plurality of training text line regions in which text is organized in different orientations, and corresponding ground truth texts of the plurality of training text line regions having different angular variations in a longitudinal orientation or a lateral orientation; And Based on the determined probability distribution information, determine the text in the target text line region.

18. The computer program product according to claim 17, wherein the ground truth texts in the plurality of training text line regions include a plurality of texts related to a plurality of predetermined languages, and each of the plurality of texts includes a single-language text or a mixed-language text.

19. The computer program product according to claim 17, wherein determining the text in the target text line region Includes: Based on the determined probability distribution information and by means of a single-language model and a predetermined dictionary, generate the text in the target text line region, Wherein the predetermined dictionary at least includes text units of a plurality of predetermined languages, each text unit including at least one character model unit, and the single-language model constrains the syntactic relationships between the text units of the plurality of predetermined languages.

20. An electronic device, Includes: A processing unit; And A memory, coupled to the processing unit and containing instructions stored thereon, the instructions, when executed by the processing unit, cause the device to perform the following actions: Determine a target text line region in an image organized in an arbitrary orientation, the target text line region being expected to have text to be recognized; Use a single-character model to determine probability distribution information of at least one character model unit presented in the target text line region organized in an arbitrary orientation without determining the orientation and language of the text in the target text line region, each character model unit including at least one character or symbol, and the single-character model being trained based on: a plurality of training text line regions in which text is organized in different orientations and corresponding ground truth texts of the plurality of training text line regions having different angular variations in a longitudinal orientation or a lateral orientation, the ground truth texts including at least texts related to Latin languages and Oriental languages; And Based on the determined probability distribution information, determine the text in the target text line region.

Citation Information

Patent Citations

  • Method and system for converting an image to text

    US20190087677A1