Font recognition methods, devices, readable media and electronic devices
By segmenting images and extracting contextual features using a pre-defined font recognition model, the problem of low font recognition accuracy and rate is solved, enabling efficient recognition of multiple categories and languages of fonts and rapid expansion of newly added fonts.
Patent Information
- Application Number
- CN202210023481.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-10
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-01-10
AI Technical Summary
Existing font recognition methods suffer from low recognition rates and poor accuracy, especially when faced with multiple fonts and complex backgrounds.
A pre-defined font recognition model is used to segment the image to be recognized, obtain sub-image features, and obtain contextual features through a multi-head attention module. Finally, the font type of the target text is determined by Euclidean distance.
It improves the accuracy and recognition rate of font recognition, enabling the recognition of multiple categories and languages of fonts without the need for real labeled data, and supports rapid expansion for new font types.
Smart Images

Figure CN114495080B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer vision processing, and more specifically, to a font recognition method, apparatus, readable medium, and electronic device. Background Technology
[0002] In font recognition, the sheer number of fonts (over 12,000 for Chinese characters alone) and the diverse features inherent in each font, coupled with the fact that multiple features of the same font may not be present in a single character, leading to the same or different features appearing on different characters, and even different fonts exhibiting many similar or identical features on the same character, makes font recognition extremely difficult. Existing font recognition methods typically suffer from low recognition rates and poor accuracy. Summary of the Invention
[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] This disclosure provides a font recognition method, apparatus, readable medium, and electronic device.
[0005] In a first aspect, this disclosure provides a font recognition method, the method comprising:
[0006] Obtain an image to be identified, wherein the image to be identified includes target text;
[0007] The image to be recognized is input into a preset font recognition model so that the preset font recognition model outputs the font type corresponding to the target text.
[0008] The preset font recognition model is used to divide the image to be recognized into multiple sub-images, obtain a first image feature corresponding to each sub-image, determine a second image feature corresponding to the image to be recognized based on the first image feature corresponding to each sub-image in the image to be recognized, the second image feature includes the contextual association features between each sub-image in the image to be recognized and other sub-images, and determine the font type corresponding to the target text based on the second image feature.
[0009] Secondly, this disclosure provides a font recognition device, the device comprising:
[0010] The acquisition module is configured to acquire an image to be recognized, wherein the image to be recognized includes target text;
[0011] The determining module is configured to input the image to be recognized into a preset font recognition model, so that the preset font recognition model outputs the font type corresponding to the target text;
[0012] The preset font recognition model is used to divide the image to be recognized into multiple sub-images, obtain a first image feature corresponding to each sub-image, determine a second image feature corresponding to the image to be recognized based on the first image feature corresponding to each sub-image in the image to be recognized, the second image feature includes the contextual association features between each sub-image in the image to be recognized and other sub-images, and determine the font type corresponding to the target text based on the second image feature.
[0013] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect above.
[0014] Fourthly, this disclosure provides an electronic device, comprising:
[0015] A storage device on which computer programs are stored;
[0016] A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect above.
[0017] The above technical solution involves inputting the image to be recognized into a preset font recognition model, which then outputs the font type corresponding to the target text. The preset font recognition model divides the image to be recognized into multiple sub-images and obtains a first image feature corresponding to each sub-image. Based on the first image feature, it determines a second image feature corresponding to the image to be recognized. The second image feature includes contextual association features between each sub-image and other sub-images. The font type corresponding to the target text is then determined based on the second image feature. This method of determining the second image feature based on the first image feature of each sub-image allows for a more comprehensive and accurate description of the image to be recognized based on the correlation between each character image and other sub-images, thereby effectively improving the accuracy of font recognition results and increasing the font recognition rate.
[0018] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0020] Figure 1 This is a flowchart illustrating a font recognition method according to an exemplary embodiment of this disclosure;
[0021] Figure 2 This is a schematic diagram illustrating the working principle of a preset font recognition model according to an exemplary embodiment of the present disclosure;
[0022] Figure 3 This is a schematic diagram illustrating the working principle of a TPA module according to an exemplary embodiment of this disclosure;
[0023] Figure 4 This is a flowchart illustrating a model training method according to an exemplary embodiment of the present disclosure;
[0024] Figure 5 This is a block diagram illustrating a font recognition device according to an exemplary embodiment of the present disclosure;
[0025] Figure 6 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0026] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0027] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0028] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0029] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0030] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0031] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0032] Before detailing the specific implementation methods of this disclosure, the application scenarios of this disclosure are explained below. This disclosure can be used to recognize the fonts in images and documents. Related font recognition methods generally fall into two categories: one is font recognition based on manually designed features, which uses human experience to manually design feature extractors to extract features for each character and then classifies the fonts accordingly. Because this method uses fixed, manually designed features to represent font features, it is easy to lose some useful font information. Furthermore, designing highly accurate feature descriptors generally requires careful engineering design and extensive domain expertise, making it difficult to obtain accurate feature descriptors and hindering the acquisition of comprehensive and accurate font features, thus also hindering the improvement of the final font recognition accuracy. The other category is font recognition based on deep learning, which uses deep neural networks to automatically extract font features and then uses the extracted features for font classification. This type of deep learning-based font recognition scheme has not overcome the limitation of numerous font categories. The deep learning models trained in this way often still suffer from low font recognition rates and poor accuracy.
[0033] To address the aforementioned technical problems, this disclosure provides a font recognition method, apparatus, readable medium, and electronic device. The font recognition method inputs an image to be recognized into a preset font recognition model, which outputs the font type corresponding to the target text. The preset font recognition model is used to divide the image to be recognized into multiple sub-images and obtain a first image feature corresponding to each sub-image. Based on the first image feature corresponding to each sub-image, a second image feature corresponding to the image to be recognized is determined. The second image feature includes contextual association features between each sub-image and other sub-images. The font type corresponding to the target text is determined based on the second image feature. Since determining the second image feature based on the first image feature corresponding to each sub-image allows for a more comprehensive and accurate description of the image to be recognized based on the correlation between each character image and other sub-images, determining the font type corresponding to the target text based on the second image feature effectively ensures the accuracy of the font recognition results and improves the font recognition rate.
[0034] The technical solution of this disclosure will be described in detail below with reference to specific embodiments.
[0035] Figure 1 This is a flowchart illustrating a font recognition method according to an exemplary embodiment of this disclosure; as shown below. Figure 1 As shown, the method may include the following steps:
[0036] Step 101: Obtain the image to be recognized, which includes the target text.
[0037] In addition to the target text, the image to be identified may also include an image background, and the image background and the target text together form the image to be identified.
[0038] Step 102: Input the image to be recognized into a preset font recognition model so that the preset font recognition model outputs the font type corresponding to the target text.
[0039] The preset font recognition model is used to divide the image to be recognized into multiple sub-images, obtain the first image feature corresponding to each sub-image, determine the second image feature corresponding to the image to be recognized based on the first image feature corresponding to each sub-image in the image to be recognized, the second image feature includes the contextual association feature between each sub-image in the image to be recognized and other sub-images, and determine the font type corresponding to the target text based on the second image feature.
[0040] It should be noted that the preset font recognition model may include an image segmentation module and a multi-head attention module. The image segmentation module is used to divide the image to be recognized into multiple sub-images, obtain a first image feature corresponding to each sub-image, and input the first image feature corresponding to each sub-image into the multi-head attention module, so that the multi-head attention module can obtain the contextual association features between each sub-image and other sub-images in the image to be recognized, thereby obtaining a second image feature corresponding to the image to be recognized. Specifically, when dividing the image to be recognized into multiple sub-images, the image segmentation module can divide the image to be recognized into multiple sub-images according to a preset image segmentation rule. This image segmentation rule can be based on dividing the image into multiple sub-images according to pixel equality, dividing the image to be recognized into multiple sub-images with a preset pixel ratio, or dividing continuous, textured parts into a single sub-image according to image texture, thereby obtaining multiple sub-images. The multi-head attention module can refer to the relevant descriptions of multi-head attention mechanisms in the prior art. Since the multi-head attention mechanism is widely used in the prior art and the relevant technology is relatively easy to obtain, it will not be elaborated upon here.
[0041] For example, Figure 2 This is a schematic diagram illustrating the working principle of a preset font recognition model as shown in an exemplary embodiment of this disclosure, such as... Figure 2 As shown, multiple datasets (each containing multiple font recognition images) are first created using the Batch module. Each dataset serves as a training dataset. Each training dataset is then input into the Backbone network to obtain image features for each font recognition image. The TPA (Temporal Pyramid Attention) mechanism then segments the font recognition image based on these image features, yielding the first image features of each segmented sub-image. These first image features are then concatenated and fed into the multi-head attention module, allowing it to capture the contextual image features reflecting the interactions between different sub-images. (The structural diagram of the TPA can be shown below.) Figure 3 As shown, Figure 3This is a schematic diagram illustrating the working principle of a TPA module according to an exemplary embodiment of this disclosure; This allows for a comprehensive and accurate description of the font recognition image through the context image features. Then, the second image features are dimensionality-reduced using an MLP (Multilayer Perceptron) to obtain low-dimensional features corresponding to each sub-image. For example, a 512-dimensional feature corresponding to a font recognition image is transformed into 64-dimensional data, i.e., a font recognition image is represented by a 64-dimensional vector (i.e., an embedding operation is performed on each font recognition image), thus obtaining the second image features of each font recognition image. Based on the embedding results corresponding to the font recognition images in each training dataset (i.e., the second image features of each font recognition image), the representative features of each recognizable font are calculated (i.e., through the embedding...). The space (embedded interval network) obtains the mean of the second image features of each font recognition image in the training dataset. The mean of the second image features is used as the representative feature of the corresponding font type in the training dataset. Similarly, representative features corresponding to multiple font types can be obtained. After obtaining the representative features corresponding to multiple font types, the image to be recognized can be obtained. The image to be recognized is then passed through the backbone network, the TPA module, and the MLP module to obtain the second image features corresponding to the image to be recognized. The sampler calculates the Euclidean distance between the representative features corresponding to multiple font types and the second image features corresponding to the image to be recognized. The font type corresponding to the target text in the image to be recognized is determined based on the Euclidean distance.
[0042] In addition, it should be noted that the above-described implementation method for determining the font type corresponding to the target text based on the second image feature may include the following steps S1 to S3:
[0043] S1, obtain the Euclidean distance between the second image feature and the representative feature of each of the multiple recognizable fonts corresponding to the preset font recognition model, so as to obtain multiple Euclidean distances between the image to be recognized and the representative features of the multiple recognizable fonts.
[0044] For example, if the preset font recognition model can recognize N fonts, then the preset font recognition model includes representative features of each of these N fonts. For instance, if the representative features corresponding to font A are [a1 a2 … a n If the representative feature corresponding to font B is [b1 b2 … b], n The second image feature corresponding to the current image to be identified is [x1 x2 … x]. n If the Euclidean distance between the current image to be recognized and the representative features of font A is [value missing], then the Euclidean distance between them is [value missing]. The Euclidean distance between the image to be recognized and the representative features of font B is currently... Similarly, the Euclidean distance between the image to be recognized and each recognizable font in the preset font recognition model can be obtained, wherein the preset font recognition model includes representative features of the recognizable font.
[0045] S2 determines the minimum target distance from multiple Euclidean distances.
[0046] For example, if the preset font recognition model includes representative features of 6 recognizable fonts, namely font A to font F, wherein the Euclidean distance between the second image feature of the image to be recognized and the representative feature of font A is less than the Euclidean distance between the image to be recognized and the representative features of fonts B, C, D, E, and F respectively, that is, the Euclidean distance between the second image feature of the image to be recognized and the representative feature of font A is the target distance.
[0047] S3, determine the font type corresponding to the target text in the image to be recognized based on the target distance.
[0048] In this step, one possible implementation is: if the target distance is less than a preset distance threshold, the target font type corresponding to the target representative feature used to calculate the target distance is taken as the font type of the target text.
[0049] Taking the steps shown in S2 above as an example, when the Euclidean distance between the second image feature of the image to be identified and the representative feature of font A is the target distance, if the target distance is less than the preset distance threshold, then font A is taken as the font type corresponding to the target text in the image to be identified.
[0050] In another possible implementation of this step, if the target distance is determined to be greater than or equal to a preset distance threshold, the font type corresponding to the target text is determined to be a newly added font.
[0051] Taking the steps shown in S2 above as an example, when the Euclidean distance between the second image feature of the image to be identified and the representative feature of font A is the target distance, if the target distance is greater than or equal to the preset distance threshold, then the font type corresponding to the target text in the image to be identified is determined to be a new font, that is, the preset font recognition model cannot identify the specific font type corresponding to the target text in the image to be identified.
[0052] The above technical solution inputs the image to be recognized into a preset font recognition model, so that the preset font recognition model outputs the font type corresponding to the target text. The preset font recognition model is used to divide the image to be recognized into multiple sub-images and obtain a first image feature corresponding to each sub-image. Based on the first image feature corresponding to each sub-image, a second image feature corresponding to the image to be recognized is determined. The second image feature includes the contextual relationship features between each sub-image and other sub-images. The font type corresponding to the target text is determined based on the second image feature. This method can more comprehensively and accurately describe the image to be recognized based on the correlation between each character image and other sub-images, thereby effectively improving the accuracy of font recognition results and increasing the font recognition rate.
[0053] Optionally, the preset font recognition model is also used to expand the recognizable font types through the steps shown in S4 to S6 below, as follows:
[0054] S4, obtain multiple first font recognition sample images of the target newly added font.
[0055] The first font recognition sample image includes a specified text sample of the target newly added font.
[0056] In this step, one possible implementation may include: obtaining a specified text sample of the target new font from a preset font corpus; obtaining a target background image from a preset background library; and combining the specified text sample and the target background image to form the first font recognition sample image.
[0057] It should be noted that, since the characters in the specified text samples in the first font recognition sample image are all obtained from the preset font corpus and are corresponding to the target newly added font, the specified text samples in the obtained first font recognition sample image are all of the target newly added font type.
[0058] S5, obtain the second image features corresponding to each first font recognition sample image.
[0059] In this step, the first font recognition sample image can be input into the image segmentation model, so that the image segmentation module divides the first font recognition sample image into multiple sub-images, obtains the first image feature corresponding to each sub-image, and inputs the first image feature corresponding to each sub-image into the multi-head attention module, so that the multi-head attention module obtains the contextual association features between each sub-image and other sub-images in the first font recognition sample image, thereby obtaining the second image feature corresponding to the first font recognition sample image.
[0060] S6, obtain the target mean of multiple second image features corresponding to the multiple first font recognition sample images, use the target mean as the target representative feature corresponding to the target newly added font, and store the target representative feature.
[0061] For example, the second image features corresponding to each of the 50 first font recognition sample images can be obtained, thus obtaining 50 second image features. The mean of these 50 second image features is used as the target representative feature corresponding to the newly added font.
[0062] The above technical solutions can not only effectively avoid the problem of difficulty in obtaining training data during the training process of font recognition models in related technologies, but also achieve the recognition of multiple categories and languages of font types without the need for real labeled data, and can also be quickly expanded for new font types that will appear in the future.
[0063] Figure 4 This is a flowchart illustrating a model training method according to an exemplary embodiment of this disclosure; as shown below. Figure 4 As shown, this preset font recognition model can be trained in the following way:
[0064] Step 401: Obtain multiple second font recognition sample images, which include annotation data for multiple first font types.
[0065] In this step, the second font recognition sample image can be generated by: obtaining characters of the first font type from a preset font corpus; obtaining a specified background image from a preset background library; and combining the characters of the first font type and the specified background image to form the second font recognition sample image. The preset font corpus can be a simplified Chinese corpus, a traditional Chinese corpus, or an English font corpus.
[0066] Step 402: Use the multiple second font recognition sample images as the first training dataset to pre-train the preset initial model to obtain the first undetermined model.
[0067] The preset initial model may include an image segmentation initial module and a multi-head attention initial module.
[0068] In this step, multiple font category recognition tasks can be established, and the models for these multiple font category recognition tasks can be trained using the first training dataset to obtain the first undetermined model.
[0069] Step 403: Obtain multiple third font recognition sample images, which include annotation data for multiple second font types.
[0070] The first font type may be the same as or different from the second font type.
[0071] It should be noted that the third font recognition sample image can be generated by obtaining the characters of the second font type from a preset font corpus; obtaining the required background image from a preset background library; and combining the characters of the second font type and the required background image to form the third font recognition sample image. This can effectively reduce the difficulty of obtaining training data during the font recognition model training process.
[0072] Step 404: Use the multiple third font recognition sample images as the second training dataset to train the first undetermined model to obtain the preset font recognition model.
[0073] It should be noted that the above model training process can refer to the meta-learning process in existing technologies. That is, through steps 401 to 402 above as the meta-training stage, multiple different classification tasks are constructed using synthetic datasets, and the model is trained using the MAML (Model-Agnostic Meta-Learning) method to obtain the first undetermined model. Then, through steps 403 to 404 above as the meta-testing stage, a second training dataset is constructed using synthetic data. The first undetermined model is then trained based on the second training dataset, which can effectively improve the convergence rate of the preset font recognition model and improve the training efficiency of the preset font recognition model.
[0074] Additionally, it should be noted that the above model training process can be conducted offline or online. After obtaining the preset font recognition model, if a new font type is added and the preset font recognition model needs to be able to recognize the new font type, it is only necessary to obtain the representative features of the new font type and save them into the preset font recognition model, which will enable the preset font recognition model to recognize the new font type.
[0075] The above technical solutions can effectively avoid the problem of difficulty in obtaining training data during the training process of font recognition models in related technologies. They can achieve the recognition of font types of multiple categories and languages without the need for real labeled data. Furthermore, the meta-learning training mechanism can obtain a preset font recognition model that can be quickly expanded for newly added font types in the future.
[0076] Figure 5 This is a block diagram illustrating a font recognition device according to an exemplary embodiment of the present disclosure; as shown below. Figure 5 As shown, the device may include:
[0077] The acquisition module 501 is configured to acquire an image to be recognized, which includes target text;
[0078] The determining module 502 is configured to input the image to be recognized into a preset font recognition model so that the preset font recognition model outputs the font type corresponding to the target text;
[0079] The preset font recognition model is used to divide the image to be recognized into multiple sub-images, obtain the first image feature corresponding to each sub-image, determine the second image feature corresponding to the image to be recognized based on the first image feature corresponding to each sub-image in the image to be recognized, the second image feature includes the contextual association feature between each sub-image in the image to be recognized and other sub-images, and determine the font type corresponding to the target text based on the second image feature.
[0080] Determining the second image feature corresponding to the image to be identified based on the first image feature corresponding to each sub-image in the image to be identified can more comprehensively and accurately describe the image to be identified based on the correlation between each character image and other sub-images. Therefore, determining the font type corresponding to the target text based on the second image feature can effectively improve the accuracy of font recognition results and also effectively improve the font recognition rate.
[0081] Optionally, this preset font recognition model is used for:
[0082] The Euclidean distance between the second image feature and the representative feature of each of the multiple recognizable fonts corresponding to the preset font recognition model is obtained, so as to obtain multiple Euclidean distances between the image to be recognized and the representative features of the multiple recognizable fonts;
[0083] Determine the minimum target distance from among multiple Euclidean distances;
[0084] The font type corresponding to the target text in the image to be identified is determined based on the target distance.
[0085] Optionally, this preset font recognition model is used for:
[0086] If the target distance is less than a preset distance threshold, the target font type corresponding to the target representative feature used to calculate the target distance will be used as the font type of the target text.
[0087] Optionally, this preset font recognition model is used for:
[0088] If the distance to the target is determined to be greater than or equal to a preset distance threshold, the font type corresponding to the target text is determined to be a newly added font.
[0089] Optionally, the preset font recognition model is also used for:
[0090] Obtain multiple first font recognition sample images of the target newly added font, wherein the first font recognition sample images include a specified text sample of the target newly added font;
[0091] Obtain the second image features corresponding to each first font recognition sample image;
[0092] Obtain the target mean of multiple second image features corresponding to the multiple first font recognition sample images, use the target mean as the target representative feature corresponding to the target newly added font, and store the target representative feature.
[0093] Optionally, this preset font recognition model is used for:
[0094] Obtain a specified text sample of the newly added font from the preset font corpus;
[0095] Obtain the target background image from the preset background library;
[0096] The specified text sample and the target background image are combined to form the first font recognition sample image.
[0097] Optionally, the device also includes a model training module 503, configured to:
[0098] Acquire multiple second font recognition sample images, wherein the multiple second font recognition sample images include annotation data of multiple first font types;
[0099] The multiple second font recognition sample images are used as the first training dataset to pre-train a preset initial model to obtain a first undetermined model. The preset initial model includes an image segmentation initial module and a multi-head attention initial module.
[0100] Acquire multiple third font recognition sample images, wherein the multiple third font recognition sample images include annotation data of multiple second font types, and the first font type may be the same as or different from the second font type;
[0101] The multiple third font recognition sample images are used as the second training dataset to train the first undetermined model, so as to obtain the preset font recognition model.
[0102] The above technical solutions can effectively avoid the problem of difficulty in obtaining training data during the training process of font recognition models in related technologies. They can achieve the recognition of font types of multiple categories and languages without the need for real labeled data. Furthermore, the meta-learning training mechanism can obtain a preset font recognition model that can be quickly expanded for newly added font types in the future.
[0103] The following is for reference. Figure 6This diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0104] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0105] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0106] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0107] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0108] In some implementations, communication can be conducted using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can be interconnected with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0109] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0110] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire an image to be recognized, the image to be recognized including target text; input the image to be recognized into a preset font recognition model, so that the preset font recognition model outputs the font type corresponding to the target text; wherein the preset font recognition model is used to divide the image to be recognized into multiple sub-images, and acquire a first image feature corresponding to each sub-image, determine a second image feature corresponding to the image to be recognized based on the first image feature corresponding to each sub-image in the image to be recognized, the second image feature including the contextual association features of each sub-image in the image to be recognized with other sub-images, and determine the font type corresponding to the target text based on the second image feature.
[0111] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0113] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the module itself; for example, an acquisition module can also be described as "acquiring an image to be recognized, wherein the image to be recognized includes target text".
[0114] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0115] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0116] According to one or more embodiments of this disclosure, Example 1 provides a font recognition method, the method comprising:
[0117] Obtain an image to be identified, wherein the image to be identified includes target text;
[0118] The image to be recognized is input into a preset font recognition model so that the preset font recognition model outputs the font type corresponding to the target text.
[0119] The preset font recognition model is used to divide the image to be recognized into multiple sub-images, obtain a first image feature corresponding to each sub-image, determine a second image feature corresponding to the image to be recognized based on the first image feature corresponding to each sub-image in the image to be recognized, the second image feature includes the contextual association features between each sub-image in the image to be recognized and other sub-images, and determine the font type corresponding to the target text based on the second image feature.
[0120] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein determining the font type corresponding to the target text based on the second image features includes:
[0121] Obtain the Euclidean distance between the second image feature and the representative feature of each of the multiple recognizable fonts corresponding to the preset font recognition model, so as to obtain multiple Euclidean distances between the image to be recognized and the representative features of the multiple recognizable fonts;
[0122] Determine the minimum target distance from among the multiple Euclidean distances;
[0123] The font type corresponding to the target text in the image to be identified is determined based on the target distance.
[0124] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein determining the font type corresponding to the target text in the image to be identified based on the target distance includes:
[0125] If the target distance is less than a preset distance threshold, the target font type corresponding to the target representative feature used to calculate the target distance is taken as the font type of the target text.
[0126] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 2, wherein determining the font type corresponding to the target text in the image to be identified based on the target distance includes:
[0127] If the target distance is determined to be greater than or equal to a preset distance threshold, the font type corresponding to the target text is determined to be a newly added font.
[0128] According to one or more embodiments of this disclosure, Example 5 provides the method described in Example 1, wherein the preset font recognition model is further used for:
[0129] Obtain multiple first font recognition sample images of the target newly added font, wherein the first font recognition sample images include a specified text sample of the target newly added font;
[0130] Obtain the second image features corresponding to each of the first font recognition sample images;
[0131] Obtain the target mean of multiple second image features corresponding to the multiple first font recognition sample images, use the target mean as the target representative feature corresponding to the target newly added font, and store the target representative feature.
[0132] According to one or more embodiments of this disclosure, Example 6 provides the method described in Example 5, wherein obtaining multiple first font recognition sample images of the target newly added font includes:
[0133] Obtain a specified text sample of the target newly added font from a preset font corpus;
[0134] Obtain the target background image from the preset background library;
[0135] The specified text sample and the target background image are combined to form the first font recognition sample image.
[0136] According to one or more embodiments of this disclosure, Example 7 provides the method described in any one of Examples 1-6, wherein the preset font recognition model is trained in the following manner:
[0137] Acquire multiple second font recognition sample images, which include annotation data for multiple first font types;
[0138] The plurality of second font recognition sample images are used as the first training dataset to pre-train a preset initial model to obtain a first undetermined model. The preset initial model includes an image segmentation initial module and a multi-head attention initial module.
[0139] Acquire multiple third font recognition sample images, wherein the multiple third font recognition sample images include annotation data of multiple second font types, and the first font type is the same as or different from the second font type;
[0140] The multiple third font recognition sample images are used as the second training dataset to train the first undetermined model, so as to obtain the preset font recognition model.
[0141] According to one or more embodiments of this disclosure, Example 8 provides a font recognition device, the device comprising:
[0142] The acquisition module is configured to acquire an image to be recognized, wherein the image to be recognized includes target text;
[0143] The determining module is configured to input the image to be recognized into a preset font recognition model, so that the preset font recognition model outputs the font type corresponding to the target text;
[0144] The preset font recognition model is used to divide the image to be recognized into multiple sub-images, obtain a first image feature corresponding to each sub-image, determine a second image feature corresponding to the image to be recognized based on the first image feature corresponding to each sub-image in the image to be recognized, the second image feature includes the contextual association features between each sub-image in the image to be recognized and other sub-images, and determine the font type corresponding to the target text based on the second image feature.
[0145] According to one or more embodiments of the present disclosure, Example 9 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-7 above.
[0146] According to one or more embodiments of this disclosure, Example 10 provides an electronic device, including:
[0147] A storage device on which computer programs are stored;
[0148] A processing device for executing the computer program in the storage device to implement the steps of any of the methods described in Examples 1-7 above.
[0149] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0150] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0151] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A font recognition method, characterized in that, The method includes: Obtain an image to be identified, wherein the image to be identified includes target text; The image to be recognized is input into a preset font recognition model so that the preset font recognition model outputs the font type corresponding to the target text. The preset font recognition model is used to divide the image to be recognized into multiple sub-images, obtain a first image feature corresponding to each sub-image, determine a second image feature corresponding to the image to be recognized based on the first image feature corresponding to each sub-image in the image to be recognized, the second image feature includes the contextual association features between each sub-image in the image to be recognized and other sub-images, and determine the font type corresponding to the target text based on the second image feature. The multiple sub-images are obtained by dividing based on image texture, and the texture in the sub-images has continuity.
2. The method according to claim 1, characterized in that, Determining the font type corresponding to the target text based on the second image features includes: Obtain the Euclidean distance between the second image feature and the representative feature of each of the multiple recognizable fonts corresponding to the preset font recognition model, so as to obtain multiple Euclidean distances between the image to be recognized and the representative features of the multiple recognizable fonts; Determine the minimum target distance from among the multiple Euclidean distances; The font type corresponding to the target text in the image to be identified is determined based on the target distance.
3. The method according to claim 2, characterized in that, Determining the font type corresponding to the target text in the image to be identified based on the target distance includes: If the target distance is less than a preset distance threshold, the target font type corresponding to the target representative feature used to calculate the target distance is taken as the font type of the target text.
4. The method according to claim 2, characterized in that, Determining the font type corresponding to the target text in the image to be recognized based on the target distance includes: If the target distance is determined to be greater than or equal to a preset distance threshold, the font type corresponding to the target text is determined to be a newly added font.
5. The method according to claim 1, characterized in that, The preset font recognition model is also used for: Obtain multiple first font recognition sample images of the target newly added font, wherein the first font recognition sample images include a specified text sample of the target newly added font; Obtain the second image features corresponding to each of the first font recognition sample images; Obtain the target mean of multiple second image features corresponding to the multiple first font recognition sample images, use the target mean as the target representative feature corresponding to the target newly added font, and store the target representative feature.
6. The method according to claim 5, characterized in that, The acquisition of multiple first font recognition sample images of the target newly added font includes: Obtain a specified text sample of the target newly added font from a preset font corpus; Obtain the target background image from the preset background library; The specified text sample and the target background image are combined to form the first font recognition sample image.
7. The method according to any one of claims 1-6, characterized in that, The preset font recognition model is trained in the following way: Acquire multiple second font recognition sample images, which include annotation data for multiple first font types; The plurality of second font recognition sample images are used as the first training dataset to pre-train a preset initial model to obtain a first undetermined model. The preset initial model includes an image segmentation initial module and a multi-head attention initial module. Acquire multiple third font recognition sample images, wherein the multiple third font recognition sample images include annotation data of multiple second font types, and the first font type is the same as or different from the second font type; The multiple third font recognition sample images are used as the second training dataset to train the first undetermined model, so as to obtain the preset font recognition model.
8. A font recognition device, characterized in that, The device includes: The acquisition module is configured to acquire an image to be recognized, wherein the image to be recognized includes target text; The determining module is configured to input the image to be recognized into a preset font recognition model, so that the preset font recognition model outputs the font type corresponding to the target text; The preset font recognition model is used to divide the image to be recognized into multiple sub-images, obtain a first image feature corresponding to each sub-image, determine a second image feature corresponding to the image to be recognized based on the first image feature corresponding to each sub-image in the image to be recognized, the second image feature includes the contextual association features between each sub-image in the image to be recognized and other sub-images, and determine the font type corresponding to the target text based on the second image feature. The multiple sub-images are obtained by dividing based on image texture, and the texture in the sub-images has continuity.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method described in any one of claims 1-7.
10. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-7.