A font recognition method, device and equipment
By integrating glyphs and font features in the font recognition process, the problem of inaccurate font recognition in the prior art is solved, and a more accurate font recognition effect is achieved.
Patent Information
- Application Number
- CN202410774122.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-06-17
AI Technical Summary
The existing font recognition algorithm ignores the impact of glyphs on font recognition results, resulting in weakening of font features and inaccurate recognition.
By obtaining the picture to be recognized for glyph recognition and font feature extraction, feature fusion processing is performed by combining the glyph recognition results with the font feature extraction results, and font feature encoding is performed to output the final font recognition results.
The accuracy of font recognition is enhanced. By considering the characteristics of the glyph, the sensitivity to glyph is weakened, making font recognition more accurate.
Smart Images

Figure CN118675186B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of character information processing, and particularly to a font recognition method, device and equipment. Background Art
[0002] Layout analysis and page layout restoration play a very important role in realizing paperless office and improving user experience. As a component of the page layout, accurate recognition of fonts is of great significance. With the development of computer and deep learning technologies, automated font recognition has become a current hot research direction.
[0003] Existing font recognition algorithms first extract text image features and send them into a neural network to parse and obtain the font recognition result. However, such algorithms to a certain extent ignore the influence of glyphs on the font recognition result. In the neural network model, features from multiple angles such as fonts and glyphs need to be comprehensively considered, and the feature extraction is extensive, resulting in the weakening of font features. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a font recognition method, device and equipment to make font recognition more accurate.
[0005] To solve the above technical problem, the technical solution of the present invention is as follows:
[0006] A font recognition method includes:
[0007] Obtain a to-be-recognized picture containing text;
[0008] Perform glyph recognition on the to-be-recognized picture to obtain a glyph recognition result;
[0009] Extract font features from the to-be-recognized picture to obtain a font feature extraction result;
[0010] Perform feature fusion processing on the font feature extraction result and the glyph recognition result to obtain a fusion processing result;
[0011] Perform font feature encoding processing on the fusion processing result to obtain the font recognition result of the text in the to-be-recognized picture.
[0012] Optionally, performing glyph recognition on the to-be-recognized picture to obtain a glyph recognition result includes:
[0013] Perform glyph recognition on the text in the to-be-recognized picture through a glyph recognition model to obtain a glyph recognition result, and the glyph recognition model is trained according to a text training set in a picture set.
[0014] Optionally, performing glyph recognition on the text in the to-be-recognized picture through a glyph recognition model to obtain a glyph recognition result includes:
[0015] Extract the glyph features of the image to be recognized through the glyph feature extraction layer of the glyph recognition model to obtain the glyph feature extraction result;
[0016] Perform glyph recognition processing on the glyph feature extraction result through the glyph recognition layer of the glyph recognition model, and output a first output vector, where the first output vector is the glyph recognition result.
[0017] Optionally, perform font feature extraction on the image to be recognized to obtain the font feature extraction result, including:
[0018] Extract the font features of the image to be recognized through the font feature extraction layer of the font recognition model, and output a third output vector, where the third output vector is the font feature extraction result; the font feature extraction result includes: at least one feature map containing font features of all the characters in the image to be recognized; the font recognition model is trained according to the font training set in the image set.
[0019] Optionally, perform feature fusion processing on the font feature extraction result and the glyph recognition result to obtain the fusion processing result, including:
[0020] Perform convolution processing on the first output vector to obtain a second output vector;
[0021] Through the feature fusion layer of the font recognition model, splice the feature data of the target dimension in the third output vector and the second output vector to obtain the fusion processing result.
[0022] Optionally, perform font feature encoding processing on the fusion processing result to obtain the font recognition result of the characters in the image to be recognized, including:
[0023] Recognize the font type in the fusion processing result through the font feature encoding layer of the font recognition model to obtain the font recognition result of the characters in the image to be recognized.
[0024] Optionally, recognize the font type in the fusion processing result through the font feature encoding layer of the font recognition model to obtain the font recognition result of the characters in the image to be recognized, including:
[0025] Perform dimensionality reduction processing on the fusion processing result to obtain a fourth output vector;
[0026] Input the fourth output vector into the convolutional layer of the font feature encoding layer, perform classification convolution on the fusion processing result, and output the font recognition result of the characters in the image to be recognized.
[0027] The present invention also provides a font recognition device, including:
[0028] An acquisition module, configured to acquire a to-be-recognized picture containing characters;
[0029] A processing module, configured to perform glyph recognition on the to-be-recognized picture to obtain a glyph recognition result; perform font feature extraction on the to-be-recognized picture to obtain a font feature extraction result; perform feature fusion processing on the font feature extraction result and the glyph recognition result to obtain a fusion processing result; and perform font feature encoding processing on the fusion processing result to obtain a font recognition result of the characters in the to-be-recognized picture.
[0030] The present invention also provides a computing device, including: a processor and a memory storing a computer program, and when the computer program is run by the processor, the method as described above is executed.
[0031] The present invention also provides a computer-readable storage medium storing instructions, and when the instructions are run on a computer, the computer is made to execute the method as described above.
[0032] The above solution of the present invention has at least the following beneficial effects:
[0033] In the above solution of the present invention, by acquiring a to-be-recognized picture containing characters; performing glyph recognition on the to-be-recognized picture to obtain a glyph recognition result; performing font feature extraction on the to-be-recognized picture to obtain a font feature extraction result; performing feature fusion processing on the font feature extraction result and the glyph recognition result to obtain a fusion processing result; and performing font feature encoding processing on the fusion processing result to obtain a font recognition result of the characters in the to-be-recognized picture. When performing font recognition, the glyph features are considered, but the glyph features are weakened, and the sensitivity to font features is enhanced, so that the font recognition is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic flowchart of the font recognition method according to an embodiment of the present invention;
[0035] Figure 2 is a schematic diagram of the font recognition model according to an embodiment of the present invention;
[0036] Figure 3 is a schematic diagram of the glyph recognition model according to an embodiment of the present invention;
[0037] Figure 4 is a schematic diagram of font classification according to an embodiment of the present invention;
[0038] Figure 5 is a schematic structural diagram of the font recognition device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to enable the scope of the present invention to be fully communicated to those skilled in the art.
[0040] like Figure 1 As shown, an embodiment of the present invention provides a font recognition method, comprising:
[0041] Step 11, obtaining a picture to be identified containing text;
[0042] Here, if Figure 2 or Figure 3 The picture to be identified shown in the figure contains text: Today's weather is really good; of course, the text here can also be text in other languages, or other characters with glyph features, etc.;
[0043] Step 12, performing font recognition on the image to be recognized to obtain a font recognition result;
[0044] Here, the character shape refers to: the appearance of a single character, such as the characters "国" and "法" have different appearances;
[0045] Step 13, extracting font features from the image to be identified to obtain a font feature extraction result;
[0046] Here, font refers to: the external form of the text, such as Figure 4 The Songti, bold, or Kaiti “国” characters shown in the figure are all different. Each font has its own unique appearance. The same character “国” will have different stroke thickness and roundness when rendered in different fonts.
[0047] Step 14, performing feature fusion processing on the font feature extraction result and the character shape recognition result to obtain a fusion processing result;
[0048] Step 15, performing font feature encoding processing on the fusion processing result to obtain the font recognition result of the text in the image to be recognized.
[0049] This embodiment of the present invention obtains a fusion processing result by performing feature fusion processing on the font feature extraction result and the glyph recognition result; and obtains a font recognition result of the text in the image to be recognized by performing font feature encoding processing on the fusion processing result; thereby, a font recognition result that takes glyph features into consideration can be output, making font recognition more accurate.
[0050] In this application, by extracting glyph features and font features, and fusing the extracted glyph features and font features. Then, the fused feature map is classified, that is, feature encoding, and finally, the font recognition result considering the glyph features is output. During the font recognition process, the glyph features are separated, so that the font recognition can more effectively focus on the font features, thereby weakening the sensitivity to glyphs and enhancing the sensitivity to fonts, making the font recognition more accurate.
[0051] In an alternative embodiment of the present invention, step 12 may include:
[0052] Step 121, performing glyph recognition on the text in the to-be-recognized picture through a glyph recognition model, obtaining a glyph recognition result, where the glyph recognition model is trained according to the text training set in the picture set.
[0053] In this embodiment, as Figure 3 shown, the glyph recognition model includes: a glyph feature extraction layer, configured to receive the input to-be-recognized picture, and perform glyph feature extraction on the text in the to-be-recognized picture respectively, and output a first sub-output vector, where the first sub-output vector includes the preliminary glyph features of multiple characters. For example, input a picture with a pixel size of 960×32, and the text in the picture is: The weather is really nice today. The glyph feature extraction layer will output the preliminary glyph features of each character in The weather is really nice today;
[0054] a function layer, configured to perform deletion, merging, or rearrangement of relevant vectors on the first sub-output vector of the preliminary glyph features of each character output by the glyph feature extraction layer, and output a second sub-output vector;
[0055] a glyph feature encoding module, configured to receive the second sub-output vector, and process the second sub-output vector by inputting it into a recurrent neural network, obtaining a third sub-output vector, and further inputting the third sub-output vector into a convolutional layer for processing, and outputting a first output vector, that is, the glyph recognition result;
[0056] The data processing process of this glyph recognition model is as follows:
[0057] Obtain a target picture containing text;
[0058] Perform glyph feature extraction on the text in the target picture, obtaining a first sub-output vector, where the first sub-output vector includes the preliminary glyph features of multiple characters in the picture;
[0059] Input the preliminary glyph features into the function layer for vector adjustment processing, obtaining a second sub-output vector; here, performing vector adjustment processing is to meet the data processing requirements of the glyph feature encoding module;
[0060] Input the second sub-output vector into the recurrent neural network of the glyph feature encoding module for processing to obtain a third sub-output vector;
[0061] Input the third sub-output vector into the convolutional layer of the glyph feature encoding module for processing to obtain the output result of the glyph recognition model, that is, the first output vector.
[0062] In this embodiment, the recurrent neural network can be an LSTM (Long Short-Term Memory Recurrent Neural Network), etc.
[0063] During the training process of the glyph recognition model, the length of the character set used is 8192, that is, the glyph recognition model can recognize these 8192 characters, and the final output classification of the model is 8192 classes, representing the probability distribution result of the recognized characters of the model in these 8192 classifications.
[0064] Specifically, step 121 may include:
[0065] Step 1211, extract the glyph features of the picture to be recognized through the glyph feature extraction layer of the glyph recognition model to obtain the glyph feature extraction result;
[0066] Step 1212, perform glyph recognition processing on the glyph feature extraction result through the glyph recognition layer of the glyph recognition model, and output the first output vector, where the first output vector is the glyph recognition result.
[0067] In this embodiment, the glyph recognition model is divided into two parts: glyph feature extraction and glyph feature encoding. For glyph feature extraction, its features are implicitly included in the weight parameters and output results of the neural network, while for glyph feature encoding, the extracted glyph features are transformed into corresponding text characters, so as to ensure that the effective character information is extracted by this network structure. Of course, it can be understood that after glyph feature extraction, the feature vector is adjusted through a function layer. The function layer includes at least one of the following functions, such as the squeeze function, which is used to delete a dimension in the tensor, and the permute function, which is used to rearrange a specified vector into an array. The glyph features are assisted and adjusted through the function layer to make the glyph feature encoding and the recognition of glyphs more accurate.
[0068] The specific network structure and parameters of the glyph feature extraction layer are as follows:
[0069] First, it is determined that the input is 1x3x32x960, that is, a three-channel picture with a width of 960 pixels and a height of 32 pixels, and the batch_size is 1;
[0070] The initial part of the network consists of three convolutional modules. The convolutional kernels are all of size 3x3, with a stride of 1 and a padding of 1. The ReLU activation function is used. The numbers of convolutional kernels are 32, 32, and 64 respectively, and the output vector is 1x64x32x960;
[0071] After passing through the max pooling layer with a stride of 2 and a kernel size of 2, the output vector is 1x64x16x480;
[0072] The output vector after passing through the max pooling layer then enters the residual calculation module BasicBlock, which consists of three convolutional layers. The parameter settings of the first two convolutional layers are both 64 convolutional kernels of size 3x3, with a stride of 1 and a padding of 1. The convolutional kernel size of the third convolutional layer is 1 while other settings remain the same. The sequential operations of the vector passing through the first two convolutional layers are added to the operations of the third convolutional layer to implement the residual operation, and then the output is realized through the ReLU activation function.
[0073] Specifically, the above residual calculation module is distributed in the following four parts, and the specific settings of each part are as follows:
[0074] The first part contains 3 BasicBlock modules, and the output is 1x64x16x480;
[0075] The second part contains 4 BasicBlock modules. The stride settings of the first convolutional layer and the third convolutional layer of the first BasicBlock module are (2, 1), and the strides of all other convolutional layers are 1. The number of convolutional kernels is 128, and other settings are the same as those of the first part. Then the output vector is 1x128x8x480;
[0076] The third part contains 6 BasicBlock modules. The stride settings of the first convolutional layer and the third convolutional layer of the first BasicBlock module are (2, 1), and the strides of all other convolutional layers are 1. The number of convolutional kernels is 256, and other settings are the same as those of the first part. Then the output vector is 1x256x4x480;
[0077] The fourth part contains 3 BasicBlock modules. The stride settings of the first convolutional layer and the third convolutional layer of the first BasicBlock module are (2, 1), and the strides of all other convolutional layers are 1. The number of convolutional kernels is 512, and other settings are the same as those of the first part. Then the output vector is 1x512x2x480;
[0078] The output vector of the fourth part passes through the max pooling layer with a stride of 2 and a kernel_size of 2, and then the output vector is 1x512x1x240;
[0079] After passing through the convolutional layer, 512 convolutional kernels with a size of (1, 2) and a stride of (1, 2) are set, and the output vector is 1x512x1x120. This vector is dimensionally rearranged to obtain the first sub-output vector. The above-mentioned first sub-output vector of 1×512×1×120 is output after glyph feature extraction, where the first 1 represents the batch size, that is, the number of samples selected for one training is 1, 512 represents 512 channels, the second 1 represents the height of the feature map, and 120 represents the width of the feature map. Since it is obtained through actual experience testing, usually each Chinese character corresponds to a width of about 8 pixels. For an input image with a width of 960 pixels, according to experience, the width of the output feature map is 120, corresponding to 120 characters. The output vector is adjusted in dimension through the function layer to obtain an output vector of 120×1×512.
[0080] The output vector is input into the glyph feature encoding part for feature classification. In this application, considering the context semantic association between texts, a long short-term memory network is added to the glyph feature encoding part. And a convolutional layer is used to replace the fully connected layer in the prior art to achieve the technical effect of reducing the number of parameters while improving the operation speed. Specifically, the output vector of 120×1×512 is first input into the long short-term memory network layer. The long short-term memory network layer includes multiple recurrent neural network units. Each neural network unit receives the cell state, hidden state output by the previous neural network unit, and the input vector input to this neural network unit at the current moment. After calculations by the forget gate, update gate, and output gate inside the neural network unit, the hidden state and cell state of the current neural network unit are output, and the hidden state and cell state are input into the next neural network unit. The neural network units are sequentially connected to form a neural network link, and finally an output vector of 120×1×256 is obtained.
[0081] The output vector of 120×1×256 is input into the convolutional layer for feature classification. According to the number of common Chinese characters, a character set of length 8192 is used in this application, that is, the final glyph output classification is 8192 classes.
[0082] First, the third sub-output vector of 120×1×256 is dimensionally expanded to 1x120×1×256, and then the dimensions are swapped to 1x256×1×120. Note that the 1 in the first position and the 1 in the third position here are swapped in dimension because although the numerical values look the same, their actual meanings are different. The source code is:
[0083] x.permute(2,3,0,1) means arranging the original vector dimensions in this order, and the output becomes 1x256×1×120;
[0084] Then, through the convolutional layer with 8,192 convolutional kernels and a kernel size of 1, the output vector obtained is 1x8,192x1x120. After compressing the 1 in the third dimension and then swapping the data dimensions, it becomes 1x120x8,192.
[0085] In the training process of the glyph recognition model of the present application, the input is a picture with an aspect ratio of 960x32, and each picture contains a text strip. To increase the richness and diversity of the training samples, a variety of data augmentation methods are used, including: rotation, cropping, warping, bolding, salt-and-pepper noise, blurring, contrast transformation, text stretching, and more than a dozen other methods, and text information from various industries is used to construct a training set and a corresponding test set with a data volume of millions. The loss function uses the connectionist temporal classification loss function, the initial learning rate is set to 0.1, and a total of 100 iterations are performed. The learning rate is decayed every 30 times, decaying to 0.1 times the original value, observing the decrease of the loss function, and conducting tests. Finally, an accuracy rate of 98% is achieved on the test set, indicating that the glyph recognition model can better distinguish the text content, and also indirectly indicating that the glyph features of the text can be extracted relatively accurately and encoded.
[0086] In an optional embodiment of the present invention, step 13 may include:
[0087] Step 131, extracting font features from the to-be-recognized picture through the font feature extraction layer of the font recognition model, and outputting a third output vector, where the third output vector is the font feature extraction result; the font feature extraction result includes: at least one feature map containing font features of all the characters in the to-be-recognized picture; the font recognition model is trained according to the font training set in the picture set.
[0088] As Figure 2 shown, the font recognition model includes:
[0089] The font feature extraction layer is used to receive the input to-be-recognized picture and extract font features from the characters in the to-be-recognized picture respectively, and output a third output vector, where the third output vector includes preliminary font features of multiple characters. For example, when inputting a picture with a pixel size of 960×32, which contains the text: The weather is really nice today, this layer will output the preliminary font features of each character in The weather is really nice today. It should be noted that when the font recognition model and the above-mentioned glyph recognition model recognize the characters in the picture, the same to-be-recognized picture is input.
[0090] A feature fusion layer, which is used to receive the third output vector and the second output vector finally output by the glyph recognition model, and perform feature fusion processing on the second output vector and the third output vector to output a fourth output vector. It should be noted that the second output vector and the third output vector include the recognition results of the same text and have the same number of dimensions for each dimension;
[0091] A font feature encoding module, which is used to receive the fourth output vector and perform convolution processing on the fourth output vector to obtain a fifth output vector, that is, the final font recognition result;
[0092] The data processing process of this font recognition model is as follows:
[0093] Obtain a target picture containing text;
[0094] Extract font features of the text in the target picture to obtain a third output vector, which contains the preliminary font features of multiple texts in the picture;
[0095] Input the preliminary font features into the feature fusion layer, perform feature fusion processing of vectors with the second output vector finally output by the glyph recognition model, obtain a fusion processing result, and perform dimensionality reduction processing on the fusion processing result to output a fourth output vector;
[0096] Input the fourth output vector into the convolution layer of the font feature encoding module for processing to obtain the final font recognition result.
[0097] In this embodiment, the font feature extraction layer still adopts the form of a convolutional neural network, sets multiple convolutional layers, the final output channel is 512, and the feature map size is 1×120. Since the architectures of the font feature extraction layer and the glyph feature extraction layer are the same, the detailed calculation process will not be elaborated here.
[0098] In an alternative embodiment of the present invention, step 14 may include:
[0099] Step 141, perform convolution processing on the first output vector to obtain a second output vector;
[0100] Step 142, through the feature fusion layer of the font recognition model, splice the feature data of the target dimension in the third output vector and the second output vector to obtain a fusion processing result.
[0101] In this embodiment, the first output vector of the glyph recognition model, which is 1×120×8192, serves as an input to the font recognition model. Here, 1 represents the batch size, that is, the number of pictures, 120 is the total number of predicted characters, and 8192 corresponds to the length of the character set. After dimension elevation, convolution, etc. to extract features, the dimension of the output vector is converted to 1x512x1 x120. The third output vector of the font feature extraction layer and the converted second output vector are merged in the second dimension to obtain a vector with a dimension of 1x1024x1 x120, which is the fusion processing result of the font feature and the glyph feature.
[0102] In an alternative embodiment of the present invention, step 15 may include:
[0103] Step 151, perform dimensionality reduction processing on the fusion processing result to obtain a fourth output vector;
[0104] Step 152, input the fourth output vector into the convolutional layer of the font feature encoding layer, perform classification convolution on the fusion processing result, and output the font recognition result of the text in the picture to be recognized.
[0105] Specifically, after reducing the dimension of the fusion processing result and inputting it into the convolutional layer, classification convolution is performed on the fusion processing result through a preset number of filters and the convolution results are concatenated to obtain a target feature vector with a preset length. The length of the target feature vector corresponds to the number of font types; the target feature vector includes the font recognition result of the text in the picture to be recognized.
[0106] In this application, the feature encoding part of the font recognition model uses a reshape layer, a convolutional layer, etc. to recognize the font. Specifically, the number of convolution kernels in the convolutional layer is 64, corresponding to 64 fonts, and the convolution kernel size is 1x1. After classification convolution processing by the convolutional layer, a prediction result of 1x64x1x120 is finally obtained, and the result is converted to 1x64x120, that is, the font type of each character is predicted.
[0107] In this application, the training data set of the font recognition model is pictures with an aspect ratio of 960x32. Different from the glyph recognition task, the font recognition task mainly recognizes font information. Therefore, during the data set generation process, data expansion methods such as text stretching, distortion, and bolding may change the font features themselves, so such data expansion methods should be avoided as much as possible. Finally, a data set of millions is generated, including 64 common font categories. The cross-entropy loss function is used in training, the parameters of the glyph recognition model are frozen, the initial learning rate is set to 0.01, and a total of 10 iterations are performed to finally make the project available.
[0108] In this application, font recognition is divided into two parts: glyph recognition and font recognition. The glyph features and font features extracted by convolution are fused to obtain a font recognition result that takes into account the glyph features, making the font recognition result more accurate.
[0109] As Figure 5 shown, an embodiment of the present invention further provides a font recognition device 50, including:
[0110] An acquisition module 51, configured to acquire a to-be-recognized picture containing text;
[0111] A processing module 52, configured to perform glyph recognition on the to-be-recognized picture to obtain a glyph recognition result; perform font feature extraction on the to-be-recognized picture to obtain a font feature extraction result; perform feature fusion processing on the font feature extraction result and the glyph recognition result to obtain a fusion processing result; perform font feature encoding processing on the fusion processing result to obtain a font recognition result of the text in the to-be-recognized picture.
[0112] Optionally, performing glyph recognition on the to-be-recognized picture to obtain a glyph recognition result includes:
[0113] Performing glyph recognition on the text in the to-be-recognized picture through a glyph recognition model to obtain a glyph recognition result, where the glyph recognition model is trained according to a text training set in a picture set.
[0114] Optionally, performing glyph recognition on the text in the to-be-recognized picture through a glyph recognition model to obtain a glyph recognition result includes:
[0115] Performing glyph feature extraction on the to-be-recognized picture through a glyph feature extraction layer of the glyph recognition model to obtain a glyph feature extraction result;
[0116] Performing glyph recognition processing on the glyph feature extraction result through a glyph recognition layer of the glyph recognition model, and outputting a first output vector, where the first output vector is the glyph recognition result.
[0117] Optionally, performing font feature extraction on the to-be-recognized picture to obtain a font feature extraction result includes:
[0118] Performing font feature extraction on the to-be-recognized picture through a font feature extraction layer of a font recognition model, and outputting a third output vector, where the third output vector is the font feature extraction result; the font feature extraction result includes: at least one feature map containing font features of all the text in the to-be-recognized picture; the font recognition model is trained according to a font training set in a picture set.
[0119] Optionally, perform feature fusion processing on the font feature extraction result and the glyph recognition result to obtain a fusion processing result, including:
[0120] Perform convolution processing on the first output vector to obtain a second output vector;
[0121] Through the feature fusion layer of the font recognition model, splice the feature data of the target dimension in the third output vector and the second output vector to obtain a fusion processing result.
[0122] Optionally, perform font feature encoding processing on the fusion processing result to obtain a font recognition result of the text in the to-be-recognized picture, including:
[0123] Through the font feature encoding layer of the font recognition model, recognize the font type in the fusion processing result to obtain a font recognition result of the text in the to-be-recognized picture.
[0124] Optionally, through the font feature encoding layer of the font recognition model, recognize the font type in the fusion processing result to obtain a font recognition result of the text in the to-be-recognized picture, including:
[0125] Perform dimensionality reduction processing on the fusion processing result to obtain a fourth output vector;
[0126] Input the fourth output vector into the convolution layer of the font feature encoding layer to perform classification convolution on the fusion processing result, and output a font recognition result of the text in the to-be-recognized picture.
[0127] It should be noted that this device corresponds to the above method, and all implementation manners in the above method embodiments are applicable to the embodiments of this device and can also achieve the same technical effects.
[0128] An embodiment of the present invention further provides a computing device, including: a processor and a memory storing a computer program. When the computer program is run by the processor, it executes the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can also achieve the same technical effects.
[0129] An embodiment of the present invention further provides a computer-readable storage medium storing instructions. When the instructions are run on a computer, the computer is made to execute the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can also achieve the same technical effects.
[0130] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0131] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0132] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0133] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0134] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0135] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0136] In addition, it should be noted that in the device and method of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations shall be regarded as equivalent solutions of the present invention. Moreover, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel or independently of each other. For those of ordinary skill in the art, it is understandable that all or any steps or components of the method and device of the present invention can be implemented in any computing device (including processors, storage media, etc.) or a network of computing devices in the form of hardware, firmware, software, or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present invention.
[0137] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing device. The computing device can be a well-known general-purpose device. Therefore, the object of the present invention can also be achieved only by providing a program product containing program code for implementing the method or device. That is to say, such a program product also constitutes the present invention, and a storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future. It should also be noted that in the device and method of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations shall be regarded as equivalent solutions of the present invention. Moreover, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel or independently of each other.
[0138] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A font recognition method, characterized in that, Comprising: Obtaining a to-be-recognized picture containing text; Performing glyph recognition on the to-be-recognized picture to obtain a glyph recognition result; Extracting font features from the to-be-recognized picture to obtain a font feature extraction result; Performing feature fusion processing on the font feature extraction result and the glyph recognition result to obtain a fusion processing result; Performing font feature encoding processing on the fusion processing result to obtain a font recognition result of the text in the to-be-recognized picture; Wherein, performing glyph recognition on the to-be-recognized picture to obtain a glyph recognition result includes: Performing glyph feature extraction on the to-be-recognized picture through a glyph feature extraction layer of a glyph recognition model to obtain a glyph feature extraction result; the glyph feature extraction layer includes: three convolutional modules, one max pooling layer, and four parts of residual calculation modules, and each part of the residual calculation module includes multiple residual calculation modules BasicBlock; Performing glyph recognition processing on the glyph feature extraction result through a glyph recognition layer of the glyph recognition model, and outputting a first output vector, where the first output vector is the glyph recognition result; the glyph recognition layer includes: a recurrent neural network and a convolutional layer; Wherein, extracting font features from the to-be-recognized picture to obtain a font feature extraction result includes: Performing font feature extraction on the to-be-recognized picture through a font feature extraction layer of a font recognition model, and outputting a third output vector, where the third output vector is the font feature extraction result; the network architecture of the font feature extraction layer is the same as that of the glyph feature extraction layer; the font feature extraction result includes: at least one feature map containing font features of all the text in the to-be-recognized picture; the font recognition model is trained according to a font training set in a picture set; Wherein, performing feature fusion processing on the font feature extraction result and the glyph recognition result to obtain a fusion processing result includes: Performing convolutional processing on the first output vector to obtain a second output vector; Through a feature fusion layer of the font recognition model, splicing the feature data of the target dimension in the third output vector and the second output vector to obtain a fusion processing result.
2. The font recognition method according to claim 1, wherein Performing font feature encoding processing on the fusion processing result to obtain a font recognition result of the text in the to-be-recognized picture includes: Recognizing the font type in the fusion processing result through a font feature encoding layer of the font recognition model to obtain a font recognition result of the text in the to-be-recognized picture.
3. The font recognition method according to claim 2, wherein Recognizing the font type in the fusion processing result through a font feature encoding layer of the font recognition model to obtain a font recognition result of the text in the to-be-recognized picture includes: Performing dimensionality reduction processing on the fusion processing result to obtain a fourth output vector; Inputting the fourth output vector into a convolutional layer of the font feature encoding layer, performing classification convolution on the fusion processing result, and outputting a font recognition result of the text in the to-be-recognized picture.
4. A font recognition device, characterized in that, Comprising: An acquisition module, configured to acquire a to-be-recognized picture containing text; A processing module, configured to perform glyph recognition on the to-be-recognized picture to obtain a glyph recognition result; Extract font features from the picture to be recognized to obtain a font feature extraction result; Perform feature fusion processing on the font feature extraction result and the glyph recognition result to obtain a fusion processing result; Perform font feature encoding processing on the fusion processing result to obtain a font recognition result of the text in the picture to be recognized; Among them, performing glyph recognition on the picture to be recognized to obtain a glyph recognition result includes: Performing glyph feature extraction on the picture to be recognized through the glyph feature extraction layer of the glyph recognition model to obtain a glyph feature extraction result; the glyph feature extraction layer includes: three convolutional modules, a max pooling layer, and four parts of residual calculation modules, and each part of the residual calculation module includes multiple residual calculation modules BasicBlock; Performing glyph recognition processing on the glyph feature extraction result through the glyph recognition layer of the glyph recognition model, and outputting a first output vector, where the first output vector is the glyph recognition result; the glyph recognition layer includes: a recurrent neural network and a convolutional layer; Among them, extracting font features from the picture to be recognized to obtain a font feature extraction result includes: Performing font feature extraction on the picture to be recognized through the font feature extraction layer of the font recognition model, and outputting a third output vector, where the third output vector is the font feature extraction result; the network architecture of the font feature extraction layer is the same as the network architecture of the glyph feature extraction layer; the font feature extraction result includes: at least one feature map containing font features of all the text in the picture to be recognized; the font recognition model is trained according to the font training set in the picture set; Among them, performing feature fusion processing on the font feature extraction result and the glyph recognition result to obtain a fusion processing result includes: Performing convolutional processing on the first output vector to obtain a second output vector; Through the feature fusion layer of the font recognition model, splicing the feature data of the target dimension in the third output vector and the second output vector to obtain a fusion processing result.
5. A computing device, characterized in that, Includes: A processor and a memory storing a computer program. When the computer program is run by the processor, it executes the method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, A storage instruction, when the instruction runs on a computer, causes the computer to execute the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Calligraphy font type and text content synchronous identification method
CN111104912A