A document font recognition method and device, computer equipment and a storage medium
Patent Information
- Application Number
- CN202611005678.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
该类方法的性能高度依赖OCR识别质量,当文档图像存在噪声、模糊或复杂背景时,OCR识别误差会被进一步放大,影响字体识别的整体准确性
本发明在传统HOG特征与SVM分类器的基础上,引入了基于SVM分类置信度差值的门控触发机制,即在低置信度样本上进一步利用笔画方向能量比进行二次判别。该方法能够有效识别SVM分类器在形近字体上的边界样本,并通过笔画方向能量比的辅助判别实现对相近字体的准确区分。与现有技术相比,本发明在不依赖大规模标注数据和复杂深度学习模型的条件下,显著提高了形近字体的识别准确率,减少了误分类现象,增强了系统的分类稳定性和可解释性,同时保持了传统方法实现复杂度低、工程调试方便的优势,可广泛应用于毕业论文格式审查、规范文档检测及自动化排版分析等场景。
Smart Images

Figure CN122821571A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image recognition, and specifically relates to a document font recognition method, apparatus, computer device, and storage medium. Background Technology
[0002] In the field of document image analysis and recognition, font recognition is a crucial step in document layout understanding and format checking. Existing font recognition methods mainly fall into three categories: methods based on hand-designed features and shallow classifiers, methods based on optical character recognition (OCR), and methods based on deep learning.
[0003] Among methods based on hand-designed features and shallow classifiers, the combination of Histogram of Oriented Gradients (HOG) features with a Support Vector Machine (SVM) classifier is a common approach. This type of method extracts HOG features from text region images to characterize the stroke direction and edge contour structure of characters, and then uses an SVM classifier to determine the font category. However, in practical applications, this type of method has significant limitations: for fonts with similar glyph structures (such as SimSun and Microsoft YaHei), the HOG features exhibit similar local gradient distributions, resulting in small differences in the category scores output by the SVM classifier, easily leading to misclassification, and lacking effective error correction methods.
[0004] OCR-based methods typically first use an OCR engine to recognize text content and character attributes, and then extract font information from them. The performance of this type of method is highly dependent on the quality of OCR recognition. When the document image contains noise, blurriness, or a complex background, the OCR recognition error will be further amplified, affecting the overall accuracy of font recognition.
[0005] Therefore, how to effectively solve the problem of confusion between similar-looking fonts while maintaining the advantages of traditional methods such as strong interpretability and low implementation cost is a technical challenge that urgently needs to be solved in the field of font recognition technology. Summary of the Invention
[0006] To address the problem of confusion in recognizing similar-looking fonts while maintaining interpretability in existing technologies, this invention provides a document font recognition method, apparatus, computer device, and storage medium.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A document font recognition method, the method comprising: Obtain an image of at least one text region from the document to be identified, obtained through text region detection. Extract the directional gradient histogram feature vector from the text region image; The directional gradient histogram feature vector is input into a pre-trained multi-class support vector machine classifier to obtain the classification score of each font category. The two font categories with the highest scores are selected and denoted as the first candidate category and its score, and the second candidate category and its score, respectively. The difference between the score of the first candidate category and the score of the second candidate category is determined. If the difference is greater than or equal to a preset score difference threshold, the first candidate category is directly output as the font recognition result of the text region image. If the difference is less than the score difference threshold, the stroke direction energy ratio is determined according to the ratio of the vertical gradient energy to the horizontal gradient energy of the text region image. The final font recognition result is determined from the first candidate category and the second candidate category based on the comparison result of the stroke direction energy ratio and the preset energy ratio separation threshold.
[0008] Optionally, extracting the directional gradient histogram feature vector from the text region image includes: Calculate the horizontal and vertical gradients of each pixel in the text region image, and obtain the gradient magnitude and gradient direction of each pixel based on this. The text region image is divided into multiple non-overlapping cell regions. The gradient direction distribution of all pixels in each cell region is counted to construct a cell direction histogram. Multiple adjacent cells are combined into a Block, and the orientation histogram vectors of all cells in each Block are concatenated and normalized. The normalized feature vectors of all blocks are concatenated sequentially to form the directional gradient histogram feature vector.
[0009] Optionally, based on the comparison result of the stroke direction energy ratio and the preset energy ratio separation threshold, the final font recognition result is determined from the first candidate category and the second candidate category and output, including: When the first candidate category and the second candidate category meet the preset font pair conditions, if the stroke direction energy ratio is less than the energy ratio separation threshold, the first candidate category is output; if the stroke direction energy ratio is greater than or equal to the energy ratio separation threshold, the second candidate category is output.
[0010] Optionally, the multi-class support vector machine classifier employs a radial basis function kernel function; before inputting the directional gradient histogram feature vector into the multi-class support vector machine classifier, a step is taken to standardize each dimension of the directional gradient histogram feature vector to balance the influence of each dimension in the classifier.
[0011] Optionally, before extracting the directional gradient histogram feature vector, the text region image is preprocessed by converting the text region image from a color image to a grayscale image and normalizing the size of all text region images to a preset uniform size specification.
[0012] A document font recognition device, the device comprising: The acquisition module is used to acquire at least one text region image obtained by text region detection in the document to be identified; and to extract the directional gradient histogram feature vector from the text region image. The classification module is used to input the directional gradient histogram feature vector into a pre-trained multi-class support vector machine classifier to obtain the classification score of each font category, and select the two font categories with the highest scores, which are respectively denoted as the first candidate category and its score, and the second candidate category and its score. The determination module is used to determine the difference between the score of the first candidate category and the score of the second candidate category. If the difference is greater than or equal to a preset score difference threshold, the first candidate category is directly output as the font recognition result of the text region image. If the difference is less than the score difference threshold, the stroke direction energy ratio is determined according to the ratio of the vertical gradient energy to the horizontal gradient energy of the text region image. The final font recognition result is determined from the first candidate category and the second candidate category according to the comparison result of the stroke direction energy ratio and the preset energy ratio separation threshold.
[0013] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned document font recognition method.
[0014] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the aforementioned document font recognition method.
[0015] The document font recognition method provided by this invention has the following beneficial effects: This invention, based on traditional HOG features and SVM classifiers, introduces a gated triggering mechanism based on the difference in SVM classification confidence. Specifically, it further utilizes the stroke direction energy ratio for secondary discrimination on low-confidence samples. This method effectively identifies boundary samples of similar-looking fonts by the SVM classifier and achieves accurate differentiation of similar fonts through the auxiliary discrimination of stroke direction energy ratio. Compared with existing technologies, this invention significantly improves the recognition accuracy of similar-looking fonts, reduces misclassification, and enhances the classification stability and interpretability of the system without relying on large-scale labeled data and complex deep learning models. It also maintains the advantages of traditional methods, such as low implementation complexity and convenient engineering debugging, and can be widely applied in scenarios such as graduation thesis format review, standardized document detection, and automated typesetting analysis. Attached Figure Description
[0016] To more clearly illustrate the embodiments and design schemes of the present invention, the accompanying drawings required for this embodiment will be briefly described below. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a document font recognition method according to an exemplary embodiment of the present invention.
[0018] Figure 2 This is a flowchart of a font recognition process provided by the present invention according to an exemplary embodiment.
[0019] Figure 3 This is a flowchart illustrating a gating mechanism provided by the present invention according to an exemplary embodiment.
[0020] Figure 4 This is a block diagram of a document font recognition device according to an exemplary embodiment of the present invention. Detailed Implementation
[0021] To enable those skilled in the art to better understand and implement the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be construed as limiting the scope of protection of the present invention.
[0022] The technical solutions provided by the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] First, this invention provides a document font recognition method, specifically as follows: Figure 1 As shown, it includes the following steps: S101. Obtain at least one text region image from the document to be identified, obtained through text region detection.
[0024] In this step, the acquired text region image is also preprocessed by converting the text region image from a color image to a grayscale image and normalizing the size of all text region images to a preset uniform size specification.
[0025] S102. Extract the directional gradient histogram feature vector from the image of the text region.
[0026] In this step, the horizontal and vertical gradients of each pixel in the text region image are calculated, and the gradient magnitude and direction of each pixel are obtained accordingly. The text region image is divided into multiple non-overlapping cell regions, and the gradient direction distribution of all pixels in each cell region is statistically analyzed to construct a cell orientation histogram. Adjacent cells are combined into blocks, and the orientation histogram vectors of all cells within each block are concatenated and normalized. The normalized feature vectors of all blocks are concatenated end-to-end in sequence to form the orientation gradient histogram feature vector. The normalization of the feature vectors within each block uses L2 norm normalization.
[0027] S103. Input the feature vector of the histogram of the directional gradient into a pre-trained multi-class support vector machine classifier to obtain the classification score of each font category, and select the two font categories with the highest scores, which are respectively denoted as the first candidate category and its score, and the second candidate category and its score.
[0028] The multi-class support vector machine classifier uses a radial basis kernel function. Before inputting the directional gradient histogram feature vector into the multi-class support vector machine classifier, a standardization process is performed on each dimension of the directional gradient histogram feature vector to balance the influence of each dimension in the classifier.
[0029] S104. Determine the difference between the score of the first candidate category and the score of the second candidate category. If the difference is less than a preset score difference threshold, determine the stroke direction energy ratio. Based on the comparison result between the stroke direction energy ratio and the preset energy ratio separation threshold, determine the final font recognition result.
[0030] In this step, if the difference is greater than or equal to a preset score difference threshold, the first candidate category is directly output as the font recognition result of the text region image; if the difference is less than the score difference threshold, the stroke direction energy ratio is determined based on the ratio of the vertical gradient energy to the horizontal gradient energy of the text region image, using the following formula: ; in, The horizontal gradient energy of the preprocessed text region image. For vertical gradient energy, , , and These are the horizontal and vertical gradients for each pixel, respectively. This represents the total number of pixels included in the statistics.
[0031] Based on the comparison between the energy ratio of the stroke direction and the preset energy ratio separation threshold, the final font recognition result is determined from the first candidate category and the second candidate category.
[0032] For example, when the first candidate category and the second candidate category meet the preset font pair conditions, if the stroke direction energy ratio is less than the energy ratio separation threshold, the first candidate category is output; if the stroke direction energy ratio is greater than or equal to the energy ratio separation threshold, the second candidate category is output; wherein, the preset font pair conditions are that the first candidate category is SimSun and the second candidate category is Microsoft YaHei, or the first candidate category is Microsoft YaHei and the second candidate category is SimSun, and the preset energy ratio separation threshold is predetermined based on the stroke direction energy ratio distribution of SimSun and Microsoft YaHei fonts on the training samples.
[0033] This invention, based on traditional HOG features and an SVM classifier, introduces a gated triggering mechanism based on the difference in SVM classification confidence. Specifically, it further utilizes the stroke direction energy ratio for secondary discrimination on low-confidence samples. This method effectively identifies boundary samples of similar-looking fonts by the SVM classifier and achieves accurate differentiation of similar fonts through the auxiliary discrimination of the stroke direction energy ratio. Compared with existing technologies, this invention significantly improves the recognition accuracy of similar-looking fonts, reduces misclassification, and enhances the classification stability and interpretability of the system without relying on large-scale labeled data and complex deep learning models. It also maintains the advantages of traditional methods, such as low implementation complexity and convenient engineering debugging, and can be widely applied in scenarios such as graduation thesis format review, standardized document detection, and automated typesetting analysis.
[0034] Based on the above method steps, the present invention also provides an embodiment for illustration.
[0035] Font recognition, a task within document formatting recognition, aims to determine the font category of a text region based on its visual features. Unlike font size and line spacing, font differences are primarily manifested in character outlines, stroke thickness, serif structure, and overall glyph style. Therefore, the key to font recognition lies in extracting features that effectively represent the glyph structure from the text region image and using a classification model to accurately distinguish between different fonts.
[0036] In PDF documents, font is one of the important attributes that distinguishes the format of the body text. Although different fonts express the same text content, they usually have obvious differences in visual appearance. For example, the overall strokes of the sans-serif font are thicker and the structure is fuller; the serif font usually has serif features and the horizontal strokes are thin and the vertical strokes are thicker; Microsoft YaHei has a smoother overall outline and a more modern style. For computers, these differences will ultimately manifest as differences in the distribution of image edges, gradient direction, and local structural statistics.
[0037] However, font recognition is not a simple image classification problem. A text region may contain multiple characters, and the structural differences between these characters can interfere with the representation of font differences. The same font can also exhibit scale variations at different font sizes. Some fonts are similar in local structure; for example, Microsoft YaHei and Songti share some similarities in certain stroke areas. Therefore, font recognition must not only capture the distribution of character stroke directions but also maximize sensitivity to font style differences.
[0038] Considering the data scale and implementation requirements, this invention does not employ a large-scale deep learning model. Instead, it uses a combination of traditional features and machine learning classifiers. This approach offers good interpretability and achieves relatively stable results even with small sample sizes.
[0039] The flowchart of the font recognition of this invention is as follows: Figure 2 As shown, for each text region detected through the main text region, grayscale and size normalization are first performed; then the gradients of the image in the horizontal and vertical directions are calculated; then, the directional gradient histogram features are constructed; then the features are input into the support vector machine classifier to obtain the font category prediction result; when necessary, a gating mechanism is used to perform secondary discrimination on the samples on the boundary, and the final font recognition result is output.
[0040] To reduce the influence of irrelevant factors on feature extraction, image preprocessing of the input text region is required. Let the original text region image be... If the input is a color image, it will first be converted to a grayscale image.
[0041] ; in, , , These are the pixel values for the three color channels: RGB. This corresponds to the grayscale value.
[0042] Because the sizes of the text regions in the samples can vary, direct extraction would result in different feature dimensions, thus affecting the training of the classification model. Therefore, this invention scales all text regions to a fixed size, assuming the normalized image size is... , can be represented as:
[0043] ; In the above formula, This indicates an image scaling operation. This represents the normalized image. After this operation, all samples are uniformly mapped to the same scale, which is beneficial for subsequent extraction of HOG feature vectors of consistent length.
[0044] HOG, or Histogram of Oriented Gradients, is a classic method for describing local features. Its core idea is to describe the edge and contour structure of an image by statistically analyzing the distribution of gradient directions in local regions. Since character strokes are essentially a series of edges with significant gray-level variations, HOG is very suitable for font recognition.
[0045] Let the normalized grayscale image be Then the gradients of the image in the horizontal and vertical directions can be expressed as follows: ; ; Based on this, the gradient magnitude and gradient direction of each pixel can be calculated: ; ; in, Indicates the edge intensity of a pixel. Indicates the direction of the edge. The gradient magnitude is used to measure the importance of the edge, and the gradient direction is used to reflect the direction of the local stroke.
[0046] After obtaining the gradient information of each pixel, this invention divides the image into several small regions, called cells. For each cell, the distribution of gradients of all pixels within it is statistically analyzed to form an orientation histogram. Assume the orientations are uniformly divided as follows: The nth angle interval, then the nth The histogram values for each directional interval can be represented as:
[0047] ; in, Indicates the first One directional interval, This is an exponential function; the value is 1 when the pixel gradient direction falls within this interval, and 0 otherwise. In this way, the cell eventually forms a... Dimensional directional histogram vector.
[0048] For text images, different fonts have different proportions of horizontal, vertical and diagonal strokes. Therefore, this directional statistics can better reflect the structural differences between fonts.
[0049] Using only cell-level histograms may still be affected by local brightness and contrast variations, therefore further block-level normalization is required. Let a block consist of several adjacent cells; the corresponding concatenated feature vector is... Then its normalization result is:
[0050] ; In the above formula, Representing vectors The 2-norm, To prevent the use of tiny constants with a denominator of zero, normalization is applied, making HOG features more robust to changes in local illumination and contrast.
[0051] Finally, the normalized features of all blocks are concatenated in order to obtain the overall HOG feature vector of the text region: ; in, Indicates the total number of blocks.
[0052] After obtaining the HOG feature vector Subsequently, this invention employs a Support Vector Machine (SVM) as the font classifier. The basic idea of an SVM is to find an optimal separating hyperplane in the feature space that minimizes the difference between samples of different classes.
[0053] In a binary classification problem, let the training samples be... ,in Then the hyperclassification plane can be represented as: ; in, For the weight vector, This is a bias term.
[0054] To achieve maximum margin classification, SVM needs to solve the following optimization problem; ; The constraints are: ; The optimization objective is to maximize the classification margin while ensuring correct sample classification, thereby improving the model's generalization ability.
[0055] In practical font recognition tasks, the HOG feature distributions of different fonts are often not linearly separable, especially when some font structures are similar, making it difficult for a single linear hyperplane to achieve the desired effect. Therefore, this invention introduces a kernel function to map the input features to a higher-dimensional space, thereby achieving the goal of learning nonlinear decision boundaries.
[0056] This invention employs a radial basis function (RBF) kernel, the expression of which is: ; in, This is the kernel function used to control the range of influence of the samples. The RBF kernel has strong nonlinear modeling capabilities and is suitable for handling complex classification tasks with small to medium-sized samples, making it suitable for use in the context of this invention.
[0057] After introducing soft margins, the optimization objective of SVM can be written as: ; The constraints are: ; in, Represents the feature mapping function. For penalty parameters, These are slack variables. By introducing slack variables, the model allows a small number of samples to appear within the interval or to be misclassified, thus achieving a balance between training error and generalization ability.
[0058] Since this invention needs to identify more than two font categories, including SimSun, Heiti, and Microsoft YaHei, a binary classification SVM needs to be extended to a multi-classification scenario. In practice, a "one-to-many" or "one-to-one" strategy can be used to construct a multi-classifier. Let the set of font categories to be identified be:
[0059] ; Then for input features The classifier ultimately outputs its class: ; in, Indicate category The corresponding classification scoring function ultimately selects the category with the highest score as the font recognition result.
[0060] Since the numerical ranges of each dimension of the HOG feature vector may differ, directly inputting it into an SVM could lead to some dimensions with larger units having an excessive impact on the classification results. Therefore, feature standardization is necessary before classification. Let the feature vector be... Weiwei Its mean and standard deviation are respectively and Then the standardized features can be expressed as:
[0061] ; After standardizing all dimensions, we obtain the new feature vector: ; in, This represents the feature dimension. After standardization, the influence of each dimension on the classifier is more balanced, which is beneficial to improving the stability of model training and classification performance.
[0062] In actual experiments, it can be observed that although some fonts have different overall styles, their gradient distributions in local regions may be quite similar. For example, Microsoft YaHei and SimSun are easily confused in certain character areas. If only the direct output of HOG+SVM is relied upon, a small number of boundary samples may be misclassified.
[0063] To further improve the recognition stability of boundary samples, this invention designs a gating correction mechanism, such as... Figure 3 As shown, when the difference between the scores of the first two candidate classes of the classifier is small, it indicates that the sample is near the classification boundary. At this time, the system further introduces a simple stroke direction energy ratio feature for auxiliary discrimination.
[0064] Let the gradients of the image in the horizontal and vertical directions be respectively. and The corresponding average absolute gradient capability can be expressed as: ; ; in, This represents the total number of pixels participating in the statistics. Based on the above formula, the directional energy ratio is constructed as follows:
[0065] ; This ratio reflects the relative relationship between the vertical stroke energy and the horizontal stroke energy in a text region.
[0066] The reason why stroke energy ratio can be used as a basis for font recognition is that it directly quantifies the differences in the core design feature of horizontal and vertical stroke thickness among different fonts. Specifically, this ratio is calculated by comparing the vertical gradient energy (reflecting the edge strength of vertical strokes) to the horizontal gradient energy (reflecting the edge strength of horizontal strokes) in the text region image, effectively capturing the "bones and flesh" structure of the font: for typical Song typefaces, the "thin horizontal and thick vertical" serif design makes the edges of vertical strokes much stronger than those of horizontal strokes, so the stroke energy ratio is significantly greater than 1; while for sans-serif fonts like Heiti and Microsoft YaHei, the thickness of horizontal and vertical strokes is basically the same, and the energy ratio is close to 1. Because different fonts have different distributions of horizontal and vertical strokes, therefore It can be used as an auxiliary criterion for judgment.
[0067] Let the scores of the top two candidate categories output by the SVM be respectively and If the following conditions are met: ; This indicates that the sample is a low-confidence sample and gating correction needs to be initiated. This is a preset threshold.
[0068] In this case, the directional energy ratio is combined The secondary discrimination can be written as: ; in, and For the two candidate categories output by the SVM, This is the energy ratio segmentation threshold. This mechanism allows for secondary correction of a small number of difficult-to-classify samples, thereby improving overall classification stability.
[0069] The HOG+SVM font recognition method used in this invention has good interpretability. HOG features are directly derived from character outlines and stroke direction distribution, while SVM classifies characters using the maximum margin principle; therefore, the model output has clear physical meaning. Compared with end-to-end deep learning methods, this method is easier to analyze the sources of error and is more suitable for providing a detailed explanation of its principles and implementation process in this invention.
[0070] Furthermore, this method demonstrates good practicality under conditions of small to medium-sized samples. HOG feature extraction has low computational overhead, and support vector machines exhibit strong generalization ability even with limited samples, thus enabling high recognition accuracy without relying on large-scale training data.
[0071] Secondly, the present invention also provides a document font recognition device, such as... Figure 4 As shown, it includes: The acquisition module 201 is used to acquire at least one text region image obtained by text region detection in the document to be identified; and to extract the directional gradient histogram feature vector from the text region image.
[0072] The classification module 202 is used to input the directional gradient histogram feature vector into a pre-trained multi-class support vector machine classifier to obtain the classification score of each font category, and select the two font categories with the highest scores, which are respectively denoted as the first candidate category and its score, and the second candidate category and its score.
[0073] The determining module 203 is used to determine the difference between the score of the first candidate category and the score of the second candidate category. If the difference is greater than or equal to a preset score difference threshold, the first candidate category is directly output as the font recognition result of the text region image. If the difference is less than the score difference threshold, the stroke direction energy ratio is determined according to the ratio of the vertical gradient energy to the horizontal gradient energy of the text region image. The final font recognition result is determined from the first candidate category and the second candidate category according to the comparison result of the stroke direction energy ratio and the preset energy ratio separation threshold.
[0074] The present invention also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 The steps of a document font recognition method are provided.
[0075] This invention also provides a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, memory, and non-volatile memory, and may also include other hardware required for various operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then executes it to achieve the above-mentioned functions. Figure 1 The steps of a document font recognition method are provided.
[0076] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0077] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0078] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0079] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0080] It should be noted that the specific embodiments described above enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although the present invention has been described in detail in this specification, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the present invention; and all technical solutions and improvements that do not depart from the spirit and scope of the present invention are covered within the protection scope of the patent of the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A document font recognition method, characterized in that, The method includes: Obtain an image of at least one text region from the document to be identified, obtained through text region detection. Extract the directional gradient histogram feature vector from the text region image; The directional gradient histogram feature vector is input into a pre-trained multi-class support vector machine classifier to obtain the classification score of each font category. The two font categories with the highest scores are selected and denoted as the first candidate category and its score, and the second candidate category and its score, respectively. The difference between the score of the first candidate category and the score of the second candidate category is determined. If the difference is greater than or equal to a preset score difference threshold, the first candidate category is directly output as the font recognition result of the text region image. If the difference is less than the score difference threshold, the stroke direction energy ratio is determined according to the ratio of the vertical gradient energy to the horizontal gradient energy of the text region image. The final font recognition result is determined from the first candidate category and the second candidate category based on the comparison result of the stroke direction energy ratio and the preset energy ratio separation threshold.
2. The method according to claim 1, characterized in that, Extracting the directional gradient histogram feature vector from the text region image includes: Calculate the horizontal and vertical gradients of each pixel in the text region image, and obtain the gradient magnitude and gradient direction of each pixel based on this. The text region image is divided into multiple non-overlapping cell regions. The gradient direction distribution of all pixels in each cell region is counted to construct a cell direction histogram. Multiple adjacent cells are combined into a Block, and the orientation histogram vectors of all cells in each Block are concatenated and normalized. The normalized feature vectors of all blocks are concatenated sequentially to form the directional gradient histogram feature vector.
3. The method according to claim 1, characterized in that, Based on the comparison result between the stroke direction energy ratio and the preset energy ratio separation threshold, the final font recognition result is determined from the first candidate category and the second candidate category and output, including: When the first candidate category and the second candidate category meet the preset font pair conditions, if the stroke direction energy ratio is less than the energy ratio separation threshold, the first candidate category is output; if the stroke direction energy ratio is greater than or equal to the energy ratio separation threshold, the second candidate category is output.
4. The method according to claim 1, characterized in that, The multi-class support vector machine classifier employs a radial basis function kernel. Before inputting the directional gradient histogram feature vector into the multi-class support vector machine classifier, a standardization process is performed on each dimension of the directional gradient histogram feature vector to balance the influence of each dimension in the classifier.
5. The method according to claim 1, characterized in that, Before extracting the directional gradient histogram feature vector, the text region image is preprocessed by converting the text region image from a color image to a grayscale image and normalizing the size of all text region images to a preset uniform size specification.
6. A document font recognition device, characterized in that, The device includes: The acquisition module is used to acquire at least one text region image obtained by text region detection in the document to be identified; and to extract the directional gradient histogram feature vector from the text region image. The classification module is used to input the directional gradient histogram feature vector into a pre-trained multi-class support vector machine classifier to obtain the classification score of each font category, and select the two font categories with the highest scores, which are respectively denoted as the first candidate category and its score, and the second candidate category and its score. The determination module is used to determine the difference between the score of the first candidate category and the score of the second candidate category. If the difference is greater than or equal to a preset score difference threshold, the first candidate category is directly output as the font recognition result of the text region image. If the difference is less than the score difference threshold, the stroke direction energy ratio is determined according to the ratio of the vertical gradient energy to the horizontal gradient energy of the text region image. The final font recognition result is determined from the first candidate category and the second candidate category according to the comparison result of the stroke direction energy ratio and the preset energy ratio separation threshold.
7. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 5.
8. A computer device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any one of claims 1 to 5.