A method and system for discriminating the category of official document fonts
Through image processing technology and training models, the font categories of the agency official documents are automatically identified, which solves the problems of low manual discrimination efficiency and easy confusion, and achieves efficient and accurate assessment of the agency official documents.
Patent Information
- Application Number
- CN202210656569.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-06-10
Smart Images

Figure CN115063818B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of language recognition, and in particular to a method and system for discriminating the font categories of official documents of government agencies. Background Art
[0002] Official documents of government agencies are documents with legal effect and standardized formats in the administration of the country's Party, government, and military. The quality of official documents is an important indicator of the administrative management level and work style, and is also an important indicator for the assessment of government agency operations. An important indicator of the quality of official documents of government agencies is whether the font of the official document conforms to national standards and specifications. For example, the title should be in small double-line Song typeface, the first-level heading should be in boldface, the second-level heading should be in regular script, the third-level and fourth-level headings should be in imitation Song typeface, and each element of the header and footer also has its own font regulations. When assessing government agency operations, it depends on manual experience to judge whether the font of official documents in paper, scanned, or PDF format conforms to the specifications. It is easy to confuse fonts with similar styles, and when there is a large amount of official documents to be evaluated, it is time-consuming and laborious. Summary of the Invention
[0003] An embodiment of the present invention provides a method and system for discriminating the font categories of official documents of government agencies. Font recognition is based on image processing, and various types of documents can be recognized with higher speed and accuracy.
[0004] To achieve the above object, on the one hand, an embodiment of the present invention provides a method for discriminating the font categories of official documents of government agencies, including:
[0005] When the official document of the government agency to be recognized is a text type, first convert the text-type official document into a PDF, and then convert it into an official document image in a preset format; when the official document of the government agency to be recognized is an image, unify the official document into an official document image in a preset format; when the official document of the government agency to be recognized is a paper type, use an official document collection terminal to photograph the paper-type official document into a picture, and perform grayscale processing, noise reduction processing, edge detection, image segmentation, and perspective transformation on the photographed picture in sequence to obtain an official document image in a preset format; use the official document image in the preset format as the official document image to be detected;
[0006] Perform binarization processing on the official document image to be detected, perform horizontal and vertical projections on the binarized official document image respectively, and segment the official document into characters through black pixels and white interval pixels to obtain an image of each character, and record the position information of each character;
[0007] According to each character image, perform font recognition on each character image through a trained font recognition model, and output a font recognition result after the recognition is completed; the font recognition result includes the font category of the character and the position information of the character;
[0008] According to the position information of each character, draw a rectangular box around each character on the official document image to be detected, mark the font category code of the corresponding character inside the rectangular box, and set a prompt for the font category code when the right mouse button is slid over.
[0009] When the right mouse button slides over the font category code on the official document image to be detected, automatically restore the font category code to the corresponding font category name and display it.
[0010] On the other hand, the embodiment of the present invention provides an official document font category discrimination system, including:
[0011] An official document preprocessing unit, which is used to convert the text-type official document into PDF first and then into an official document image in a preset format when the official document to be recognized is of the text type; when the official document to be recognized is an image, unify the official document into an official document image in a preset format; when the official document to be recognized is of the paper type, use an official document acquisition terminal to photograph the paper-type official document into a picture, and perform gray processing, noise reduction processing, edge detection, image segmentation, and perspective transformation on the photographed picture in sequence to obtain an official document image in a preset format; use the official document image in the preset format as the official document image to be detected;
[0012] An official document segmentation unit, which is used to perform binarization processing on the official document image to be detected, perform horizontal and vertical projections on the binarized official document image respectively, and segment the official document into characters through black pixels and white interval pixels to obtain the image of each character, and record the position information of each character;
[0013] A character category recognition unit, which is used to perform font recognition on each character image through a trained font recognition model according to each character image, and output a font recognition result after the recognition; the font recognition result includes the font category of the character and the position information of the character;
[0014] A character category marking unit, which is used to draw a rectangular box around each character on the official document image to be detected according to the position information of each character, mark the font category code of the corresponding character inside the rectangular box, and set a prompt for the font category code when the right mouse button is slid over.
[0015] A display unit, which is used to automatically restore the font category code to the corresponding font category name and display it when the right mouse button slides over the font category code on the official document image to be detected.
[0016] The above technical solution has the following beneficial effects: Font recognition is based on image processing, can recognize various types of documents, and has higher speed and accuracy. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 is a flowchart of the method for discriminating the category of official document fonts in the embodiments of the present invention;
[0019] Figure 2 is a structural diagram of the system for discriminating the category of official document fonts in the embodiments of the present invention;
[0020] Figure 3 is a general structural diagram of the embodiments of the present invention;
[0021] Figure 4 Structural diagram of the official document collection module in the embodiments of the present invention;
[0022] Figure 5 Statistical chart of the training indicators of the font recognition model in the embodiments of the present invention;
[0023] Figure 6 Test result diagram of the font recognition model in the embodiments of the present invention;
[0024] Figure 7 Effect diagram of official document font recognition in the embodiments of the present invention. Detailed implementation manners
[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0026] As Figure 1 shown, in combination with the embodiments of the present invention, a method for discriminating the category of official document fonts is provided, including:
[0027] S101: When the official document to be recognized is a text type, first convert the text-type official document into a PDF, and then convert it into an official document image in a preset format; when the official document to be recognized is an image, unify the official document into an official document image in a preset format; when the official document to be recognized is a paper type, use an official document collection terminal to photograph the paper-type official document into a picture, and perform gray processing, noise reduction processing, edge detection, image segmentation, and perspective transformation on the photographed picture in sequence to obtain an official document image in a preset format; use the official document image in the preset format as the official document image to be detected;
[0028] S102: Binarize the official document image to be detected, perform horizontal and vertical projections on the binarized official document image respectively, segment the characters of the official document through black pixels and white interval pixels to obtain the image of each character, and record the position information of each character;
[0029] S103: According to each character image, perform font recognition on each character image through a trained font recognition model, and output the font recognition result after recognition; the font recognition result includes the font category of the character and the position information of the character;
[0030] S104: According to the position information of each character, draw a rectangular frame for each character on the official document image to be detected, label the font category code of the corresponding character within the rectangular frame, and set a prompt for the font category code when the right mouse button slides over;
[0031] S105: When the right mouse button slides over the font category code on the official document image to be detected, automatically restore the font category code to the corresponding font category name for display.
[0032] Preferably, in step 101, when the official document to be recognized is a paper document, use an official document collection terminal to photograph the paper official document into a picture, and perform gray processing, noise reduction processing, edge detection, image segmentation, and perspective transformation on the photographed picture in sequence to obtain an official document image in a preset format; use the official document image in the preset format as the official document image to be detected, specifically including:
[0033] The high-definition camera automatically photographs the paper official document to be recognized according to the collected image acquisition instruction, and outputs the original image P of the paper official document to be recognized;
[0034] Create a copy P1 of the original official document image P, and perform gray processing on P1; after the gray processing is completed, use Gaussian blur to filter out noise to obtain the edge image E of the original official document image P; specifically:
[0035] Use the Sobel filter to determine the gradient G(x, y) and direction θ of the edge of P1 M ; where G x is the vertical edge, which refers to the sudden change of the gradient in the x direction, and G y is the horizontal edge, which refers to the sudden change of the gradient in the y direction; perform non-maximum suppression on the gradient amplitude in the gradient direction. For each pixel point i of P1, compare the magnitudes of the surrounding 8 neighborhood values along the four types of gradient directions of 0°, 45°, 90°, and 135° in the 3*3 area. If pixel i is the maximum value, keep the pixel point, otherwise set it to 0; combine the double-threshold algorithm to detect and connect the edges to obtain the edge image E of the original official document image P;
[0036] Create a copy E1 of E, obtain the set of edge-closed contours in E1, determine the quadrilateral edges from the set of edge-closed contours, and use the quadrilateral edges as the edges of the original official document image P; and segment the target official document image from the original official document image P according to the position information of the quadrilateral edges.
[0037] Map the target official document image into an image with the size ratio of A4 paper for standard official documents through the coordinates of 4 vertices, and correct each character in the target official document image into a visually proportionally coordinated style through perspective transformation. After correction, an official document image in a preset format is formed.
[0038] Preferably, in step 102, perform binarization processing on the official document image to be detected, perform horizontal and vertical projections on the binarized official document image respectively, and segment the characters of the official document through the black pixels and white interval pixels to obtain the image of each character, specifically including:
[0039] Perform binarization processing on the official document image to be detected to obtain a binarized official document image; perform a horizontal projection on the binarized official document image on the y-axis to obtain the binary image of each row, perform a vertical projection on the binary image of each row on the x-axis, and judge the start position and end position of each character in the official document image to be detected based on the black pixels and white interval pixels obtained by the projection, and obtain the position coordinates of the character according to the start position and end position of each character.
[0040] Draw a segmentation rectangle frame for each character in the official document image to be detected according to the position coordinates of each character, segment the characters through the rectangle frame to obtain the image and position of each character, and at the same time number each character image in a preset form, and record the numbers of each character and the corresponding character position information.
[0041] Preferably, in step 103, according to each character image, perform font recognition on each character image through a trained font recognition model, and output the font recognition result after the recognition, specifically including:
[0042] Input each character image into the trained font recognition model. Use Focus of the backbone network Backbone to obtain 4 sub-images of the same size for each sliced character image respectively. Integrate the width and height of each sub-image through Concat, increasing the number of channels of the input image to 64. Use the Conv convolutional block to perform a convolution operation with a convolution kernel of 3 and a stride of 2 on the sub-images after Concat integration, and output the first feature image. After the first feature image passes through 3 BottleneckCSP modules and Conv convolutional blocks and is output, it becomes the second feature image. Perform max pooling operations on the second feature image at four ratios through the SSP module. Integrate the pooling results through the Concat connection layer, and perform convolution and connection operations on the integrated pooling results through 14 layers of networks in the neck Neck and the head Head to output the font category of each character. After the recognition is completed, output the font recognition result;
[0043] Step 103, according to the position information of each character, draw a rectangular box for each character on the document image to be detected, label the font category code of the corresponding character inside the rectangular box, and set a prompt for the font category code when the right mouse button slides over, specifically including:
[0044] Use the rectangle() function of the python cv2 package to draw a font position rectangular box for each character in the document image to be detected according to the position information of each character. Use the plt.text(x,y,s) function called by the matplotlib package to label the font category code inside the character position rectangular box, where x and y are the abscissa and ordinate of the midpoint of the character position rectangular box respectively, and s is the font category code of the character, and set to display the font category code according to the font category code when the right mouse button slides over the character position rectangular box.
[0045] Preferably, in S106, the following method is used to train the font recognition model:
[0046] Obtain Chinese single characters, Arabic numerals, punctuation marks, and mathematical symbols used in official documents of government agencies, and make font sample pictures using rectangular boxes respectively according to the font categories used in official documents of government agencies. The background color of the font sample pictures is pure white, the text color is black, and the font style is not bold; each font sample picture contains one Chinese character, number or symbol; among them, the font categories used include at least one of the following: Fangzheng Xiaobiao Song Simplified, Fangsong, Fangsong_GB2312, Boldface, Kaiti, Kaiti_GB2312, Songti;
[0047] Use the labelImg annotation software to annotate the character categories in each font sample image with character category codes, and output a label annotation result file in txt format; each label annotation result file corresponds to its homonymous font sample image; use each label annotation result file and its homonymous font sample image as a dataset, and divide the dataset into a training set and a validation set; among them, the data in each annotation result file includes: cls, x, y, w, h, where cls is the font category, and x and y are the horizontal and vertical coordinates of the center point of the rectangle respectively, and w and h are the width value and height value of the rectangle frame;
[0048] For the configuration of the font recognition model, set the depth control parameter depth_multiple of the font recognition model to 0.33 and the width control parameter width_multiple to 0.50; at 8x, 18x, and 32x downsampling, the sizes of the 3 prior boxes are set to (10, 13), (16, 30), (33, 23), (30, 61), (62, 45), (59, 119), (116, 90), (156, 198), (373, 326) respectively, and the weight file is set to yolov5s.pt; set the number of categories to 7 in the data configuration file VOC.yaml, set the font category names according to the character category codes, and configure the addresses of the training set and validation set of the font recognition model;
[0049] For the training set, segment the official document images of government agencies into single-character images, and adjust each single-character image to a preset size and input it into the YOLOV5 model; the YOLOV5 model includes a backbone network Backbone, a neck Neck, and a head Head; the backbone network Backbone includes Focus, Concat, Conv convolutional blocks, SSP, and BottleneckCSP;
[0050] Use Focus to slice each single-character image to obtain 4 sub-images of the same size respectively; use Concat to integrate the width and height of the sub-images and increase the number of channels of the input image to 64; use the Conv convolutional block to perform a convolutional operation with a convolutional kernel of 3 and a stride of 2 on the sub-images integrated by Concat, and output the first feature image; after the first feature image passes through 3 BottleneckCSP modules and Conv convolutional blocks and is output, it becomes the second feature image; use the SSP module to perform maximum pooling operations on the second feature image in four ratios respectively; use the Concat connection layer to integrate the pooling results, and perform convolutional and connection operations on the integrated pooling results through 14 layers of networks in the neck Neck and the head Head to output the single-character border and the font category;
[0051] And use the validation set for verification to obtain a trained font recognition model.
[0052] As Figure 2 shown, in combination with the embodiments of the present invention, a discriminant system for the font category of official documents of an institution is provided, including:
[0053] An official document preprocessing unit 21, which is used to convert a text-type official document into a PDF first and then into an official document image in a preset format when the official document of the institution to be recognized is of the text type; when the official document of the institution to be recognized is an image, unify the official document of the institution into an official document image in a preset format; when the official document of the institution to be recognized is of the paper type, use an official document acquisition terminal to photograph the paper-type official document of the institution into a picture, and perform gray-scale processing, noise reduction processing, edge detection, image segmentation, and perspective transformation on the photographed picture in sequence to obtain an official document image in a preset format; and use the official document image in the preset format as the official document image to be detected;
[0054] An official document segmentation unit 22, which is used to perform binarization processing on the official document image to be detected, perform horizontal and vertical projections on the binarized official document image respectively, and segment the official document by black pixels and white interval pixels to obtain an image of each character, and record the position information of each character;
[0055] A character category recognition unit 23, which is used to perform font recognition on each character image through a trained font recognition model according to each character image, and output a font recognition result after the recognition; the font recognition result includes the font category of the character and the position information of the character;
[0056] A character category annotation unit 24, which is used to draw a rectangular frame for each character on the official document image to be detected according to the position information of each character, label the corresponding font category code within the rectangular frame, and set a prompt for the font category code when the right mouse button slides over;
[0057] A display unit 25, which is used to automatically restore the font category code to the corresponding font category name and display it when the right mouse button slides over the font category code on the official document image to be detected.
[0058] Preferably, the official document preprocessing unit 21 includes a paper-type official document preprocessing subunit 211, and the paper-type official document preprocessing subunit 211 is specifically used for:
[0059] A high-definition camera automatically photographs the paper-type official document of the institution to be recognized according to the collected image acquisition instruction, and outputs the original image P of the paper-type official document of the institution to be recognized;
[0060] Create a copy P1 of the original official document image P, and perform gray-scale processing on P1; after the gray-scale processing is completed, use Gaussian blur to filter out noise to obtain the edge image E of the original official document image P; specifically:
[0061] Determine the gradient G(x, y) and direction θ of the P1 edge using the Sobel filter M ; where G x is the vertical edge, which refers to the mutation of the gradient in the x direction, and G y is the horizontal edge, which refers to the mutation of the gradient in the y direction; perform non-maximum suppression on the gradient magnitude in the gradient direction. For each pixel point i of P1, compare the magnitudes of the surrounding 8 neighborhood values along the four types of gradient directions of 0°, 45°, 90°, and 135° within the 3*3 region. If pixel i is the maximum value, retain the pixel point, otherwise set it to 0; combine the double-threshold algorithm to detect and connect the edges to obtain the edge image E of the original official document image P;
[0062] Create a copy E1 of E, obtain the set of edge closed contours in E1, determine the quadrilateral edges from the set of edge closed contours, and use the quadrilateral edges as the edges of the original official document image P; and segment the target official document image from the original official document image P according to the position information of the quadrilateral edges;
[0063] Map the target official document image into an image with the size ratio of the A4 paper of the standard official document through the 4 vertex coordinates, and correct each character of the target official document image into a visually proportionally coordinated style through perspective transformation to form an official document image in a preset format.
[0064] Preferably, the official document segmentation unit 22 is specifically used for:
[0065] Perform binarization processing on the official document image to be detected to obtain a binarized official document image; perform horizontal projection on the binarized official document image on the y-axis to obtain the binary image of each row, perform vertical projection on the binary image of each row on the x-axis, and judge the start position and end position of each character in the official document image to be detected based on the black pixels and white interval pixels obtained by the projection, and obtain the position coordinates of the character according to the start position and end position of each character;
[0066] Draw a segmentation rectangle frame for each character in the official document image to be detected according to the position coordinates of each character, segment the characters through the rectangle frame to obtain the image and position of each character, and simultaneously number each character image in a preset form, and record the numbers of each character and the corresponding character position information.
[0067] Preferably, the character category recognition unit 23 is specifically used for:
[0068] Input each character image into the trained font recognition model. Through the Focus of the Backbone, 4 sub-images of the same size are obtained by slicing each character image respectively. Integrate the width and height of each sub-image through Concat, and increase the number of channels of the input image to 64. Use the Conv convolution block to perform a convolution operation with a convolution kernel of 3 and a stride of 2 on the sub-images integrated by Concat, and output the first feature image. After the first feature image passes through 3 BottleneckCSP modules and Conv convolution blocks and is output, it becomes the second feature image. Perform max-pooling operations on the second feature image in four ratios through the SSP module. Integrate the pooling results through the Concat connection layer, and perform convolution and connection operations on the integrated pooling results through the 14-layer network of the Neck and Head to output the font category of each character. After the recognition is completed, output the font recognition result;
[0069] The character category annotation unit 24 is specifically used for:
[0070] Use the rectangle() function of the python cv2 package to draw a font position rectangle frame for each character in the document image to be detected according to the position information of each character. Use the plt.text(x,y,s) function of the matplotlib package to label the font category code inside the character position rectangle frame, where x and y are the abscissa and ordinate of the midpoint of the character position rectangle frame respectively, and s is the font category code of the character, and set to display the font category code according to the font category code when the mouse right button slides over the character position rectangle frame.
[0071] Preferably, it further includes a font recognition model training unit 26. The font recognition model training unit 26 includes:
[0072] Obtain Chinese single characters, Arabic numerals, punctuation marks, and mathematical symbols used in official documents of government agencies, and make font sample pictures in the form of rectangular frames respectively according to the font categories used in official documents of government agencies. The background color of the font sample pictures is pure white, the text color is black, and the font style is not bold. Each font sample picture contains a Chinese character, number or symbol. Among them, the font categories used include at least one of the following: Fangzheng Xiaobiao Song Simplified, Fangsong, Fangsong_GB2312, Heiti, Kaiti, Kaiti_GB2312, Songti;
[0073] Use the labelImg annotation software to label the character categories in each font sample picture with character category codes, and output a label annotation result file in txt format; each label annotation result file corresponds to its homonymous font sample picture; use each label annotation result file and its homonymous font sample picture as a dataset, and divide the dataset into a training set and a validation set; among them, the data in each annotation result file includes: cls, x, y, w, h, where cls is the font category, and x and y are the horizontal and vertical coordinates of the center point of the rectangle respectively, and w and h are the width value and height value of the rectangle frame;
[0074] For the configuration of the font recognition model, set the depth control parameter depth_multiple of the font recognition model to 0.33 and the width control parameter width_multiple to 0.50; at 8x, 18x, and 32x downsampling, the sizes of the 3 prior boxes are set to (10, 13), (16, 30), (33, 23), (30, 61), (62, 45), (59, 119), (116, 90), (156, 198), (373, 326) respectively, and the weight file is set to yolov5s.pt; set the number of categories to 7 in the data configuration file VOC.yaml, set the font category names according to the character category codes, and configure the addresses of the training set and validation set of the font recognition model;
[0075] For the training set, segment the official document images of government agencies into single-character images, and adjust each single-character image to a preset size and input it into the YOLOV5 model; the YOLOV5 model includes a backbone network Backbone, a neck Neck, and a head Head; the backbone network Backbone includes Focus, Concat, Conv convolutional blocks, SSP, and BottleneckCSP;
[0076] Use Focus to slice each single-character image to obtain 4 sub-images of the same size; use Concat to integrate the width and height of the sub-images, and increase the number of channels of the input image to 64; use the Conv convolutional block to perform a convolutional operation on the sub-images integrated by Concat with a convolutional kernel of 3 and a stride of 2, and output the first feature image; after the first feature image passes through 3 BottleneckCSP modules and Conv convolutional blocks and is output, it becomes the second feature image; use the SSP module to perform maximum pooling operations on the second feature image in four ratios respectively; use the Concat connection layer to integrate the pooling results, and perform convolutional and connection operations on the integrated pooling results through 14 layers of networks in the neck Neck and the head Head to output the single-character border and the font category;
[0077] And use the validation set for verification to obtain a trained font recognition model.
[0078] The above technical solution of the embodiment of the present invention is described in detail below in conjunction with specific application examples. For technical details not introduced during the implementation process, please refer to the relevant description in the previous text.
[0079] A method and device for distinguishing fonts of official documents of government agencies based on image processing technology, involving the fields of natural language processing and computer machine vision, and mainly used for quality assessment and evaluation of official documents of party, government, military and other agencies.
[0080] The present invention is aimed at the problem that when evaluating the quality of official documents of government agencies, manual judgment and word-by-word verification of whether the fonts of each element are correct are time-consuming and laborious, similar fonts are easily confused, and the use of machine analysis to read fonts is faced with the situation that paper documents, scanned documents, PDF and other main image documents cannot be distinguished. The present invention can quickly and efficiently determine the fonts of official documents of government agencies, and provide an efficient solution for assisting the quality evaluation of official documents of government agencies. Among them, official documents of government agencies refer to the format of official documents of national administrative agencies, which conform to the provisions of the national standard GB-T 9704-2012 of the People's Republic of China.
[0081] According to the characteristics of each font detail image feature, the present invention uses the deep learning target detection method based on the image to distinguish fonts. It is divided into three parts: the document collection terminal, the central processing system, and the display and judgment terminal. Figure 3 shown.
[0082] 1. Official document collection terminal
[0083] It consists of three parts: image acquisition module, document import interface, and recognition conversion module. The structure is as follows Figure 4 shown.
[0084] (1) Image acquisition module: One high-definition camera is set up, with a resolution of no less than 1080P (1920*1080), video compression method: Motion-JPEG, signal system: PAL or NTSC, frame rate: greater than 25fps, interface: USB3.0 high-speed interface, driver-free; after receiving the image acquisition command, the image acquisition module automatically shoots paper documents and outputs 1920*1080 original images.
[0085] (2) Document import interface: Configure the USB data interface to directly import data from electronic documents such as doc, docx, wps, and scanned or copied images such as bmp, jpg, png, titf.
[0086] (3) Recognition and conversion module: After receiving the original document image P taken by the image acquisition module, it creates a copy of P P1, and then performs grayscale processing on P1 to remove the image color. Grayscale processing calculation method is as follows:
[0087]
[0088] Among them, f(x, y) is the image after grayscale processing, and r(x, y), g(x, y), and b(x, y) represent the R, G, and B values at that place.
[0089] Then, perform Gaussian blur on the image P1 to filter out noise for a more accurate edge image E. Specifically: Use the Sobel filter to determine the gradient G(x, y) and direction θ of the edges of the image P1 M , defined as in formulas (2) and (3). Among them, G x is the vertical edge, which refers to the abrupt change of the gradient in the x direction, and G y is the horizontal edge, which refers to the abrupt change of the gradient in the y direction.
[0090]
[0091]
[0092] Then, perform non-maximum suppression on the gradient magnitude in the gradient direction. For each pixel point i of the original official document image P1, in the 3*3 area, along the four types of gradient directions of 0°, 45°, 90°, and 135°, compare the magnitudes of the surrounding 8 neighborhood values. If pixel i is the maximum value, then retain the pixel point; otherwise, set it to 0. Perform non-maximum suppression in this way, and finally use the double-threshold algorithm to detect and connect the edges to obtain the edge image E of the original official document image P.
[0093] Create a copy E1 of E, find the set of edge closed contours from E1, find the quadrilateral edges from the contour set, determine them as the edges of the original official document image P, and then segment the target official document image from the original official document image P through the quadrilateral edge position information.
[0094] Map the target official document image into an image with the size ratio of A4 paper of the standard official document through 4 vertex coordinates, with a resolution of 1120*790, and then perform perspective transformation. The perspective transformation method is as in formula (4).
[0095]
[0096] Among them, the 3*3 matrix on the right side of the formula is the mapping matrix, obtained from the OpenCV library, and x, y are the coordinates of the target official document image before perspective transformation. The image after perspective transformation is defined as in formula (5).
[0097]
[0098] The image f after perspective transformation t(x, y) is saved as JPEG format as the image to be detected of the paper document. When shooting the image of the paper document, the image is deformed due to the incorrect angle. The purpose of perspective transformation is to correct the deformed image to an A4 size image.
[0099] The recognition and conversion module receives documents such as doc, docx, wps, etc. imported from the document import interface and converts them into PDF first, and then into JPEG format images with a resolution of 1120*790 for the next processing; among them, scanned and copied image documents such as bmp, jpg, png, titf, etc. are also uniformly converted into JPEG format images with a resolution of 1120*790 for the next processing.
[0100] 2. Central processing system
[0101] (1) Construction of font recognition model. The font recognition model is built on the basis of the YOLOv5 model. The YOLOv5 model is improved on the basis of the YOLOv3 model. Based on the Pytorch framework, it is easier to configure and use in practice, and has faster training speed, higher accuracy, and an object recognition speed of 140FPS. The font recognition model consists of three parts: the backbone network, the neck, and the head.
[0102] The backbone network Backbone includes Focus, Conv convolutional blocks, SSP, BottleneckCSP and other modules. When the model recognizes official document fonts, it first divides the document image with a resolution of 1120*790 into single word images, and then adjusts each single word image to 640*640 size for input into the model. The model slices the single word image through Focus, adjusts the single word image into 4 sub-images of 320*320 size, and then integrates the width and height of the sub-image through Concat, increasing the number of channels of the input image to 64. For the image integrated by Concat, the Conv convolutional block is used to perform a convolution operation with a convolution kernel of 3 and a stride of 2, and the output result is a feature image of 160*160*128. After the feature image is output by the BottleneckCSP module and the Conv convolutional block three times, it becomes a 20*20*1024 image. The SSP module then performs maximum pooling operations on the 20*20 image, dividing it into four groups of 1*1, 5*5, 9*9, and 13*13 to improve the accuracy of the model. Finally, the Concat layer is used to integrate the pooling results together, and the 14-layer network of the neck and head performs convolution and connection operations to output single-word bounding boxes and font categories.
[0103] The main metric for model loss calculation is the rectangular box loss, which is calculated using CIOU and is defined as in formula (6).
[0104]
[0105] In formula (6), ρ 2 (b, b gt ) is the geometric distance between the midpoints of the predicted and the true bounding boxes, c is the area of the smallest enclosing box covering the predicted and the true boxes minus the union of the predicted and the true boxes, IOU is the intersection over union of the predicted and the true boxes, and v is the degree of fit of the lengths and widths of the predicted and the true boxes, which is defined as in formula (7).
[0106]
[0107] α is a tuning parameter, which is defined as in formula (8).
[0108]
[0109] (2) Training of the font recognition model. After the model is constructed, a training set is made and the model is trained.
[0110] Step 1. Dataset making. Collect all the commonly used single Chinese characters, Arabic numerals, punctuation marks, mathematical symbols, etc. in official document drafting, and make square sample pictures of the fonts respectively according to 7 types of fonts, namely Fangzheng Xiaobiao Song Simplified, Fang Song, Fang Song_GB2312, Hei Ti, Kai Ti, Kai Ti_GB2312, and Song Ti, which are commonly used in official documents of government agencies. The resolution is not less than 32*32. Each font sample picture contains only one Chinese character, numeral or symbol, the background color is pure white, the text color is black, and the font style is not bold.
[0111] Step 2. Sample annotation. Use the labelImg annotation software to annotate each font sample picture, and the output format is the YOLO format. The category codes of the 7 types of fonts, namely Fangzheng Xiaobiao Song Simplified, Fang Song, Fang Song_GB2312, Hei Ti, Kai Ti, Kai Ti_GB2312, and Song Ti, are respectively recorded as: "xiaobiaosong", "fangsong", "fangsong_gb", "heiti", "kaiti", "kaiti_gb", "songti". After annotation, output the label file in txt format. Each txt format annotation result file corresponds to 1 sample picture with the same name. The data storage format of the annotation result file is: cls, x, y, w, h, where cls is the target category, x and y are the horizontal and vertical coordinates of the center point of the annotation box, and w and h are the width and height values of the annotation box.
[0112] step3. Model configuration. The model type is selected as yolov5s with low complexity. The model depth control parameter depth_multiple is set to 0.33, and the model width control parameter width_multiple is set to 0.50. At 8x, 18x, and 32x downsampling, the sizes of the 3 prior boxes are set to (10,13), (16,30), (33,23), (30,61), (62,45), (59,119), (116,90), (156,198), (373,326) respectively. The weight file is set to yolov5s.pt. In the data configuration file VOC.yaml, the number of classes is set to 7, and the class names are set according to the font class codes. The dataset and validation set file addresses of the model are configured.
[0113] step4. Model training. After the model configuration is completed, the data samples are input into the model for training. Considering the relatively small overall scale of the dataset, the sample images and annotation results are both divided into a training set and a validation set in a ratio of 8:2. Each file is associated by file name. The sample images and sample annotation results are stored in the images and labels folders respectively. Under the images and labels folders, create 1 train folder and 1 val folder each to store their respective validation sets and training sets.
[0114] Start the model training. After the model training is completed, start the official document font recognition.
[0115] (3) Single-character segmentation of official documents. Before the official document font recognition, the official document needs to be segmented into individual character images. The segmentation first performs binarization processing on the official document image to be recognized, and then performs horizontal projection on the official document image (binary image) on the y-axis to obtain the binary image of each row. Then, perform vertical projection on the binary image of each row on the x-axis. Finally, by counting the number of black pixels and white interval pixels in the projection, the start and end positions of each character are determined, and then the position coordinates of each character are obtained, a rectangular box is drawn for character segmentation, and the number and position information of the character are recorded.
[0116] (4) Font discrimination. Input the segmented single-character images into the model to obtain the font recognition results and position information.
[0117] 3. Display and interpretation terminal
[0118] According to the font recognition results and the position information of individual characters, the rectangle() function of the python cv2 package is used to draw a frame (font position rectangular frame) for each character on the image to be recognized, and the recognized font category code is marked. An auxiliary prompt for the right-click font name is set. When the right mouse button slides over the font mark information, the font category code is automatically restored to the original font name of "Fang Zheng Xiao Biao Song Simplified, Fang Song, Fang Song_GB2312, Bold, Kaiti, Kaiti_GB2312, Songti", which is convenient for official document examiners to read.
[0119] The beneficial effects achieved by the present invention are as follows:
[0120] The font recognition of the present invention is based on image processing technology, which overcomes the limitation that traditional text processing technology cannot process paper-based scans, photos, PDF and other image-based documents when interpreting fonts; using machine automatic identification instead of manual word-by-word verification avoids misjudgment of similar and easily confused fonts, with higher speed and accuracy, and improves the efficiency of the quality assessment of government documents. The effect of the present invention is not limited to official text font recognition, but can also help users identify font categories through font effect images during graphic design and multimedia production, thereby laying a foundation for accurate font search and efficient creation.
[0121] This embodiment takes the 7 types of fonts involved in commonly used official documents, including Fangzheng Xiaobiao Song Simplified, Fangsong, Fangsong_GB2312, Heiti, Kaiti, Kaiti_GB2312, and Songti, as examples. Experimental environment: Windows 10 operating environment, Python 3.9.7 programming language, Anaconda3 package manager and environment manager, CPU: Intel (R) Core (TM) i7-10875H, graphics card: NVIDIA GeForce RTX 2070, memory: 32G.
[0122] 1. First, for paper documents, the document collection terminal converts them into JPEG format images through operations such as shooting, edge detection, image segmentation, and perspective transformation for the next step of processing; for electronic document-type documents such as doc, docx, and wps, they are first converted into PDF and then converted into JPEG image format for the next step of processing; for scanned and copied image format documents such as bmp, jpg, png, and titf, they are uniformly converted into JPEG image format for the next step of processing.
[0123] 2. Construct a training dataset for the font recognition model. Collect 7,889 characters commonly used in official documents and store them in a single docx document. Then copy this document 7 times and name each copy after one of 7 font names. Set the font of all characters in the document to the font corresponding to the document name. Set the page size of these 7 docx documents to 5 cm wide, 5 cm high, with a white background and a 0.5 cm margin on all sides. Set the font size to 100, black, and not bold, so that each page contains exactly one character.
[0124] 3. Use the functions in the client package of the win32com module in the Python language to open the character set document. Use the ExportAsFixedFormat() function to convert the doc document into a pdf format document. Then use the fitz module to open the pdf document and export the image of each page of the document. The image name is "font name + page number", and the format is JPEG, with a resolution of 189 * 189.
[0125] Given the large number of character samples, to improve the sample training speed and make the sample text close to the real text size, the 189 * 189 page images are compressed proportionally to 30 * 30. Considering that manual annotation is time-consuming and laborious, and the annotation target is single and the size is fixed, a method of automatically generating samples by the program is adopted. Referring to the output xml format of the labelImg annotation software, <folder> <filename> <path> <source> <size>Attribute tags such as these are written according to the attribute information of the sample images. The upper-left and upper-right coordinates of the characters in the sample images are set as (5,5) and (25,25) respectively. Batch generate annotation files in xml annotation format, and then batch convert them into YOLO format annotation result files.
[0126] 4. Store 7 types of 55,223 font sample images and 7 types of 55,223 YOLO format annotation result files in the same directory respectively. Each directory consists of 43,563 files to form a training set and 11,660 files to form a validation set. The names of the sample images and the annotation result files correspond one by one. After configuring the model, input the sample and annotation result files into the model. The training parameters are batch-size: 16, workers: 0, epochs: 50. The training results are as Figure 5 。
[0127] Table 1 Statistical Table of Model Training Metrics
[0128]
[0129] Through 50 rounds of iterative training, the model accuracy is 96.7%, the recall rate is 96.2%, and mAP_0.5 is 98.9%. For details, see Table 1. train / box_loss is the training set box loss, train / obj_loss is the training set confidence loss, train / cls_loss is the training set class loss, val / box_loss is the validation set box loss, val / obj_loss is the validation set confidence loss, and val / cls_loss is the validation set class loss. The model has achieved good metric results. Use the trained model to test a set of 16-character font test images including 13 regular script fonts and 3 Fangzheng Xiaobiao Song fonts. The results are as Figure 6 shown. The model accurately predicts the font category and gives the confidence value of the corresponding font category.
[0130] 5. Binarize 1 test official document image, then perform horizontal and vertical projections, and then perform character segmentation. The segmented characters are numbered in the form of "c(sequence number)", and the position coordinates (x1, y1, x2, y2) of each character are recorded. Input the character images into the trained model for font recognition. The model output results are as Figure 7 shown. On the left is the single character image after segmentation of the test official document image, and on the right is the result after the model recognizes the font. Among them, the red is the recognized Fangzheng Xiaobiao Song font, the light yellow is the recognized regular script font, and the pink is the recognized FangSong_GB2312 font.
[0131] 6. For the recognized font position and category information, use the rectangle() function in the python cv2 package to draw a rectangular box for the font position on the official document image, and use the matplotlib package to call the plt.text(x, y, s) function to label the font category information, where x and y are the horizontal and vertical coordinates of the midpoint of the label box respectively, and s is the category information of the font. Then display the labeled official document on the terminal display device to assist the official document evaluation personnel of the organization to quickly conduct the quality assessment of official documents.
[0132] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of the present disclosure. The appended method claims present the elements of the various steps in an exemplary order and are not limited to the specific order or hierarchy recited.
[0133] In the above detailed description, various features are combined in a single embodiment to simplify the present disclosure. This method of disclosure should not be construed as reflecting an intention that the embodiments of the claimed subject matter require more features than are clearly recited in each claim. On the contrary, as reflected in the appended claims, the present invention lies in a state less than all the features of the disclosed single embodiment. Therefore, the appended claims are hereby expressly incorporated into the detailed description, where each claim stands alone as a separate preferred embodiment of the present invention.
[0134] In order to enable any person skilled in the art to implement or use the present invention, the disclosed embodiments have been described above. For those skilled in the art, various modification methods of these embodiments are obvious, and the general principles defined herein can also be applied to other embodiments without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.
[0135] The foregoing description includes examples of one or more embodiments. Of course, it is not possible to describe all possible combinations of components or methods for the purpose of describing the above embodiments, but those of ordinary skill in the art should recognize that the various embodiments can be further combined and arranged. Therefore, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, this term is covered in a manner similar to the term "including", as interpreted when "including" is used as a transitional word in a claim. In addition, any use of the term "or" in the specification or claims is intended to mean "non-exclusive or".
[0136] Those skilled in the art can also appreciate that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention can be implemented by electronic hardware, computer software, or a combination of both. To clearly show the interchangeability of hardware and software, the above various illustrative components, units, and steps have been generally described in terms of their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the overall system. Those skilled in the art can use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present invention.
[0137] The various illustrative logical blocks or units described in the embodiments of the present invention can be implemented or operated to perform the described functions by a general-purpose processor, a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of the above designs. The general-purpose processor can be a microprocessor, and optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other similar configuration.
[0138] In the embodiments of the present invention, the steps of the methods or algorithms described can be directly embedded in hardware, software modules executed by a processor, or a combination of the two. The software modules can be stored in a RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, register, hard disk, removable disk, CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be provided in an ASIC, and the ASIC can be provided in a user terminal. Optionally, the processor and the storage medium can also be provided in different components of the user terminal.
[0139] In one or more exemplary designs, the above-described functions in the embodiments of the present invention can be implemented in hardware, software, firmware, or any combination of the three. If implemented in software, these functions can be stored on a computer-readable medium or transmitted on a computer-readable medium in the form of one or more instructions or codes. A computer-readable medium includes a computer storage medium and a communication medium that facilitates the transfer of a computer program from one place to another. The storage medium can be any available medium accessible by a general or special computer. For example, such a computer-readable medium can include, but is not limited to, RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions or data structures and other forms readable by a general or special computer, or a general or special processor. In addition, any connection can be appropriately defined as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote resource via a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless means such as infrared, wireless, and microwave, it is also included in the defined computer-readable medium. The disks and discs mentioned above include compact disks, laser disks, optical discs, DVDs, floppy disks, and Blu-ray discs. Disks usually replicate data magnetically, while discs usually replicate data optically by laser. The above combinations can also be included in the computer-readable medium.
[0140] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.< / size> < / path> < / filename> < / folder>
Claims
1. A method for discriminating the type of official document font, characterized in that, Including: When the official document of the organization to be recognized is a text type, first convert the text official document into PDF, and then convert it into an official document image in a preset format; When the official document of the organization to be recognized is an image, unify the official document of the organization into an official document image in a preset format; when the official document of the organization to be recognized is a paper type, use an official document collection terminal to photograph the paper official document of the organization into a picture, and perform gray-scale processing, denoising processing, edge detection, image segmentation, and perspective transformation on the photographed picture in sequence to obtain an official document image in a preset format; use the official document image in the preset format as the official document image to be detected; Perform binarization processing on the official document image to be detected, perform horizontal and vertical projections on the binarized official document image respectively, and perform character segmentation on the official document through black pixels and white interval pixels to obtain an image of each character, and record the position information of each character; According to each character image, perform font recognition on each character image through a trained font recognition model, and output the font recognition result after the recognition; the font recognition result includes the font category of the character and the position information of the character; According to the position information of each character, draw a rectangular box for each character on the official document image to be detected, mark the font category code of the corresponding character in the rectangular box, and set a prompt for the font category code when the right mouse button slides over; When the right mouse button slides over the font category code on the official document image to be detected, automatically restore the font category code to the corresponding font category name and display it; When the official document of the organization to be recognized is a paper type, use an official document collection terminal to photograph the paper official document of the organization into a picture, and perform gray-scale processing, denoising processing, edge detection, image segmentation, and perspective transformation on the photographed picture in sequence to obtain an official document image in a preset format, specifically including: The high-definition camera automatically photographs the paper official document of the organization to be recognized according to the collected image acquisition instruction, and outputs the original image P of the paper official document of the organization to be recognized; Create a copy P1 of the original official document image P, and perform gray-scale processing on P1; after the gray-scale processing is completed, use Gaussian blur to filter out the noise to obtain the edge image E of the original official document image P; specifically: Determine the gradient G(x, y) and direction θ of the P1 edge using the Sobel filter M ; where G x is a vertical edge, which refers to the mutation of the gradient in the x direction, and G y is a horizontal edge, which refers to the mutation of the gradient in the y direction; perform non-maximum suppression on the gradient magnitude in the gradient direction. For each pixel point i of P1, compare the magnitudes of the surrounding 8 neighborhood values along the four types of gradient directions of 0°, 45°, 90°, and 135° in the 3*3 region. If pixel i is the maximum value, retain the pixel point; otherwise, set it to 0; combine the double-threshold algorithm to detect and connect the edges to obtain the edge image E of the original official document image P; Create a copy E1 of E, obtain the set of edge closed contours in E1, determine the quadrilateral edge from the set of edge closed contours, and use the quadrilateral edge as the edge of the original official document image P; and segment the target official document image from the original official document image P according to the position information of the quadrilateral edge; Map the target official document image into an image with the size ratio of the A4 paper of the standard official document through 4 vertex coordinates, and correct each character of the target official document image into a visually proportionally coordinated style through perspective transformation, and form an official document image in a preset format after correction.
2. The method for discriminating the category of official document fonts according to claim 1, wherein The performing binarization processing on the official document image to be detected, performing horizontal and vertical projections on the binarized official document image respectively, and performing character segmentation on the official document through black pixels and white interval pixels to obtain an image of each character, specifically including: Perform binarization processing on the official document image to be detected to obtain a binarized official document image; perform horizontal projection on the binarized official document image on the y-axis to obtain a binary image for each row, perform vertical projection on the binary image for each row on the x-axis, and determine the start position and end position of each character in the official document image to be detected based on the black pixel and white interval pixel obtained from the projection, and obtain the position coordinates of the character according to the start position and end position of each character; Draw a segmentation rectangle frame for each character in the official document image to be detected according to each character position coordinate, segment the characters through the rectangle frame to obtain the image and position of each character, and at the same time number each character image in a preset form, and record the numbers of each character and the corresponding character position information.
3. The method for discriminating the font category of official documents of an institution according to claim 2, wherein, According to each character image, perform font recognition on each character image through a trained font recognition model, and output the font recognition result after recognition, specifically including: Input each character image into a trained font recognition model, and use the Focus of the backbone network Backbone to slice each character image to obtain 4 sub-images of the same size; use Concat to integrate the width and height of each sub-image, and increase the number of channels of the input image to 64; use the Conv convolution block to perform a convolution operation with a convolution kernel of 3 and a stride of 2 on the sub-image integrated by Concat, and output the first feature image; after the first feature image passes through 3 BottleneckCSP modules and the output of the Conv convolution block, it becomes the second feature image; perform a maximum pooling operation on the second feature image through the SSP module; use the Concat connection layer to integrate the pooling results, and perform convolution and connection operations on the integrated pooling results through 14 layers of networks in the Neck and Head to output the font category of each character, and output the font recognition result after recognition; According to the position information of each character, draw a rectangle frame for each character on the official document image to be detected, label the font category code of the corresponding character within the rectangle frame, and set a prompt for the font category code when the right mouse button slides over, specifically including: Use the rectangle() function of the python cv2 package to draw a character position rectangle frame for each character in the official document image to be detected according to the position information of each character; use the matplotlib package to call the plt.text(x,y,s) function to label the font category code within the character position rectangle frame, where x and y are the abscissa and ordinate of the center point of the character position rectangle frame respectively, and s is the font category code of the character, and set to display the font category code according to the font category code when the right mouse button slides over the character position rectangle frame.
4. The method for discriminating the type of official document font according to claim 1, characterized in that, Train the font recognition model by the following method: Obtain single Chinese characters, Arabic numerals, punctuation marks, and mathematical symbols used in official documents of the agency. Make font sample pictures in the form of rectangular frames according to the font categories used in official documents of the agency. The background color of the font sample pictures is pure white, the text color is black, and the font style is not bold; each font sample picture contains a Chinese character, numeral, or symbol; among them, the font categories used include at least one of the following: Founder Small Title Song Simplified, FangSong, FangSong_GB2312, HeiTi, KaiTi, KaiTi_GB2312, SongTi. Use the labelImg annotation software to annotate the character categories in each font sample picture with character category codes, and output a label annotation result file in txt format; each label annotation result file corresponds to its font sample picture with the same name; regard each label annotation result file and its font sample picture with the same name as a data set, and divide the data set into a training set and a validation set; among them, the data in each annotation result file includes: cls, x, y, w, h, where cls is the font category, x and y are the horizontal and vertical coordinates of the center point of the rectangular frame respectively, and w and h are the width value and height value of the rectangular frame. Configuration of the font recognition model: Set the depth control parameter depth_multiple of the font recognition model to 0.33 and the width control parameter width_multiple to 0.50; at 8, 18, and 32 times, set the sizes of 3 prior boxes to (10, 13), (16, 30), (33, 23), (30, 61), (62, 45), (59, 119), (116, 90), (156, 198), (373, 326) respectively, and set the weight file to yolov5s.pt; in the data configuration file VOC.yaml, set the number of categories to 7, set the font category names according to the font category codes, and configure the addresses of the training set and validation set of the font recognition model. For the training set, segment the official document images into single-character images, and adjust each single-character image to a preset size and input it into the YOLOV5 model; the YOLOV5 model includes a backbone network Backbone, a neck Neck, and a head Head; the backbone network Backbone includes Focus, Concat, Conv convolutional blocks, SSP, and BottleneckCSP. Four sub-images of the same size are obtained by Focus for each single-character image slice; the width and height of the sub-images are integrated by Concat, and the number of channels of the input image is increased to 64; a convolution operation with a convolution kernel of 3 and a stride of 2 is performed on the sub-images integrated by Concat using a Conv convolution block, and the first feature image is output; after the first feature image passes through 3 BottleneckCSP modules and Conv convolution blocks and is output, it becomes the second feature image; the second feature image is subjected to max pooling operation by the SSP module; the pooling results are integrated through the Concat connection layer, and the integrated pooling results are subjected to convolution and connection operations through a 14-layer network of the Neck and Head, and the single-character bounding box and font category are output; And a validation set is used for validation to obtain a trained font recognition model.
5. A discriminant system for official document font categories of an institution, characterized in that, Including: A document preprocessing unit, configured to convert a text-type official document into a PDF and then into an official document image in a preset format when the official document to be recognized is of the text type; When the official document to be recognized is an image, the official document is unified into an official document image in a preset format; when the official document to be recognized is of the paper type, a document acquisition terminal is used to photograph the paper-type official document into a picture, and the photographed picture is sequentially subjected to grayscale processing, denoising processing, edge detection, image segmentation, and perspective transformation to obtain an official document image in a preset format; the official document image in the preset format is used as the official document image to be detected; A document segmentation unit, configured to binarize the official document image to be detected, perform horizontal and vertical projections on the binarized official document image respectively, and segment the characters of the official document through black pixels and white interval pixels to obtain an image of each character, and record the position information of each character; A character category recognition unit, configured to perform font recognition on each character image through a trained font recognition model according to each character image, and output a font recognition result after the recognition; the font recognition result includes the font category of the character and the position information of the character; A character category annotation unit, configured to draw a rectangular box for each character on the official document image to be detected according to the position information of each character, annotate the font category code of the corresponding character within the rectangular box, and set a prompt for the font category code when the right mouse button slides over; A display unit, configured to automatically restore the font category code to the corresponding font category name and display it when the right mouse button slides over the font category code on the official document image to be detected; The document preprocessing unit includes a paper-type official document preprocessing subunit, and the paper-type official document preprocessing subunit is specifically configured to: A high-definition camera automatically photographs the paper-type official document to be recognized according to the collected image acquisition instruction, and outputs the original image P of the paper-type official document to be recognized; Create a copy P1 of the original official document image P, and perform grayscale processing on P1; after the grayscale processing is completed, use Gaussian blur to filter out the noise to obtain the edge image E of the original official document image P; specifically: Determine the gradient G(x, y) and direction θ of the P1 edge using the Sobel filter M ; where G x is the vertical edge, which refers to the mutation of the gradient in the x direction, and G y is the horizontal edge, which refers to the mutation of the gradient in the y direction; perform non-maximum suppression on the gradient magnitude in the gradient direction. For each pixel point i of P1, compare the magnitudes of the surrounding 8 neighborhood values along the four types of gradient directions of 0°, 45°, 90°, and 135° in the 3*3 region. If pixel i is the maximum value, keep the pixel point; otherwise, set it to 0; combine the double-threshold algorithm to detect and connect the edges to obtain the edge image E of the original official document image P; Create a copy E1 of E, obtain the set of edge-closed contours in E1, determine the quadrilateral edges from the set of edge-closed contours, and use the quadrilateral edges as the edges of the original official document image P; and segment the target official document image from the original official document image P according to the position information of the quadrilateral edges. Map the target official document image into an image with the size ratio of A4 paper of the standard official document through 4 vertex coordinates, and correct each character of the target official document image into a visually proportionally coordinated style through perspective transformation, and form an official document image in a preset format after correction.
6. The official document font category discrimination system according to claim 5, characterized in that, The official document segmentation unit is specifically used for: Perform binarization processing on the official document image to be detected to obtain a binarized official document image; perform horizontal projection on the binarized official document image on the y-axis to obtain the binary image of each row, perform vertical projection on the binary image of each row on the x-axis, and determine the start position and end position of each character in the official document image to be detected according to the black pixel and white interval pixel obtained by the projection, and obtain the position coordinates of the character according to the start position and end position of each character. Draw a segmentation rectangle frame for each character in the official document image to be detected according to each character position coordinate, segment the characters through the rectangle frame to obtain the image and position of each character, and at the same time number each character image in a preset form, and record the number of each character and the corresponding character position information.
7. The official document font category discrimination system according to claim 6, characterized in that The character category recognition unit is specifically used for: Input each character image into the trained font recognition model, and respectively obtain 4 sub-images of the same size by slicing each character image through the Focus of the backbone network Backbone. Integrate the width and height of each sub-image through Concat, and increase the number of channels of the input image to 64; perform a convolution operation with a convolution kernel of 3 and a stride of 2 on the sub-image integrated by Concat using the Conv convolution block, and output the first feature image; after the first feature image passes through 3 BottleneckCSP modules and the output of the Conv convolution block, it becomes the second feature image; perform a maximum pooling operation on the second feature image through the SSP module; integrate the pooling results through the Concat connection layer, and perform convolution and connection operations on the integrated pooling results through the 14-layer network of the Neck and Head to output the font category of each character, and output the font recognition result after recognition. The character category annotation unit is specifically used for: Use the rectangle() function of the python cv2 package to draw a character position rectangle frame for each character in the official document image to be detected according to the position information of each character. Use the plt.text(x,y,s) function called by the matplotlib package to label the font category code in the character position rectangle frame, where x and y are the abscissa and ordinate of the center point of the character position rectangle frame respectively, s is the font category code of the character, and set to display the font category code according to the font category code when the mouse right button slides over the character position rectangle frame.
8. The mechanism for discriminating official document font categories according to claim 5, wherein, It also includes: The font recognition model training unit, and the font recognition model training unit is specifically used for: Obtain single Chinese characters, Arabic numerals, punctuation marks, and mathematical symbols used in official documents of the institution, and make font sample images in the form of rectangular frames according to the font categories used in official documents of the institution. The background color of the font sample images is pure white, the text color is black, and the font style is not bold; each font sample image contains a Chinese character, numeral, or symbol; among them, the font categories used include at least one of the following: Founder Small Title Song Simplified, FangSong, FangSong_GB2312, HeiTi, KaiTi, KaiTi_GB2312, SongTi. Use the labelImg annotation software to annotate the character categories in each font sample image with character category codes, and output the label annotation result file in txt format; each label annotation result file corresponds to its font sample image with the same name; use each label annotation result file and its font sample image with the same name as the data set, and divide the data set into a training set and a validation set; among them, the data in each annotation result file includes: cls, x, y, w, h, where cls is the font category, x and y are the horizontal and vertical coordinates of the center point of the rectangular frame respectively, and w and h are the width value and height value of the rectangular frame. Configure the font recognition model: set the depth control parameter depth_multiple of the font recognition model to 0.33 and the width control parameter width_multiple to 0.50; at 8, 18, and 32 times, the sizes of 3 prior boxes are set to (10, 13), (16, 30), (33, 23), (30, 61), (62, 45), (59, 119), (116, 90), (156, 198), (373, 326) respectively, and the weight file is set to yolov5s.pt; in the data configuration file VOC.yaml, set the number of categories to 7, set the font category names according to the font category codes, and configure the addresses of the training set and validation set of the font recognition model. For the training set, segment the official document images of the institution into single-character images, and adjust each single-character image to a preset size and input it into the YOLOV5 model; the YOLOV5 model includes a backbone network Backbone, a neck Neck, and a head Head; the backbone network Backbone includes Focus, Concat, Conv convolutional blocks, SSP, and BottleneckCSP. Four sub-images of the same size are obtained by Focus for each single-character image slice; the width and height of the sub-images are integrated by Concat, and the number of channels of the input image is increased to 64; the integrated sub-images by Concat are subjected to a convolution operation with a convolution kernel of 3 and a stride of 2 using a Conv convolution block, and the first feature image is output; after the first feature image is output through 3 BottleneckCSP modules and Conv convolution blocks, it becomes the second feature image; the second feature image is subjected to a maximum pooling operation through the SSP module; the pooling results are integrated through the Concat connection layer, and the integrated pooling results are subjected to convolution and connection operations through a 14-layer network of the Neck and Head, and the single-character bounding box and font category are output; And a validation set is used for validation to obtain a trained font recognition model.
Citation Information
Patent Citations
Character recognition method
CN109190630A
Electronic official document recognition and reproduction method and system based on domestic CPU
CN112949471A