A document classification method and system based on template matching
By adopting a document classification method based on template matching in the procuratorial industry, using text box detection and improved SSIM calculation structure similarity index, the problem of poor document classification effect in the prior art is solved, and fast and accurate document classification and reduced maintenance costs are achieved.
Patent Information
- Application Number
- CN202510180212.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The prior art has poor document classification effect in procuratorial industry scenarios, rules-based classification methods are costly to maintain and difficult to process new types of documents, and content-based classification methods cannot understand the semantic relationship between words, resulting in incorrect classification.
Using a document classification method based on template matching, the text box detection, hierarchical clustering method combines connected components, and the improved SSIM to calculate the structural similarity index, determine the template with the highest similarity score, and establish a mapping relationship to complete the document classification.
It realizes the rapid and accurate document classification in the procuratorial industry scenario, reduces maintenance costs, can effectively process new types of documents, and improves the accuracy of classification.
Smart Images

Figure CN119672744B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of document classification, and in particular relates to a document classification method and system based on template matching. Background Art
[0002] In the prosecution industry, it mainly includes criminal prosecution, civil prosecution, administrative prosecution and public interest litigation prosecution, etc., which contain a large number of file documents that need to be classified. Among them, these file documents are different from general documents and have certain particularities. It is understandable that the file documents in the prosecution industry have strict format requirements, and some documents also require evidence presentation.
[0003] Document classification is the process of dividing a collection of documents into different categories or groups according to certain rules and standards. Its purpose is to better organize, manage and retrieve documents and improve the efficiency of information processing.
[0004] Currently, common document classification methods include rule-based classification methods and content-based classification methods. Among them, rule-based classification methods require manual writing of rules. When the document type is complex or there are many rules, the maintenance cost is high, and it is difficult to flexibly handle newly emerging document types; content-based classification methods mainly use algorithms to extract keywords from documents, such as the TF-IDF algorithm, which classifies documents according to the importance and frequency of keywords. This method only relies on keywords and cannot understand the semantic relationship between words, which may lead to misclassification. Summary of the invention
[0005] Based on this, an embodiment of the present invention provides a document classification method and system based on template matching, aiming to solve the problem of poor classification effect of document classification methods in industry scenarios in the prior art.
[0006] A first aspect of an embodiment of the present invention provides a document classification method based on template matching, which is applied in a procuratorial industry scenario. The method includes:
[0007] Perform text box detection on the document to be classified. Specifically, obtain an image of the document to be classified, and convert the image of the document to be classified into a grayscale image;
[0008] The grayscale image is processed by an adaptive threshold method to obtain a binary layout map, wherein a suitable threshold is automatically determined according to the grayscale characteristics of a local area of the grayscale image to cope with different lighting conditions;
[0009] Using connected component analysis to identify potential text regions of the binary layout graph, and removing noise according to a preset region aspect ratio and area to obtain a target connected component;
[0010] According to the hierarchical clustering method, target connected components with intervals less than a preset distance are combined to obtain text blocks;
[0011] Get the border of the text block and the coordinates and size of the border;
[0012] Obtain the binary layout diagram of the document to be classified and the binary layout diagram of each template, and perform layout similarity calculation to obtain a similarity score. The improved SSIM is used to calculate the structural similarity index, and the calculation formula is:
[0013]
[0014] in, Indicated by ( i, j ) is the center of the window layout density comparison, Indicated by ( i, j ) is the layout contrast comparison of the window centered on Indicated by ( i, j ) is the center of the window structure comparison, represents the spatial weighting function, represents the structural similarity index, α represents the weight of layout density comparison, β represents the weight of layout contrast comparison, and λ represents the weight of structure comparison;
[0015] Determine the score with the highest similarity score and the corresponding template, and determine whether the score exceeds a threshold;
[0016] If so, the template with the highest similarity score is determined as the target template, and a mapping relationship between the document to be classified and the matching area of the target template is established to complete the document classification.
[0017] Furthermore, the calculation formula for the layout density comparison is:
[0018]
[0019] in, Indicated by ( i, j ) is the center of the window layout density comparison, Indicates that the document A to be classified is ( i, j ) is the average layout density of windows centered on Indicates that the template is ( i, j ) is the average layout density of the window centered at 1 Represents the first constant.
[0020] Furthermore, the calculation formula for the layout contrast comparison is:
[0021]
[0022] in, Indicated by ( i, j ) is the layout contrast comparison of the window centered on Indicates that the document A to be classified is ( i, j ) is the variance of the window centered at Indicates that the template is ( i, j ) is the variance of the window centered at Indicates that the document A to be classified is ( i, j ) is the standard deviation of the window centered at Indicates that the template is ( i, j ) is the standard deviation of the window centered at 2 Represents the second constant.
[0023] Furthermore, the calculation formula for the structure comparison is:
[0024]
[0025] in, Indicated by ( i, j ) is the center of the window structure comparison, Represents the document A to be classified and the template ( i, j ) is the covariance of the window centered at Indicates that the document A to be classified is ( i, j ) is the standard deviation of the window centered at Indicates that the template is ( i, j ) is the standard deviation of the window centered at 3 Represents the third constant.
[0026] Furthermore, the calculation formula of the spatial weighting function is:
[0027]
[0028] in, represents the spatial weighting function, ( i 0 ,j 0 ) is the center of the document to be classified, represents an exponential function with the natural constant e as the base, Represents the standard deviation of the Gaussian function.
[0029] Furthermore, the step of obtaining the binary layout diagram of the document to be classified and the binary layout diagram of each template, and performing layout similarity calculation to obtain a similarity score includes:
[0030] Calculating a normalization factor according to the spatial weighting function;
[0031] Calculating a first similarity score according to the normalization factor and the structural similarity index;
[0032] A text line spacing consistency parameter and a page margin alignment parameter are obtained, and a second similarity score is calculated according to the first similarity score, the text line spacing consistency parameter and the page margin alignment parameter.
[0033] Furthermore, in the step of establishing a mapping relationship between the document to be classified and the target template matching area, the borders of the text block of the document to be classified and the borders of the text block of the target template are sorted from top to bottom and from left to right according to the coordinates of the borders, and correspond one to one.
[0034] A second aspect of an embodiment of the present invention provides a document classification system based on template matching, which is used to implement the document classification method based on template matching described in the first aspect. The system includes:
[0035] Detection module, used to detect text boxes on documents to be classified;
[0036] The detection module comprises:
[0037] A conversion unit, used for acquiring an image of a document to be classified, and converting the image of the document to be classified into a grayscale image;
[0038] A processing unit, configured to process the grayscale image using an adaptive threshold method to obtain a binary layout map, wherein a suitable threshold is automatically determined according to the grayscale characteristics of a local area of the grayscale image to cope with different lighting conditions;
[0039] An analysis unit, configured to identify potential text regions of the binary layout graph using connected component analysis, and remove noise according to a preset region aspect ratio and area to obtain a target connected component;
[0040] A combining unit, used for combining target connected components whose intervals are less than a preset distance according to a hierarchical clustering method to obtain a text block;
[0041] An acquisition unit, used for acquiring the border of the text block and the coordinates and size of the border;
[0042] The calculation module is used to obtain the binary layout diagram of the document to be classified and the binary layout diagram of each template, and perform layout similarity calculation to obtain a similarity score. The improved SSIM is used to calculate the structural similarity index, and the calculation formula is:
[0043]
[0044] in, Indicated by ( i, j ) is the center of the window layout density comparison, Indicated by ( i, j ) is the layout contrast comparison of the window centered on Indicated by ( i, j ) is the center of the window structure comparison, represents the spatial weighting function, represents the structural similarity index, α represents the weight of layout density comparison, β represents the weight of layout contrast comparison, and λ represents the weight of structure comparison;
[0045] A judgment module, used to determine the score with the highest similarity score and the corresponding template, and to judge whether the score exceeds a threshold;
[0046] The mapping relationship establishment module is used to determine the template with the highest similarity score as the target template when it is determined that the score exceeds the threshold, and establish a mapping relationship between the document to be classified and the matching area of the target template to complete the document classification.
[0047] A third aspect of an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the document classification method based on template matching provided in the first aspect.
[0048] A fourth aspect of an embodiment of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the document classification method based on template matching provided in the first aspect is implemented.
[0049] A document classification method and system based on template matching is provided in an embodiment of the present invention. The method performs text box detection on the document to be classified, determines the border of the text block and the coordinates and size of the border, obtains the binary layout diagram of the document to be classified and the binary layout diagram of each template, and uses the improved SSIM to calculate the layout similarity to obtain a similarity score; determines the score with the highest similarity score and the corresponding template, and judges whether the score exceeds a threshold; if so, determines the template with the highest similarity score as the target template, and establishes a mapping relationship between the document to be classified and the target template matching area, so as to quickly and accurately complete document classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A flow chart of an implementation of a document classification method based on template matching provided in the first embodiment of the present invention;
[0051] Figure 2 A structural block diagram of a document classification system based on template matching provided in Embodiment 2 of the present invention;
[0052] Figure 3This is a structural block diagram of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION
[0053] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0054] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0056] Embodiment 1
[0057] According to an embodiment of the present invention, a document classification method embodiment based on template matching is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0058] In the first embodiment, a document classification method based on template matching is provided, which can be used in electronic devices, such as computers. Figure 1 , Figure 1 A flowchart of an implementation of a document classification method based on template matching provided in the first embodiment of the present invention is shown, which specifically includes steps S01 to S04.
[0059] Step S01, performing text box detection on the document to be classified.
[0060] Specifically, obtain the image of the document to be classified, convert the image of the document to be classified into a grayscale image, and use the weighted average method to convert the RGB value of each pixel into a grayscale value. The formula is:
[0061]
[0062] in, Gray Represents a grayscale image, R represents the red pixel value, G represents the green pixel value, and B represents the blue pixel value;
[0063] The grayscale image is processed by an adaptive threshold method to obtain a binary layout map, wherein a suitable threshold is automatically determined according to the grayscale characteristics of the local area of the grayscale image to cope with different lighting conditions. It is understandable that since the input document may be under different lighting conditions, the brightness of the local area will vary greatly. The use of adaptive threshold processing can automatically determine a suitable threshold according to the grayscale characteristics of the local area of the image, and divide the image into foreground and background, which is helpful to extract the text area. It should be noted that a mean-based adaptive threshold method can be used, which uses the average grayscale value of the pixel neighborhood plus or minus a constant as the threshold. For example, in OpenCV, the cv2.adaptiveThreshold function can be used, where cv2.ADAPTIVE_THRESH_MEAN_C is a mean-based adaptive threshold type;
[0064] Connected component analysis is used to identify potential text areas in the binary layout graph, and noise is removed according to a preset area aspect ratio and area to obtain a target connected component. In this embodiment, the number of connected components, the label matrix of each component, statistical information (including the area of each component, the bounding rectangle, etc.), and centroid information are returned through the cv2.connectedComponentsWithStats function in OpenCV;
[0065] According to the hierarchical clustering method, the target connected components with intervals less than the preset distance are combined to obtain text blocks. It should be noted that, first, the center position of each connected component is calculated, that is, the centroid information or the center of the circumscribed rectangle is used to calculate the distance matrix between components, and the Euclidean distance or Manhattan distance can be used. Then, the linkage function in the scipy.cluster.hierarchy library is used to cluster the components, and the components with similar distances are merged according to the clustering results;
[0066] Get the border of the text block and the coordinates and size of the border
[0067] Step S02: obtaining the binary layout diagram of the document to be classified and the binary layout diagram of each template, and performing layout similarity calculation to obtain a similarity score.
[0068] In this embodiment, it is assumed that L A and L T They are the binary layout diagrams of the document A to be classified and the template, , where 1 represents the text area, the size is standardized to N×M pixels, and further, the local statistics are calculated in the sliding window. i, j ) for each w×w window centered at:
[0069] The average layout density is:
[0070]
[0071] Indicates that the document A to be classified contains ( i, j ) is the average layout density of windows centered on Indicates that the template ends with ( i, j ) as the average layout density of windows centered on the
[0072] The variance is:
[0073]
[0074] Indicates that the document A to be classified is ( i, j ) is the variance of the window centered at Indicates that the template is ( i, j ) is the variance of the window centered at .
[0075] The covariance is:
[0076]
[0077] Represents the document A to be classified and the template ( i, j ) is the covariance of the window centered at .
[0078] It should be noted that layout density comparison, layout contrast comparison and structure comparison are performed on each local window, wherein the calculation formula for layout density comparison is:
[0079]
[0080] in, Indicated by ( i, j ) is the center of the window layout density comparison, Indicates that the document A to be classified is ( i, j ) is the average layout density of windows centered on Indicates that the template is ( i, j ) is the average layout density of the window centered at 1 Represents the first constant, and the calculation formula for layout contrast comparison is:
[0081]
[0082] in, Indicated by ( i, j ) is the layout contrast comparison of the window centered on Indicates that the document A to be classified is ( i, j ) is the variance of the window centered at Indicates that the template is ( i, j ) is the variance of the window centered at Indicates that the document A to be classified is ( i, j ) is the standard deviation of the window centered at Indicates that the template is ( i, j ) is the standard deviation of the window centered at 2 Represents the second constant, and the calculation formula for structural comparison is:
[0083]
[0084] in, Indicated by ( i, j ) is the center of the window structure comparison, Represents the document A to be classified and the template ( i, j ) is the covariance of the window centered at Indicates that the document A to be classified is ( i, j ) is the standard deviation of the window centered at Indicates that the template is ( i, j ) is the standard deviation of the window centered at 3 represents the third constant, which is understandable, C 1 , C 2 and C 3 is a small constant that avoids division by zero.
[0085] The calculation formula of the spatial weighting function is:
[0086]
[0087] in, represents the spatial weighting function, ( i 0 ,j 0 ) is the center of the document to be classified, represents an exponential function with the natural constant e as the base, Represents the standard deviation of the Gaussian function.
[0088] In this embodiment, the improved SSIM is used to calculate the structural similarity index, and the calculation formula is:
[0089]
[0090] in, Indicated by ( i, j ) is the center of the window layout density comparison, Indicated by ( i, j ) is the layout contrast comparison of the window centered on Indicated by ( i, j ) is the center of the window structure comparison, represents the spatial weighting function, represents the structural similarity index, α represents the weight of layout density comparison, β represents the weight of layout contrast comparison, λ represents the weight of structure comparison, and the sum of α, β and λ is 1.
[0091] In order to make the similarity calculation result more accurate, additional layout-specific items are considered, including text line spacing consistency and page margin alignment. Specifically, according to the spatial weighting function, the normalization factor is calculated, which can be expressed as:
[0092]
[0093] Z is the normalization factor;
[0094] According to the normalization factor and the structural similarity index, a first similarity score is calculated, which can be expressed as:
[0095]
[0096] S is the first similarity score;
[0097] Obtain a text line spacing consistency parameter and a page margin alignment parameter, and calculate a second similarity score according to the first similarity score, the text line spacing consistency parameter, and the page margin alignment parameter, wherein the expression of the text line spacing consistency parameter is:
[0098]
[0099] LS Represents the text line spacing consistency parameter, represents the average line spacing of the document A to be classified, Indicates the average line spacing of the template. represents the standard deviation of the Gaussian function, i.e. a constant;
[0100] The expression for the margin alignment parameter is:
[0101]
[0102] MA Indicates the page margin alignment parameters, represents the margin measurement value of the document A to be classified, Represents the margin measurements of the template, represents the standard deviation of the Gaussian function, i.e. a constant;
[0103] The expression for the second similarity score is:
[0104]
[0105] Score represents the second similarity score, that is, the final similarity score.
[0106] Step S03, determining the score with the highest similarity score and the corresponding template, and judging whether the score exceeds a threshold, if so, executing step S04.
[0107] The threshold value ranges from 0.75 to 0.85. In this embodiment, the threshold value is 0.75.
[0108] Step S04: determine the template with the highest similarity score as the target template, and establish a mapping relationship between the document to be classified and the matching area of the target template to complete the document classification.
[0109] Specifically, since the frame and its coordinates and size are known, the frame of the text block of the document to be classified and the frame of the text block of the target template are sorted from top to bottom and from left to right according to the coordinates of the frame, and correspond one to one.
[0110] In summary, the document classification method based on template matching in the above-mentioned embodiment of the present invention performs text box detection on the document to be classified, determines the border of the text block as well as the coordinates and size of the border, and then obtains the binary layout map of the document to be classified and the binary layout map of each template, and uses the improved SSIM to calculate the layout similarity to obtain a similarity score; determines the score with the highest similarity score and the corresponding template, and judges whether the score exceeds the threshold; if so, determines the template with the highest similarity score as the target template, and establishes a mapping relationship between the document to be classified and the target template matching area, so as to quickly and accurately complete document classification.
[0111] Embodiment 2
[0112] See also Figure 2 , Figure 2 This is a structural block diagram of a document classification system based on template matching provided in Embodiment 2 of the present invention. The document classification system 200 based on template matching is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" may be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation in hardware, or a combination of software and hardware, is also possible and conceivable.
[0113] Specifically, the document classification system 200 based on template matching includes: a detection module 21, a calculation module 22, a judgment module 23 and a mapping relationship establishment module 24, wherein:
[0114] A detection module 21, used for performing text box detection on the document to be classified;
[0115] The detection module 21 includes:
[0116] A conversion unit, used for acquiring an image of a document to be classified, and converting the image of the document to be classified into a grayscale image;
[0117] A processing unit, configured to process the grayscale image using an adaptive threshold method to obtain a binary layout map, wherein a suitable threshold is automatically determined according to the grayscale characteristics of a local area of the grayscale image to cope with different lighting conditions;
[0118] An analysis unit, configured to identify potential text regions of the binary layout graph using connected component analysis, and remove noise according to a preset region aspect ratio and area to obtain a target connected component;
[0119] A combining unit, used for combining target connected components whose intervals are less than a preset distance according to a hierarchical clustering method to obtain a text block;
[0120] An acquisition unit, used for acquiring the border of the text block and the coordinates and size of the border;
[0121] The calculation module 22 is used to obtain the binary layout diagram of the document to be classified and the binary layout diagram of each template, and perform layout similarity calculation to obtain a similarity score, wherein the improved SSIM is used to calculate the structural similarity index, and the calculation formula is:
[0122]
[0123] in, Indicated by ( i, j ) is the center of the window layout density comparison, Indicated by ( i, j ) is the layout contrast comparison of the window centered at Indicated by ( i, j ) is the center of the window structure comparison, represents the spatial weighting function, represents the structural similarity index, α represents the weight of layout density comparison, β represents the weight of layout contrast comparison, and λ represents the weight of structure comparison. The calculation formula of layout density comparison is:
[0124]
[0125] in, Indicated by ( i, j ) is the center of the window layout density comparison, Indicates that the document A to be classified is ( i, j ) is the average layout density of windows centered on Indicates that the template is ( i, j ) is the average layout density of the window centered at 1 represents the first constant, and the calculation formula for the layout contrast comparison is:
[0126]
[0127] in, Indicated by ( i, j ) is the layout contrast comparison of the window centered on Indicates that the document A to be classified is ( i, j ) is the variance of the window centered at Indicates that the template is ( i, j ) is the variance of the window centered at Indicates that the document A to be classified is ( i, j ) is the standard deviation of the window centered at Indicates that the template is ( i, j ) is the standard deviation of the window centered at 2 represents the second constant, and the calculation formula for the structure comparison is:
[0128]
[0129] in, Indicated by ( i, j ) is the center of the window structure comparison, Represents the document A to be classified and the template ( i, j ) is the covariance of the window centered at Indicates that the document A to be classified is ( i, j ) is the standard deviation of the window centered at Indicates that the template is ( i, j ) is the standard deviation of the window centered at 3 represents the third constant, and the calculation formula of the spatial weighting function is:
[0130]
[0131] in, represents the spatial weighting function, ( i 0 ,j 0 ) is the center of the document to be classified, represents an exponential function with the natural constant e as the base, represents the standard deviation of the Gaussian function;
[0132] A judgment module 23 is used to determine the score with the highest similarity score and the corresponding template, and to judge whether the score exceeds a threshold;
[0133] The mapping relationship establishment module 24 is used to determine the template with the highest similarity score as the target template when it is determined that the score exceeds the threshold, and establish a mapping relationship between the document to be classified and the matching area of the target template to complete the document classification, wherein the border of the text block of the document to be classified and the border of the text block of the target template are sorted from top to bottom and from left to right according to the coordinates of the border, and correspond one to one.
[0134] Furthermore, in some optional embodiments of the present invention, the calculation module 22 includes:
[0135] A first calculation unit, configured to calculate a normalization factor according to the spatial weighting function;
[0136] a second calculation unit, configured to calculate a first similarity score according to the normalization factor and the structural similarity index;
[0137] The third calculation unit is used to obtain a text line spacing consistency parameter and a page margin alignment parameter, and calculate a second similarity score according to the first similarity score, the text line spacing consistency parameter and the page margin alignment parameter.
[0138] Embodiment 3
[0139] Another aspect of the present invention provides an electronic device, see Figure 3 , shown is an electronic device in Embodiment 3 of the present invention, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor, wherein when the processor 10 executes the computer program 30, the document classification method based on template matching as described above is implemented.
[0140] In some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor or other data processing chip, used to run program codes or process data stored in the memory 20, such as executing access restriction programs.
[0141] Among them, the memory 20 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 20 can be an internal storage unit of an electronic device, such as a hard disk of the electronic device. In other embodiments, the memory 20 can also be an external storage device of an electronic device, such as a plug-in hard disk equipped on the electronic device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. Further, the memory 20 can also include both an internal storage unit of the electronic device and an external storage device. The memory 20 can not only be used to store application software and various types of data of the electronic device, but also can be used to temporarily store data that has been output or is to be output.
[0142] It should be pointed out that Figure 3 The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than those shown in the figure, or combine certain components, or arrange the components differently.
[0143] The embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the document classification method based on template matching as described above is implemented.
[0144] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For purposes of this specification, "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0145] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0146] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0147] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0148] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the attached claims.
Claims
1. A document classification method based on template matching, characterized in that: Applied in the procuratorial industry scenario, the method includes: Perform text box detection on the document to be classified. Specifically, obtain an image of the document to be classified, and convert the image of the document to be classified into a grayscale image; The grayscale image is processed by an adaptive threshold method to obtain a binary layout map, wherein a suitable threshold is automatically determined according to the grayscale characteristics of a local area of the grayscale image to cope with different lighting conditions; Using connected component analysis to identify potential text regions of the binary layout graph, and removing noise according to a preset region aspect ratio and area to obtain a target connected component; According to the hierarchical clustering method, target connected components with intervals less than a preset distance are combined to obtain text blocks; Get the border of the text block and the coordinates and size of the border; Obtain the binary layout diagram of the document to be classified and the binary layout diagram of each template, and perform layout similarity calculation to obtain a similarity score. The improved SSIM is used to calculate the structural similarity index, and the calculation formula is: in, Indicated by ( i, j ) is the center of the window layout density comparison, Indicated by ( i, j ) is the layout contrast comparison of the window centered at Indicated by ( i, j ) is the center of the window structure comparison, represents the spatial weighting function, represents the structural similarity index, α represents the weight of layout density comparison, β represents the weight of layout contrast comparison, and λ represents the weight of structure comparison; Determine the score with the highest similarity score and the corresponding template, and determine whether the score exceeds a threshold; If so, the template with the highest similarity score is determined as the target template, and a mapping relationship between the document to be classified and the matching area of the target template is established to complete the document classification; The step of obtaining the binary layout diagram of the document to be classified and the binary layout diagram of each template, and calculating the layout similarity to obtain the similarity score includes: Calculating a normalization factor according to the spatial weighting function; Calculating a first similarity score according to the normalization factor and the structural similarity index; A text line spacing consistency parameter and a page margin alignment parameter are obtained, and a second similarity score is calculated according to the first similarity score, the text line spacing consistency parameter and the page margin alignment parameter, where the second similarity score is the final similarity score.
2. The document classification method based on template matching according to claim 1, characterized in that: The calculation formula for the layout density comparison is: in, Indicated by ( i, j ) is the center of the window layout density comparison, Indicates that the document A to be classified is ( i, j ) is the average layout density of windows centered on Indicates that the template is ( i, j ) is the average layout density of windows centered on , and C1 represents the first constant.
3. The document classification method based on template matching according to claim 2 is characterized in that: The calculation formula for the layout contrast comparison is: in, Indicated by ( i, j ) is the layout contrast comparison of the window centered at Indicates that the document A to be classified is ( i, j ) is the variance of the window centered at Indicates that the template is ( i, j ) is the variance of the window centered at Indicates that the document A to be classified is ( i, j ) is the standard deviation of the window centered at Indicates that the template is ( i, j ) is the standard deviation of the window centered on , and C2 represents the second constant.
4. The document classification method based on template matching according to claim 3 is characterized in that: The calculation formula for the structure comparison is: in, Indicated by ( i, j ) is the center of the window structure comparison, Represents the document A to be classified and the template ( i, j ) is the covariance of the window centered at Indicates that the document A to be classified is ( i, j ) is the standard deviation of the window centered at Indicates that the template is ( i, j ) is the standard deviation of the window centered on , and C3 represents the third constant.
5. The document classification method based on template matching according to claim 4 is characterized in that: The calculation formula of the spatial weighting function is: in, represents the spatial weighting function, ( i 0 ,j 0) is the center of the document to be classified, represents an exponential function with the natural constant e as the base, Represents the standard deviation of the Gaussian function.
6. The document classification method based on template matching according to claim 5 is characterized in that: In the step of establishing a mapping relationship between the document to be classified and the target template matching area, the borders of the text block of the document to be classified and the borders of the text block of the target template are sorted from top to bottom and from left to right according to the coordinates of the borders, and correspond one to one.
7. A document classification system based on template matching, characterized in that: The system for implementing the document classification method based on template matching according to any one of claims 1 to 6 comprises: Detection module, used to detect text boxes on documents to be classified; The detection module comprises: A conversion unit, used for acquiring an image of a document to be classified, and converting the image of the document to be classified into a grayscale image; A processing unit, configured to process the grayscale image using an adaptive threshold method to obtain a binary layout map, wherein a suitable threshold is automatically determined according to the grayscale characteristics of a local area of the grayscale image to cope with different lighting conditions; An analysis unit, configured to identify potential text regions of the binary layout graph using connected component analysis, and remove noise according to a preset region aspect ratio and area to obtain a target connected component; A combining unit, used for combining target connected components whose intervals are less than a preset distance according to a hierarchical clustering method to obtain a text block; An acquisition unit, used for acquiring the border of the text block and the coordinates and size of the border; The calculation module is used to obtain the binary layout diagram of the document to be classified and the binary layout diagram of each template, and perform layout similarity calculation to obtain a similarity score. The improved SSIM is used to calculate the structural similarity index, and the calculation formula is: in, Indicated by ( i, j ) is the center of the window layout density comparison, Indicated by ( i, j ) is the layout contrast comparison of the window centered at Indicated by ( i, j ) is the center of the window structure comparison, represents the spatial weighting function, represents the structural similarity index, α represents the weight of layout density comparison, β represents the weight of layout contrast comparison, and λ represents the weight of structure comparison; A judgment module, used to determine the score with the highest similarity score and the corresponding template, and to judge whether the score exceeds a threshold; A mapping relationship establishment module, which is used to determine the template with the highest similarity score as the target template when it is determined that the score exceeds the threshold, and establish a mapping relationship between the document to be classified and the matching area of the target template to complete the document classification; The calculation module comprises: A first calculation unit, configured to calculate a normalization factor according to the spatial weighting function; a second calculation unit, configured to calculate a first similarity score according to the normalization factor and the structural similarity index; The third calculation unit is used to obtain a text line spacing consistency parameter and a page margin alignment parameter, and calculate a second similarity score according to the first similarity score, the text line spacing consistency parameter and the page margin alignment parameter, where the second similarity score is a final similarity score.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the document classification method based on template matching as described in any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the document classification method based on template matching as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Detection method and system for certificate document image tampering
CN114419633A
Joint processing of spectral data to obtain contrast distribution in material basis maps
WO2025019291A2