Letter association law and regulation automatic identification method and identification system
Through image preprocessing and chapter recognition technology, the problem of inefficient identification of correspondence and regulations in the existing technology is solved, and efficient, accurate identification and clear display of diverse legal and regulatory documents is achieved, thereby improving user experience.
Patent Information
- Application Number
- CN202510879287.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The prior art is inefficient in the identification of correspondence between letters and regulations, and its accuracy is difficult to guarantee. Especially when dealing with diverse formats of legal and regulatory documents, it is impossible to accurately identify the structure and content information of the chapter catalog, resulting in confusion and disorderly information extraction and display.
Image preprocessing technology is used to divide text lines and character columns into grayscale, noise removal, tilt correction, horizontal and vertical projection methods, combine standard character template libraries and regular expression matching, determine chapter levels through font features and numbering structures, and form a tree structure to display the chapter catalogs and contents of legal and regulatory documents.
It achieves compatibility with multiple file formats, improves text recognition and the accuracy of the catalog structure, ensures reasonable extraction and clear display of content information, and improves user experience and work efficiency.
Smart Images

Figure CN120388388A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document recognition, and particularly to a method and a system for automatically recognizing relevant regulations of correspondence. Background Art
[0002] There are many deficiencies in the existing technology for the associated recognition of correspondence and regulations, resulting in low work efficiency and difficult to guarantee accuracy. The traditional method for processing laws and regulations documents mostly relies on manual access and analysis. Staff need to spend a lot of time and energy in a mountain of paper documents or electronic documents, word by word to find the regulations related to the correspondence. This method is not only extremely inefficient, but also very easy to miss important information due to human negligence. With the continuous increase in the number of documents, the difficulty of manual processing increases exponentially and it is difficult to meet the growing work needs.
[0003] In terms of digital recognition technology, although OCR technology has been widely used, there are still significant defects in processing laws and regulations documents. On the one hand, the sources of laws and regulations documents are diverse and the formats are complex, including different forms such as PDF, pictures and scanned documents. The existing OCR systems have poor compatibility with these diverse formats and often have recognition errors or cannot recognize. On the other hand, even if the text can be successfully recognized, for the complex chapter and section directory structure in the document, ordinary OCR technology lacks effective recognition and parsing capabilities and cannot accurately sort out the hierarchical relationship of the regulations content, resulting in great difficulties in subsequent information extraction and association work.
[0004] In terms of the content analysis after text recognition, the existing technology lacks the intelligent processing ability for the characteristics of laws and regulations documents. It cannot accurately recognize the chapter and section directory structure in the document, and cannot accurately judge the chapter level according to key elements such as title prefix, font characteristics and numbering structure, resulting in chaotic extraction and organization of the regulations content. At the same time, in the process of content information extraction and association, it is also difficult to reasonably match the chapter and section directory with the corresponding text content and cannot provide users with clear and organized display of regulations information. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and a system for automatically recognizing relevant regulations of correspondence, and solve the technical problems proposed in the background art.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A method for automatically recognizing relevant regulations of correspondence includes the following steps:
[0008] File uploading: receiving the laws and regulations documents uploaded by users;
[0009] Preprocessing: Perform image preprocessing on laws and regulations documents;
[0010] Text recognition: Use OCR technology to recognize the text content in laws and regulations documents;
[0011] Chapter recognition: Based on pre-set chapter rules, recognize the chapter and section directory structure in laws and regulations documents;
[0012] Content extraction: Extract corresponding content information according to the chapter and section directory structure;
[0013] Tree structure construction: Organize the extracted chapter and section directories and content information into a tree structure;
[0014] Result display: Display the tree structure to the user.
[0015] As a further solution of the present invention: The laws and regulations documents uploaded by the user are any one of the PDF, picture, and scanned copy formats.
[0016] As a further solution of the present invention: The image preprocessing method is as follows:
[0017] Step Y1, Image grayscale conversion:
[0018] Convert the color image to a grayscale image;
[0019] The grayscale conversion formula is: H = 0.299×R + 0.587×G + 0.114×B;
[0020] where R, G, and B are the red, green, and blue components of the pixel point respectively, and H is the grayscale value;
[0021] Step Y2, Noise removal:
[0022] Use median filtering or mean filtering to remove noise in the grayscale image;
[0023] First, for each pixel point, take the grayscale values of all pixels within its 3×3 neighborhood;
[0024] For median filtering, sort the grayscale values of all pixels in ascending order and take the median value as the new value of this pixel;
[0025] For mean filtering, calculate the average value of all pixel grayscale values and take this average value as the new value of this pixel;
[0026] Step Y3, Binarization processing:
[0027] Convert the grayscale image to a black and white binary image, and the method is as follows:
[0028] First, through: ;
[0029] Calculate the binarization threshold HY;
[0030] Where Hi is the gray value of the i-th pixel, i = 1, 2, …… n, and n is the total number of pixels in the image;
[0031] Traverse all pixels. If Hi≥HY, then adjust the gray value of this pixel to 255, which is white; if Hi<HY, then adjust the gray value of this pixel to 0, which is black;
[0032] Step Y4, skew correction:
[0033] Detect the skew angle of the text in the black-and-white binary image, and perform rotation correction on the text in the black-and-white binary image according to it;
[0034] The skew angle detection method is as follows:
[0035] First, use the Canny operator to perform edge detection on the binarized image;
[0036] Then, use the Hough transform to detect straight lines and obtain the skew angles of all text lines;
[0037] Finally, take the average value of the skew angles of all text lines as the final skew angle;
[0038] The rotation correction method is as follows:
[0039] According to the final skew angle, determine the rotation matrix, and then perform an affine transformation on the black-and-white binary image according to the rotation matrix to make the text lines horizontally aligned;
[0040] The rotation matrix is: .
[0041] As a further solution of the present invention: the text content recognition method is as follows:
[0042] Step W1, line segmentation:
[0043] Use the horizontal projection method to perform horizontal projection on the black-and-white binary image, then count the number of black pixels in each row and mark it as H(y);
[0044] Where y represents the row number variable of the pixels in the black-and-white binary image;
[0045] Compare the Hy value corresponding to each row with the pre-determined threshold PH:
[0046] When H(y)>PH, it is determined that this row of the black-and-white binary image is a text line;
[0047] When H(y)≤PH, it is determined that this row of the black-and-white binary image is a line spacing;
[0048] Then, according to the line number variable, the lines that are continuously determined to be text lines are recorded as text line regions;
[0049] Step W2, Character Segmentation:
[0050] Use the vertical projection method to perform vertical projection on the text line region in the black and white binary image, then count the number of black pixels in each column and mark it as L(x);
[0051] Among them, x represents the column number variable of the pixels in the text line region;
[0052] Compare the Lx value corresponding to each line with the pre-determined judgment threshold PL:
[0053] When L(x) > PL, it is determined that this column of the text line region is a character column;
[0054] When L(x) ≤ PL, it is determined that this column of the text line region is a character gap column;
[0055] Then, according to the column number variable, the columns that are continuously determined to be text lines are recorded as character regions;
[0056] Step W3, Character Recognition:
[0057] Convert each segmented character region into the corresponding text symbol, and the specific method is as follows:
[0058] Use common Chinese characters, numbers, and punctuation marks to establish a standard character template library;
[0059] Then perform scaling processing on the character region to be recognized, specifically by adjusting the size of the character region to be the same as the size of the characters in the standard character template library;
[0060] Then through: ;
[0061] Calculate the similarity SX between the character region to be recognized and each standard character template in the standard character template library:
[0062] In the formula, E1 r1 is the pixel value of the r1-th pixel coordinate of the standard character template, E2 r2 is the pixel value of the r2-th pixel coordinate of the character region to be recognized, r1 = 1, 2,..., u1, u1 is the total number of pixels of the standard character template, r2 = 1, 2,..., u2, u2 is the total number of pixels of the character region to be recognized;
[0063] After that, select the standard character template with a similarity SX higher than the pre-set similarity threshold and the highest similarity SX as the recognition result of the character region to be recognized.
[0064] As a further solution of the present invention, the recognition method of the chapter and section directory structure is as follows:
[0065] Step Z1: Match the title prefix based on regular expressions;
[0066] First, pre-define a rule pattern set using a regular expression library;
[0067] Then, match the character recognition result of each text line area with each rule pattern in the rule pattern set;
[0068] When the character recognition result of a text line area matches any rule pattern, mark this text line area as a candidate title line;
[0069] When the character recognition result of a text line area does not match all rule patterns, mark this text line area as a non-title line, also denoted as a body text line;
[0070] Step Z2: Verify the title through font features;
[0071] On the text line area corresponding to the candidate title line, select a character area, and obtain its maximum pixel height and maximum pixel width, and mark them as H1 and W1 respectively;
[0072] At the same time, on the text line area corresponding to the body text line, select a character area, and obtain its maximum pixel height and maximum pixel width, and mark them as H2 and W2 respectively;
[0073] Then compare the ratio of H1 to H2 and the ratio of W1 to W2 with the preset judgment ratio PB respectively:
[0074] When either H1 / H2≥PB or W1 / W2≥PB holds, determine the text line area corresponding to this candidate title line as a title line;
[0075] When neither H1 / H2≥PB nor W1 / W2≥PB holds, do not determine the text line area corresponding to this candidate title line as a title line;
[0076] Step Z3: Determine the level according to font features and numbering structure;
[0077] Step Z3.1: Determine the font feature level
[0078] First, extract all determined title lines, select a character area from them, and obtain its maximum pixel height and maximum pixel width;
[0079] Then, sort the maximum pixel height values or maximum pixel width values of the selected character areas corresponding to each title line in descending order, and determine the level of the title line according to the sorting result;
[0080] Among them, when the maximum pixel height or maximum pixel width of any two or more title lines is the same, they are sorted side by side and are at the same level;
[0081] When the sorting order of a title line according to the maximum pixel height is different from its sorting order according to the maximum pixel width, the level with a higher sorting order is taken as the level of the title line;
[0082] Step Z3.2, Determination of the hierarchical structure of the numbering
[0083] Extract the preset numbering structure, which is reflected by the number separator;
[0084] Then, determine the level of the title line according to the number of number separators contained in the text line area corresponding to the title line;
[0085] Specifically:
[0086] When there are 0 number separators in the text line area corresponding to the title line, the title line is recorded as the chapter title of the first level;
[0087] When there is 1 number separator in the text line area corresponding to the title line, the title line is recorded as the chapter title of the second level;
[0088] And so on, when there are N - 1 number separators in the text line area corresponding to the title line, the title line is recorded as the chapter title of the Nth level.
[0089] As a further solution of the present invention: The maximum pixel width is defined as the maximum pixel value occupied by the selected character area in the vertical direction; the maximum pixel height is defined as the maximum pixel value occupied by the selected character area in the horizontal direction.
[0090] As a further solution of the present invention: When determining the font feature level, other character areas are also re - selected in each title line for verification.
[0091] As a further solution of the present invention: The method for determining the hierarchical structure of the numbering is to perform a secondary hierarchical division on the title lines of the same level on the basis of determining the font feature level.
[0092] As a further solution of the present invention: The method for extracting content information is as follows:
[0093] First, all the text lines between two adjacent title lines of the same level are used as the text content in the upper title line among the two adjacent title lines of the same level, and they are associated to form a "title - content" pair.
[0094] As a further solution of the present invention: The method for constructing the tree - shaped structure is as follows:
[0095] First, according to the font feature hierarchy, tree nodes are respectively constructed for the top-level chapter title lines;
[0096] For each lower-level chapter title under each top-level title, it is used as a direct child node of the corresponding top-level tree node;
[0097] And so on, continuously setting the subsequent-level chapter titles as subordinate child nodes of the previous-level child nodes until all chapter title levels are determined in the tree structure;
[0098] Subsequently, according to the numbering structure hierarchy, the remaining hierarchical structure is determined;
[0099] Among the established direct child nodes and subordinate child nodes, refined nodes are further divided according to the same child node division rule;
[0100] Meanwhile, each refined node is combined with its corresponding "title-content" pair, and finally a complete tree structure is formed.
[0101] A correspondence regulation automatic recognition system for letters, which is used to execute a correspondence regulation automatic recognition method. The system includes:
[0102] A file upload module for users to upload laws and regulations files;
[0103] A preprocessing module for image preprocessing of laws and regulations files;
[0104] An OCR recognition module for performing OCR recognition on the uploaded files and extracting text content;
[0105] A chapter extraction module for extracting chapter headings and corresponding content information from the text content recognized by OCR;
[0106] A tree structure construction module for automatically constructing a tree structure according to the extracted chapter headings and content information;
[0107] A user interface module for providing an interface for users to view the tree structure.
[0108] Advantages of the present invention:
[0109] Strong file format compatibility: The laws and regulations files uploaded by users can be any one of PDF, picture, and scanned copy formats, which can adapt to a variety of common file types, meet the file storage and upload habits of different users, and improve the applicability and flexibility of the system.
[0110] Good image preprocessing effect: By adopting a scientific grayscale formula, it can accurately convert a color image into a grayscale image, laying a foundation for subsequent processing. It provides two methods, median filtering and mean filtering, to remove noise, and the appropriate method can be selected according to the actual situation, effectively improving the image quality. The grayscale image is binarized by calculating the binarization threshold, converting the image into a black-and-white binary image, which is beneficial to the recognition and processing of text. The Canny operator and Hough transform are used to detect the text tilt angle and perform rotation correction to align the text lines horizontally, ensuring the accuracy of text recognition and reducing recognition errors caused by text tilt.
[0111] Accurate and efficient text recognition: The horizontal projection method is used to count the number of black pixels in each row. By comparing with the determination threshold, the text lines and line spacing are accurately segmented to determine the text line area. The vertical projection method is used to process the text line area, and the character columns and character gap columns can be accurately segmented to determine the character area. A standard character template library is established. Through scaling processing and similarity calculation, the standard character template with the highest similarity and higher than the threshold is selected as the recognition result, improving the accuracy of character recognition.
[0112] Accurate recognition of chapter and section directory structure: It can quickly and accurately screen out candidate title lines, improving the efficiency of title recognition. Combining the pixel height and width characteristics of the font for verification further ensures the accuracy of the title lines and avoids misjudgment. The hierarchy of the title lines is comprehensively determined from two aspects: font characteristics and numbering structure. It takes into account both the differences in font display and the rules of the numbering structure, making the recognition of the chapter and section directory structure more accurate. At the same time, the method for determining the numbering structure hierarchy performs a secondary hierarchy division on the basis of the determination of the font characteristics hierarchy, further refining the hierarchy structure.
[0113] Reasonable extraction of content information: All the text lines between two adjacent title lines of the same level are used as the text content of the previous title line and associated to form a "title - content" pair. This content information extraction method conforms to the structural characteristics of legal and regulatory documents and can accurately extract the content information corresponding to the title.
[0114] Scientific construction of tree structure: The tree structure is gradually constructed according to the font characteristics hierarchy and the numbering structure hierarchy, organically combining the chapter titles and content information to form a complete and clear tree structure, which is convenient for users to intuitively view the structure and content of legal and regulatory documents and helps users quickly locate and understand the required regulatory information.
[0115] The system has comprehensive functions and good user experience: The automatic identification system for correspondence-related regulations includes multiple modules such as file upload, preprocessing, OCR recognition, section extraction, tree structure construction, and user interface. These modules work together to achieve full-process automated processing from file upload to result display, and provide a user interface for viewing the tree structure, which is convenient to operate and improves the user experience. Brief Description of the Drawings
[0116] The present invention will be further described below in conjunction with the accompanying drawings.
[0117] Figure 1 It is a system block diagram of an automatic identification method and system for correspondence-related regulations of the present invention.
[0118] Figure 2 It is a schematic flow diagram of an automatic identification method and system for correspondence-related regulations of the present invention. Detailed Embodiments
[0119] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0120] Embodiment 1
[0121] Please refer to Figure 1 and Figure 2 As shown, the present invention is an automatic identification method for correspondence-related regulations, including:
[0122] The first step, file upload:
[0123] Receive the laws and regulations files uploaded by the user;
[0124] Among them, the laws and regulations files uploaded by the user are any one of the formats of PDF, pictures, and scanned documents;
[0125] The second step, text recognition:
[0126] Use OCR technology to recognize the text content in the file, and the specific method is as follows:
[0127] Step W1, line segmentation:
[0128] Use the horizontal projection method to perform horizontal projection on the black and white binary image, then count the number of black pixels in each row and mark it as H(y);
[0129] Among them, y represents the row serial number variable of the pixels in the black and white binary image;
[0130] Compare the Hy value corresponding to each line with the preset determination threshold PH:
[0131] When H(y) > PH, it is determined that this line of the black-and-white binary image is a text line;
[0132] When H(y) ≤ PH, it is determined that this line of the black-and-white binary image is a line spacing;
[0133] In this embodiment, the determination threshold PH takes a value of 1% - 3% of the width of the black-and-white binary image;
[0134] Then, according to the line number variable, the lines continuously determined as text lines are recorded as the text line area;
[0135] Step W2, Character segmentation:
[0136] Use the vertical projection method to perform vertical projection on the text line area in the black-and-white binary image, then count the number of black pixels in each column and mark it as L(x);
[0137] Among them, x represents the column number variable of the pixels in the text line area;
[0138] Compare the Lx value corresponding to each line with the preset determination threshold PL:
[0139] When L(x) > PL, it is determined that this column in the text line area is a character column;
[0140] When L(x) ≤ PL, it is determined that this column in the text line area is a character gap column;
[0141] In this embodiment, the determination threshold PH takes a value of 5% - 10% of the height of the text line area;
[0142] Then, according to the column number variable, the columns continuously determined as text lines are recorded as the character area;
[0143] Step W3, Character recognition:
[0144] Convert each segmented character area into the corresponding text symbol, and the specific method is as follows:
[0145] Use common Chinese characters, numbers, and punctuation marks to establish a standard character template library;
[0146] Then perform scaling processing on the character area to be recognized, specifically by adjusting the size of the character area to be the same as the size of the characters in the standard character template library;
[0147] Then through:
[0148] Calculate the similarity SX between the character area to be recognized and each standard character template in the standard character template library:
[0149] Wherein, E1 r1 is the pixel value of the r1-th pixel coordinate of the standard character template, and E2 r2 is the pixel value of the r2-th pixel coordinate of the character region to be recognized, where r1 = 1, 2,..., u1, u1 is the total number of pixels of the standard character template, r2 = 1, 2,..., u2, and u2 is the total number of pixels of the character region to be recognized;
[0150] After that, select the standard character template with a similarity SX higher than the pre-set similarity threshold and the highest similarity SX as the recognition result of the character region to be recognized;
[0151] Step 3, chapter recognition:
[0152] Based on the pre-set chapter rules, recognize the chapter and section directory structure in the laws and regulations documents. The specific method is as follows:
[0153] Step Z1, match the title prefix based on the regular expression;
[0154] First, pre-define a rule pattern set using the regular expression library;
[0155] In this embodiment, the rule pattern set includes: the rule pattern corresponding to "(i)+th chapter / section / article", where i is a quantifier, which can be "1 / 2 / 3 / 4...", or "one / two / three / four...";
[0156] Then, match the character recognition result of each text line region with each rule pattern in the rule pattern set;
[0157] When the character recognition result of a text line region matches any rule pattern, mark this text line region as a candidate title line;
[0158] When the character recognition result of a text line region does not match all rule patterns, mark this text line region as a non-title line, also denoted as a text line;
[0159] Step Z2, verify the title through font features;
[0160] On the text line region corresponding to the candidate title line, select a character region, and obtain its maximum pixel height and maximum pixel width, and mark them as H1 and W1 respectively;
[0161] Among them, the maximum pixel width refers to the maximum pixel value occupied by the selected character region in the vertical direction; the maximum pixel height refers to the maximum pixel value occupied by the selected character region in the horizontal direction;
[0162] At the same time, a character area is selected in the text line area corresponding to the text line, and its maximum pixel height and maximum pixel width are obtained, and they are marked as H2 and W2 respectively;
[0163] Then, the ratio of H1 to H2 and the ratio of W1 to W2 are compared with the preset judgment ratio PB:
[0164] When any one of H1 / H2≥PB and W1 / W2≥PB is true, the text row area corresponding to the candidate title row is determined as the title row;
[0165] When H1 / H2≥PB and W1 / W2≥PB are both not true, the text row area corresponding to the candidate title row is not determined as the title row;
[0166] Step Z3: Determine the level based on the font characteristics and numbering structure;
[0167] Step Z3.1: Determine font feature level
[0168] First, extract all the determined title rows, select a character area from them, and obtain its maximum pixel height and maximum pixel width;
[0169] Then, the maximum pixel height value or the maximum pixel width value of the selected character area corresponding to each title row is sorted in descending order, and the level of the title row is determined according to the sorting result;
[0170] Among them, when any two or more title rows have the same maximum pixel height or maximum pixel width, they are sorted in parallel and are at the same level;
[0171] When the sorting order of a title row by maximum pixel height is different from the sorting order by maximum pixel width, the first-order level is used as the level of the title row;
[0172] In this embodiment, verification is also performed by reselecting other character areas in each title row;
[0173] For example:
[0174] Assume there are 5 heading rows (H1, H2, H3, H4, H5), and extract their maximum pixel height and maximum pixel width;
[0175] The details are as follows:
[0176] The maximum pixel height of the title row H1 is 40, and the maximum pixel width is 500;
[0177] The maximum pixel height of the header line H2 is 32, and the maximum pixel width is 450;
[0178] The maximum pixel height of the header line H3 is 32, and the maximum pixel width is 400;
[0179] The maximum pixel height of the H4 header row is 28, and the maximum pixel width is 450;
[0180] The maximum pixel height of the title row H5 is 24, and the maximum pixel width is 380;
[0181] Step 1: Sort by height in descending order
[0182] H1 corresponds to the first level, H2 and H3 correspond to the second level, H4 corresponds to the third level, and H5 corresponds to the fourth level;
[0183] Step 2: Sort by width in descending order
[0184] H1 corresponds to the first level, H2 and H4 correspond to the second level, H3 corresponds to the third level, and H5 corresponds to the fourth level;
[0185] Step 3: Comprehensive comparison to determine the final level
[0186] H1: It is ranked first in both height and width sorting and is ultimately determined to be the first level;
[0187] H2: It is ranked 2nd in height sorting, 2nd in width sorting, and finally ranked 2nd;
[0188] H3: It is the second level in the height sorting and the third level in the width sorting. The higher level is the second level.
[0189] H4: In the height sorting, it is the third level, and in the width sorting, it is the second level. The higher level is the second level.
[0190] H5: It is ranked 4th in both height and width. If there is no 3rd level among H1, H2, H3, and H4, it will eventually be promoted from 4th level to 3rd level.
[0191] Step Z3.2: Determine the numbering structure level
[0192] The numbering structure level is determined by dividing the title lines of the same level into two levels based on the font feature level.
[0193] Extract the preset numbering structure, which is reflected by the number separator;
[0194] In this embodiment, the number separators are such as "." and "-", which shall be determined according to the actual document conventions;
[0195] Then, the level of the title row is determined according to the number of number separators in the text row area corresponding to the title row;
[0196] Specifically:
[0197] When there are 0 number separators in the text line area corresponding to the title line, the title line is recorded as the chapter title of the first level;
[0198] When there is 1 number separator in the text line area corresponding to the title line, the title line is recorded as the chapter title of the second level;
[0199] And so on, when there are N - 1 number separators in the text line area corresponding to the title line, the title line is recorded as the chapter title of the Nth level;
[0200] Illustrative examples:
[0201] Chapter titles of the first level: "Chapter 1" or "Overview", without separators;
[0202] Chapter titles of the second level: "1.1", with the separator ".";
[0203] Chapter titles of the third level: "1.1.1", with 2 separators ".";
[0204] Chapter titles of the fourth level: "1.1.1.1", with 3 separators ".";
[0205] Fourth step, content extraction:
[0206] Extract the corresponding content information according to the chapter and section directory structure;
[0207] The specific steps are as follows:
[0208] First, take all the text lines between two adjacent title lines of the same level as the text content in the upper title line among the two adjacent title lines of the same level, and associate them to form a "title - content" pair;
[0209] Fifth step, tree - structure construction:
[0210] Construct a tree - structure with the extracted chapter and section directory and content information;
[0211] Specifically:
[0212] First, construct multiple tree nodes according to the chapter title lines of the first level determined by the font - feature level. Then, take the chapter title lines of the second level in the first level as the first - level child nodes of the corresponding tree nodes. And so on, take the chapter title lines of the subsequent levels as the subordinate child nodes of the corresponding first - level child nodes until all subsequent levels are determined;
[0213] Next, determine other levels according to the hierarchical structure of the numbering, and divide out refined nodes in the corresponding subordinate child nodes according to the child node division method. At the same time, combine the "title-content" pairs corresponding to each refined node to form a tree structure;
[0214] Step 6, Result display:
[0215] Display the tree structure to the user.
[0216] Embodiment 1 provides a method for automatically identifying relevant regulations for correspondence, which has many significant advantages. The file upload link supports multiple formats such as PDF, pictures, and scanned documents, greatly facilitating the processing of files from different sources. In the text recognition stage, OCR technology is used to perform line and character segmentation through horizontal and vertical projection methods, and characters are recognized in combination with a standard character template library, effectively improving the accuracy and efficiency of text recognition. The chapter recognition is based on regular expressions and font features, determining the title lines and levels from multiple dimensions, making the recognition of the chapter and section directory structure more reasonable and reliable. The content extraction forms "title-content" pairs based on the chapter and section directory and constructs a tree structure, clearly showing the overall architecture and specific content of the file, facilitating users to consult and understand.
[0217] Embodiment 2
[0218] As Embodiment 2 of the present invention, in the specific implementation of this application, compared with Embodiment 1, the difference between the technical solution of this embodiment and that of Embodiment 1 is only that in this embodiment, it further includes a preprocessing step:
[0219] This step is used to perform image preprocessing on the laws and regulations documents, and the image preprocessing method is as follows:
[0220] Step Y1, Image grayscale conversion:
[0221] Convert the color image into a grayscale image;
[0222] The grayscale conversion formula is: H = 0.299×R + 0.587×G + 0.114×B;
[0223] Among them, R, G, and B are the red, green, and blue components of the pixel point respectively, and H is the grayscale value;
[0224] Step Y2, Noise removal:
[0225] Use median filtering or mean filtering to remove the noise in the grayscale image;
[0226] First, for each pixel point, take the grayscale values of all pixels within its surrounding 3×3 neighborhood;
[0227] For median filtering, sort the grayscale values of all pixels in ascending order and take the median value as the new value of this pixel;
[0228] For mean filtering, calculate the average value of the grayscale values of all pixels and use this average value as the new value of the pixel.
[0229] Step Y3, Binarization processing:
[0230] Convert the grayscale image into a black-and-white binary image in the following way:
[0231] First, calculate the binarization threshold HY through: ;
[0232] where Hi is the grayscale value of the i-th pixel, i = 1, 2,..., n, and n is the total number of pixels in the image;
[0233] Traverse all pixels. If Hi ≥ HY, adjust the grayscale value of this pixel to 255, which is white; if Hi < HY, adjust the grayscale value of this pixel to 0, which is black.
[0234] Traverse all pixels. If Hi ≥ HY, adjust the grayscale value of this pixel to 255, which is white; if Hi < HY, adjust the grayscale value of this pixel to 0, which is black.
[0235] Step Y4, Skew correction:
[0236] Detect the skew angle of the text in the black-and-white binary image and perform rotation correction on the text in the black-and-white binary image according to it.
[0237] The skew angle detection method is as follows:
[0238] First, use the Canny operator to perform edge detection on the binary image.
[0239] Next, use the Hough transform to detect lines and obtain the skew angles of all text lines.
[0240] Among them, the basic principle of the Hough line detection is to represent and detect the lines in the image space by converting them to the parameter space. This is an existing technology and will not be elaborated here.
[0241] Finally, take the average value of the skew angles of all text lines as the final skew angle.
[0242] The rotation correction method is as follows:
[0243] According to the final skew angle Dp, determine the rotation matrix, and then perform an affine transformation on the black-and-white binary image according to the rotation matrix to make the text lines horizontally aligned.
[0244] The rotation matrix is: .
[0245] Example 2 adds a preprocessing step on the basis of Example 1 to perform image preprocessing on legal and regulatory documents. Simplify the image information through image grayscale to lay the foundation for subsequent processing; use median or mean filtering to remove noise, improve the image quality, and reduce interference with text recognition; perform binarization processing to highlight the text information for easy text segmentation; perform skew correction to detect and correct the text skew angle to make the text lines horizontally aligned. These preprocessing operations effectively improve the image quality, ensure that the OCR technology can play a better role, and further improve the accuracy of text recognition.
[0246] Example 3
[0247] As Example 3 of the present invention, in the specific implementation of this application, compared with Example 1 and Example 2, the technical solution of this example is to combine the solutions of the above Example 1 and Example 2. The difference between the technical solution of this example and Example 1 and Example 2 is only that in this example, the result display also supports the following functions:
[0248] Function 1: Tree expansion / collapse:
[0249] The user can expand or collapse the tree nodes to view or hide the content;
[0250] Function 2: Content retrieval:
[0251] The user inputs keywords, and the matching content is highlighted.
[0252] Example 3 combines the solutions of Example 1 and Example 2 and adds practical functions in the result display session. The tree expansion / collapse function allows the user to flexibly view or hide the content according to their own needs, facilitating quick positioning and browsing of the interested parts. The content retrieval function allows the user to input keywords, and the system will highlight the matching content, greatly improving the efficiency of the user to find specific information, enhancing the practicality and interactivity of the system, and providing a better user experience for the user.
[0253] Example 4
[0254] As Example 4 of the present invention, in the specific implementation of this application, compared with Example 1, Example 2, and Example 3, the technical solution of this example is to combine the solutions of the above Example 1, Example 2, and Example 3.
[0255] Example 4 combines the technical solutions of Examples 1, 2, and 3, giving full play to the advantages of each example. It has the advantages of Example 1 in terms of file format compatibility, character recognition, chapter recognition, and content structuring. It also has the guarantee of image preprocessing in Example 2 to improve the recognition effect, and incorporates the interactive function of Example 3. This comprehensive solution provides users with a comprehensive, efficient, and convenient automatic identification service for letter-related regulations, meets the diverse needs of users, and improves the overall system performance and user experience.
[0256] The present invention also provides a system for automatically identifying letter-related regulations. This system is used to execute a method for automatically identifying letter-related regulations. The system includes:
[0257] A file upload module for users to upload laws and regulations files;
[0258] A preprocessing module for performing image preprocessing on laws and regulations files;
[0259] An OCR recognition module for performing OCR recognition on the uploaded files and extracting text content;
[0260] A chapter extraction module for extracting chapter headings and corresponding content information from the text content recognized by OCR;
[0261] A tree structure construction module for automatically constructing a tree structure based on the extracted chapter headings and content information;
[0262] A user interface module for providing a user interface to view the tree structure.
[0263] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0264] The above is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.
Claims
1. A method for automatically identifying relevant regulations of correspondence, characterized in that, It includes the following steps: File upload: Receive the laws and regulations files uploaded by users; Preprocessing: Perform image preprocessing on the laws and regulations files; Text recognition: Use OCR technology to recognize the text content in the laws and regulations files; Chapter recognition: Based on the pre-set chapter rules, recognize the chapter and section directory structure in the laws and regulations files; Content extraction: Extract the corresponding content information according to the chapter and section directory structure; Tree structure construction: Construct the extracted chapter and section directory and content information into a tree structure; Result display: Display the tree structure to the user.
2. The automatic identification method for correspondence-related regulations according to claim 1, wherein The image preprocessing method is as follows: Step Y1, Image grayscale conversion: Convert the color image to a grayscale image; Step Y2, Noise removal: Use median filtering or mean filtering to remove the noise in the grayscale image; Step Y3, Binarization processing: Convert the grayscale image to a black and white binary image, and the method is as follows: First, calculate the average value of the grayscale values of all pixels in the grayscale image and determine it as the binarization threshold; Traverse all pixels. If Hi≥HY, adjust the grayscale value of this pixel to 255; if Hi<HY, adjust the grayscale value of this pixel to 0; Where, Hi is the grayscale value of the i-th pixel, i = 1, 2,... n, n is the total number of pixels in the image, and HY is the binarization threshold; Step Y4, Skew correction: Detect the skew angle of the text in the black and white binary image and perform rotation correction on the text in the black and white binary image according to it; The skew angle detection method is as follows: First, use the Canny operator to perform edge detection on the binary image; Then, use the Hough transform to detect lines and obtain the skew angles of all text lines; Finally, take the average value of the skew angles of all text lines as the final skew angle; The rotation correction method is as follows: According to the final skew angle, determine the rotation matrix, and then perform an affine transformation on the black and white binary image according to the rotation matrix to make the text lines horizontally aligned.
3. The automatic identification method for correspondence-related regulations according to claim 1, characterized in that The text content recognition method is as follows: Step W1, Line segmentation: Use the horizontal projection method to perform horizontal projection on the black and white binary image, and then count the number of black pixels in each row and mark it as H(y); Where, y represents the row number variable of the pixels in the black and white binary image; Compare the corresponding Hy value of each row with the pre-set judgment threshold PH: When H(y)>PH, it is determined that this row of the black and white binary image is a text line; When H(y)≤PH, it is determined that this row of the black and white binary image is a line spacing; Then, according to the row number variable, record the rows continuously determined as text lines as the text line area; Step W2, Character segmentation: Use the vertical projection method to perform vertical projection on the text line area in the black and white binary image, and then count the number of black pixels in each column and mark it as L(x); Where, x represents the column number variable of the pixels in the text line area; Compare the corresponding Lx value of each row with the pre-set judgment threshold PL: When L(x)>PL, it is determined that this column of the text line area is a character column; When L(x)≤PL, it is determined that this column of the text line area is a character gap column; Then, according to the column number variable, record the columns continuously determined as text lines as the character area; Step W3, Character recognition: Convert each segmented character region into corresponding text symbols in the following specific way: Establish a standard character template library using common Chinese characters, numbers, and punctuation marks; Subsequently, perform scaling processing on the character regions to be recognized, specifically by adjusting the size of the character regions to be the same as the size of the characters in the standard character template library; Then by: ; Calculate the similarity SX between the character regions to be recognized and each standard character template in the standard character template library: where E1 r1 is the pixel value of the r1-th pixel coordinate of the standard character template, and E2 r2 is the pixel value of the r2-th pixel coordinate of the character region to be recognized, r1 = 1, 2, …… u1, u1 is the total number of pixels of the standard character template, r2 = 1, 2, …… u2, u2 is the total number of pixels of the character region to be recognized; After that, select the standard character template with a similarity SX higher than the pre-set similarity threshold and the highest similarity SX as the recognition result of the character region to be recognized.
4. The automatic identification method for correspondence-related regulations according to claim 1, wherein The recognition method for the chapter and section directory structure is as follows: Step Z1: Match the title prefix based on regular expressions; First, pre-define a set of rule patterns using a regular expression library; Then, match the character recognition result of each text line region with each rule pattern in the set of rule patterns; When the character recognition result of a text line region matches any rule pattern, mark the text line region as a candidate title line; When the character recognition result of a text line region does not match all rule patterns, mark the text line region as a non-title line, also denoted as a text line; Step Z2: Verify the title through font features; On the text line region corresponding to the candidate title line, select a character region, and obtain its maximum pixel height and maximum pixel width, and mark them as H1 and W1 respectively; At the same time, on the text line region corresponding to the text line, select a character region, and obtain its maximum pixel height and maximum pixel width, and mark them as H2 and W2 respectively; Then compare the ratio of H1 to H2 and the ratio of W1 to W2 with the pre-set judgment ratio PB respectively: When either H1 / H2 ≥ PB or W1 / W2 ≥ PB holds, determine the text line region corresponding to the candidate title line as a title line; When neither H1 / H2 ≥ PB nor W1 / W2 ≥ PB holds, do not determine the text line region corresponding to the candidate title line as a title line; Step Z3: Determine the level according to font features and numbering structure; Step Z3.1: Determine the font feature level First, extract all determined title lines, select a character region from them, and obtain its maximum pixel height and maximum pixel width; Subsequently, sort the maximum pixel height values or maximum pixel width values of the selected character regions corresponding to each title line in descending order, and determine the level of the title line according to the sorting result; Among them, when the maximum pixel height or maximum pixel width of any two or more title lines is the same, arrange them side by side in the sorting, and they are at the same level; When the sorting order of a title line according to the maximum pixel height is different from its sorting order according to the maximum pixel width, take the level with the higher sorting position as the level of the title line; Step Z3.2: Determine the numbering structure level Extract the pre-set numbering structure, which is reflected by the number separator; Subsequently, determine the level of the title line according to the number of number separators contained in the text line region corresponding to the title line; Specifically: When the text line region corresponding to the title line contains 0 number separators, mark the title line as the chapter title of the first level; When there is one number separator in the text line area corresponding to the title line, the title line is recorded as the chapter title of the second level; And so on, when there are N-1 number separators in the text line area corresponding to the title line, the title line is recorded as the chapter title of the Nth level.
5. A method for automatically identifying relevant regulations of correspondence according to claim 4, characterized in that, The maximum pixel width refers to the maximum pixel value occupied by the selected character area in the vertical direction; the maximum pixel height refers to the maximum pixel value occupied by the selected character area in the horizontal direction.
6. The automatic identification method of correspondence-related regulations according to claim 5, characterized in that The method for determining the hierarchical structure of the numbering is to perform a secondary hierarchical division on the title lines of the same level on the basis of the determination of the font feature level.
7. The automatic identification method for correspondence-related regulations according to claim 5, characterized in that, The specific method for extracting content information is as follows: First, all the text lines between two adjacent title lines of the same level are taken as the text content in the upper title line among the two adjacent title lines of the same level, and they are associated to form a "title-content" pair.
8. A method for automatically identifying relevant regulations of correspondence according to claim 7, characterized in that, The specific method for constructing the tree structure is as follows: First, according to the font feature level, tree nodes are constructed for the chapter title lines at the top level respectively; For each chapter title at a lower level under each top-level title, it is used as the direct child node of the corresponding top-level tree node; And so on, continuously setting the subsequent-level chapter titles as the subordinate child nodes of the previous-level child nodes until all the chapter title levels are determined in the tree structure; Subsequently, according to the hierarchical structure of the numbering, the remaining hierarchical structure is determined; Among the established direct child nodes and subordinate child nodes, refined nodes are further divided according to the same child node division rule; At the same time, each refined node is combined with its corresponding "title-content" pair to finally form a complete tree structure.
9. A method for automatically identifying relevant regulations of correspondence according to claim 1, characterized in that The laws and regulations documents uploaded by the user are in any one of the PDF, picture, and scanned copy formats.
10. A correspondence-related regulation automatic recognition system, which is used to execute a correspondence-related regulation automatic recognition method according to any one of claims 1-9, and is characterized in that The system includes: A file upload module for the user to upload laws and regulations documents; A preprocessing module for performing image preprocessing on the laws and regulations documents; An OCR recognition module for performing OCR recognition on the uploaded file and extracting text content; A chapter extraction module for extracting the chapter headings and corresponding content information from the text content recognized by OCR; A tree structure construction module for automatically constructing a tree structure according to the extracted chapter headings and content information; A user interface module for providing an interface for the user to view the tree structure.
Citation Information
Patent Citations
Method and device for identifying association laws and regulations of letters
CN112199466A
Intelligent letter distribution method and application system for nuclear power plant
CN117349441A
Automatic identification and import method based on file table
CN117973334A