A method for segmenting articles based on clustering
By processing OCR-recognized text using a clustering-based method and removing redundant line breaks, the problem of incorrect text paragraph segmentation was solved, and the accuracy of text paragraph reconstruction and data mining was improved.
Patent Information
- Application Number
- CN202210827775.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-14
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-07-14
AI Technical Summary
In paperless offices, existing technologies often result in redundant line breaks in the text after OCR recognition, leading to incorrect paragraph segmentation and affecting the accuracy of subsequent data mining.
A clustering-based approach is used to calculate the character length of a text line and map it to coordinate points. Cluster analysis is then performed to identify line breaks that need to be deleted. Variance control is used to reduce false positives and the paragraphs are reorganized.
It effectively removes redundant line breaks, restores the correctness of text paragraphs, and improves the accuracy of data mining.
Smart Images

Figure CN115311662B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer application technologies, and particularly to an article paragraphing method based on clustering. Background Art
[0002] In the process of paperless office, a common scenario is that the paper materials submitted by the parties need to be scanned and converted into double-layer PDF, and then the text information is extracted from the PDF. At this time, each line of the text information will have a line break character, so that the original paragraph information is split into information for each line. This format may split some characteristic keywords onto different lines, bringing certain difficulties to subsequent data mining. The present invention can process such text into normal paragraphs, delete redundant line break characters, and lay a good foundation for the subsequent utilization of the text. Summary of the Invention
[0003] In view of the problems existing in the prior art, the present invention provides an article paragraphing method based on clustering, which can delete redundant line break characters from the documents recognized by OCR and reorganize the paragraphs of the documents; and can reduce misjudgment situations.
[0004] The purpose of the present invention is achieved through the following technical solutions.
[0005] An article paragraphing method based on clustering, the steps include:
[0006] 1) Calculate the length of the half-width characters of each line of the document to be processed and store them in a one-dimensional array;
[0007] 2) Map the values in the one-dimensional array into coordinate points on a plane: map them into coordinate points on the X-axis or Y-axis;
[0008] 3) Perform clustering analysis with the mapped coordinate points CustomerPoint as the input and an initial distance x, that is, for the coordinate points on the plane, the points with a maximum radius not exceeding x are grouped into one set, and several sets are obtained: the clustering set clusterResult;
[0009] 4) Find the clustering sets with the largest and the second largest coordinate values in the clustering set clusterResult;
[0010] 5) Calculate the number of elements in the clustering sets with the largest and the second largest coordinate values, and the clustering set with the larger number is used as the set to be processed clusterReturn;
[0011] 6) Calculate the variance of the set to be processed to obtain the set variance diff;
[0012] 7) When the set variance diff > a, recluster with the distance as x - 1, and then return to execute steps 3) - 6);
[0013] 8) When the set variance diff > a, recluster with the distance as x - 2, and then return to execute steps 3) - 6), and so on until the set variance diff ≤ a, to obtain the finally required set clusterReturn;
[0014] 9) Process the document to be processed line by line, and determine whether the length of the half - width characters in each line is in the finally required set clusterReturn;
[0015] 10) Determine that if the current line length does not match or ends with a punctuation mark, the line break character is not deleted, otherwise the line break character is deleted, and finally the text after the paragraph recombination is obtained.
[0016] The initial distance x is an integer greater than or equal to 3.
[0017] The value range of a is 2.5 - 3.5, preferably 3.
[0018] The punctuation marks include exclamation marks, full stops, semicolons and colons.
[0019] Compared with the prior art, the advantages of the present invention are as follows: for the document recognized by OCR, the present invention uses a clustering algorithm to find out the lines that need to delete the line break characters; deletes the redundant line break characters to recombine the paragraphs of the document; and uses variance to control the results to reduce misjudgment situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flowchart for paperless office.
[0021] Figure 2 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present invention will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0023] As Figure 1 shown, in the paperless process, the original paper - based text materials are scanned and converted into PDF, and then extracted into txt text. At this time, there are line break characters in each line, and the original article paragraphs are separated by line break characters, resulting in incorrect final results in the subsequent data mining process because keywords may be separated by line break characters. This algorithm can restore such text into the original paragraphs, laying a foundation for the smooth implementation of paperless office.
[0024] As Figure 2As shown in the figure, the present invention provides a clustering-based article paragraphing algorithm, which maps the character length of each line into points on the coordinate axis, and then through clustering, finds the lines where line breaks need to be deleted, and finally restores the segmented paragraphs to the original lines. The specific process is as follows:
[0025] 1. Calculate the length of the half-width characters of each line of the document to be processed and store it in a one-dimensional array: text length lineLength[];
[0026] 2. Map the values in this array into coordinate points on the plane. In the present invention, it is a one-dimensional array, which can be mapped into coordinate points on the X-axis or Y-axis: mapped coordinate point CustomerPoint;
[0027] 3. Perform clustering analysis with the mapped coordinate point CustomerPoint as the input and an initial distance of 3, that is, for the coordinate points on the plane, points with a maximum radius not exceeding 3 are grouped into one set. Finally, several sets are obtained: clustering set clusterResult;
[0028] 4. Find the classes with the largest (take the first element in the set and read the coordinate value) and the second largest values in the clustering set clusterResult to obtain two sets.
[0029] 5. Calculate the number of elements in these two sets and keep the set with the larger number: the set to be processed clusterReturn;
[0030] 6. Calculate the variance of this set to obtain a value: set variance diff;
[0031] 7. Judge that if the set variance diff > 3.0, it means that the included range is relatively large, re-cluster with a distance of 2, obtain the set according to step 4, calculate the variance again according to step 5, and obtain the new set variance diff;
[0032] 8. Judge that if the set variance diff > 3.0, it means that the included range is still relatively large, re-cluster with a distance of 1, and execute steps 4 and 5 again to obtain the final set to be processed clusterReturn;
[0033] 9. Process each line of the document to be processed and judge whether the length of the half-width characters of each line is in the clustering result;
[0034] 10. Judge whether the current line has an inconsistent length or ends with a punctuation mark. The punctuation marks can be defined in advance, and by default, they are exclamation marks, full stops, semicolons, and colons. If so, do not delete the line break, otherwise delete the line break;
[0035] 11. Finally, obtain the text after the paragraphs are reorganized.
Claims
1. A method for segmenting article paragraphs based on clustering, characterized by the steps Including: 1) Calculate the length of the half-width characters in each line of the document to be processed and store them in a one-dimensional array; 2) Map the values in the one-dimensional array into coordinate points on a plane: map them into coordinate points on the X-axis or Y-axis; 3) Perform clustering analysis with the mapped coordinate point CustomerPoint as the input and the initial distance x, that is, for the coordinate points on the plane, points with a maximum radius not exceeding x are grouped into one set, obtaining several sets: the clustering set clusterResult; 4) Find the clustering sets with the largest and the second largest coordinate values in the clustering set clusterResult; 5) Calculate the number of elements in the clustering sets with the largest and the second largest coordinate values, and the clustering set with a larger number of elements is used as the set to be processed clusterReturn; 6) Calculate the variance of the set to be processed to obtain the set variance diff; 7) When the set variance diff > a, re-cluster with a distance of x - 1 and re-return to execute steps 3) - 6); 8) When the set variance diff > a, re-cluster with a distance of x - 2 and re-return to execute steps 3) - 6), and so on, until the set variance diff ≤ a, obtaining the final set to be processed clusterReturn; 9) Process the document to be processed line by line and determine whether the length of the half-width characters in each line is in the final set to be processed clusterReturn; 10) Determine that if the current line length does not match or ends with a punctuation mark, the line break is not deleted, otherwise the line break is deleted, finally obtaining the text after the reorganized paragraph; the a is 3.
2. The method for segmenting an article based on clustering according to claim 1, wherein The initial distance x is an integer greater than or equal to 3.
3. A method for segmenting an article based on clustering according to claim 1 or 2, characterized in that The punctuation marks include exclamation marks, full stops, semicolons and colons.
Citation Information
Patent Citations
E-book typesetting method and e-book typesetting system
CN102081600A
Typesetting processing method and device based on original document
CN102681980A