Paragraph division method and system based on text semantic information fusion

By combining heuristic rules and semantic analysis model paragraph division method, OCR technology solves the problem of accuracy and efficiency when dealing with complex documents, and achieves fast and highly accurate paragraph division.

CN120220174APending Publication Date: 2025-06-27XIAN TPRI THERMAL CONTROL TECH +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510276660.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When existing OCR technology processes documents with variable formats and complex layout, the accuracy of identification results and processing convenience are limited. The existing paragraph division method has high computational complexity, long processing time and insufficient accuracy.

Method used

The paragraph division method based on text semantic information fusion is adopted, text fusion and paragraph division are quickly performed through heuristic rules, and paragraph information is deeply processed in combination with semantic analysis models to supplement the missing paragraph division and improve accuracy.

Benefits of technology

It significantly reduces processing time, improves the accuracy of paragraph division, can effectively identify implicit paragraph boundaries in text, and adapts to non-standard format text and complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220174A_ABST
    Figure CN120220174A_ABST
Patent Text Reader

Abstract

The invention discloses a paragraph division method, system and device based on text semantic information fusion, a medium and a program. The method comprises the following steps: identifying a character image to be identified to obtain textboxes, traversing each textbox, and merging the textboxes into lines according to relative positions to obtain a position information merged text; according to the distance between the lines in the position information merging text, identifying the spatial position of the text, and according to an identification result, performing paragraph merging on the text lines to obtain a paragraph information merging text; performing text semantic information fusion processing on the paragraph information merging text based on a semantic analysis model to obtain paragraph text information; and traversing each row in the paragraph character information, and carrying out paragraph calculation layout to obtain divided paragraphs. According to the method, text fusion and paragraph division can be rapidly carried out, missing paragraph division is supplemented by combining semantic information of the text needing paragraph division, and the accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a paragraph division method, system, device, medium and program based on text semantic information fusion. Background Art

[0002] In today's era of rapid information development, optical character recognition (OCR) technology has undoubtedly become a shining star in the field of digital document processing. With its unique ability to convert text information in images into digital text that can be edited and retrieved by computers, it has greatly broadened the boundaries of information acquisition and enabled massive amounts of paper documents and image materials to be quickly transformed into digital knowledge treasures.

[0003] However, OCR technology also faces many challenges. Especially when dealing with documents with varied formats and complex layouts, the accuracy of the recognition results and the convenience of subsequent processing are often severely tested. Users expect not only isolated text characters, but also paragraphed text that can maintain the original document structure and semantic coherence. This requires further fine paragraph restoration or division of the recognized text on the basis of OCR recognition.

[0004] However, in practice, existing paragraph segmentation methods use layout analysis models or large deep learning models to segment paragraphs. Although the accuracy has been improved, the high computational cost and long processing time are prohibitive. These models require a comprehensive analysis of multiple factors such as the document layout, font, font size, line spacing, etc., which requires a lot of computing resources and memory support. In the case of limited computing resources or time constraints, these methods are obviously difficult to meet actual needs. In addition, layout analysis is easily affected by multiple factors such as document quality, shooting angle, lighting conditions, etc., resulting in unstable analysis results.

[0005] On the other hand, directly setting heuristic rules for paragraph division has fast processing speed and low resource consumption, but its accuracy is greatly reduced. Heuristic rules usually merge or divide paragraphs based on the spatial position information of text recognition, but they are often unable to cope with documents with special layouts or complex formats. For example, when there are line indents, column layouts, or fonts of different sizes in the document, heuristic rules are prone to errors. In addition, heuristic rules are even more difficult to deal with complex situations such as cross-page paragraphs, headers and footers. Summary of the invention

[0006] In view of the problems existing in the prior art that the OCR paragraph division technology faces high computational complexity, long processing time, and insufficient accuracy, the present invention provides a paragraph division method based on text semantic information fusion, which uses heuristic rules to quickly perform text fusion and paragraph division, and combines the semantic information of the text to be divided into paragraphs to supplement the missing paragraph division and improve its accuracy. While improving the text segmentation processing speed, it also ensures a high division accuracy.

[0007] To achieve the above object, the present invention provides the following technical solutions.

[0008] In a first aspect, the present invention provides a paragraph division method based on text semantic information fusion, including: Performing recognition on the text image to be recognized to obtain text boxes, traversing each text box, and merging the text boxes into lines according to the relative positions to obtain a position information merged text; Identifying the spatial position of the text according to the distance between the lines in the position information merged text, and merging the text lines according to the recognition result to obtain a paragraph information merged text; Performing text semantic information fusion processing on the paragraph information merged text based on a semantic analysis model to obtain paragraph text information; Traversing each line in the paragraph text information to perform paragraph calculation layout to obtain the divided paragraphs.

[0009] As a further improvement of the present invention, the performing recognition on the text image to be recognized to obtain text boxes, traversing each text box, and merging the text boxes into lines according to the relative positions to obtain a position information merged text includes: Performing recognition on the text image to be recognized to obtain the text and text boxes; Traversing each text box, and merging each text box into a line group according to set conditions to obtain a merged line group; Filtering the abnormal text with misrecognition from the merged line group to obtain a position information merged text.

[0010] As a further improvement of the present invention, the identifying the spatial position of the text according to the distance between the lines in the position information merged text, and merging the text lines according to the recognition result to obtain a paragraph information merged text includes: Traversing the text lines in the position information merged text, and analyzing the distance between the current line and the previous line; Identifying the spatial position of the text according to the analyzed distance between the current line and the previous line, and judging whether the current line and the previous line belong to the same natural paragraph to obtain a recognition result; Merging the text lines according to the recognition result to obtain a paragraph information merged text.

[0011] As a further improvement of the present invention, the semantic analysis model performs text semantic information fusion processing on the paragraph information merged text to obtain paragraph text information, including: Obtain the first sentence text1 in the previous line and the last sentence text2 in the previous line in the paragraph information merged text; Use the semantic analysis model to perform parsing processing on the first sentence text1 in the previous line and the last sentence text2 in the previous line to obtain paragraph text information.

[0012] As a further improvement of the present invention, the use of the semantic analysis model to perform parsing processing on the first sentence text1 in the previous line and the last sentence text2 in the previous line to obtain paragraph text information includes: Use the semantic analysis model to perform parsing processing on the first sentence text1 in the previous line and the last sentence text2 in the previous line, and analyze whether the first sentence text1 in the previous line and the last sentence text2 in the previous line are complete; If both the first sentence text1 in the previous line and the last sentence text2 in the previous line are complete, then merge the first sentence text1 in the previous line and the last sentence text2 in the previous line into text3, and use the semantic analysis model to analyze text3. If there are word segments in the analysis result that span text1 and text2, then the first sentence text1 in the previous line and the last sentence text2 in the previous line belong to the same paragraph and are merged; if there are no word segments in the analysis result that span text1 and text2, then the first sentence text1 in the previous line and the last sentence text2 in the previous line do not belong to the same paragraph and are not merged; If one of the first sentence text1 in the previous line and the last sentence text2 in the previous line is incomplete, then the first sentence text1 in the previous line and the last sentence text2 in the previous line belong to the same paragraph and are merged.

[0013] As a further improvement of the present invention, traversing each line in the paragraph text information to perform paragraph calculation and layout to obtain the divided paragraphs includes: Traverse each line in the paragraph text information, and calculate the width and height of the area, as well as the center point of the area, according to the set page width, paragraph list, and area box; Perform paragraph layout according to the calculated width and height of the area and the center point of the area to obtain the divided paragraph information; Perform paragraph division according to the divided paragraph information to obtain the divided paragraphs.

[0014] In a second aspect, the present invention provides a paragraph division system based on text semantic information fusion, including: Position information merging module: used to recognize the text image to be recognized, obtain text boxes, traverse each text box, merge the text boxes into lines according to their relative positions, and obtain the position information merged text; Paragraph information merging module: used to recognize the spatial position of text according to the distance between lines in the position information merged text, and merge the text lines into paragraphs according to the recognition result to obtain the paragraph information merged text; Paragraph text information module: used to perform text semantic information fusion processing on the paragraph information merged text based on a semantic analysis model to obtain paragraph text information; Finished paragraph division module: used to traverse each line in the paragraph text information, perform paragraph calculation layout, and obtain the divided paragraphs.

[0015] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the paragraph division method based on text semantic information fusion are implemented.

[0016] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the paragraph division method based on text semantic information fusion are implemented.

[0017] In a fifth aspect, the present invention provides a computer program product including computer instructions, and when the computer instructions are executed by a processor, the steps of the paragraph division method based on text semantic information fusion are implemented.

[0018] Compared with the prior art, the present invention has the following beneficial effects: A paragraph division method based on text semantic information fusion. The present invention performs preliminary recognition on the text image to be recognized, quickly obtains text boxes, and uses the relative position information of the text boxes to merge adjacent text boxes into lines, simplifying the preliminary processing steps and reducing unnecessary computational overhead. On this basis, by further analyzing the distance between text lines and combining the recognition results of the spatial positions of the characters, rapid paragraph merging of the text lines is achieved, thereby effectively reducing the overall processing time. This heuristic merging strategy based on position information not only maintains the simplicity of the processing but also significantly improves the efficiency, enabling the OCR technology to handle large-scale text data more proficiently. Secondly, the present invention performs in-depth processing on the paragraph information merged text through a semantic analysis model. By fusing the semantic information of the text, the logical structure and semantic relationship of the text content can be understood more accurately. It can not only effectively identify the implicit paragraph boundaries in the text, such as the theme conversion between paragraphs and the change of logical relationships, but also supplement and divide the paragraphs that are not correctly recognized due to format problems or typesetting errors, improving the accuracy of paragraph division. In addition, the application of semantic analysis also enables the method to maintain a high recognition accuracy when dealing with complex scenarios such as non-standard format texts and handwritten texts, further expanding the application scope of the OCR technology. Therefore, the present invention has achieved remarkable results in optimizing the computational complexity, shortening the processing time, and improving the accuracy. It solves the problem of paragraph division in the existing OCR technology and improves the overall performance of the OCR technology BRIEF DESCRIPTION OF THE DRAWINGS The drawings described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure of the present invention in any way. In the drawings: Figure 1 is a schematic flow chart of a paragraph division method based on text semantic information fusion of the present invention; Figure 2 is a specific schematic flow chart of a paragraph division method based on text semantic information fusion of the present invention; Figure 3 is a diagram of the OCR recognition result before processing; Figure 4 is a diagram of the paragraph division result based only on text and position information; Figure 5 is a diagram of the recognition result in an embodiment of the present invention; Figure 6 is a schematic structural diagram of a paragraph division system based on text semantic information fusion of the present invention; Figure 7 is a schematic diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0021] In view of the problems in the prior art that the OCR paragraph division technology faces high computational complexity, long processing time and insufficient accuracy. The present invention provides a paragraph division method based on text semantic information fusion, as Figure 1 shown, the method includes: S100: Identify the text image to be recognized to obtain text boxes, traverse each text box, and merge the text boxes into lines according to the relative positions to obtain a position information merged text; S200: Identify the spatial positions of the characters according to the distances between the lines in the position information merged text, and merge the text lines according to the recognition results to obtain a paragraph information merged text; S300: Perform text semantic information fusion processing on the paragraph information merged text based on a semantic analysis model to obtain paragraph text information; S400: Traverse each line in the paragraph text information to perform paragraph calculation layout to obtain the divided paragraphs.

[0022] The present invention uses heuristic rules to quickly perform text fusion and paragraph division, and combines the semantic information of the text to be divided into paragraphs to supplement the missing paragraph division and improve its accuracy. While improving the text segmentation processing speed, it also ensures a high division accuracy.

[0023] The following further explains the present invention with reference to specific drawings.

[0024] As Figure 2 shown, a paragraph division method based on text semantic information fusion specifically includes: S1: Text line merging and filtering Identify the text image to be recognized to obtain the text and text boxes.

[0025] S11: Traverse each text box and try to add it to the existing line group. The basis for determining whether it belongs to the same line is as follows: a) Determine whether it is on the same line based on the Y-axis: Calculate the vertical boundary difference between lines. If the difference is within the threshold alpha, it is considered to be on the same line. Alpha can be set artificially according to actual needs. Usually, alpha can be set as a proportion of the text line height. The specific representation is as follows: If |top1 - bottom2| < alpha, it is considered that they are on the same line.

[0026] Among them, top1 is the upper boundary of two adjacent text boxes; bottom2 is the lower boundary of two adjacent text boxes.

[0027] b) Determine whether the X-axis overlaps: Determine whether there is obvious overlap by the height and width of the vertical projection wireframe.

[0028] Determine whether there is an overlapping area by comparing the horizontal projections of the text boxes. If their overlapping part exceeds a threshold, it is considered that they belong to the same line. The specific representation is as follows: Calculate the overlapping area of two text boxes on the X-axis. If the overlapping area is greater than a certain threshold, it is considered that they are on the same line. The overlapping area overlap_area of the two boxes needs to be greater than a certain threshold.

[0029] right1 > left2 and left1 < right2, then there is overlap.

[0030] Among them, right1 is the right endpoint of the first interval; left1 is the left endpoint of the first interval; right2 is the right endpoint of the second interval; left2 is the left endpoint of the second interval.

[0031] right1 > left2 means that the right endpoint of the first interval is on the right side of the left endpoint of the second interval, which means that the first interval may extend to the left side of the second interval.

[0032] left1 < right2 means that the left endpoint of the first interval is on the left side of the right endpoint of the second interval, which means that the first interval may start on the right side of the second interval.

[0033] These two conditions together indicate that there is an overlapping part between the first interval and the second interval.

[0034] For each text box, merge it into the existing line group according to the above Y-axis and X-axis conditions. If there is no matching line, create a new line group.

[0035] S12: After the merging is completed, the misrecognized abnormal text needs to be filtered. The conditions for judging abnormalities are as follows: a) The distance between text boxes is too large Judge whether the horizontal distance between the left boundary of the last detection box and the right boundary of the penultimate detection box is greater than the text line height text_h * 4; If the horizontal distance between the left boundary of the last detection box and the right boundary of the penultimate detection box is greater than the text line height text_h * 4, then the distance between the last character and the previous characters is too large; If the horizontal distance between the left boundary of the last detection box and the right boundary of the penultimate detection box is less than the text line height text_h * 4, then the distance between the last character and the previous characters meets the requirements.

[0036] Specifically, it is as follows: If left_last - right_second_last > text_h * 4, then it is considered that these two boxes do not belong to the same line.

[0037] Among them: left_last is the leftmost position of the last detection box, that is, the x coordinate of the left edge of the detection box; right_second_last is the right edge position of the penultimate detection box, that is, the x coordinate of the right edge of the detection box; text_h is the height of the text box.

[0038] b) Noisy text near the page edge Judge whether the left boundary of the last detection box is close to the right boundary of the page box region_box. If the left boundary of the last detection box is greater than region_box[2] * 0.95, it means that this character may be noise at the edge of the detection area.

[0039] c) Short characters or isolated words If the length of the last character is less than or equal to 2, it means that it may be a short character, such as punctuation or an isolated word. It should be filtered out. Punctuation and single words may be the result of misrecognition.

[0040] S2: Merge texts based on text and location information Traverse each text line, analyze the current line and the previous line, and calculate the horizontal distance in the x direction between the beginning of the current line and the previous line (dif_head_x), the horizontal distance in the x direction between the end of the current line and the previous line (dif_tail_x), the vertical distance between the current line and the previous line (dif_y), and the minimum line height between the current line and the previous line (min_h). These values are used as the basic conditions for paragraph judgment.

[0041] First, set broad rules to determine that two lines belong to the same natural paragraph and should be merged. Then, set detailed rules to determine whether two lines are not in the same paragraph and modify the merge flag.

[0042] Among them, the criteria for judging the same paragraph are as follows: a) Head spacing limit: The x-distance at the beginning of the line is within a reasonable range.

[0043] For example, -min_h * 4.0 < dif_head_x < min_h * 4. The limiting conditions are all centered around min_h, and the limiting conditions are automatically adjusted according to the actual situation of the document. b) Tail alignment check: The x-distance at the end of the line cannot be too small. c) Overall box limit: The head spacing between the right boundary of the previous paragraph and the current line cannot be too large.

[0044] d) Vertical spacing limit: The y-spacing between the current line and the previous line cannot be too large.

[0045] e) Use text matching. If the previous line ends with a comma, it is directly considered that the current line is a continuation of the same paragraph.

[0046] Among them, the criteria for judging different paragraphs are as follows: a) If the previous paragraph has only one line: If the previous line is indented, the indentation of the current line is similar to that of the previous line, or there is an obvious indentation in the current line itself, it may be a new paragraph. The text of the current line and the previous line need to be passed into the text merging module based on semantic information for further judgment.

[0047] If the previous paragraph starts with a number or a serial number, such as "(1)", "(2)", etc., it may be a subheading. If the right distance of the previous line is very close to the right border of the overall recognition box, and the indentation of this line is less than that of the previous line, it is considered to be in the same paragraph. Otherwise, the text of the current line and the previous line need to be passed into the text merging module based on semantic information for further judgment.

[0048] If the right boundary of the previous paragraph is very far from the right boundary of the overall area, it is usually considered that this paragraph is relatively short, and the current line is considered not to belong to the previous paragraph, not merged, and a new paragraph is started.

[0049] If the length of the previous line is significantly shorter than expected (less than 70% of the maximum line length), it is considered that the current line does not belong to the previous paragraph, not merged, and a new paragraph is started.

[0050] b) If the previous line has multiple lines: If the distance between the current line and the previous paragraph is much larger than the line spacing of the previous paragraph, it is considered that the current line does not belong to the previous paragraph, not merged, and a new paragraph is started.

[0051] If the distance between the current line and the previous paragraph is much greater than the paragraph distance, it is considered that the current line does not belong to the previous paragraph, and they are not merged. Instead, a new paragraph is started.

[0052] c) If each line on the current page has "........", it is the table of contents page and they are not merged.

[0053] d) If the current line starts with a number, such as one, two, 1, 2, 2.1.2, 9.4, 1), the third article, etc., it is considered that the current line does not belong to the previous paragraph, and they are not merged. Instead, a new paragraph is started.

[0054] e) If the previous line ends with ".", and the difference in the x direction between the beginning of the current line and the previous line exceeds 1.5 times the line height, and the indentation of the current line is significantly more than that of the previous line, it is considered that the current line does not belong to the previous paragraph, and they are not merged. Instead, a new paragraph is started.

[0055] f) If the previous line ends with ".", and the difference in the x direction between the beginning of the current line and the previous line exceeds 1.5 times the line height, and the indentation of the current line is significantly more than that of the previous line, it is considered that the current line does not belong to the previous paragraph, and they are not merged. Instead, a new paragraph is started.

[0056] g) If the previous line ends with ":", and the indentation of the current line exceeds that of the previous line, it is considered that the current line does not belong to the previous paragraph, and they are not merged. Instead, a new paragraph is started.

[0057] S3: Merge the text based on semantic information S31: Receive the current line text and the previous line text passed in by S2. First, respectively intercept the first sentence (text1) of the current line and the last sentence (text2) of the previous line as the core for processing.

[0058] S32: Call the semantic analysis model to parse text1 and text2 to obtain their word segmentation, part-of-speech, and dependency relationships. According to the part-of-speech and dependency relationships of the word segmentation, determine whether there are a subject, a predicate (linking verb), and an object (predicative) in the text. If so, record their positions for subsequent processing reference.

[0059] S33: Analyze whether the sentence is complete based on the presence and positions of the subject, predicate (linking verb), and object (predicative): a) The subject, predicate (linking verb), and object (predicative) all exist: Subject > Predicate (linking verb) > Object (predicative), subject-predicate-object or subject-linking verb-predicative structure, the sentence structure is complete.

[0060] Subject > Object > Predicate (linking verb), ba-sentence, the sentence structure is complete.

[0061] Object > Subject > Predicate (linking verb), passive sentence, the sentence structure is complete.

[0062] All other cases are considered to have an incomplete sentence structure.

[0063] b) When the subject and the predicate (linking verb) exist but the object (predicative) does not, it is considered that the sentence structure is incomplete.

[0064] c) When the predicate (linking verb) and the object (predicative) exist but the subject does not, it is an imperative sentence and the sentence structure is considered complete.

[0065] d) In other cases, the core component of the sentence - the predicate (linking verb) - is missing, and the sentence structure is considered incomplete.

[0066] S4: After the sentence integrity of text1 and text2 is judged, discuss in different cases to obtain the paragraph text information.

[0067] a) Both sentences are complete: It is still possible that the two sentences belong to the same paragraph, and further judgment is made using participles. Combine text2 and text1 into text3, call the semantic analysis model to analyze text3. If there is a participle in the participle result that spans text1 and text2, that is, a part of the participle is in text1 and a part is in text2, then it is considered to belong to the same paragraph and should be combined.

[0068] b) If any one of the two sentences is incomplete, it is considered to belong to the same paragraph and should be combined.

[0069] S5: Traverse each line of the paragraph text information, and combine them into paragraphs or start new paragraphs according to whether they are marked for combination with the previous line. The first line is defaulted to a new paragraph.

[0070] Calculate and set the alignment, indentation, font size and other styles of the paragraph according to the set page width, paragraph list, and region box.

[0071] Calculate the width (region_width) and height (region_height) of the region, as well as the center point (region_center_x) of the region.

[0072] Traverse each paragraph and calculate the layout for each paragraph.

[0073] a) Alignment setting: If the distance between the current paragraph and the previous paragraph is very close and the line start positions are similar, it will be set to justified alignment.

[0074] If the width of the paragraph is very small (less than twice the height of the paragraph), it may be set to distributed alignment.

[0075] The center point of the paragraph is very close to the center of the area, and the left and right margins of the paragraph are similar, usually set to be centered. If the difference in the starting positions of the previous line and the current line is small, and the alignment of the previous line is not centered, then it remains non-centered. If the text line is particularly long, or the text has a large inclination, or the left and right spacing of the table cell is very small, then avoid centered alignment.

[0076] If it is a table of contents paragraph, then set it to be justified.

[0077] b) Indentation setting: The left indentation is determined by the ratio of the left side of the paragraph to the left boundary of the area.

[0078] The first line indentation is calculated by the distance from the second line in a multi-line paragraph, and in a single-line paragraph, it inherits the indentation setting of the previous paragraph.

[0079] If both the previous paragraph and the current paragraph are justified, and the difference in their first line positions is not significant, the current paragraph will inherit the indentation setting of the previous paragraph.

[0080] c) Font setting: According to the width and height of the page and the width and height of the text block, considering factors such as the number of lines, the number of Chinese characters, and the character width, calculate the font size relative to A4 paper.

[0081] If it is about the same size as the font of the previous paragraph, set it to the font size of the previous paragraph.

[0082] S6: Output the segmented paragraph information.

[0083] According to the segmented paragraph information, perform paragraph segmentation to obtain the segmented paragraphs.

[0084] In summary, the present invention combines heuristic rules with semantic analysis to form an efficient and accurate paragraph segmentation mechanism. The application of heuristic rules makes the process of text fusion and paragraph segmentation more flexible and fast, and can quickly respond to changes in various text features; while the addition of semantic analysis provides strong intellectual support for this process, making paragraph segmentation no longer rely solely on the surface text arrangement, but able to penetrate into the essence of the text content to achieve more accurate and intelligent segmentation. This fusion strategy not only improves the intelligent level of the OCR technology, but also provides users with a more reliable and convenient text processing experience.

[0085] Compared with the method of simply merging paragraphs according to the heuristic rules of position information, the average time consumption per page of full-text documents only increases by less than 0.1 s, which is much less than the time consumption of calling large models or open-source layout analysis models. However, the accuracy rate on the test dataset has increased from 86% correct rate of text merging by pure position judgment to 96% correct rate. Even in the case of time requirements and insufficient computing resources, it can still be used and ensure a paragraph segmentation correct rate of more than 95%, which is sufficient to cover the vast majority of document OCR text paragraph segmentation scenarios.

[0086] The present invention will be further explained and illustrated below in conjunction with specific embodiments.

[0087] As Figure 3 shown, the unprocessed OCR recognition results lack a deep understanding of the text context. Therefore, when encountering ambiguous or fuzzy characters, the system may not be able to make a correct judgment based on the context. Manual proofreading or post-processing is required to correct these errors. This increases the workload and time cost.

[0088] As Figure 4 shown, when dividing paragraphs only based on text and position information, although relevant paragraphs can also be found, for those paragraphs containing indentation formats, the division method based only on text and position information often cannot accurately identify them. When the text is extracted from different sources or formats, the indentation information may be lost or become inconsistent. Due to the insufficient recognition ability of this division method for indented paragraphs, in actual applications, it is often found that some paragraphs are wrongly divided or merged together. This not only affects the readability of the text but also may lead to misunderstandings or omissions of information. To correct these errors, manual proofreading or post-processing is required.

[0089] Manual proofreading is a time-consuming and cumbersome process that requires checking the text paragraph by paragraph to ensure that each paragraph is correctly divided. Post-processing may require the assistance of some professional text processing tools or algorithms to automatically identify and correct errors in paragraph division. No matter which method is adopted, additional time and effort are required to ensure the accuracy of text processing.

[0090] Therefore, when processing text containing indentation formats, it is necessary to more carefully select the paragraph division method and consider combining other information or technologies to improve the division accuracy. At the same time, for paragraphs that cannot be automatically recognized, we should also be prepared for manual proofreading or post-processing to ensure the final quality of text processing.

[0091] As Figure 5As shown, according to the method in the present application for recognition, text fusion and paragraph division are quickly carried out through heuristic rules, greatly simplifying the processing flow, reducing the computational complexity, and thus significantly shortening the processing time. It not only considers the spatial position information of the text but also combines the semantic information of the text for paragraph division. This multi-dimensional information fusion processing method significantly improves the accuracy of paragraph division. Even in the face of complex or ambiguous text layouts, the present invention can accurately identify the paragraph boundaries, avoiding omission or incorrect division. Moreover, through the introduction of the semantic analysis model, the present invention enhances the understanding and adaptability to complex text structures, making the paragraph division more flexible and accurate.

[0092] The second object of the present invention is to propose a paragraph division system based on text semantic information fusion, as Figure 6 shown, including: Position information merging module 100: used to recognize the text image to be recognized, obtain text boxes, traverse each text box, and merge the text boxes into lines according to the relative positions to obtain position information merged text; Paragraph information merging module 200: used to recognize the spatial positions of the text according to the distances between the lines in the position information merged text, and merge the text lines into paragraphs according to the recognition results to obtain paragraph information merged text; Paragraph text information module 300: used to perform text semantic information fusion processing on the paragraph information merged text based on the semantic analysis model to obtain paragraph text information; Completed paragraph division module 400: used to traverse each line in the paragraph text information to perform paragraph calculation layout to obtain the divided paragraphs.

[0093] As Figure 7 shown, the third object of the present invention is to provide an electronic device, which includes: a processor 501, a memory 502, and a display screen 503. Among them, the memory 502 and the display screen 503 are both connected to the processor 501, such as through a bus 504. Optionally, the electronic device may further include a transceiver 505. It should be noted that in practical applications, the transceiver 505 is not limited to one, and the structure of this electronic device does not constitute a limitation to the embodiments of the present application.

[0094] The processor 501 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 501 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0095] The bus 504 may include a path for transmitting information between the above components. The bus 504 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 504 may be divided into an address bus, a data bus, a control bus, etc.

[0096] The memory 502 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0097] The memory 502 is used to store the application program code for executing the solution of this application and is controlled by the processor 501 for execution. The processor 501 is used to execute the application program code stored in the memory 502 to implement the content shown in the foregoing method embodiments.

[0098] Figure 7 The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0099] The fourth object of the present invention is to provide a computer-readable storage medium storing a computer program, on which a computer program is stored. When the program is executed by a processor, it implements each process of the method embodiment as described above. Figure 1 For example, a memory including instructions that can be executed by a processor of an electronic device to complete the above method.

[0100] A computer-readable storage medium may be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium may be, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above. Specifically, a computer-readable storage medium may be a portable computer disk, a hard disk, a USB flash drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, an optical disk, a magnetic disk, a mechanical encoding device, and any combination of the above.

[0101] The fifth object of the present invention is to provide a computer program product including computer instructions that, when executed by a processor, implement each process of the method embodiment as described above. Figure 1 And can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0102] By reading the above description, many embodiments and many applications beyond the provided examples will be obvious to those skilled in the art. Therefore, the scope of this teaching should not be determined with reference to the above description, but should be determined with reference to the full scope of the foregoing claims and the equivalents of these claims. For the sake of comprehensiveness, all articles and references, including patent applications and published announcements, are incorporated herein by reference. The omission of any aspect of the subject matter disclosed herein in the foregoing claims is not intended to abandon such subject matter, nor should it be considered that the applicant has not considered such subject matter as part of the disclosed inventive subject matter.

[0103] The above content is a further detailed description of the present invention. It cannot be determined that the specific implementation manners of the present invention are limited thereto. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as falling within the protection scope determined by the claims submitted for the present invention.

Claims

1. A paragraph division method based on text semantic information fusion, characterized in that: include: Recognize the text image to be recognized, obtain the text box, traverse each text box, merge the text boxes into rows according to the relative positions, and obtain the position information merged text; Merge the distance between lines in the text according to the position information, identify the spatial position of the text, and merge the text lines into paragraphs according to the recognition result to obtain a paragraph information merged text; Based on the semantic analysis model, the paragraph information is merged into the text to perform text semantic information fusion processing to obtain the paragraph text information; Traverse each line of the paragraph text information, perform paragraph calculation and layout, and obtain the divided paragraphs.

2. A paragraph division method based on text semantic information fusion according to claim 1, characterized in that: The method of identifying the text image to be identified, obtaining text boxes, traversing each text box, merging the text boxes into rows according to relative positions, and obtaining position information merged text includes: Recognize the text image to be recognized and obtain the text and text box; Traverse each text box, merge each text box into a row group according to the set conditions, and obtain a merged row group; The merged rows are grouped, and the misidentified abnormal texts are filtered to obtain the merged text of the location information.

3. The paragraph division method based on text semantic information fusion according to claim 1 is characterized in that: The method of combining the distance between lines in the text according to the position information, identifying the spatial position of the text, and combining the text lines into paragraphs according to the recognition result to obtain the paragraph information combined text includes: Traverse the text lines in the position information merge text and analyze the distance between the current line and the previous line; According to the analysis of the distance between the current line and the previous line, the spatial position of the text is identified, and it is determined whether the current line and the previous line belong to the same natural paragraph to obtain the recognition result; According to the results, the text lines are merged into paragraphs to obtain the paragraph information merged text.

4. The paragraph division method based on text semantic information fusion according to claim 1 is characterized in that: The text semantic information fusion processing is performed on the paragraph information merged text based on the semantic analysis model to obtain the paragraph text information, including: Get paragraph information and merge the first sentence text1 of the previous line and the last sentence text2 of the previous line; Use the semantic analysis model to parse the first sentence text1 of the previous line and the last sentence text2 of the previous line to obtain the paragraph text information.

5. The paragraph division method based on text semantic information fusion according to claim 4 is characterized in that: The semantic analysis model is used to parse the first sentence text1 of the previous line and the last sentence text2 of the previous line to obtain paragraph text information, including: Use the semantic analysis model to parse the first sentence of the previous line, text1, and the last sentence of the previous line, text2, to analyze whether the first sentence of the previous line, text1, and the last sentence of the previous line, text2, are complete; If the first sentence of the previous line, text1, and the last sentence of the previous line, text2, are both complete, then the first sentence of the previous line, text1, and the last sentence of the previous line, text2, are merged into text3. The semantic analysis model is used to analyze text3. If there is a segmentation word that spans text1 and text2 in the analysis result, then the first sentence of the previous line, text1, and the last sentence of the previous line, text2, belong to the same paragraph and are merged. If there is no segmentation word that spans text1 and text2 in the analysis result, then the first sentence of the previous line, text1, and the last sentence of the previous line, text2, do not belong to the same paragraph and are not merged. If one of the first sentence text1 in the previous line and the last sentence text2 in the previous line is incomplete, then the first sentence text1 in the previous line and the last sentence text2 in the previous line belong to the same paragraph and are merged.

6. The paragraph division method based on text semantic information fusion according to claim 1 is characterized in that: The traversing each line of the paragraph text information, performing paragraph calculation layout, and obtaining the divided paragraphs includes: Traverse each line of the paragraph text information, and calculate the width and height of the area, as well as the center point of the area according to the set page width, paragraph list, and area box; Perform paragraph layout based on the width and height of the calculated area and the center point of the area to obtain the divided paragraph information; According to the divided paragraph information, paragraph division is performed to obtain the divided paragraphs.

7. A paragraph segmentation system based on text semantic information fusion, characterized in that: include: Position information merging module: used to recognize the text image to be recognized, obtain text boxes, traverse each text box, merge the text boxes into rows according to relative positions, and obtain position information merged text; Paragraph information merging module: used to merge the distance between lines in the text according to the position information, identify the spatial position of the text, merge the text lines into paragraphs according to the recognition results, and obtain the paragraph information merged text; Paragraph text information module: used to perform text semantic information fusion processing on paragraph information merged text based on semantic analysis model to obtain paragraph text information; Complete the paragraph division module: used to traverse each line of the paragraph text information, perform paragraph calculation layout, and obtain the divided paragraphs.

8. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method implements the steps of a paragraph division method based on text semantic information fusion as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the paragraph division method based on text semantic information fusion according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of a paragraph division method based on text semantic information fusion as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Text detection and recognition method and system for complex layout image

    CN122135352A

  • Broken line merging method and system based on natural language word segmentation and position information fusion

    CN122412582A