Method and device for identifying number modification area in document

By combining text line association detection, number modification detection and alignment detection, and combining semantic and visual features, the number modification areas in the document are identified, which solves the recognition difficulties in the existing technology and improves the recognition accuracy and document processing efficiency.

CN120612709APending Publication Date: 2025-09-09ZHUHAI KINGSOFT OFFICE SOFTWARE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410253357.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively identifying numbered modified areas in documents, which affects document structuring and reading comprehension.

Method used

By determining the area to be identified in the target document, text line association detection, number modification detection and number alignment detection are performed, and semantic and visual features are combined to determine whether the area is a number modification area.

Benefits of technology

The accuracy of identifying number modification areas has been improved, and number modifications at the line and paragraph levels can be identified, supporting subsequent document reading and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612709A_ABST
    Figure CN120612709A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for recognizing a number modification area in a document, and relates to the technical field of document recognized.The method for recognizing the number modification area in the document comprises the steps that a to-be-recognized area in a target document is determined, and the to-be-recognized area comprises at least one text line; performing text line association detection on each text line to obtain an association detection result indicating whether the to-be-identified region meets a preset line association condition; performing serial number modification detection on each text line to obtain a serial number modification detection result indicating whether the to-be-recognized region meets a preset serial number modification condition; performing serial number alignment detection on each text line to obtain a serial number alignment detection result indicating whether the to-be-identified area meets a preset serial number alignment condition; and according to the association detection result, the number modification detection result and the number alignment detection result, determining whether the to-be-identified area is a number modification area. According to the method, each number modification area in the document can be quickly and accurately identified, and the identification accuracy of the number modification areas is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document recognition, and in particular to a method and device for recognizing numbered modified areas in a document. Background Art

[0002] List numbers can be thought of as a decoration or marking method used to identify and distinguish a series of items. In formal documents, numbering can also help create a structured and orderly content layout, making it easier for readers to understand the organization of the text.

[0003] Whether it is a flow document, a layout document, or an image document, it can contain areas modified with numbers. Therefore, how to detect numbered areas in documents has become an urgent problem to be solved. Summary of the Invention

[0004] In view of the problems existing in the prior art, embodiments of the present invention provide a method and apparatus for identifying number-modified areas in a document.

[0005] The present invention provides a method for identifying numbered modified areas in a document, comprising:

[0006] Determine a region to be recognized in a target document, wherein the region to be recognized includes at least one text line;

[0007] Performing a text line association test on each of the text lines to obtain an association test result indicating whether the area to be identified meets a preset line association condition;

[0008] Performing a number modification detection on each of the text lines to obtain a number modification detection result indicating whether the area to be identified meets a preset number modification condition;

[0009] Performing a number alignment test on each of the text lines to obtain a number alignment test result indicating whether the area to be identified meets a preset number alignment condition;

[0010] It is determined whether the to-be-identified region is a number-modified region according to the association detection result, the number-modification detection result, and the number-alignment detection result.

[0011] According to a method for identifying number modification areas in a document provided by the present invention, the number modification detection includes first character detection and / or number item detection, the preset number modification condition includes a first character condition and / or a number item condition, and the number modification detection result includes a first character detection result and / or a number item detection result;

[0012] The performing of number modification detection on each of the text lines to obtain a number modification detection result indicating whether the to-be-identified area meets a preset number modification condition includes:

[0013] Performing a first character detection on each of the text lines to obtain a first character detection result indicating whether the area to be identified meets the first character condition; and / or,

[0014] A numbering item detection is performed on each of the text lines to obtain a numbering item detection result indicating whether the area to be identified meets the numbering item condition.

[0015] According to a method for identifying numbered modified areas in a document provided by the present invention, the first character condition is that the first characters of each of the text lines belong to the same numbered character set;

[0016] The first character detection is performed on each of the text lines to obtain a first character detection result indicating whether the area to be identified meets the first character condition, including:

[0017] Obtaining the first character of each text line and searching for the Unicode of each first character;

[0018] Determining whether each of the first characters belongs to the same numbered character set according to the unified character encoding of each of the first characters;

[0019] If so, determining that the first character detection result is that the area to be recognized meets the first character condition;

[0020] If not, it is determined that the first character detection result is that the area to be recognized does not meet the first character condition.

[0021] According to a method for identifying numbered modified areas in a document provided by the present invention, the numbered item condition is that each of the text lines has a numbered item;

[0022] The performing of numbered item detection on each of the text lines to obtain a numbered item detection result indicating whether the area to be identified meets the numbered item condition includes:

[0023] For each of the text lines, determining whether a numbered item exists in the text line based on characters between the first character and a set character in the text line, where the set character is the character preceding the first specified character, and the specified character is not a numeric character or a separator;

[0024] If a numbered item exists in each of the text lines, determining the numbered item detection result is that the area to be identified meets the numbered item condition;

[0025] If a specific text line exists in each of the text lines, it is determined that the numbered item detection result is that the area to be identified does not meet the numbered item condition, and the specific text line is a text line in which no numbered item exists.

[0026] According to a method for identifying numbered modified areas in a document provided by the present invention, determining whether a numbered item exists in the text line based on characters between the first character and a set character in the text line includes:

[0027] Recognize the first character of the text line to obtain a recognition result; and determine whether the text line meets a first condition based on the recognition result, wherein the first condition indicates that the first character is a numeric character;

[0028] Performing a delimiter search on characters between the first character and the set character in the text line to obtain a search result; and determining whether the text line meets a second condition based on the search result, the second condition indicating that a delimiter exists between the first character and the set character;

[0029] Obtaining, from characters between a second character and a set character in the text line, a first number of characters of at least one numeric character combination, the numeric character combination being at least one consecutive numeric character, the first number of characters being the number of numeric characters in the numeric character combination; and determining, based on each first number of characters, whether the text line meets a third condition, the third condition indicating that each first number of characters is less than or equal to a first value;

[0030] Obtaining, from characters between a second character and a set character in the text line, a second number of characters of at least one separator combination, the separator combination being at least one consecutive separator, the second number of characters being the number of separators in the separator combination; and determining whether the text line meets a fourth condition based on each second number of characters, the fourth condition being that each second number of characters is less than or equal to a second value;

[0031] If the text line satisfies the first condition, the second condition, the third condition, and the fourth condition at the same time, determining that a numbered item exists in the text line;

[0032] If the text line does not meet at least one of the first condition, the second condition, the third condition, and the fourth condition, it is determined that no numbered item exists in the text line.

[0033] According to a method for identifying numbered modified areas in a document provided by the present invention, after determining that a numbered item exists in the text line, the method further includes:

[0034] The numbering level of the numbered items in the text line is determined based on the first combination quantity of the numeric character combinations and the second combination quantity of the separator combinations in the text line, the first combination quantity being the number of the numeric character combinations contained in the text line, and the second combination quantity being the number of the separator combinations contained in the text line.

[0035] According to a method for identifying number-modified areas in a document provided by the present invention, performing number alignment detection on each of the text lines to obtain a number alignment detection result indicating whether the area to be identified meets a preset number alignment condition includes:

[0036] Calculating the average font size corresponding to the area to be identified based on the font size of each character in all the text lines;

[0037] Constructing a mapping table, wherein the mapping table records the first left boundary value corresponding to each numbering level;

[0038] For each of the text lines, obtaining a numbering level corresponding to the text line and a second left boundary value of a circumscribed rectangular box; searching the mapping table for a first left boundary value corresponding to the text line based on the numbering level corresponding to the text line; determining an absolute value of a difference between the first left boundary value and the second left boundary value corresponding to the text line; and determining that the text line is a set text line if the absolute value is greater than the average font size;

[0039] In a case where the set text line exists in each of the text lines, determining that the number alignment detection result is that the to-be-identified area does not meet the preset number alignment condition;

[0040] In a case where the set text line does not exist in each of the text lines, it is determined that the number alignment detection result is that the to-be-identified area meets the preset number alignment condition.

[0041] According to a method for identifying numbered modified areas in a document provided by the present invention, the text line association detection includes content type detection, text line type detection, and color alignment detection, the preset line association conditions include content type conditions, text line type conditions, and color alignment conditions, and the association detection results include content type detection results, text line type detection results, and color alignment detection results;

[0042] The performing text line association detection on each of the text lines to obtain an association detection result of whether the area to be identified meets a preset line association condition includes:

[0043] Performing content type detection on each of the text lines to obtain a content type detection result of whether the area to be identified meets the content type condition;

[0044] Performing text line type detection on each of the text lines to obtain a text line type detection result indicating whether the area to be identified meets the text line type condition;

[0045] A color alignment test is performed on each of the text lines to obtain a color alignment test result indicating whether the area to be identified meets the color alignment condition.

[0046] According to a method for identifying numbered modified areas in a document provided by the present invention, the content type detection includes text content detection and text type element detection, the content type condition includes a text content condition and a text type element condition, the text content condition is that each of the text lines does not contain a first text line without text content, the text type element condition is that each of the text lines does not contain a second text line without a text type element, and the content type detection result includes a text content detection result and a text type element result;

[0047] The performing of content type detection on each of the text lines to obtain a content type detection result of whether the area to be identified meets the content type condition includes:

[0048] Performing text content detection on each of the text lines to determine whether each of the text lines is a first text line without text content; if the first text line exists in each of the text lines, determining that the text content detection result indicates that the area to be identified does not meet the text content condition; if the first text line does not exist in each of the text lines, determining that the text content detection result indicates that the area to be identified meets the text content condition;

[0049] Perform text type element detection on each of the text lines to determine whether each of the text lines is a second text line containing a text type element; if a second text line exists in each of the text lines, determine that the text type element detection result is that the area to be identified does not meet the text type element condition; if no second text line exists in each of the text lines, determine that the text type element detection result is that the area to be identified meets the text type element condition.

[0050] According to a method for identifying numbered modified areas in a document provided by the present invention, the text line type detection includes writing direction detection, left boundary detection, and font style detection; the text line type condition includes a writing direction condition, a left boundary condition, and a font style condition; the writing direction condition is that the writing directions of the text lines are the same; the left boundary condition is that when the area to be identified contains a text line, the coordinates of the left boundary of the text line are less than or equal to the coordinates of the right boundary of the text line above the text line; the font style condition is that there is no first target text line pair in each text line pair; the text line pair is two connected text lines in each text line; and the font thickness ranges between the text lines in the first target text line pair do not overlap; the text line type detection result includes a writing direction detection result, a left boundary detection result, and a font style detection result;

[0051] The performing text line type detection on each of the text lines to obtain a text line type detection result indicating whether the area to be identified meets the text line type condition includes:

[0052] Performing writing direction detection on each of the text lines to determine the writing direction of each of the text lines; if the writing directions are the same, determining that the writing direction detection result indicates that the area to be identified meets the writing direction condition; if the writing directions are different, determining that the writing direction detection result indicates that the area to be identified does not meet the writing direction condition;

[0053] In the case where the area to be identified includes a text line, performing left boundary detection on the text line to determine whether the coordinates of the left boundary of the text line are greater than the coordinates of the right boundary of the text line above the text line; if not, determining that the left boundary detection result indicates that the area to be identified meets the left boundary condition; if so, determining that the left boundary detection result indicates that the area to be identified does not meet the left boundary condition;

[0054] Perform font style detection on each of the text line pairs to determine whether there is an intersection in the font thickness ranges between the text lines in each of the text line pairs; if so, determine that the font style detection result indicates that the area to be identified meets the font style conditions; if not, determine that the font style detection result indicates that the area to be identified does not meet the font style conditions.

[0055] According to a method for identifying numbered modified areas in a document provided by the present invention, the color alignment condition is that the color ranges of two adjacent text lines intersect;

[0056] The performing color alignment detection on each of the text lines to obtain a color alignment detection result of whether the to-be-identified area meets the color alignment condition includes:

[0057] For each of the text lines, identifying the color corresponding to each text object in the text line under the target color mode, and obtaining the color range corresponding to each text object in the text line;

[0058] Determining whether each text line pair is a second target text line pair, wherein the text line pair is two connected text lines in each of the text lines, and the color ranges corresponding to the two text lines in the second target text line pair have an intersection;

[0059] If so, determining that the color alignment detection result is that the to-be-identified area meets the color alignment condition;

[0060] If not, it is determined that the color alignment detection result is that the to-be-identified area does not meet the color alignment condition.

[0061] According to a method for identifying a number-modified region in a document provided by the present invention, determining whether the region to be identified is a number-modified region based on the association detection result, the number-modified detection result, and the number alignment detection result includes:

[0062] If the association detection result, the number modification detection result, and the number alignment detection result all indicate that a condition is met, determining that the area to be identified is a number modification area;

[0063] If any one of the association detection result, the number modification detection result, and the number alignment detection result does not meet the condition, it is determined that the area to be identified is not a number modification area.

[0064] According to a method for identifying a number-modified area in a document provided by the present invention, after determining that the area to be identified is not a number-modified area, the method further includes:

[0065] Obtain the number of text lines contained in the area to be identified;

[0066] When the number of rows is not equal to the third value, subtracting the fourth value from the number of rows to obtain an updated number of rows;

[0067] Determine each target sub-region in the to-be-recognized region, wherein the target sub-region contains the updated number of connected text lines;

[0068] The target sub-regions are respectively used as updated regions to be identified, and the step of performing text line association detection on each text line is continued to obtain an association detection result of whether the region to be identified meets a preset line association condition.

[0069] The present invention also provides a device for identifying numbered modified areas in a document, comprising:

[0070] A first determining module is configured to determine a region to be identified in a target document, wherein the region to be identified includes at least one text line;

[0071] A first detection module is configured to perform text line association detection on each of the text lines to obtain an association detection result indicating whether the area to be identified meets a preset line association condition;

[0072] A second detection module is configured to perform a number modification detection on each of the text lines to obtain a number modification detection result indicating whether the area to be identified meets a preset number modification condition;

[0073] A third detection module is configured to perform a number alignment detection on each of the text lines to obtain a number alignment detection result indicating whether the area to be identified meets a preset number alignment condition;

[0074] The second determining module is configured to determine whether the to-be-identified area is a number-modified area according to the association detection result, the number-modification detection result, and the number-alignment detection result.

[0075] The present invention also provides an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for identifying numbered modified areas in a document as described above is implemented.

[0076] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above methods for identifying numbered modified areas in a document.

[0077] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above methods for identifying numbered modified areas in a document.

[0078] The present invention provides a method and device for identifying number-modified areas in a document. The method and device determine an area to be identified in a target document, wherein the area to be identified includes at least one text line; perform text line association detection on each of the text lines to obtain an association detection result of whether the area to be identified meets a preset line association condition; perform number modification detection on each of the text lines to obtain a number modification detection result of whether the area to be identified meets a preset number modification condition; perform number alignment detection on each of the text lines to obtain a number alignment detection result of whether the area to be identified meets a preset number alignment condition; and determine whether the area to be identified is a number-modified area based on the association detection result, the number modification detection result, and the number alignment detection result. The present invention determines a number-modified area that may have number modification based on the features within and between text lines. It can not only identify number modifications at the line level, but can also be extrapolated to number modifications at the paragraph level. The number-modified areas modified by different numbers are combined with semantic features and visual features to comprehensively improve the determination accuracy, which is convenient for subsequent reading, conversion, and other processing of the document. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0080] Figure 1 1 is a flow chart of a method for identifying numbered modified areas in a document provided by the present invention;

[0081] Figure 2 is a schematic diagram of the numbered modified regions provided by the present invention;

[0082] Figure 3 It is a schematic structural diagram of the device for identifying numbered modified areas in a document provided by the present invention;

[0083] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0084] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0085] In order to facilitate a clearer understanding of the embodiments of the present invention, some relevant knowledge is first introduced as follows.

[0086] Number level (Level): refers to the specific level corresponding to each content item (ListItem) modified by the number. For example, 1., 1.2.1 are level 1 and level 3 respectively

[0087] A text object (TextObject) is one of the basic elements in a layout document (such as a PDF document). It contains the character code (CharCode) of the text content in the font, font information (Font), character coordinate box (CBox or BBox), and Unicode.

[0088] Font Weight: Used to specify the thickness of characters.

[0089] Writing mode is used to describe the direction of text layout, including LrTb (row direction is from left to right, paragraph direction is from top to bottom, adopted by most writing systems), RlTb (row direction is from right to left, paragraph direction is from top to bottom, usually used in Arabic / Hebrew writing systems), and TbRl (row direction is from top to bottom, paragraph direction is from right to left, usually used in writing systems such as ancient Chinese and Japanese).

[0090] Figure 1 This is one of the flow charts of the method and apparatus for identifying numbered modified areas in a document provided by the present invention. Figure 1 As shown, it includes steps 101 to 105, wherein:

[0091] Step 101: Determine a region to be recognized in a target document, where the region to be recognized includes at least one text line.

[0092] First of all, it should be noted that the execution subject of the present invention can be any electronic device that can identify the numbered modified area in the document, for example, it can be any one of a smart phone, a smart watch, a desktop computer, a laptop computer, etc.

[0093] Specifically, see Figure 2 , Figure 2 Schematic diagram of the number modification area provided by the present invention: the number modification area refers to the paragraph or text line modified by the number. Figure 2 The numbered modification area shown uses the number "1.1.1" at the beginning of the paragraph to introduce the paragraph content (that is, the text part after 1.1.1).

[0094] The area to be identified refers to the area in the target document that needs to be completely numbered; the target document can be a layout document, such as a PDF (Portable Document Format) document, an OFD (Open Fixed-layout Document) document, a streaming document, such as a Word document, a TXT document, a presentation document, such as a slide document, or a picture document, and the present invention does not impose any restrictions on this.

[0095] In practical applications, before determining the area to be identified in the target document, the target document must be acquired first. There are many methods for acquiring the target document, which are not limited in the present invention.

[0096] Exemplarily, a user uploads a target document through a document upload page, and accordingly, the execution entity obtains the target document.

[0097] Exemplarily, the execution subject receives a document acquisition instruction or a number modification area identification instruction, and accordingly, the execution subject acquires the target document from the storage area pointed to by the document acquisition instruction or the number modification area identification instruction.

[0098] On the basis of acquiring the target document, the area to be identified in the target document is further determined: the entire display area of ​​the target document can be used as the area to be identified; when a selection instruction is received, the selection instruction is mapped to the area of ​​the target document as the area to be identified.

[0099] It should be noted that a line recognizer can be used to identify text lines and paragraphs. If the target document is a layout document, the determination is based on the position of the TextObject in the document. If the target document is an image document, text feature extraction is performed using technologies such as OCR (Optical Character Recognition). If the target document is a streaming document, no special recognition is required.

[0100] Step 102: Perform text line association detection on each of the text lines to obtain an association detection result indicating whether the area to be identified meets a preset line association condition.

[0101] Specifically, text line association detection refers to detecting whether text lines within the area to be recognized are associated, that is, detecting whether the text lines within the area to be recognized meet preset line association conditions. The preset line association conditions are conditions that indicate whether the features between the text lines are associated. The association detection result refers to the result of the text line association detection.

[0102] In actual applications, after determining the area to be identified, text line association detection is performed on each text line in the area to be identified. For example, a line identifier can be used to perform text line association detection to obtain the association detection result of the area to be identified, that is, to determine whether the area to be identified meets the preset line association conditions.

[0103] Step 103: performing a number modification detection on each of the text lines to obtain a number modification detection result indicating whether the to-be-identified area meets a preset number modification condition.

[0104] Specifically, number modification detection refers to detecting whether text lines within the area to be identified have number modification features. A preset number modification condition indicates that each text line has number modification features. The number modification detection result is the result of the number modification detection on the text line, that is, the result of whether the area to be identified meets the number modification condition.

[0105] In practical applications, after determining the area to be identified, the significant semantics (Unicode) features of number modification can be utilized, that is, number modification detection can be performed on each text line according to the preset number modification conditions to determine whether the area to be identified meets the number modification detection results of the preset number modification conditions.

[0106] Step 104: performing a number alignment test on each of the text lines to obtain a number alignment test result indicating whether the area to be identified meets a preset number alignment condition.

[0107] Specifically, number alignment testing involves checking whether the numbered items of text lines within the area to be recognized are aligned, that is, whether the text lines within the area to be recognized meet the preset number alignment conditions. The preset number alignment conditions specify whether the numbered items of each text line are consistent. The number alignment test result refers to the result of the number alignment test on the text lines, that is, the result of whether the area to be recognized meets the preset number alignment conditions.

[0108] In practical applications, after determining the area to be identified, visual features can be combined, that is, number alignment detection can be performed on each text line according to preset number alignment conditions to determine whether the area to be identified meets the number alignment detection results of the preset number alignment conditions.

[0109] Step 105: Determine whether the to-be-identified region is a number-modified region based on the association detection result, the number-modification detection result, and the number-alignment detection result.

[0110] In practical applications, based on the determination of the association detection results, number modification detection results and number alignment detection results, it can be determined whether the area to be identified is a number modification area based on the association detection results, number modification detection results and number alignment detection results.

[0111] Exemplarily, the association detection results, number modification detection results and number alignment detection results are compared with their corresponding set results or conditions respectively. If they are the same, the area to be identified is determined to be a number modification area. If there are differences, it is determined that the area to be identified in the target document is not a number modification area.

[0112] The method for identifying number-modified areas in a document provided by the present invention is as follows: determining an area to be identified in a target document, wherein the area to be identified includes at least one text line; performing text line association detection on each of the text lines to obtain an association detection result of whether the area to be identified meets a preset line association condition; performing number modification detection on each of the text lines to obtain a number modification detection result of whether the area to be identified meets a preset number modification condition; performing number alignment detection on each of the text lines to obtain a number alignment detection result of whether the area to be identified meets a preset number alignment condition; and determining whether the area to be identified is a number-modified area based on the association detection result, the number modification detection result, and the number alignment detection result. The present invention determines a number-modified area that may have number modification based on features within and between text lines, and can not only identify number modifications at the line level, but can also be extrapolated to number modifications at the paragraph level. Number-modified areas modified by different numbers are combined with semantic features and visual features to comprehensively improve the accuracy of determination, facilitating subsequent reading, conversion, and other processing of the document.

[0113] Optionally, the number modification detection includes first character detection and / or number item detection, the preset number modification condition includes a first character condition and / or a number item condition, and the number modification detection result includes a first character detection result and / or a number item detection result; performing number modification detection on each of the text lines to obtain a number modification detection result indicating whether the area to be identified meets the preset number modification condition includes:

[0114] Performing a first character detection on each of the text lines to obtain a first character detection result indicating whether the area to be identified meets the first character condition; and / or,

[0115] A numbering item detection is performed on each of the text lines to obtain a numbering item detection result indicating whether the area to be identified meets the numbering item condition.

[0116] In actual applications, number modification detection includes at least one of first character detection and number item detection; when number modification detection includes first character detection, the preset number modification condition includes the first character condition, and the number modification detection result includes the first character detection result; when number modification detection includes number item detection, the preset number modification condition includes the number item condition, and the number modification detection result includes the number item detection result.

[0117] Specifically, when the number modification detection includes the first character detection, the first character detection can be performed on each text line according to the first character condition to determine whether the area to be identified meets the first character condition.

[0118] When the number modification detection includes number item detection, the number item detection can be performed on each text line according to the number item condition to determine whether the area to be identified meets the number item condition. In this way, the efficiency and accuracy of determining the number modification detection result can be improved.

[0119] It should be noted that when the number modification detection result includes the first character detection result and the number item detection result, if any one of the representations of the first character detection result and the number item detection result meets the conditions, the number modification detection result is that the area to be identified meets the number modification conditions.

[0120] Optionally, the first character condition is that the first characters of each text line belong to the same numbered character set;

[0121] The first character detection is performed on each of the text lines to obtain a first character detection result indicating whether the area to be identified meets the first character condition, including:

[0122] Obtaining the first character of each text line and searching for the Unicode of each first character;

[0123] Determining whether each of the first characters belongs to the same numbered character set according to the unified character encoding of each of the first characters;

[0124] If so, determining that the first character detection result is that the area to be recognized meets the first character condition;

[0125] If not, it is determined that the first character detection result is that the area to be recognized does not meet the first character condition.

[0126] Specifically, there are multiple numbered character sets, each of which contains a type of numbered characters. For example, the numbered characters in the first numbered character set include ①, ②, and ③, etc., the numbered characters in the second numbered character set include Ⅰ, Ⅱ, and Ⅲ, etc., and the numbered characters in the third numbered character set include 1, 2, and 3, etc.

[0127] In practice, for each text line, you can find the first TextObject in that text line and obtain the first Unicode value in that TextObject, that is, the Unicode code point of the first character in that text line. Then, you can traverse each text line and find the Unicode code point of the first character in each text line.

[0128] Furthermore, it is checked whether the Unicode of the first character in each text line belongs to the same numbered character set. If so, it means that the first character detection result is that the area to be recognized meets the first character condition; if not, it means that the first character detection result is that the area to be recognized does not meet the first character condition.

[0129] Since the number modification area has a numbered item with unified numbered characters, the unified character encoding of the first character can be used to determine whether the first character belongs to the same numbered character set, so that the first character detection result can be obtained simply, quickly and accurately, thereby improving the efficiency and accuracy of determining the first character detection result.

[0130] Optionally, the numbered item condition is that each of the text lines has a numbered item; performing numbered item detection on each of the text lines to obtain a numbered item detection result indicating whether the area to be identified meets the numbered item condition includes:

[0131] For each of the text lines, determining whether a numbered item exists in the text line based on characters between the first character and a set character in the text line, where the set character is the character preceding the first specified character, and the specified character is not a numeric character or a separator;

[0132] If a numbered item exists in each of the text lines, determining the numbered item detection result is that the area to be identified meets the numbered item condition;

[0133] If a specific text line exists in each of the text lines, it is determined that the numbered item detection result is that the area to be identified does not meet the numbered item condition, and the specific text line is a text line in which no numbered item exists.

[0134] Specifically, a numbered item is also a number, such as "1.1.1", "1)", etc.

[0135] In practical applications, for each text line, the first character to the set characters in the text line can be used to determine whether the text line contains the numbered item. If not, the text line is a specific text line. Each text line is traversed to determine whether each text line is a specific text line.

[0136] Further, determine whether there are specific text lines in each text line: if there are no specific text lines, it means that the beginning of each text line contains a numbered item, and at this time it can be determined that the numbered item detection result is that the area to be identified meets the numbered item conditions; if there are specific text lines, it means that not all text lines in the area to be identified contain numbered items at the beginning of the line, and at this time it can be determined that the numbered item detection result is that the area to be identified does not meet the numbered item conditions.

[0137] In this way, the numbered item detection result can be accurately obtained through the numeric characters and separators at the beginning of the text line, which improves the determination efficiency and accuracy of the numbered item detection result.

[0138] Optionally, determining whether a numbered item exists in the text line based on characters between the first character and a set character in the text line includes:

[0139] Recognize the first character of the text line to obtain a recognition result; and determine whether the text line meets a first condition based on the recognition result, wherein the first condition indicates that the first character is a numeric character;

[0140] Performing a delimiter search on characters between the first character and the set character in the text line to obtain a search result; and determining whether the text line meets a second condition based on the search result, the second condition indicating that a delimiter exists between the first character and the set character;

[0141] Obtaining, from characters between a second character and a set character in the text line, a first number of characters of at least one numeric character combination, the numeric character combination being at least one consecutive numeric character, the first number of characters being the number of numeric characters in the numeric character combination; and determining, based on each first number of characters, whether the text line meets a third condition, the third condition indicating that each first number of characters is less than or equal to a first value;

[0142] Obtaining, from characters between a second character and a set character in the text line, a second number of characters of at least one separator combination, the separator combination being at least one consecutive separator, the second number of characters being the number of separators in the separator combination; and determining whether the text line meets a fourth condition based on each second number of characters, the fourth condition being that each second number of characters is less than or equal to a second value;

[0143] If the text line satisfies the first condition, the second condition, the third condition, and the fourth condition at the same time, determining that a numbered item exists in the text line;

[0144] If the text line does not meet at least one of the first condition, the second condition, the third condition, and the fourth condition, it is determined that no numbered item exists in the text line.

[0145] In actual applications, the first character of the current text line is identified. If the first character is a numeric character, the text line meets the first condition. If the first character is not a numeric character, the text line does not meet the first condition.

[0146] It is determined whether there is a separator between the first character and the set character in the text line. If there is a separator, it means that the text line meets the second condition. If there is no separator, the text line does not meet the second condition.

[0147] Determine all numeric character combinations in the text line. For example, if the text line contains "11..2.35 Who am I?", the numeric character combinations include "11," "2," and "35." Then, count the number of numeric characters in each numeric character combination, i.e., the number of first characters. For example, the number of first characters in "11" and "35" is 2, while the number of first characters in "2" is 1. Furthermore, determine whether the number of first characters in each numeric character combination is less than or equal to the first value. If so, the text line meets the third condition. If not, the text line does not meet the third condition.

[0148] Determine all delimiter combinations in the text line. For example, if the text line contains "11..2.35 Who am I?", the delimiter combinations include ".." and ".". Then, count the number of delimiters in each segmentation combination, i.e., the number of second characters. For example, the number of second characters in ".." is always 2, while the number of first characters in "." is 1. Furthermore, determine whether the number of second characters in each delimiter combination is less than or equal to the second value. If so, the text line meets the fourth condition. If not, the text line does not meet the fourth condition.

[0149] If the text line does not meet at least one of the first to fourth conditions, it means that there is no numbered item in the text line; if the text line meets all of the first to fourth conditions, it means that there is a numbered item in the text line.

[0150] In this way, through numeric characters and separators, it is possible to determine whether a text line has elements that form a numbered item, and through the number of first characters and the number of second characters, it is possible to determine whether a text line has a structure that forms a numbered item, thereby improving the reliability and accuracy of determining whether a text line contains a numbered item.

[0151] It should be noted that the determination of whether a text line meets the first to fourth conditions may be performed simultaneously or in a preset order, and the present invention does not impose any limitation on this.

[0152] Exemplarily, the TextObject of the current text line is traversed, and the Unicode of each TextObject is obtained in turn. If the first Unicode representation is a numeric character (representing a number from 0 to 9), that is, the first character is a numeric character, then the text line meets the first condition and the judgment continues. If the subsequent characters are still numbers, the search continues, but if the numbers appear more than 2 times in a row (the first value), the judgment fails, that is, the text line does not meet the third condition. If a separator is found, such as ".", the text line is considered to meet the second condition, and the judgment continues. If numbers and separators appear and the numbers appear no more than 2 times in a row, and the separators appear no more than 1 time in a row (the second value), they are all legal, that is, the text line meets the third and fourth conditions. When the Unicode is guaranteed to be neither a numeric character nor a separator, the judgment ends. If the text line meets the first to fourth conditions at the end of the judgment, then it can be determined that the current text line contains numbered items.

[0153] In addition, after determining that a numbered item exists in the text line, the item level of the numbered item can also be obtained. The specific process is as follows:

[0154] The numbering level of the numbered items in the text line is determined based on the first combination quantity of the numeric character combinations and the second combination quantity of the separator combinations in the text line, the first combination quantity being the number of the numeric character combinations contained in the text line, and the second combination quantity being the number of the separator combinations contained in the text line.

[0155] In practical applications, the maximum value of the number of occurrences of the numeric character combination (first combination number) and the number of occurrences of the separator combination (second combination number) can be used as the Level level (numbering level) of the numbered item and saved.

[0156] For example, the numbered item is a combination of "number + separator + number", such as the form of "1.1.2". It is necessary to ensure that the separator "." appears at least once, and the last character can be "." or a number. There cannot be more than 2 consecutive numbers in a numbered item, for example, 1.231.2 is not compliant. The larger the number of occurrences of the numeric character combination and the separator combination, the more it indicates that 1.2.1 is a level 3 serial number (numbered item), and 1.2.1. is also a level 3 serial number (numbered item).

[0157] Optionally, performing number alignment detection on each of the text lines to obtain a number alignment detection result indicating whether the area to be identified meets a preset number alignment condition includes:

[0158] Calculating an average font size corresponding to the area to be identified based on the font size of each character in all the text lines;

[0159] Constructing a mapping table, wherein the mapping table records the first left boundary value corresponding to each numbering level;

[0160] For each of the text lines, obtaining a numbering level corresponding to the text line and a second left boundary value of a circumscribed rectangular box; searching the mapping table for a first left boundary value corresponding to the text line based on the numbering level corresponding to the text line; determining an absolute value of a difference between the first left boundary value and the second left boundary value corresponding to the text line; and determining that the text line is a set text line if the absolute value is greater than the average font size;

[0161] In a case where the set text line exists in each of the text lines, determining that the number alignment detection result is that the to-be-identified area does not meet the preset number alignment condition;

[0162] In a case where the set text line does not exist in each of the text lines, it is determined that the number alignment detection result is that the to-be-identified area meets the preset number alignment condition.

[0163] In practice, we first count the FontSize of each character in all TextObjects in each text line, then calculate the average FontSize of the entire text, which is the average font size of the area to be recognized. This is used as the threshold Thres. The average font size can be a fixed floating-point value greater than 0.

[0164] Then, a mapping table is constructed to record the BBox left edge value of the numbering level and the corresponding line, that is, the first left edge value corresponding to each numbering level. Among them, if it is a paragraph-level numbering modification, the corresponding line is the first text line of the paragraph.

[0165] Traverse each text line in the area to be identified in turn. For example, if the numbering level corresponding to the current text line is level 2, then if there is no first left boundary value corresponding to level 2 in the map, then insert the BBox left boundary value (second boundary value) corresponding to the current text line into the map, recorded as map[2] = BBox.left; if there is a first left boundary value corresponding to level 2 in the map, then calculate the absolute value of the difference between the current BBox.left (second boundary value) and the BBox.left (first boundary value) stored in map[2]. If the difference is less than or equal to Thres, the text content at the same level has the same indentation distance, and the difference in the left boundary value is small; if the difference is greater than Thres, then the text line is the set text line.

[0166] If the area to be identified contains a set text line, the number alignment test result indicates that the area to be identified does not meet the preset number alignment conditions, that is, not all of the area to be identified has number modifications and is not a number-modified area. If the area to be identified does not contain a set text line, the number alignment test result indicates that the area to be identified meets the preset number alignment conditions.

[0167] In this way, based on the principle that text content at the same level should have the same indentation distance and the left boundary of the BBox will not differ greatly, number alignment detection is performed on the identified area to improve the reliability and accuracy of the number alignment detection results.

[0168] Optionally, the text line association detection includes content type detection, text line type detection and color alignment detection, the preset line association condition includes a content type condition, a text line type condition and a color alignment condition, and the association detection result includes a content type detection result, a text line type detection result and a color alignment detection result;

[0169] The performing text line association detection on each of the text lines to obtain an association detection result of whether the area to be identified meets a preset line association condition includes:

[0170] Performing content type detection on each of the text lines to obtain a content type detection result of whether the area to be identified meets the content type condition;

[0171] Performing text line type detection on each of the text lines to obtain a text line type detection result indicating whether the area to be identified meets the text line type condition;

[0172] A color alignment test is performed on each of the text lines to obtain a color alignment test result indicating whether the area to be identified meets the color alignment condition.

[0173] Specifically, content type detection refers to the detection of content within a text line. The content type condition indicates whether the content type within a text line is consistent. Content types include blank line types (no text content) and text line types (no text type elements). Text line types indicate the presence of text content but no text type elements, such as image lines and vector path element lines. Blank line types indicate the absence of substantive content, meaning no text content.

[0174] Text line type detection refers to the detection of relationships between text lines or text line types. Relationships between text lines include whether the writing methods of the text lines are the same, whether the left margin of the current text line is larger than the right margin of the previous text line, and whether the font styles are the same. Text line type refers to the type of text line containing text, including writing direction, left margin, and font style. The text line type condition represents the condition for whether the types of text lines are relatively consistent. The text line type detection result refers to the result of text line type detection on the text line, that is, whether the area to be identified meets the text line type condition.

[0175] Color alignment testing involves checking the color range corresponding to each text line. A color range refers to the range of color values ​​and can be in either ARGB or RGB color mode. The color alignment condition determines whether the color ranges corresponding to each text line are consistent. The color alignment test result indicates whether the area to be identified meets the color alignment condition.

[0176] In actual applications, a line identifier can be used to detect the content type of text lines, and then determine whether the area to be identified meets the content type conditions, that is, determine the content type detection result; the color range of text objects in each text line can be identified, such as the line identifier can be used to identify the color range of text objects in each text line, obtain the color range corresponding to each text line, and then perform color alignment detection based on each color range to determine whether the area to be identified meets the color detection result of the color detector conditions; at the same time, the line identifier can be used to perform text line type detection on each text line, and then determine whether the area to be identified meets the text line type conditions, that is, determine the text line type detection result.

[0177] In this way, the text association detection results are determined in a more fine-grained manner from three dimensions: content type, color alignment, and text line type, thereby improving the accuracy and reliability of the text line association detection results.

[0178] Optionally, the content type detection includes text content detection and text type element detection, the content type condition includes a text content condition and a text type element condition, the text content condition is that each of the text lines does not contain a first text line without text content, the text type element condition is that each of the text lines does not contain a second text line without a text type element, and the content type detection result includes a text content detection result and a text type element detection result;

[0179] The performing of content type detection on each of the text lines to obtain a content type detection result of whether the area to be identified meets the content type condition includes:

[0180] Performing text content detection on each of the text lines to determine whether each of the text lines is a first text line without text content; if the first text line exists in each of the text lines, determining that the text content detection result indicates that the area to be identified does not meet the text content condition; if the first text line does not exist in each of the text lines, determining that the text content detection result indicates that the area to be identified meets the text content condition;

[0181] Perform text type element detection on each of the text lines to determine whether each of the text lines is a second text line containing a text type element; if a second text line exists in each of the text lines, determine that the text type element detection result is that the area to be identified does not meet the text type element condition; if no second text line exists in each of the text lines, determine that the text type element detection result is that the area to be identified meets the text type element condition.

[0182] Specifically, the first text line, also known as the blank line, refers to a text line without text content, such as a line without any characters, a line with only spaces, a line with only lines, etc. The second text line refers to a text line with text content but no text type element.

[0183] In actual applications, for each text line, the text line can be detected for main content based on the content or characters in the text line, that is, to identify whether the text line is the first text line without main content. For example, if the current text line only contains underlines or spaces modified by underlines, then there is no specific text content, and the text line is the first text line. If the first text line exists in each text line, the main content detection result is determined to be that the area to be identified does not meet the main content conditions. If the first text line does not exist in each text line, the main content detection result is determined to be that the area to be identified meets the main content conditions. In this way, the efficiency and accuracy of determining the main content detection results can be improved.

[0184] For each text line, a text type element detection can be performed on the text line based on the content or characters in the text line, that is, it is identified whether the text line is a second text line that does not contain a text type element detection. For example, if there are only pictures and / or an appropriate amount of path elements in the current text line, then there is no text type element, and the text line is the second text line. If there is a second text line in each text line, the text type element detection result is determined to be that the area to be identified does not meet the text type element conditions. If there is no second text line in each text line, the text type element detection result is determined to be that the area to be identified meets the text type element conditions. In this way, the efficiency and accuracy of determining the text type element detection results can be improved.

[0185] Optionally, the text line type detection includes writing direction detection, left boundary detection and font style detection, the text line type condition includes a writing direction condition, a left boundary condition and a font style condition, the writing direction condition is that the writing directions of the text lines are the same, the left boundary condition is that when the area to be identified contains a text line, the coordinates of the left boundary of the text line are less than or equal to the coordinates of the right boundary of the text line above the text line, the font style condition is that there is no first target text line pair in each text line pair, the text line pair is two connected text lines in each text line, and the font thickness ranges between the text lines in the first target text line pair do not overlap, and the text line type detection result includes a writing direction detection result, a left boundary detection result and a font style detection result;

[0186] The performing text line type detection on each of the text lines to obtain a text line type detection result indicating whether the area to be identified meets the text line type condition includes:

[0187] Performing writing direction detection on each of the text lines to determine the writing direction of each of the text lines; if the writing directions are the same, determining that the writing direction detection result indicates that the area to be identified meets the writing direction condition; if the writing directions are different, determining that the writing direction detection result indicates that the area to be identified does not meet the writing direction condition;

[0188] In the case where the area to be identified includes a text line, performing left boundary detection on the text line to determine whether the coordinates of the left boundary of the text line are greater than the coordinates of the right boundary of the text line above the text line; if not, determining that the left boundary detection result indicates that the area to be identified meets the left boundary condition; if so, determining that the left boundary detection result indicates that the area to be identified does not meet the left boundary condition;

[0189] Perform font style detection on each of the text line pairs to determine whether there is an intersection in the font thickness ranges between the text lines in each of the text line pairs; if so, determine that the font style detection result indicates that the area to be identified meets the font style conditions; if not, determine that the font style detection result indicates that the area to be identified does not meet the font style conditions.

[0190] In practical applications, a line identifier can be used to detect the writing direction of each text line, identifying the writing direction of the text line. This process then determines whether the writing directions of all text lines are the same. If so, the writing direction detection result indicates that the area to be identified meets the writing direction condition. If not, the writing direction detection result indicates that the area to be identified does not meet the writing direction condition. This improves the efficiency and accuracy of determining the writing direction detection results.

[0191] Preferably, determining whether the writing directions of all text lines are the same can be as follows: determining whether the writing directions of all text lines are specified writing directions, the specified writing direction is LrTb, that is, the sorting within the line is from left to right and the sorting between lines is from top to bottom.

[0192] For the case where the area to be identified contains a text line, if there is a previous text line for the text line, the left boundary detection can be performed on the text line: obtain the BBox of the text line to determine the coordinates of the left boundary of the text line, and obtain the BBox of the text line before the text line to obtain the coordinates of the right boundary of the previous text line; further determine whether the coordinates of the left boundary are greater than the coordinates of the right boundary; if not, the left boundary detection result is determined to be that the area to be identified meets the left boundary condition; if greater, it means that the previous text line is the end line of the previous paragraph, and the text line is a new text line. If only looking at these two lines, it is obvious that the text line cannot independently become a numbered modification, and the left boundary detection result can be determined to be that the area to be identified does not meet the left boundary condition. In this way, the efficiency and accuracy of the left boundary detection can be improved.

[0193] It should be noted that if the area to be identified contains a text line which is the first text line of the target document, then there is no need to perform left boundary detection.

[0194] For each text line pair, the font of each text line in the pair is identified, and the corresponding font weight ranges for each text line in the pair are obtained. A determination is then made as to whether the font weight ranges for each text line in each text line pair intersect. If so, the font style detection result is determined to indicate that the area to be identified meets the font style conditions; if not, the font style detection result is determined to indicate that the area to be identified does not meet the font style conditions. This improves the efficiency and accuracy of determining font style detection results.

[0195] By determining the text line type detection results in a more fine-grained manner from three dimensions: writing direction detection, left boundary detection, and font style detection, the accuracy and reliability of the text line type detection results are improved.

[0196] Optionally, the color alignment condition is that the color ranges of two adjacent text lines intersect; performing color alignment detection on each of the text lines to obtain a color alignment detection result of whether the area to be identified meets the color alignment condition includes:

[0197] For each of the text lines, identifying the color corresponding to each text object in the text line under the target color mode, and obtaining the color range corresponding to each text object in the text line;

[0198] A color alignment detection result is used to determine whether the to-be-identified area meets the color alignment condition according to the color range.

[0199] Specifically, a color model is a way to represent color in the digital world. To represent various colors in the digital world, colors are typically divided into several components. Color models include RGB, CMYK, HSB, LAB, and ARGB. Different color models correspond to different colors and, accordingly, different color ranges. The target color model is a pre-defined color model. The color alignment condition indicates that the text lines in the area to be recognized are consistent in color.

[0200] In practical applications, for each text line in the area to be recognized, the color corresponding to each text object in the text line is identified under the target color mode, thus obtaining the color range corresponding to each text object in the text line. The color range corresponding to each text object in each text line is then used to determine whether the area to be recognized meets the color alignment conditions, i.e., determine the color alignment test result. This can improve the speed of color alignment test results.

[0201] Optionally, the color alignment detection result of determining whether the to-be-identified area meets the color alignment condition according to the color range includes:

[0202] Determining whether each text line pair is a second target text line pair, wherein the text line pair is two connected text lines in each of the text lines, and the color ranges corresponding to the two text lines in the second target text line pair have an intersection;

[0203] If so, determining that the color alignment detection result is that the to-be-identified area meets the color alignment condition;

[0204] If not, it is determined that the color alignment detection result is that the to-be-identified area does not meet the color alignment condition.

[0205] In practice, for each text line pair (two connected text lines), if the color ranges of the corresponding colors of any two consecutive text lines (text line pair) do not intersect or overlap in value, the color alignment test result indicates that the area to be identified does not meet the color alignment conditions; otherwise, the color alignment test result indicates that the area to be identified meets the color alignment conditions. In this way, by determining the color alignment test result based on whether there is overlap in the color range, it can avoid subtle color differences in text lines in the numbered modified area that may lead to recognition failure. This improves the efficiency and accuracy of identifying numbered modified areas in the document while ensuring the accuracy of the color alignment test results.

[0206] Optionally, the color mode is an ARGB color mode; the color corresponding to each text object in the text line under the target color mode is identified, and the color range corresponding to each text object in the text line is obtained. The specific implementation process may be as follows:

[0207] Identify the transparency, red, green, and blue corresponding to each text object in the text line in the ARGB color mode, and obtain the transparency color range, red color range, green color range, and blue color range corresponding to each text object in the text line.

[0208] In practice, the ARGB color model is based on the principle of color luminescence and has four channels: transparency (Alpha, A), red (Red, R), green (Green, G), and blue (Blue, B), each corresponding to a color. Therefore, in the ARGB color model, for each text line, the transparency color range, red color range, green color range, and blue color range corresponding to each text object in that text line are obtained. This more fine-grained determination of the color range from multiple colors can improve the accuracy of the color range.

[0209] Optionally, when the color mode is ARGB, the determination of whether each text line pair is the second target text line pair can be performed by determining, for each text line pair, whether the transparency color ranges, red color ranges, green color ranges, and blue color ranges of the two text lines in the text line pair overlap. That is, only when it is determined that there is overlap in the transparency color range, red color range, green color range, and blue color range can the text line pair be determined to be the second text line pair. This improves the accuracy of determining the second target text line pair. If there is no overlap in any of the transparency color range, red color range, green color range, and blue color range, there is no need to determine the remaining undetermined text line pairs, and the text line pair is not the second text line pair. This can reduce the amount of data processing to a certain extent and improve determination efficiency.

[0210] Optionally, determining whether the to-be-identified region is a numbered modification region according to the association detection result, the numbered modification detection result, and the numbered alignment detection result includes:

[0211] When the association detection result, the numbered modification detection result, and the numbered alignment detection result all indicate compliance with the conditions, determine that the to-be-identified region is a numbered modification region;

[0212] When any one of the association detection result, the numbered modification detection result, and the numbered alignment detection result indicates non-compliance with the conditions, determine that the to-be-identified region is not a numbered modification region.

[0213] In practical applications, compare the association detection result, the numbered modification detection result, and the numbered alignment detection result with the set identification conditions for numbered modification regions respectively, where the set identification conditions for numbered modification regions are that the association detection result, the alignment method detection result, and the color detection result are all compliant.

[0214] The association detection result can be compared with the set identification conditions for numbered modification regions, the numbered modification detection result can be compared with the set identification conditions for numbered modification regions, and the numbered alignment detection result can be compared with the set identification conditions for numbered modification regions to obtain three comparison results. Determine whether each comparison result is compliant. If so, determine that the to-be-identified region in the target document is a numbered modification region; if not, determine that the to-be-identified region in the target document is not a numbered modification region.

[0215] The association detection result, the numbered modification detection result, and the numbered alignment detection result can be constructed into a matrix, and determine whether the matrix meets the set identification conditions for numbered modification regions. If so, determine that the to-be-identified region in the target document is a numbered modification region; if not, determine that the to-be-identified region in the target document is not a numbered modification region. In this way, by setting the identification conditions for numbered modification regions, it is possible to quickly and accurately determine whether the to-be-identified region is a numbered modification region.

[0216] For example, assume that a total of N rows are currently recognized. Then a two-dimensional matrix matrix[m][n] can be constructed, and each element in it is initialized to False, where matrix[m][n] represents whether all rows from the m-th row to the n-th row can form a numbered modification region, 1 ≤ m < n ≤ N. Update each element in matrix[m][n] according to the association detection result, the numbered modification detection result, and the numbered alignment detection result. If all elements of matrix[m][n] meet the set identification conditions for numbered modification regions, that is, all are compliant, then matrix[m][n] = True, that is, the to-be-identified region in the target document is a numbered modification region.

[0217] It should be noted that when judging whether each element of matrix[m][n] meets the set number modification area recognition conditions, it can be judged in sequence. If any element does not meet the conditions, there is no need to judge the remaining elements. At this time, it can be determined that the area to be identified in the target document is not a number modification area, that is, matrix[m][n]=False.

[0218] Optionally, after determining that the area to be identified is not a number-modified area, the method further includes:

[0219] Obtain the number of text lines contained in the area to be identified;

[0220] When the number of rows is not equal to the third value, subtracting the fourth value from the number of rows to obtain an updated number of rows;

[0221] Determine each target sub-region in the to-be-recognized region, wherein the target sub-region contains the updated number of connected text lines;

[0222] The target sub-regions are respectively used as updated regions to be identified, and the step of performing text line association detection on each text line is continued to obtain an association detection result of whether the region to be identified meets a preset line association condition.

[0223] Specifically, the fourth preset value and the fifth preset value are both pre-set values ​​used to adjust the accuracy of paragraph recognition or to recognize paragraphs of a set length. The fourth preset value and the fifth preset value can be the same or different. Preferably, to improve accuracy, the fourth preset value and the fifth preset value are both one.

[0224] In practical applications, the order of determination should be from longest to shortest line length. That is, the longest interval (the area to be recognized) should be determined first, and then the sub-intervals (the target sub-areas within the area to be recognized) should be determined based on the result. This can avoid splitting a paragraph that occupies multiple text lines into multiple paragraphs, thereby improving the accuracy and completeness of paragraph recognition.

[0225] Optionally, the step of using each target sub-region as an updated region to be identified, continuing to perform text line association detection on each text line, and obtaining an association detection result of whether the region to be identified meets a preset line association condition may be specifically implemented as follows:

[0226] The first target sub-region is used as the updated area to be identified, and the step of performing text line association detection on each of the text lines to obtain an association detection result of whether the area to be identified meets the preset line association condition is continued, wherein the first target sub-region is any target sub-region;

[0227] If the first target subregion is a paragraph, it is determined that the second target subregion is not a paragraph, wherein the combination of the second target subregion and the first target subregion is any designated area to be identified before updating, and the designated area to be identified is not a paragraph.

[0228] In practical applications, each target sub-region can be treated as an updated region to be identified, and the numbered modified region identification process can be continued for each target sub-region. If the first target sub-region is a numbered modified region, the second target sub-region can be directly determined not to be a numbered modified region. In this way, the numbered modified region identification process can be performed on the second target sub-region, greatly reducing the amount of data processing and improving the efficiency of numbered modified region identification while ensuring the accuracy of numbered modified region identification.

[0229] Continuing with the above example, if matrix[i][j] = True, and (i,j) contains (k,l), then there is no need to evaluate matrix[k][l]. For example, assuming there are 10 rows, matrix[1]

[10] should be evaluated first. Next, matrix[1][9] and matrix[2]

[10] , which have 9 rows, should be evaluated. Then matrix[1][8], matrix[2][8], and matrix[3]

[10] should be evaluated, and so on. If matrix[1][9] is False and matrix[1][8] is True, since the union of matrix[2][9] and matrix[1][8] is matrix[1][9], there is no need to judge matrix[2][9] based on the above process, and matrix[2][9] can be directly determined to be False. Similarly, if the union of matrix[1][8] and matrix[3]

[10] is matrix[1]

[10] , there is no need to judge matrix[3]

[10] based on the above process, and matrix[2][9] can be directly determined to be False. That is, when matrix[1][9] is False and matrix[1][8] is True, since both matrix[2][9] and matrix[3]

[10] have an intersection with matrix[1][9], there is no need to judge either of them, and only the result of matrix[9]

[10] needs to be determined. That is, if matrix[1][9] is False and matrix[1][8] is True, then matrix[2][9] must be False, because if both are True, then matrix[1][9] should be True, which is a contradiction. Therefore, the present invention does not need to worry about missing the optimal result and can accurately identify each numbered modified area.

[0230] Exemplarily, the method for identifying the numbered modified area in the target document is as follows:

[0231] Precondition: Line recognizer. If it is a picture type, it is necessary to first extract the text content therein using technologies such as OCR. If it is a layout document, it is necessary to determine according to the position where the TextObject appears in the document. If it is a streaming document, no special recognition is required.

[0232] 1. For a given text block area (target document), disassemble the lines therein into all possible combinations. For example, for an n-line text area, we can select the start from the i-th line and end at the j-th line, as long as 0 < i < j ≤ n is ensured, then the area to be recognized is obtained. The given lines mentioned later refer to any i to j lines of text here.

[0233] If it is at the line level, the text block area is a piece of text content, with lines as the unit; if it is at the paragraph level, the text block area is the full text, with paragraphs as the unit.

[0234] 2. For a given number of text lines (area to be recognized), first determine whether all the currently given text lines have the numbered modification attribute through the following determination conditions. If any one of the conditions is met, it is considered that not all the currently given text lines have the numbered modification attribute (for example, if 3 - 5 lines are selected, even if lines 3 - 4 have the numbered modification and line 5 does not, it is considered that lines 3 - 5 do not all have the numbered modification), and quickly filter out the intervals that do not all have the numbered modification through the attribute characteristics.

[0235] (1) There is no body content in any line (corresponding to the body content condition that there is body content in each line).

[0236] (2) In any line, although there is body content, there is no text type element (corresponding to the text type element condition that each line is body content containing text type elements).

[0237] (3) If the given line has only one line, and the left coordinate of the current line coordinate box BBox is greater than the right coordinate of the previous line (if any) coordinate box BBox (corresponding to the left boundary condition that the left boundary coordinate of the text line is less than or equal to the right boundary coordinate of the previous text line of the text line).

[0238] (4) According to the method for quickly identifying the writing direction of text content in PDF, determine the writing direction of each line. Among all the given lines, there is a situation where the writing direction is not LrTb (corresponding to the writing direction condition that the writing direction of each line is LrTb).

[0239] (5) If there are any two consecutive lines within all given lines, there is no intersection between the Font Weights given in the first TextObject in the line, or the Italic attributes are not equal (corresponding to the font style condition, the font style condition is that the first target text line pair does not exist in each text line pair, and the font weight ranges between the text lines in the first target text line pair do not overlap).

[0240] (6) Count the value ranges of the ARGB four channels of all TextObject colors in each given row; if there are any two consecutive rows where the color ranges of any channel do not intersect (corresponding to the color alignment condition, the color ranges of the two connected text rows all intersect).

[0241] 3. For the unidentified region filtered through step 1 above, we use the significant semantic (Unicode) features of number modification to determine whether all rows in the current range have number modification features. If every row in the given range meets all of the following criteria, then all rows in the current range are considered to have number modification features.

[0242] (1) Find the first TextObject in each row and get the first Unicode. If they are all the same character set in the labeled character set (corresponding to the first character condition),

[0243] (2) Sequentially traverse the TextObjects of each line and obtain the Unicode of each TextObject in turn. If the first Unicode representation is a numeric character (representing a number from 0 to 9), that is, the first character is a numeric character, then the text line meets the first condition and the judgment continues. If the subsequent characters are still numeric characters, continue to search, but if the number appears more than 2 times in a row (the first value), the judgment fails, that is, the text line does not meet the third condition. If a separator is found, such as ".", the text line is considered to meet the second condition and the judgment continues. If a number and a separator appear and the number appears no more than 2 times in a row and the separator appears no more than 1 time in a row (the second value), they are all legal, that is, the text line meets the third and fourth conditions. When the Unicode is guaranteed to be neither a numeric character nor a separator, the judgment ends. If the text line meets the first to fourth conditions at the end of the judgment, then it can be determined that the current text line contains a numbered item (corresponding to the numbered item condition).

[0244] 4. For the to-be-recognized area in 2 above, it can be highly probable that there are number decorations in all lines within the to-be-recognized area. However, to distinguish it from mathematical calculation formulas (such as 1 + 2 = 3 which belongs to a mathematical calculation formula) and IP addresses (IPs like 1.0.0.1 are prone to confusion with Level 4), it is still necessary to combine visual features (the BBox of the line). If the given interval fully meets the following conditions (preset number alignment conditions), it can be considered that the entire given interval has the feature of number decoration.

[0245] (1) First, count the FontSize of each text in all TextObjects within each line to obtain the overall average FontSize as the threshold Thres.

[0246] (2) Construct a mapping table map to record the level hierarchy and the left boundary value of the BBox of the corresponding line. Traverse each line in turn. For example, if the current line is level2, then if there is no left boundary value of the line BBox corresponding to level2 in map, insert it, denoted as map[2] = BBox.left. If it exists in map, take it out and compare the absolute value of the difference between the current BBox.left and the BBox.left stored in map[2]. If the difference is less than Thres, it is legal; otherwise, it is considered illegal, that is, not all lines in the current interval have number decoration.

[0247] 5. Assume that a total of N lines are currently recognized. Then a two-dimensional matrix matrix[m][n] can be constructed, and each element in it is initialized to False, where matrix[m][n] represents whether all lines from the m-th line to the n-th line can form a number decoration area, and 1 ≤ m < n ≤ N. Update each element in matrix[m][n] according to the association detection result, number decoration detection result, and number alignment detection result. If all elements in matrix[m][n] meet the set number decoration area recognition conditions, that is, all are compliant, then matrix[m][n] = True, that is, the to-be-recognized area in the target document is a number decoration area.

[0248] The order of determination should be from long to short. If matrix[i][j] = True, and (i,j) contains (k,l), then matrix[k][l] does not need to be determined. For example, if there are 10 rows in total, matrix[1]

[10] should be determined first. Next, matrix[1][9] and matrix[2]

[10] with 9 rows should be determined. Then matrix[1][8], matrix[2][8], and matrix[3]

[10] should be determined, and so on. If matrix[1][9] is False and matrix[1][8] is True, since the union of matrix[2][9] and matrix[1][8] is matrix[1][9], there is no need to judge matrix[2][9] based on the above process, and matrix[2][9] can be directly determined to be False. Similarly, if the union of matrix[1][8] and matrix[3]

[10] is matrix[1]

[10] , there is no need to judge matrix[3]

[10] based on the above process, and matrix[2][9] can be directly determined to be False. That is, when matrix[1][9] is False and matrix[1][8] is True, since both matrix[2][9] and matrix[3]

[10] have an intersection with matrix[1][9], there is no need to judge either of them, and only the result of matrix[9]

[10] needs to be determined. That is, if matrix[1][9] is False and matrix[1][8] is True, then matrix[2][9] must be False, because if both are True, then matrix[1][9] should be True, which is a contradiction. Therefore, the present invention does not need to worry about missing the optimal result and can accurately identify each numbered modified area.

[0249] 6. Based on the matrix result, obtain the specific row number and region with the numbered modifier. For example, if there are 10 rows, the region to be recognized is rows 1-10. First, look at the 10-row region 1-10. If it is False, the results for the 9-row region 1-9 and 2-10 are False. If both are False, look at the 8-row region. If 1-8 is True, then rows 1 to 8 are considered to form a single region. Update the range to 9-10 and repeat the above process, first looking at the 2-row region. And so on.

[0250] (1) First, obtain the number of rows that are True in the long interval segment, and obtain several number-modified areas with number modifications.

[0251] (2) Obtain the results that do not intersect with the above interval segment but the matrix result is True, and obtain several numbered modified areas with numbered modifications.

[0252] (3) Repeat the above (1) and (2) steps until all rows are traversed.

[0253] 7. Label each numbered row interval with a corresponding label to obtain the final recognition result.

[0254] The embodiments provided by the present invention can be applied to layout documents (such as PDF) and streaming documents, as well as images whose text content is identified through technologies such as OCR. It can support the recognition of line-level numbering modifications and paragraph-level numbering modifications. Then, when the reader typeset the document, it can add additional indentation alignment based on the numbering modifications to facilitate user reading; when copying the document content, its structural features can be retained; when performing document type conversion, the entire document can be processed together to unify the text style with the same hierarchical numbering. It can accurately identify all areas with numbering modifications in the document, namely, numbering modification areas. It can not only recognize the line level, but also the paragraph level. At the same time, it can also distinguish between connected numbering modification areas with different numbering modifications. Combining semantic features and visual features, the judgment accuracy is comprehensively improved. It facilitates document reading, conversion, etc. In addition, it can also recognize line-level lists and paragraph-level lists, calculate the specific numbering level of each list item, and distinguish between different consecutive lists.

[0255] The present invention can be applied to scenarios such as layout document conversion, layout document rearrangement, layout document editing, and streaming document display.

[0256] The following describes the device for identifying numbered modified areas in a document provided by the present invention. The device for identifying numbered modified areas in a document described below and the method for identifying numbered modified areas in a document described above can be referenced to each other.

[0257] Figure 3 Schematic diagram of the structure of the device for identifying numbered modified areas in a document provided by the present invention. Figure 3 As shown, the device 300 for identifying numbered modified areas in a document includes:

[0258] A first determining module 301 is configured to determine a region to be identified in a target document, where the region to be identified includes at least one text line;

[0259] The first detection module 302 is configured to perform text line association detection on each of the text lines to obtain an association detection result indicating whether the area to be identified meets a preset line association condition;

[0260] The second detection module 303 is configured to perform a number modification detection on each of the text lines to obtain a number modification detection result indicating whether the area to be identified meets a preset number modification condition;

[0261] The third detection module 304 is configured to perform a number alignment detection on each of the text lines to obtain a number alignment detection result indicating whether the area to be identified meets a preset number alignment condition;

[0262] The second determination module 305 is configured to determine whether the to-be-identified region is a number-modified region according to the association detection result, the number-modification detection result, and the number-alignment detection result.

[0263] The present invention provides a device for identifying number-modified areas in a document. The device determines an area to be identified in a target document, wherein the area to be identified includes at least one text line; performs a text line association test on each of the text lines to obtain an association test result of whether the area to be identified meets a preset line association condition; performs a number modification test on each of the text lines to obtain a number modification test result of whether the area to be identified meets a preset number modification condition; performs a number alignment test on each of the text lines to obtain a number alignment test result of whether the area to be identified meets a preset number alignment condition; and determines whether the area to be identified is a number-modified area based on the association test result, the number modification test result, and the number alignment test result. The present invention determines a number-modified area that may have number modification based on the features within and between text lines. It can not only identify number modifications at the line level, but can also be extrapolated to number modifications at the paragraph level. The number-modified areas modified by different numbers are combined with semantic features and visual features to comprehensively improve the accuracy of determination, facilitating subsequent reading, conversion, and other processing of the document.

[0264] Optionally, the number modification detection includes first character detection and / or number item detection, the preset number modification condition includes a first character condition and / or a number item condition, and the number modification detection result includes a first character detection result and / or a number item detection result;

[0265] The second detection module 303 is further configured to:

[0266] Performing a first character detection on each of the text lines to obtain a first character detection result indicating whether the area to be identified meets the first character condition; and / or,

[0267] A numbering item detection is performed on each of the text lines to obtain a numbering item detection result indicating whether the area to be identified meets the numbering item condition.

[0268] Optionally, the first character condition is that the first characters of each text line belong to the same numbered character set;

[0269] The second detection module 303 is further configured to:

[0270] Obtaining the first character of each text line and searching for the Unicode of each first character;

[0271] Determining whether each of the first characters belongs to the same numbered character set according to the unified character encoding of each of the first characters;

[0272] If so, determining that the first character detection result is that the area to be recognized meets the first character condition;

[0273] If not, it is determined that the first character detection result is that the area to be recognized does not meet the first character condition.

[0274] Optionally, the numbered item condition is that each of the text lines has a numbered item;

[0275] The second detection module 303 is further configured to:

[0276] For each of the text lines, determining whether a numbered item exists in the text line based on characters between the first character and a set character in the text line, where the set character is the character preceding the first specified character, and the specified character is not a numeric character or a separator;

[0277] If a numbered item exists in each of the text lines, determining the numbered item detection result is that the area to be identified meets the numbered item condition;

[0278] If a specific text line exists in each of the text lines, it is determined that the numbered item detection result is that the area to be identified does not meet the numbered item condition, and the specific text line is a text line in which no numbered item exists.

[0279] Optionally, the second detection module 303 is further configured to:

[0280] Recognize the first character of the text line to obtain a recognition result; and determine whether the text line meets a first condition based on the recognition result, wherein the first condition indicates that the first character is a numeric character;

[0281] Performing a delimiter search on characters between the first character and the set character in the text line to obtain a search result; and determining whether the text line meets a second condition based on the search result, the second condition indicating that a delimiter exists between the first character and the set character;

[0282] Obtaining, from characters between a second character and a set character in the text line, a first number of characters of at least one numeric character combination, the numeric character combination being at least one consecutive numeric character, the first number of characters being the number of numeric characters in the numeric character combination; and determining, based on each first number of characters, whether the text line meets a third condition, the third condition indicating that each first number of characters is less than or equal to a first value;

[0283] Obtaining, from characters between a second character and a set character in the text line, a second number of characters of at least one separator combination, the separator combination being at least one consecutive separator, the second number of characters being the number of separators in the separator combination; and determining whether the text line meets a fourth condition based on each second number of characters, the fourth condition being that each second number of characters is less than or equal to a second value;

[0284] If the text line satisfies the first condition, the second condition, the third condition, and the fourth condition at the same time, determining that a numbered item exists in the text line;

[0285] If the text line does not meet at least one of the first condition, the second condition, the third condition, and the fourth condition, it is determined that no numbered item exists in the text line.

[0286] Optionally, the second detection module 303 is further configured to:

[0287] The numbering level of the numbered items in the text line is determined based on the first combination quantity of the numeric character combinations and the second combination quantity of the separator combinations in the text line, the first combination quantity being the number of the numeric character combinations contained in the text line, and the second combination quantity being the number of the separator combinations contained in the text line.

[0288] Optionally, the third detection module 304 is further configured to:

[0289] Calculating the average font size corresponding to the area to be identified based on the font size of each character in all the text lines;

[0290] Constructing a mapping table, wherein the mapping table records the first left boundary value corresponding to each numbering level;

[0291] For each of the text lines, obtaining a numbering level corresponding to the text line and a second left boundary value of a circumscribed rectangular box; searching the mapping table for a first left boundary value corresponding to the text line based on the numbering level corresponding to the text line; determining an absolute value of a difference between the first left boundary value and the second left boundary value corresponding to the text line; and determining that the text line is a set text line if the absolute value is greater than the average font size;

[0292] In a case where the set text line exists in each of the text lines, determining that the number alignment detection result is that the to-be-identified area does not meet the preset number alignment condition;

[0293] In a case where the set text line does not exist in each of the text lines, it is determined that the number alignment detection result is that the to-be-identified area meets the preset number alignment condition.

[0294] Optionally, the text line association detection includes content type detection, text line type detection and color alignment detection, the preset line association condition includes a content type condition, a text line type condition and a color alignment condition, and the association detection result includes a content type detection result, a text line type detection result and a color alignment detection result;

[0295] The first detection module 302 is further configured to:

[0296] Performing content type detection on each of the text lines to obtain a content type detection result of whether the area to be identified meets the content type condition;

[0297] Performing text line type detection on each of the text lines to obtain a text line type detection result indicating whether the area to be identified meets the text line type condition;

[0298] A color alignment test is performed on each of the text lines to obtain a color alignment test result indicating whether the area to be identified meets the color alignment condition.

[0299] Optionally, the content type detection includes text content detection and text type element detection, the content type condition includes a text content condition and a text type element condition, the text content condition is that each of the text lines does not contain a first text line without text content, the text type element condition is that each of the text lines does not contain a second text line without a text type element, and the content type detection result includes a text content detection result and a text type element detection result;

[0300] The first detection module 302 is further configured to:

[0301] Performing text content detection on each of the text lines to determine whether each of the text lines is a first text line without text content; if the first text line exists in each of the text lines, determining that the text content detection result indicates that the area to be identified does not meet the text content condition; if the first text line does not exist in each of the text lines, determining that the text content detection result indicates that the area to be identified meets the text content condition;

[0302] Perform text type element detection on each of the text lines to determine whether each of the text lines is a second text line containing a text type element; if a second text line exists in each of the text lines, determine that the text type element detection result is that the area to be identified does not meet the text type element condition; if no second text line exists in each of the text lines, determine that the text type element detection result is that the area to be identified meets the text type element condition.

[0303] Optionally, the text line type detection includes writing direction detection, left boundary detection and font style detection, the text line type condition includes a writing direction condition, a left boundary condition and a font style condition, the writing direction condition is that the writing directions of the text lines are the same, the left boundary condition is that when the area to be identified contains a text line, the coordinates of the left boundary of the text line are less than or equal to the coordinates of the right boundary of the text line above the text line, the font style condition is that there is no first target text line pair in each text line pair, the text line pair is two connected text lines in each text line, and the font thickness ranges between the text lines in the first target text line pair do not overlap, and the text line type detection result includes a writing direction detection result, a left boundary detection result and a font style detection result;

[0304] The first detection module 302 is further configured to:

[0305] Performing writing direction detection on each of the text lines to determine the writing direction of each of the text lines; if the writing directions are the same, determining that the writing direction detection result indicates that the area to be identified meets the writing direction condition; if the writing directions are different, determining that the writing direction detection result indicates that the area to be identified does not meet the writing direction condition;

[0306] In the case where the area to be identified includes a text line, performing left boundary detection on the text line to determine whether the coordinates of the left boundary of the text line are greater than the coordinates of the right boundary of the text line above the text line; if not, determining that the left boundary detection result indicates that the area to be identified meets the left boundary condition; if so, determining that the left boundary detection result indicates that the area to be identified does not meet the left boundary condition;

[0307] Perform font style detection on each of the text line pairs to determine whether there is an intersection in the font thickness ranges between the text lines in each of the text line pairs; if so, determine that the font style detection result indicates that the area to be identified meets the font style conditions; if not, determine that the font style detection result indicates that the area to be identified does not meet the font style conditions.

[0308] Optionally, the color alignment condition is that the color ranges of two adjacent text lines intersect;

[0309] The first detection module 302 is further configured to:

[0310] For each of the text lines, identifying the color corresponding to each text object in the text line under the target color mode, and obtaining the color range corresponding to each text object in the text line;

[0311] Determining whether each text line pair is a second target text line pair, wherein the text line pair is two connected text lines in each of the text lines, and the color ranges corresponding to the two text lines in the second target text line pair have an intersection;

[0312] If so, determining that the color alignment detection result is that the to-be-identified area meets the color alignment condition;

[0313] If not, it is determined that the color alignment detection result is that the to-be-identified area does not meet the color alignment condition.

[0314] Optionally, the second determining module 305 is further configured to:

[0315] If the association detection result, the number modification detection result, and the number alignment detection result all indicate that a condition is met, determining that the area to be identified is a number modification area;

[0316] If any one of the association detection result, the number modification detection result, and the number alignment detection result does not meet the condition, it is determined that the area to be identified is not a number modification area.

[0317] Optionally, the apparatus 300 for identifying number-modified regions in a document further includes an acquisition module configured to:

[0318] Obtain the number of text lines contained in the area to be identified;

[0319] When the number of rows is not equal to the third value, subtracting the fourth value from the number of rows to obtain an updated number of rows;

[0320] Determine each target sub-region in the to-be-recognized region, wherein the target sub-region contains the updated number of connected text lines;

[0321] The target sub-regions are respectively used as updated regions to be identified, and the step of performing text line association detection on each text line is continued to obtain an association detection result of whether the region to be identified meets a preset line association condition.

[0322] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call logic instructions in the memory 430 to execute a method for identifying a number-modified region in a document, the method comprising: determining a region to be identified in a target document, the region to be identified comprising at least one text line; performing a text line association test on each of the text lines to obtain an association test result indicating whether the region to be identified meets a preset line association condition; performing a number modification test on each of the text lines to obtain a number modification test result indicating whether the region to be identified meets a preset number modification condition; performing a number alignment test on each of the text lines to obtain a number alignment test result indicating whether the region to be identified meets a preset number alignment condition; and determining whether the region to be identified is a number-modified region based on the association test result, the number modification test result, and the number alignment test result.

[0323] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0324] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for identifying number-modified areas in documents provided by the above methods, the method including: determining an area to be identified in a target document, the area to be identified containing at least one text line; performing text line association detection on each of the text lines to obtain an association detection result of whether the area to be identified meets a preset line association condition; performing number modification detection on each of the text lines to obtain a number modification detection result of whether the area to be identified meets the preset number modification condition; performing number alignment detection on each of the text lines to obtain a number alignment detection result of whether the area to be identified meets the preset number alignment condition; and determining whether the area to be identified is a number-modified area based on the association detection result, the number modification detection result and the number alignment detection result.

[0325] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the method for identifying number-modified areas in documents provided by the above-mentioned methods, the method comprising: determining an area to be identified in a target document, the area to be identified containing at least one text line; performing a text line association detection on each of the text lines to obtain an association detection result of whether the area to be identified meets a preset line association condition; performing a number modification detection on each of the text lines to obtain a number modification detection result of whether the area to be identified meets the preset number modification condition; performing a number alignment detection on each of the text lines to obtain a number alignment detection result of whether the area to be identified meets the preset number alignment condition; and determining whether the area to be identified is a number-modified area based on the association detection result, the number modification detection result and the number alignment detection result.

[0326] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0327] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0328] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for identifying numbered modified areas in a document, characterized in that: include: Determine a region to be recognized in a target document, wherein the region to be recognized includes at least one text line; Performing a text line association test on each of the text lines to obtain an association test result indicating whether the area to be identified meets a preset line association condition; Performing a number modification detection on each of the text lines to obtain a number modification detection result indicating whether the area to be identified meets a preset number modification condition; Performing a number alignment test on each of the text lines to obtain a number alignment test result indicating whether the area to be identified meets a preset number alignment condition; It is determined whether the to-be-identified region is a number-modified region according to the association detection result, the number-modification detection result, and the number-alignment detection result.

2. The method for identifying numbered modified areas in a document according to claim 1, wherein: The number modification detection includes first character detection and / or number item detection, the preset number modification condition includes first character condition and / or number item condition, and the number modification detection result includes first character detection result and / or number item detection result; The performing of number modification detection on each of the text lines to obtain a number modification detection result indicating whether the to-be-identified area meets a preset number modification condition includes: Performing a first character detection on each of the text lines to obtain a first character detection result indicating whether the area to be identified meets the first character condition; and / or, A numbering item detection is performed on each of the text lines to obtain a numbering item detection result indicating whether the area to be identified meets the numbering item condition.

3. The method for identifying numbered modified areas in a document according to claim 2, wherein: The first character condition is that the first characters of each text line belong to the same numbered character set; The first character detection is performed on each of the text lines to obtain a first character detection result indicating whether the area to be identified meets the first character condition, including: Obtaining the first character of each text line and searching for the Unicode of each first character; Determining whether each of the first characters belongs to the same numbered character set according to the unified character encoding of each of the first characters; If so, determining that the first character detection result is that the area to be recognized meets the first character condition; If not, it is determined that the first character detection result is that the area to be recognized does not meet the first character condition.

4. The method for identifying numbered modified areas in a document according to claim 2 or 3, wherein: The numbered item condition is that each of the text lines has a numbered item; The performing of numbered item detection on each of the text lines to obtain a numbered item detection result indicating whether the area to be identified meets the numbered item condition includes: For each of the text lines, determining whether a numbered item exists in the text line based on characters between the first character and a set character in the text line, where the set character is the character preceding the first specified character, and the specified character is not a numeric character or a separator; If a numbered item exists in each of the text lines, determining the numbered item detection result is that the area to be identified meets the numbered item condition; If a specific text line exists in each of the text lines, it is determined that the numbered item detection result is that the area to be identified does not meet the numbered item condition, and the specific text line is a text line in which no numbered item exists.

5. The method for identifying numbered modified areas in a document according to claim 4, wherein: The determining whether a numbered item exists in the text line according to characters between the first character and the set character in the text line includes: Recognize the first character of the text line to obtain a recognition result; and determine whether the text line meets a first condition based on the recognition result, wherein the first condition indicates that the first character is a numeric character; Performing a delimiter search on characters between the first character and the set character in the text line to obtain a search result; and determining whether the text line meets a second condition based on the search result, the second condition indicating that a delimiter exists between the first character and the set character; Obtaining, from characters between a second character and a set character in the text line, a first number of characters of at least one numeric character combination, the numeric character combination being at least one consecutive numeric character, the first number of characters being the number of numeric characters in the numeric character combination; and determining, based on each first number of characters, whether the text line meets a third condition, the third condition indicating that each first number of characters is less than or equal to a first value; Obtaining, from characters between a second character and a set character in the text line, a second number of characters of at least one separator combination, the separator combination being at least one consecutive separator, the second number of characters being the number of separators in the separator combination; and determining whether the text line meets a fourth condition based on each second number of characters, the fourth condition being that each second number of characters is less than or equal to a second value; If the text line satisfies the first condition, the second condition, the third condition, and the fourth condition at the same time, determining that a numbered item exists in the text line; If the text line does not meet at least one of the first condition, the second condition, the third condition, and the fourth condition, it is determined that no numbered item exists in the text line.

6. The method for identifying numbered modified areas in a document according to claim 5, wherein: After determining that the text line has a numbered item, the method further includes: The numbering level of the numbered items in the text line is determined based on the first combination quantity of the numeric character combinations and the second combination quantity of the separator combinations in the text line, the first combination quantity being the number of the numeric character combinations contained in the text line, and the second combination quantity being the number of the separator combinations contained in the text line.

7. The method for identifying numbered modified areas in a document according to claim 1, wherein: The performing of number alignment detection on each of the text lines to obtain a number alignment detection result indicating whether the area to be identified meets a preset number alignment condition includes: Calculating an average font size corresponding to the area to be identified based on the font size of each character in all the text lines; Constructing a mapping table, wherein the mapping table records the first left boundary value corresponding to each numbering level; For each of the text lines, obtaining a numbering level corresponding to the text line and a second left boundary value of a circumscribed rectangular box; searching the mapping table for a first left boundary value corresponding to the text line based on the numbering level corresponding to the text line; determining an absolute value of a difference between the first left boundary value and the second left boundary value corresponding to the text line; and determining that the text line is a set text line if the absolute value is greater than the average font size; In a case where the set text line exists in each of the text lines, determining that the number alignment detection result is that the to-be-identified area does not meet the preset number alignment condition; In a case where the set text line does not exist in each of the text lines, it is determined that the number alignment detection result is that the to-be-identified area meets the preset number alignment condition.

8. The method for identifying numbered modified areas in a document according to claim 1, wherein: The text line association detection includes content type detection, text line type detection and color alignment detection, the preset line association conditions include content type conditions, text line type conditions and color alignment conditions, and the association detection results include content type detection results, text line type detection results and color alignment detection results; The performing text line association detection on each of the text lines to obtain an association detection result of whether the area to be identified meets a preset line association condition includes: Performing content type detection on each of the text lines to obtain a content type detection result of whether the area to be identified meets the content type condition; Performing text line type detection on each of the text lines to obtain a text line type detection result indicating whether the area to be identified meets the text line type condition; A color alignment test is performed on each of the text lines to obtain a color alignment test result indicating whether the area to be identified meets the color alignment condition.

9. The method for identifying numbered modified areas in a document according to claim 8, wherein: The content type detection includes text content detection and text type element detection, the content type condition includes a text content condition and a text type element condition, the text content condition is that each of the text lines does not contain a first text line without text content, the text type element condition is that each of the text lines does not contain a second text line without a text type element, and the content type detection result includes a text content detection result and a text type element result; The performing of content type detection on each of the text lines to obtain a content type detection result of whether the area to be identified meets the content type condition includes: Performing text content detection on each of the text lines to determine whether each of the text lines is a first text line without text content; if the first text line exists in each of the text lines, determining that the text content detection result indicates that the area to be identified does not meet the text content condition; if the first text line does not exist in each of the text lines, determining that the text content detection result indicates that the area to be identified meets the text content condition; Perform text type element detection on each of the text lines to determine whether each of the text lines is a second text line containing a text type element; if a second text line exists in each of the text lines, determine that the text type element detection result is that the area to be identified does not meet the text type element condition; if no second text line exists in each of the text lines, determine that the text type element detection result is that the area to be identified meets the text type element condition.

10. The method for identifying numbered modified areas in a document according to claim 8, wherein: The text line type detection includes writing direction detection, left boundary detection and font style detection, the text line type condition includes writing direction condition, left boundary condition and font style condition, the writing direction condition is that the writing directions of the text lines are the same, the left boundary condition is that when the area to be identified contains a text line, the coordinates of the left boundary of the text line are less than or equal to the coordinates of the right boundary of the text line above the text line, the font style condition is that there is no first target text line pair in each text line pair, the text line pair is two connected text lines in each text line, and the font thickness ranges between the text lines in the first target text line pair do not overlap, and the text line type detection result includes the writing direction detection result, the left boundary detection result and the font style detection result; The performing text line type detection on each of the text lines to obtain a text line type detection result indicating whether the area to be identified meets the text line type condition includes: Performing writing direction detection on each of the text lines to determine the writing direction of each of the text lines; if the writing directions are the same, determining that the writing direction detection result indicates that the area to be identified meets the writing direction condition; if the writing directions are different, determining that the writing direction detection result indicates that the area to be identified does not meet the writing direction condition; In the case where the area to be identified includes a text line, performing left boundary detection on the text line to determine whether the coordinates of the left boundary of the text line are greater than the coordinates of the right boundary of the text line above the text line; if not, determining that the left boundary detection result indicates that the area to be identified meets the left boundary condition; if so, determining that the left boundary detection result indicates that the area to be identified does not meet the left boundary condition; Perform font style detection on each of the text line pairs to determine whether there is an intersection in the font thickness ranges between the text lines in each of the text line pairs; if so, determine that the font style detection result indicates that the area to be identified meets the font style conditions; if not, determine that the font style detection result indicates that the area to be identified does not meet the font style conditions.

11. The method for identifying numbered modified areas in a document according to claim 8, wherein: The color alignment condition is that the color ranges of two adjacent text lines intersect; The performing color alignment detection on each of the text lines to obtain a color alignment detection result of whether the to-be-identified area meets the color alignment condition includes: For each of the text lines, identifying the color corresponding to each text object in the text line under the target color mode, and obtaining the color range corresponding to each text object in the text line; Determining whether each text line pair is a second target text line pair, wherein the text line pair is two connected text lines in each of the text lines, and the color ranges corresponding to the two text lines in the second target text line pair have an intersection; If so, determining that the color alignment detection result is that the to-be-identified area meets the color alignment condition; If not, it is determined that the color alignment detection result is that the to-be-identified area does not meet the color alignment condition.

12. The method for identifying numbered modified areas in a document according to claim 1, wherein: The determining, based on the association detection result, the number modification detection result, and the number alignment detection result, whether the to-be-identified area is a number modification area includes: If the association detection result, the number modification detection result, and the number alignment detection result all indicate that a condition is met, determining that the area to be identified is a number modification area; If any one of the association detection result, the number modification detection result, and the number alignment detection result does not meet the condition, it is determined that the area to be identified is not a number modification area.

13. The method for identifying numbered modified areas in a document according to claim 12, wherein: After determining that the area to be identified is not a number-modified area, the method further includes: Obtain the number of text lines contained in the area to be identified; When the number of rows is not equal to the third value, subtracting the fourth value from the number of rows to obtain an updated number of rows; Determine each target sub-region in the to-be-recognized region, wherein the target sub-region contains the updated number of connected text lines; The target sub-regions are respectively used as updated regions to be identified, and the step of performing text line association detection on each text line is continued to obtain an association detection result of whether the region to be identified meets a preset line association condition.

14. A device for identifying numbered modified areas in a document, characterized in that: include: A first determining module is configured to determine a region to be identified in a target document, wherein the region to be identified includes at least one text line; A first detection module is configured to perform text line association detection on each of the text lines to obtain an association detection result indicating whether the area to be identified meets a preset line association condition; A second detection module is configured to perform a number modification detection on each of the text lines to obtain a number modification detection result indicating whether the area to be identified meets a preset number modification condition; A third detection module is configured to perform a number alignment detection on each of the text lines to obtain a number alignment detection result indicating whether the area to be identified meets a preset number alignment condition; The second determining module is configured to determine whether the to-be-identified area is a number-modified area according to the association detection result, the number-modification detection result, and the number-alignment detection result.