Method and device for detecting first character sinking paragraph
By extracting non-modifying target elements and counting their types and heights when text line features meet certain conditions, the problem of detecting drop-cap paragraphs in documents is solved, and the accuracy of document format conversion and rearrangement is improved.
Patent Information
- Application Number
- CN202410253306.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-09-09
AI Technical Summary
It is difficult to accurately detect drop cap paragraphs in documents with existing technologies, which affects the accuracy of document format conversion and rearrangement.
By extracting a set number of non-modifying target elements when the text line features of the paragraph to be detected meet the preset conditions, counting the element type features, and determining the first-letter drop paragraph when the element type and height meet the preset conditions.
Improves the detection accuracy of drop cap paragraphs, helping to make document conversion and re-layout more accurate.
Smart Images

Figure CN120611698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document recognition, and in particular to a method and device for detecting a paragraph with a drop cap. Background Art
[0002] A drop cap is a decorative typographic element, typically used at the beginning of a capitalized word or paragraph, to enhance the readability and visual appeal of text.
[0003] Whether it is a flow document, a layout document, or an image document, all can be typed using the drop cap method. Therefore, how to detect the drop cap paragraphs in the document has become an urgent problem to be solved. Summary of the Invention
[0004] In view of the problems existing in the prior art, an embodiment of the present invention provides a method and device for detecting a paragraph with a dropped first character.
[0005] The present invention provides a method for detecting a paragraph with a drop cap, comprising:
[0006] When the text line feature of the paragraph to be detected meets the preset text line condition, extracting a set number of non-modifying target elements from the paragraph to be detected;
[0007] Performing element type statistics on each of the target elements to obtain element type characteristics;
[0008] When the element type feature meets the preset element condition, obtaining the first height of the first element in the paragraph to be detected;
[0009] When the first height meets a preset height condition, the paragraph to be detected is determined to be a first-character drop-head paragraph.
[0010] According to a method for detecting a drop cap paragraph provided by the present invention, extracting a set number of non-modifying target elements from the paragraph to be detected includes:
[0011] Extracting each non-modifying initial element in the paragraph to be detected;
[0012] Obtaining a first circumscribed rectangular frame of each of the initial elements and a second circumscribed rectangular frame of the paragraph to be detected;
[0013] Determining the distance between each of the first circumscribed rectangular frames and the second circumscribed rectangular frames;
[0014] According to each of the distances, a set number of target elements are selected from each of the initial elements.
[0015] According to a method for detecting a drop cap paragraph provided by the present invention, performing element type statistics on each target element to obtain element type features includes:
[0016] Counting a first number of path types in each of the target elements, and counting a second number of picture types in each of the target elements;
[0017] For each text element in the target element, count a third number of characters, a fourth number of punctuation marks, a fifth number of numbers, and a sixth number of other types;
[0018] The first quantity, the second quantity, the third quantity, the fourth quantity, the fifth quantity and the sixth quantity are determined as the element type characteristics.
[0019] According to a method for detecting a drop cap paragraph provided by the present invention, the preset element conditions include a text type condition and a comprehensive type condition, wherein the comprehensive type condition represents conditions corresponding to a path type and an image type;
[0020] Before obtaining the first height of the first element in the paragraph to be detected when the element type feature meets the preset element condition, the method further includes:
[0021] Comparing the element type feature with the text type condition and the comprehensive type condition respectively to obtain a text comparison result and a comprehensive comparison result;
[0022] If both the text comparison result and the comprehensive comparison result indicate a mismatch, determining that the element type feature does not meet the preset element condition;
[0023] If there is a comparison result indicating compliance between the text comparison result and the comprehensive comparison result, it is determined that the element type feature complies with the preset element condition.
[0024] According to a method for detecting a drop cap paragraph provided by the present invention, the element type feature includes a first quantity, a second quantity, a third quantity, a fourth quantity, a fifth quantity, and a sixth quantity;
[0025] The text type condition includes that the first quantity, the second quantity, and the sixth quantity are all first preset values, the first element is a character type or a punctuation type, and the fourth quantity is less than a first set value or the ratio of the fourth quantity to the total quantity is less than a second set value, and the total quantity is the sum of the third quantity, the fourth quantity, and the sixth quantity;
[0026] The comprehensive type condition includes that the sum of the first quantity and the second quantity is a second preset value, and the sixth quantity is a first preset value.
[0027] According to a method for detecting a drop cap paragraph provided by the present invention, obtaining a first height of a first element in the paragraph to be detected includes:
[0028] Obtain the third circumscribed rectangular frame of the first element in the paragraph to be detected;
[0029] The height of the third circumscribed rectangular frame is used as the first height of the first element.
[0030] According to a method for detecting a drop cap paragraph provided by the present invention, before determining that the paragraph to be detected is a drop cap paragraph when the first height meets a preset height condition, the method further includes:
[0031] Obtaining a fourth circumscribed rectangular frame of each designated element in the first text line of the paragraph to be detected, wherein the designated elements are elements other than the first element in the first text line;
[0032] Using the height of the union of the fourth circumscribed rectangular frames as the second height of the first text line;
[0033] For each designated text line in the paragraph to be detected, obtaining a fifth circumscribed rectangular box of each element in the designated text line, wherein the designated text line is a text line in the paragraph to be detected except the first text line; and using the height of the union of the fifth circumscribed rectangular boxes as the third height of the designated text line;
[0034] A preset height condition is constructed according to the first height, the second height and each of the third heights.
[0035] According to a method for detecting a drop cap paragraph provided by the present invention, the preset height condition includes:
[0036] The first height is greater than a product of a first value and the second height, and a product of each of the third heights and the second value is less than the first height.
[0037] According to a method for detecting a drop cap paragraph provided by the present invention, before extracting a set number of non-modifying target elements from the paragraph to be detected when the text line features of the paragraph to be detected meet the preset text line conditions, the method further includes:
[0038] Obtain the number of text lines contained in the paragraph to be detected;
[0039] Identifying the writing direction of each text line in the paragraph to be detected;
[0040] Determining whether each text line in the paragraph to be detected has a first text feature of the text content;
[0041] Determining whether a second text feature of a jump target anchor exists in each text line in the paragraph to be detected;
[0042] Determining whether a third text feature of a text type element exists in each text line in the paragraph to be detected;
[0043] The number of text lines, the writing direction, the first text feature, the second text feature, and the third text feature are determined as the text line feature.
[0044] According to a method for detecting a drop cap paragraph provided by the present invention, the preset text line conditions include:
[0045] The number of text lines is greater than a third value, the writing directions are all set writing directions, the first text feature is that there is text content in each text line, the second text feature is that there is no jump target anchor in each text line, and the third text feature is that there is a text type element in each text line.
[0046] The present invention also provides a device for detecting a paragraph with a dropped first character, comprising:
[0047] an extraction module configured to extract a set number of non-modifying target elements from the paragraph to be detected when the text line features of the paragraph to be detected meet the preset text line conditions;
[0048] a statistics module configured to perform element type statistics on each target element to obtain element type characteristics;
[0049] An acquisition module configured to acquire a first height of the first element in the paragraph to be detected if the element type feature meets a preset element condition;
[0050] The determination module is configured to determine that the paragraph to be detected is a drop cap paragraph and the first element is a drop cap element if the first height meets a preset height condition.
[0051] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for detecting a paragraph with a dropped first character as described above is implemented.
[0052] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for detecting a paragraph with a drop cap as described above is implemented.
[0053] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned methods for detecting a drop cap paragraph.
[0054] The present invention provides a method and device for detecting a drop cap paragraph. When the text line characteristics of the paragraph to be detected meet the preset text line conditions, a set number of non-modifying target elements are extracted from the paragraph to be detected; element type statistics are performed on each of the target elements to obtain element type characteristics; when the element type characteristics meet the preset element conditions, the first height of the first element in the paragraph to be detected is obtained; when the first height meets the preset height conditions, the paragraph to be detected is determined to be a drop cap paragraph. The present invention can identify whether there is a drop cap element in a paragraph and the corresponding specific element modified, that is, detect drop cap paragraphs and drop cap elements in a document; it is conducive to more accurate analysis of paragraph layout, helping subsequent format document conversion, re-layout, etc. to be more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0056] Figure 1 This is one of the flow charts of the method for detecting a drop cap paragraph provided by the present invention;
[0057] Figure 2 Schematic diagram of the drop cap paragraph provided by the present invention;
[0058] Figure 3 This is the second flow chart of the method for detecting a drop cap paragraph provided by the present invention;
[0059] Figure 4 1 is a schematic structural diagram of a device for detecting a drop cap paragraph provided by the present invention;
[0060] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0061] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0062] In order to facilitate a clearer understanding of the embodiments of the present invention, some relevant knowledge is first introduced as follows.
[0063] A drop cap is a larger letter, often larger than normal text, extending across multiple lines. It can help separate paragraphs and sections, making text easier to read, especially in longer articles. This typographic element is frequently used in a variety of publications, including printed magazines, books, and religious texts.
[0064] The basic element, namely the text element (Text), includes the text's Unicode, character code (CharCode), font information, color, and Bounding Box (BBox) coordinates.
[0065] Path elements (Path elements or vector elements) usually contain drawing instructions, colors, BBox coordinates, etc.
[0066] Image element, which contains RGB color information and BBox coordinates, etc.
[0067] Writing mode is used to describe the direction of text layout, including LrTb (row direction is from left to right, paragraph direction is from top to bottom, adopted by most writing systems), RlTb (row direction is from right to left, paragraph direction is from top to bottom, usually used in Arabic / Hebrew writing systems), and TbRl (row direction is from top to bottom, paragraph direction is from right to left, usually used in writing systems such as ancient Chinese, Japanese or Mongolian).
[0068] The present invention provides a method and device for detecting a drop cap paragraph. When the text line characteristics of the paragraph to be detected meet the preset text line conditions, a set number of non-modifying target elements are extracted from the paragraph to be detected; element type statistics are performed on each of the target elements to obtain element type characteristics; when the element type characteristics meet the preset element conditions, the first height of the first element in the paragraph to be detected is obtained; when the first height meets the preset height conditions, the paragraph to be detected is determined to be a drop cap paragraph. The present invention can identify whether there is a drop cap element in a paragraph and the corresponding specific element modified, that is, detect drop cap paragraphs and drop cap elements in a document; it is conducive to more accurate analysis of paragraph layout, helping subsequent format document conversion, re-layout, etc. to be more accurate.
[0069] The following combination Figure 1-Figure 5 The present invention describes a method and device for detecting a drop cap paragraph.
[0070] Figure 1 This is one of the flow charts of the method for detecting the first-letter drop paragraph provided by the present invention, see Figure 1 As shown, it includes steps 101 to 104, wherein:
[0071] Step 101: When the text line features of the paragraph to be detected meet the preset text line conditions, a set number of non-modifying target elements are extracted from the paragraph to be detected.
[0072] First, it should be noted that the present invention can be implemented by any electronic device capable of detecting drop cap paragraphs, such as a smartphone, smartwatch, desktop computer, laptop computer, etc. The present invention can be applied to scenarios requiring drop cap paragraph detection, such as layout document conversion, layout document reformatting, and layout document editing.
[0073] Specifically, if Figure 2 The schematic diagram of the drop cap paragraph provided by the present invention is shown: a drop cap paragraph refers to a paragraph that is typeset using a drop cap; in a drop cap paragraph, the first character (such as a letter, text, etc.) at the beginning of the paragraph, that is, the drop cap element, is in large size, and the layout style of the paragraph is changed by sinking or hanging, occupying the first few lines of the paragraph.
[0074] The paragraph to be detected refers to the paragraph that needs to be detected whether the first word drop is used, and can be any paragraph in the target document, where the target document can be a layout document, such as a PDF (Portable Document Format) document, an OFD (Open Fixed-layout Document) document, a streaming document, such as a Word document, a TXT document, a presentation document, such as a slide document, or a picture document, and the present invention does not impose any limitation on this.
[0075] The text line features refer to features that characterize the text lines in the paragraph to be detected and / or features corresponding to the text lines.
[0076] The preset text line condition refers to a pre-set condition used to measure whether the text line characteristics meet the requirements, and can be used to preliminarily detect whether the paragraph to be detected is a drop-first-word paragraph.
[0077] The set number refers to a pre-set value used to limit the number of extracted target elements. For example, the set number is a natural number not less than 5.
[0078] Non-descriptive representations are not descriptive. For example, TextObject text elements, PathObject path elements, and ImageObject image elements are not descriptive. Descriptive basic elements may include underline, strikethrough, etc.
[0079] In practical applications, when the text line features of the paragraph to be detected meet the preset text line conditions, the paragraph to be detected may be determined before extracting a set number of non-modifying target elements from the paragraph to be detected.
[0080] First, the target document is obtained. For example, if a user uploads the target document through the document upload page, the execution entity will obtain the target document accordingly. Alternatively, if the execution entity receives a document acquisition instruction, the execution entity will obtain the target document from the storage area pointed to by the document acquisition instruction.
[0081] Furthermore, the paragraphs to be detected can be determined from the target document. For example, in the scenario of detecting drop cap paragraphs in the target document, each paragraph in the target document can be used as the paragraph to be detected; for example, in the scenario of detecting whether a specified paragraph is a drop cap paragraph, a paragraph selection instruction can be received, and the paragraph selected by the paragraph selection instruction in the target document can be used as the paragraph to be detected.
[0082] After determining the paragraph to be detected, the text line features of the paragraph to be detected are further extracted and a preset text line condition is obtained. The text line features are compared with the preset text line condition. If the text line features do not meet the preset text line condition, the paragraph to be detected is determined not to be a drop cap paragraph. If the text line features meet the preset text line condition, it is necessary to further determine whether the paragraph to be detected is a drop cap paragraph, that is, it is necessary to extract a set number of non-modifying target elements from the paragraph to be detected.
[0083] Specifically, a set number of non-modifying target elements are extracted from the paragraph to be detected. The modifying basic elements in the paragraph to be detected can be filtered to leave the non-modifying basic elements, and then a set number of basic elements are extracted from the non-modifying basic elements as target elements.
[0084] It should be noted that a line recognizer can be used to identify text lines and paragraphs. If the target document is a layout document, the determination is based on the position of the TextObject in the document. If the target document is an image document, text feature extraction is performed using technologies such as OCR (Optical Character Recognition).
[0085] Step 102: performing element type statistics on each target element to obtain element type characteristics.
[0086] Specifically, the element type refers to the type corresponding to the basic element. The element type may include path type (the type corresponding to the path element), image type (the type corresponding to the image element) and text type (the type corresponding to the text element), wherein the text type includes character type, punctuation type, number type and other types.
[0087] In practical applications, on the basis of obtaining a set number of target elements, the number of each element type in the target elements is further counted to obtain the element type characteristics.
[0088] Step 103: When the element type feature meets the preset element condition, obtain the first height of the first element in the paragraph to be detected.
[0089] Specifically, the preset element condition refers to a pre-set condition used to measure whether the element type characteristics meet the requirements. It can further detect whether the paragraph to be detected is a drop cap paragraph. The first element refers to the first basic element in the paragraph to be detected. The first height is used to represent the height of the first element.
[0090] In actual applications, on the basis of obtaining the element type features, the element type features are further compared with the preset element conditions. If the preset element features do not meet the preset element conditions, it is determined that the paragraph to be detected is not a first-drop paragraph; if the element type features meet the preset element conditions, it is necessary to further determine whether the paragraph to be detected is a first-drop paragraph, that is, it is necessary to obtain the first height of the first element in the paragraph to be detected.
[0091] Specifically, to obtain the first height of the first element in the paragraph to be detected, the first element in the paragraph to be detected may be first determined, and then the first height may be determined based on the bounding rectangle or font size of the first element.
[0092] Step 104: When the first height meets a preset height condition, determining that the paragraph to be detected is a first-character drop-head paragraph.
[0093] Specifically, the preset height condition refers to a pre-set condition for measuring whether the first height meets the requirements, and can ultimately detect whether the paragraph to be detected is a drop cap paragraph. The drop cap element refers to the basic element of the drop cap.
[0094] In actual application, based on the first height obtained, the first height is further compared with the preset height condition. If the first height feature does not meet the preset element condition, it is determined that the paragraph to be detected is not a drop-cap paragraph; if the first height meets the preset height condition, it is determined that the paragraph to be detected is a drop-cap paragraph, and the first element is the drop-cap element of the drop-cap paragraph. Figure 2 As shown, the paragraph is a drop cap paragraph, and the first element "F" is a drop cap element.
[0095] The present invention provides a method for detecting a drop cap paragraph. When the text line feature of the paragraph to be detected meets a preset text line condition, a set number of non-modifying target elements are extracted from the paragraph to be detected; element type statistics are performed on each of the target elements to obtain an element type feature; when the element type feature meets the preset element condition, the first height of the first element in the paragraph to be detected is obtained; when the first height meets the preset height condition, the paragraph to be detected is determined to be a drop cap paragraph. The present invention can identify whether there is a drop cap element in a paragraph and the corresponding specific element modified, that is, detect drop cap paragraphs and drop cap elements in a document; it is conducive to more accurate analysis of paragraph layout, helping subsequent format document conversion, rearrangement, etc. to be more accurate.
[0096] In one or more optional embodiments of the present invention, the step of extracting a set number of non-modifying target elements from the paragraph to be detected may be specifically implemented as follows:
[0097] Extracting each non-modifying initial element in the paragraph to be detected;
[0098] Obtaining a first circumscribed rectangular frame of each of the initial elements and a second circumscribed rectangular frame of the paragraph to be detected;
[0099] Determining the distance between each of the first circumscribed rectangular frames and the second circumscribed rectangular frames;
[0100] According to each of the distances, a set number of target elements are selected from each of the initial elements.
[0101] Specifically, the first circumscribed rectangular frame refers to the circumscribed rectangular frame of the initial element; the second circumscribed rectangular frame refers to the circumscribed rectangular frame of the paragraph to be detected.
[0102] In practical applications, the modifying basic elements in the paragraph to be detected can be filtered to leave the non-modifying basic elements, and each non-modifying basic element can be used as the initial element. Then, the first bounding rectangle of each initial element and the second bounding rectangle of the paragraph to be detected can be obtained.
[0103] Furthermore, based on the coordinates of each first circumscribed rectangular frame and the coordinates of the second circumscribed rectangular frame, the distance between each first circumscribed rectangular frame and the second circumscribed rectangular frame is calculated. For example, the distance between the upper left corner of each first circumscribed rectangular frame and the upper left corner of the second circumscribed rectangular frame can be calculated, and the distance between the center point of each first circumscribed rectangular frame and the upper left corner of the second circumscribed rectangular frame can be calculated.
[0104] Then, based on the distances, a set number of target elements are selected from the initial elements. For example, the initial elements may be arranged in ascending order of distance, and the set number of initial elements with the closest distances are selected as target elements.
[0105] In this way, by screening target elements based on the distances between each first circumscribed rectangular frame and the second circumscribed rectangular frame, the accuracy and reliability of the selected target elements can be guaranteed, and the accuracy of detecting the first-letter drop-head paragraph can be further improved.
[0106] It should be noted that obtaining the second bounding rectangle of the paragraph to be detected can be specifically implemented by obtaining the bounding rectangles of each basic element in the paragraph to be detected; and determining the union of the bounding rectangles of each basic element as the second bounding rectangle of the paragraph to be detected. This can improve the efficiency and accuracy of obtaining the second bounding rectangle.
[0107] In one or more optional embodiments of the present invention, extracting a set number of non-modifying target elements from the paragraph to be detected may be specifically implemented by extracting each non-modifying initial element in the paragraph to be detected; and randomly extracting a set number of target elements from each of the initial elements. This can improve the efficiency of obtaining target elements.
[0108] In one or more optional embodiments of the present invention, performing element type statistics on each target element to obtain element type characteristics may be specifically implemented as follows:
[0109] Counting a first number of path types in each of the target elements, and counting a second number of picture types in each of the target elements;
[0110] For each text element in the target element, count a third number of characters, a fourth number of punctuation marks, a fifth number of numbers, and a sixth number of other types;
[0111] The first quantity, the second quantity, the third quantity, the fourth quantity, the fifth quantity and the sixth quantity are determined as the element type characteristics.
[0112] Specifically, the path type refers to the type corresponding to the path element, for example, vector objects (SVG, Scalable Vector Graphics), flowcharts, and arrows are all path types; the image type refers to the type corresponding to the image element; the text element refers to the basic element whose element type is the text type, and the text type includes character (Letter) type, punctuation (Punctuation) type, number (Number) type, and other (Other) types.
[0113] In actual applications, the first number of path types in each target element (PathCount) is counted, that is, PathCount is recorded as the number of Path type target elements; the second number of image types in each target element (ImageCount) is counted, that is, ImageCount is recorded as the number of Image type target elements.
[0114] For the text elements in each target element, the third number (L) of the character type, the fourth number (P) of the punctuation type, the fifth number (N) of the number type, and the sixth number (O) of the other types are counted. That is, for all target elements of the Text type, according to the Unicode character category classification standard, the number of text elements of the character type, punctuation type, number, and other types are counted, and recorded as L, P, N, and O respectively.
[0115] Among them, the first quantity, the second quantity, the third quantity, the fourth quantity, the fifth quantity and the sixth quantity constitute the element type feature, that is, the first quantity, the second quantity, the third quantity, the fourth quantity, the fifth quantity and the sixth quantity are determined as the element type feature.
[0116] In this way, by counting the first number, the second number, the third number, the fourth number, the fifth number and the sixth number, the element type feature can be made more fine-grained and more accurate, which can further improve the accuracy of detecting the first-letter drop paragraph.
[0117] In one or more optional embodiments of the present invention, the preset element condition includes a text type condition and a comprehensive type condition, the comprehensive type condition representing conditions corresponding to a path type and an image type; and when the element type feature meets the preset element condition, before obtaining the first height of the first element in the paragraph to be detected, the method further includes:
[0118] Comparing the element type feature with the text type condition and the comprehensive type condition respectively to obtain two text comparison results and a comprehensive comparison result;
[0119] If both the text comparison result and the comprehensive comparison result indicate a mismatch, determining that the element type feature does not meet the preset element condition;
[0120] If there is a comparison result indicating compliance between the text comparison result and the comprehensive comparison result, it is determined that the element type feature complies with the preset element condition.
[0121] Specifically, the text type condition refers to a pre-set condition used to measure whether the element type characteristics meet the text type requirements; the comprehensive type condition refers to a pre-set condition used to measure whether the element type characteristics meet the path type and image type requirements; the text comparison result refers to the result of comparing the element type characteristics with the text type condition; the comprehensive comparison result refers to the result of comparing the element type characteristics with the comprehensive type condition.
[0122] In practical applications, the element type feature can be compared with the text type condition and the comprehensive type condition respectively. If the element type feature does not meet the text type condition, nor does it meet the comprehensive type condition, that is, both the text comparison result and the comprehensive comparison result indicate non-compliance, then it means that the element type feature does not meet the preset element condition; if the element type feature meets at least one of the text type condition and the comprehensive type condition, that is, there is a comparison result that indicates compliance in the text comparison result and the comprehensive comparison result, then it means that the element type feature meets the preset element condition. In this way, the accuracy of determining whether the element type feature meets the preset element condition can be improved, thereby improving the accuracy of detecting paragraphs with the first word dropped.
[0123] In one or more optional embodiments of the present invention, the element type characteristics include a first quantity, a second quantity, a third quantity, a fourth quantity, a fifth quantity, and a sixth quantity;
[0124] The text type condition includes that the first quantity, the second quantity, and the sixth quantity are all first preset values, the first element is a character type or a punctuation type, and the fourth quantity is less than a first set value or the ratio of the fourth quantity to the total quantity is less than a second set value, and the total quantity is the sum of the third quantity, the fourth quantity, and the sixth quantity;
[0125] The comprehensive type condition includes that the sum of the first quantity and the second quantity is a second preset value, and the sixth quantity is a first preset value.
[0126] Specifically, the first number is the number of path types in each target element; the second number is the number of picture types in each target element; the third number is the number of character types of text elements in each target element; the fourth number is the number of punctuation types of text elements in each target element; the fifth number is the number of numeric types of text elements in each target element; and the sixth number is the number of other types of text elements in each target element.
[0127] It should be noted that since the drop cap modifier must modify Latin-based text, only basic elements of the character series, punctuation series, and number series can appear. If the sixth number O > 0, it indicates that there is no Drop Cap in the current paragraph to be detected, that is, the paragraph to be detected is not a drop cap paragraph. Therefore, both the text type condition and the comprehensive type condition include the sixth number being the first preset value, and the first preset value is zero.
[0128] When the sixth quantity is the first preset value, if there are no basic elements of the image type and the path type in each target element, it indicates that the element with the first word dropped (the first word dropped element) may be of text type. At this time, it is necessary to ensure that the remaining text types are all in legal form, that is, the beginning of the text must be a character type or a numeric type text element, and the number of symbol type target elements (the fourth quantity) must account for a smaller proportion or the fourth quantity must be smaller. Therefore, the text type conditions include: the first quantity, the second quantity, and the sixth quantity are all the first preset value; the first element is a character type or a punctuation type; the fourth quantity is less than the first set value (such as 3) or the ratio of the fourth quantity to the total amount is less than the second set value, and the total amount is the sum of the third quantity, the fourth quantity, and the sixth quantity.
[0129] When the sixth number is the first preset value, if an image type or path type exists among the target elements, then the element indicating the drop cap is the target element corresponding to that image type or path type, so the total number of image types and path types cannot exceed one. Therefore, the comprehensive type condition includes: the sum of the first number and the second number is the second preset value, and the second preset value is one; and the sixth number is the first preset value.
[0130] In this way, the reliability of the text type condition and the comprehensive type condition can be improved, thereby improving the accuracy and reliability of detecting the first-letter drop paragraph.
[0131] In one or more optional embodiments of the present invention, the obtaining of the first height of the first element in the paragraph to be detected may be specifically implemented as follows:
[0132] Obtain the third circumscribed rectangular frame of the first element in the paragraph to be detected;
[0133] The height of the third circumscribed rectangular frame is used as the first height of the first element.
[0134] Specifically, the third circumscribed rectangular frame refers to the circumscribed rectangular frame of the first element.
[0135] In practice, obtain the third bounding box of the first element of the paragraph to be detected, and then use the height of the third bounding box as the first height of the first element (DropCapBBox.Height). This method can improve the accuracy of the first height by using the third bounding box of the first element.
[0136] In one or more optional embodiments of the present invention, when the first height meets a preset height condition, determining that the paragraph to be detected is a drop cap paragraph and the first element is a drop cap element further includes:
[0137] Obtaining a fourth circumscribed rectangular frame of each designated element in the first text line of the paragraph to be detected, wherein the designated elements are elements other than the first element in the first text line;
[0138] Using the height of the union of the fourth circumscribed rectangular frames as the second height of the first text line;
[0139] For each designated text line in the paragraph to be detected, obtaining a fifth circumscribed rectangular box of each element in the designated text line, wherein the designated text line is a text line in the paragraph to be detected except the first text line; and using the height of the union of the fifth circumscribed rectangular boxes as the third height of the designated text line;
[0140] A preset height condition is constructed according to the first height, the second height and each of the third heights.
[0141] Specifically, the fourth bounding rectangle refers to the bounding rectangle of the specified element in the first text line. The second height refers to the height of the first text line. The fifth bounding rectangle refers to the bounding rectangle of the basic element in the specified text line. The third height refers to the height of the specified text line.
[0142] In practical applications, the fourth bounding rectangles of all specified elements in the first text line, excluding the first element, can be obtained. Furthermore, the union of the fourth bounding rectangles is used as the bounding rectangle of the first text line (OtherBBox1), and the height of OtherBBox1 is used as the second height of the first text line (OtherBBox1.Height).
[0143] For the non-first text line in the paragraph to be detected, that is, the specified text line, obtain the fifth circumscribed rectangular box of each element (basic element) in the text line; further, use the union of the fifth circumscribed rectangular boxes related to the specified text line as the circumscribed rectangular box of the specified text line (OtherBBox2), and use the height of OtherBBox2 as the third height (OtherBBox2.Height) of the specified text line.
[0144] Based on the first height, the second height, and each third height, a preset height condition is constructed based on the first height, the second height, and each third height. In this way, the matching degree and correlation between the preset height condition and the paragraph to be detected can be ensured, thereby ensuring the accuracy of detecting the paragraph with the first letter drop.
[0145] In one or more optional embodiments of the present invention, the preset height condition includes:
[0146] The first height is greater than a product of a first value and the second height, and a product of each of the third heights and the second value is less than the first height.
[0147] Specifically, the first value may be greater than or equal to 1.5, and the second value may also be greater than or equal to 1.5. The first value may be the same as or different from the second value.
[0148] Exemplarily, the preset height conditions include: DropCapBBox.Height>1.6(first value)*OtherBBox1.Height, and DropCapBBox.Height>1.55(second value)*OtherBBox2.Height(i), i=1,2,…,n, where n is the number of specified text lines in the paragraph to be detected.
[0149] In one or more optional embodiments of the present invention, before extracting a set number of non-modifying target elements from the paragraph to be detected when the text line features of the paragraph to be detected meet the preset text line conditions, the method further includes:
[0150] Obtain the number of text lines contained in the paragraph to be detected;
[0151] Identifying the writing direction of each text line in the paragraph to be detected;
[0152] Determining whether each text line in the paragraph to be detected has a first text feature of the text content;
[0153] Determining whether a second text feature of a jump target anchor exists in each text line in the paragraph to be detected;
[0154] Determining whether a third text feature of a text type element exists in each text line in the paragraph to be detected;
[0155] The number of text lines, the writing direction, the first text feature, the second text feature, and the third text feature are determined as the text line feature.
[0156] Specifically, the jump target anchor refers to the location corresponding to the link jump.
[0157] In practical applications, the number of text lines contained in the paragraph to be detected, that is, the number of text lines, can be determined. Since DropCap modification must be implemented with two or more text lines, it is necessary to identify the number of text lines contained in the paragraph to be detected.
[0158] The writing direction of each text line in the paragraph to be detected can be identified.
[0159] It is possible to identify whether each text line in the paragraph to be detected contains text content, thereby obtaining the first text feature; for example, if there is only an underline in the text line, and the underline modifies a space, then there is no specific text content in the text line.
[0160] The second text feature can be determined by identifying whether there is a bookmark jump target anchor in each text line of the paragraph to be detected. Since bookmarks, hyperlinks, and jumps between anchors are allowed in documents, the jump location is usually the title. If a jump target anchor exists in a text line, it means that the text line is in the title area, and the drop cap mechanism is impossible.
[0161] The third text feature can be obtained by identifying whether each text line in the paragraph to be detected contains a text type element. Since an image may be included as the text content of the current line, and when an image occupies a text line, it is impossible to use a DropCap decoration, so it is necessary to identify whether each text line contains a text type element.
[0162] Furthermore, the number of text lines, writing direction, first text feature, second text feature, and third text feature constitute a text line feature. This improves the accuracy of the text line feature and, in turn, the accuracy of detecting drop cap paragraphs.
[0163] In one or more optional embodiments of the present invention, the preset text line condition includes: the number of text lines is greater than a third value, the writing directions are all set writing directions, the first text feature is that each text line contains text content, the second text feature is that each text line does not contain a jump target anchor, and the third text feature is that each text line contains a text type element. The third value is a natural number greater than 1.
[0164] The following combination Figure 3 The method for detecting the drop cap paragraph provided by the present invention is further described. Figure 3 This is the second flow chart of the method for detecting the first-character drop-head paragraph provided by the present invention.
[0165] Step 301: Extract text line features of the paragraph to be detected.
[0166] Specifically, the number of text lines contained in the paragraph to be detected is obtained; the writing direction of each text line in the paragraph to be detected is identified; a first text feature is determined as to whether the main text content exists in each text line in the paragraph to be detected; a second text feature is determined as to whether a jump target anchor exists in each text line in the paragraph to be detected; a third text feature is determined as to whether a text type element exists in each text line in the paragraph to be detected; the text line features include the number of text lines, the writing direction, the first text feature, the second text feature and the third text feature.
[0167] Step 302: Determine whether the text line feature meets the preset text line conditions.
[0168] If yes, go to step 303; if no, go to step 310.
[0169] Among them, the preset text line conditions include: the number of text lines is greater than a third value, the writing direction is the set writing direction, the first text feature is that there is text content in each text line, the second text feature is that there is no jump target anchor in each text line, and the third text feature is that there is a text type element in each text line.
[0170] For example, a basic feature determination is performed on a text line in a text region (paragraph to be detected). If any of the following conditions are met, it is considered that there is no drop cap modification in a given text line in the current text region:
[0171] The number of text lines in the text area is less than 2;
[0172] There is no text content in any line;
[0173] There is a jump target anchor for the bookmark in any line;
[0174] Based on the method of quickly identifying the writing direction of text content in PDF, the writing direction of each line is determined. Within all given lines, there are cases where the writing direction is not LrTb (the writing direction is from left to right within the line and from top to bottom between lines);
[0175] In any row, although there is text content, there is no text type element.
[0176] Step 303: extracting a set number of non-modifying target elements from the paragraph to be detected.
[0177] Specifically, each non-modified initial element in the paragraph to be detected is extracted; a first circumscribed rectangular box of each initial element and a second circumscribed rectangular box of the paragraph to be detected are obtained; the distance between each first circumscribed rectangular box and the second circumscribed rectangular box is determined; and a set number of target elements are screened from each initial element based on the distances.
[0178] For example, all non-modifying basic elements (initial elements) in the current text area are taken out and sorted from smallest to largest based on the distance from the center of each initial element's BBox to the upper left corner of the text block (paragraph to be detected). The five basic elements with the closest distance are selected as the target elements.
[0179] Step 304: Perform element type statistics on each target element to obtain element type characteristics.
[0180] Specifically, the first number of path types in each target element is counted, and the second number of picture types in each target element is counted; for text elements in each target element, the third number of character types, the fourth number of punctuation types, the fifth number of number types, and the sixth number of other types are counted; the element type features include the first number, the second number, the third number, the fourth number, the fifth number, and the sixth number.
[0181] For example, the target elements extracted are judged one by one: (1) PathCount is recorded as the number of Path type elements in the basic elements; (2) ImageCount is recorded as the number of Image type elements in the basic elements; (3) For all basic elements of Text type, according to the Unicode character category classification standard, the number of character (L) type, punctuation (P) type, number (N) and other (O) types are counted, and recorded as L, P, N, O respectively.
[0182] Step 305: Determine whether the element type feature meets the preset element conditions.
[0183] If yes, go to step 306; if no, go to step 310.
[0184] Specifically, the preset element conditions include text type conditions and comprehensive type conditions; the text type conditions include that the first quantity, the second quantity and the sixth quantity are all zero, the first element is a character type or a punctuation type, and the fourth quantity is less than the first set value or the ratio of the fourth quantity to the total amount is less than the second set value, and the total amount is the sum of the third quantity, the fourth quantity and the sixth quantity; the comprehensive type conditions include that the sum of the first quantity and the second quantity is one, and the sixth quantity is zero.
[0185] For example, since the drop cap modification must modify Latin language text, only elements of the Letter type, punctuation P type, and number N type can appear. If O>0, it indicates that there is no DropCap in the current language; if O=0, continue to judge:
[0186] If there are no image or path type target elements among the target elements, the element with the drop cap may be of text type. In this case, you need to ensure that the remaining text types are all legal. That is, the text must start with a letter or a number, and the number of symbols must be small (P less than 3);
[0187] If any target element has an image or path type, then the element to be used for the drop cap is an element of that image or path type. Therefore, the number of image and path type elements cannot exceed one.
[0188] Step 306: Obtain the first height of the first element in the paragraph to be detected.
[0189] Specifically, the third circumscribed rectangular frame of the first element in the paragraph to be detected is obtained; and the height of the third circumscribed rectangular frame is used as the first height of the first element.
[0190] Exemplarily, the element closest to the upper left corner in the paragraph to be detected (the first element) is recorded as DropCapObject, the BBox of the first element is recorded as DropCapBBox, and the first height is recorded as DropCapBBox.Height.
[0191] Step 307: Constructing a preset height condition.
[0192] Specifically, the fourth circumscribed rectangular box of each specified element in the first text line in the paragraph to be detected is obtained, where the specified elements are the elements other than the first element in the first text line; the height of the union of each fourth circumscribed rectangular box is used as the second height of the first text line; for each specified text line in the paragraph to be detected, the fifth circumscribed rectangular box of each element in the specified text line is obtained, where the specified text line is the text line other than the first text line in the paragraph to be detected; the height of the union of each fifth circumscribed rectangular box is used as the third height of the specified text line; and a preset height condition is constructed based on the first height, the second height and each third height.
[0193] The preset height condition includes: the first height is greater than the product of the first value and the second height, and the product of each third height and the second value is less than the first height.
[0194] Step 308: Determine whether the first height meets a preset height condition.
[0195] If yes, go to step 309; if no, go to step 310.
[0196] If the first height satisfies all of the following conditions, then it is considered that there is a drop cap modification in the text area, and the element of the drop cap modification is a DropCapObject:
[0197] (1) In the first text line, except for the first element, the union of all the remaining basic element BBoxes is called OtherBBox1;
[0198] DropCapBBox.Height>1.6*OtherBBox1.Height;
[0199] (2) Except for the first text line, the union of all basic element BBoxes of each text line in all remaining text lines is called OtherBBox2;
[0200] DropCapBBox.Height>1.55*OtherBBox2.Height.
[0201] Step 309: Determine whether the paragraph to be detected is a drop cap paragraph and the first element is a drop cap element.
[0202] Specifically, if the DropCapObject is found, the attribute DropCap=True is assigned to its corresponding element.
[0203] Step 310: Determine whether the paragraph to be detected is not a drop cap paragraph.
[0204] The method for detecting a paragraph with a drop cap provided in this embodiment determines whether there is a drop cap modification based on the features between a number of text elements and lines at the beginning of a paragraph, and determines the specific text element corresponding to the drop cap character through BBox. This method can be used to detect whether there is a drop cap element at the beginning of a paragraph, including layout documents (such as PDF, OFD, etc.) and pictures whose text content is identified by OCR technology. The typesetting modifications of the current document are thereby identified and correctly divided. When performing document conversion (for example, conversion into a streaming document), the drop cap modification can be added to a supported editor, or directly converted to a normal font size for easy reading. In technologies such as PDF reformatting, the first character can be treated as ordinary text.
[0205] The following describes a device for detecting a drop-cap paragraph provided by the present invention. The device for detecting a drop-cap paragraph described below and the method for detecting a drop-cap paragraph described above can be referenced to each other.
[0206] Figure 4 : is a schematic diagram of the structure of the detection device for the first-letter drop paragraph provided by the present invention, such as Figure 4 As shown, the apparatus 400 for detecting a drop cap paragraph includes: an extraction module 401, a statistics module 402, an acquisition module 403, and a determination module 404, wherein:
[0207] The extraction module 401 is configured to extract a set number of non-modifying target elements from the paragraph to be detected when the text line features of the paragraph to be detected meet the preset text line conditions;
[0208] A statistics module 402 is configured to perform element type statistics on each target element to obtain element type characteristics;
[0209] The acquisition module 403 is configured to acquire a first height of the first element in the paragraph to be detected if the element type feature meets the preset element condition;
[0210] The determination module 404 is configured to determine that the paragraph to be detected is a first-character drop-head paragraph when the first height meets a preset height condition.
[0211] The device for detecting a drop cap paragraph provided by the present invention extracts a set number of non-modifying target elements from the paragraph to be detected when the text line characteristics of the paragraph to be detected meet the preset text line conditions; performs element type statistics on each of the target elements to obtain element type characteristics; obtains the first height of the first element in the paragraph to be detected when the element type characteristics meet the preset element conditions; and determines that the paragraph to be detected is a drop cap paragraph when the first height meets the preset height conditions. The present invention can identify whether there is a drop cap element in a paragraph and the corresponding specific element modified, that is, detects drop cap paragraphs and drop cap elements in a document; is conducive to more accurate analysis of paragraph layout, and helps to make subsequent format document conversion, re-layout, etc. more accurate.
[0212] In one or more optional embodiments of the present invention, the extraction module 401 is further configured to:
[0213] Extracting each non-modifying initial element in the paragraph to be detected;
[0214] Obtaining a first circumscribed rectangular frame of each of the initial elements and a second circumscribed rectangular frame of the paragraph to be detected;
[0215] Determining the distance between each of the first circumscribed rectangular frames and the second circumscribed rectangular frames;
[0216] According to each of the distances, a set number of target elements are selected from each of the initial elements.
[0217] In one or more optional embodiments of the present invention, the statistics module 402 is further configured to:
[0218] Counting a first number of path types in each of the target elements, and counting a second number of picture types in each of the target elements;
[0219] For each text element in the target element, count a third number of characters, a fourth number of punctuation marks, a fifth number of numbers, and a sixth number of other types;
[0220] The first quantity, the second quantity, the third quantity, the fourth quantity, the fifth quantity and the sixth quantity are determined as the element type characteristics.
[0221] In one or more optional embodiments of the present invention, the preset element condition includes a text type condition and a comprehensive type condition, and the comprehensive type condition represents conditions corresponding to the path type and the image type;
[0222] The apparatus 400 for detecting a drop cap paragraph further includes a comparison module configured to:
[0223] Comparing the element type feature with the text type condition and the comprehensive type condition respectively to obtain a text comparison result and a comprehensive comparison result;
[0224] If both the text comparison result and the comprehensive comparison result indicate a mismatch, determining that the element type feature does not meet the preset element condition;
[0225] If there is a comparison result indicating compliance between the text comparison result and the comprehensive comparison result, it is determined that the element type feature complies with the preset element condition.
[0226] In one or more optional embodiments of the present invention, the element type characteristics include a first quantity, a second quantity, a third quantity, a fourth quantity, a fifth quantity, and a sixth quantity;
[0227] The text type condition includes that the first quantity, the second quantity, and the sixth quantity are all first preset values, the first element is a character type or a punctuation type, and the fourth quantity is less than a first set value or the ratio of the fourth quantity to the total quantity is less than a second set value, and the total quantity is the sum of the third quantity, the fourth quantity, and the sixth quantity;
[0228] The comprehensive type condition includes that the sum of the first quantity and the second quantity is a second preset value, and the sixth quantity is a first preset value.
[0229] In one or more optional embodiments of the present invention, the acquisition module 403 is further configured to:
[0230] Obtain the third circumscribed rectangular frame of the first element in the paragraph to be detected;
[0231] The height of the third circumscribed rectangular frame is used as the first height of the first element.
[0232] In one or more optional embodiments of the present invention, the apparatus 400 for detecting a drop cap paragraph further includes a construction module configured to:
[0233] Obtaining a fourth circumscribed rectangular frame of each designated element in the first text line of the paragraph to be detected, wherein the designated elements are elements other than the first element in the first text line;
[0234] Using the height of the union of the fourth circumscribed rectangular frames as the second height of the first text line;
[0235] For each designated text line in the paragraph to be detected, obtaining a fifth circumscribed rectangular box of each element in the designated text line, wherein the designated text line is a text line in the paragraph to be detected except the first text line; and using the height of the union of the fifth circumscribed rectangular boxes as the third height of the designated text line;
[0236] A preset height condition is constructed according to the first height, the second height and each of the third heights.
[0237] In one or more optional embodiments of the present invention, the preset height condition includes:
[0238] The first height is greater than a product of a first value and the second height, and a product of each of the third heights and the second value is less than the first height.
[0239] In one or more optional embodiments of the present invention, the apparatus 400 for detecting a drop cap paragraph further includes a recognition module configured to:
[0240] Obtain the number of text lines contained in the paragraph to be detected;
[0241] Identifying the writing direction of each text line in the paragraph to be detected;
[0242] Determining whether each text line in the paragraph to be detected has a first text feature of the text content;
[0243] Determining whether a second text feature of a jump target anchor exists in each text line in the paragraph to be detected;
[0244] Determining whether a third text feature of a text type element exists in each text line in the paragraph to be detected;
[0245] The number of text lines, the writing direction, the first text feature, the second text feature, and the third text feature are determined as the text line feature.
[0246] In one or more optional embodiments of the present invention, the preset text line condition includes:
[0247] The number of text lines is greater than a third value, the writing directions are all set writing directions, the first text feature is that there is text content in each text line, the second text feature is that there is no jump target anchor in each text line, and the third text feature is that there is a text type element in each text line.
[0248] Figure 5An example of a physical structure diagram of an electronic device is shown below. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute a method for detecting a paragraph with a first-letter drop, the method comprising: extracting a set number of non-modifying target elements from the paragraph to be detected when the text line feature of the paragraph to be detected meets the preset text line condition; performing element type statistics on each of the target elements to obtain an element type feature; obtaining the first height of the first element in the paragraph to be detected when the element type feature meets the preset element condition; and determining that the paragraph to be detected is a paragraph with a first-letter drop when the first height meets the preset height condition.
[0249] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0250] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the detection method of the first-letter-drop paragraph provided by the above-mentioned methods, and the method includes: when the text line feature of the paragraph to be detected meets the preset text line condition, extracting a set number of non-modifying target elements from the paragraph to be detected; performing element type statistics on each of the target elements to obtain element type features; when the element type feature meets the preset element condition, obtaining the first height of the first element in the paragraph to be detected; when the first height meets the preset height condition, determining that the paragraph to be detected is a first-letter-drop paragraph.
[0251] On the other hand, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for detecting paragraphs with a first-letter drop provided by the above-mentioned methods, the method comprising: when the text line feature of the paragraph to be detected meets the preset text line condition, extracting a set number of non-modifying target elements from the paragraph to be detected; performing element type statistics on each of the target elements to obtain element type features; when the element type features meet the preset element conditions, obtaining the first height of the first element in the paragraph to be detected; when the first height meets the preset height condition, determining that the paragraph to be detected is a paragraph with a first-letter drop.
[0252] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0253] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0254] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for detecting a paragraph with a drop cap, characterized in that: include: When the text line feature of the paragraph to be detected meets the preset text line condition, extracting a set number of non-modifying target elements from the paragraph to be detected; Performing element type statistics on each of the target elements to obtain element type characteristics; When the element type feature meets the preset element condition, obtaining the first height of the first element in the paragraph to be detected; When the first height meets a preset height condition, the paragraph to be detected is determined to be a first-character drop-head paragraph.
2. The method for detecting a drop cap paragraph according to claim 1, wherein: The step of extracting a set number of non-modifying target elements from the paragraph to be detected includes: Extracting each non-modifying initial element in the paragraph to be detected; Obtaining a first circumscribed rectangular frame of each of the initial elements and a second circumscribed rectangular frame of the paragraph to be detected; Determining the distance between each of the first circumscribed rectangular frames and the second circumscribed rectangular frames; According to each of the distances, a set number of target elements are selected from each of the initial elements.
3. The method for detecting a drop cap paragraph according to claim 1, wherein: The performing element type statistics on each target element to obtain element type characteristics includes: Counting a first number of path types in each of the target elements, and counting a second number of picture types in each of the target elements; For each text element in the target element, count a third number of characters, a fourth number of punctuation marks, a fifth number of numbers, and a sixth number of other types; The first quantity, the second quantity, the third quantity, the fourth quantity, the fifth quantity and the sixth quantity are determined as the element type characteristics.
4. The method for detecting a drop cap paragraph according to any one of claims 1 to 3, wherein: The preset element conditions include text type conditions and comprehensive type conditions, and the comprehensive type conditions represent conditions corresponding to the path type and the image type; Before obtaining the first height of the first element in the paragraph to be detected when the element type feature meets the preset element condition, the method further includes: Comparing the element type feature with the text type condition and the comprehensive type condition respectively to obtain a text comparison result and a comprehensive comparison result; If both the text comparison result and the comprehensive comparison result indicate a mismatch, determining that the element type feature does not meet the preset element condition; If there is a comparison result indicating compliance between the text comparison result and the comprehensive comparison result, it is determined that the element type feature complies with the preset element condition.
5. The method for detecting a drop cap paragraph according to claim 4, wherein: The element type characteristics include a first quantity, a second quantity, a third quantity, a fourth quantity, a fifth quantity, and a sixth quantity; The text type condition includes that the first quantity, the second quantity, and the sixth quantity are all first preset values, the first element is a character type or a punctuation type, and the fourth quantity is less than a first set value or the ratio of the fourth quantity to the total quantity is less than a second set value, and the total quantity is the sum of the third quantity, the fourth quantity, and the sixth quantity; The comprehensive type condition includes that the sum of the first quantity and the second quantity is a second preset value, and the sixth quantity is a first preset value.
6. The method for detecting a drop cap paragraph according to claim 1, wherein: The obtaining of the first height of the first element in the paragraph to be detected includes: Obtain the third circumscribed rectangular frame of the first element in the paragraph to be detected; The height of the third circumscribed rectangular frame is used as the first height of the first element.
7. The method for detecting a drop cap paragraph according to claim 1, wherein: Before determining that the paragraph to be detected is a drop cap paragraph when the first height meets a preset height condition, the method further includes: Obtaining a fourth circumscribed rectangular frame of each designated element in the first text line of the paragraph to be detected, wherein the designated elements are elements other than the first element in the first text line; Using the height of the union of the fourth circumscribed rectangular frames as the second height of the first text line; For each designated text line in the paragraph to be detected, obtaining a fifth circumscribed rectangular box of each element in the designated text line, wherein the designated text line is a text line in the paragraph to be detected except the first text line; and using the height of the union of the fifth circumscribed rectangular boxes as the third height of the designated text line; A preset height condition is constructed according to the first height, the second height and each of the third heights.
8. The method for detecting a drop cap paragraph according to claim 7, wherein: The preset height conditions include: The first height is greater than a product of a first value and the second height, and a product of each of the third heights and the second value is less than the first height.
9. The method for detecting a drop cap paragraph according to claim 1, wherein: Before extracting a set number of non-modifying target elements from the paragraph to be detected when the text line features of the paragraph to be detected meet the preset text line conditions, the method further includes: Obtain the number of text lines contained in the paragraph to be detected; Identifying the writing direction of each text line in the paragraph to be detected; Determining whether each text line in the paragraph to be detected has a first text feature of the text content; Determining whether a second text feature of a jump target anchor exists in each text line in the paragraph to be detected; Determining whether a third text feature of a text type element exists in each text line in the paragraph to be detected; The number of text lines, the writing direction, the first text feature, the second text feature, and the third text feature are determined as the text line feature.
10. The method for detecting a drop cap paragraph according to claim 9, wherein: The preset text line conditions include: The number of text lines is greater than a third value, the writing directions are all set writing directions, the first text feature is that there is text content in each text line, the second text feature is that there is no jump target anchor in each text line, and the third text feature is that there is a text type element in each text line.
11. A device for detecting a paragraph with a dropped first character, characterized in that: include: an extraction module configured to extract a set number of non-modifying target elements from the paragraph to be detected when the text line features of the paragraph to be detected meet the preset text line conditions; a statistics module configured to perform element type statistics on each target element to obtain element type characteristics; An acquisition module configured to acquire a first height of the first element in the paragraph to be detected if the element type feature meets a preset element condition; The determination module is configured to determine that the paragraph to be detected is a drop cap paragraph and the first element is a drop cap element if the first height meets a preset height condition.