Text extraction method, apparatus, device, and storage medium
Patent Information
- Application Number
- CN202610952833.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本申请的主要目的在于提供了一种文本提取方法、装置、设备及存储介质,旨在解决PDF文本提取过程中由于乱码原因难以提取有效文本的技术问题
[0015]本申请提出的一个或多个技术方案,至少具有以下技术效果:本申请的文本提取方法包括:确定待提取文件对应的文件类型,并在所述文件类型为预设文本类型的情况下,从所述待提取文件中提取初始文本;对所述初始文本进行检测,确定当前乱码比例和有效文本长度;基于所述当前乱码比例和所述有效文本长度,对所述待提取文件进行降级提取,得到对应的乱码替换文本;根据所述乱码替换文本对所述初始文本进行替换,得到所述待提取文件对应的文本提取结果。
Smart Images

Figure CN122819149A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text processing technology, and in particular to a text extraction method, apparatus, device and storage medium. Background Technology
[0002] Currently, in scenarios involving the automated processing of enterprise documents such as vouchers, contracts, and reports, the method of directly extracting text from Portable Document Format (PDF) files to support subsequent Artificial Intelligence (AI) analysis is generally achieved by parsing the electronic text within the PDF's internal text layer.
[0003] Existing PDF text extraction solutions mainly fall into two categories: those containing a text layer use direct extraction from the electronic text (such as tools like Apache PDFBox and iText), while those without a text layer use Optical Character Recognition (OCR). However, due to issues like abnormal font embedding, missing encoding mapping, or the use of proprietary fonts, some PDFs containing a text layer still contain a large number of garbled characters when directly extracted (such as Unicode replacement characters U+FFFD, proprietary characters, etc.). Furthermore, the extraction strategy cannot be adaptively adjusted based on the degree of garbled characters, making it difficult to identify such garbled characters and switch to an effective extraction method, thus rendering subsequent AI analysis completely ineffective. Summary of the Invention
[0004] The main purpose of this application is to provide a text extraction method, apparatus, device and storage medium, which aims to solve the technical problem that it is difficult to extract valid text due to garbled characters during PDF text extraction.
[0005] To achieve the above objectives, this application proposes a text extraction method, the method comprising: Determine the file type corresponding to the file to be extracted, and extract the initial text from the file to be extracted if the file type is a preset text type; The initial text is inspected to determine the current garbled text ratio and the length of the effective text; Based on the current garbled text ratio and the effective text length, the file to be extracted is downgraded for extraction to obtain the corresponding garbled text replacement text; The initial text is replaced with the garbled text replacement text to obtain the text extraction result corresponding to the file to be extracted.
[0006] In one embodiment, the step of detecting the initial text and determining the current garbled text ratio and the effective text length includes: The number of first invalid characters is determined based on the number of times the replacement character appears in the initial text; Detect private zone characters in the initial text, and determine the number of second invalid characters based on the number of occurrences of the private zone characters; Detect control characters in the initial text, and determine the number of third invalid characters based on the number of times the control characters appear; Based on the number of the first invalid characters, the number of the second invalid characters, the number of the third invalid characters, and the total number of characters in the initial text, determine the current garbled character ratio and the effective text length corresponding to the initial text.
[0007] In one embodiment, the step of determining the current garbled character ratio and valid text length corresponding to the initial text based on the number of the first invalid characters, the number of the second invalid characters, the number of the third invalid characters, and the total number of characters in the initial text includes: Detect candidate characters located in the preset proxy character area in the initial text, and determine the number of fourth invalid characters corresponding to invalid proxy pairs among the candidate characters; Code point detection is performed on the initial text to obtain the number of fifth invalid characters corresponding to non-characters; The number of sixth invalid characters is determined based on the number of occurrences of whitespace characters in the initial text; The sum of the first invalid character count, the second invalid character count, the third invalid character count, the fourth invalid character count, the fifth invalid character count, and the sixth invalid character count is used to obtain the current invalid character count. Based on the total number of characters in the initial text and the current number of invalid characters, determine the corresponding current garbled character ratio and valid text length.
[0008] In one embodiment, the step of downgrading the extraction of the file to be extracted based on the current garbled text ratio and the effective text length to obtain the corresponding garbled text replacement text includes: The file to be extracted is detected based on the current garbled character ratio and the effective text length to determine whether it meets the preset degradation conditions. The preset degradation conditions include the current garbled character ratio reaching a first preset ratio and / or the effective text length being less than a preset length. If the file to be extracted meets the preset downgrade conditions, the file to be extracted is downgraded and converted to obtain the corresponding image to be extracted; Optical character recognition is performed on the image to be extracted to obtain the current valid text; Determine the corresponding garbled text replacement text from the currently valid text.
[0009] In one embodiment, after the step of detecting whether the file to be extracted meets the preset downgrade conditions based on the current garbled text ratio and the effective text length, the method further includes: If the current garbled character ratio does not reach the first preset ratio, it is determined whether the current garbled character ratio reaches the second preset ratio, where the first preset ratio is higher than the second preset ratio; If the current garbled text ratio reaches the second preset ratio, then the text pages containing garbled text in the file to be extracted are marked as candidate downgraded texts, and the candidate downgraded texts are reviewed. If the verification result is a garbled result, the candidate downgraded text is downgraded and extracted to obtain the corresponding garbled replacement text.
[0010] In one embodiment, the step of replacing the initial text with the garbled replacement text to obtain the text extraction result corresponding to the file to be extracted includes: The garbled text replacement text is uploaded to a distributed object storage system, and a corresponding first storage path is obtained. The distributed object storage system is used to store the text extraction results. Determine the second storage path of the initial text in a preset relational database; The second storage path is updated to the first storage path in the preset relational database to obtain the text extraction result corresponding to the file to be extracted.
[0011] In one embodiment, the step of determining the file type corresponding to the file to be extracted includes: The text layer detection of the file to be extracted is performed to obtain the first detection result; If a text layer exists in the file to be extracted, extract the font name field and font descriptor from the text layer; The second detection result is determined based on the symbol characteristics of the font name field and / or the flag characteristics of the font descriptor; Based on the first detection result and the second detection result, the file type corresponding to the file to be extracted is determined.
[0012] Furthermore, to achieve the above objectives, this application also proposes a text extraction device, the device comprising: The extraction module is used to determine the file type of the file to be extracted, and extract the initial text from the file to be extracted if the file type is a preset text type; The detection module is used to detect the initial text and determine the current garbled text ratio and the length of the effective text. The downgrade module is used to downgrade the file to be extracted based on the current garbled character ratio and the effective text length, so as to obtain the corresponding garbled character replacement text. The update module is used to replace the initial text with the garbled text replacement text to obtain the text extraction result corresponding to the file to be extracted.
[0013] In addition, to achieve the above objectives, this application also proposes a text extraction device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the text extraction method as described above.
[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the text extraction method described above.
[0015] One or more technical solutions proposed in this application have at least the following technical effects: The text extraction method of this application includes: determining the file type corresponding to the file to be extracted, and extracting initial text from the file to be extracted when the file type is a preset text type; detecting the initial text to determine the current garbled character ratio and the effective text length; performing downgrade extraction on the file to be extracted based on the current garbled character ratio and the effective text length to obtain the corresponding garbled character replacement text; and replacing the initial text according to the garbled character replacement text to obtain the text extraction result corresponding to the file to be extracted.
[0016] This application first determines the file type of the file to be extracted and extracts the initial text when it is not plain text. Then, it detects the current garbled text ratio and the length of valid text in the initial text. Based on the detection results, it performs downgraded extraction to obtain garbled text replacement text and updates the initial text. This application can automatically judge the extraction quality and perform downgraded extraction by using the garbled text ratio and the length of valid text, reducing the occurrence of invalid text output and improving the reliability of text processing. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating an embodiment of the text extraction method of this application. Figure 2 This is an overall flowchart of document extraction provided in Embodiment 1 of this application; Figure 3 This is a flowchart illustrating Embodiment 2 of the text extraction method of this application; Figure 4 This is a flowchart illustrating Embodiment 3 of the text extraction method of this application; Figure 5 This is a block diagram of the module structure of the text extraction device according to an embodiment of this application; Figure 6 This is a schematic diagram of the hardware operating environment involved in the text extraction device in this application embodiment.
[0020] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0022] The main solution proposed in this application is as follows: Currently, in scenarios involving the automated processing of enterprise documents such as vouchers, contracts, and reports, in order to extract text from portable document format files to support subsequent artificial intelligence analysis, the method of directly extracting electronic text by parsing the internal text layer of PDF is generally adopted.
[0023] Existing PDF text extraction solutions mainly fall into two categories: those containing a text layer use direct extraction from the electronic text, while those without a text layer use optical character recognition (OCR). However, due to issues such as abnormal font embedding, missing encoding mapping, or the use of proprietary fonts, the text extracted directly from some PDF files containing a text layer still contains a large number of garbled characters. It is difficult to identify such garbled characters and switch to an effective extraction method, causing subsequent AI analysis to completely fail.
[0024] To address the aforementioned issues, this application provides a text extraction method. This method first determines the file type of the file to be extracted and extracts the initial text when it is not plain text. Then, it detects the current garbled text ratio and the length of valid text in the initial text. Based on the detection results, it performs downgraded extraction to obtain garbled text replacement text and updates the initial text. This application can automatically determine the extraction quality and perform downgraded extraction based on the garbled text ratio and the length of valid text, reducing invalid text output and improving the reliability of text processing.
[0025] It should be noted that the executing entity of this application embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a personal computer, server, cloud server, or other text extraction device capable of executing the text extraction method of this application. This embodiment does not limit this. The following uses a text extraction device (hereinafter referred to as the device) as an example to describe this embodiment and the following embodiments.
[0026] Based on this, this application proposes a text extraction method according to a first embodiment, referring to... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the text extraction method of this application. In this embodiment, the text extraction method may include steps S10 to S40: Step S10: Determine the file type corresponding to the file to be extracted, and if the file type is a preset text type, extract the initial text from the file to be extracted.
[0027] It should be noted that the file to be extracted can be an electronic document from which text content needs to be extracted, such as a portable document format file, an image format file, or a scanned document.
[0028] Understandably, the file type can be a classification result of the structural features of the file to be extracted, including plain text, scanned, and mixed types. The preset text type can be a pre-defined PDF file type that may result in garbled characters during electronic text extraction, i.e., scanned, mixed, or other non-plain text types.
[0029] Plain text files can be files containing a text layer that can be parsed normally and does not have garbled characters. For example, a PDF file directly exported from word processing software has complete character encoding and correct font embedding in its text layer, and can be directly extracted to obtain readable text. However, scanned and mixed text types may have garbled characters during the text extraction process, so subsequent garbled character detection or quality assessment is required.
[0030] It should be understood that the initial text may be the raw character sequence obtained directly from the file to be extracted by parsing the internal text layer, but has not yet undergone garbled character detection or quality assessment.
[0031] In practical use, the aforementioned device first acquires the file to be extracted, then determines the file type, such as checking whether a text layer exists in the file, whether the font embedding information in the text layer is complete, and whether the character encoding mapping is normal, to determine the file type. Next, it determines whether the determined file type is plain text. If it is not plain text, it performs text extraction from the file to obtain the initial text, and then performs subsequent garbled character detection or quality assessment.
[0032] If the file type is plain text, the subsequent garbled character detection and downgrade extraction process can be skipped, and the extracted regular text result can be used directly as the final output.
[0033] Step S20: Detect the initial text to determine the current garbled text ratio and the effective text length.
[0034] It should be noted that the current garbled text ratio can be the ratio between the number of invalid characters in the initial text and the total number of characters in the initial text, used to quantitatively assess the severity of garbled text in the initial text. Invalid characters can include characters that cannot be correctly displayed as readable text due to missing encoding mapping, abnormal font embedding, or character set mismatch.
[0035] Understandably, effective text length can be the number of readable characters remaining in the initial text after removing invalid characters, reflecting the actual amount of effective information available in the initial text.
[0036] For example, the device extracts the initial text "Contract XX terms total 10X" from the file to be extracted. Traversing this initial text, it finds three garbled characters (represented as "X") that cannot be displayed correctly, while the total number of characters in the initial text is 10 (including Chinese characters, numbers, and garbled characters). The device counts 3 invalid characters, and the total number of characters is 10. The current garbled character ratio is calculated to be 3 / 10 = 0.3 (i.e., 30%), and the effective text length is 10 - 3 = 7 characters. That is, the current garbled character ratio is 30%, and the effective text length is 7 characters.
[0037] In practical use, after obtaining the initial text, the device performs a garbled character detection operation. First, it iterates through each character in the initial text, identifying whether each character is invalid. Next, it counts the total number of invalid characters in the initial text, while simultaneously counting the total number of characters. Then, based on the ratio between the total number of invalid characters and the total number of characters, it calculates the current garbled character ratio, and based on the difference between the total number of characters and the total number of invalid characters, it determines the length of the valid text.
[0038] Step S30: Based on the current garbled text ratio and the effective text length, perform downgraded extraction on the file to be extracted to obtain the corresponding garbled text replacement text.
[0039] It should be understood that the garbled text replacement text can be obtained through a downgrade extraction method, and can replace the original garbled text with valid text content.
[0040] Specifically, when the initial text obtained through conventional electronic text extraction methods is of substandard quality, the approach can be downgraded from directly parsing the internal text layer of the PDF to performing text recognition on the page image using optical character recognition technology. Alternatively, other methods may be used; this embodiment does not limit the specific approach.
[0041] In practical use, after obtaining the current garbled text ratio and the length of the valid text, if the current garbled text ratio is too high or the length of the valid text is too short, the degradation extraction process is started. The file to be extracted is converted into a form suitable for degradation processing (e.g., rendering the page as an image), and the valid text content is identified from the rendered image to obtain the garbled text replacement text.
[0042] Step S40: Replace the initial text with the garbled text replacement text to obtain the text extraction result corresponding to the file to be extracted.
[0043] It should be noted that the text extraction result can be valid text content that can be used for subsequent analysis or display after garbled character detection and downgrade replacement processing.
[0044] For example, the device extracts the initial text "Contract Terms XX" from a PDF file, then performs a downgraded extraction to obtain the garbled replacement text "Contract Terms (5 Articles)". The garbled replacement text "Contract Terms (5 Articles)" is then overwritten into the location where the initial text was originally stored. Subsequent applications or users will then read this readable text and will no longer see the garbled characters.
[0045] In practical use, after obtaining the garbled text replacement text through downgrade extraction, the initial text can be replaced with the garbled text replacement text, so that any subsequent reference or access to the text extraction results will return the garbled text replacement text instead of the original garbled text.
[0046] Furthermore, in order to determine the above-mentioned file type, in this embodiment, the step of determining the file type corresponding to the file to be extracted includes: Step S11: Perform text layer detection on the file to be extracted to obtain the first detection result.
[0047] It should be noted that the first detection result can be a Boolean value or status flag obtained after performing a text layer existence check on the file to be extracted, indicating whether the file to be extracted contains a parsable text layer. The text layer can be an internal data structure in the file to be extracted that stores character content and positional information, such as the page content stream in a PDF file, which includes character encoding, font references, and coordinate information.
[0048] Step S12: If a text layer exists in the file to be extracted, extract the font name field and font descriptor from the text layer.
[0049] Understandably, the font name field can be a string attribute in the font object used to identify the font name. The font descriptor can be a data structure in the font object used to describe font attributes, including information such as font type, embedding flags, and font encoding method.
[0050] Step S13: Determine the second detection result based on the symbol characteristics of the font name field and / or the flag characteristics of the font descriptor.
[0051] It should be understood that symbolic features can be special symbolic patterns in the font name field, such as whether it contains a "+" symbol or a random string prefix, used to determine whether the font is an embedded subset font. Flag features can refer to the values of specific flag bits in the font descriptor, such as the "all font embeddings" flag bit or the "font subset embeddings" flag bit, used to indicate the completeness of font embedding.
[0052] It should be noted that the second detection result can be a font quality status indicator obtained based on the analysis of the font name field and font descriptor, which is used to help determine the file type.
[0053] Step S14: Based on the first detection result and the second detection result, determine the file type corresponding to the file to be extracted.
[0054] For ease of understanding, the following examples are provided for illustration, but no specific limitations are imposed on this embodiment. (Reference) Figure 2 , Figure 2This is an overall flowchart of document extraction provided in Embodiment 1 of this application. After obtaining the PDF file, the device can first load the PDF file and traverse each page to parse the file, such as detecting whether a text layer exists on the page and whether the font embedding information in the text layer is complete. Finally, based on the above detection results, the file type is comprehensively detected: if there is no text layer, it indicates that the file type is scanned, that is, the PDF file is a scanned PDF; if there is a text layer and the font is complete, it indicates that the file type is electronic, that is, the PDF file is an electronic PDF; if there is a text layer but the font is missing or partially missing, it indicates that the file type is a mixed type, that is, the PDF file is a mixed PDF.
[0055] In this embodiment, the device first parses the page structure of the file to be extracted and checks whether there is any non-empty text layer content. If text layer content exists, the first detection result is "text layer exists"; if no text layer content exists, the first detection result is "text layer does not exist". If a text layer exists in the file to be extracted, each font object used in the text layer is further traversed, the string value corresponding to the font name field is read, and the flag information in the font descriptor is read. Then, analysis is performed based on the symbol characteristics of the font name field, such as detecting whether the font name field contains a "+" symbol or a random alphanumeric prefix, to determine whether the font is an embedded subset font; or analysis is performed based on the flag characteristics of the font descriptor, such as checking whether the embedding flag is true, to determine whether the font is fully embedded, thereby determining the second detection result. This second detection result may include states such as "font fully embedded", "font subset embedded but mapping normal", or "font embedding abnormal". Finally, the file type corresponding to the file to be extracted is determined based on the first and second detection results. If the first detection result indicates the absence of a text layer, the file type is a scanned file. If the first detection result indicates the presence of a text layer and the second detection result indicates complete font embedding, the file type is a plain text file. If the first detection result indicates the presence of a text layer but the second detection result indicates abnormal font embedding or a mapping problem in font subset embedding, the file type is a mixed file. Because the symbolic features in the font name and the subset embedding flag in the font descriptor accurately reflect the completeness of font embedding, it can precisely distinguish between electronic, mixed, and scanned file types, achieving refined classification of PDF file types. This provides a basis for subsequent processing strategies and avoids garbled output caused by misjudgment based solely on the presence of a text layer.
[0056] Furthermore, in order to replace the initial text, in this embodiment, the step of replacing the initial text with the garbled replacement text to obtain the text extraction result corresponding to the file to be extracted includes: Step S41: Upload the garbled text replacement text to the distributed object storage system and obtain the corresponding first storage path. The distributed object storage system is used to store the text extraction results.
[0057] It should be noted that a distributed object storage system can be a storage system used to store unstructured data (such as files, text content, and images), such as Amazon S3, MinIO, and other common distributed object storage systems.
[0058] Understandably, the first storage path can be a unique access address or key value obtained after the garbled text is stored in the distributed object storage system, which is used to locate and read the text content later.
[0059] Step S42: Determine the second storage path of the initial text in the preset relational database.
[0060] It should be understood that the default relational database can be a pre-configured relational database used to store document metadata and text storage path mappings. The second storage path can be the storage location information of the initial text recorded in the default relational database, pointing to the location of the initial text in a distributed object storage system or local storage, and usually exists in the form of the original value of a field in a database table.
[0061] Step S43: Update the second storage path to the first storage path in the preset relational database to obtain the text extraction result corresponding to the file to be extracted.
[0062] For example, after the downgrade extraction is completed, the garbled text replacement text can be obtained. This replacement text can then be uploaded to the MinIO distributed object storage system. Next, the second storage path of the initial text in the preset relational database is updated, pointing to the first storage path corresponding to the garbled text replacement text. It should be noted that if the downgrade extraction fails, the initial text can be retained and a quality warning marked, thus preventing any data loss.
[0063] Furthermore, in resource-constrained scenarios, the device can only perform the above-mentioned file type detection and garbled character ratio determination processes; then, the garbled character ratio is used as the document quality score output, without performing downgraded extraction. The upper-layer business decides whether to trigger the downgraded extraction operation, which can also achieve a quantitative assessment of text quality.
[0064] In this embodiment, after obtaining the garbled text replacement through downgrade extraction, the device can establish a network connection with the distributed object storage system and store the content of the garbled text replacement as an object in a designated storage bucket via a write operation. After the upload is complete, the distributed object storage system returns a unique identifier to the device and records it as the first storage path. Then, a query request is sent to a preset relational database to read the path field value of the initial stored text, obtaining the second storage path. Subsequently, the second storage path field value is changed to the first storage path in the preset relational database. Because the distributed object storage and the relational database are separate, the result replacement can be completed through atomic updates of the path in the database, without moving or overwriting large file objects. This avoids data inconsistency issues caused by intermediate states such as successful upload but failed update, achieving atomic replacement of the text extraction result.
[0065] This application provides a text extraction method. First, the file type of the file to be extracted is determined. If the file is not plain text, initial text is extracted. Then, the current garbled text ratio and effective text length are detected in the initial text. Based on the detection results, downgraded extraction is performed to obtain garbled text replacement text, which is then updated to update the initial text. This embodiment can automatically determine the extraction quality and perform downgraded extraction based on the garbled text ratio and effective text length, reducing invalid text output and improving the reliability of text processing.
[0066] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the above embodiment can be referred to the above description, and will not be repeated hereafter. Based on this, refer to... Figure 3 , Figure 3 This is a flowchart illustrating Embodiment 2 of the text extraction method of this application. To obtain the aforementioned current garbled character ratio and effective text length, as shown... Figure 3 As shown, in this embodiment, the step of detecting the initial text and determining the current garbled text ratio and effective text length includes: Step S21: Determine the number of first invalid characters based on the number of times the replacement character appears in the initial text.
[0067] It should be noted that the replacement character can be a special character in the Universal Coded Character Set (Unicode) specifically used to represent characters with missing encoding mappings or those that are unrecognizable. The code point corresponding to the replacement character is U+FFFD, which is usually displayed as a diamond symbol with a question mark in the text. When the character encoding in the PDF file cannot be mapped to a valid Unicode code point, the parsing tool will output this character. The first invalid character count can be the total number of times the replacement character appears in the initial text.
[0068] Step S22: Detect private zone characters in the initial text, and determine the number of second invalid characters based on the number of occurrences of the private zone characters.
[0069] It should be understood that private region characters can be characters within a character region reserved for font-specific mappings in the Unicode character set. The code point range for private region characters is from U+E000 to U+F8FF. There is no unified standard definition for characters within this region, and different fonts can define their own glyph mappings, leading to garbled characters when crossing font or system boundaries. The second invalid character count can be the total number of occurrences of private region characters in the initial text.
[0070] Step S23: Detect control characters in the initial text and determine the number of third invalid characters based on the number of times the control characters appear.
[0071] Understandably, control characters can be non-printable characters in the ASCII character set with code points ranging from 0x00 to 0x1F, such as the null character (NULL), horizontal tab (HT), newline (LF), and carriage return (CR). Control characters are typically used to control device behavior rather than represent readable text. Apart from common whitespace control characters such as tabs, newlines, and carriage returns, the appearance of other control characters in text content often indicates data corruption or extraction errors. The third invalid character count can be the total number of times control characters appear in the initial text. It's important to note that the first, second, and third counts are only distinguished and do not have any order of priority.
[0072] Step S24: Determine the current garbled character ratio and valid text length corresponding to the initial text based on the number of the first invalid characters, the number of the second invalid characters, the number of the third invalid characters, and the total number of characters in the initial text.
[0073] In practical use, the device first determines the number of invalid characters based on the frequency of occurrence of replacement characters in the initial text. Then, it detects private area characters and control characters to determine the number of invalid characters for the second and third categories, respectively. Finally, it combines the number of these three types of invalid characters with the total number of characters to determine the current garbled text ratio and the length of the valid text. Since replacement characters, private area characters, and control characters are the three most common and representative types of abnormal encoding in PDF garbled text, detecting these three types of characters simultaneously can cover the vast majority of garbled text scenarios, thus avoiding the missed detection problem of single-dimensional detection.
[0074] Furthermore, considering special cases such as invalid proxy pairs or non-characters, in this embodiment, the step of determining the current garbled text ratio and effective text length corresponding to the initial text based on the number of the first invalid characters, the number of the second invalid characters, the number of the third invalid characters, and the total number of characters in the initial text includes: Step S241: Detect candidate characters located in the preset proxy character area in the initial text, and determine the number of fourth invalid characters corresponding to invalid proxy pairs in the candidate characters.
[0075] It should be noted that the aforementioned preset surrogate character regions can be the high and low surrogate regions specifically reserved for surrogate pairs in the Unicode encoding standard. The code point range of the high surrogate region is from U+D800 to U+DBFF, and the code point range of the low surrogate region is from U+DC00 to U+DFFF, with the entire surrogate region ranging from U+D800 to U+DFFF. The surrogate characters themselves do not represent valid characters; they are only used to represent auxiliary plane characters in UTF-16 encoding.
[0076] It should also be noted that invalid proxy pairs can be either high or low proxy characters that are located within the proxy area but cannot be paired with a corresponding proxy character to form a valid proxy pair. A proxy character is considered invalid when it is not followed by a valid low proxy character, or when it is not preceded by a valid high proxy character. For example, the single character U+D800 is an invalid proxy pair. The fourth invalid character count can be the total number of times invalid proxy pairs appear in the initial text.
[0077] Step S242: Perform code point detection on the initial text to obtain the number of fifth invalid characters corresponding to non-characters.
[0078] Understandably, non-characters can be code points explicitly marked by the Unicode standard as unusable for interchangeable characters, such as U+FFFE (non-character flag), U+FFFF (non-character flag), and code points in the range U+FDD0 to U+FDEF. Non-characters should not appear in normal text; their appearance indicates data anomaly. The fifth invalid character count represents the total number of non-character occurrences in the initial text.
[0079] Step S243: Determine the number of sixth invalid characters based on the number of times whitespace characters appear in the initial text.
[0080] It should be understood that whitespace characters can be spaces, tabs, newlines, carriage returns, form feeds, or other characters that represent blanking or separation. In this embodiment, whitespace characters are considered invalid because when the proportion of whitespace characters in the initial text is too high (e.g., exceeding 90%), it indicates that the text content is minimal or that valid information is missing. The sixth invalid character count can refer to the total number of times whitespace characters appear in the initial text.
[0081] Step S244: Sum the number of the first invalid characters, the second invalid characters, the third invalid characters, the fourth invalid characters, the fifth invalid characters, and the sixth invalid characters to obtain the current number of invalid characters.
[0082] Step S245: Determine the corresponding current garbled character ratio and valid text length based on the total number of characters in the initial text and the current number of invalid characters.
[0083] Specifically, by counting the total number of the above six types of invalid characters, we can calculate: Current garbled character ratio = Current number of invalid characters / Total number of characters.
[0084] In this embodiment, after determining the third invalid character, the device can further detect candidate characters located in the preset proxy character area in the initial text and determine the number of fourth invalid characters corresponding to invalid proxy pairs. Then, code point detection is performed on the initial text to obtain the number of fifth invalid characters corresponding to non-characters. Next, the number of sixth invalid characters is determined based on the occurrence frequency of whitespace characters. Then, the number of invalid characters from the first to the sixth invalid characters is summed to obtain the current number of invalid characters. Finally, the current garbled character ratio and valid text length are determined based on the total number of characters and the current number of invalid characters. Because the detection of three dimensions—invalid proxy pairs, non-characters, and whitespace characters—is incorporated simultaneously, and the garbled character ratio is calculated by summing the six types of invalid characters, a comprehensive coverage of various garbled character scenarios, such as encoding errors, illegal reserved characters, and sparse content, is achieved, realizing all-round detection of PDF garbled character scenarios and further reducing the false negative rate.
[0085] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to that in embodiments one and two above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 , Figure 4 This is a flowchart illustrating Embodiment 3 of the text extraction method of this application. To obtain the aforementioned garbled text replacement text, as... Figure 4 As shown, in this embodiment, the step of downgrading the extraction of the file to be extracted based on the current garbled text ratio and the effective text length to obtain the corresponding garbled text replacement text includes: Step S31: Detect whether the file to be extracted meets the preset downgrade conditions based on the current garbled character ratio and the effective text length. The preset downgrade conditions include the current garbled character ratio reaching a first preset ratio and / or the effective text length being less than a preset length.
[0086] It should be noted that the preset degradation conditions can be pre-set rules used to determine whether to switch from electronic text extraction to optical character recognition extraction. The preset degradation conditions include two scenarios: the current garbled text ratio reaches a first preset ratio, and the effective text length is less than a preset length. These two scenarios can occur individually or simultaneously.
[0087] Understandably, the first preset ratio can be a pre-set threshold for the proportion of garbled text, used to determine whether the quality of the initial text is unacceptable. The preset length can be a pre-set threshold for the minimum effective text length, used to determine whether there is too little readable content in the initial text.
[0088] For example, such as Figure 2 As shown, when the device extracts the initial text, it needs to perform further quality checks. This process is divided into two types of checks: garbled character ratio and text length. The first preset ratio can be set to 10%, and the preset length can be set to 100 characters to determine whether downgraded extraction is required.
[0089] Step S32: If the file to be extracted meets the preset downgrade conditions, the file to be extracted is downgraded to obtain the corresponding image to be extracted.
[0090] It should be understood that the image to be extracted can be an image file containing the page content of the file to be extracted, obtained by converting the file to be extracted from its original format to an image format.
[0091] For example, each page of the PDF file can be rendered as a Portable Network Graphics (PNG) image or a Joint Photographic Experts Group (JPEG) image, or other formats, which are not limited in this embodiment.
[0092] Step S33: Perform optical character recognition on the image to be extracted to obtain the current valid text.
[0093] Step S34: Determine the corresponding garbled text replacement text from the currently valid text.
[0094] It should also be understood that optical character recognition (OCR) is a technique that uses an optical character recognition engine to identify and convert text regions in an image into editable text. The currently valid text can be a complete sequence of text output after the image to be extracted is recognized by optical character recognition.
[0095] For example, such as Figure 2As shown, assuming that after detecting the initial text, the current garbled text ratio is 12% and the effective text length is 80 characters. If the first preset ratio is 10% and the preset length is 100 characters, the device finds that the current garbled text ratio of 12% has reached the 10% threshold, and the effective text length of 80 is also less than the 100 threshold. Therefore, the initial text is deemed unacceptable and needs to be downgraded for extraction. Subsequently, the PDF file to be extracted is rendered page by page as PNG images, resulting in a total of 5 images to be extracted. Then, text recognition is performed on each image to be extracted, and the recognition results are concatenated to obtain the current effective text. All of the current effective text is used as the garbled text replacement text for subsequent result updates.
[0096] If the current garbled text ratio is less than 10%, and the valid text length is greater than 100 characters, then the initial text is considered acceptable and will be directly used for subsequent result updates. For example, if the PDF type is scanned, OCR extraction can be performed directly.
[0097] This embodiment first checks whether the file to be extracted meets preset degradation conditions based on the current garbled text ratio and the length of the effective text. If the conditions are met, the file to be extracted is downgraded to obtain the image to be extracted. Then, optical character recognition is performed on the image to be extracted to obtain the current effective text. Finally, the garbled text replacement text is determined from the current effective text. Because it uses dual judgment conditions of garbled text ratio and effective text length, and directly links the degradation conditions to the quality of the electronic extraction result, the time-consuming OCR degradation process is only triggered when the quality of the electronic extraction is substandard, avoiding the waste of resources caused by unconditionally performing OCR on all files.
[0098] Furthermore, to improve the accuracy of whether downgraded extraction is triggered, in this embodiment, after the step of detecting whether the file to be extracted meets the preset downgrade conditions based on the current garbled character ratio and the effective text length, the method further includes: Step S32': If the current garbled character ratio does not reach the first preset ratio, determine whether the current garbled character ratio reaches the second preset ratio, where the first preset ratio is higher than the second preset ratio.
[0099] It should be noted that the second preset ratio can be a pre-set threshold for the proportion of garbled text, lower than the first preset ratio, used to determine whether the quality of the initial text falls within the "suspected garbled text" gray area. When the current garbled text ratio has reached the second preset ratio but not the first preset ratio, it indicates that the initial text may have localized garbled text issues, but the severity has not yet reached the point where automatic downgrade extraction will be triggered. For example, the first preset ratio can be set to 10%, and the second preset ratio can be set to 5%.
[0100] Step S33': If the current garbled text ratio reaches the second preset ratio, then the text pages containing garbled text in the file to be extracted are marked as candidate downgraded texts, and the candidate downgraded texts are reviewed.
[0101] Understandably, candidate degradation text can be the text content corresponding to the pages in the file to be extracted that are marked as potentially requiring degradation extraction. It's important to note that candidate degradation text can be a portion of the entire file to be extracted (e.g., only pages with a high proportion of garbled characters), not all pages.
[0102] Specifically, the review can be a manual or assisted review of the candidate downgraded text to confirm whether the candidate downgraded text actually has garbled text issues, in order to avoid automatic misjudgment. The review can be performed by human reviewers or by the system calling auxiliary tools for secondary verification; this embodiment does not impose any restrictions on this.
[0103] Step S34': If the verification result is a garbled result, the candidate downgraded text is downgraded and extracted to obtain the corresponding garbled replacement text.
[0104] It should be understood that when the review result is a garbled result, it means that the candidate downgraded text does indeed have a garbled character problem and needs to be downgraded for extraction.
[0105] For example, in scenarios requiring high precision, when the proportion of garbled characters is between 5% and 10%, text pages containing garbled characters in the file to be extracted can be marked as "suspected garbled characters" and pushed to a manual review queue. After manual confirmation, a decision can be made on whether to trigger OCR downgrade extraction, thereby further improving the processing accuracy.
[0106] In this embodiment, the device first determines whether a second preset ratio has been reached if the current garbled text ratio is below the first preset ratio. If it is, the text page containing garbled text is marked as candidate downgraded text and reviewed. If the review result is garbled text, the candidate downgraded text is extracted to obtain the garbled text replacement text. Because a second preset ratio lower than the first preset ratio is set as a fuzzy range, and the suspected garbled text within this range is confirmed a second time by manual review, the erroneous downgrade or missed downgrade that may be caused by a single hard threshold near the critical value is avoided, further improving the quality assurance of text extraction.
[0107] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the text extraction method of this application. Any simple modifications based on this technical concept are within the scope of protection of this application. Furthermore, all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection regulations of the country where the application is located and with authorization from the owner of the relevant device.
[0108] This application also provides a text extraction device, please refer to... Figure 5 , Figure 5 This is a block diagram of the module structure of the text extraction device according to an embodiment of this application; in this embodiment, the text extraction device includes: Extraction module 101 is used to determine the file type corresponding to the file to be extracted, and extract initial text from the file to be extracted if the file type is a preset text type; Detection module 102 is used to detect the initial text and determine the current garbled text ratio and effective text length; The downgrade module 103 is used to downgrade the file to be extracted based on the current garbled character ratio and the effective text length, so as to obtain the corresponding garbled character replacement text. The update module 104 is used to replace the initial text according to the garbled text replacement text to obtain the text extraction result corresponding to the file to be extracted.
[0109] This embodiment first determines the file type of the file to be extracted and extracts the initial text when it is not plain text. Then, it detects the current garbled text ratio and the length of the valid text in the initial text. Based on the detection results, it performs downgraded extraction to obtain garbled text replacement text and updates the initial text. This embodiment can automatically judge the extraction quality and perform downgraded extraction by using the garbled text ratio and the length of the valid text, reducing the occurrence of invalid text output and improving the reliability of text processing.
[0110] In one implementation, the detection module 102 is further configured to: determine the first invalid character count based on the number of times the replacement character appears in the initial text; detect private area characters in the initial text and determine the second invalid character count based on the number of times the private area characters appear; detect control characters in the initial text and determine the third invalid character count based on the number of times the control characters appear; and determine the current garbled character ratio and effective text length corresponding to the initial text based on the first invalid character count, the second invalid character count, the third invalid character count, and the total number of characters in the initial text.
[0111] In one implementation, the detection module 102 is further configured to detect candidate characters located in the preset proxy character area in the initial text, and determine the number of fourth invalid characters corresponding to invalid proxy pairs among the candidate characters; perform code point detection on the initial text to obtain the number of fifth invalid characters corresponding to non-characters; determine the number of sixth invalid characters based on the number of occurrences of whitespace characters in the initial text; sum the number of first invalid characters, second invalid characters, third invalid characters, fourth invalid characters, fifth invalid characters, and sixth invalid characters to obtain the current number of invalid characters; and determine the corresponding current garbled character ratio and effective text length based on the total number of characters in the initial text and the current number of invalid characters.
[0112] In one implementation, the downgrade module 103 is further configured to detect whether the file to be extracted meets preset downgrade conditions based on the current garbled character ratio and the effective text length, wherein the preset downgrade conditions include the current garbled character ratio reaching a first preset ratio and / or the effective text length being less than a preset length; if the file to be extracted meets the preset downgrade conditions, the file to be extracted is downgraded to obtain a corresponding image to be extracted; optical character recognition is performed on the image to be extracted to obtain the current effective text; and the corresponding garbled character replacement text is determined from the current effective text.
[0113] In one implementation, the downgrade module 103 is further configured to, when the current garbled character ratio does not reach the first preset ratio, determine whether the current garbled character ratio reaches a second preset ratio, wherein the first preset ratio is higher than the second preset ratio; if the current garbled character ratio reaches the second preset ratio, mark the text pages containing garbled characters in the file to be extracted as candidate downgraded texts, and review the candidate downgraded texts; if the review result is a garbled character result, downgrade the candidate downgraded texts to obtain the corresponding garbled character replacement text.
[0114] In one implementation, the update module 104 is further configured to upload the garbled text replacement text to a distributed object storage system and obtain a corresponding first storage path, wherein the distributed object storage system is used to store the text extraction results; determine the second storage path of the initial text in a preset relational database; and update the second storage path to the first storage path in the preset relational database to obtain the text extraction results corresponding to the file to be extracted.
[0115] In one implementation, the extraction module 101 is further configured to perform text layer detection on the file to be extracted to obtain a first detection result; if a text layer exists in the file to be extracted, extract a font name field and a font descriptor from the text layer; determine a second detection result based on the symbol features of the font name field and / or the flag features of the font descriptor; and determine the file type corresponding to the file to be extracted based on the first detection result and the second detection result.
[0116] Other embodiments or specific implementations of the text extraction device of this application can be found in the above-described method embodiments, and will not be repeated here.
[0117] The text extraction device provided in this application, employing the text extraction method described in the above embodiments, can solve the technical problem of difficulty in extracting valid text during PDF text extraction due to garbled characters. Compared with the prior art, the beneficial effects of the text extraction device provided in this application are the same as those of the text extraction method provided in the above embodiments, and other technical features in the text extraction device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0118] This application provides a text extraction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the text extraction methods in the above embodiments.
[0119] The following is for reference. Figure 6 , Figure 6 This is a schematic diagram of the hardware operating environment involved in the text extraction device in the embodiments of this application, showing a structural schematic diagram suitable for implementing the text extraction device in the embodiments of this application. The text extraction device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The text extraction device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0120] like Figure 6As shown, the text extraction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the text extraction device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the text extraction device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show text extraction devices with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.
[0121] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0122] The text extraction device provided in this application, employing the text extraction method described in the above embodiments, can solve the technical problem of difficulty in extracting valid text during PDF text extraction due to garbled characters. Compared with the prior art, the beneficial effects of the text extraction device provided in this application are the same as those of the text extraction method provided in the above embodiments, and other technical features of this text extraction device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0123] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0124] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0125] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the text extraction method in the above embodiments.
[0126] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.
[0127] The aforementioned computer-readable storage medium may be included in the text extraction device; or it may exist independently and not be assembled into the text extraction device.
[0128] The aforementioned computer-readable storage medium carries one or more programs. When the one or more programs are executed by a text extraction device, the text extraction device: determines the file type corresponding to the file to be extracted, and if the file type is a preset text type, extracts initial text from the file to be extracted; detects the initial text to determine the current garbled character ratio and the effective text length; performs downgraded extraction on the file to be extracted based on the current garbled character ratio and the effective text length to obtain corresponding garbled character replacement text; and replaces the initial text according to the garbled character replacement text to obtain the text extraction result corresponding to the file to be extracted.
[0129] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation that may be implemented in systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0131] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0132] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described text extraction method. This solves the technical problem of difficulty in extracting valid text during PDF text extraction due to garbled characters. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the text extraction method provided in the above embodiments, and will not be repeated here.
[0133] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the text extraction method described above.
[0134] The computer program product provided in this application can solve the technical problem of difficulty in extracting valid text due to garbled characters during PDF text extraction. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the text extraction method provided in the above embodiments, and will not be repeated here.
[0135] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.
Claims
1. A text extraction method, characterized in that, The method includes: Determine the file type corresponding to the file to be extracted, and extract the initial text from the file to be extracted if the file type is a preset text type; The initial text is inspected to determine the current garbled text ratio and the length of the effective text; Based on the current garbled text ratio and the effective text length, the file to be extracted is downgraded for extraction to obtain the corresponding garbled text replacement text; The initial text is replaced with the garbled text replacement text to obtain the text extraction result corresponding to the file to be extracted.
2. The method as described in claim 1, characterized in that, The step of detecting the initial text and determining the current garbled text ratio and effective text length includes: The number of first invalid characters is determined based on the number of times the replacement character appears in the initial text; Detect private zone characters in the initial text, and determine the number of second invalid characters based on the number of occurrences of the private zone characters; Detect control characters in the initial text, and determine the number of third invalid characters based on the number of times the control characters appear; Based on the number of the first invalid characters, the number of the second invalid characters, the number of the third invalid characters, and the total number of characters in the initial text, determine the current garbled character ratio and the effective text length corresponding to the initial text.
3. The method as described in claim 2, characterized in that, The step of determining the current garbled character ratio and effective text length corresponding to the initial text based on the number of the first invalid characters, the number of the second invalid characters, the number of the third invalid characters, and the total number of characters in the initial text includes: Detect candidate characters located in the preset proxy character area in the initial text, and determine the number of fourth invalid characters corresponding to invalid proxy pairs among the candidate characters; Code point detection is performed on the initial text to obtain the number of fifth invalid characters corresponding to non-characters; The number of sixth invalid characters is determined based on the number of occurrences of whitespace characters in the initial text; The sum of the first invalid character count, the second invalid character count, the third invalid character count, the fourth invalid character count, the fifth invalid character count, and the sixth invalid character count is used to obtain the current invalid character count. Based on the total number of characters in the initial text and the current number of invalid characters, determine the corresponding current garbled character ratio and valid text length.
4. The method as described in claim 1, characterized in that, The step of performing downgraded extraction on the file to be extracted based on the current garbled character ratio and the effective text length to obtain the corresponding garbled character replacement text includes: The file to be extracted is detected based on the current garbled character ratio and the effective text length to determine whether it meets the preset degradation conditions. The preset degradation conditions include the current garbled character ratio reaching a first preset ratio and / or the effective text length being less than a preset length. If the file to be extracted meets the preset downgrade conditions, the file to be extracted is downgraded and converted to obtain the corresponding image to be extracted; Optical character recognition is performed on the image to be extracted to obtain the current valid text; Determine the corresponding garbled text replacement text from the currently valid text.
5. The method as described in claim 4, characterized in that, After the step of detecting whether the file to be extracted meets the preset downgrade conditions based on the current garbled text ratio and the effective text length, the method further includes: If the current garbled character ratio does not reach the first preset ratio, it is determined whether the current garbled character ratio reaches the second preset ratio, where the first preset ratio is higher than the second preset ratio; If the current garbled text ratio reaches the second preset ratio, then the text pages containing garbled text in the file to be extracted are marked as candidate downgraded texts, and the candidate downgraded texts are reviewed. If the verification result is a garbled result, the candidate downgraded text is downgraded and extracted to obtain the corresponding garbled replacement text.
6. The method as described in claim 1, characterized in that, The step of replacing the initial text with the garbled replacement text to obtain the text extraction result corresponding to the file to be extracted includes: The garbled text replacement text is uploaded to a distributed object storage system, and a corresponding first storage path is obtained. The distributed object storage system is used to store the text extraction results. Determine the second storage path of the initial text in a preset relational database; The second storage path is updated to the first storage path in the preset relational database to obtain the text extraction result corresponding to the file to be extracted.
7. The method according to any one of claims 1 to 6, characterized in that, The step of determining the file type corresponding to the file to be extracted includes: The text layer detection of the file to be extracted is performed to obtain the first detection result; If a text layer exists in the file to be extracted, extract the font name field and font descriptor from the text layer; The second detection result is determined based on the symbol characteristics of the font name field and / or the flag characteristics of the font descriptor; Based on the first detection result and the second detection result, the file type corresponding to the file to be extracted is determined.
8. A text extraction device, characterized in that, The device includes: The extraction module is used to determine the file type of the file to be extracted, and extract the initial text from the file to be extracted if the file type is a preset text type; The detection module is used to detect the initial text and determine the current garbled text ratio and the length of the effective text. The downgrade module is used to downgrade the file to be extracted based on the current garbled character ratio and the effective text length, so as to obtain the corresponding garbled character replacement text. The update module is used to replace the initial text with the garbled text replacement text to obtain the text extraction result corresponding to the file to be extracted.
9. A text extraction device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the text extraction method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the text extraction method as described in any one of claims 1 to 7.